Training method of multi-parameter collaborative regulation model of electric pump separate production well based on bias features

CN122548272BActive Publication Date: 2026-09-29CHINA UNIV OF PETROLEUM (BEIJING)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611041611.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29
Estimated Expiration
2046-07-14

AI Technical Summary

Technical Problem

[0003]目前油藏层面的分层配产方案优化已形成成熟体系,但电泵分采井生产端分层流量调控仍高度依赖人工经验判断,现有技术中电泵分采井生产参数(动作向量)调控存在人工经验调控主观性强、调配效率低的问题,导致各产层产液量的调控精度较低,因此,现有技术缺乏能够实现电泵分采井各产层产液量的准确调控的方法

Benefits of technology

[0018]上述技术方案,获取基于预设油井机理模型确定的电泵分采井在当前状态的第一状态向量和在上一状态的第二状态向量,可以避免仅依靠单时刻工况决策带来的滞后误判,为后续模型训练打下数据基础。基于待训练的预设初始电泵分采井多参协同调控模型,根据当前状态的第一状态向量确定当前状态的动作向量以及动作向量的时间差分误差值,并在预设油井机理模型中执行动作向量得到电泵分采井在下一状态的第三状态向量,可以在预设油井机理模型仿真环境内下发动作向量、推演后续状态向量,全程无需改动现场真实油井生产参数,杜绝了阀门、电泵频繁调试造成的设备损耗,同步计算时间差分误差,以达到量化单条样本的学习价值的技术效果。基于第一状态向量、第二状态向量和第三状态向量各自对应的执行偏差量,确定当前状态的奖励值,可以以油田现场最关心的电泵分采井在各状态的实际产量和预期产量之间的偏差量作为奖励构造依据,使用三个状态的执行偏差量弱化了瞬时状态扰动干扰。基于当前状态的第一状态向量、动作向量、动作向量的时间差分误差值、奖励值、下一状态的第三状态向量构建当前状态对应的模型训练样本,可见训练样本完整包含观测状态、调控动作、时序误差、奖惩反馈、后续新状态全部关键信息,样本特征维度完备。将下一状态更新为当前状态,重复上述迭代步骤以重新确定当前状态对应的模型训练样本,直至得到预设数量的模型训练样本,也即,本申请采用滚动迭代方式批量自动生成海量多样化训练样本,可遍历各类状态,无需人工逐组采集现场数据,样本构建效率大幅提升,同时保证样本覆盖面足够全面。本申请根据预设数量的模型训练样本,训练预设初始电泵分采井多参协同调控模型,以得到训练好的电泵分采井多参协同调控模型,可以依托完备样本完成模型训练,训练完成后的模型可自主输出流量控制阀开度(ICV开度)、潜油电泵的电泵频率、电泵分采井的井口油压整套连续调控指令,摆脱人工反复整定调控参数的依赖,以实现电泵分采井动作向量的准确调控,进而实现电泵分采井各产层产液量的准确调控,另外,本申请可以全程以实际产量与预期配产的执行偏差作为奖励计算依据,同时,双模式自适应切换的奖励机制也能够显著提升奖励值的确定速度与精准度,满足预设初始电泵分采井多参协同调控模型的训练样本中的奖励值的计算需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548272B_ABST
    Figure CN122548272B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a training method of an electric pump separate production well multi-parameter collaborative regulation model based on deviation characteristics, and belongs to the technical field of oil and gas development. The training method comprises the following steps: based on a preset initial electric pump separate production well multi-parameter collaborative regulation model to be trained, determining an action vector of a current state and a time difference error value of the action vector according to a first state vector of the current state, and executing the action vector in a preset oil well mechanism model to obtain a third state vector of the electric pump separate production well in a next state; determining a reward value of the current state based on respective execution deviation amounts of the first state vector, the second state vector and the third state vector; updating the next state to the current state, and determining a model training sample corresponding to the current state; training the model according to a preset number of model training samples, so as to obtain a trained electric pump separate production well multi-parameter collaborative regulation model. The application can realize accurate regulation of liquid production of each production layer of the electric pump separate production well.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of oil and gas development technology, and specifically to a training method for a multi-parameter collaborative control model of an electric pump production well based on deviation characteristics. Background Technology

[0002] Offshore high-water-cut oilfields have generally entered a stage of high recovery and high water-cut development, with prominent inter-layer conflicts and significant challenges in stabilizing oil production and controlling water. Refined stratified injection-production technology can effectively alleviate the problem of uneven displacement between layers. Among these technologies, electric pump-assisted production (EPP) wells, with their strong lifting capacity and wide displacement adaptability range, have become the core process for the accurate development of multi-layer offshore reservoirs. Therefore, the efficient application of EPP wells is crucial.

[0003] At present, the optimization of stratified production schemes at the reservoir level has formed a mature system. However, the stratified flow rate control at the production end of electric pump-propelled wells still relies heavily on human experience. The existing technology for controlling the production parameters (action vectors) of electric pump-propelled wells suffers from the problems of strong subjectivity and low efficiency of human experience in control, resulting in low control accuracy of the fluid production of each production layer. Therefore, the existing technology lacks a method to accurately control the fluid production of each production layer in electric pump-propelled wells. Summary of the Invention

[0004] The purpose of this application is to provide a training method for a multi-parameter collaborative control model of an electric pump-propelled well based on deviation characteristics, which is used to provide a method for accurately controlling the fluid production of each producing layer in an electric pump-propelled well.

[0005] To achieve the above objectives, the first aspect of this application provides a training method for a multi-parameter collaborative control model of an electric pump production well, the training method comprising: Obtain the first state vector of the electric pump production well in the current state and the second state vector in the previous state, which are determined based on the preset oil well mechanism model; Based on the preset initial multi-parameter collaborative control model of the electric pump production well to be trained, the action vector of the current state and the time difference error value of the action vector are determined according to the first state vector of the current state. The action vector is then executed in the preset oil well mechanism model to obtain the third state vector of the electric pump production well in the next state. The action vector includes the wellhead oil pressure of the electric pump production well, the electric pump frequency of the submersible electric pump installed in the electric pump production well, and the control valve opening of the flow control valve installed in the electric pump production well. Based on the execution deviations corresponding to the first state vector, the second state vector, and the third state vector, the reward value of the current state is determined. The execution deviation is the deviation between the actual production and the expected production of the electric pump well in each state. The model training samples corresponding to the current state are constructed based on the first state vector, action vector, time difference error value of action vector, reward value, and third state vector of the next state. Update the next state to the current state, and repeat the above iterative steps to redetermine the model training samples corresponding to the current state until a preset number of model training samples are obtained. Based on a preset number of model training samples, a preset initial multi-parameter collaborative control model for electric pump-separated wells is trained to obtain a well-trained multi-parameter collaborative control model for electric pump-separated wells.

[0006] Specifically, determining the reward value of the current state based on the execution deviations corresponding to the first, second, and third state vectors includes: if the execution deviation corresponding to the third state vector is greater than a preset deviation threshold, determining the reward value of the current state based on the execution deviations corresponding to the first and second state vectors; if the execution deviation corresponding to the third state vector is less than or equal to the preset deviation threshold, determining the reward value of the current state based on the execution deviations corresponding to the first and third state vectors.

[0007] In this embodiment of the application, when the execution deviation corresponding to the third state vector is greater than a preset deviation threshold, the reward value of the current state is determined based on the execution deviation corresponding to the first state vector and the execution deviation corresponding to the second state vector, including determining the reward value of the current state according to the following formula:

[0008] in, The reward value for the current state. This is the execution deviation corresponding to the first state vector. This is the execution deviation corresponding to the second state vector. a This is the preset reward coefficient.

[0009] In this embodiment of the application, when the execution deviation corresponding to the third state vector is less than or equal to a preset deviation threshold, the reward value of the current state is determined based on the execution deviation corresponding to the first state vector and the execution deviation corresponding to the third state vector. This includes: when the execution deviation corresponding to the first state vector is greater than the execution deviation corresponding to the third state vector, a preset fixed positive reward value is used as the reward value of the current state, wherein the preset fixed positive reward value is a value greater than zero; when the execution deviation corresponding to the first state vector is less than or equal to the execution deviation corresponding to the third state vector, a preset fixed negative reward value is used as the reward value of the current state, wherein the preset fixed negative reward value is a value less than zero.

[0010] In this embodiment, the first state vector, the second state vector, and the third state vector include the total pumping volume deviation and the product volume deviation of each layer. The execution deviation is determined according to the following formula:

[0011] in, To implement the deviation amount, To preset the balance coefficient, To preset the pump efficiency penalty coefficient, This is due to the deviation in total pumped liquid volume. For the first The preset production penalty coefficient corresponding to the layer To account for the deviation in liquid production at each layer, This represents the total number of floors.

[0012] A second aspect of this application provides a method for determining the action vector of an electric submersible pump (ESP) well. The method includes: acquiring a predicted state vector of the ESP well in its current state, the predicted state vector including the wellhead oil pressure, the operating frequency of the ESP in the ESP well, the opening degree of the flow control valve in the ESP well, the production deviation of each layer in the reservoir where the ESP well is located, and the total pumping volume deviation of the ESP; and determining the target action vector of the ESP well based on a trained multi-parameter collaborative control model of the ESP well, according to the predicted state vector, wherein the target action vector includes the expected wellhead oil pressure, the expected ESP frequency, and the expected control valve opening degree, and the trained multi-parameter collaborative control model of the ESP well is trained according to the training method of the ESP well multi-parameter collaborative control model described above.

[0013] In this embodiment of the application, the method further includes: determining whether the target action vector of the electric pump submersible well meets the preset constraint conditions, wherein the preset constraint conditions include at least the control valve erosion constraint, the control valve opening constraint, and the submersible electric pump operating frequency constraint. The erosion constraint of the control valve satisfies the following formula:

[0014] in, The actual flow rate passing through the flow control valve. The valve flow area of ​​the flow control valve is The corresponding erosion flow rate at that time For valve flow area, The erosion flow coefficient is... For the first Density of the layer produced liquid The pressure drop of the control valve for the flow control valve; Among them, the control valve opening constraint is that the expected control valve opening is within the range of the preset maximum control valve opening and the preset minimum control valve opening; the submersible pump operating frequency constraint is that the expected pump frequency is within the range of the preset maximum pump frequency and the preset minimum pump frequency; when the target action vector of the submersible pump production well does not meet the preset constraint conditions, the target action vector is corrected based on the multi-parameter collaborative control model of the submersible pump production well.

[0015] In this embodiment of the application, the method further includes: obtaining the current production volume of each layer in the reservoir determined based on a preset oil well mechanism model; determining the absolute value of the difference between the current production volume of each layer and the preset target production volume corresponding to each layer, so as to obtain the production volume deviation of each layer in the reservoir where the electric submersible pump production well is located; obtaining the current total pumping volume of the submersible electric pump determined based on the preset oil well mechanism model; and determining the absolute value of the difference between the current total pumping volume and the preset target total pumping volume, so as to obtain the total pumping volume deviation of the submersible electric pump.

[0016] A third aspect of this application provides a training apparatus for a multi-parameter collaborative control model of an electric pump-propelled well based on deviation characteristics, comprising: a memory configured to store instructions; and a processor configured to retrieve instructions from the memory and, when executing the instructions, to implement the aforementioned training method for the multi-parameter collaborative control model of an electric pump-propelled well based on deviation characteristics.

[0017] A fourth aspect of this application provides an apparatus for determining the motion vector of an electric pump-operated well, comprising: a memory configured to store instructions; and a processor configured to retrieve instructions from the memory and, when executing the instructions, to implement the apparatus for determining the motion vector of an electric pump-operated well as described above.

[0018] The above technical solution obtains the first state vector of the electric pump production well in the current state and the second state vector in the previous state, determined based on a preset oil well mechanism model. This avoids the lag and misjudgment caused by relying solely on single-moment operating condition decisions, laying a data foundation for subsequent model training. Based on the preset initial multi-parameter collaborative control model of the electric pump production well to be trained, the action vector and the time difference error value of the action vector in the current state are determined according to the first state vector of the current state. The action vector is then executed in the preset oil well mechanism model to obtain the third state vector of the electric pump production well in the next state. The action vector can be issued and the subsequent state vector can be deduced within the preset oil well mechanism model simulation environment. The entire process does not require modification of the actual oil well production parameters on site, eliminating equipment wear caused by frequent valve and electric pump adjustments. The time difference error is calculated synchronously to achieve the technical effect of quantifying the learning value of a single sample. Based on the execution deviations corresponding to the first, second, and third state vectors, the reward value for the current state is determined. The deviation between the actual and expected production of the electric pump-propelled wells in each state, which is of most concern to the oilfield, can be used as the basis for reward construction. Using the execution deviations of the three states weakens the interference of instantaneous state disturbances. Based on the first state vector, action vector, time difference error value of the action vector, reward value, and the third state vector of the next state, a model training sample corresponding to the current state is constructed. It can be seen that the training sample completely contains all key information of the observed state, control actions, time-series errors, reward and punishment feedback, and subsequent new states, with complete sample feature dimensions. The next state is updated to the current state, and the above iterative steps are repeated to redetermine the model training sample corresponding to the current state until a preset number of model training samples are obtained. That is, this application uses a rolling iteration method to automatically generate massive and diverse training samples in batches, which can traverse various states without the need for manual collection of field data group by group, greatly improving sample construction efficiency while ensuring sufficient sample coverage. This application trains a pre-set initial multi-parameter collaborative control model for electric submersible pump (ESP) wells based on a pre-set number of model training samples to obtain a well-trained ESP well multi-parameter collaborative control model. The model training can be completed using complete samples. After training, the model can autonomously output a complete set of continuous control commands for the flow control valve opening (ICV opening), the ESP frequency, and the wellhead oil pressure of the ESP well, eliminating the reliance on repeated manual tuning of control parameters. This enables accurate control of the ESP well's action vector, thereby achieving accurate control of the fluid production of each production layer in the ESP well. Furthermore, this application can use the deviation between actual production and expected production as the basis for reward calculation throughout the process. Simultaneously, the dual-mode adaptive switching reward mechanism can significantly improve the speed and accuracy of reward value determination, meeting the reward value calculation requirements of the pre-set initial ESP well multi-parameter collaborative control model training samples.

[0019] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0020] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 The illustration shows a flowchart of a training method for a multi-parameter collaborative control model of an electric pump production well according to an embodiment of this application; Figure 2 The diagram illustrates the framework of a training method for a multi-parameter collaborative control model of an electric pump production well according to an embodiment of this application. Figure 3 The diagram schematically illustrates the execution framework of the strategy network for a multi-parameter collaborative control model for an electric pump production well according to an embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the embodiments of this application and are not intended to limit the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0022] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0023] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0024] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0025] Figure 1 The illustration shows a flowchart of a training method for a multi-parameter collaborative control model of an electric pump production well according to an embodiment of this application. Figure 2 This diagram schematically illustrates the framework of a training method for a multi-parameter collaborative control model of an electric pump-propelled well according to an embodiment of this application. Figure 1 and Figure 2 As shown in the embodiment of this application, a training method for a multi-parameter collaborative control model of an electric pump production well is provided. This method may include the following steps: Step S101: Obtain the first state vector of the electric pump production well in the current state and the second state vector in the previous state, which are determined based on the preset oil well mechanism model. It is understandable that the preset oil well mechanism model is a pre-defined oil well mechanism model. An electric submersible centrifugal pump (ESP) well refers to an oil well that uses an electric submersible centrifugal pump as the lifting power and employs tools such as inflow control valves to independently and controllably extract different oil layers in layers. The first state vector is the state vector of the ESP well in the preset oil well mechanism model corresponding to the current state, and the second state vector is the state vector of the ESP well in the preset oil well mechanism model corresponding to the previous state.

[0026] Step S102: Based on the preset initial electric pump production well multi-parameter collaborative control model to be trained, determine the action vector of the current state and the time difference error value of the action vector according to the first state vector of the current state, and execute the action vector in the preset oil well mechanism model to obtain the third state vector of the electric pump production well in the next state. It is understandable that the preset initial multi-parameter collaborative control model for the electric submersible pump (ESP) well is a pre-configured model that has not yet been trained. This model can take a state vector as input and output an action vector for controlling the ESP well after multi-parameter collaborative computation. The action vector includes the wellhead oil pressure of the ESP well, the pump frequency of the ESP installed in the ESP well, and the control valve opening of the flow control valve installed in the ESP well. The Temporal Difference Error (TD error) is an error metric used to measure the difference between the "current estimate" and the "more credible target".

[0027] Specifically, the preset initial electric pump-based multi-parameter collaborative control model for production wells in this application is a model to be trained. Inputting state vectors of different states into the preset initial electric pump-based multi-parameter collaborative control model yields corresponding action vectors. These action vectors are then used to execute within a preset well mechanism model to obtain the state vector under the execution of the action vectors. For example, in step S102, this application inputs the first state vector of the current state into the preset initial electric pump-based multi-parameter collaborative control model to be trained, and then outputs the action vector and time difference error value of the current state. The time difference error value... In the formula, The value network output value in the preset initial electric pump production well multi-parameter collaborative control model is also known as the "current estimate". The target value is preset, that is, a "more credible target". Then, the action vector of the current state is input into the preset oil well mechanism model, and the third state vector of the next state under the action vector is output.

[0028] Step S103: Based on the execution deviations corresponding to the first state vector, the second state vector, and the third state vector, determine the reward value of the current state. The execution deviation is the deviation between the actual production and the expected production of the electric pump well in each state. It can be understood that the third state vector is the state vector of the electro-pumped well in the preset oil well mechanism model corresponding to the next state. The reward value is used to evaluate the impact of the action vector on the electro-pumped well in the preset oil well mechanism model.

[0029] Specifically, this application determines the execution deviation amount corresponding to the first state vector, the second state vector, and the third state vector respectively, and determines the reward value of the current state according to the execution deviation amount corresponding to the first state vector, the second state vector, and the third state vector respectively, so as to prepare for the subsequent formation of training samples for training the preset initial electric pump production well multi-parameter collaborative control model.

[0030] Specifically, determining the reward value of the current state based on the execution deviations corresponding to the first, second, and third state vectors includes: if the execution deviation corresponding to the third state vector is greater than a preset deviation threshold, determining the reward value of the current state based on the execution deviations corresponding to the first and second state vectors; if the execution deviation corresponding to the third state vector is less than or equal to the preset deviation threshold, determining the reward value of the current state based on the execution deviations corresponding to the first and third state vectors.

[0031] It is understood that the preset deviation threshold is a pre-defined deviation threshold, which can be 50. The state vector includes... In the formula, for The iteration step Layer-by-layer production deviation (production deviation of each layer segment) ,in For the first The actual liquid production of the layer at the current moment. For the first (Target liquid production rate of the current stage). for The iteration step The actual opening degree of the flow control valve ICV; for Pump efficiency deviation at any time (pump efficiency deviation of submersible electric pump) ,in This represents the current total lift volume. Current operating frequency (corresponding to the highest pump efficiency flow rate). for The operating frequency of the submersible electric pump at any given time; for The wellhead oil pressure of the electric pump-propelled well at all times.

[0032] Specifically, this application switches between two sets of reward calculation logics based on whether the execution deviation of the third state exceeds a preset deviation threshold, achieving differentiated reward value determination in stages: the coarse adjustment stage (the execution deviation corresponding to the third state vector is greater than the preset deviation threshold) and the steady-state fine maintenance stage (the execution deviation corresponding to the third state vector is less than or equal to the preset deviation threshold). For example, when the execution deviation of the third state is greater than the preset threshold (large deviation, in the rapid optimization coarse adjustment stage), the reward value is calculated based on the execution deviations corresponding to the current state and the previous state, capturing the continuous trend of deviation changes and achieving rapid coarse adjustment. When the execution deviation of the third state is less than or equal to the preset deviation threshold (deviation meets the standard, entering the steady-state fine control stage), the reward is constructed using the execution deviations corresponding to the current state and the next state, focusing on evaluating whether the state vector determined by the current action vector can continuously maintain a small deviation steady state or further optimize the operating conditions. Therefore, this application can use the deviation between actual production and expected production as the basis for reward calculation throughout the process. At the same time, the dual-mode adaptive switching reward mechanism can significantly improve the speed and accuracy of reward value determination, and meet the reward value calculation requirements in the training samples of the preset initial electric pump production well multi-parameter collaborative control model.

[0033] Step S104: Construct model training samples corresponding to the current state based on the first state vector of the current state, the action vector, the time difference error value of the action vector, the reward value, and the third state vector of the next state. Specifically, the training samples can be expressed in the following forms: ,in, The first state vector of the current state. The action vector for the current state. The reward value for the current state. Let be the third state vector of the next state. This represents the time difference error value for the current state. (In this training sample) Training samples with higher or lower time difference error values ​​are more valuable for training and learning, as they are used to characterize the specificity of the samples.

[0034] Step S105: Update the next state to the current state, and repeat the above iterative steps to redetermine the model training samples corresponding to the current state until a preset number of model training samples are obtained. Specifically, this application updates the next state to the current state and the previous current state to the previous state. Based on the preset initial multi-parameter collaborative control model of the electric pump production well to be trained, the action vector of the current state is determined according to the state vector corresponding to the current state. The action vector is then executed in the preset oil well mechanism model to obtain the state vector corresponding to the electric pump production well in the next state, so as to obtain training samples to update the next state to the current state and the previous current state to the previous state, until a preset number of model training samples are obtained.

[0035] Step S106: Train a preset initial multi-parameter collaborative control model for electric pump-separated wells based on a preset number of model training samples to obtain a trained multi-parameter collaborative control model for electric pump-separated wells.

[0036] It can be understood that a well-trained multi-parameter collaborative control model for electro-hydraulic pump-based production wells is a model that can be used to generate action vectors. Specifically, the initial preset multi-parameter collaborative control model for electro-hydraulic pump-based production wells includes a policy network and a value network, such as... Figure 3 As shown, the policy network employs a two-layer hidden layer structure: (1) Input layer: dimension is ( Production deviation + (ICV opening degree + pump efficiency deviation + frequency + oil pressure); (2) First hidden layer: 24 neurons, activation function is LeakyReLU ( To avoid gradient vanishing; (3) Second hidden layer: 36 neurons, activation function is LeakyReLU ( ); (4) Output layer: dimension is ( (ICV opening degree + frequency + oil pressure), activation function is tanh; (5) Learning rate: 0.03, using Adam optimizer; For value networks, a single hidden layer structure is adopted: (1) Input layer: dimension is (State vector dimension) +Motion Vector Dimension ); (2) Hidden layer: 60 neurons, activation function is ReLU; (3) Output layer: 1 neuron, outputs the value estimate of state-action pairs. ; (4) Learning rate: 0.05, using Adam optimizer.

[0037] The complete process of training a multi-parameter collaborative control model for a pre-set initial electric pump production well is exemplarily described here: (1) Initialization: Policy Network Value Network and its corresponding target network , ,make , Initialize the priority experience replay pool Capacity is 10000; set discount factor. Soft update coefficient Batch size Maximum step size in a single round Initial noise variance attenuation coefficient ; (2) Reset the simulation environment (preset oil well mechanism model), randomly initialize the reservoir IPR curve, ICV opening (0~100% uniform distribution), production target (randomly generated within the allowable range), electric pump frequency (30~60Hz uniform distribution) and wellhead oil pressure (1~3MPa uniform distribution) to obtain the initial state. ; (3) For each iteration step arrive : Select an action based on the current strategy. ,in ; Execute action Calculate the next state through a simulation environment With the comprehensive objective function value ; Calculate rewards , sample Store in the experience replay pool ,in ; like ,from Sampling by priority For each set of samples, calculate the target Q value:

[0038] Update the value network by minimizing the mean squared error loss:

[0039] In the formula, Importance sampling weights are used to correct biases introduced by prioritizing sampling.

[0040] The importance sampling index is set to 0.4; The policy network is updated using a deterministic gradient strategy.

[0041] Soft update target network parameters:

[0042] Update noise variance: ; (4) Repeat steps (2)-(3) above until the single-round reward is stable in the range of 30~45 for 100 consecutive times, and the overall objective function value is If the value is less than 10 for 50 consecutive times, the multi-parameter collaborative control model for electric pump-separated wells is considered to have converged.

[0043] The above technical solution obtains the first state vector of the electric pump production well in the current state and the second state vector in the previous state, determined based on a preset oil well mechanism model. This avoids the lag and misjudgment caused by relying solely on single-moment operating condition decisions, laying a data foundation for subsequent model training. Based on the preset initial multi-parameter collaborative control model of the electric pump production well to be trained, the action vector and the time difference error value of the action vector in the current state are determined according to the first state vector of the current state. The action vector is then executed in the preset oil well mechanism model to obtain the third state vector of the electric pump production well in the next state. The action vector can be issued and the subsequent state vector can be deduced within the preset oil well mechanism model simulation environment. The entire process does not require modification of the actual oil well production parameters on site, eliminating equipment wear caused by frequent valve and electric pump adjustments. The time difference error is calculated synchronously to achieve the technical effect of quantifying the learning value of a single sample. Based on the execution deviations corresponding to the first, second, and third state vectors, the reward value for the current state is determined. The deviation between the actual and expected production of the electric pump-propelled wells in each state, which is of most concern to the oilfield, can be used as the basis for reward construction. Using the execution deviations of the three states weakens the interference of instantaneous state disturbances. Based on the first state vector, action vector, time difference error value of the action vector, reward value, and the third state vector of the next state, a model training sample corresponding to the current state is constructed. It can be seen that the training sample completely contains all key information of the observed state, control actions, time-series errors, reward and punishment feedback, and subsequent new states, with complete sample feature dimensions. The next state is updated to the current state, and the above iterative steps are repeated to redetermine the model training sample corresponding to the current state until a preset number of model training samples are obtained. That is, this application uses a rolling iteration method to automatically generate massive and diverse training samples in batches, which can traverse various states without the need for manual collection of field data group by group, greatly improving sample construction efficiency while ensuring sufficient sample coverage. This application trains a pre-set initial multi-parameter collaborative control model for electric submersible pump (ESP) wells based on a pre-set number of model training samples to obtain a well-trained ESP well multi-parameter collaborative control model. The model training can be completed using complete samples. After training, the ESP well multi-parameter collaborative control model can autonomously output a complete set of continuous control commands for the flow control valve opening (ICV opening), the ESP frequency, and the wellhead oil pressure of the ESP well, eliminating the reliance on repeated manual tuning of control parameters. This enables accurate control of the ESP well action vector, thereby achieving accurate control of the fluid production of each production layer in the ESP well. Furthermore, this application can use the deviation between actual production and expected production as the basis for reward calculation throughout the process. Simultaneously, the dual-mode adaptive switching reward mechanism can significantly improve the speed and accuracy of reward value determination, meeting the reward value calculation requirements of the pre-set initial ESP well multi-parameter collaborative control model training samples.

[0044] In this embodiment of the application, when the execution deviation corresponding to the third state vector is greater than a preset deviation threshold, the reward value of the current state is determined based on the execution deviation corresponding to the first state vector and the execution deviation corresponding to the second state vector, including determining the reward value of the current state according to the following formula:

[0045] in, The reward value for the current state. This is the execution deviation corresponding to the first state vector. This is the execution deviation corresponding to the second state vector. a This is the preset reward coefficient.

[0046] It is understandable that the preset reward coefficient is a pre-set reward coefficient, which can be 0.2.

[0047] Specifically, when the execution deviation corresponding to the third state vector is greater than 50 ( (This can be done by calculating reward and penalty values ​​based on the change in execution deviation before and after, for example: according to the formula...) Calculate reward value A coefficient of 0.2 is used to balance the reward magnitude. If This indicates that the action is approaching the optimal solution, and a positive reward is given; if This indicates that the action deviates from the optimal solution, and a negative penalty is imposed.

[0048] In existing technologies, many methods for determining reward values ​​can only provide a qualitative "good or bad" judgment, or use a fixed reward value. This non-quantitative or fixed feedback is rather coarse, making it difficult for the model to determine the degree of "good" and thus hindering fine-tuning. However, in the embodiments of this application, the reward value... The magnitude of the reward directly reflects the extent of progress or regression. Greater progress results in greater rewards; greater regression incurs heavier penalties. This "distribution according to work" quantitative mechanism guides the multi-parameter collaborative control model of the electric pump-separated well to take more refined and optimal adjustment actions in pursuit of higher rewards when approaching the optimal solution. More importantly, this application can classify the determination of reward values ​​based on the execution deviation corresponding to the third state vector. Here, when the execution deviation corresponding to the third state vector is greater than a preset deviation threshold, the change in execution deviation is significant. This application can then "distribute according to work" to effectively guide the multi-parameter collaborative control model of the electric pump-separated well to maintain and further optimize the optimal operating conditions.

[0049] In this embodiment of the application, when the execution deviation corresponding to the third state vector is less than or equal to a preset deviation threshold, the reward value of the current state is determined based on the execution deviation corresponding to the first state vector and the execution deviation corresponding to the third state vector. This includes: when the execution deviation corresponding to the first state vector is greater than the execution deviation corresponding to the third state vector, a preset fixed positive reward value is used as the reward value of the current state, wherein the preset fixed positive reward value is a value greater than zero; when the execution deviation corresponding to the first state vector is less than or equal to the execution deviation corresponding to the third state vector, a preset fixed negative reward value is used as the reward value of the current state, wherein the preset fixed negative reward value is a value less than zero.

[0050] It is understandable that the preset fixed positive reward value is a pre-set reward value greater than zero, which can be 10. The preset fixed negative reward value is a pre-set reward value less than zero, which can be -15.

[0051] Specifically, when The magnitude of the deviation narrows significantly, indicating that the basic reward and punishment rules are unable to effectively guide the preset initial electric pump production well multi-parameter collaborative control model to maintain and further optimize the optimal operating conditions. If the action vector output by the preset initial electric pump production well multi-parameter collaborative control model is executed, the system maintains the optimal operating conditions or further approaches the global optimal solution, and a preset fixed positive reward value is given. 10. If the system deviates from the optimal solution after executing the action vector output by the preset initial electric pump production well multi-parameter collaborative control model, a preset fixed negative reward value will be given. -15, enhances the ability of the preset initial electric pump production well multi-parameter collaborative control model to maintain and optimize the optimal operating conditions.

[0052] In existing technologies, reward values ​​are often not differentiated based on the execution deviation value of the next state. This application employs state-based classification rewards and penalties. However, when the system approaches the optimal solution (the execution deviation value of the next state has decreased to below 50), the change in execution deviation value becomes extremely small. At this point, the basic reward and penalty based on the minute change in execution deviation value becomes extremely weak (e.g., 0.02). These signals are buried in random environmental noise, causing the multi-parameter collaborative control model of the electric pump-separated well to "spin in place" near the optimal solution, making convergence difficult. This application actively cuts off the dependence on minute changes by providing larger preset fixed positive and negative reward values. It no longer focuses on how much the multi-parameter collaborative control model of the electric pump-separated well has "improved," but only on whether the model maintains or approaches the optimal region. This provides the AI ​​with a clear and strong "good / bad" signal (+10 / -15) in the optimal region, directly solving the problem of weak or even absent learning signals in the near-optimal region.

[0053] In this embodiment, the first state vector, the second state vector, and the third state vector include the total pumping volume deviation and the product volume deviation of each layer. The execution deviation is determined according to the following formula:

[0054] in, To implement the deviation amount, To preset the balance coefficient, To preset the pump efficiency penalty coefficient, This is due to the deviation in total pumped liquid volume. For the first The preset production penalty coefficient corresponding to the layer To account for the deviation in liquid production at each layer, This represents the total number of floors.

[0055] It can be understood that the preset balance coefficient is a pre-set coefficient used to balance the deviation of the total pumped fluid volume and the deviation of the fluid production volume of each layer, and can be 0.4. The preset pump efficiency penalty coefficient is a pre-set penalty coefficient used to characterize whether the total fluid production volume is within or outside the working displacement constraint range of the submersible electric pump. The preset production allocation penalty coefficient is a pre-set penalty coefficient used to characterize whether the fluid production volume of the layer is within or outside the production allocation constraint range or the erosion constraint range of the ICV control valve. The total pumped fluid volume deviation is the absolute value of the difference between the current total pumped fluid volume obtained according to the preset well mechanism model and the preset target total pumped fluid volume. The fluid production volume deviation of each layer is the absolute value of the difference between the current fluid production volume of each layer obtained according to the preset well mechanism model and the corresponding preset target fluid production volume of each layer.

[0056] Specifically, The execution deviation to be minimized; The balance coefficient between the pump efficiency target and the production target can be taken as 0.4; To pre-determine the pump efficiency penalty coefficient, when the total fluid production is outside the operating displacement range of the submersible electric pump, ,on the contrary ; To pre-determine the production penalty coefficient, when the liquid production of the layer is outside the production constraint range or the control valve ICV erosion constraint range. on the contrary This application can calculate the execution deviation of the first state vector, the second state vector, and the third state vector using the above formula, which is then used to calculate the reward value.

[0057] In existing technologies, traditional stratified oil recovery control typically employs a step-by-step, cyclical control strategy. For example, the frequency of the electric pump is first adjusted to meet the total fluid volume requirement, and then the production distributors for each stratum are adjusted one by one to reduce single-stratum errors. This "divide and conquer" approach is prone to causing problems where one aspect is addressed at the expense of another, potentially resulting in situations where the total fluid volume meets the target but the single-stratum error is large, or where the single-stratum error is small but the electric pump deviates from its high-efficiency zone.

[0058] This application uses the above formula to perform a weighted summation, thus binding these two originally separate objectives together. It also considers the total liquid volume deviation. (Representing the overall lifting efficiency of the electric pump) and the deviation of liquid volume in each layer (This represents the accuracy of stratified production allocation). This enables the multi-parameter collaborative control model of the electric pump production well to be trained to learn a globally optimal coordination strategy, rather than a locally optimal compromise.

[0059] A second aspect of this application provides a method for determining the action vector of an electric submersible pump (ESP) well. The method includes: acquiring a predicted state vector of the ESP well in its current state, the predicted state vector including the wellhead oil pressure, the operating frequency of the ESP in the ESP well, the opening degree of the flow control valve in the ESP well, the production deviation of each layer in the reservoir where the ESP well is located, and the total pumping volume deviation of the ESP; and determining the target action vector of the ESP well based on a trained multi-parameter collaborative control model of the ESP well, according to the predicted state vector, wherein the target action vector includes the expected wellhead oil pressure, the expected ESP frequency, and the expected control valve opening degree, and the trained multi-parameter collaborative control model of the ESP well is trained according to the training method of the ESP well multi-parameter collaborative control model described above.

[0060] It can be understood that the state vector to be predicted is the state vector collected from the electric submersible pump (ESP) well under the current state, and the target action vector is the action vector obtained by inputting the state vector to be predicted into the trained multi-parameter collaborative control model of the ESP well. The state vector to be predicted includes the wellhead oil pressure of the ESP well, the operating frequency of the ESP in the ESP well, the opening degree of the flow control valve in the ESP well, the production deviation of each layer in the reservoir where the ESP well is located, and the total pumping volume deviation of the ESP. The specific formula is as follows:

[0061] In the formula, for The state step Deviation in liquid production of the layer; for The state step The actual opening degree of the layer control valve; for Total pumped fluid volume deviation at any given time; for The operating frequency of the electric pump at any given time; for The wellhead oil pressure at any given time.

[0062] The deviation of liquid production in each stage is as follows: ,in For the first The actual fluid production of the current fluid layer. The total pumping rate deviation of the submersible electric pump is... ,in This represents the current total lift volume. Current operating frequency The corresponding maximum pump efficiency flow rate is calculated by the submersible electric pump lifting characteristic model in the preset oil well mechanism model.

[0063] Specifically, this application inputs the state vector to be predicted into the trained multi-parameter collaborative control model of the electric pump production well, which can determine the target action vector of the electric pump production well. The trained multi-parameter collaborative control model of the electric pump production well can be used to determine the action vector corresponding to the state vector, so that the action vector can be used for actual production control, so as to achieve accurate control of the fluid production of each production layer of the electric pump production well.

[0064] In this embodiment of the application, the method further includes: determining whether the target action vector of the electric pump submersible well meets the preset constraint conditions, wherein the preset constraint conditions include at least the control valve erosion constraint, the control valve opening constraint, and the submersible electric pump operating frequency constraint. The erosion constraint of the control valve satisfies the following formula:

[0065] in, The actual flow rate passing through the flow control valve. The valve flow area of ​​the flow control valve is The corresponding erosion flow rate at that time For valve flow area, The erosion flow coefficient is... For the first Density of the layer produced liquid The pressure drop of the control valve for the flow control valve; Among them, the control valve opening constraint is that the expected control valve opening is within the range of the preset maximum control valve opening and the preset minimum control valve opening; the submersible pump operating frequency constraint is that the expected pump frequency is within the range of the preset maximum pump frequency and the preset minimum pump frequency; when the target action vector of the submersible pump production well does not meet the preset constraint conditions, the target action vector is corrected based on the multi-parameter collaborative control model of the submersible pump production well.

[0066] Specifically, the formulaic expressions for the above-mentioned preset constraints (control valve erosion constraint, control valve opening constraint, and submersible pump operating frequency constraint) are as follows: (1) Control valve opening constraint: expected control valve opening The valve must satisfy its own physical adjustment range constraints, the mathematical expression of which is:

[0067] When ICV is At that time, the valve is fully closed; ICV is At that time, the valve was fully open.

[0068] (2) Control valve erosion constraint: Excessive flow rate can easily lead to erosion failure of the ICV throttle valve, or even valve failure. To suppress the erosion effect caused by fluid acceleration at the valve orifice, the following erosion resistance safety constraints are set for the control valve:

[0069] in, The valve flow area is At that time, the erosion flow rate corresponding to the control valve, and the produced fluid of the layer. Should be smaller . The flow area of ​​the control valve can be calculated using the valve opening degree and structural parameters. The erosion flow coefficient, according to industry standards, typically ranges from 0.6 to 0.8. This represents the actual flow rate through the control valve.

[0070] (3) Operating frequency constraint of submersible electric pump: that is, the operating frequency of submersible electric pump is 30Hz~60Hz.

[0071]

[0072] f This refers to the expected operating frequency of the electric pump.

[0073] Additionally, it is worth noting that the preset constraints also include production allocation constraints and submersible electric pump displacement range constraints, for example: (1) Production allocation constraints: Based on the tiered production allocation scheme, i.e., given tiered production targets. Actual output liquid of each stage Production must meet the matching requirements, and the matching error must not exceed [a certain threshold]. ,Right now:

[0074] In the formula, and These are the upper and lower limits for production allocation; To allow for the maximum relative production error, Oilfield production units typically require production matching errors to be no more than 15%.

[0075] (2) Submersible ESP displacement range constraints: The operating point of the submersible ESP must be within the allowable operating range between the upper and lower flow limits. Within this range, the ESP can maintain high operating efficiency. If the operating point exceeds this range, the pump efficiency will be significantly reduced, and long-term operation will easily lead to equipment failure. Based on the operating frequency and pumping flow rate, the submersible ESP displacement range constraints are as follows:

[0076] in, Indicates the operating frequency is The lifting fluid volume (pumping fluid volume) of the submersible electric pump. and Operating frequency The corresponding lower and upper limits of allowed traffic.

[0077] It is worth noting that the application of preset constraints can first be used for the validity verification of action vectors. For example, the state vector can be used to... Input the converged bias feature pre-extraction DDPG policy network and perform the following inference and verification process: 1. Model Reasoning: Input the state vector into the preset initial electric pump production well action vector. The output model outputs normalized action vectors. The initial control action is obtained by mapping the control action to the actual control range through linear transformation. ; Specifically, the output model outputs normalized action vectors. Mapping to the actual control range through linear transformation includes introducing Gaussian noise into the initial control action. To achieve behavioral exploration, the noise variance decreases exponentially with the training process: In the formula, The initial noise variance can be set to 0.2; The attenuation coefficient can be set to 0.001. The training iterations are specified; the final policy network action output is: ,in, This is the mapping function for the strategy network in the preset initial electric pump-based well action vector output model. These are the parameters of the policy network. The output layer of the policy network uses an activation function to normalize the actions to the [-1, 1] interval to obtain... Then, a linear transformation is used to map it to the actual control range: In the formula, The normalized action value; , These are the upper and lower limits of the corresponding control parameters.

[0078] 2. Hard constraint verification: Verify one by one whether the control action meets all the preset constraints, including control valve opening constraints, electric pump operating frequency constraints, control valve erosion constraints, and submersible electric pump displacement range constraints; if there are any violations, correct them to within the constraint boundaries and mark the violation type.

[0079] In addition, the aforementioned preset constraints can also be used for pre-constraint verification to detect the input state vector to the preset initial electric pump production well action vector. Does the preset constraint condition meet? For example, if the current state has triggered the control valve erosion constraint, the system will immediately enter the emergency control mode, prioritize adjusting the parameters to the safe range, and then execute the subsequent optimization process.

[0080] Optionally, the model, after real-time production data is input and converged, outputs the optimal joint adjustment scheme for ICV opening, electric pump operating frequency, and wellhead oil pressure. After execution, production data is collected to update the environmental status, realizing dynamic closed-loop control. The specific process is as follows: (1) Synchronous acquisition of multi-source real-time data Three types of dynamic and static data are collected synchronously through downhole permanent intelligent sensors, surface SCADA systems, and reservoir management platforms, with a sampling frequency of once per minute and a data transmission delay of no more than 10 seconds: Downhole production data: fluid production, water cut, temperature, and pressure for each formation; lift system operation data: submersible electric pump operating frequency, input power, outlet pressure, pump inlet pressure, and motor temperature; wellhead and surface data: wellhead oil pressure, casing pressure, wellhead temperature, wellhead flow rate, and ambient temperature; simultaneously, the currently effective stratified production targets are obtained from the reservoir management platform. Rated characteristic parameters of submersible electric pumps.

[0081] (2) Standardized data preprocessing The collected raw data were processed sequentially as follows to ensure the data quality of the input model: Outlier removal: The sliding 3σ criterion was used to identify outlier data points, with a sliding window size of 5 sampling points, and outliers exceeding the mean ± 3σ range were removed; Noise filtering: A 5-point moving average filtering method was used to smooth the data and suppress high-frequency noise interference from the sensors; Missing value completion: For single missing sampling points, linear interpolation was used to complete the missing data. When three or more consecutive sampling points were missing, the weighted average of the first five valid sampling points was used to complete the missing data, and the data quality level was marked; Unit standardization: All physical quantities were uniformly converted to the International System of Units (SI) to maintain consistency with the unit system used in the simulation environment and model training.

[0082] (3) Perform a validity check based on the preset constraints; (4) Automatic execution of the multi-parameter joint control scheme. The control scheme is executed by the ground intelligent control system in the order of downhole first, then hoisting, and then wellhead. This can reduce the impact of wellbore flow fluctuations on the control effect. The execution interval is 30 seconds: First step: Adjust the ICV opening of each section to the target value in sequence. This is achieved through a downhole electric actuator; the second step is to adjust the operating frequency of the submersible pump to the target value. Stepless adjustment is achieved through a frequency converter; the third step is to adjust the wellhead oil pressure to the target value. This is achieved by adjusting the opening of the ground throttle valve; feedback is executed: after each parameter adjustment is completed, the actual value of the parameter is collected and compared with the target value. If the deviation exceeds the allowable range (ICV opening deviation > 2%, frequency deviation > 0.1Hz, oil pressure deviation > 0.05MPa), a second fine adjustment is performed.

[0083] (5) Multi-dimensional verification of control effect. After all control schemes are executed, wait 10 minutes (the dynamic response lag time of matching the multiphase flow in the wellbore) and re-collect production data to verify the control effect from three dimensions: production allocation accuracy verification: calculate the actual production allocation deviation of each layer, requiring that the deviation of all layers does not exceed 5%; lift efficiency verification: calculate the actual operating pump efficiency of the submersible electric pump, requiring that the pump efficiency is not less than 60%; equipment safety verification: calculate the ICV erosion risk index of each layer, requiring that the erosion flow rate is lower than the threshold. The electric pump operates within the allowable displacement range.

[0084] (6) Dynamic closed-loop triggering and periodic update. Different closed-loop logics are executed according to the verification results: if all verification indicators meet the requirements, enter steady-state monitoring mode, wait for the next regular sampling cycle (1 minute), and repeat steps (1)-(5); if any verification indicator does not meet the requirements, immediately trigger the model to re-optimize, recalculate the state vector according to the target production of the i-th layer, and output a new control scheme; if the requirements cannot be met after 3 consecutive optimizations, trigger manual intervention warning; when the reservoir management platform issues a new stratified production target, there is no need to retrain the model, and the target production of the i-th layer is directly updated. This allows for the output of adjustment actions adapted to the new production plan within 5 iterations.

[0085] It is worth noting that traditional methods typically involve setting a fixed maximum flow rate for the control valve. This approach is overly simplistic because it fails to consider the effects of valve opening and actual pressure drop. Under high pressure differentials, the risk of erosion can be extremely high even if the flow rate does not reach the fixed threshold; conversely, the risk can also be high if the flow rate does not reach the fixed threshold. However, this application, by introducing a precise physical formula, transforms the complex erosion mechanism into a quantifiable, dynamic safety boundary. This boundary can be automatically adjusted according to the real-time operating status of the valve, thus providing a more accurate and reliable basis for safety judgment than the static threshold.

[0086] Furthermore, traditional methods tend to be passive. When multiple constraints exist simultaneously and conflict with each other (e.g., closing the valve to meet erosion constraints but opening it to meet production targets), traditional methods struggle to make global trade-offs and coordination, potentially leading to suboptimal compromises or even failing to satisfy all constraints simultaneously. This application, however, incorporates multiple core constraints, such as control valve erosion constraints, control valve opening constraints, and electric pump operating frequency constraints, as conditions that the multi-parameter collaborative control model for electric pump production wells must consider simultaneously during decision-making (outputting the target action vector). This enables the multi-parameter collaborative control model for electric pump production wells to automatically perform multi-objective, multi-constraint joint optimization at the moment of action output, intelligently selecting an optimal solution that maximizes production while ensuring all safety and equipment red lines are not violated.

[0087] In this embodiment of the application, the method further includes: obtaining the current production volume of each layer in the reservoir determined based on a preset oil well mechanism model; determining the absolute value of the difference between the current production volume of each layer and the preset target production volume corresponding to each layer, so as to obtain the production volume deviation of each layer in the reservoir where the electric submersible pump production well is located; obtaining the current total pumping volume of the submersible electric pump determined based on the preset oil well mechanism model; and determining the absolute value of the difference between the current total pumping volume and the preset target total pumping volume, so as to obtain the total pumping volume deviation of the submersible electric pump.

[0088] It is understandable that the pre-defined oil well mechanism model includes: (1) Calculation model of multiphase flow in wellbore

[0089] in, The pressure of the gas-liquid two-phase fluid inside the wellbore, in Pa; The temperature of the gas-liquid two-phase fluid inside the wellbore, in K; The distance traveled axially by the fluid in the wellbore, in meters (m). , For the density of the fluid in the gaseous and liquid phases, in kg / m³; , , where represents the apparent flow velocity of gaseous and liquid fluids in the wellbore, in m / s; , The gas holdup and liquid holdup of the wellbore cross section are given. . and Let kg / s represent the conversion rates from liquid to gas and from gas to liquid, respectively. Let be the cross-sectional area of ​​the well shaft, in m2. Let be the wellbore inclination angle, in rad; Gas-liquid two-phase slippage liquid carrying capacity of wellbore cross section, % The inner diameter of the well shaft is in meters (m). The flow rate of the mixture is expressed in m / s. ρ is the density of the mixture, in kg / m³. The Fanning friction coefficient of the wellbore is calculated using the Chen relation. The ambient temperature, in K; Joule-Thomson coefficient, K / Pa; , Let J be the specific heat capacity of the gas phase and liquid phase, respectively, in kg / (kg·K). The gas content of the wellbore cross-section is %. The overall heat transfer coefficient of the wellbore is W / (m2·K); Where is the outer diameter of the well shaft, in meters; The total mass flow rate is expressed in kg / s. is the specific heat capacity of the mixture, J / (kg·K). The hydraulic gradient is measured in m / m.

[0090] (2) Calculation of pressure drop of downhole flow control valve ICV

[0091] in, For the first The flow resistance coefficient of the ICV valve orifice is related to the valve opening and the flow area, and can be determined by indoor flow resistance experiments. For the first ICV opening degree,%. For the first Density of the layer's produced liquid, kg / m3; The flow area of ​​the ICV valve port is determined by the structure of the ICV valve body.

[0092] (3) The stratified inflow model uses a two-parameter equation to describe the relationship between the product flow rate and the production pressure difference in each stage: No. The capacity description for each segment is as follows:

[0093] in, The product volume is m³ / d. The product flow rate is 0 m³ / d when the flow pressure is 0. The pressure is the flow pressure, in MPa. The static pressure of the formation is MPa. , This is the production capacity coefficient, dimensionless.

[0094] (4) Submersible electric pump lifting characteristic model

[0095] in, Head at the reference frequency ; For power, ; For pump efficiency, . To increase the total liquid volume by the pump, ; , c To fit parameters using experimental or field production data, after providing the pump characteristic curve at the rated frequency, the pump characteristic parameters at other frequencies can be calculated using the formula.

[0096]

[0097] The operating frequency of the submersible electric pump, . As the reference frequency, .

[0098] Specifically, this application determines the current production rate of each layer in the reservoir by inputting pre-processed parameters for determining production rate into a preset well mechanism model. Then, based on the absolute value of the difference between the current production rate of each layer and the corresponding preset target production rate, the production rate deviation of each layer in the reservoir where the electric submersible pump (ESP) well is located is obtained. Similarly, by inputting pre-processed parameters for determining the total pumping rate into the preset well mechanism model, the current total pumping rate of the ESP is determined. Then, based on the absolute value of the difference between the current total pumping rate and the preset target total pumping rate, the total pumping rate deviation of the ESP is obtained.

[0099] In this embodiment, the construction of the pre-set oil well mechanism model avoids repeated adjustments to actual production parameters. By determining the production rate deviation and total pumping rate deviation of each layer in the reservoir, a massive amount of training samples can be generated to train the pre-set initial electric pump-propelled well multi-parameter collaborative control model. Therefore, the advantages and positive effects of this application include at least the following: (1) High control precision: The global coordinated optimization of ICV opening, electric pump frequency and wellhead oil pressure is achieved. Field verification of the four-layer electric pump production well in Bohai Oilfield shows that the production deviation of each layer after optimization is controlled within 3%, and the difference between the pump efficiency of the submersible electric pump and the optimal pump efficiency is no more than 1.5%.

[0100] (2) Low equipment wear: The optimal ICV opening is determined in one go through global optimization, avoiding repeated fine-tuning, reducing the risk of valve port erosion by more than 35%, extending the average service life of ICV by 1.2 years, and reducing the number of well repair operations and costs.

[0101] (3) Strong generalization ability: The model can be adapted to electric pump production wells with 3 to 6 layers. When the production allocation scheme is dynamically adjusted, there is no need to retrain. Only the production allocation deviation characteristics need to be updated to output a new control scheme within 5 iterations.

[0102] (4) High degree of automation: It replaces the traditional manual experience control, shortens the single well allocation cycle from 5-8 hours to less than 30 minutes, and significantly improves the level of refined mining of sub-mining wells.

[0103] A third aspect of this application provides a training apparatus for a multi-parameter collaborative control model of an electric pump-propelled well based on deviation characteristics, comprising: a memory configured to store instructions; and a processor configured to retrieve instructions from the memory and, when executing the instructions, to implement the aforementioned training method for the multi-parameter collaborative control model of an electric pump-propelled well based on deviation characteristics.

[0104] A fourth aspect of this application provides an apparatus for determining the motion vector of an electric pump-operated well, comprising: a memory configured to store instructions; and a processor configured to retrieve instructions from the memory and, when executing the instructions, to implement the apparatus for determining the motion vector of an electric pump-operated well as described above.

[0105] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0106] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A training method for a multi-parameter collaborative control model of an electric pump production well based on deviation characteristics, characterized in that, The training method includes: Obtain the first state vector of the electric pump production well in the current state and the second state vector in the previous state, which are determined based on the preset oil well mechanism model; Based on the preset initial multi-parameter collaborative control model of the electric pump production well to be trained, the action vector of the current state and the time difference error value of the action vector are determined according to the first state vector of the current state. The action vector is then executed in the preset oil well mechanism model to obtain the third state vector of the electric pump production well in the next state. The action vector includes the wellhead oil pressure of the electric pump production well, the electric pump frequency of the submersible electric pump installed in the electric pump production well, and the control valve opening of the flow control valve installed in the electric pump production well. Based on the execution deviation amount corresponding to the first state vector, the second state vector, and the third state vector, the reward value of the current state is determined. The execution deviation amount is the deviation between the actual production and the expected production of the electric pump well in each state. The first state vector, the second state vector, and the third state vector include the total pumping volume deviation and the production volume deviation of each layer. The model training samples corresponding to the current state are constructed based on the first state vector of the current state, the action vector, the time difference error value of the action vector, the reward value, and the third state vector of the next state. Update the next state to the current state, and repeat the above iterative steps to redetermine the model training samples corresponding to the current state until a preset number of model training samples are obtained. Based on the preset number of model training samples, train the preset initial electric pump production well multi-parameter collaborative control model to obtain the trained electric pump production well multi-parameter collaborative control model. The step of determining the reward value of the current state based on the execution deviations corresponding to the first state vector, the second state vector, and the third state vector includes: If the execution deviation corresponding to the third state vector is greater than a preset deviation threshold, the reward value of the current state is determined based on the execution deviation corresponding to the first state vector and the execution deviation corresponding to the second state vector. If the execution deviation corresponding to the third state vector is less than or equal to the preset deviation threshold, the reward value of the current state is determined based on the execution deviation corresponding to the first state vector and the execution deviation corresponding to the third state vector.

2. The training method according to claim 1, characterized in that, When the execution deviation corresponding to the third state vector is greater than a preset deviation threshold, the reward value of the current state is determined based on the execution deviation corresponding to the first state vector and the execution deviation corresponding to the second state vector, including determining the reward value of the current state according to the following formula: in, The reward value for the current state. This is the execution deviation corresponding to the first state vector. This is the execution deviation corresponding to the second state vector. a This is the preset reward coefficient.

3. The training method according to claim 1, characterized in that, When the execution deviation corresponding to the third state vector is less than or equal to the preset deviation threshold, the reward value of the current state is determined based on the execution deviation corresponding to the first state vector and the execution deviation corresponding to the third state vector, including: If the execution deviation corresponding to the first state vector is greater than the execution deviation corresponding to the third state vector, a preset fixed positive reward value is used as the reward value of the current state, wherein the preset fixed positive reward value is a value greater than zero. If the execution deviation corresponding to the first state vector is less than or equal to the execution deviation corresponding to the third state vector, a preset fixed negative reward value is used as the reward value for the current state, wherein the preset fixed negative reward value is a value less than zero.

4. The training method according to claim 1, characterized in that, The execution deviation is determined according to the following formula: in, The execution deviation amount, To preset the balance coefficient, To preset the pump efficiency penalty coefficient, This refers to the deviation in the total pumped liquid volume. For the first The preset production penalty coefficient corresponding to the layer The deviation in the liquid production of each layer, This represents the total number of floors.

5. A method for determining the motion vector of an electric pump-propelled well, characterized in that, The method includes: Obtain the predicted state vector of the electric pump production well in its current state. The predicted state vector includes the wellhead oil pressure of the electric pump production well, the electric pump operating frequency of the submersible electric pump in the electric pump production well, the control valve opening of the flow control valve in the electric pump production well, the production volume deviation of each layer in the reservoir where the electric pump production well is located, and the total pumping volume deviation of the submersible electric pump. Based on the trained multi-parameter collaborative control model of the electric pump production well, the target action vector of the electric pump production well is determined according to the state vector to be predicted. The target action vector includes the expected wellhead oil pressure, the expected electric pump frequency, and the expected control valve opening. The trained multi-parameter collaborative control model of the electric pump production well is trained according to the training method of the multi-parameter collaborative control model of the electric pump production well based on deviation characteristics as described in any one of claims 1 to 4.

6. The method according to claim 5, characterized in that, The method further includes: Determine whether the target action vector of the electric pump submersible well meets the preset constraints, wherein the preset constraints include at least the control valve erosion constraint, the control valve opening constraint, and the submersible pump operating frequency constraint. The erosion constraint of the control valve satisfies the following formula: in, The actual flow rate passing through the flow control valve. The valve flow area of ​​the flow control valve is The corresponding erosion flow rate at that time For valve flow area, The erosion flow coefficient is... For the first Density of the layer produced liquid The pressure drop of the control valve for the flow control valve; Wherein, the control valve opening constraint is that the expected control valve opening is within the range of a preset maximum control valve opening and a preset minimum control valve opening; The operating frequency constraint of the submersible pump is that the expected pump frequency is within the range of a preset maximum pump frequency and a preset minimum pump frequency. If the target action vector of the electric pump production well does not meet the preset constraint conditions, the target action vector is corrected based on the multi-parameter collaborative control model of the electric pump production well.

7. The method according to claim 5, characterized in that, The method further includes: Obtain the current fluid production of each layer in the reservoir, as determined by a preset oil well mechanism model; Determine the absolute value of the difference between the current production volume of each layer and the preset target production volume of each layer, so as to obtain the production volume deviation of each layer in the reservoir where the electric pump production well is located; Obtain the current total pumping volume of the submersible electric pump determined based on a preset oil well mechanism model; The absolute value of the difference between the current total pumping volume and the preset target total pumping volume is determined to obtain the total pumping volume deviation of the submersible electric pump.

8. A training device for a multi-parameter collaborative control model of an electric pump production well based on deviation characteristics, characterized in that, include: The memory is configured to store instructions; as well as A processor is configured to retrieve the instructions from the memory and, when executing the instructions, to implement the training method for a multi-parameter collaborative control model of an electric pump production well based on deviation characteristics, according to any one of claims 1 to 4.

9. A device for determining the motion vector of an electric pump-propelled well, characterized in that, include: The memory is configured to store instructions; as well as The processor is configured to retrieve the instructions from the memory and, when executing the instructions, to implement the method for determining the action vector of an electric pumped well according to any one of claims 5 to 7.

Citation Information

Patent Citations

  • Real-time optimization regulation and control method for oil reservoir layered injection-production scheme

    CN120020332A

  • Regulation and control method, device and equipment of electric submersible pump layered producing well simulation model and medium

    CN120124491A