Multi-device collaborative control method and device for drilling mud processing
By establishing a multi-equipment collaborative control model, using deep reinforcement learning and closed-loop feedback control, the problem of multi-equipment collaborative control in drilling mud treatment is solved, dynamic adjustment and adaptive control are realized, and processing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510164161.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-02-14
AI Technical Summary
Traditional drilling mud treatment methods lack multi-equipment collaborative control mechanism, resulting in low processing efficiency and high energy consumption, making it difficult to achieve precise control under complex working conditions.
By obtaining the parameters of the mixing tank, precipitation tank and oil-water separation equipment, establishing steady-state and dynamic control equations, using deep reinforcement learning to build models, implementing parallel branch processing and dynamic adjustment of the equipment, and designing a closed-loop feedback control mechanism.
Dynamic adjustment and adaptive control of drilling mud equipment are realized, processing efficiency is improved, energy consumption is reduced, and parameter adjustment is ensured.
Smart Images

Figure CN119620625B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-device collaborative control, and in particular to a multi-device collaborative control method and device for drilling mud processing. Background Art
[0002] Drilling mud treatment systems are a critical component of oil drilling projects, and their effectiveness directly impacts the safety and efficiency of drilling operations. Traditional drilling mud treatment methods rely primarily on the independent control of individual devices, lacking a coordinated control mechanism for mixing tanks, settling tanks, and oil-water separation equipment. This results in low treatment efficiency and high energy consumption.
[0003] As drilling projects move toward complex formations and deeper wells, mud handling systems face even more challenging challenges. Existing parameter adjustment solutions rely primarily on manual adjustments based on experience. These adjustments require simultaneous consideration of the operating status of multiple devices and changes in mud properties. This makes precise control difficult under complex real-world conditions and fails to meet the demands of real-time decision-making. Summary of the Invention
[0004] The present invention provides a multi-device collaborative control method and device for drilling mud processing, which solves the coupling control problem between multiple devices and realizes dynamic adjustment and adaptive control of multiple drilling mud devices.
[0005] In a first aspect, the present invention provides a multi-device collaborative control method for drilling mud processing, the multi-device collaborative control method for drilling mud processing comprising:
[0006] Obtain the stirring parameters of the mixing tank equipment, the settling time parameters of the settling tank equipment, and the separation efficiency parameters of the oil-water separation equipment, and obtain the initial collaborative control parameter set by simultaneously solving the steady-state control equation and the dynamic adjustment equation;
[0007] Based on the initial collaborative control parameter set, the mixing tank device, the sedimentation tank device, and the oil-water separation device are divided into branch control architectures and processed in parallel to obtain a branch parallel control parameter set;
[0008] Performing deep reinforcement learning training based on the branch parallel control parameter set, constructing a reinforcement learning model including a state space, an action space, and a reward function, and generating an optimal parameter adjustment strategy;
[0009] Based on the optimal parameter adjustment strategy, the mud density data and mud viscosity data at the outlet of the mixing tank equipment, the sand content data at the outlet of the sedimentation tank equipment, and the oil content data at the outlet of the oil-water separation equipment are monitored and threshold-triggered, and dynamic adjustment control instructions are output.
[0010] In a second aspect, the present invention provides a multi-device collaborative control device for drilling mud processing, the multi-device collaborative control device for drilling mud processing comprising:
[0011] An acquisition module is used to obtain the stirring parameters of the mixing tank equipment, the settling time parameters of the sedimentation tank equipment, and the separation efficiency parameters of the oil-water separation equipment, and obtain the initial collaborative control parameter set by simultaneously solving the steady-state control equation and the dynamic adjustment equation;
[0012] a parallel processing module for performing branch control architecture division and parallel processing on the mixing tank device, the sedimentation tank device, and the oil-water separation device based on the initial collaborative control parameter set to obtain a branch parallel control parameter set;
[0013] A reinforcement learning module is used to perform deep reinforcement learning training based on the branch parallel control parameter set, build a reinforcement learning model including a state space, an action space, and a reward function, and generate an optimal parameter adjustment strategy;
[0014] The output module is used to monitor and perform threshold triggering processing on the mud density data and mud viscosity data at the outlet of the mixing tank equipment, the sand content data at the outlet of the sedimentation tank equipment, and the oil content data at the outlet of the oil-water separation equipment based on the optimal parameter adjustment strategy, and output dynamic adjustment control instructions.
[0015] The third aspect of the present invention provides a multi-device collaborative control device for drilling mud processing, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the multi-device collaborative control device for drilling mud processing executes the above-mentioned multi-device collaborative control method for drilling mud processing.
[0016] A fourth aspect of the present invention provides a computer-readable storage medium having instructions stored therein, which, when executed on a computer, enables the computer to execute the above-mentioned multi-device collaborative control method for drilling mud processing.
[0017] In the technical solution provided by the present invention, a mathematical model for collaborative control of multiple devices is established, and the Laplace transform and deconvolution methods are used to jointly solve the steady-state control equations and dynamic adjustment equations, thereby solving the coupling control problem between multiple devices; through the design of a branch control architecture, parallel processing of parameters of the mixing tank, sedimentation tank and oil-water separation equipment is realized, reducing the computational complexity of the system; a deep reinforcement learning method is used to construct a reinforcement learning model containing state space, action space and reward function, and pre-verification is performed in combination with digital twin technology to ensure the real-time and accuracy of parameter adjustment; a control mechanism based on closed-loop feedback is designed, and dynamic adjustment and adaptive control of the system are realized through real-time monitoring of export parameters and threshold triggering processing; a deep deterministic policy gradient algorithm is used to iteratively optimize the neural network parameters, thereby improving the generalization ability and control accuracy of the model.
[0018] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic diagram of an embodiment of a multi-device collaborative control method for drilling mud processing according to an embodiment of the present invention;
[0021] Figure 2 A schematic diagram of an embodiment of a multi-device coordinated control device for drilling mud processing according to an embodiment of the present invention;
[0022] Figure 3 Schematic diagram of an embodiment of a multi-device collaborative control device for drilling mud processing in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0024] The terms "including," "having," and any variations thereof, as used in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or device.
[0025] To facilitate understanding of this embodiment, a multi-device collaborative control method for drilling mud processing disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, this method includes the following steps:
[0026] 101. Obtain the stirring parameters of the mixing tank equipment, the settling time parameters of the settling tank equipment, and the separation efficiency parameters of the oil-water separation equipment, and obtain the initial coordinated control parameter set by simultaneously solving the steady-state control equation and the dynamic adjustment equation;
[0027] It is understandable that the execution subject of the present invention may be a multi-device collaborative control device for drilling mud processing, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking the server as the execution subject as an example.
[0028] Specifically, by sampling the stirring speed and stirring time data of the mixing tank equipment, a mixing tank parameter sequence describing the mixing process is obtained. This sequence contains key characteristic data of the mixing tank's operating status, ensuring the uniformity of the mixing process and the consistency of mud properties. Similarly, by sampling the settling time and discharge time data of the sedimentation tank equipment, a sedimentation tank parameter sequence is generated. This sequence reflects the dynamic process of particle settling and discharge in the sedimentation tank, providing fundamental support for solid-liquid separation of the mud. By collecting separation speed and separation temperature data of the oil-water separation equipment, an oil-water separation parameter sequence is obtained. Laplace transforms are performed on the mixing tank parameter sequence, the sedimentation tank parameter sequence, and the oil-water separation parameter sequence. The dynamic data in the time domain is mapped to the complex domain to extract the steady-state characteristics of the equipment operation, and then the steady-state control equation is constructed. The steady-state control equation targets the equilibrium state of the system and ensures that the synergistic relationship between the various equipment is stable during long-term operation. Furthermore, a deconvolution operation is performed on these three parameter sequences to extract the dynamic response characteristics of the equipment under different input disturbances, thereby establishing a dynamic adjustment equation. The dynamic adjustment equation reflects the transient behavior of the equipment and forms the theoretical basis for its rapid response and adjustment to external disturbances during actual operation. The steady-state control equations and the dynamic adjustment equations are input into a simultaneous solver for calculation, generating a matrix of equipment operating parameters. The matrix is then subjected to parameter range constraints to form a parameter constraint space. By defining the physical and process boundaries of the equipment parameters, the parameter constraint space ensures that the resulting parameter set meets both theoretical requirements and actual operating conditions. The equipment operating parameter matrix and the parameter constraint space are then matched and optimized. This process, centered on a multi-objective optimization algorithm, ensures coordinated equipment operation while maximizing efficiency and minimizing energy consumption. After matching and optimization, an initial coordinated control parameter set is generated, comprising the initial agitation parameters for the mixing tank, the initial settling parameters for the sedimentation tank, and the initial separation parameters for the oil-water separator.
[0029] A Laplace forward transform is performed on the stirring speed and stirring time data in the mixing tank parameter sequence, mapping the stirring parameters in the time domain to the complex domain. This allows the steady-state characteristics of the mixing tank to be extracted, resulting in a steady-state transfer function for the mixing tank. This transfer function clarifies the input-output relationship of the mixing tank and its long-term response to specific disturbances. Similarly, a Laplace forward transform is performed on the settling time and blowdown time data in the settling tank parameter sequence to generate a steady-state transfer function for the settling tank. This function describes the dynamic characteristics of particle settling during the settling process and the impact of the blowdown cycle on the system's steady state. A Laplace forward transform is performed on the separation speed and separation temperature data in the oil-water separation parameter sequence to generate a steady-state transfer function for the oil-water separation system. This function reflects the long-term impact of oil-water interface changes and temperature effects on separation efficiency within the separation equipment. The steady-state transfer function for the mixing tank, the settling tank, and the oil-water separation system are coupled to obtain the steady-state control equation for the entire drilling mud treatment system. This equation integrates the parameter relationships between different equipment under steady-state operation and clarifies the coupling effects and synergistic characteristics between the various components in the system. The parameter sequences for the mixing tank, settling tank, and oil-water separation systems are discretized in the time domain, converting the continuous time domain data into discrete data points to generate a discrete time data set. This discrete data set represents the real-time state of the equipment as a set of processable discrete data points based on the time step. The discrete time data set is then fed into a deconvolution algorithm for dynamic characteristic extraction. The deconvolution algorithm inverts the input-output relationship in the time domain to isolate the dynamic response characteristics of the equipment and generate a dynamic response function that describes the equipment's transient behavior. The dynamic response function uses time as a variable to accurately represent the dynamic changes of the equipment under various disturbance conditions. The equipment dynamic response function is then reconstructed in state space. This state space reconstruction transforms the dynamic response function into a state variable matrix. The state variable matrix, centered on the system state, mathematically organizes the dynamic characteristics of the equipment into a set of closely related state variables, facilitating subsequent system modeling and analysis. Based on the state variable matrix, the system's dynamic equations are constructed, integrating the core characteristics of the equipment's dynamic behavior and describing the system's behavior under unsteady conditions. The dynamic equation is linearized and simplified into a linear form, which is suitable for most numerical solutions and optimization algorithms to obtain the dynamic adjustment equation.
[0030] 102. Based on the initial collaborative control parameter set, the mixing tank equipment, the sedimentation tank equipment, and the oil-water separation equipment are divided into branch control architectures and processed in parallel to obtain a branch parallel control parameter set;
[0031] Specifically, based on the initial collaborative control parameter set for the mixing tank, initial settling tank parameters, and initial oil-water separation parameters, a branched control architecture is constructed for the system, forming the mixing tank control branch, the settling tank control branch, and the oil-water separation control branch, respectively. Parameter collection units are constructed for each control branch to capture key data from equipment operation. In the mixing tank control branch, the stirring speed collection module and the stirring time collection module are combined. By jointly collecting these two core parameters, a mixing tank parameter collection matrix is constructed to characterize the operating status of the mixing tank equipment, including the impact of stirring intensity and time on mud properties. Similarly, in the settling tank control branch, the settling time collection module and the blowdown time collection module are combined. The settling time data reflects particle settling efficiency, while the blowdown time data reflects the particle discharge period, forming the settling tank parameter collection matrix. In the oil-water separation control branch, the separation speed collection module and the separation temperature collection module are combined to generate an oil-water separation parameter collection matrix, which describes the effects of speed and temperature on oil-water separation efficiency during the equipment separation process. By comprehensively analyzing the parameter acquisition matrices for the mixing tank, settling tank, and oil-water separator, the operating status of each device is assessed in real time, generating a device status database. Based on the device status database, an adjustment execution unit is constructed to independently and parallelize the mixing tank's stirring parameters, the settling tank's settling parameters, and the oil-water separator's separation parameters, generating branch control sequences. These control sequences represent the optimized results of each device's independent operation, ensuring that each branch achieves the desired performance when executing its control tasks. The branch control sequences are input into a parallel processor for efficient computation through multi-task parallel computing. The parallel processor simultaneously processes control tasks for multiple branches, effectively reducing computational latency and improving overall operational efficiency. The parallel computation results are output as a parallel processing result set. Parameter synchronization and data integration are performed on this result set to coordinate and unify the control results of the independent branches, ensuring information consistency and control coordination across branches. After parameter synchronization and data integration are complete, a branch parallel control parameter set is established. This parameter set includes the parallel parameters for the mixing tank, settling tank, and oil-water separator.
[0032] 103. Conduct deep reinforcement learning training based on the branch parallel control parameter set, build a reinforcement learning model that includes state space, action space, and reward function, and generate the optimal parameter adjustment strategy;
[0033] Specifically, the parallel parameters for the mixing tank, settling tank, and oil-water separator in the branch parallel control parameter set are encoded to establish a deep learning basic model. Based on this deep learning basic model, a state space is constructed. The stirring speed and stirring time states of the mixing tank, the settling time and discharge time states of the settling tank, and the separation speed and temperature states of the oil-water separator are vectorized to obtain device state feature vectors. These device state feature vectors not only contain static information about the current device operation but also capture the dynamic coordination relationships between devices. An action space is constructed based on the device state feature vectors to clarify the range and possibilities of parameter adjustment. By setting adjustment ranges for the mixing tank stirring parameters, the settling time adjustment range for the settling tank, and the oil-water separator speed adjustment range, these adjustment ranges are vectorized to form parameter adjustment action vectors. These action vectors represent the specific adjustment measures that the model can take during the learning process and represent the dynamic operability of the system. To drive the optimization of the reinforcement learning model, a reasonable reward function is constructed. The equipment state feature vector and parameter adjustment action vector are input into the reward function construction unit. A reward metric function model is constructed based on key mud treatment indicators such as density, viscosity, sand content, and oil content. This reward function, through a positive incentive and negative penalty mechanism, ensures that the model gradually learns the optimal strategy for optimizing mud treatment performance during training. The design of the reward function directly impacts the convergence and efficiency of the reinforcement learning process, and is therefore continuously optimized through experiments and simulations. The equipment state feature vector, parameter adjustment action vector, and reward metric function model are input into the digital twin module. Deep reinforcement learning training is then carried out through the equipment operation state simulation unit, mud property prediction unit, and parameter optimization unit. During this process, the digital twin module provides a rich set of training samples for the reinforcement learning model through high-precision simulation of the system's operating state and mud treatment results. These samples include not only static parameters but also feedback data from equipment operation during dynamic adjustments, ensuring the model's robustness under complex operating conditions. Using this reinforcement learning training set, the neural network parameters of the deep learning base model are optimized using policy gradient calculation methods. During this process, a deep deterministic policy gradient algorithm is used to iteratively optimize the target policy network parameters through multiple iterations, allowing the model to gradually approach the optimal adjustment strategy. At the same time, based on the target policy network parameters, a temporal difference learning model is constructed to iteratively calculate the state-action value function. The Bellman equation is introduced to approximate and optimize the value function to obtain the target value network parameters. Parameter decision optimization is performed based on the target policy network parameters and the target value network parameters, combined with the matching relationship between the device state feature vector and the parameter adjustment action vector. The optimal parameter adjustment strategy is generated by calculating the action probability distribution. As the core output of the entire system, the optimal parameter adjustment strategy can dynamically guide the adjustment of the operating parameters of the mixing tank, sedimentation tank, and oil-water separation equipment, so that the equipment coordination efficiency and mud treatment effect reach the optimal level.
[0034] Batch sampling is performed on the reinforcement learning training sample set to ensure data diversity and stability during model training. A policy gradient training sample matrix is constructed by combining the current state sample, the executed action sample, and the next state sample in chronological order into a time-series training sequence. This matrix, centered on the state changes and action responses of the device, comprehensively reflects the dynamic characteristics of the system and provides accurate data support for policy optimization. The policy gradient training sample matrix is input into the target network, and the target Q-value is calculated through forward propagation. The target Q-value is a key metric used in reinforcement learning to evaluate the long-term value of a specific state-action pair, reflecting the performance of the current policy under the long-term optimization goal. Simultaneously, the target Q-value is input into the current network, and action value calculation is performed to obtain the current Q-value. The current Q-value, as an immediate assessment of the value of the current state and action combination, together with the target Q-value, constitutes the optimization target of the reinforcement learning model. A temporal difference error is calculated between the target Q-value and the current Q-value. The error between the target Q-value and the current Q-value is calculated to quantify the optimization direction of the model. This error value, as the policy network loss, is input into the backpropagation algorithm to calculate the gradient vector of the neural network parameters. The gradient vector not only indicates the direction in which network parameters need to be adjusted but also provides a basis for the magnitude of the update. Based on the calculated parameter update direction, a soft update is performed on the network weights. Soft update is a commonly used weight update method in reinforcement learning. It introduces an update increment to avoid drastic parameter changes and improve model stability. During the soft update calculation process, the weight update increment is calculated based on the gradient direction. This increment is then superimposed on the current network parameters to obtain the updated network parameters. The updated network parameters reflect the degree of improvement in the current policy's optimization direction. The updated network parameters are then clipped to limit the parameter range and avoid excessively large gradients that can lead to unstable learning. The corrected network parameters are then input into the target network for a synchronous parameter update, ensuring parameter consistency between the target network and the current network. Through these steps, the target policy network parameters are obtained.
[0035] 104. Based on the optimal parameter adjustment strategy, the mud density data and mud viscosity data at the outlet of the mixing tank equipment, the sand content data at the outlet of the sedimentation tank equipment, and the oil content data at the outlet of the oil-water separation equipment are monitored and threshold-triggered, and dynamic adjustment control instructions are output.
[0036] Specifically, the density and viscosity of the mud at the outlet of the mixing tank are measured. Density and viscosity sensors are used to precisely sample the mud at the outlet, and the sampling results are recorded as a monitoring sequence for the mixing tank outlet. This monitoring sequence dynamically captures the changing characteristics of the mud's physical properties during the mixing tank's operation, using a time axis. Similarly, the mud at the outlet of the sedimentation tank is measured, and a sand content sensor is used to sample the particle concentration of the mud at the outlet. The test results are recorded as a monitoring sequence for the sedimentation tank outlet, which reflects the actual treatment performance of the sedimentation tank during the solid-liquid separation process. The oil content of the liquid at the outlet of the oil-water separator is measured, and a highly sensitive oil content sensor is used to collect data on the oil-water ratio of the separated liquid. This generates an oil-water separator outlet parameter monitoring sequence, which provides real-time data for evaluating oil-water separation performance. The monitoring sequence for the mixing tank outlet, the sedimentation tank outlet, and the oil-water separator outlet are input into a closed-loop feedback control system. The feedback control system then performs numerical calculations on the mud density, viscosity, sand content, and oil content to generate a real-time monitoring data matrix. The real-time monitoring data matrix presents the dynamic changes of all key parameters in a unified format, enabling the system to capture the output status of each device in real time during operation. Based on the threshold ranges preset in the optimal parameter adjustment strategy, the real-time monitoring data matrix is compared with the mud density, viscosity, sand content, and oil content thresholds. The calculation results are presented as parameter deviation vectors. This parameter deviation vector quantifies the deviation between the device output parameter and the target value, providing a direct optimization direction for dynamic adjustment. Each element in the parameter deviation vector reflects the performance deviation of a specific device. To ensure efficient adjustment, the parameter deviation vectors are ranked and prioritized according to importance. By ranking the importance of each deviation element, the key parameters with the greatest impact on overall system performance are identified, and an adjustment priority sequence is generated based on their priority. This adjustment priority sequence is then input into the compensation controller, which then generates dynamic adjustment control instructions based on the optimal parameter adjustment strategy. By comprehensively analyzing the deviation vectors, adjustment priorities, and the optimal parameter adjustment strategy, the compensation controller dynamically optimizes key control parameters such as the mixing tank stirring speed, the settling tank settling time, and the oil-water separator speed, and generates specific control instructions. The generated dynamic adjustment control commands are sent via the communication interface to the corresponding actuators in the mixing tank, settling tank, and oil-water separator. Upon receiving the stirring speed command, the mixing tank actuator adjusts the stirring speed to optimize mud mixing. The settling tank actuator adjusts the particle settling time based on the settling time command to improve solid-liquid separation efficiency. The oil-water separator actuator adjusts the operating speed of the separation equipment based on the separation speed command, thereby enhancing oil-water separation. This entire process continuously optimizes equipment operating status through a closed-loop feedback mechanism, ensuring the efficiency and stability of the mud treatment process.
[0037] In an embodiment of the present invention, a mathematical model for collaborative control of multiple devices is established, and the Laplace transform and deconvolution methods are used to jointly solve the steady-state control equations and dynamic adjustment equations, thereby solving the coupling control problem between multiple devices; through the design of a branch control architecture, parallel processing of parameters of the mixing tank, sedimentation tank and oil-water separation equipment is achieved, thereby reducing the computational complexity of the system; a deep reinforcement learning method is used to construct a reinforcement learning model containing state space, action space and reward function, and pre-verification is performed in combination with digital twin technology to ensure the real-time and accuracy of parameter adjustment; a control mechanism based on closed-loop feedback is designed, and dynamic adjustment and adaptive control of the system are achieved through real-time monitoring of outlet parameters and threshold triggering processing; a deep deterministic policy gradient algorithm is used to iteratively optimize the neural network parameters, thereby improving the generalization ability and control accuracy of the model.
[0038] In a specific embodiment, the process of executing step 101 may specifically include the following steps:
[0039] The stirring speed data and stirring time data of the mixing tank equipment are sampled to obtain a mixing tank parameter sequence, and the settling time data and sewage discharge time data of the settling tank equipment are sampled to obtain a settling tank parameter sequence, and the separation speed data and separation temperature data of the oil-water separation equipment are sampled to obtain an oil-water separation parameter sequence;
[0040] Performing Laplace transform on the mixing tank parameter sequence, the sedimentation tank parameter sequence, and the oil-water separation parameter sequence to obtain the steady-state control equation, and performing deconvolution operation on the mixing tank parameter sequence, the sedimentation tank parameter sequence, and the oil-water separation parameter sequence to obtain the dynamic adjustment equation;
[0041] Input the steady-state control equation and the dynamic adjustment equation into the simultaneous solver for simultaneous calculation to obtain the equipment operation parameter matrix, and then impose parameter range constraints on the equipment operation parameter matrix to obtain the parameter constraint space;
[0042] The equipment operation parameter matrix and the parameter constraint space are matched and optimized to obtain the initial collaborative control parameter set, which includes the initial stirring parameters of the mixing tank, the initial sedimentation parameters of the sedimentation tank, and the initial separation parameters of the oil-water separation.
[0043] Specifically, the stirring speed of the mixing tank equipment and stirring time Sampling is performed and the stirring data in the continuous time domain is recorded as a mixing tank parameter sequence These data reflect key dynamic characteristics of the mixing process, such as higher Speeds up mixing but increases energy consumption, while shorter This results in uneven mixing of the slurry. and discharge time Take samples and generate a settling tank parameter sequence The sedimentation time determines the adequacy of particle settling, while the discharge time reflects the discharge efficiency and the continuity of mud flow. and separation temperature Sampling is performed to obtain the oil-water separation parameter sequence The speed and temperature directly affect the oil-water separation effect and equipment energy consumption. After collecting these sequences, they are Laplace transformed to obtain the steady-state control equation. Taking the mixing tank parameter sequence as an example, the Laplace transform formula is applied:
[0044] ;
[0045] in, is the steady-state transfer function of the mixing tank, is a complex frequency domain variable, and are time functions of stirring speed and time respectively. Similarly, the steady-state transfer function of the sedimentation tank parameters is expressed as:
[0046] ;
[0047] The steady-state transfer function of the oil-water separation parameters is:
[0048] ;
[0049] The above steady-state equations are combined with the operating characteristics of the equipment to establish the steady-state control equations for the entire system. At the same time, based on the dynamic input and output relationships of the mixing tank, sedimentation tank, and oil-water separation equipment, deconvolution operations are performed to extract dynamic characteristics. For example, the dynamic adjustment equation for the mixing tank is expressed as:
[0050] ;
[0051] in, represents the dynamic response of the mixing tank, is the convolution operator, is the impulse response function of the mixing tank. Similarly, the dynamic adjustment equations of the sedimentation tank and oil-water separation equipment are:
[0052] ;
[0053] ;
[0054] in, and These are the impulse response functions of the sedimentation tank and the oil-water separation equipment, respectively. Input the steady-state control equation and the dynamic adjustment equation into the simultaneous solver to solve the equipment operation parameter matrix:
[0055] ;
[0056] in, It is the equipment operating parameter matrix, which includes multiple operating parameters of the mixing tank, sedimentation tank and oil-water separation equipment, specifically expressed as and are the speed and time parameters of the mixing tank, and is the sedimentation and discharge time parameter of the sedimentation tank, and The speed and temperature parameters of the oil-water separation equipment are defined. The operating parameter matrix is constrained by parameter ranges, and these constraints are constructed into a parameter constraint space. The operating parameter matrix and the parameter constraint space are then matched and optimized using a multi-objective optimization algorithm, such as a genetic algorithm or a particle swarm optimization algorithm, to ultimately determine the initial collaborative control parameter set.
[0057] In a specific embodiment, the execution step of performing Laplace transform on the mixing tank parameter sequence, the settling tank parameter sequence, and the oil-water separation parameter sequence to obtain a steady-state control equation, and performing a deconvolution operation based on the mixing tank parameter sequence, the settling tank parameter sequence, and the oil-water separation parameter sequence to obtain a dynamic adjustment equation may specifically include the following steps:
[0058] Performing Laplace forward transformation on the stirring speed data and stirring time data in the mixing tank parameter sequence to obtain the mixing tank steady-state transfer function, performing Laplace forward transformation on the settling time data and sewage discharge time data in the settling tank parameter sequence to obtain the settling tank steady-state transfer function, and performing Laplace forward transformation on the separation speed data and separation temperature data in the oil-water separation parameter sequence to obtain the oil-water separation steady-state transfer function;
[0059] The steady-state transfer function of the mixing tank, the steady-state transfer function of the sedimentation tank and the steady-state transfer function of the oil-water separation are systematically coupled to obtain the steady-state control equation.
[0060] The mixing tank parameter sequence, sedimentation tank parameter sequence, and oil-water separation parameter sequence are discretized in the time domain to obtain a time domain discrete data set. The time domain discrete data set is then input into the deconvolution algorithm to extract dynamic characteristics and obtain the dynamic response function of the equipment.
[0061] The state space of the equipment dynamic response function is reconstructed to obtain the state variable matrix, and the system dynamic equation is constructed and linearized based on the state variable matrix to obtain the dynamic adjustment equation.
[0062] Specifically, for the stirring speed in the mixing tank parameter sequence and stirring time Perform Laplace forward transform to construct the steady-state transfer function of the mixing tank. The formula for Laplace transform is:
[0063] ;
[0064] in, is the Laplace domain representation of the stirring velocity, is a complex frequency variable. Similarly, the stirring time The Laplace transform of is:
[0065] ;
[0066] The steady-state transfer function of the mixing tank is described by combining these two parameters:
[0067] ;
[0068] in, Indicates that the mixing tank is under input disturbance Lower pair output Similarly, the settling time in the settling tank parameter sequence is and discharge time Perform Laplace transform to obtain the steady-state transfer function of the sedimentation tank:
[0069] ;
[0070] in, It represents the steady-state behavior of the sedimentation tank under the parameters of sedimentation and discharge time. and separation temperature Perform Laplace transform to obtain the steady-state transfer function of oil-water separation:
[0071] ;
[0072] in, The influence of separation speed and temperature on separation effect in oil-water separation equipment is described. 、 、 Perform system coupling calculations and construct steady-state control equations. Assuming that there is a certain linear coupling relationship between the mixing tank, sedimentation tank, and oil-water separation equipment, the overall steady-state control equation of the system is expressed as:
[0073] ;
[0074] in, It is the steady-state transfer function of the entire system, describing the steady-state coupling characteristics between multiple devices. After completing the steady-state modeling, the parameter sequences of the mixing tank, sedimentation tank and oil-water separation equipment are discretized in the time domain. For example, the stirring speed data of the mixing tank Discretized into , sedimentation time data Discretized into , separate speed data Discretized into ,in Represents a discrete time step. Through discretization, a time-domain discrete data set is formed:
[0075] ;
[0076] These data sets capture the dynamic characteristics of the equipment's operation and are used in subsequent deconvolution operations. The time-domain discrete data sets are fed into the deconvolution algorithm to extract the equipment's dynamic characteristics. For example, the dynamic response function of a mixing tank can be expressed as follows using the deconvolution formula:
[0077] ;
[0078] Among them, Deconv represents the deconvolution operation. Similarly, the dynamic response functions of the sedimentation tank and oil-water separation equipment are:
[0079] ;
[0080] ;
[0081] These dynamic response functions 、 、 Describe the transient behavior characteristics of the equipment at different time steps. Reconstruct the state space of the equipment dynamic response function. For example, the state variables of the mixing tank are expressed as , the state variables of the sedimentation tank and the oil-water separation equipment are and Combined with the dynamic response function, the state space model is expressed as:
[0082] ;
[0083] ;
[0084] in, is the state variable matrix, is the input matrix (such as 、 、 ), is the output matrix (such as 、 ), 、 、 is the system matrix that describes the relationship between state, input, and output. Through linearization, the system dynamic equation is converted into a computable linear form, resulting in the dynamic adjustment equation. This dynamic adjustment equation, combined with the steady-state control equation, constitutes a comprehensive control model for the system, providing a theoretical basis for equipment collaborative optimization. For example, the final dynamic adjustment equation is expressed as:
[0085] ;
[0086] in, and is the linearized system matrix, which can be used more efficiently for numerical calculations and optimization.
[0087] In a specific embodiment, the process of executing step 102 may specifically include the following steps:
[0088] A branch control architecture is established based on the initial stirring parameters of the mixing tank, the initial sedimentation parameters of the sedimentation tank, and the initial separation parameters of the oil-water separation in the initial collaborative control parameter set, and a mixing tank control branch, a sedimentation tank control branch, and an oil-water separation control branch are obtained;
[0089] A parameter acquisition unit is constructed for the mixing tank control branch, and the stirring speed acquisition module and the stirring time acquisition module are combined to obtain the mixing tank parameter acquisition matrix. A parameter acquisition unit is constructed for the sedimentation tank control branch, and the sedimentation time acquisition module and the sewage discharge time acquisition module are combined to obtain the sedimentation tank parameter acquisition matrix. A parameter acquisition unit is constructed for the oil-water separation control branch, and the separation speed acquisition module and the separation temperature acquisition module are combined to obtain the oil-water separation parameter acquisition matrix.
[0090] Based on the mixing tank parameter acquisition matrix, the sedimentation tank parameter acquisition matrix, and the oil-water separation parameter acquisition matrix, the equipment operation status is evaluated to obtain the equipment status database;
[0091] An adjustment execution unit is built based on the equipment status database to independently and parallelly calculate the mixing tank stirring parameters, sedimentation tank sedimentation parameters, and oil-water separation parameters to obtain a branch control sequence.
[0092] The branch control sequence is input into the parallel processor for multi-task parallel calculation to obtain a parallel processing result set, and the parallel processing result set is subjected to parameter synchronization and data integration to establish a branch parallel control parameter set. The branch parallel control parameter set includes the mixing tank parallel parameters, the sedimentation tank parallel parameters and the oil-water separation parallel parameters.
[0093] Specifically, based on the initial stirring parameters of the mixing tank in the initial collaborative control parameter set , initial sedimentation parameters of the sedimentation tank , and the initial separation parameters of oil-water separation , respectively construct the mixing tank control branch, sedimentation tank control branch and oil-water separation control branch. The mixing tank control branch describes the stirring speed of the equipment and stirring time Impact on slurry uniformity, sedimentation tank control branch reflects the sedimentation time and discharge time The effect of solid-liquid separation on mud, while the oil-water separation control branch reflects the separation speed and separation temperature After the branch is built, a parameter acquisition unit is established for each branch to collect operating data. In the mixing tank control branch, the mixing tank parameter acquisition matrix is formed by combining the stirring speed acquisition module and the stirring time acquisition module. , which is defined as:
[0094] ;
[0095] in, and Respectively represent Similarly, in the sedimentation tank control branch, the sedimentation time acquisition module and the sewage discharge time acquisition module are combined to form the sedimentation tank parameter acquisition matrix :
[0096] ;
[0097] in, and Respectively In the oil-water separation control branch, the oil-water separation parameter acquisition matrix is generated by combining the separation speed acquisition module and the separation temperature acquisition module. :
[0098] ;
[0099] in, and Respectively The separation speed and separation temperature collected are input into the equipment operation status evaluation module, and the equipment status database is generated through data processing and analysis. , whose structure is:
[0100] ;
[0101] The state database records the parameter changes of each device under different operating conditions. Based on the device state database, an adjustment execution unit is constructed to independently and concurrently calculate the control parameters of the mixing tank, sedimentation tank, and oil-water separation device. The adjustment process is defined as an optimization problem, where the goal is to minimize the operating deviation. For the mixing tank control branch, the optimization objective is:
[0102] ;
[0103] in, is the adjustment vector of the control parameters of the mixing tank. Similarly, for the sedimentation tank and oil-water separation control branches, they are:
[0104] ;
[0105] ;
[0106] By solving these optimization problems, we can obtain the branch control sequences for the mixing tank, sedimentation tank, and oil-water separation equipment. The branch control sequences are input into the parallel processor, and the parallel processing result set is obtained through multi-task parallel computing. The result set includes the operating optimization parameters of each equipment, such as the optimized value of the stirring speed. , the optimized value of sedimentation time , Optimized value of separation speed Etc. The parallel processing result set is represented as:
[0107] ;
[0108] Perform parameter synchronization and data integration on the parallel processing result set to generate the branch parallel control parameter set ;
[0109] in, .
[0110] In a specific embodiment, the process of executing step 103 may specifically include the following steps:
[0111] Parameter encoding is performed on the mixing tank parallel parameters, sedimentation tank parallel parameters, and oil-water separation parallel parameters in the branch parallel control parameter set to establish a deep learning basic model;
[0112] Based on the deep learning basic model, the state space is constructed, and the stirring speed state and stirring time state of the mixing tank equipment, the settling time state and sewage discharge time state of the sedimentation tank equipment, and the separation speed state and separation temperature state of the oil-water separation equipment are vectorized to obtain the equipment state feature vector;
[0113] The action space of the equipment state feature vector is constructed to set the adjustment range of the mixing tank stirring parameters, the adjustment range of the sedimentation tank settling time, and the adjustment range of the oil-water separation speed to obtain the parameter adjustment action vector.
[0114] The equipment state feature vector and parameter adjustment action vector are input into the reward function construction unit, and a reward measurement function model is constructed based on the processing indicators of mud density, viscosity, sand content and oil content;
[0115] The equipment state feature vector, parameter adjustment action vector, and reward metric function model are input into the digital twin module. Deep reinforcement learning training is performed through the equipment operation state simulation unit, mud property prediction unit, and parameter optimization unit to obtain a reinforcement learning training sample set.
[0116] Perform policy gradient calculation on the reinforcement learning training sample set, and use the deep deterministic policy gradient algorithm to iteratively optimize the neural network parameters of the deep learning basic model to obtain the target policy network parameters;
[0117] A temporal difference learning model is constructed based on the target policy network parameters. The state-action value function is iteratively calculated and the value function is approximated by the Bellman equation to obtain the target value network parameters.
[0118] Parameter decisions are made based on the target strategy network parameters and the target value network parameters, and the optimal parameter adjustment strategy is generated through action probability distribution calculation.
[0119] Specifically, the parallel parameters of the mixing tank are extracted from the branch parallel control parameter set , sedimentation tank parallel parameters and oil-water separation parallel parameters These parameters are standardized and the parameters of different devices are expressed as a unified vector form through parameter encoding. For example, the mixing tank parameter is encoded as ,in and Represent the normalized values of stirring speed and stirring time respectively. Similarly, other equipment parameters are coded as and These encoding results constitute the input of the deep learning basic model. Based on the deep learning basic model, the state space of the system is constructed, and the stirring speed state of the mixing tank equipment is converted to and stirring time status , sedimentation time status of sedimentation tank equipment and discharge time status , separation speed status of oil-water separation equipment and separation temperature state Perform vectorization processing to generate device state feature vector
[0120] The construction of state feature vector unifies the key operating states of each device into a multi-dimensional vector form, providing high-dimensional input for the reinforcement learning model. After the state space is constructed, the action space needs to be further defined by setting the adjustment range of the mixing tank stirring parameters. , Adjustment range of sedimentation tank sedimentation time , and the adjustment range of oil-water separation speed , generate parameter-adjusted action vector ,in 、 and Represent the adjustment values of stirring speed, settling time and separation speed respectively. The action vector defines the optimization dimension and control space of the reinforcement learning model. and parameter adjustment action vector Input reward function building block, based on mud density , viscosity , sand content and oil content etc., and construct a reward measurement function model. For example, the reward function is expressed as:
[0121] ;
[0122] in, are the target mud density, viscosity, sand content and oil content respectively, is the weight of each indicator. By minimizing the deviation of each indicator, the reward function drives the model towards the target optimization direction. , parameter adjustment action vector and the reward function The digital twin module is input and a reinforcement learning training sample set is generated through the equipment operation status simulation unit, the mud property prediction unit, and the parameter optimization unit. The simulation unit generates equipment status data under various working conditions through simulation. The prediction unit uses a physical model to predict mud treatment performance. The optimization unit selects the optimal action through the reward function, thereby generating rich training data. Using the reinforcement learning training sample set, a deep deterministic policy gradient algorithm is used to optimize the neural network parameters of the deep learning basic model. At each training step, the policy gradient calculation formula is:
[0123] ;
[0124] in, It is a strategic network. For its parameters, is the state-action value function. Through multiple iterative optimizations, the target strategy network parameters are obtained. . Based on the target strategy network parameters Construct a temporal difference learning model to learn the state-action value function Perform iterative calculations. Using the Bellman equation:
[0125] ;
[0126] in, is the discount factor, For the next state, the target value network parameters are obtained through approximation iteration. Combined with the target strategy network parameters and target value network parameters, and generates the optimal parameter adjustment strategy through action probability distribution calculation .
[0127] In a specific embodiment, the execution step performs policy gradient calculation on the reinforcement learning training sample set, and uses a deep deterministic policy gradient algorithm to iteratively optimize the neural network parameters of the deep learning basic model to obtain the target policy network parameters. The process can specifically include the following steps:
[0128] Batch sampling is performed on the reinforcement learning training sample set, and the current state sample, the executed action sample, and the next state sample are combined into a time-series training sequence to obtain the policy gradient training sample matrix;
[0129] Input the policy gradient training sample matrix into the target network, calculate the target Q value through forward propagation, and input the target Q value into the current network for action value calculation to obtain the current Q value;
[0130] Perform temporal difference error calculation on the target Q value and the current Q value to obtain the policy network loss value, and input the policy network loss value into the back propagation algorithm to calculate the gradient vector of the network parameters and obtain the parameter update direction;
[0131] Perform soft update calculation on the network weights based on the parameter update direction to obtain the weight update increment, and superimpose the weight update increment with the current network parameters to obtain the updated network parameters;
[0132] The updated network parameters are trimmed to obtain the corrected network parameters, and the corrected network parameters are input into the target network for parameter synchronization update to obtain the target strategy network parameters.
[0133] Specifically, for the reinforcement learning training sample set Perform batch sampling. The sample set contains three types of data: current state samples , execute action samples and the next state sample Batch sampling randomly selects from the sample set records, combined into a time series training sequence ,in It is The current state vector of the sample, is the corresponding action vector, is the next state obtained after executing the action, is the immediate reward value. For example, in a mud processing system, Indicates the equipment operating status such as stirring speed and stirring time, Indicates the corresponding action adjustment. Input the target policy network for forward propagation. The target policy network is a deep neural network whose structure includes an input layer, several hidden layers and an output layer, and is designed to approximate the target state-action value function. Through forward propagation calculation, the target policy network is calculated based on the input Output Target value:
[0134] ;
[0135] in, It’s an instant reward. Is a discount factor that balances the weights of current rewards and future gains. The value describes the long-term benefits of the system after executing actions under the current strategy. The value is input into the current policy network and the current value The structure of the current policy network is the same as the target network, but the parameters are independent and used for real-time optimization of the policy. Value and current Value, calculate the timing difference error:
[0136] ;
[0137] in, is the temporal difference error, which indicates the degree of deviation of the current network from the target strategy. As a basis, define the loss function of the policy network:
[0138] ;
[0139] in, are the parameters of the current network, is the batch size. The loss function measures the deviation between the network output and the target value. Enter the back propagation algorithm and calculate the gradient vector of the network parameters through the chain rule:
[0140] ;
[0141] The gradient vector indicates the direction and magnitude of the parameter update. Based on the calculated gradient vector, the network weights are soft-updated. The soft-update strategy uses exponential smoothing to gradually update the network parameters. The formula is:
[0142] ;
[0143] in, are the target network parameters, is the update coefficient, which controls the step size of parameter update. In this way, it can avoid network instability caused by large parameter changes. Superimpose with the current network parameters to obtain the updated network parameters:
[0144] ;
[0145] To prevent parameter values from exceeding a reasonable range, the updated network parameters are clipped. For example, the parameter values are limited to between [-1, 1]. The specific formula is:
[0146] ;
[0147] Input the trimmed corrected network parameters into the target policy network and perform parameter synchronization update:
[0148] ;
[0149] The synchronously updated target policy network parameters better approximate the optimal state-action value function, thereby improving the convergence of reinforcement learning.
[0150] In a specific embodiment, the process of executing step 104 may specifically include the following steps:
[0151] The density sensor and viscosity sensor are used to sample and detect the outlet mud of the mixing tank equipment to obtain the monitoring sequence of the mixing tank outlet parameters; the sand content sensor is used to sample and detect the outlet mud of the sedimentation tank equipment to obtain the monitoring sequence of the sedimentation tank outlet parameters; the oil content sensor is used to sample and detect the outlet liquid of the oil-water separation equipment to obtain the monitoring sequence of the oil-water separation outlet parameters;
[0152] Input the monitoring sequence of the mixing tank outlet parameter, the sedimentation tank outlet parameter monitoring sequence and the oil-water separation outlet parameter monitoring sequence into the closed-loop feedback control system, perform numerical calculations on the mud density, viscosity, sand content and oil content, and obtain a real-time monitoring data matrix;
[0153] Based on the preset threshold range in the optimal parameter adjustment strategy, the real-time monitoring data matrix is compared and calculated with the mud density threshold, viscosity threshold, sand content threshold, and oil content threshold to obtain the parameter deviation vector;
[0154] The parameter deviation vector is sorted and prioritized to obtain an adjustment priority sequence, which is then input into the compensation controller to generate dynamic adjustment control instructions based on the optimal parameter adjustment strategy.
[0155] The mixing tank stirring speed instruction, the sedimentation tank sedimentation time instruction and the oil-water separation speed instruction in the dynamic adjustment control instruction are sent to the corresponding actuators through the communication interface.
[0156] Specifically, the density and viscosity of the outlet mud of the mixing tank equipment are sampled in real time, and the density of the mud is obtained using a density sensor. and viscosity sensor to obtain mud viscosity These data constitute the monitoring sequence of mixing tank outlet parameters:
[0157] ;
[0158] in, It's mud in time The density, The dynamic viscosity of the mud is obtained by sampling the sand content of the mud at the outlet of the sedimentation tank equipment and obtaining the sand content through the sand content sensor. , which is the volume percentage of solid particles in the mud, in units of The monitoring sequence of sedimentation tank outlet parameters is defined as:
[0159] ;
[0160] The oil content of the outlet liquid of the oil-water separation equipment is sampled in real time, and the oil content sensor is used to obtain , that is, the percentage of oil in the liquid. The oil-water separation outlet parameter monitoring sequence is defined as:
[0161] ;
[0162] Integrate the monitoring sequence of outlet parameters of mixing tanks, sedimentation tanks and oil-water separation equipment into a real-time monitoring data matrix , whose structure is:
[0163] ;
[0164] Through a closed-loop feedback control system, the real-time monitoring data matrix is used to calculate and analyze the mud density, viscosity, sand content, and oil content. Assume that the threshold ranges of these parameters are preset in the system's optimal parameter adjustment strategy, for example:
[0165] ;
[0166] ;
[0167] Compare the real-time monitoring data matrix with these thresholds and calculate the deviation vector of each parameter:
[0168] ;
[0169] The deviation vectors are ranked and prioritized, and the adjustment priority sequence is determined based on the degree of influence of the parameters and the deviation range, for example: This sequence serves as the input to the compensation controller to guide dynamic adjustments. The adjustment priority sequence is input to the compensation controller, which generates dynamic adjustment control instructions based on the optimal parameter adjustment strategy. For example, if the adjustment strategy is to reduce the mud density deviation by increasing the mixing tank stirring speed, reduce the sand content deviation by extending the settling time, and reduce the oil content deviation by reducing the separation speed, the adjustment instruction is expressed as:
[0170] ;
[0171] in, 、 、 These represent the adjustment amounts for agitation speed, settling time, and separation speed, respectively. Dynamic adjustment control commands are sent to the corresponding actuators via the communication interface. This closed-loop feedback and dynamic adjustment process optimizes equipment operating status in real time, ensuring that mud treatment performance meets target requirements.
[0172] The above describes the multi-device collaborative control method for drilling mud processing in an embodiment of the present invention. The following describes the multi-device collaborative control device for drilling mud processing in an embodiment of the present invention. Figure 2 In one embodiment of the present invention, a multi-device coordinated control device for drilling mud processing includes:
[0173] An acquisition module 201 is used to obtain the stirring parameters of the mixing tank equipment, the settling time parameters of the settling tank equipment, and the separation efficiency parameters of the oil-water separation equipment, and obtain an initial collaborative control parameter set by simultaneously solving the steady-state control equation and the dynamic adjustment equation;
[0174] The parallel processing module 202 is used to perform branch control architecture division and parallel processing on the mixing tank equipment, the sedimentation tank equipment, and the oil-water separation equipment based on the initial collaborative control parameter set to obtain a branch parallel control parameter set;
[0175] Reinforcement learning module 203, configured to perform deep reinforcement learning training based on the branch parallel control parameter set, construct a reinforcement learning model including a state space, an action space, and a reward function, and generate an optimal parameter adjustment strategy;
[0176] The output module 204 is used to monitor and perform threshold triggering processing on the mud density data and mud viscosity data at the outlet of the mixing tank equipment, the sand content data at the outlet of the sedimentation tank equipment, and the oil content data at the outlet of the oil-water separation equipment based on the optimal parameter adjustment strategy, and output dynamic adjustment control instructions.
[0177] Through the collaborative cooperation of the above-mentioned components, by establishing a mathematical model for collaborative control of multiple devices, the Laplace transform and deconvolution methods are used to jointly solve the steady-state control equations and dynamic adjustment equations, thereby solving the coupling control problem between multiple devices; through the design of a branch control architecture, parallel processing of the parameters of the mixing tank, sedimentation tank and oil-water separation equipment is achieved, reducing the computational complexity of the system; a deep reinforcement learning method is used to construct a reinforcement learning model containing state space, action space and reward function, and pre-verification is carried out in combination with digital twin technology to ensure the real-time and accuracy of parameter adjustment; a control mechanism based on closed-loop feedback is designed, and dynamic adjustment and adaptive control of the system are achieved through real-time monitoring of outlet parameters and threshold triggering processing; a deep deterministic policy gradient algorithm is used to iteratively optimize the neural network parameters, thereby improving the generalization ability and control accuracy of the model.
[0178] above Figure 2 The multi-device collaborative control device for drilling mud processing in an embodiment of the present invention is described in detail from the perspective of modular functional entities. The multi-device collaborative control device for drilling mud processing in an embodiment of the present invention is described in detail from the perspective of hardware processing.
[0179] Figure 3The figure is a schematic diagram of the structure of a multi-device collaborative control device for drilling mud processing provided by an embodiment of the present invention. The multi-device collaborative control device 300 for drilling mud processing may vary significantly depending on configuration or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors), a memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) storing application programs 333 or data 332. The memory 320 and storage medium 330 may be either transient or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown), each of which may include a series of instruction operations within the multi-device collaborative control device 300 for drilling mud processing. Furthermore, the processor 310 may be configured to communicate with the storage medium 330, executing the series of instruction operations stored in the storage medium 330 on the multi-device collaborative control device 300 for drilling mud processing, thereby implementing the steps of the multi-device collaborative control method for drilling mud processing described above.
[0180] The multi-device collaborative control device 300 for drilling mud processing may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The structure of the multi-device collaborative control device for drilling mud processing shown does not constitute a limitation on the multi-device collaborative control device for drilling mud processing provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0181] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the multi-device collaborative control method for drilling mud processing.
[0182] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0184] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-device collaborative control method for drilling mud processing, characterized in that: The method comprises: Obtain the stirring parameters of the mixing tank equipment, the settling time parameters of the sedimentation tank equipment and the separation efficiency parameters of the oil-water separation equipment, and obtain the initial collaborative control parameter set by jointly solving the steady-state control equation and the dynamic adjustment equation; specifically, sampling the stirring speed data and stirring time data of the mixing tank equipment to obtain the mixing tank parameter sequence, sampling the sedimentation time data and sewage discharge time data of the sedimentation tank equipment to obtain the sedimentation tank parameter sequence, and sampling the separation speed data and separation temperature data of the oil-water separation equipment to obtain the oil-water separation parameter sequence; pulling the mixing tank parameter sequence, the sedimentation tank parameter sequence and the oil-water separation parameter sequence Plas transform is performed to obtain a steady-state control equation, and a deconvolution operation is performed based on the mixing tank parameter sequence, the sedimentation tank parameter sequence, and the oil-water separation parameter sequence to obtain a dynamic adjustment equation; the steady-state control equation and the dynamic adjustment equation are input into a simultaneous solver for simultaneous calculation to obtain an equipment operation parameter matrix, and parameter range constraints are performed on the equipment operation parameter matrix to obtain a parameter constraint space; the equipment operation parameter matrix is matched with the parameter constraint space for optimization calculation to obtain an initial coordinated control parameter set, wherein the initial coordinated control parameter set includes initial stirring parameters of the mixing tank, initial sedimentation parameters of the sedimentation tank, and initial separation parameters of the oil-water separation; Based on the initial collaborative control parameter set, the mixing tank device, the sedimentation tank device, and the oil-water separation device are divided into branch control architectures and processed in parallel to obtain a branch parallel control parameter set; Performing deep reinforcement learning training based on the branch parallel control parameter set, constructing a reinforcement learning model including a state space, an action space, and a reward function, and generating an optimal parameter adjustment strategy; Based on the optimal parameter adjustment strategy, the mud density data and mud viscosity data at the outlet of the mixing tank equipment, the sand content data at the outlet of the sedimentation tank equipment, and the oil content data at the outlet of the oil-water separation equipment are monitored and threshold-triggered, and dynamic adjustment control instructions are output.
2. The multi-device coordinated control method for drilling mud processing according to claim 1, characterized in that: The Laplace transform is performed on the mixing tank parameter sequence, the sedimentation tank parameter sequence, and the oil-water separation parameter sequence to obtain a steady-state control equation, and a deconvolution operation is performed based on the mixing tank parameter sequence, the sedimentation tank parameter sequence, and the oil-water separation parameter sequence to obtain a dynamic adjustment equation, including: Performing a Laplace forward transform on the stirring speed data and stirring time data in the mixing tank parameter sequence to obtain a mixing tank steady-state transfer function, performing a Laplace forward transform on the settling time data and sewage discharge time data in the settling tank parameter sequence to obtain a settling tank steady-state transfer function, and performing a Laplace forward transform on the separation speed data and separation temperature data in the oil-water separation parameter sequence to obtain an oil-water separation steady-state transfer function; Performing system coupling calculation on the mixing tank steady-state transfer function, the settling tank steady-state transfer function and the oil-water separation steady-state transfer function to obtain a steady-state control equation; Performing time-domain discretization processing on the mixing tank parameter sequence, the sedimentation tank parameter sequence, and the oil-water separation parameter sequence to obtain a time-domain discrete data set, and inputting the time-domain discrete data set into a deconvolution algorithm to extract dynamic characteristics to obtain a dynamic response function of the equipment; The state space of the dynamic response function of the device is reconstructed to obtain a state variable matrix, and the system dynamic equation is constructed and linearized based on the state variable matrix to obtain a dynamic adjustment equation.
3. The multi-device coordinated control method for drilling mud processing according to claim 2, characterized in that: Based on the initial collaborative control parameter set, the mixing tank device, the sedimentation tank device, and the oil-water separation device are subjected to branch control architecture division and parallel processing to obtain a branch parallel control parameter set, including: A branch control architecture is established based on the initial stirring parameters of the mixing tank, the initial settling parameters of the settling tank, and the initial separation parameters of the oil-water separation in the initial collaborative control parameter set, to obtain a mixing tank control branch, a settling tank control branch, and an oil-water separation control branch; A parameter acquisition unit is constructed for the mixing tank control branch, and the stirring speed acquisition module and the stirring time acquisition module are combined to obtain a mixing tank parameter acquisition matrix; a parameter acquisition unit is constructed for the sedimentation tank control branch, and the sedimentation time acquisition module and the sewage discharge time acquisition module are combined to obtain a sedimentation tank parameter acquisition matrix; a parameter acquisition unit is constructed for the oil-water separation control branch, and the separation speed acquisition module and the separation temperature acquisition module are combined to obtain an oil-water separation parameter acquisition matrix; Performing equipment operation status evaluation based on the mixing tank parameter acquisition matrix, the sedimentation tank parameter acquisition matrix, and the oil-water separation parameter acquisition matrix to obtain an equipment status database; An adjustment execution unit is constructed based on the equipment status database to independently and parallelly calculate the mixing tank stirring parameters, the sedimentation tank sedimentation parameters, and the oil-water separation parameters to obtain a branch control sequence; The branch control sequence is input into a parallel processor for multi-task parallel computing to obtain a parallel processing result set, and parameter synchronization and data integration are performed on the parallel processing result set to establish a branch parallel control parameter set, wherein the branch parallel control parameter set includes mixing tank parallel parameters, sedimentation tank parallel parameters and oil-water separation parallel parameters.
4. The multi-device coordinated control method for drilling mud processing according to claim 3, characterized in that: The deep reinforcement learning training is performed according to the branch parallel control parameter set, a reinforcement learning model including a state space, an action space and a reward function is constructed, and an optimal parameter adjustment strategy is generated, including: Parameter encoding is performed on the mixing tank parallel parameters, the sedimentation tank parallel parameters, and the oil-water separation parallel parameters in the branch parallel control parameter set to establish a deep learning basic model; Based on the deep learning basic model, a state space is constructed, and the stirring speed state and stirring time state of the mixing tank equipment, the settling time state and sewage discharge time state of the sedimentation tank equipment, and the separation speed state and separation temperature state of the oil-water separation equipment are vectorized to obtain the equipment state feature vector; Constructing an action space for the equipment state feature vector, setting an adjustment interval for the mixing tank stirring parameter, an adjustment interval for the settling tank settling time, and an adjustment interval for the oil-water separation speed, and obtaining a parameter adjustment action vector; Inputting the equipment state feature vector and the parameter adjustment action vector into a reward function construction unit, and constructing a reward metric function model based on processing indicators of mud density, viscosity, sand content, and oil content; Inputting the equipment state feature vector, the parameter adjustment action vector, and the reward metric function model into the digital twin module, performing deep reinforcement learning training through the equipment operation state simulation unit, the mud property prediction unit, and the parameter optimization unit to obtain a reinforcement learning training sample set; Performing policy gradient calculation on the reinforcement learning training sample set, and iteratively optimizing the neural network parameters of the deep learning basic model using a deep deterministic policy gradient algorithm to obtain target policy network parameters; A temporal difference learning model is constructed based on the target strategy network parameters, the state-action value function is iteratively calculated, and the value function is approximated by the Bellman equation to obtain the target value network parameters; Parameter decision is made according to the target strategy network parameters and the target value network parameters, and an optimal parameter adjustment strategy is generated through action probability distribution calculation.
5. The multi-device coordinated control method for drilling mud processing according to claim 4, characterized in that: The policy gradient calculation is performed on the reinforcement learning training sample set, and the neural network parameters of the deep learning basic model are iteratively optimized using a deep deterministic policy gradient algorithm to obtain target policy network parameters, including: Batch sampling is performed on the reinforcement learning training sample set, and the current state sample, the executed action sample, and the next state sample are combined into a time series training sequence to obtain a policy gradient training sample matrix; Input the policy gradient training sample matrix into the target network, calculate the target Q value through forward propagation, and input the target Q value into the current network to calculate the action value to obtain the current Q value; Performing a temporal difference error calculation on the target Q value and the current Q value to obtain a policy network loss value, and inputting the policy network loss value into a back propagation algorithm to calculate the gradient vector of the network parameters and obtain a parameter update direction; Performing a soft update calculation on the network weights based on the parameter update direction to obtain a weight update increment, and superimposing the weight update increment with the current network parameters to obtain updated network parameters; The updated network parameters are subjected to parameter clipping to obtain modified network parameters, and the modified network parameters are input into the target network for parameter synchronization update to obtain target strategy network parameters.
6. The multi-device coordinated control method for drilling mud processing according to claim 5, characterized in that: Based on the optimal parameter adjustment strategy, the mud density data and mud viscosity data at the outlet of the mixing tank equipment, the sand content data at the outlet of the sedimentation tank equipment, and the oil content data at the outlet of the oil-water separation equipment are monitored and threshold-triggered, and dynamic adjustment control instructions are output, including: The density sensor and viscosity sensor are used to sample and detect the outlet mud of the mixing tank equipment to obtain the monitoring sequence of the mixing tank outlet parameters; the sand content sensor is used to sample and detect the outlet mud of the sedimentation tank equipment to obtain the monitoring sequence of the sedimentation tank outlet parameters; the oil content sensor is used to sample and detect the outlet liquid of the oil-water separation equipment to obtain the monitoring sequence of the oil-water separation outlet parameters; Inputting the mixing tank outlet parameter monitoring sequence, the sedimentation tank outlet parameter monitoring sequence, and the oil-water separation outlet parameter monitoring sequence into a closed-loop feedback control system, performing numerical calculations on mud density, viscosity, sand content, and oil content to obtain a real-time monitoring data matrix; Based on the preset threshold range in the optimal parameter adjustment strategy, the real-time monitoring data matrix is compared and calculated with the mud density threshold, viscosity threshold, sand content threshold, and oil content threshold to obtain a parameter deviation vector; Sorting the importance and prioritizing the parameter deviation vectors to obtain an adjustment priority sequence, inputting the adjustment priority sequence into a compensation controller, and generating a dynamic adjustment control instruction according to the optimal parameter adjustment strategy; The mixing tank stirring speed instruction, the sedimentation tank sedimentation time instruction and the oil-water separation speed instruction in the dynamic adjustment control instruction are respectively sent to the corresponding actuators through the communication interface.
7. A multi-device coordinated control device for drilling mud processing, characterized in that: The device is used to execute the multi-device coordinated control method for drilling mud processing according to any one of claims 1 to 6, comprising: An acquisition module is used to obtain the stirring parameters of the mixing tank equipment, the settling time parameters of the sedimentation tank equipment, and the separation efficiency parameters of the oil-water separation equipment, and obtain the initial collaborative control parameter set by simultaneously solving the steady-state control equation and the dynamic adjustment equation; a parallel processing module for performing branch control architecture division and parallel processing on the mixing tank device, the sedimentation tank device, and the oil-water separation device based on the initial collaborative control parameter set to obtain a branch parallel control parameter set; A reinforcement learning module is used to perform deep reinforcement learning training based on the branch parallel control parameter set, build a reinforcement learning model including a state space, an action space, and a reward function, and generate an optimal parameter adjustment strategy; The output module is used to monitor and perform threshold triggering processing on the mud density data and mud viscosity data at the outlet of the mixing tank equipment, the sand content data at the outlet of the sedimentation tank equipment, and the oil content data at the outlet of the oil-water separation equipment based on the optimal parameter adjustment strategy, and output dynamic adjustment control instructions.
8. A multi-device coordinated control device for drilling mud processing, characterized in that: The multi-device collaborative control device for drilling mud processing includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory to enable the multi-device collaborative control device for drilling mud processing to execute the multi-device collaborative control method for drilling mud processing according to any one of claims 1-6.
9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the multi-device collaborative control method for drilling mud processing according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Electric heating drying room energy consumption control method with self-adaptive capability
CN118349049A
Shield slurry flocculation filter pressing system based on machine learning, optimization method and computer readable storage medium
CN119219291A