Dimethyl sulfoxide production parameter control method and system based on reinforcement learning

Through a reinforcement learning-based method, using historical process data and digital twin simulation models, a dimethyl sulfoxide production parameter control system was constructed, which solved the adaptability problems of traditional control methods under nonlinear and complex dynamic conditions, achieved high-fit modeling and dynamic response, and improved the stability and adaptability of the control strategy.

CN120686751APending Publication Date: 2025-09-23JIANGSU YUEHUA PETROCHEMICAL ENG CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510833263.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Traditional control methods have weak adaptability to nonlinear and complex dynamic conditions and are prone to control errors under load fluctuations or disturbances. Model-based control methods are highly dependent on accurate modeling of the system, and the DMSO production process involves multiple uncertain factors, which makes it difficult for the model to fully reflect the actual working conditions and the control effect is unstable.

Method used

By obtaining historical process operation data of the dimethyl sulfoxide production process, performing correlation analysis to screen key control variables, building a digital twin simulation model, defining the state space, action space and reward function, and using the PPO algorithm to train and optimize the reinforcement learning model, the optimal control strategy is output to achieve intelligent control of production parameters.

Benefits of technology

Achieve continuous self-learning and adjustment in complex dynamic environments, improve the system's responsiveness and adaptability to nonlinear working conditions and load disturbances, reduce dependence on precise physical models, and improve the generalization ability and control stability of control strategies under various uncertainties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120686751A_ABST
    Figure CN120686751A_ABST
Patent Text Reader

Abstract

The invention provides a dimethyl sulfoxide production parameter control method and system based on reinforcement learning, and relates to the technical field of production parameter control, and the method comprises the steps: obtaining historical process operation data in a dimethyl sulfoxide production process; performing correlation analysis on the historical process operation data, and screening key control variables influencing the yield and the energy consumption; based on the historical process operation data and the key control variables, constructing a digital twinborn simulation model for providing an interaction environment for reinforcement learning; defining a state space, an action space and a reward function, and constructing a dimethyl sulfoxide production parameter control model based on reinforcement learning; based on the reward function and the digital twinborn simulation model, training and optimizing the dimethyl sulfoxide production parameter control model through a PPO algorithm, and outputting an optimal control strategy; and controlling the production parameters of dimethyl sulfoxide according to the optimal control strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of production parameter control, and in particular to a dimethyl sulfoxide production parameter control method and system based on reinforcement learning. Background Art

[0002] With the rapid development of modern manufacturing, CNC machine tools, as core equipment for automated production, have been widely used in machining, precision manufacturing, and other fields. However, during complex machining processes, machine tool collisions are prone to occur due to improper operation, program errors, or equipment failures. This not only causes equipment damage but can also seriously impact workpiece quality and production efficiency. Therefore, how to monitor machine tool operating status in real time, predict potential collision risks, and take preventive measures has become a pressing issue in modern manufacturing.

[0003] Currently, DMSO production process control primarily relies on traditional PID (proportional-integral-derivative) controllers, fuzzy control systems, or expert rule systems. These methods are typically based on empirical formulas or parameters set by process engineers, and are adjusted based on specific feedback signals. Some high-end plants are beginning to introduce model-based optimization control (MPC) methods, which dynamically adjust operating parameters by building mathematical models to predict system responses.

[0004] However, traditional control methods have limited adaptability to nonlinear and complex dynamic conditions, and are prone to control errors under load fluctuations or disturbances. Furthermore, model-based control methods rely heavily on accurate system modeling. The DMSO production process involves numerous uncertainties, making it difficult for models to fully reflect actual operating conditions, resulting in unstable control results. Summary of the Invention

[0005] In view of the above shortcomings of the existing technology, the purpose of the embodiments of the present invention is to provide a method for controlling dimethyl sulfoxide production parameters based on reinforcement learning. This method can address the technical issues that traditional control methods have limited adaptability to nonlinear and complex dynamic conditions and are prone to control errors under load fluctuations or disturbances. Furthermore, model-based control methods rely heavily on accurate system modeling. However, the DMSO production process involves multiple uncertainties, making it difficult for the model to fully reflect actual operating conditions, resulting in unstable control effects.

[0006] A first aspect of an embodiment of the present invention provides a method for controlling dimethyl sulfoxide production parameters based on reinforcement learning, comprising:

[0007] S1: Obtain historical process operation data during the production of dimethyl sulfoxide;

[0008] S2: performing correlation analysis on the historical process operation data to screen key control variables that affect yield and energy consumption;

[0009] S3: Based on the historical process operation data and the key control variables, construct a digital twin simulation model for providing an interactive environment for reinforcement learning;

[0010] S4: Define the state space, action space, and reward function, and build a DMSO production parameter control model based on reinforcement learning;

[0011] S5: Based on the reward function and the digital twin simulation model, the dimethyl sulfoxide production parameter control model is trained and optimized by the PPO algorithm to output an optimal control strategy;

[0012] S6: Controlling the production parameters of dimethyl sulfoxide according to the optimal control strategy.

[0013] A second aspect of an embodiment of the present invention provides a dimethyl sulfoxide production parameter control system based on reinforcement learning, comprising: a processor and a memory;

[0014] The memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the dimethyl sulfoxide production parameter control method based on reinforcement learning as described in the first aspect are implemented.

[0015] In a third aspect of an embodiment of the present invention, a readable storage medium is proposed, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the dimethyl sulfoxide production parameter control method based on reinforcement learning as described in the first aspect are implemented.

[0016] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0017] In an embodiment of the present invention, a dimethyl sulfoxide production parameter control model based on reinforcement learning is constructed by defining the state space, action space, and reward function, and the model is trained and optimized in combination with the PPO algorithm to output the optimal control strategy, which can achieve continuous self-learning and adjustment in a complex dynamic environment, significantly improving the system's responsiveness and adaptability to nonlinear working conditions and load disturbances. At the same time, key variable screening is carried out based on historical process data of the actual production process, and a digital twin simulation environment is constructed to provide interactive support for reinforcement learning. Without relying on precise physical models, high-fit modeling and dynamic response simulation of the production process are achieved, reducing dependence on mechanism models, and improving the generalization ability and control stability of the control strategy under various uncertain factors. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are only for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference symbols represent the same components. Obviously, the drawings described below are only some embodiments of the present invention. It is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0019] Figure 1 1 is a flow chart of a method for controlling dimethyl sulfoxide production parameters based on reinforcement learning provided by an embodiment of the present invention;

[0020] Figure 2 This is a structural diagram of a dimethyl sulfoxide production parameter control system based on reinforcement learning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the scope of protection of the present invention.

[0022] The following describes in detail the dimethyl sulfoxide production parameter control method based on reinforcement learning provided by the embodiment of the present invention through specific embodiments and application scenarios in conjunction with the accompanying drawings.

[0023] Reference Manual Figure 1 , shows a flow chart of a dimethyl sulfoxide production parameter control method based on reinforcement learning provided by an embodiment of the present invention.

[0024] An embodiment of the present invention provides a method for controlling dimethyl sulfoxide production parameters based on reinforcement learning, which may include the following steps:

[0025] S1: Obtain historical process operation data for the dimethyl sulfoxide production process.

[0026] Historical process operation data refers to the collection of data such as process parameters, state variables, and output indicators collected and recorded during actual industrial production operations. These data reflect the actual response of the production system under different operating conditions and serve as a fundamental data source for building digital twin models, performing variable screening, and conducting reinforcement learning training.

[0027] In one possible implementation, the historical process operation data includes:

[0028] Heating temperature, feed flow rate, stirring speed, feed concentration, cooling rate, initial reactor charge, reactor internal temperature, reactor pressure, reactor volume, DMSO concentration, reaction time, final DMSO yield, energy consumption per unit of output, and total reaction time.

[0029] In the embodiment of the present invention, obtaining historical process operation data provides a logical starting point for data support, enhances practical feasibility, and lays a solid foundation for subsequent modeling, reinforcement learning training, and strategy optimization steps.

[0030] S2: Conduct correlation analysis on historical process operation data to screen key control variables that affect yield and energy consumption.

[0031] Among them, correlation analysis is a statistical method used to measure the strength of the linear relationship between two variables.

[0032] Key control variables are adjustable parameters that significantly affect the quality of the final product (e.g., yield, energy consumption, etc.) during the production process. They are variables selected for optimization goals (e.g., increasing yield, reducing energy consumption, etc.) and process control.

[0033] In a possible implementation, S2 specifically includes:

[0034] S201: Set the target variables to DMSO yield and unit energy consumption.

[0035] S202: Setting control variables including reaction temperature, oxygen flow rate, raw material concentration, pressure, reaction time and addition ratio.

[0036] S203: Calculate the correlation evaluation values ​​of each control variable with the DMSO yield and the unit energy consumption respectively through the Pearson correlation coefficient to obtain a plurality of yield correlation evaluation values ​​and unit energy consumption correlation evaluation values.

[0037] Among them, the Pearson Correlation Coefficient is a statistic that measures the degree of linear correlation between two variables and is used to determine whether the two variables are positively correlated, negatively correlated, or uncorrelated.

[0038] In the embodiment of the present invention, the Pearson correlation coefficient is used to numerically measure the impact of each variable on the target. It is not as subjective as manual experience judgment, and has strong comparability and high scientificity.

[0039] S204: Calculate a comprehensive evaluation value based on the yield correlation evaluation value and the unit energy consumption correlation evaluation value corresponding to each control variable.

[0040] In a possible implementation, the comprehensive evaluation value is calculated as follows:

[0041] P i =λ|r i (y) |+(1-λ)|r i (e) |

[0042] Among them, P i represents the comprehensive evaluation value of the i-th control variable, r i (y) represents the correlation evaluation value of the i-th control variable with respect to the yield, r i (e) represents the correlation evaluation value of the i-th control variable relative to unit energy consumption, || represents the absolute value sign, λ represents the weight coefficient, λ∈(0,1).

[0043] In an embodiment of the present invention, a weighted comprehensive evaluation value is introduced to achieve a balance evaluation among multiple objectives. For example, the productivity weight can be increased or the energy consumption weight can be reduced according to actual needs, which is flexible and practical.

[0044] S205: Filter control variables whose comprehensive evaluation values ​​are greater than preset comprehensive evaluation values ​​as key control variables.

[0045] It should be noted that those skilled in the art can preset the size of the evaluation value according to actual needs, and the present invention is not limited thereto.

[0046] In an embodiment of the present invention, through correlation analysis, combined with the two optimization objectives of yield and unit energy consumption, key control variables that significantly affect system performance are quantitatively screened, which not only improves the scientificity and pertinence of the model input characteristics, but also improves the computational efficiency and interpretability of subsequent modeling and optimization.

[0047] S3: Based on historical process operation data and key control variables, a digital twin simulation model is constructed to provide an interactive environment for reinforcement learning.

[0048] In one possible implementation, the digital twin simulation model is trained based on historical process operation data, using supervised learning to construct a state transition function model. Specifically, key control variables and system state variables are selected as input features, and a feedforward neural network is used to model the mapping between the current state and control actions and the next state (or yield, energy consumption), thereby realizing a high-fidelity process simulator for a reinforcement learning interactive environment.

[0049] In an embodiment of the present invention, a digital twin simulation model trained based on historical data is constructed, which provides an interactive, high-fidelity, safe and controllable training environment for reinforcement learning, avoids high-risk trial and error operations in real systems, and significantly improves the efficiency and generalization ability of policy optimization.

[0050] S4: Define the state space, action space, and reward function, and build a dimethyl sulfoxide production parameter control model based on reinforcement learning.

[0051] The state space is the complete set of descriptions of the state of the environment at each point in time by the reinforcement learning agent. Each "state" contains the current operating status of the system and serves as the basis for the agent's decision-making.

[0052] Among them, the action space refers to the set of operations that the agent can choose in each state, that is, the control parameters that can be adjusted.

[0053] Among them, the reward function is a mechanism used to measure the quality of actions in reinforcement learning. It represents the contribution of each action to the system goal and is the "driving force" of the reinforcement learning agent's optimization strategy.

[0054] In a possible implementation, S4 specifically includes:

[0055] S401: Define the state space of dimethyl sulfoxide during the production process:

[0056] S=[T,P,C feed ,F O2 ,Y DMSO ,E unit ,t]

[0057] Where S represents the state space, T represents the current reactor temperature, P represents the current system pressure, and C feed Indicates feed concentration, F O2 Indicates oxygen flow rate, Y DMOS represents the yield of dimethyl sulfoxide, E unit represents unit energy consumption, and t represents the current reaction time.

[0058] S402: Define the action space of dimethyl sulfoxide in the production process:

[0059] A=[ΔT,ΔC feed ,ΔF O2 ]

[0060] Among them, A represents the action space, ΔT represents the temperature adjustment range, ΔC feed Indicates the adjustment range of raw material concentration, ΔF O2 Indicates the oxygen flow adjustment range.

[0061] S403: Define a reward function for dimethyl sulfoxide during the production process.

[0062] In one possible implementation, the reward function is determined by:

[0063] In order to balance the relationship between maximizing yield and minimizing energy consumption, a weighted approach is used to determine the main reward function:

[0064] R main =α·Y DMSO -β·E unit

[0065] Among them, R main represents the main reward function, α represents the weight coefficient of the yield, Y DMSO represents the yield of dimethyl sulfoxide, β represents the weight coefficient of unit energy consumption, Table E unit Indicates unit energy consumption.

[0066] Introduce a penalty term based on the logarithmic barrier function and determine the constraint penalty function:

[0067]

[0068] Among them, C i represents the constraint penalty function of the i-th control variable, ε represents the constraint penalty adjustment coefficient, log represents the logarithmic function, represents the upper limit of the i-th control variable, x (i) represents the i-th control variable, represents the lower limit of the i-th control variable.

[0069] It should be noted that the introduction of a penalty term based on a logarithmic barrier function can prevent the agent from exceeding the operation boundary during the exploration process.

[0070] Based on the main reward function and the constraint penalty function, construct the reward function:

[0071]

[0072] Where R represents the reward function.

[0073] In this embodiment of the present invention, by introducing weighted coefficients for yield and energy consumption into the main reward function, the intelligent agent can simultaneously optimize multiple objectives, avoiding optimization bias caused by a single objective. Furthermore, by introducing a logarithmic barrier function penalty term, the intelligent agent is prevented from exceeding operational boundaries during learning, ensuring that the production process proceeds within a safe and reasonable range, avoiding action selection that is inconsistent with actual process conditions, and improving the reliability and stability of the algorithm.

[0074] S404: Based on the state space, action space, and reward function, a Markov decision process is used to determine a dimethyl sulfoxide production parameter control model based on reinforcement learning:

[0075] M=(S,A,B,R,γ)

[0076] Where M represents the dimethyl sulfoxide production parameter control model based on reinforcement learning, B represents the state transition function, R represents the reward function, and γ represents the discount factor.

[0077] The Markov Decision Process (MDP), a mathematical framework for modeling the interactive decision-making process between an intelligent agent and its environment, is widely used in reinforcement learning. It describes how a system makes optimal decisions based on its current state in an uncertain environment to maximize long-term returns.

[0078] It should be noted that the state transfer function is approximated by the digital twin model, which is based on process operation data and trained using supervised learning methods to achieve the ability to predict the next state from the current state and action.

[0079] In this embodiment, the Markov decision process (MDP) is used to clarify the state transition and reward feedback mechanisms of reinforcement learning, facilitating the construction of a more rigorous control model. Furthermore, a digital twin model of the state transition function is approximated and can be trained based on historical process data, providing a response closer to the actual production process.

[0080] S5: Based on the reward function and digital twin simulation model, the PPO algorithm is used to train and optimize the dimethyl sulfoxide production parameter control model and output the optimal control strategy.

[0081] Among them, PPO (Proximal Policy Optimization) is a reinforcement learning policy optimization algorithm, an improved form of the Policy Gradient method. It significantly improves the stability and efficiency of the training process while maintaining optimization capabilities. It is currently widely used in tasks such as robotic control, industrial optimization, and game agents.

[0082] In a possible implementation, S5 specifically includes:

[0083] S501: Establish a training environment for reinforcement learning interaction based on the state space, action space, reward function and digital twin simulation model.

[0084] S502: Initialize the policy function, the value function, and the hyperparameters for training, wherein the hyperparameters include the learning rate, the discount factor, the shear coefficient, and the maximum number of training rounds.

[0085] It should be noted that both the policy function and the value function are constructed by a multi-layer feedforward neural network. The output of the policy network is the action probability distribution, and the output of the value function network is the expected cumulative reward value of the current state.

[0086] In this embodiment of the present invention, the policy network and value function network required for reinforcement learning are initialized, and training hyperparameters including the learning rate, discount factor, and shear coefficient are configured. By properly designing the network structure (such as a feedforward neural network) and parameter range, the convergence speed and stability of the training process can be effectively improved, while providing a good foundation for subsequent strategy and value evaluation, thereby enhancing training efficiency and the expressiveness of control strategies.

[0087] S503: Construct a clipping agent objective function for training the policy network through the policy gradient optimization principle.

[0088] In a possible implementation, S503 specifically includes:

[0089] S5031: Based on the policy gradient optimization principle, determine the optimization direction of the policy function:

[0090]

[0091] Among them, θ represents the parameters of the policy function, ▽ represents the gradient operator, J represents the expected cumulative reward, represents the gradient of the parameter θ, t represents the time step index, s t represents the state of the tth time step, a t represents the action at the t-th time step, log represents the logarithmic function, π θ (a t |s t ) means in state s t Next, perform action a t The strategy probability, Q(s t ,a t ) means in state s t Next, perform action a t The expected cumulative reward obtained after

[0092] S5032: Based on the optimization direction of the policy function, construct the clipping agent objective function for training the policy network:

[0093]

[0094] Among them, J surr represents the shear proxy objective function, rj represents the strategy probability ratio, A j represents the advantage function, μ represents the clipping coefficient, clip(r j ,1-μ,1+μ) means cutting the strategy ratio.

[0095] Alternatively, the advantage function can be calculated by means of a generalized advantage estimate:

[0096] A j =R j -V0(s j )

[0097] Among them, R j represents the actual cumulative reward of the jth sample, V0(s j ) represents the predicted value of the state value function.

[0098] It should be noted that in order to improve the stability and security of policy training, a shear proxy objective function is constructed based on the obtained optimization direction for actual training of the policy network.

[0099] In this embodiment of the present invention, a clipping mechanism is introduced to construct a proxy objective function based on the policy gradient direction. By limiting the update amplitude between the old and new policies, the stability and robustness of the training process are enhanced. Furthermore, the combination of advantage functions for evaluation improves sample utilization efficiency, enabling the policy network to more effectively optimize the control strategy, ultimately achieving better process control results.

[0100] S504: Construct a loss function for determining the cumulative reward of the regression state:

[0101]

[0102] Among them, J value Represents the loss of the value function, j represents the training sample number, V θ (s j ) represents the state s given by the value function network j The estimated return, R j Indicates the actual accumulated rewards.

[0103] In this embodiment of the present invention, a loss function is constructed to fit the cumulative reward, guiding the value function network to learn the expected long-term payoff for each state. Accurate value function estimation provides more reliable reference information for the policy network, improving the calculation accuracy of the advantage function and indirectly improving the quality of policy training.

[0104] S505: Obtain the control action of the policy function in the training environment.

[0105] S506: Apply the control action to the digital twin simulation model to simulate the production process of dimethyl sulfoxide to obtain the next state and immediate reward.

[0106] S507: Update the policy network using the cut proxy objective function.

[0107] S508: Update the value function network using the loss function.

[0108] S509: Repeat steps S506 to S508, and when the number of training rounds reaches the maximum number of training rounds or the fluctuation of the average reward in a plurality of consecutive rounds is less than a threshold, stop training and output the optimal control strategy.

[0109] It should be noted that those skilled in the art can set the threshold value according to actual needs, and the present invention does not limit this.

[0110] In this embodiment of the present invention, the reinforcement learning training process constructed in this step enables efficient iterative learning of intelligent agents in a simulation environment, generating optimal control strategies with practical application value. The clipping of the proxy objective function improves training stability, the value function network enhances state assessment accuracy, and the training environment is driven by a digital twin model, avoiding the trial-and-error risks of a real system. This approach possesses strong generalization, deployability, and energy optimization potential, making it suitable for reinforcement learning optimization tasks in complex industrial process control.

[0111] S6: Control the production parameters of dimethyl sulfoxide according to the optimal control strategy.

[0112] Specifically, the trained optimal control strategy is deployed to an actual or simulated production system. By collecting the current process state in real time as input, the strategy function outputs control actions to adjust key control parameters such as reactor temperature, feed concentration, and oxygen flow rate. The system continuously executes the strategy after receiving new state feedback, achieving closed-loop intelligent optimization control of the dimethyl sulfoxide production process.

[0113] In the embodiment of the present invention, closed-loop intelligent control is achieved by deploying the optimal control strategy, which not only improves the stability and response efficiency of the production process, but also effectively optimizes the balance between yield and energy consumption, and enhances the system's adaptability and intelligence level under complex working conditions.

[0114] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0115] In an embodiment of the present invention, a dimethyl sulfoxide production parameter control model based on reinforcement learning is constructed by defining the state space, action space, and reward function, and the model is trained and optimized in combination with the PPO algorithm to output the optimal control strategy, which can achieve continuous self-learning and adjustment in a complex dynamic environment, significantly improving the system's responsiveness and adaptability to nonlinear working conditions and load disturbances. At the same time, key variable screening is carried out based on historical process data of the actual production process, and a digital twin simulation environment is constructed to provide interactive support for reinforcement learning. Without relying on precise physical models, high-fit modeling and dynamic response simulation of the production process are achieved, reducing dependence on mechanism models, and improving the generalization ability and control stability of the control strategy under various uncertain factors.

[0116] Reference Manual Figure 2 , which shows a structural schematic diagram of a dimethyl sulfoxide production parameter control system based on reinforcement learning provided by an embodiment of the present invention.

[0117] The embodiment of the present invention provides a dimethyl sulfoxide production parameter control system 20 based on reinforcement learning, comprising: a processor 201 and a memory 202;

[0118] The memory 202 stores a program or instruction that can be run on the processor 201. When the program or instruction is executed by the processor 201, the steps of the above-mentioned dimethyl sulfoxide production parameter control method based on reinforcement learning are implemented, and the same technical effect can be achieved. To avoid repetition, the present invention will not be repeated.

[0119] It should be understood that the processor 201 in the embodiment of the present invention may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0120] It should also be understood that the memory 202 in the embodiment of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0121] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0122] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0123] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0124] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0125] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0126] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0127] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0128] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0129] An embodiment of the present invention provides a readable storage medium including: a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by the processor, the steps of the above-mentioned dimethyl sulfoxide production parameter control method based on reinforcement learning are implemented, and the same technical effect can be achieved. To avoid repetition, the present invention will not be described in detail.

[0130] Finally, it should be noted that the above embodiments are merely illustrative of the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be covered by the scope of protection of the present invention.

Claims

1. A dimethyl sulfoxide production parameter control method based on reinforcement learning, characterized in that: include: S1: Obtain historical process operation data during the production of dimethyl sulfoxide; S2: performing correlation analysis on the historical process operation data to screen key control variables that affect yield and energy consumption; S3: Based on the historical process operation data and the key control variables, construct a digital twin simulation model for providing an interactive environment for reinforcement learning; S4: Define the state space, action space, and reward function, and build a DMSO production parameter control model based on reinforcement learning; S5: Based on the reward function and the digital twin simulation model, the dimethyl sulfoxide production parameter control model is trained and optimized by the PPO algorithm to output an optimal control strategy; S6: Controlling the production parameters of dimethyl sulfoxide according to the optimal control strategy.

2. The method for controlling dimethyl sulfoxide production parameters based on reinforcement learning according to claim 1, characterized in that: The historical process operation data includes: Heating temperature, feed flow rate, stirring speed, feed concentration, cooling rate, initial reactor charge, reactor internal temperature, reactor pressure, reactor volume, DMSO concentration, reaction time, final DMSO yield, energy consumption per unit of output, and total reaction time.

3. The dimethyl sulfoxide production parameter control method based on reinforcement learning according to claim 1, characterized in that: The S2 specifically includes: S201: Set the target variables to DMSO yield and specific energy consumption; S202: Setting control variables including reaction temperature, oxygen flow rate, raw material concentration, pressure, reaction time and addition ratio; S203: calculating the correlation evaluation value between each of the control variables and the DMSO yield and the unit energy consumption using the Pearson correlation coefficient, to obtain a plurality of yield correlation evaluation values ​​and unit energy consumption correlation evaluation values; S204: Calculating a comprehensive evaluation value based on the yield correlation evaluation value and the unit energy consumption correlation evaluation value corresponding to each control variable; S205: Filter the control variables corresponding to the comprehensive evaluation values ​​that are greater than the preset comprehensive evaluation values ​​as the key control variables.

4. The dimethyl sulfoxide production parameter control method based on reinforcement learning according to claim 3, characterized in that: The comprehensive evaluation value is calculated as follows: P i =λ|r i (y) |+(1-λ)|r i (e) |; Among them, P i represents the comprehensive evaluation value of the i-th control variable, r i (y) represents the correlation evaluation value of the i-th control variable with respect to the yield, r i (e) represents the correlation evaluation value of the i-th control variable relative to unit energy consumption, || represents the absolute value sign, λ represents the weight coefficient, λ∈(0,1).

5. The method for controlling dimethyl sulfoxide production parameters based on reinforcement learning according to claim 1, characterized in that: The S4 specifically includes: S401: defining the state space of dimethyl sulfoxide during the production process; S402: defining the action space of dimethyl sulfoxide during the production process; S403: defining the reward function during the production process of dimethyl sulfoxide; S404: Based on the state space, the action space, and the reward function, a dimethyl sulfoxide production parameter control model based on reinforcement learning is determined using a Markov decision process.

6. The method for controlling dimethyl sulfoxide production parameters based on reinforcement learning according to claim 5, characterized in that: The reward function is determined by: In order to balance the relationship between maximizing yield and minimizing energy consumption, a weighted approach is used to determine the main reward function: R main =α·Y DMSO -b·E unit ; Among them, R main represents the main reward function, α represents the weight coefficient of the yield, Y DMSO represents the yield of dimethyl sulfoxide, β represents the weight coefficient of unit energy consumption, E unit Indicates unit energy consumption; Introduce a penalty term based on the logarithmic barrier function and determine the constraint penalty function: Among them, C i represents the constraint penalty function of the i-th control variable, ε represents the constraint penalty adjustment coefficient, log represents the logarithmic function, represents the upper limit of the i-th control variable, x (i) represents the i-th control variable, represents the lower limit of the i-th control variable; Based on the main reward function and the constraint penalty function, the reward function is constructed: Where R represents the reward function.

7. The method for controlling dimethyl sulfoxide production parameters based on reinforcement learning according to claim 1, characterized in that: The S5 specifically includes: S501: Establishing a training environment for reinforcement learning interaction based on the state space, the action space, the reward function, and the digital twin simulation model; S502: Initializing a policy function, a value function, and hyperparameters for training, wherein the hyperparameters include a learning rate, a discount factor, a clipping coefficient, and a maximum number of training rounds; S503: Constructing a clipping agent objective function for training the policy network through the policy gradient optimization principle; S504: Constructing a loss function for determining the cumulative reward of the regression state; S505: Obtaining control actions of the policy function in the training environment; S506: Applying the control action to the digital twin simulation model to simulate the production process of dimethyl sulfoxide to obtain a next state and an immediate reward; S507: Update the policy network using the cut proxy objective function; S508: Update the value function network using the loss function; S509: Repeat steps S506 to S508, stop training when the number of training rounds reaches the maximum number of training rounds or the fluctuation of the average reward in multiple consecutive rounds is less than a threshold, and output the optimal control strategy.

8. The method for controlling dimethyl sulfoxide production parameters based on reinforcement learning according to claim 7, characterized in that: The S503 specifically includes: S5031: Determine the optimization direction of the policy function based on the policy gradient optimization principle; S5032: Based on the optimization direction of the policy function, construct the clipping agent objective function for training the policy network.

9. A dimethyl sulfoxide production parameter control system based on reinforcement learning, characterized in that: include: processor and memory; The memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the dimethyl sulfoxide production parameter control method based on reinforcement learning as described in any one of claims 1 to 8 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the dimethyl sulfoxide production parameter control method based on reinforcement learning as described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Method and system for monitoring technological process of polyester net

    CN120891808A

  • Industrial control method and system for glyceryl triacetate production process

    CN121254788A