Method and device for optimizing polymerization reaction temperature based on generative AI

Through the temperature optimization method based on generative AI, real-time data is used to predict future temperature changes and dynamically adjust the opening of the thermally conductive oil valve, the problem of large temperature fluctuations in the polymerization reaction is solved, and more accurate temperature control and higher production efficiency are achieved.

CN120581085APending Publication Date: 2025-09-02HANGZHOU TRANSFAR FINE CHEM CO LTD +3

Patent Information

Application Number
CN202511088104.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

It is difficult to adapt to the rapid changes in the temperature in the reactor during the polymerization reaction, resulting in large temperature fluctuations, affecting the polymer molecular weight distribution and product performance.

Method used

The temperature optimization method based on generative AI is adopted, and the generational AI model with pre-trained acquisition of multi-source data input in real time is used to output future temperature change trajectories, and the thermal oil valve opening is dynamically adjusted in combination with reinforcement learning strategies to optimize temperature control.

Benefits of technology

More precise temperature control is achieved, reducing temperature fluctuations, improving product quality and production efficiency, and reducing energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581085A_ABST
    Figure CN120581085A_ABST
Patent Text Reader

Abstract

The invention discloses a polymerization reaction temperature optimization method and device based on generative AI, and a server side method comprises the steps: collecting and preprocessing multi-source data in a reaction kettle for polymerization reaction in real time through a sensor, and obtaining real-time state information in the reaction kettle; inputting the real-time state information into a pre-trained generative AI model, and outputting a temperature change track in a future preset time period; according to the temperature change track, a preset reinforcement learning strategy is used for dynamically adjusting the opening parameter of the heat conduction oil valve, and a heat conduction oil valve opening instruction is generated; and sending a conduction oil valve opening instruction to a control system for controlling the conduction oil valve so as to optimize the polymerization reaction temperature. According to the embodiment of the invention, the system can know the temperature change trend in advance through the pre-trained generative AI model and the reinforcement learning strategy, thereby achieving the more precise temperature control, facilitating the reduction of temperature fluctuation, and improving the product quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent manufacturing technology, and in particular to a polymerization reaction temperature optimization method and device based on generative AI. Background Art

[0002] Polymerization is a core step in the production of polymer materials, widely used in the manufacture of materials such as plastics, rubber, and fibers. Polymerization reactions are typically carried out in industrial-scale reactors, where temperature control is crucial to ensuring product quality and production efficiency. For example, when producing high-performance engineering plastics, the reactor temperature must be precisely controlled within ±1°C of the set value to ensure uniform molecular weight distribution of the polymer, thereby maintaining the product's mechanical properties, thermal stability, and optical properties.

[0003] Currently, polymerization reaction temperature control primarily relies on the traditional PID (proportional-integral-derivative) control method. PID control achieves temperature feedback control by adjusting three parameters: proportional, integral, and differential. In practice, the PID controller calculates a deviation signal based on real-time data collected by the temperature sensor and adjusts the heating or cooling system to maintain a stable temperature within the reactor.

[0004] However, while PID control is widely used in industry, it has significant limitations when dealing with complex dynamic systems. PID control struggles to adapt to rapid temperature changes within reactors, leading to large temperature fluctuations, typically up to 3°C. These large temperature fluctuations can lead to a broadened polymer molecular weight distribution (PDI > 2.5), thus degrading product performance. Summary of the Invention

[0005] The present embodiments provide a method and apparatus for optimizing polymerization reaction temperature based on generative AI. To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is provided below. This summary is not intended to be a comprehensive review, identify key or important elements, or delineate the scope of these embodiments. Its sole purpose is to present some concepts in a simplified form, serving as a prelude to the detailed description that follows.

[0006] In a first aspect, an embodiment of the present application provides a polymerization reaction temperature optimization method based on generative AI, which is applied to a server. The method includes: The sensors collect and pre-process multi-source data in the reactor used for polymerization reaction in real time to obtain real-time status information in the reactor; Input real-time status information into a pre-trained generative AI model to output the temperature change trajectory within a preset time period in the future; Based on the temperature change trajectory, the preset reinforcement learning strategy is used to dynamically adjust the opening parameters of the thermal oil valve and generate the thermal oil valve opening instruction; Send thermal oil valve opening instructions to the control system used to control the thermal oil valve to optimize the polymerization reaction temperature.

[0007] Optionally, generate a pre-trained generative AI model by following these steps, including: Acquire and pre-process historical multi-source data collected in advance for the reactor during a historical time period to obtain historical status information and historical real temperature at each historical moment in the reactor; The historical real temperature at each historical moment is used as the label of the historical status information at each historical moment to annotate the data and obtain the training data set; Use the mean square error expression to construct the loss function; Use loss functions and neural network algorithms to create generative AI models; Based on the training data set, machine learning is performed on the generative AI model to obtain a pre-trained generative AI model.

[0008] Optionally, use loss functions and neural network algorithms to create generative AI models, including: Construct an input layer for receiving the training dataset, where the dimension parameter of the input layer is equal to the number of data categories of the historical multi-source data; A neural network algorithm is used to construct a hidden layer to capture the long-term and short-term dependencies in the time series of the training dataset; A neural network algorithm is used to construct an output layer for predicting the temperature change trajectory within a preset time period in the future; The input layer, hidden layer, output layer, and loss function are connected in sequence to form a generative AI model.

[0009] Optionally, perform machine learning on the generative AI model based on the training dataset to obtain a pre-trained generative AI model, including: Input the training dataset into the generative AI model and output the model loss value; When the model loss value reaches a minimum, a pre-trained generative AI model is generated; or, if the model loss value does not reach a minimum, the model loss value is back-propagated to update the model parameters of the generative AI model, and the step of inputting the training dataset into the generative AI model is continued until the model loss value reaches a minimum; Among them, the expression of the loss function is:

[0010] in, is the model loss value, is the total number of training data in the training dataset, The training data set training data, when predicting the future preset time period containing 10 time steps, is the first time step within 10 time steps time steps, For the The training data in The predicted temperature for each time step, It is The training data in The actual temperature of each time step.

[0011] Optionally, the generative AI model includes an input layer, a hidden layer, an output layer, and a loss function; Input the training dataset into the generative AI model and output the model loss value, including: The input layer obtains the historical state information in the training data set; The hidden layer captures the long-term and short-term dependencies in the time series of each historical state information in the training data set, and outputs the long-term and short-term dependencies in the time series corresponding to each historical state information; The output layer predicts the first temperature change trajectory within a preset future time period corresponding to each historical state information based on the long-term and short-term dependencies in the time series corresponding to each historical state information; The first temperature change trajectory within the future preset time period corresponding to each historical state information and the actual temperature marked by each historical state information are input into the loss function for calculation to obtain the model loss value.

[0012] Optionally, based on the temperature change trajectory, a preset reinforcement learning strategy is used to dynamically adjust the opening parameters of the thermal oil valve to generate a thermal oil valve opening instruction, including: Divide the temperature change trajectory into multiple moments in a preset time period and obtain the temperature prediction value at each moment; The temperature prediction value at each moment, the current temperature of the reactor, and the set value are spliced ​​into a decision vector; The decision vector is input into the preset reinforcement learning strategy, and the score values ​​corresponding to different opening parameters of the thermal oil valve are output; The opening parameter of the thermal oil valve is replaced by the target opening parameter with the highest score value to obtain the actual opening parameter of the thermal oil valve. The target opening parameter with the highest score value is the valve opening parameter that can make the temperature closest to the set value; The actual opening parameters are encoded according to the industrial protocol, added with a timestamp and a check code, and encapsulated as a thermal oil valve opening instruction.

[0013] Optionally, follow these steps to generate a preset reinforcement learning policy, including: The real-time temperature, monomer conversion rate, and stirring power in the reactor are defined as state variables of the environment. The different opening adjustments of the thermal oil valve are defined as executable action variables. A reward function is defined to minimize the cumulative time that the temperature deviates from the set value. Encapsulate state variables, action variables, and reward functions into an interactive sequential decision-making environment; Determine the target state variable at the target time from the temporal decision environment and randomly select the target action variable; Execute the target action variable in conjunction with the reinforcement learning algorithm to obtain the new state variable and target reward value of the preset time step after executing the target action variable; The target state variable, the target action variable, the new state variable, and the target reward value are stored as a four-tuple, and the step of determining the target state variable at the target moment from the sequential decision environment is continued until the number of four-tuples reaches a preset number, thereby obtaining a plurality of experience samples; Based on multiple experience samples, a preset reinforcement learning strategy is generated; the functional expression of the reward function is:

[0014] Among them, R is the reward value, T is the preset time step, The first time steps, is the actual temperature, To set the temperature.

[0015] Optionally, generate a preset reinforcement learning strategy based on multiple experience samples, including: A reinforcement learning algorithm is used to construct a reinforcement learning component, which includes a Q network and a target network. The Q network is used to estimate the value of each state variable-action variable pair, and the target network is used to provide updated targets. Initially, the parameters are the same as those of the Q network. Input each experience sample into the Q network and output the predicted Q value of each experience sample; For each experience sample, the target Q value is calculated using the Q value calculation function of the target network; Calculate the mean squared error loss between the predicted Q value of each experience sample and the target Q value of each experience sample; When the mean square error loss has not reached the minimum, the network parameters of the Q network are updated using the mean square error loss, and the step of inputting each experience sample into the Q network is continued until the mean square error loss reaches the minimum; The network parameters of the Q network are copied to the target network to obtain the preset reinforcement learning strategy.

[0016] Optionally, the function expression of the Q value calculation function of the target network is:

[0017] in, is the target Q value, is the reward value in each experience sample, is the discount factor, is the new state variable, is the action variable taken under the new state variable, Is in state Take action variables The expected return, is to select one from all action variables such that The action with the largest value .

[0018] In a second aspect, an embodiment of the present application provides a polymerization reaction temperature optimization device based on generative AI, the device comprising: A data preprocessing module is used to collect and preprocess multi-source data in the reactor used for polymerization reaction in real time through sensors to obtain real-time status information in the reactor; The temperature change trajectory output module is used to input real-time status information into a pre-trained generative AI model and output the temperature change trajectory within a preset time period in the future; The instruction generation module is used to dynamically adjust the opening parameters of the thermal oil valve according to the temperature change trajectory using a preset reinforcement learning strategy to generate the thermal oil valve opening instruction; The instruction sending module is used to send a thermal oil valve opening instruction to a control system used to control the thermal oil valve to optimize the polymerization reaction temperature.

[0019] The technical solutions provided by the embodiments of the present application may have the following beneficial effects: In an embodiment of the present application, on the one hand, by inputting multi-source data collected and pre-processed in real time into a pre-trained generative AI model, the model can output the temperature change trajectory within a preset time period in the future. The prediction results provide important input information for the reinforcement learning strategy, allowing the system to understand the temperature change trend in advance. The reinforcement learning strategy can more effectively calculate the optimal opening of the thermal oil valve, thereby achieving more precise temperature control, helping to reduce temperature fluctuations and improve product quality. On the other hand, the reinforcement learning strategy can not only respond to the predictions of the generative AI model, but also achieve closed-loop control based on real-time feedback. This dynamic adjustment mechanism significantly improves the flexibility of temperature control, helps to maintain temperature stability in the reactor, reduce energy consumption, and improve production efficiency.

[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0022] Figure 1 1 is a schematic diagram of a method flow of a polymerization reaction temperature optimization method based on generative AI provided in an embodiment of the present application; Figure 2 This is an architectural diagram of a generative AI model provided in an embodiment of the present application; Figure 3 This is an implementation monitoring UI interface for the administrator backend provided in the embodiment of the present application; Figure 4 This is a schematic diagram of the interaction between a server and a client provided in an embodiment of the present application; Figure 5 This is a flow chart of a model training method for a generative AI model provided in an embodiment of the present application; Figure 6 This is a UI interface diagram of a generative AI model training process provided by an embodiment of the present application; Figure 7 This is a flow chart of a method for generating a preset reinforcement learning strategy provided in an embodiment of the present application; Figure 8 Schematic diagram of a polymerization reaction temperature optimization device based on generative AI provided in an embodiment of the present application; Figure 9 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following description and the drawings sufficiently illustrate specific embodiments of the application to enable those skilled in the art to practice them.

[0024] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0025] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0026] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances. In addition, in the description of this application, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0027] Currently, polymerization reaction temperature control primarily relies on the traditional PID (proportional-integral-derivative) control method. PID control achieves temperature feedback control by adjusting three parameters: proportional, integral, and differential. In practice, the PID controller calculates a deviation signal based on real-time data collected by the temperature sensor and adjusts the heating or cooling system to maintain a stable temperature within the reactor.

[0028] The applicants of this application recognize that PID control, while widely used in industry, has significant limitations when dealing with complex dynamic systems. PID control struggles to adapt to rapid temperature changes within reactors, resulting in large temperature fluctuations, typically up to 3°C. These large temperature fluctuations can cause a broadening of the polymer molecular weight distribution (PDI>2.5), thereby reducing product performance.

[0029] In order to solve the above problems, the present application provides a polymerization reaction temperature optimization method and device based on generative AI to solve the problems existing in the above-mentioned related technical problems. In an embodiment of the present application, on the one hand, by inputting multi-source data collected and pre-processed in real time into a pre-trained generative AI model, the model can output the temperature change trajectory within a preset time period in the future. The prediction results provide important input information for the reinforcement learning strategy, so that the system can understand the temperature change trend in advance. The reinforcement learning strategy can more effectively calculate the optimal opening of the thermal oil valve, thereby achieving more accurate temperature control, helping to reduce temperature fluctuations and improve product quality. On the other hand, the reinforcement learning strategy can not only respond to the prediction of the generative AI model, but also realize closed-loop control based on real-time feedback. This dynamic adjustment mechanism significantly improves the flexibility of temperature control, helps to maintain the temperature stability in the reactor, reduce energy consumption, and improve production efficiency. The following exemplary embodiments are used for detailed description.

[0030] The following will be combined with the Figure 1 -Attached Figure 7This article details the generative AI-based polymerization reaction temperature optimization method provided in the embodiments of this application. This method can be implemented using a computer program and run on a generative AI-based polymerization reaction temperature optimization device based on a von Neumann architecture. This computer program can be integrated into an application or run as a standalone tool application.

[0031] See Figure 1 , provides a flow chart of a polymerization reaction temperature optimization method based on generative AI for an embodiment of the present application, which is applied to the server. Figure 1 As shown, the method of the embodiment of the present application includes the following steps: S101, collecting and preprocessing multi-source data in a reactor used for polymerization reaction in real time through sensors to obtain real-time status information in the reactor; A sensor is a detection device that senses the information being measured and, according to a specific pattern, converts that information into an electrical signal or other desired form of information output. A polymerization reaction is a type of reaction in chemistry in which monomers (small molecules) combine to form polymers (large molecules). This reaction is crucial in the manufacture of materials such as plastics, rubber, and fibers. Multi-source data refers to data from different sources or types. In this application, multi-source data includes temperature readings, pressure readings, stirring speed, and so on.

[0032] In some embodiments of the present application, a variety of sensors are installed on the polymerization reactor. These sensors can detect different parameters, such as temperature and pressure. The sensors monitor various parameters in the reactor in real time and convert these data into electrical signals to form a raw data stream. The collected raw data stream is transmitted to the server via wired or wireless means. The server preprocesses the raw data stream, and the preprocessing methods include filtering, normalization, missing value processing, and outlier detection. Finally, the preprocessed data is integrated into a unified data set to obtain real-time status information in the reactor.

[0033] Filtering removes noise and interference from the data. Normalization scales the data to a uniform range, such as 0 to 1, for easier processing and comparison. Missing value handling interpolates or uses other methods to fill in missing data points. Outlier detection identifies and handles outliers to prevent them from affecting subsequent analysis.

[0034] The real-time status information in the reactor is shown in Table 1.

[0035] Table 1

[0036] S102, inputting the real-time status information into a pre-trained generative AI model to output the temperature change trajectory within a preset time period in the future; The generative AI model is an artificial intelligence model that can be trained to predict future temperature changes. This model uses real-time status information to predict temperatures over a period of time. The future preset time period is the time period for which the model predicts temperature changes. This time period is pre-set, such as the next 10 minutes. The temperature trajectory is the predicted path of the reactor temperature over time within the preset time period. This trajectory allows the server to understand future temperature trends.

[0037] For example Figure 2 ,The pre-trained generative AI model includes an input layer, a hidden layer, and an output layer.

[0038] In one possible implementation, the real-time status information is input into a pre-trained generative AI model, and the specific process of outputting the temperature change trajectory within a preset time period in the future includes: the input layer obtains the real-time status information; the hidden layer captures the long-term and short-term dependencies in the time series of the real-time status information, and outputs the long-term and short-term dependencies in the time series corresponding to the real-time status information; the output layer predicts the temperature change trajectory within a preset time period in the future corresponding to the real-time status information based on the long-term and short-term dependencies in the time series.

[0039] Specifically, the input layer receives real-time state information from sensors, which may include temperature, pressure, stirring speed, and other information. It formats this state information into a form that the neural network can process, typically a vector where each element represents a specific state parameter. The hidden layer uses a neural network algorithm (such as a long short-term memory (LSTM) network) to analyze the input data. LSTM is particularly well-suited for processing and predicting time series data because it can capture long-term dependencies. The hidden layer learns the time series patterns in the input data, identifying which historical data points influence the current and future states, thereby capturing long-term and short-term dependencies. The hidden layer transmits these analysis results, which encode the key dependencies in the time series, to the output layer. Based on the time series dependencies provided by the hidden layer and the historical state information, the output layer predicts the future state. The output layer generates a temperature trajectory for a preset future time period. The temperature trajectory is a continuous sequence of values ​​that predicts the temperature at each point in time.

[0040] S103, dynamically adjusting the opening parameter of the thermal oil valve using a preset reinforcement learning strategy based on the temperature change trajectory to generate a thermal oil valve opening instruction; The preset reinforcement learning strategy is a pre-designed and trained model based on the principles of reinforcement learning. Reinforcement learning is a machine learning method that learns to make optimal decisions through interaction with the environment. In this application, the strategy is used to determine how to adjust the opening of the thermal oil valve. The opening parameter of the thermal oil valve guides the degree of opening of the thermal oil valve and is expressed as a percentage. The opening parameter affects the flow rate of the thermal oil, which in turn affects the temperature within the reactor.

[0041] In some embodiments of the present application, according to the temperature change trajectory, a preset reinforcement learning strategy is used to dynamically adjust the opening parameters of the thermal oil valve, and the specific process of generating the thermal oil valve opening instruction includes: evenly dividing the temperature change trajectory according to multiple moments in a preset time period to obtain the temperature prediction value at each moment; splicing the temperature prediction value at each moment, the current temperature of the reactor, and the set value into a decision vector; inputting the decision vector into the preset reinforcement learning strategy, and outputting the score values ​​corresponding to different opening parameters of the thermal oil valve; replacing the opening parameter of the thermal oil valve with the target opening parameter with the highest score value to obtain the actual opening parameter of the thermal oil valve, and the target opening parameter with the highest score value is the valve opening parameter that can make the temperature closest to the set value; encoding the actual opening parameter according to the industrial protocol, adding a timestamp and a check code, and encapsulating it into a thermal oil valve opening instruction.

[0042] For example, a generative AI model predicts the temperature trajectory over the next 10 minutes, with the following predictions (one per minute): 1 minute later: 98°C, 2 minutes later: 99°C, 3 minutes later: 101°C, ..., 10 minutes later: 100°C. These predictions are concatenated with the reactor's current temperature (e.g., 97°C) and setpoint (100°C) to form a decision vector: decision vector = [97°C (current temperature), 98°C (prediction after 1 minute), 99°C (prediction after 2 minutes), ..., 100°C (prediction after 10 minutes), 100°C (setpoint)]. This decision vector is then fed into a pre-set reinforcement learning strategy (e.g., a trained Q-network). The Q-network outputs a score (Q-value) for each possible thermal oil valve opening parameter: 30% opening: Q-value = 3, 40% opening: Q-value = 5, and 50% opening: Q-value = 7. The target opening parameter with the highest score, a 50% thermal oil valve opening, is selected because this brings the temperature closest to the setpoint of 100°C. The actual opening parameter (50%) is encoded according to the industrial protocol, timestamped, and verified, and packaged as a thermal oil valve opening command: command = {opening: 50%, timestamp: 2023-12-18T12:00:00Z, verification code: 123456}.

[0043] S104: Sending a thermal oil valve opening instruction to a control system for controlling the thermal oil valve to optimize the polymerization reaction temperature.

[0044] The thermal oil valve opening command refers to a specific instruction calculated using a reinforcement learning strategy for adjusting the thermal oil valve opening. This instruction contains the desired valve opening value. A control system is a system used to manage and regulate various parameters (such as temperature, pressure, and flow) in an industrial process. In this application, the control system is responsible for receiving the thermal oil valve opening command and performing corresponding actions to adjust the valve. The opening refers to the degree of opening of the valve. A larger opening increases the fluid flow through the valve, while a smaller opening reduces the flow.

[0045] In some embodiments of the present application, encoded instructions are sent from a computer to a control system via a network (possibly Ethernet, wireless, or other industrial communication network). Upon receiving the instructions, the control system parses the instructions and extracts the opening parameters of the thermal oil valve. Based on the parsed opening parameters, the control system drives an actuator (such as an electric actuator or a pneumatic actuator) to adjust the thermal oil valve to the desired opening. The control system monitors the valve position feedback signal to ensure that the valve has been correctly adjusted to the commanded opening.

[0046] For example Figure 3 As shown, Figure 3 This is the administrator's backend monitoring UI interface. This console displays the reactor's real-time status. A pre-trained generative AI model predicts the temperature for the next 10 minutes, and based on this predicted temperature, a reinforcement learning strategy generates values. It shows that the optimal valve opening is 50% when the Q value is 8.7. Therefore, the valve instruction encapsulates 50%, along with a timestamp and checksum. Administrators can refresh the data in this UI by clicking Refresh Data to obtain the latest parameter values.

[0047] For example Figure 4 As shown, server 110 uses sensors to collect and preprocess multi-source data from a polymerization reactor in real time, obtaining real-time status information within the reactor. Server 110 then inputs this real-time status information into a pre-trained generative AI model, which outputs a temperature trajectory for a preset future time period. Based on this temperature trajectory, server 110 uses a preset reinforcement learning strategy to dynamically adjust the opening parameters of the thermal oil valve, generating a thermal oil valve opening instruction. Server 110 then sends this thermal oil valve opening instruction to the control system used to control the thermal oil valve to optimize the polymerization reaction temperature. Server 110 transmits the real-time status information, the temperature trajectory for the preset future time period, and the thermal oil valve opening instruction to client 120.

[0048] In an embodiment of the present application, on the one hand, by inputting multi-source data collected and pre-processed in real time into a pre-trained generative AI model, the model can output the temperature change trajectory within a preset time period in the future. The prediction results provide important input information for the reinforcement learning strategy, allowing the system to understand the temperature change trend in advance. The reinforcement learning strategy can more effectively calculate the optimal opening of the thermal oil valve, thereby achieving more precise temperature control, helping to reduce temperature fluctuations and improve product quality. On the other hand, the reinforcement learning strategy can not only respond to the predictions of the generative AI model, but also achieve closed-loop control based on real-time feedback. This dynamic adjustment mechanism significantly improves the flexibility of temperature control, helps to maintain temperature stability in the reactor, reduce energy consumption, and improve production efficiency.

[0049] See Figure 5 , provides a flow chart of a model training method for a generative AI model according to an embodiment of the present application. Figure 5 As shown, the method of the embodiment of the present application may include the following steps: S201, acquiring and preprocessing historical multi-source data collected in advance for the reactor during a historical time period to obtain historical state information and historical real temperature at each historical moment in the reactor; The historical time period refers to a specific time range in the past, such as the past year, month, or day. Data collected within this time range will be used for model training. Historical multi-source data refers to historical data on the reactor's operating status collected from multiple sources. This data includes temperature, pressure, stirring speed, material concentration, etc. Historical actual temperature refers to the actual temperature reached by the reactor during the historical time period. This data is measured and recorded by sensors.

[0050] S202, using the historical real temperature at each historical moment as a label of the historical state information at each historical moment to perform data labeling to obtain a training data set; In some embodiments of the present application, it is necessary to collect multi-source data of the reactor within a historical time period, which include temperature, pressure, stirring speed, material concentration, etc. The collected data is preprocessed, including noise removal, missing value filling, outlier processing, etc., to ensure the quality of the data. The real temperature value of each historical moment is used as a label and paired with the corresponding historical state information. This step is to associate the historical real temperature with other state information to form a sample for training. The labeled data is organized into a training data set, which usually includes a feature matrix and a label vector. Each row of the feature matrix corresponds to a sample (state information of a historical moment), and the label vector contains the corresponding real temperature value.

[0051] S203, constructing a loss function using a mean square error expression; Among them, the expression of the loss function is:

[0052] in, is the model loss value, is the total number of training data in the training dataset, The training data set training data, when predicting the future preset time period containing 10 time steps, is the first time step within 10 time steps time steps, For the The training data in The predicted temperature for each time step, It is The training data in The actual temperature of each time step.

[0053] S204, using loss function and neural network algorithm to create a generative AI model; In some embodiments of the present application, the specific process of creating a generative AI model using a loss function and a neural network algorithm includes: constructing an input layer for receiving a training data set, where the dimension parameter of the input layer is equal to the number of data categories of the historical multi-source data; using a neural network algorithm to construct a hidden layer for capturing the long-term and short-term dependencies in the time series of the training data set; using a neural network algorithm to construct an output layer for predicting the temperature change trajectory within a preset time period in the future; and connecting the input layer, hidden layer, output layer, and loss function in sequence to form a generative AI model.

[0054] Among them, the neural network algorithm is a multi-layer long short-term memory network structure, each layer of the long short-term memory network structure includes 128 neurons, and the activation function is the ReLU function.

[0055] S205: Perform machine learning on the generative AI model based on the training data set to obtain a pre-trained generative AI model.

[0056] In some embodiments of the present application, the specific process of performing machine learning on the generative AI model based on the training data set to obtain a pre-trained generative AI model includes: inputting the training data set into the generative AI model and outputting the model loss value; when the model loss value reaches the minimum, generating a pre-trained generative AI model; or, when the model loss value does not reach the minimum, backpropagating the model loss value to update the model parameters of the generative AI model, and continuing to execute the step of inputting the training data set into the generative AI model until the model loss value reaches the minimum.

[0057] Among them, the generative AI model includes input layer, hidden layer, output layer and loss function.

[0058] In some embodiments of the present application, a training data set is input into a generative AI model, and a specific process of outputting a model loss value includes: the input layer obtains each historical state information in the training data set; the hidden layer captures the long-term and short-term dependencies in the time series of each historical state information in the training data set, and outputs the long-term and short-term dependencies in the time series corresponding to each historical state information; the output layer predicts the first temperature change trajectory within a future preset time period corresponding to each historical state information based on the long-term and short-term dependencies in the time series corresponding to each historical state information; the first temperature change trajectory within a future preset time period corresponding to each historical state information and the actual temperature marked by each historical state information are input into the loss function for calculation to obtain the model loss value.

[0059] For example Figure 6 As shown, Figure 6 This is a UI interface diagram of the generative AI model training process, which schematically displays the relevant training data set, network structure, and training process parameters.

[0060] In an embodiment of the present application, on the one hand, by inputting multi-source data collected and pre-processed in real time into a pre-trained generative AI model, the model can output the temperature change trajectory within a preset time period in the future. The prediction results provide important input information for the reinforcement learning strategy, allowing the system to understand the temperature change trend in advance. The reinforcement learning strategy can more effectively calculate the optimal opening of the thermal oil valve, thereby achieving more precise temperature control, helping to reduce temperature fluctuations and improve product quality. On the other hand, the reinforcement learning strategy can not only respond to the predictions of the generative AI model, but also achieve closed-loop control based on real-time feedback. This dynamic adjustment mechanism significantly improves the flexibility of temperature control, helps to maintain temperature stability in the reactor, reduce energy consumption, and improve production efficiency.

[0061] See Figure 7 , which is a flow chart of a method for generating a preset reinforcement learning strategy provided in the embodiment of the present application. Figure 7 As shown, the method of the embodiment of the present application may include the following steps: S301: Define the real-time temperature, monomer conversion rate, and stirring power in the reactor as state variables of the environment, define the different opening adjustments of the thermal oil valve as executable action variables, and define a reward function to minimize the cumulative time that the temperature deviates from the set value; Among them, the function expression of the reward function is:

[0062] Among them, R is the reward value, T is the preset time step, The first time steps, is the actual temperature, To set the temperature.

[0063] S302, encapsulating state variables, action variables, and reward functions into an interactive sequential decision-making environment; S303, determining the target state variable at the target moment from the temporal decision environment, and randomly selecting the target action variable; S304, executing the target action variable in combination with the reinforcement learning algorithm, obtaining a new state variable and a target reward value for a preset time step after executing the target action variable; S305, storing the target state variable, target action variable, new state variable, and target reward value as a four-tuple, and continuing to perform the step of determining the target state variable at the target moment from the temporal decision environment until the number of four-tuples reaches a preset number, thereby obtaining a plurality of experience samples; S306, generating a preset reinforcement learning strategy based on multiple experience samples; In some embodiments of the present application, the specific process of generating a preset reinforcement learning strategy based on multiple experience samples includes: using a reinforcement learning algorithm to construct a reinforcement learning component, the reinforcement learning component includes a Q network and a target network, the Q network is used to estimate the value of each state variable-action variable pair, and the target network is used to provide an update target, which is initially the same as the Q network parameters; inputting each experience sample into the Q network and outputting the predicted Q value of each experience sample; for each experience sample, using the Q value calculation function of the target network to calculate the target Q value; calculating the mean square error loss between the predicted Q value of each experience sample and the target Q value of each experience sample; when the mean square error loss has not reached the minimum, using the mean square error loss to update the network parameters of the Q network, and continuing to execute the step of inputting each experience sample into the Q network until the mean square error loss reaches the minimum; copying the network parameters of the Q network to the target network to obtain the preset reinforcement learning strategy.

[0064] Among them, the function expression of the Q value calculation function of the target network is:

[0065] in, is the target Q value, is the reward value in each experience sample, is the discount factor, is the new state variable, is the action variable taken under the new state variable, Is in state Take action variables The expected return, is to select one from all action variables such that The action with the largest value .

[0066] In an embodiment of the present application, on the one hand, by inputting multi-source data collected and pre-processed in real time into a pre-trained generative AI model, the model can output the temperature change trajectory within a preset time period in the future. The prediction results provide important input information for the reinforcement learning strategy, allowing the system to understand the temperature change trend in advance. The reinforcement learning strategy can more effectively calculate the optimal opening of the thermal oil valve, thereby achieving more precise temperature control, helping to reduce temperature fluctuations and improve product quality. On the other hand, the reinforcement learning strategy can not only respond to the predictions of the generative AI model, but also achieve closed-loop control based on real-time feedback. This dynamic adjustment mechanism significantly improves the flexibility of temperature control, helps to maintain temperature stability in the reactor, reduce energy consumption, and improve production efficiency.

[0067] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0068] See Figure 8 , which shows a schematic diagram of the structure of a polymerization reaction temperature optimization device based on generative AI, provided by an exemplary embodiment of the present application. This polymerization reaction temperature optimization device based on generative AI can be implemented as all or part of an electronic device through software, hardware, or a combination of both. The device 1 includes a data preprocessing module 10, a temperature change trajectory output module 20, an instruction generation module 30, and an instruction sending module 40.

[0069] The data preprocessing module 10 is used to collect and preprocess multi-source data in the reactor used for the polymerization reaction in real time through sensors to obtain real-time status information in the reactor; The temperature change trajectory output module 20 is used to input real-time status information into a pre-trained generative AI model and output the temperature change trajectory within a preset time period in the future; The instruction generation module 30 is used to dynamically adjust the opening parameter of the thermal oil valve according to the temperature change trajectory using a preset reinforcement learning strategy to generate a thermal oil valve opening instruction; The instruction sending module 40 is used to send a thermal oil valve opening instruction to a control system for controlling the thermal oil valve, so as to optimize the polymerization reaction temperature.

[0070] It should be noted that the polymerization reaction temperature optimization device based on generative AI provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when executing the polymerization reaction temperature optimization method based on generative AI. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the polymerization reaction temperature optimization device based on generative AI provided in the above embodiment and the polymerization reaction temperature optimization method based on generative AI belong to the same concept. The implementation process thereof is detailed in the method embodiment and will not be repeated here.

[0071] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0072] In an embodiment of the present application, on the one hand, by inputting multi-source data collected and pre-processed in real time into a pre-trained generative AI model, the model can output the temperature change trajectory within a preset time period in the future. The prediction results provide important input information for the reinforcement learning strategy, allowing the system to understand the temperature change trend in advance. The reinforcement learning strategy can more effectively calculate the optimal opening of the thermal oil valve, thereby achieving more precise temperature control, helping to reduce temperature fluctuations and improve product quality. On the other hand, the reinforcement learning strategy can not only respond to the predictions of the generative AI model, but also achieve closed-loop control based on real-time feedback. This dynamic adjustment mechanism significantly improves the flexibility of temperature control, helps to maintain temperature stability in the reactor, reduce energy consumption, and improve production efficiency.

[0073] The present application also provides a computer-readable medium having program instructions stored thereon, which, when executed by a processor, implements the polymerization reaction temperature optimization method based on generative AI provided in the above-mentioned various method embodiments.

[0074] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the polymerization reaction temperature optimization method based on generative AI according to each of the above-mentioned method embodiments.

[0075] See Figure 9 , is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 9 As shown, the electronic device 1000 may include: at least one processor 1001 , at least one network interface 1004 , a user interface 1003 , a memory 1005 , and at least one communication bus 1002 .

[0076] The communication bus 1002 is used to implement the connection and communication between these components.

[0077] The user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface.

[0078] The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0079] The processor 1001 may include one or more processing cores. The processor 1001 utilizes various interfaces and circuits to connect various components within the electronic device 1000. It executes instructions, programs, code sets, or instruction sets stored in the memory 1005, and accesses data stored in the memory 1005 to perform various functions and process data within the electronic device 1000. Optionally, the processor 1001 may be implemented in hardware using at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 1001 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the processor 1001 and implemented on a separate chip.

[0080] Among them, the memory 1005 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 1005 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 1005 may also be optionally at least one storage system located away from the aforementioned processor 1001. As Figure 9 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a polymerization reaction temperature optimization application based on generative AI.

[0081] exist Figure 9 In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an input interface for the user and obtain user input data; and the processor 1001 can be used to call the generative AI-based polymerization reaction temperature optimization application stored in the memory 1005 and specifically perform the following operations: The sensors collect and pre-process multi-source data in the reactor used for polymerization reaction in real time to obtain real-time status information in the reactor; Input real-time status information into a pre-trained generative AI model to output the temperature change trajectory within a preset time period in the future; Based on the temperature change trajectory, the preset reinforcement learning strategy is used to dynamically adjust the opening parameters of the thermal oil valve and generate the thermal oil valve opening instruction; Send thermal oil valve opening instructions to the control system used to control the thermal oil valve to optimize the polymerization reaction temperature.

[0082] In one embodiment, when executing the generation of a pre-trained generative AI model, the processor 1001 specifically performs the following operations: Acquire and pre-process historical multi-source data collected in advance for the reactor during a historical time period to obtain historical status information and historical real temperature at each historical moment in the reactor; The historical real temperature at each historical moment is used as the label of the historical status information at each historical moment to annotate the data and obtain the training data set; Use the mean square error expression to construct the loss function; Use loss functions and neural network algorithms to create generative AI models; Based on the training data set, machine learning is performed on the generative AI model to obtain a pre-trained generative AI model.

[0083] In one embodiment, when processor 1001 creates a generative AI model using a loss function and a neural network algorithm, it specifically performs the following operations: Construct an input layer for receiving the training dataset, where the dimension parameter of the input layer is equal to the number of data categories of the historical multi-source data; A neural network algorithm is used to construct a hidden layer to capture the long-term and short-term dependencies in the time series of the training dataset; A neural network algorithm is used to construct an output layer for predicting the temperature change trajectory within a preset time period in the future; The input layer, hidden layer, output layer, and loss function are connected in sequence to form a generative AI model.

[0084] In one embodiment, when the processor 1001 performs machine learning on the generative AI model based on the training data set to obtain a pre-trained generative AI model, the processor 1001 specifically performs the following operations: Input the training dataset into the generative AI model and output the model loss value; When the model loss value reaches the minimum, a pre-trained generative AI model is generated; alternatively, when the model loss value does not reach the minimum, the model loss value is back-propagated to update the model parameters of the generative AI model, and the step of inputting the training data set into the generative AI model is continued until the model loss value reaches the minimum.

[0085] In one embodiment, when the processor 1001 inputs the training data set into the generative AI model and outputs the model loss value, it specifically performs the following operations: The input layer obtains the historical state information in the training data set; The hidden layer captures the long-term and short-term dependencies in the time series of each historical state information in the training data set, and outputs the long-term and short-term dependencies in the time series corresponding to each historical state information; The output layer predicts the first temperature change trajectory within a preset future time period corresponding to each historical state information based on the long-term and short-term dependencies in the time series corresponding to each historical state information; The first temperature change trajectory within the future preset time period corresponding to each historical state information and the actual temperature marked by each historical state information are input into the loss function for calculation to obtain the model loss value.

[0086] In one embodiment, when the processor 1001 dynamically adjusts the opening parameter of the thermal oil valve according to the temperature change trajectory using a preset reinforcement learning strategy and generates a thermal oil valve opening instruction, the processor 1001 specifically performs the following operations: Divide the temperature change trajectory into multiple moments in a preset time period and obtain the temperature prediction value at each moment; The temperature prediction value at each moment, the current temperature of the reactor, and the set value are spliced ​​into a decision vector; The decision vector is input into the preset reinforcement learning strategy, and the score values ​​corresponding to different opening parameters of the thermal oil valve are output; The opening parameter of the thermal oil valve is replaced by the target opening parameter with the highest score value to obtain the actual opening parameter of the thermal oil valve. The target opening parameter with the highest score value is the valve opening parameter that can make the temperature closest to the set value; The actual opening parameters are encoded according to the industrial protocol, added with a timestamp and a check code, and encapsulated as a thermal oil valve opening instruction.

[0087] In one embodiment, when executing the generation of the preset reinforcement learning strategy, the processor 1001 specifically performs the following operations: The real-time temperature, monomer conversion rate, and stirring power in the reactor are defined as state variables of the environment. The different opening adjustments of the thermal oil valve are defined as executable action variables. A reward function is defined to minimize the cumulative time that the temperature deviates from the set value. Encapsulate state variables, action variables, and reward functions into an interactive sequential decision-making environment; Determine the target state variable at the target time from the temporal decision environment and randomly select the target action variable; Execute the target action variable in conjunction with the reinforcement learning algorithm to obtain the new state variable and target reward value of the preset time step after executing the target action variable; The target state variable, the target action variable, the new state variable, and the target reward value are stored as a four-tuple, and the step of determining the target state variable at the target moment from the sequential decision environment is continued until the number of four-tuples reaches a preset number, thereby obtaining a plurality of experience samples; Generate preset reinforcement learning strategies based on multiple experience samples.

[0088] In one embodiment, when the processor 1001 generates a preset reinforcement learning strategy based on multiple experience samples, it specifically performs the following operations: A reinforcement learning algorithm is used to construct a reinforcement learning component, which includes a Q network and a target network. The Q network is used to estimate the value of each state variable-action variable pair, and the target network is used to provide updated targets. Initially, the parameters are the same as those of the Q network. Input each experience sample into the Q network and output the predicted Q value of each experience sample; For each experience sample, the target Q value is calculated using the Q value calculation function of the target network; Calculate the mean squared error loss between the predicted Q value of each experience sample and the target Q value of each experience sample; When the mean square error loss has not reached the minimum, the network parameters of the Q network are updated using the mean square error loss, and the step of inputting each experience sample into the Q network is continued until the mean square error loss reaches the minimum; The network parameters of the Q network are copied to the target network to obtain the preset reinforcement learning strategy.

[0089] In an embodiment of the present application, on the one hand, by inputting multi-source data collected and pre-processed in real time into a pre-trained generative AI model, the model can output the temperature change trajectory within a preset time period in the future. The prediction results provide important input information for the reinforcement learning strategy, allowing the system to understand the temperature change trend in advance. The reinforcement learning strategy can more effectively calculate the optimal opening of the thermal oil valve, thereby achieving more precise temperature control, helping to reduce temperature fluctuations and improve product quality. On the other hand, the reinforcement learning strategy can not only respond to the predictions of the generative AI model, but also achieve closed-loop control based on real-time feedback. This dynamic adjustment mechanism significantly improves the flexibility of temperature control, helps to maintain temperature stability in the reactor, reduce energy consumption, and improve production efficiency.

[0090] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program for optimizing polymerization reaction temperature based on generative AI can be stored in a computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. The storage medium for the program for optimizing polymerization reaction temperature based on generative AI can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.

[0091] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A polymerization reaction temperature optimization method based on generative AI, characterized in that: Applied to the server, the method includes: The multi-source data in the reactor used for the polymerization reaction are collected and pre-processed in real time by sensors to obtain real-time status information in the reactor; Input the real-time status information into a pre-trained generative AI model to output the temperature change trajectory within a preset time period in the future; According to the temperature change trajectory, a preset reinforcement learning strategy is used to dynamically adjust the opening parameter of the thermal oil valve to generate a thermal oil valve opening instruction; The thermal oil valve opening instruction is sent to a control system for controlling the thermal oil valve to optimize the polymerization reaction temperature.

2. The method according to claim 1, characterized in that Follow these steps to generate a pre-trained generative AI model, including: Acquire and pre-process historical multi-source data collected in advance for the reactor within a historical time period to obtain historical state information and historical real temperature at each historical moment in the reactor; The historical real temperature at each historical moment is used as a label of the historical state information at each historical moment to perform data annotation to obtain a training data set; Use the mean square error expression to construct the loss function; Using the loss function and neural network algorithm to create a generative AI model; Based on the training data set, machine learning is performed on the generative AI model to obtain a pre-trained generative AI model.

3. The method according to claim 2, characterized in that The use of the loss function and the neural network algorithm to create a generative AI model includes: Constructing an input layer for receiving the training data set, wherein a dimension parameter of the input layer is equal to the number of data categories of the historical multi-source data; A neural network algorithm is used to construct a hidden layer for capturing the long-term and short-term dependencies in the time series of the training data set; A neural network algorithm is used to construct an output layer for predicting the temperature change trajectory within a preset time period in the future; The input layer, the hidden layer, the output layer, and the loss function are connected in sequence to form a generative AI model.

4. The method according to claim 2, characterized in that The performing machine learning on the generative AI model according to the training data set to obtain a pre-trained generative AI model includes: Input the training data set into the generative AI model and output the model loss value; When the model loss value reaches a minimum, generating a pre-trained generative AI model; or, when the model loss value does not reach a minimum, backpropagating the model loss value to update the model parameters of the generative AI model, and continuing to perform the step of inputting the training data set into the generative AI model until the model loss value reaches a minimum; Among them, the expression of the loss function is: in, is the model loss value, is the total number of training data in the training dataset, is the first training data, when predicting the future preset time period containing 10 time steps, is the first time step within 10 time steps time steps, For the The training data in The predicted temperature for each time step, It is The training data in The actual temperature of each time step.

5. The method according to claim 4, characterized in that The generative AI model includes an input layer, a hidden layer, an output layer, and a loss function; Inputting the training data set into the generative AI model and outputting a model loss value includes: The input layer obtains each historical state information in the training data set; The hidden layer captures the long-term and short-term dependencies in the time series of each historical state information in the training data set, and outputs the long-term and short-term dependencies in the time series corresponding to each historical state information; The output layer predicts the first temperature change trajectory within a future preset time period corresponding to each historical state information based on the long-term and short-term dependencies in the time series corresponding to each historical state information; The first temperature change trajectory within a future preset time period corresponding to each historical state information and the actual temperature marked by each historical state information are input into the loss function for calculation to obtain a model loss value.

6. The method according to claim 1, characterized in that The method of dynamically adjusting the opening parameter of the thermal oil valve using a preset reinforcement learning strategy according to the temperature change trajectory to generate a thermal oil valve opening instruction includes: Dividing the temperature change trajectory into multiple moments in a preset time period to obtain a temperature prediction value at each moment; The temperature prediction value at each moment, the current temperature of the reactor, and the set value are combined into a decision vector; Inputting the decision vector into the preset reinforcement learning strategy, and outputting score values ​​corresponding to different opening parameters of the thermal oil valve; Replacing the opening parameter of the thermal oil valve with the target opening parameter with the highest score value to obtain the actual opening parameter of the thermal oil valve, wherein the target opening parameter with the highest score value is the valve opening parameter that can make the temperature closest to the set value; The actual opening parameter is encoded according to the industrial protocol, added with a timestamp and a check code, and encapsulated as a thermal oil valve opening instruction.

7. The method according to claim 1, characterized in that Follow these steps to generate a preset reinforcement learning policy, including: The real-time temperature, monomer conversion rate, and stirring power in the reactor are defined as state variables of the environment, the different opening adjustments of the thermal oil valve are defined as executable action variables, and a reward function is defined to minimize the cumulative time that the temperature deviates from the set value; Encapsulating the state variables, the action variables, and the reward function into an interactive sequential decision environment; determining a target state variable at a target time from the sequential decision environment and randomly selecting a target action variable; Executing the target action variable in combination with a reinforcement learning algorithm to obtain a new state variable and a target reward value for a preset time step after executing the target action variable; The target state variable, the target action variable, the new state variable, and the target reward value are stored as a four-tuple, and the step of determining the target state variable at the target moment from the temporal decision environment is continued until the number of the four-tuples reaches a preset number, thereby obtaining a plurality of experience samples; Based on the multiple experience samples, a preset reinforcement learning strategy is generated; wherein the function expression of the reward function is: Among them, R is the reward value, T is the preset time step, The first time steps, is the actual temperature, To set the temperature.

8. The method according to claim 7, characterized in that Generating a preset reinforcement learning strategy based on the multiple experience samples includes: A reinforcement learning component is constructed using a reinforcement learning algorithm. The reinforcement learning component includes a Q network and a target network. The Q network is used to estimate the value of each state variable-action variable pair. The target network is used to provide an update target and initially has the same parameters as the Q network. Input each experience sample into the Q network and output the predicted Q value of each experience sample; For each experience sample, the target Q value is calculated using the Q value calculation function of the target network; Calculating the mean square error loss between the predicted Q value of each experience sample and the target Q value of each experience sample; When the mean square error loss has not reached a minimum, updating the network parameters of the Q network using the mean square error loss, and continuing to perform the step of inputting each experience sample into the Q network until the mean square error loss reaches a minimum; The network parameters of the Q network are copied to the target network to obtain a preset reinforcement learning strategy.

9. The method according to claim 8, characterized in that The function expression of the Q value calculation function of the target network is: in, is the target Q value, is the reward value in each experience sample, is the discount factor, is the new state variable, is the action variable taken under the new state variable, Is in state Take action variables The expected return, is to select one from all action variables such that The action with the largest value .

10. A polymerization reaction temperature optimization device based on generative AI, characterized in that: The device comprises: A data preprocessing module is used to collect and preprocess multi-source data in the reactor used for polymerization reaction in real time through sensors to obtain real-time status information in the reactor; A temperature change trajectory output module is used to input the real-time status information into a pre-trained generative AI model and output the temperature change trajectory within a preset time period in the future; An instruction generation module is used to dynamically adjust the opening parameter of the thermal oil valve according to the temperature change trajectory using a preset reinforcement learning strategy to generate a thermal oil valve opening instruction; The instruction sending module is used to send the thermal oil valve opening instruction to a control system used to control the thermal oil valve to optimize the polymerization reaction temperature.

Citation Information

Patent Citations

  • Reaction kettle temperature control method based on fuzzy control

    CN116700393A

  • Deep learning three-dimensional ocean temperature forecasting method based on physical guidance

    CN117706657A

  • Electric vehicle charging station operation management method and system based on virtual power plant

    CN119090549A

  • Multi-model-based intermittent reaction kettle intelligent control method and system

    CN119356092A

  • Method and apparatus for constructing dispatching model of integrated energy system, medium, and electronic device

    WO2022160705A1

Cited By

  • Accurate parking control method for polymerization reaction heating temperature based on generative AI

    CN121764256A