A chemical reaction dynamic optimization method and system based on reinforcement learning and online spectral feedback

By combining online spectral sensing and digital twin dynamic simulation with reinforcement learning to create a closed-loop optimization control system, the problem of balancing control precision and safety in chemical reaction processes has been solved. This enables real-time monitoring and optimized control of chemical reaction processes, thereby improving product yield and selectivity.

CN122219104APending Publication Date: 2026-06-16SHAOXING MIAOXIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAOXING MIAOXIN TECHNOLOGY CO LTD
Filing Date
2026-04-14
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing chemical reaction process control methods are unable to directly reflect the dynamic changes of key components within the reaction system, have limited control precision, result in large fluctuations in product yield and selectivity, and make it difficult to balance safety and production efficiency.

Method used

Online spectral sensing technology is used to acquire real-time molecular-level information. Combined with digital twin dynamic simulation and reinforcement learning autonomous decision-making, a closed-loop optimization control system is constructed. The system optimizes decisions through a reinforcement learning agent and combines safety constraints and online update mechanisms to achieve real-time monitoring and control of chemical reaction processes.

Benefits of technology

It significantly improves the observability of chemical reaction processes and the accuracy of control strategies, enhances the stability and robustness of reaction processes, reduces safety risks, and improves production efficiency and product yield.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122219104A_ABST
    Figure CN122219104A_ABST
Patent Text Reader

Abstract

The application relates to a chemical reaction dynamic optimization method and system based on reinforcement learning and online spectrum feedback. The method obtains real-time spectrum data of a reaction system through an online spectrum acquisition device, extracts spectrum characteristic variables representing a reaction process, and fuses the spectrum characteristic variables with process parameters such as temperature, pressure and feed flow to construct a composite state vector. Based on the state vector, a reinforcement learning intelligent agent outputs an optimized control action and acts on a reaction execution mechanism, simultaneously obtains state feedback and a reward signal, realizes online updating of the model and dynamic correction of a digital twin model, and corrects the control action in combination with a safety constraint mechanism. The system comprises a spectrum sensing module, a data fusion module, a digital twin module, a reinforcement learning decision module and a control execution module, and forms a closed-loop optimization control structure. The application has real-time sensing and autonomous optimization control on a chemical reaction process, and improves the effects of reaction yield, selectivity and safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent chemical engineering and process control, and in particular to a method and system for dynamic optimization of chemical reactions based on reinforcement learning and online spectral feedback. Background Technology

[0002] Chemical reaction processes are widely used in fine chemicals, pharmaceutical synthesis, and new material preparation. Batch or semi-batch reactions (such as nitration, hydrogenation, and polymerization) place high demands on the precision and safety of process control due to their complex reaction pathways, significant exothermic characteristics, and sensitive operating conditions. Traditional chemical production process control typically relies on feedback regulation of macroscopic variables such as temperature, pressure, and flow rate, combined with operational experience or preset control curves for process optimization. However, these methods struggle to directly reflect the dynamic changes of key components within the reaction system, exhibiting a "black box" characteristic that results in limited control precision and significant fluctuations in product yield and selectivity.

[0003] To improve process observability, existing technologies have gradually introduced online spectroscopic analysis methods, such as Raman spectroscopy, near-infrared spectroscopy, and mid-infrared spectroscopy, to achieve real-time monitoring of reactants, products, and intermediates by acquiring molecular vibrational information. However, in practical applications, online spectroscopic data is usually only used for qualitative analysis or simple quantitative regression, lacking deep integration with the control system and failing to fully realize its potential in process optimization.

[0004] On the other hand, while advanced control methods based on mechanistic or empirical models (such as model predictive control) can achieve a certain degree of optimization, the complexity of chemical reaction mechanisms and the susceptibility of kinetic parameters to fluctuations in raw materials and changes in catalyst state make it difficult to maintain model accuracy over the long term, leading to a decline in control performance. Furthermore, in the case of highly exothermic or high-risk reactions, conservative operating strategies are typically adopted to ensure safe operation, keeping reaction conditions far from the optimal range, thereby sacrificing production efficiency and economic benefits.

[0005] In recent years, reinforcement learning, as a data-driven adaptive optimization method, has shown potential advantages in the control of complex systems. However, its application in chemical production processes still faces challenges such as high training costs and safety risks during the exploration process, making direct deployment in actual production environments difficult. Furthermore, current technologies lack a systematic solution that organically combines online real-time spectral sensing, digital twin simulation, and reinforcement learning decision-making, making it impossible to construct a closed-loop control system that balances safety and optimization performance. Regarding the aforementioned technologies, the inventors believe that a technology is lacking that deeply integrates online spectral molecular-level real-time sensing information, digital twin dynamic extrapolation capabilities, and reinforcement learning autonomous decision-making mechanisms, while simultaneously embedding safety constraints to achieve closed-loop optimization control. Summary of the Invention

[0006] To address the technical challenge of deeply integrating online spectral molecular-level real-time sensing information, digital twin dynamic extrapolation capabilities, and reinforcement learning autonomous decision-making mechanisms, while simultaneously embedding safety constraints to achieve closed-loop optimization control, this application provides a method and system for dynamic optimization of chemical reactions based on reinforcement learning and online spectral feedback.

[0007] This application provides a dynamic optimization method for chemical reactions based on reinforcement learning and online spectral feedback, employing the following technical solution: Firstly, a dynamic optimization method for chemical reactions based on reinforcement learning and online spectral feedback includes the following steps: S1. The reaction system is monitored in real time by an online spectral acquisition device to obtain raw spectral data, and the raw spectral data is preprocessed and feature extracted to obtain at least one spectral feature variable characterizing the reaction process. S2. The spectral feature variables are fused with the process parameters in the reaction process to construct a composite state vector containing reaction state information; S3. Input the composite state vector into the reinforcement learning agent, and the reinforcement learning agent outputs the corresponding optimized control action based on the current state; S4. Apply the optimized control action to the reaction actuator to adjust the reaction process, and obtain the status feedback and reward signal after execution; S5. Based on the aforementioned state feedback and reward signals, the reinforcement learning agent is updated online, and the digital twin model is dynamically corrected simultaneously. The training process of the reinforcement learning agent includes fusing virtual interaction data generated by a digital twin model with actual reaction data for training.

[0008] By adopting the above technical solutions, the control of chemical reaction processes has been transformed from traditional macroscopic parameter-driven to "molecular-level perception-intelligent decision-making-dynamic optimization control." Through online spectral acquisition and feature extraction, the dynamic changes of key components and intermediates in the reaction system can be obtained in real time, overcoming the limitation of traditional control that cannot directly observe the internal state of the reaction, thus improving the observability and information integrity of the process from the source. On this basis, spectral features are integrated with process parameters such as temperature, pressure, and flow rate to construct a composite state vector, enabling the reinforcement learning agent to make decisions based on more comprehensive and refined state information, significantly improving the pertinence and accuracy of the control strategy. At the same time, by introducing a digital twin model and adopting a virtual-real fusion training mechanism, reinforcement learning can conduct efficient and safe strategy exploration in a virtual environment, and continuously correct the model and strategy by combining actual operating data. This reduces the risk of trial and error in the real reaction system and ensures the real-time adaptability of the model. In addition, the system continuously optimizes the strategy through an online update mechanism, enabling it to dynamically respond to disturbances such as raw material fluctuations and environmental changes, improving the stability and robustness of the reaction process.

[0009] Secondly, a dynamic optimization system for chemical reactions based on reinforcement learning and online spectral feedback includes: an online spectral sensing module, which is used to collect real-time spectral data of the reaction system; It includes a spectral processing and feature extraction module, which is used to perform dimensionality reduction and feature variable extraction on the spectral data; It includes a data fusion and state construction module, which is used to fuse spectral feature variables and process parameters to generate a composite state vector; It includes a digital twin module, which is used to construct a mechanism-data hybrid model of the reaction process and provide a virtual simulation environment; It includes a reinforcement learning decision module, which is used to output an optimized control strategy based on the composite state vector; It includes a control execution module, which is used to convert the control strategy into execution instructions and apply them to the reaction device; It includes an online learning and model update module, which is used to dynamically update the reinforcement learning model and digital twin model based on actual operating data; It includes a security constraint module, which is used to perform security verification and constraint correction on the control strategy.

[0010] By adopting the above technical solutions, a closed-loop optimization control system for chemical reactions integrating "real-time perception, intelligent decision-making, safe execution, and continuous learning" is constructed. The online spectral perception module and the spectral processing module work together to acquire and analyze information on key components and intermediates in the reaction system in real time at the molecular level, improving process observability. The data fusion and state construction module expresses spectral features and process parameters such as temperature, pressure, and flow rate into a composite state vector, providing high-quality input for subsequent intelligent decision-making. The digital twin module constructs a virtual environment that evolves synchronously with the real reaction process through mechanism and data fusion modeling, providing support for strategy optimization and risk prediction. The reinforcement learning decision-making module generates dynamic optimization control strategies based on the state information, realizing adaptive regulation of complex nonlinear reaction processes. The control execution module transforms the strategies into specific control commands and applies them to the field devices. At the same time, the safety constraint module verifies and corrects the control actions in real time to ensure that the operation process is always within the safety boundary. The online learning and model update module continuously optimizes the reinforcement learning model and the digital twin model using actual operating data, enabling the system to have the ability to adapt to raw material fluctuations and changes in the chemical reaction process.

[0011] Optionally, the spectral feature extraction in step S1 includes: Adaptive peak identification and tracking based on sliding window; Extract information on the peak position, peak intensity, and full width at half maximum (FWHM) variation of characteristic peaks; Alternatively, high-dimensional spectral data can be mapped to low-dimensional latent variables through principal component analysis (PCA) or partial least squares regression (PLS).

[0012] By adopting the above technical solutions, the high-dimensional and redundant original spectral data can be effectively reduced in dimensionality and key information can be extracted, enabling accurate characterization of changes in key components during the reaction process. At the same time, through adaptive peak identification and dynamic tracking, the real-time performance and robustness of spectral feature extraction are improved, making the obtained feature variables more stable and reliable, thereby providing high-quality input data for subsequent state construction and intelligent decision-making.

[0013] Optionally, the composite state vector includes: temperature, pressure, feed flow rate, spectral latent variable, the rate of change and integral term of the spectral latent variable, and reaction time information.

[0014] By adopting the above technical solution, macroscopic process parameters and spectral latent variables, their changing trends, and cumulative effects are integrated and characterized, enabling the reaction state to be comprehensively described from both instantaneous and dynamic evolutionary dimensions. This improves the completeness and sensitivity of the state description, provides more time-series-specific and predictive input information for reinforcement learning decision-making, and ultimately enhances the accuracy and stability of the control strategy.

[0015] Optionally, the reinforcement learning agent described in step S3 may employ a policy gradient algorithm or an actor-critic structure.

[0016] By adopting the above technical solutions, reinforcement learning agents can perform efficient policy optimization in continuous action space, improving their adaptability to complex nonlinear response processes. At the same time, through the synergistic mechanism of value evaluation and policy update, the learning stability and convergence efficiency are improved, thereby achieving precise adjustment and continuous optimization of response control strategies.

[0017] Optionally, the reinforcement learning agent includes: Upper-level decision network, which is used to generate control strategies based on long-term optimization objectives; A lower-level security execution network is used to modify the control strategy and output the final control action under the condition of satisfying security constraints; The lower-level secure execution network includes a secure projection layer, which is used to map control actions that do not meet the constraints to a preset secure boundary range.

[0018] By adopting the above technical solution, the hierarchical decoupling of optimization decision-making and safety constraints is achieved. The upper-level decision network focuses on learning the global optimal strategy, while the lower-level safety execution network performs real-time correction of control actions. At the same time, the safety projection layer maps actions that do not meet the constraints to the safety boundary, thereby improving control flexibility and optimization effect while ensuring the safety of the response.

[0019] Optionally, the digital twin model is a hybrid model combining a mechanistic model and a data-driven model, including: a reaction kinetic model based on mass and energy conservation; and a parameter correction model trained based on historical data. The parameters of the digital twin model are dynamically updated using an online Bayesian inference method.

[0020] By adopting the above technical solutions, the advantages of mechanistic models and data-driven models are complemented, improving the ability to express complex reaction dynamics while ensuring physical consistency such as mass conservation and energy conservation. At the same time, online Bayesian inference is used to dynamically update model parameters, enabling the digital twin model to continuously adapt to uncertain disturbances such as raw material fluctuations and catalyst changes, thereby improving the accuracy and real-time performance of simulation predictions.

[0021] Optionally, the training of the reinforcement learning agent includes: generating virtual data through policy exploration in a digital twin environment; obtaining real data through constrained policy verification in a real reaction system; and using the fused virtual data and real data for policy updates. The reward signal is constructed based on the following factors: the change in target product concentration, wherein the target product concentration and by-product concentration are obtained through online spectral analysis; Byproduct generation amount; The degree to which the temperature deviates from the target value; Control the range of motion changes.

[0022] By adopting the above technical solutions, a reinforcement learning training mechanism combining virtual and real data is realized, enabling efficient strategy exploration in a digital twin environment, and combining it with constrained verification in a real system to improve strategy reliability. At the same time, the model is updated by fusing virtual and real data to improve learning efficiency and generalization ability. By introducing product and byproduct information obtained from online spectral analysis into the reward function, and by integrating temperature deviation and control action fluctuations, the synergistic optimization and constrained control of reaction yield, selectivity and stability are achieved.

[0023] Optionally, it also includes: triggering a safety control strategy and restricting control actions when abnormal peaks or key components in the spectral signal exceed a threshold; When an unknown impurity peak is detected, the system automatically switches to conservative control mode.

[0024] By adopting the above technical solutions, an active safety protection mechanism based on online spectral information is realized. When abnormal spectral peaks or key components exceed limits, a safety control strategy can be triggered in a timely manner to limit or correct the control actions. At the same time, when an unknown impurity peak appears, it automatically switches to a conservative control mode, thereby effectively reducing the risk of reaction runaway and improving the safety and robustness of the chemical production process.

[0025] Optionally, the online spectral sensing module includes at least one of a Raman spectrometer, a near-infrared spectrometer, or a mid-infrared spectrometer; The reinforcement learning decision-making module is deployed on industrial edge computing devices; The control execution module is connected to a distributed control system (DCS) or a programmable logic controller (PLC). The safety constraint module includes: a temperature constraint unit, a reaction rate constraint unit, and a spectral safety threshold determination unit; The digital twin module includes virtual sensors for predicting the concentrations of key components that cannot be directly measured; The system constitutes a closed-loop optimized control structure of "spectral sensing - state construction - reinforcement learning decision-execution control - feedback update".

[0026] By adopting the above technical solutions, a closed-loop intelligent optimization control system for industrial sites was constructed, enabling real-time perception of the reaction process by various types of online spectroscopic devices and rapid decision-making and response on edge computing devices. Through integration with DCS or PLC systems, the engineering execution of control strategies was realized. At the same time, by combining safety constraint units and virtual sensors, the observability of key components and safety boundary control capabilities were improved, achieving real-time optimization and safe and stable operation of the reaction process.

[0027] In summary, this application includes at least one of the following beneficial technical effects: 1. By introducing online real-time spectral sensing and feature extraction technology, dynamic monitoring of key components and intermediates within the reaction system can be achieved, significantly improving the observability and state acquisition accuracy of the chemical reaction process; 2. By constructing a composite state vector that integrates spectral feature variables and process parameters, and combining it with a reinforcement learning agent for decision optimization, the real-time performance, accuracy, and adaptability of the reaction process control strategy are improved. 3. By introducing a digital twin model and a virtual-real fusion training mechanism, and combining safety constraints with online model updates, continuous optimization control of the reaction process is achieved under the premise of ensuring safe operation, thereby improving yield and selectivity and enhancing operational stability. Attached Figure Description

[0028] Figure 1 This is an architecture diagram of a dynamic optimization method and system for chemical reactions based on reinforcement learning and online spectral feedback, according to an embodiment of this application. Detailed Implementation

[0029] The following is in conjunction with the appendix Figure 1 This application will be described in further detail.

[0030] This application discloses a method for dynamic optimization of chemical reactions based on reinforcement learning and online spectral feedback, referring to... Figure 1 , Example 1

[0031] This embodiment provides a dynamic optimization method for chemical reactions based on reinforcement learning and online spectral feedback. The method includes the following steps: S1 Real-time Spectral Acquisition and Feature Extraction The reaction system is monitored in real time by an online spectral acquisition device, which can be one or more of Raman spectrometers, near-infrared spectrometers, or mid-infrared spectrometers. Preferably, in high reaction complexity scenarios, a multi-spectral fusion method is used to improve information integrity. The acquired raw spectral data is first preprocessed, including but not limited to: Baseline correction, noise filtering (such as Savitzky-Golay filtering), spectral normalization, and scattering correction; After preprocessing, feature extraction is performed on the spectral data to obtain at least one spectral feature variable. Specifically, feature extraction includes: 1) An adaptive peak identification and tracking method based on a sliding window is used to locate dynamically changing absorption peaks in the spectrum in real time; 2) Extract information on changes in peak position, peak intensity, and full width at half maximum (FWHM) to reflect the consumption of reactants and the formation of products; 3) Principal component analysis (PCA) or partial least squares regression (PLS) are used to map high-dimensional spectral data to a low-dimensional latent variable space to reduce redundant information and enhance modeling stability. Through the above processing, a set of spectral characteristic variables that can characterize the reaction process is output.

[0032] S2 State Construction and Multi-Source Information Fusion The spectral feature variables obtained in step S1 are fused with the process parameters to construct a composite state vector; Process parameters include, but are not limited to: reaction temperature, reaction pressure, feed flow rate, stirring speed, and reaction time; The composite state vector further includes: spectral latent variables, the first-order rate of change of the spectral latent variables, the integral term of the spectral latent variables (used to characterize the cumulative effect), and time series information; This multidimensional state construction method enables the system to simultaneously represent the instantaneous state and dynamic evolution trend of the response, thereby improving the expressive power of reinforcement learning input information.

[0033] S3 Reinforcement Learning Agent Decision Making The composite state vector is input into the reinforcement learning agent, which then outputs optimized control actions. In a preferred embodiment, the reinforcement learning agent employs a policy gradient algorithm or an actor-critic structure to adapt to scenarios with continuous control variables. Reinforcement learning agents include: Upper-level decision network: used to generate global control strategies based on the goal of maximizing long-term returns; Lower-level security execution network: used to constrain and modify the upper-level policies and output the final control actions; Safety projection layer: used to map actions that exceed safety boundaries to a preset safe and feasible domain; The upper-layer network focuses on optimizing response efficiency and economy, while the lower-layer network is responsible for enforcing engineering safety constraints, thus achieving a decoupled control structure of "optimization-safety".

[0034] S4 Execution Control and Feedback Acquisition The optimized control actions are applied to the reaction actuators, including but not limited to: adjusting the opening of the feed valve, adjusting the heating / cooling power, adjusting the pressure control valve, and controlling the stirring rate.

[0035] After execution, the system obtains real-time feedback on the reaction status and reward signals; The reward signal is constructed based on the following factors: changes in target product concentration (obtained by online spectral analysis), by-product generation, temperature deviation from the target value, and the magnitude of changes in control actions (to suppress drastic fluctuations). A comprehensive reward function is formed through a multi-objective weighted approach.

[0036] S5 Online Updates and Digital Twin Fixes Based on state feedback and reward signals, the reinforcement learning agent is updated online, and the digital twin model is dynamically corrected. Digital twin models are hybrid models that integrate mechanistic models and data-driven models, including: A reaction kinetics model based on the conservation of mass and energy; Parameter correction model trained based on historical data; The model parameters are updated using an online Bayesian inference method to enable a rapid response to changes in the chemical reaction process; Furthermore, reinforcement learning training combines virtual interactive data generated by digital twins with real data collected by actual response systems for fusion training, thereby improving the model's generalization ability and sample efficiency. Example 2

[0037] This embodiment provides a dynamic optimization system for chemical reactions based on reinforcement learning and online spectral feedback, including: 1) Online Spectral Sensing Module For acquiring real-time spectral data of the reaction system, Raman, near-infrared, or mid-infrared spectrometers are preferred. 2) Spectral processing and feature extraction module Used for dimensionality reduction and feature variable extraction of spectral data to extract key reaction information; 3) Data Fusion and State Construction Module Used to fuse spectral feature variables and process parameters to generate a composite state vector; 4) Digital Twin Module It is used to build mechanism-data hybrid models and provides a virtual simulation environment and virtual sensor functions for predicting the concentration of key components that cannot be directly measured. 5) Reinforcement Learning Decision Module Deployed in industrial edge computing devices for outputting control strategies based on state vectors; 6) Control Execution Module It connects to a distributed control system (DCS) or a programmable logic controller (PLC) to execute control commands; 7) Online learning and model update module Used for online updates of reinforcement learning models and digital twin models; 8) Safety constraint module It includes a temperature constraint unit, a reaction rate constraint unit, and a spectral safety threshold determination unit to ensure the safe operation of the system. The overall system constitutes a closed-loop optimized control structure of "spectral sensing - state construction - reinforcement learning decision-making - execution control - feedback update". Example 3

[0038] In a preferred embodiment, the system further includes an anomaly protection mechanism: When abnormal peaks or key components in the spectral signal exceed the threshold, the system automatically triggers a safety control strategy to restrict or block control actions. When an unknown impurity peak is detected, the system automatically switches to conservative control mode to reduce reaction risk. This mechanism can effectively prevent runaway reactions or escalation of side reactions, thus improving the safety of industrial applications. Example 4

[0039] The training process for a reinforcement learning agent includes: 1) Explore strategies and generate virtual interactive data within a digital twin environment; 2) Conduct constrained strategy verification in a real reaction system to obtain real data; 3) Integrate virtual data with real data for policy updates; By combining virtual and real training methods, training efficiency can be improved and the cost of industrial trial and error can be reduced. Example 5

[0040] The following explanation will be based on an esterification reaction process. The reaction system involves the reaction of alcohol and acid to produce ester products under catalytic conditions. This process has significant side reactions and exothermic effects, making it difficult for traditional PID control to achieve optimal control. After adopting the method of this application: 1) Real-time monitoring of changes in characteristic peaks of ester products in the reaction system using near-infrared spectroscopy; 2) Obtain 3D latent variables as state inputs through PCA dimensionality reduction; 3) The reinforcement learning agent outputs feed ratio and temperature control strategies; 4) Digital twin models are used to predict yield changes under different control strategies; 5) The system executes control strategies in real time on edge computing devices; Experimental results show that under stable operating conditions: The yield of the target product is increased by approximately 8% to 15%; The amount of by-products generated is reduced by approximately 10% to 20%; Temperature fluctuations decreased by approximately 30%; The system response time has been significantly reduced; Meanwhile, even under abnormal disturbances (fluctuations in raw materials), the system can still maintain stable operation and has not experienced any loss of control.

[0041] The technical effect of this application, through the deep integration of online spectral sensing, digital twin modeling, and reinforcement learning decision-making, achieves intelligent, dynamic, and safe control of chemical reaction processes, and has the following advantages: Improve process observability; Enhance the adaptive capability of control strategies; Improve reaction yield and selectivity; Improve the safety and robustness of industrial operations; Reduce the cost of manual intervention and experimentation.

[0042] The implementation principle of a dynamic optimization method and system for chemical reactions based on reinforcement learning and online spectral feedback in this application is as follows: The reaction system is monitored in real time using online spectral technology. The continuously acquired spectral signals are preprocessed and feature extracted, then transformed into low-dimensional feature variables that characterize the reaction process. These variables are then fused with process parameters such as temperature, pressure, flow rate, and time to construct a composite state vector containing multi-dimensional information. This transforms the chemical reaction process from "unobservable" to "computable state space." This state vector is input into a reinforcement learning agent, which performs decision modeling under a policy gradient or actor-commentator framework to generate optimal control actions. Boundary corrections to the actions are then performed using a safety constraint structure to ensure that control commands meet industrial safety requirements before being applied to the reaction actuator. This system enables dynamic control of the reaction process. By introducing a digital twin model to construct a virtual reaction environment that integrates mechanism and data, the system performs strategy pre-training and data expansion within this environment. It then combines this with feedback data from the real reaction process for joint updates, allowing the reinforcement learning model to continuously optimize in the virtual-real interaction. Furthermore, online Bayesian inference is used to dynamically correct the parameters of the digital twin model, enabling it to reflect changes in the actual chemical reaction process in real time, improving prediction accuracy and simulation reliability. Finally, through a closed-loop control mechanism of "spectral sensing - state construction - intelligent decision-making - safe execution - feedback learning," the system achieves real-time optimized control of the chemical reaction process. This improves the yield and selectivity of the target product while effectively reducing by-product generation and energy consumption fluctuations, and significantly enhances the stability and safety of system operation.

[0043] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A dynamic optimization method for chemical reactions based on reinforcement learning and online spectral feedback, characterized in that, Includes the following steps: S1. The reaction system is monitored in real time by an online spectral acquisition device to obtain raw spectral data, and the raw spectral data is preprocessed and feature extracted to obtain at least one spectral feature variable characterizing the reaction process. S2. The spectral feature variables are fused with the process parameters in the reaction process to construct a composite state vector containing reaction state information; S3. Input the composite state vector into the reinforcement learning agent, and the reinforcement learning agent outputs the corresponding optimized control action based on the current state; S4. Apply the optimized control action to the reaction actuator to adjust the reaction process, and obtain the status feedback and reward signal after execution; S5. Based on the aforementioned state feedback and reward signals, the reinforcement learning agent is updated online, and the digital twin model is dynamically corrected simultaneously. The training process of the reinforcement learning agent includes fusing virtual interaction data generated by a digital twin model with actual reaction data for training.

2. A dynamic optimization system for chemical reactions based on reinforcement learning and online spectral feedback, characterized in that, include: An online spectral sensing module is used to collect real-time spectral data of the reaction system; It includes a spectral processing and feature extraction module, which is used to perform dimensionality reduction and feature variable extraction on the spectral data; It includes a data fusion and state construction module, which is used to fuse spectral feature variables and process parameters to generate a composite state vector; It includes a digital twin module, which is used to construct a mechanism-data hybrid model of the reaction process and provide a virtual simulation environment; It includes a reinforcement learning decision module, which is used to output an optimized control strategy based on the composite state vector; It includes a control execution module, which is used to convert the control strategy into execution instructions and apply them to the reaction device; It includes an online learning and model update module, which is used to dynamically update the reinforcement learning model and digital twin model based on actual operating data; It includes a security constraint module, which is used to perform security verification and constraint correction on the control strategy.

3. The method according to claim 1, characterized in that, The spectral feature extraction in step S1 includes: Adaptive peak identification and tracking based on sliding window; Extract information on the peak position, peak intensity, and full width at half maximum (FWHM) variation of characteristic peaks; Alternatively, high-dimensional spectral data can be mapped to low-dimensional latent variables through principal component analysis (PCA) or partial least squares regression (PLS).

4. The method according to claim 1, characterized in that, The composite state vector includes: Temperature, pressure, feed flow rate, spectral latent variable, rate of change and integral term of the spectral latent variable, and reaction time information.

5. The method according to claim 1, characterized in that, The reinforcement learning agent described in step S3 employs a policy gradient algorithm or an actor-critic structure.

6. The method according to claim 1, characterized in that, The reinforcement learning agent includes: Upper-level decision network, which is used to generate control strategies based on long-term optimization objectives; A lower-level security execution network is used to modify the control strategy and output the final control action under the condition of satisfying security constraints; The lower-level secure execution network includes a secure projection layer, which is used to map control actions that do not meet the constraints to a preset secure boundary range.

7. The method according to claim 1, characterized in that, The digital twin model is a hybrid model combining a mechanistic model and a data-driven model, including: a reaction kinetic model based on mass and energy conservation; and a parameter correction model trained on historical data. The parameters of the digital twin model are dynamically updated using an online Bayesian inference method.

8. The method according to claim 1, characterized in that, The training of the reinforcement learning agent includes: generating virtual data through policy exploration in a digital twin environment; obtaining real data through constrained policy verification in a real reaction system; and using the fused virtual and real data for policy updates. The reward signal is constructed based on the following factors: the change in target product concentration, wherein the target product concentration and by-product concentration are obtained through online spectral analysis; Byproduct generation amount; The degree to which the temperature deviates from the target value; Control the range of motion changes.

9. The method according to claim 1, characterized in that, Also includes: When abnormal peaks or key components exceed thresholds in the spectral signal, a safety control strategy is triggered and control actions are restricted. When an unknown impurity peak is detected, the system automatically switches to conservative control mode.

10. The system according to claim 2, characterized in that, The online spectral sensing module includes at least one of a Raman spectrometer, a near-infrared spectrometer, or a mid-infrared spectrometer; The reinforcement learning decision-making module is deployed on industrial edge computing devices; The control execution module is connected to a distributed control system (DCS) or a programmable logic controller (PLC). The safety constraint module includes: a temperature constraint unit, a reaction rate constraint unit, and a spectral safety threshold determination unit; The digital twin module includes virtual sensors for predicting the concentrations of key components that cannot be directly measured; The system constitutes a closed-loop optimized control structure of "spectral sensing - state construction - reinforcement learning decision-making - execution control - feedback update".