Continuous flow reactor multi-modal ai optimization system and method for high value chemical synthesis
Patent Information
- Application Number
- CN202511810901.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-01-31
AI Technical Summary
这种方法通常难以全面捕捉到复杂连续流反应器中物理变化与化学转化之间的深层次、瞬态的因果影响关系
本发明通过深度融合多模态数据与因果推断机制,构建了对化学反应过程更为深刻和可解释的认知模型。该方法不局限于分析过程参数和光谱数据表面的相关性,而是通过构建动态因果关系图并识别关键的因果屏障层,深入挖掘导致系统性能漂移的根本性因果链条,从而使AI控制系统能够像领域专家一样,基于内在机理而非表面现象进行判断,这为后续的精准预测与决策奠定了坚实可靠的基础,提升了控制系统的感知智能水平;
Smart Images

Figure CN121613739B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of automation control and artificial intelligence, and in particular to a multimodal AI optimization system and method for continuous flow reactors used in the synthesis of high-value chemicals. Background Technology
[0002] Continuous flow reactor technology in the synthesis of high-value chemicals has become an important development direction in the modern chemical industry due to its precise temperature control, efficient mass and heat transfer, and excellent reproducibility. However, continuous flow reaction processes typically involve complex multivariate coupling and nonlinear dynamic changes. Achieving real-time monitoring, accurate prediction, and robust optimization control of the reaction process remains a key technical challenge in this field. In the field of chemical process control and optimization, related intelligent technologies mainly focus on using data models for prediction and condition optimization.
[0003] Among related technologies, Chinese invention patent CN120356568A discloses an AI-driven chemical experiment simulation and result prediction system. This system collects multimodal data through an experimental data acquisition module, utilizes spectral feature analysis and reaction kinetic modeling to construct a coupled kinetic model. This model is used to predict product distribution probabilities and optimize reaction conditions, and includes an experimental parameter feedback module to adjust reaction parameters in real time. Furthermore, the system also includes a reaction safety assessment model to identify risks and provide safety protection measures.
[0004] Regarding the aforementioned technologies, the inventors believe that the existing technologies primarily focus on model-based simulation and prediction, and their core still relies on the establishment and solution of coupled dynamic models. This method typically struggles to fully capture the deep-seated, transient causal relationships between physical changes and chemical transformations in complex continuous flow reactors. Furthermore, its optimization and adjustment mechanisms are based on feedback or feedforward of prediction results, lacking the ability to intervene in potential performance drift, and failing to employ an end-to-end autonomous evolutionary decision-making mechanism such as multi-objective deep reinforcement learning. Consequently, its robustness and adaptability in dealing with multi-objective conflicts and unforeseen operating conditions during long-term operation still need improvement. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides a multimodal AI optimization system and method for continuous flow reactors used in the synthesis of high-value chemicals. It employs a control paradigm that deeply integrates causal inference with multi-objective deep reinforcement learning, enabling proactive, adaptive, and optimal control of the chemical synthesis process.
[0006] To achieve the above objectives, this application adopts the following technical solution: The first aspect provides a multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals, including: S1. Simultaneously acquire real-time process parameters and online spectral data of the continuous flow reactor, and perform preliminary screening of causal characteristics on the real-time process parameters and online spectral data to generate screened process parameter characteristics and spectral data characteristics. S2. Based on the filtered process parameter features and spectral data features, construct an initial dynamic causal relationship diagram, define a causal barrier layer in the initial dynamic causal relationship diagram, and generate a dynamic causal relationship diagram containing the causal barrier layer. S3. Integrate the screened process parameter features, spectral data features, and topological structure features of the dynamic causal relationship diagram to generate a causal-physicochemical fusion state descriptor; S4. Based on the causal-physicochemical fusion state descriptor, use a preset dynamic evolution prediction model to predict the future trajectory and reverse target trajectory of key performance indicators, and generate the predicted trajectory of key performance indicators and the predicted trajectory of reverse targets. S5. Obtain the optimized objective function for defining the trade-off between product yield and purity, and construct the causal intervention reward function based on the optimized objective function and the reverse objective prediction trajectory. S6. Using the causal-physicochemical fusion state descriptor as the state input and the causal intervention reward function as the optimization objective, a multi-objective deep reinforcement learning algorithm is used to train the control policy network and output the control parameter adjustment action. S7. Adjust the operation parameters of the continuous flow reactor according to the control parameters, and dynamically update the dynamic causal relationship diagram and control strategy network based on the operation feedback.
[0007] Based on the above technical solution, the multimodal AI optimization method for continuous flow reactors for the synthesis of high-value chemicals provided in this application adopts a control paradigm that deeply integrates causal inference with multi-objective deep reinforcement learning, which can achieve active, adaptive and optimal control of the chemical synthesis process.
[0008] In conjunction with the first aspect above, in one possible implementation, the simultaneous acquisition of real-time process parameters and online spectral data of the continuous flow reactor, and the preliminary screening of causal characteristics of the real-time process parameters and online spectral data, includes: Acquire real-time process parameters and online spectral data of a continuous flow reactor; The real-time process parameters are subjected to time-series feature extraction to generate a physical state feature vector; The online spectral data is subjected to dimensionality reduction processing to generate reaction kinetic feature vectors; The physical state feature vector and reaction kinetic feature vector are analyzed using a causal discovery algorithm to select a subset of features with causal driving relationships, and to generate the selected process parameter features and spectral data features.
[0009] In conjunction with the first aspect above, in one possible implementation, based on the filtered process parameter characteristics and spectral data characteristics, an initial dynamic causal relationship graph is constructed, and a causal barrier layer is defined in the initial dynamic causal relationship graph, including: Identify the causal paths between variables in the filtered process parameter features and spectral data features, and generate an initial dynamic causal relationship diagram; Extract the negative causal chains that cause performance drift from the initial dynamic causal relationship graph, quantify their influence strength, and generate causal barrier layer definition parameters; The parameters defining the causal barrier layer are embedded into the initial dynamic causal relationship graph, and the graph structure is dynamically updated to generate a dynamic causal relationship graph containing the causal barrier layer.
[0010] In conjunction with the first aspect described above, in one possible implementation, the integration of the filtered process parameter features, spectral data features, and the topological features of the dynamic causal relationship graph includes: Extract the state information of causal barrier nodes and the strength information of causal paths from the dynamic causal relationship graph to generate a causal topological feature vector; The filtered process parameter features, spectral data features, and causal topological feature vectors are input into a multi-layer neural network for nonlinear combination to generate intermediate fusion features. The intermediate fusion features are normalized to generate a causal-physicochemical fusion state descriptor.
[0011] In conjunction with the first aspect above, in one possible implementation, the prediction of the future trajectory and reverse target trajectory of key performance indicators using a preset dynamic evolution prediction model includes: The causal-physicochemical fusion state descriptor is input into a preset dynamic evolution prediction model to calculate the values of several future time points of key performance indicators and generate the prediction trajectory of key performance indicators. Based on the causal barrier information embedded in the causal-physicochemical fusion state descriptor, the performance degradation trend under no-intervention conditions is predicted, and a reverse target prediction trajectory is generated. Verify the consistency between the predicted trajectory of the key performance indicators and the predicted trajectory of the reverse target, and output the final prediction result.
[0012] In conjunction with the first aspect above, in one possible implementation, constructing the causal intervention reward function based on the optimization objective function and the reverse target prediction trajectory includes: The performance drift index is extracted from the predicted trajectory of the reverse target, and the deviation from the ideal interval defined by the optimization objective function is calculated. Obtain the safety constraint function that defines the safe operating range, and combine it with the deviation to construct a multi-objective composite reward function that includes positive rewards and negative penalties, thereby generating a causal intervention reward function.
[0013] In conjunction with the first aspect above, in one possible implementation, the causal-physicochemical fusion state descriptor is used as the state input, the causal intervention reward function is used as the optimization objective, and a multi-objective deep reinforcement learning algorithm is employed to train the control policy network, including: Define a reinforcement learning environment with the causal-physicochemical fusion state descriptor as the state space and the adjustment range of the control parameters of the continuous flow reactor as the action space; Using a multi-objective deep reinforcement learning algorithm, with the causal intervention reward function as the optimization objective, the control policy network is iteratively trained in a reinforcement learning environment. When the control strategy network converges, it outputs control parameter adjustment actions based on the current state.
[0014] In conjunction with the first aspect above, in one possible implementation, adjusting the operating parameters of the continuous flow reactor based on the control parameters, and dynamically updating the dynamic causal relationship diagram and control strategy network based on operational feedback includes: Perform the control parameter adjustment action to modify the operating parameters of the continuous flow reactor; Monitor changes in key performance indicators after execution and generate operational feedback data; The dynamic causal relationship graph and control strategy network are updated using the operational feedback data.
[0015] In conjunction with the first aspect above, in one possible implementation, the method further includes: initiating an autonomous evolution mechanism upon detecting a new operating condition, wherein: Real-time comparison of the current causal-physicochemical fusion state descriptor with historical patterns to identify unlearned operating condition features; When a new working condition is identified, an incremental learning process is triggered to update the definition of the causal barrier layer and the reward function; By retraining the control policy network online, it can adapt to new operating conditions and maintain optimized performance.
[0016] Secondly, a multimodal AI optimization system for continuous flow reactors used in the synthesis of high-value chemicals is provided, including: a data acquisition and causal feature initial screening module, a dynamic causal graph construction module, a state descriptor generation module, a trajectory prediction module, a reward function construction module, a reinforcement learning decision-making module, and a parameter adjustment and evolution module; among which: The data acquisition and causal feature screening module is used to simultaneously acquire real-time process parameters and online spectral data of the continuous flow reactor, and to perform causal feature screening on the real-time process parameters and online spectral data to generate screened process parameter features and spectral data features. The dynamic causal graph construction module is used to construct an initial dynamic causal relationship graph based on the filtered process parameter features and spectral data features, and to define a causal barrier layer in the initial dynamic causal relationship graph to generate a dynamic causal relationship graph containing the causal barrier layer. The state descriptor generation module is used to fuse the screened process parameter features, spectral data features, and topological structure features of the dynamic causal relationship graph to generate a causal-physicochemical fusion state descriptor. The trajectory prediction module is used to predict the future trajectory and reverse target trajectory of key performance indicators based on the causal-physicochemical fusion state descriptor and using a preset dynamic evolution prediction model, and to generate the predicted trajectory of key performance indicators and the predicted trajectory of reverse targets. The reward function construction module is used to obtain an optimized objective function for defining the trade-off between product yield and purity, and to construct a causal intervention reward function based on the optimized objective function and the reverse objective prediction trajectory. The reinforcement learning decision module is used to take the causal-physicochemical fusion state descriptor as the state input and the causal intervention reward function as the optimization objective, and to train the control policy network using a multi-objective deep reinforcement learning algorithm to output the control parameter adjustment action. The parameter adjustment and evolution module is used to adjust the action according to the control parameters, adjust the operating parameters of the continuous flow reactor, and dynamically update the dynamic causal relationship diagram and control strategy network based on the operating feedback.
[0017] Compared with the prior art, the present invention has the following advantages: This invention constructs a more profound and interpretable cognitive model of chemical reaction processes by deeply integrating multimodal data with causal inference mechanisms. This method goes beyond analyzing the surface correlations of process parameters and spectral data; instead, it constructs dynamic causal relationship graphs and identifies key causal barriers to delve into the fundamental causal chains that cause system performance drift. This enables AI control systems to make judgments based on intrinsic mechanisms rather than surface phenomena, much like domain experts. This lays a solid and reliable foundation for subsequent accurate prediction and decision-making, enhancing the perceptual intelligence level of the control system. The control method proposed in this invention is forward-looking and strategic, enabling proactive preventative optimization control. By predicting the future trajectory of key performance indicators and the performance degradation trend under no-intervention conditions, and constructing a unique causal intervention reward function, this method transforms the control objective from the traditional passive response to deviation to actively suppressing identified negative causal paths. This allows the reinforcement learning agent to learn how to anticipate and avoid potential future problems during training, achieving a leap from "post-event remediation" to "pre-event prevention," effectively overcoming the performance drift problem caused by catalyst deactivation, raw material fluctuations, and other factors during long-term operation. This invention establishes a complete closed-loop adaptive and autonomous evolutionary control framework, endowing the system with powerful long-term autonomous optimization capabilities. This method integrates reinforcement learning decision-making, physical system execution, and operational feedback learning, enabling the control strategy network to continuously learn from actual operational results and iterate itself. Furthermore, by introducing an autonomous detection and evolution mechanism for new operating conditions, it can proactively identify and adapt to significant disturbances such as raw material batch changes or process condition variations, dynamically updating its internal model and control strategy. This ensures the optimized system remains highly efficient and stable throughout its entire lifecycle, achieving truly intelligent autonomous operation of complex industrial processes. It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A structural architecture diagram of a multimodal AI optimization system for a continuous flow reactor used in the synthesis of high-value chemicals, provided in an embodiment of this application; Figure 2 A schematic diagram of the process for a multimodal AI optimization method for the synthesis of high-value chemicals in a continuous flow reactor, provided in an embodiment of this application; Figure 3 This is a schematic diagram illustrating the quantification of the influence intensity of the negative causal chain provided in the embodiments of this application.
[0020] Figure 4 This is a comparison chart of the product yield auto-evolution control effect provided in the embodiments of this application. Detailed Implementation
[0021] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0022] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0023] The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals provided in this application embodiment can be applied to, for example... Figure 1 In the multimodal AI optimization system 100 for continuous flow reactors used in the synthesis of high-value chemicals shown, such as Figure 1 As shown, the system includes: a data acquisition and initial screening module for causal features, a dynamic causal graph construction module, a state descriptor generation module, a trajectory prediction module, a reward function construction module, a reinforcement learning decision-making module, and a parameter adjustment and evolution module; among which: The data acquisition and causal feature screening module is used to simultaneously acquire real-time process parameters and online spectral data of the continuous flow reactor, and to perform causal feature screening on the real-time process parameters and online spectral data to generate screened process parameter features and spectral data features. The dynamic causal graph construction module is used to construct an initial dynamic causal relationship graph based on the filtered process parameter features and spectral data features, and to define a causal barrier layer in the initial dynamic causal relationship graph to generate a dynamic causal relationship graph containing the causal barrier layer. The state descriptor generation module is used to fuse the screened process parameter features, spectral data features, and topological structure features of the dynamic causal relationship graph to generate a causal-physicochemical fusion state descriptor. The trajectory prediction module is used to predict the future trajectory and reverse target trajectory of key performance indicators based on the causal-physicochemical fusion state descriptor and using a preset dynamic evolution prediction model, and to generate the predicted trajectory of key performance indicators and the predicted trajectory of reverse targets. The reward function construction module is used to obtain an optimized objective function for defining the trade-off between product yield and purity, and to construct a causal intervention reward function based on the optimized objective function and the reverse objective prediction trajectory. The reinforcement learning decision module is used to take the causal-physicochemical fusion state descriptor as the state input and the causal intervention reward function as the optimization objective, and to train the control policy network using a multi-objective deep reinforcement learning algorithm to output the control parameter adjustment action. The parameter adjustment and evolution module is used to adjust the action according to the control parameters, adjust the operating parameters of the continuous flow reactor, and dynamically update the dynamic causal relationship diagram and control strategy network based on the operating feedback.
[0024] like Figure 2 As shown in the embodiments of this application, a multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals is provided, including: S1. Simultaneously acquire real-time process parameters and online spectral data of the continuous flow reactor, and perform preliminary screening of causal characteristics on the real-time process parameters and online spectral data to generate screened process parameter characteristics and spectral data characteristics. S2. Based on the filtered process parameter features and spectral data features, construct an initial dynamic causal relationship diagram, define a causal barrier layer in the initial dynamic causal relationship diagram, and generate a dynamic causal relationship diagram containing the causal barrier layer. S3. Integrate the screened process parameter features, spectral data features, and topological structure features of the dynamic causal relationship diagram to generate a causal-physicochemical fusion state descriptor; S4. Based on the causal-physicochemical fusion state descriptor, use a preset dynamic evolution prediction model to predict the future trajectory and reverse target trajectory of key performance indicators, and generate the predicted trajectory of key performance indicators and the predicted trajectory of reverse targets. S5. Obtain the optimized objective function for defining the trade-off between product yield and purity, and construct the causal intervention reward function based on the optimized objective function and the reverse objective prediction trajectory. S6. Using the causal-physicochemical fusion state descriptor as the state input and the causal intervention reward function as the optimization objective, a multi-objective deep reinforcement learning algorithm is used to train the control policy network and output the control parameter adjustment action. S7. Adjust the operation parameters of the continuous flow reactor according to the control parameters, and dynamically update the dynamic causal relationship diagram and control strategy network based on the operation feedback.
[0025] It is worth noting that by simultaneously collecting and fusing information from two modalities—process parameters characterizing the macroscopic physical state and online spectral data revealing microscopic chemical dynamics—a feature space comprehensively describing the system state is constructed. Based on this, causal inference techniques are introduced to construct and dynamically maintain a dynamic causal relationship graph containing a "causal barrier layer." This graph not only reveals the causal relationships between variables but also accurately identifies and quantifies the root cause chains leading to system performance drift. Physicochemical characteristics are deeply integrated with this deep causal topological structure to generate a "causal-physicochemical fusion state descriptor." This descriptor is then used to drive a dynamic evolution prediction model, proactively predicting the future trajectory of system performance and its degradation trend under no-intervention conditions. Based on a multi-objective optimization function that balances yield and purity, and combined with the prediction of potential performance drift, a unique "causal intervention reward function" is constructed to guide a multi-objective deep reinforcement learning algorithm to train a control policy network capable of outputting optimal control parameter adjustment actions. This forms a closed loop, dynamically updating the causal graph and control policy through continuous operational feedback, thereby achieving adaptive optimization.
[0026] In one possible implementation of the embodiments of this application, combined with Figure 2 The above S1 can be implemented through the following S11, S12, S13 and S14, which are explained in detail below: S11. Obtain real-time process parameters and online spectral data of the continuous flow reactor; In some implementations, two types of data are simultaneously acquired by various sensors integrated into the continuous flow reactor. One type is real-time process parameters, such as reaction temperature, feed rate, reactor pressure, and residence time; these parameters collectively describe the macroscopic physical environment in which the reaction occurs. The other type is online spectral data, obtained through analytical techniques such as online near-infrared spectroscopy, Raman spectroscopy, or ultraviolet-visible spectroscopy. This data reflects the real-time concentration changes of each component within the reaction system, revealing microscopic chemical reaction kinetics. Both types of data are assigned a unified timestamp during acquisition to ensure temporal alignment.
[0027] For example, real-time process parameters of the continuous flow reactor are acquired, including reaction temperature, feed rate of material A, reactor pressure, and residence time, at a frequency of 1 Hz. Online spectral data are acquired using a near-infrared (NIR) spectrometer, covering a wavelength range... A complete raw spectrum is acquired every 5 seconds. Both types of data are assigned a consistent timestamp during acquisition to ensure time alignment.
[0028] S12. Extract time-series features from the real-time process parameters to generate a physical state feature vector; In some implementations, time-series features are extracted from the acquired real-time process parameters. Since process parameters are sequential data that changes over time, a fixed time window is selected to capture their dynamic characteristics. Statistical features characterizing the physical state within this window are extracted, such as the mean, standard deviation, rate of change, or slope of temperature. These statistical features are then combined to form a fixed-dimensional physical state feature vector. This vector encapsulates the dynamic trends and stability information of the process parameters over a period of time.
[0029] For example, a fixed time window of 5 minutes is selected, and the reaction temperature data within that window are recorded. Calculate its mean, standard deviation, and slope of change. Calculate the feed flow rate. Its maximum value and volatility. Combining these statistical characteristics, such as This forms a fixed dimension, such as a 10-dimensional physical state feature vector. .
[0030] S13. Perform dimensionality reduction processing on the online spectral data to generate a reaction kinetic feature vector; In some implementations, dimensionality reduction is performed on high-dimensional online spectral data. The original spectral data typically contains hundreds or even thousands of wavelengths, exhibiting high information redundancy and noise. Dimensionality reduction algorithms such as Principal Component Analysis (PCA) or Partial Least Squares (PLS) are used to project the original spectral matrix onto a few principal components or latent variables. These principal components can retain, to the greatest extent possible, the variance information related to chemical changes such as reactant consumption, intermediate formation, and target product accumulation. The resulting low-dimensional vector is the reaction kinetic eigenvector, representing the core chemical information of the reaction process.
[0031] For example, the partial least squares (PLS) algorithm is used to reduce the dimensionality of the acquired high-dimensional near-infrared spectral data. LLS is used to identify latent variables or principal components with the largest covariance to the target product concentration (KPI). The original spectral matrix is then projected onto these few latent variables to extract five latent variables. The resulting 5-dimensional low-dimensional vector is the reaction kinetic feature vector. It characterizes the core chemical information in the reaction process.
[0032] S14. Use a causal discovery algorithm to analyze the physical state feature vector and reaction kinetic feature vector, filter out the feature subset with causal driving relationship, and generate the filtered process parameter features and spectral data features.
[0033] In some implementations, causal discovery algorithms are used to jointly analyze the physical state feature vector and the reaction kinetics feature vector. For example, constraint-based PC or FCI algorithms can be employed. These algorithms analyze the conditional independence between different variables, thereby inferring whether a direct causal driving relationship exists between them, rather than a simple correlation. Through this analysis, feature variables with significant causal associations are selected from the two feature vectors, forming a feature subset. For example, the algorithm might find that a change in feed flow rate is the direct cause of changes in the intensity of a certain spectral feature peak and the concentration of the final product. The final output, the selected process parameter features and spectral data features, is a set of core variables with the most informational value, verified for causal relationships, providing high-quality input for subsequently constructing an accurate dynamic causal relationship diagram.
[0034] For example, using a constraint-based PC algorithm (Peter and Clark algorithm) to... and The algorithm performs joint analysis on all feature variables. For example, it infers a certain feature. It is the cause of a certain latent variable in the 5-dimensional features The direct cause of the change, and related to another feature The relationships between them are merely correlational. The algorithm ultimately selects a subset of feature variables that have significant causal driving relationships, for example: These constitute the final filtered process parameter characteristics and spectral data characteristics.
[0035] In one possible implementation of the embodiments of this application, combined with Figure 2 The above S2 can be implemented through the following S21, S22 and S23, which are explained in detail below:
[0036] S21. Identify the causal paths between variables in the filtered process parameter features and spectral data features, and generate an initial dynamic causal relationship diagram; In some implementations, the selected feature variables are used as nodes, and dynamic causal discovery algorithms are employed to identify causal paths between these variables, generating an initial dynamic causal relationship graph. Dynamic Bayesian networks or variant algorithms based on Granger causality tests can be used here. These algorithms are specifically designed for processing time-series data and can construct directed graphs by analyzing time delays and conditional dependencies between variables. In this graph, an arrow pointing from node A to node B indicates that a change in variable A is one of the reasons why variable B changes at subsequent time points, thus initially revealing the interaction mechanism between multimodal features within a continuous flow reactor.
[0037] For example, the Dynamic Bayesian Network (DBN) algorithm is used to analyze the selected feature subset. The relationships between variables in the data are analyzed by examining the time steps. and time step The conditional dependencies between them are used to generate a directed graph. Within this initial dynamic causal graph, a causal path is identified: Changes in feed flow rate lead to changes in a certain spectral characteristic, and This change in spectral characteristics leads to a change in the average temperature of the reactor.
[0038] S22. Extract the negative causal chains that cause performance drift from the initial dynamic causal relationship graph, quantify their influence intensity, and generate causal barrier layer definition parameters.
[0039] In some implementations, negative causal chains leading to performance drift are extracted from the initial dynamic causal graph. Performance drift refers to the phenomenon where key performance indicators, such as product yield or purity, gradually deviate from their optimal state over time. Path search algorithms in graph theory, such as performing a reverse depth-first search starting from nodes representing key performance indicators, can identify all upstream causal paths that can affect these indicators. Combining prior chemical knowledge or analyzing historical data, paths leading to performance degradation are identified as negative causal chains. For example, a path starting from a spectral feature node representing catalyst deactivation, passing through several intermediate variable nodes, and finally pointing to a node representing a decrease in product yield is a typical negative causal chain. The influence strength of these negative causal chains can be quantified using methods such as structural equation modeling to calculate the path coefficients on each causal path. The overall influence strength of a negative causal chain is defined as: ; in, negative causal chain The overall intensity of the impact; Represents the root cause node To the performance indicator node A path; For the nodes on this path To the node The causal influence coefficient, obtained by analyzing historical operating data, quantifies... A unit change The direct impact. This includes identifying all negative causal chains and their corresponding impact strengths. This summary constitutes the parameters for defining the causal barrier layer. Here, the causal barrier layer is not a physical entity, but rather a collective description of all critical causal paths leading to performance degradation in the system; it serves as strategic information to guide subsequent intervention and control. For example... Figure 3 As shown, the three main negative causal chains (P1, P2, P3) identified and their quantified overall influence strength are illustrated. This diagram visually reflects the potential negative impact of each negative causal chain on product yield, and serves as the basis for defining causal barrier layers and constructing intervention strategies.
[0040] For example, a reverse depth-first search can be performed on key performance indicator nodes, such as product yield. Identify all factors that lead to low yield. A descending causal path. A negative causal chain was extracted. In other words, excessively high flow rates lead to temperature fluctuations, which in turn reduce yield. Structural equation modeling (SEM) was used to calculate the causal influence coefficients along each path. Assuming the flow rate is obtained through historical data analysis... Temperature Causal influence coefficient It has a positive impact. Temperature For yield Causal influence coefficient It has a negative impact; excessively high temperatures lead to side reactions. According to the formula... This negative causal chain Overall impact strength The calculation is as follows: All identified negative causal chains and their corresponding influence strengths Summarize and generate the parameters for defining the causal barrier layer.
[0041] S23. Embed the causal barrier layer definition parameters into the initial dynamic causal relationship graph, dynamically update the graph structure, and generate a dynamic causal relationship graph containing the causal barrier layer.
[0042] In some implementations, the causal barrier layer definition parameters are embedded into the initial dynamic causal graph, dynamically updating the graph structure. This is done by adding attributes to nodes and edges in the graph, rather than changing its basic connection topology. Nodes belonging to the identified negative causal chains are marked, and the quantified influence strength is then used. or path coefficients for each segment The weights of the corresponding edges are labeled. After this embedding and updating, a dynamic causal relationship graph containing a causal barrier layer is finally generated. This graph not only shows the causal relationships between variables, but also clearly highlights the critical paths that lead to system performance degradation and their severity.
[0043] For example, the quantified intensity of the impact Embedded into the corresponding paths of the initial dynamic causal relationship graph. The paths... and The edges above are marked negatively, and... or piecewise coefficient It is labeled as its edge weight attribute. In addition, nodes belonging to negative causal chains... Key nodes marked as causal intervention points. After embedding and updating, a dynamic causal relationship graph containing causal barrier layers is finally obtained.
[0044] In one possible implementation, combining Figure 2 The above-mentioned S3 can be implemented through the following S31, S32 and S33, which are explained in detail below: S31. Extract the state information of the causal barrier nodes and the strength information of the causal path in the dynamic causal relationship graph to generate a causal topological feature vector;
[0045] In some implementations, the topological features of the dynamic causal graph are extracted. Based on the dynamic causal graph containing causal barrier layers, two types of key information are extracted. The first type is the state information of the causal barrier nodes, i.e., the current values of nodes identified as being on negative causal chains causing performance drift. The second type is the strength information of the causal paths, i.e., the weights of the edges connecting the nodes in these negative causal chains, which represent the strength of the causal influence. The current values of these nodes and the weights of the paths are collected and arranged into a vector, thus generating the causal topological feature vector. This vector not only contains the current state, but more importantly, it embeds diagnostic information for potential problems.
[0046] For example, based on the generated dynamic causal relationship graph containing causal barrier layers, the following information is extracted: Node state information: Extracting negative causal chains. upper node The current instantaneous value and node The current average value, for example for ; for Path strength information: Extracts the overall impact strength of this negative causal chain. These values are arranged in a predetermined order to form a 3-dimensional causal topological feature vector. .
[0047] S32. Input the filtered process parameter features, spectral data features and causal topological feature vectors into a multi-layer neural network for nonlinear combination to generate intermediate fusion features;
[0048] In some implementations, the filtered process parameter features, spectral data features, and newly generated causal topological feature vectors are nonlinearly combined. These three independent feature vectors are then concatenated along their vector dimensions to form a longer combined feature vector. This combined vector is then fed into a pre-designed multi-layer neural network. Through its multi-layered structure and nonlinear activation function, this network learns and extracts the complex, nonlinear relationships between these three different source features, mapping them to a more representative latent space to generate intermediate fused features.
[0049] For example, suppose the filtered process parameter characteristics 10-dimensional spectral data features 5-dimensional, causal topological features It is 3-dimensional. These three vectors are concatenated to form a... The 18-dimensional combined feature vector is used as input to a pre-trained multilayer perceptron (MLP). This MLP contains three layers: for example, an input layer with 18 neurons, a first hidden layer with 128 neurons using ReLU activation, a second hidden layer with 64 neurons using Tanh activation, and an output layer with 32 neurons. The 32-dimensional vector from the output layer represents the intermediate fused features.
[0050] S33. Normalize the intermediate fusion features to generate a causal-physicochemical fusion state descriptor.
[0051] In some implementations, the intermediate fused features are normalized. Since the feature values output by the neural network may vary over a wide range, they need to be scaled to a uniform interval, such as between 0 and 1, to facilitate subsequent model training and improve its stability. The min-max normalization method can be used, and its calculation process is as follows: ;
[0052] in, These are the normalized eigenvalues. These are the original values in the intermediate fusion features. and These are the minimum and maximum values of the feature in the training dataset, respectively. By performing this operation on each dimension of the intermediate fused features, the causal-physicochemical fused state descriptor is finally obtained. This is a comprehensive state vector with fixed dimensions, normalized values, and a deep fusion of physical, chemical, and causal logic.
[0053] For example, the min-max normalization method is used to process the 32-dimensional intermediate fusion features. Each dimension is scaled. Assume intermediate fusion features... The Original values The current value is 50. Analysis of the historical training dataset reveals the minimum value for this feature dimension. It is 10, the maximum value. It is 70. According to the formula Normalized eigenvalues The calculation is as follows: right Perform this operation on all 32 dimensions, ultimately generating a causal-physicochemical fusion state descriptor with fixed dimensions and values ranging from [0,1]. .
[0054] In one possible implementation, combining Figure 2 The above S4 can be implemented through the following S41, S42 and S43, which are explained in detail below: S41. Input the causal-physicochemical fusion state descriptor into a preset dynamic evolution prediction model, calculate the values of several future time points of key performance indicators, and generate the prediction trajectory of key performance indicators. In some implementations, a causal-physicochemical fusion state descriptor is used as the current state input and fed into a pre-defined dynamic evolution prediction model. This model typically employs a network structure capable of processing and predicting time-series data, such as a Long Short-Term Memory (LSTM) network or a Gated Recurrent Unit (GRU). After receiving the current state descriptor, the model iteratively calculates forward using the temporal dynamics it has learned internally, generating state descriptors for one or more future time steps. From these predicted future state descriptors, components corresponding to key performance indicators (KPIs) such as product yield or purity are extracted. Connecting these values at consecutive future time points constitutes the KPI prediction trajectory. This trajectory depicts the expected evolution path of the system performance over a future period, given the current state.
[0055] For example, the 32-dimensional causal-physicochemical fusion state descriptor generated at the current moment As input, the data is fed into a pre-defined network model based on a gated recurrent unit (GRU). This GRU model has been trained on a large amount of historical data and is capable of capturing temporal dynamics. The model iteratively calculates the predicted state descriptor sequence for the next 10 time steps. From this sequence, the components corresponding to product yield and purity are extracted, generating two continuous prediction curves. and This constitutes the prediction trajectory of key performance indicators.
[0056] S42. Based on the causal barrier information embedded in the causal-physicochemical fusion state descriptor, predict the performance degradation trend under no-intervention conditions and generate the reverse target prediction trajectory. In some implementations, causal barrier information deeply embedded in the causal-physicochemical fusion state descriptor is used to predict the reverse target trajectory. This is based on the counterfactual assumption of "no intervention." The model focuses specifically on the current strength and state of negative causal chains leading to performance degradation, as identified by the causal barrier layer. The simulation demonstrates how these negative causal chains will continue to affect the system and drive key performance indicators gradually away from the ideal target without any active control intervention. The resulting trajectory, the reverse target prediction trajectory, quantifies the system's inherent performance degradation trend and reveals the potential consequences of "inaction."
[0057] For example, the model specifically focuses on Information about causal barriers, such as negative causal chains. intensity The model simulates a counterfactual scenario: assuming that no intervention is taken to stop the current situation. The negative impact, The yield will be driven by its current strength. The yield and purity continue to decline. Based on this no-intervention assumption, the model predicts the deterioration trend of yield and purity over the next 10 time steps. This trajectory is the expected path of "natural decline" under inherent defects, which is the reverse target prediction trajectory.
[0058] S43. Verify the consistency between the predicted trajectory of the key performance indicators and the predicted trajectory of the reverse target, and output the final prediction result.
[0059] In some implementations, consistency verification is performed on the generated key performance indicator prediction trajectory and the reverse target prediction trajectory. In the absence of external control input, these two trajectories should theoretically exhibit high consistency, both pointing towards the natural evolution direction of performance. Significant deviations between the two may indicate uncertainty in the model prediction or that they have been affected by unmodeled external disturbances. This cross-validation allows for the assessment of the reliability of the prediction results and the marking or correction of predictions with low confidence. After verification, the reliable prediction trajectory is output as the final prediction result, providing reliable future scenario information for subsequent construction of the reward function and formulation of control strategies.
[0060] For example, calculate the expected evolution and deterioration trend The mean squared error (MSE) between the two is used. If the MSE is below a preset threshold, such as 0.05, it is considered that the two have high consistency, indicating that the current evolution is mainly determined by the internal inherent dynamics and the prediction result is highly reliable. If the MSE is significantly higher than the threshold, such as 0.2, the prediction result is marked as having low confidence and an early warning is triggered. After verification, the predicted trajectory is determined by the key performance indicators. and reverse target predicted trajectory This will be output as the final prediction result.
[0061] In one possible implementation, combining Figure 2 The above S5 can be implemented through the following S51 and S52, which are explained in detail below: S51. Extract the performance drift index from the predicted trajectory of the reverse target and calculate the deviation from the ideal interval defined by the optimization objective function; In some implementations, a performance drift metric is extracted from the reverse target prediction trajectory. This trajectory depicts the trend of key performance indicators, such as product yield and purity, deteriorating over time under the influence of a negative causal chain revealed by a causal barrier layer, without intervention. Future predictions on this trajectory are used as the performance drift metric. Simultaneously, an optimized objective function is obtained to define the trade-off between product yield and purity. This function typically defines a multi-dimensional ideal range, such as requiring a product yield between 95% and 97% while maintaining a purity of at least 99%. The deviation of the performance drift metric from this ideal range is calculated; this deviation quantifies the potential performance loss caused by the inherent drift trend.
[0062] For example, obtaining the ideal interval: Suppose the optimization objective function defines the yield. ideal range for ,purity ideal range for Drift index extraction and deviation calculation: predicting trajectory from reverse target. Extract the yield drift index for the first future time step. Purity drift index Assuming This is below the lower limit of the ideal range. Deviation quantification: Calculating yield deviation For example, using a squared penalty. . This deviation The performance loss that will be suffered in the short term due to the inherent drift trend under no intervention conditions was quantified.
[0063] S52. Obtain the safety constraint function that defines the safe operating range, and combine it with the deviation to construct a multi-objective composite reward function that includes positive rewards and negative penalties, and generate a causal intervention reward function.
[0064] In some implementations, a safety constraint function is obtained that defines the safe operating range. This function sets inviolable upper and lower limits for key parameters in the reaction process, such as temperature and pressure, to ensure the safety of the entire synthesis process. Combining the calculated performance drift deviation and safety constraints, a multi-objective composite reward function containing positive rewards and negative penalties is constructed. This function is designed to guide the control policy network to generate actions that actively counteract performance drift while ensuring operational safety. Specifically, for the state... Next action The new state after transition Its reward function It can be defined as: ; in It is an immediate reward gained from the new state; It is an indicator function that takes the value 1 when the condition is true and 0 otherwise; This indicates that the key performance indicators of the new state fall within the ideal interval defined by the optimization objective function, at which point the weights are obtained. Positive rewards for decisions; This is the deviation of the key performance indicators from the ideal range under the new conditions, multiplied by the weight. This constitutes a penalty for deviating from the target; This indicates that one or more operating parameters of the new state exceed the range defined by the safety constraint function, and are therefore subject to the weights. The huge negative consequences of the decision. , and These are all preset non-negative weighting coefficients used to adjust the relative importance of different objectives. This is typically set to a very large value to ensure that safety is the highest priority. This composite reward function ultimately serves as the causal intervention reward function, used for subsequent reinforcement learning training.
[0065] For example, the safety constraint function is obtained by defining the safe operating range of the reactor pressure (M). for reaction temperature of Construction of the composite reward function: incorporating bias Design causal intervention reward function : Assume the preset weights are... , , Consider the action to be performed. Then, enter a new state. Its indicator is: yield ,pressure Ideal reward items: Not here Inside. Therefore. Deviation penalty items: Safety penalties: Exceeding the safe range .so Instant rewards calculate: Despite small performance deviations, the penalty term However, due to safety violations, they received a huge negative reward. This severely penalizes the action, ensuring that reinforcement learning strategies prioritize safety above all else.
[0066] In one possible implementation, combining Figure 2 The above-mentioned S6 can be implemented through the following S61, S62 and S63, which are explained in detail below: S61. Define a reinforcement learning environment with the causal-physicochemical fusion state descriptor as the state space and the adjustment range of the control parameters of the continuous flow reactor as the action space. In some implementations, a reinforcement learning environment is defined that can interact with deep reinforcement learning algorithms. The state space of this environment is defined as a causal-physicochemical fusion state descriptor, a standardized vector that comprehensively characterizes the reactor's physics, chemistry, and causal logic. The action space of the environment is defined as the adjustment range of controllable parameters of the continuous flow reactor, such as the feed pump flow rate, the temperature setpoint of the reaction section, and the pressure value of the back pressure valve. These parameters constitute the set of actions that the agent can execute. After the agent executes an action, that action is actually applied to the continuous flow reactor. After running for a short period, a new state is generated, and the corresponding reward is calculated based on the changes in system performance.
[0067] For example, state space Defined as a 32-dimensional causal-physicochemical fusion state descriptor Action space Defined as a continuous, controllable parameter adjustment. It includes three dimensions: feed pump flow rate. The adjustment range is Reaction section temperature The adjustment range of the set point is and back pressure valve pressure The adjustment range is The reinforcement learning environment was constructed as a high-fidelity continuous flow reactor simulation model to enable safe and rapid interaction with the agent during training, generating new states. and instant rewards .
[0068] S62. Using a multi-objective deep reinforcement learning algorithm, with the causal intervention reward function as the optimization objective, iteratively train the control policy network in a reinforcement learning environment. In some implementations, multi-objective deep reinforcement learning algorithms are used to iteratively train the control policy network within a defined reinforcement learning environment. Algorithms suitable for handling continuous action spaces, such as Soft Actor-Critic (SAC), can be selected. The training process is a loop. At each time step, the control policy network, acting as an agent, receives the causal-physical-chemical fusion state descriptor of the current environment as input and outputs specific control parameter adjustment actions. After this action is executed, it evolves to the next state and calculates a reward value based on the constructed causal intervention reward function. This reward value reflects the comprehensive contribution of the action to achieving the optimization objective, avoiding performance drift, and complying with safety constraints. The state, action, reward, and new state generated by this interaction constitute an experience sample, which is stored in the experience replay pool. The algorithm periodically randomly selects a batch of experience samples from the experience replay pool to update the internal parameters of the control policy network. The goal of this update is to maximize the expected value of the long-term accumulated causal intervention reward function.
[0069] For example, suppose the Soft Actor-Critic algorithm is chosen, which is a multi-objective deep reinforcement learning algorithm suitable for handling continuous action spaces and efficient exploration. The control policy network and the Critic network are initialized. The training process iterates for 1 million time steps. At each step, the control policy network receives the current state. Output control parameter adjustment action After the action is executed in the simulation environment, the reward value is calculated based on the causal intervention reward function. For example, for Interaction samples The data is stored in an experience replay pool, and a batch of samples is periodically drawn from it to update the parameters of the control policy network. The goal is to maximize the expected value of the long-term cumulative reward.
[0070] S63. When the control strategy network converges, output the control parameter adjustment action based on the current state.
[0071] In some implementations, the network is considered to have converged and the training process ends when the performance metrics of the control policy network, such as the average cumulative reward obtained during the testing period, stabilize and reach a preset threshold. At this point, the trained control policy network has learned the nonlinear mapping relationship between complex system states and optimal control actions. In actual operation, this network can instantly output a set of optimal control parameters to adjust actions based on the real-time input causal-physicochemical fusion state descriptor, thereby directly regulating the operation of the continuous flow reactor.
[0072] For example, when the average cumulative reward obtained during the testing period, i.e., the average reward over 100 consecutive test rounds, tends to stabilize and reaches a preset performance threshold, the control policy network is considered to have converged, and training ends. The trained control policy network is then deployed to the online optimization system. In actual operation, when a new... At that time, the network will output a set of optimal adjustment actions in real time, for example: , It is used to proactively intervene and maintain optimal operating conditions.
[0073] In one possible implementation, combining Figure 2 The above-mentioned S7 can be implemented through the following S71, S72 and S73, which are explained in detail below:
[0074] S71. Execute the control parameter adjustment action to modify the operating parameters of the continuous flow reactor; In some implementations, the trained and converged control policy network outputs specific control parameter adjustments based on the current input causal-physicochemical fusion state descriptor. These adjustments are typically a series of numerical values, such as raising the reaction temperature setpoint by 0.5 degrees Celsius or lowering the feed rate of feedstock A by 0.1 mL / min. These values are sent via a control interface to the underlying process control system of the continuous flow reactor, such as a PLC or DCS, which then drives the corresponding actuators, such as heaters and metering pumps, to precisely execute these modifications, thereby adjusting the operating parameters of the continuous flow reactor in real time.
[0075] For example, a converged control policy network outputs a control parameter adjustment action. For example, raising the reaction temperature setpoint. Reduce the feed rate of raw material A. The adjustment is sent via a control interface to the underlying distributed control system (DCS) of the continuous flow reactor. The DCS then drives the corresponding actuators, such as heaters and metering pumps, to precisely execute these modifications, changing the operating parameters from... Adjust to .
[0076] S72. Monitor changes in key performance indicators after execution and generate operational feedback data; In some implementations, process parameters and spectral data are continuously collected using sensors and online spectrometers. By analyzing the data over a period of time after the action is executed, the actual changes in key performance indicators, such as product yield and purity, can be observed. The state before the control action, the executed control action, the new state obtained after execution, and the causal intervention reward value calculated based on the changes in key performance indicators are packaged together into a complete set of operational feedback data.
[0077] For example, in the next control cycle after the adjustment action is performed, such as Minutes, continuously collecting new status. Actual changes in key performance indicators were monitored: product yield. from Upgraded to The state before execution. Actions to be performed The new state obtained and according to Calculated causal intervention reward value Together, they are packaged to generate a complete sample of runtime feedback data. .
[0078] S73. Update the dynamic causal relationship graph and control strategy network using the aforementioned operational feedback data.
[0079] In some implementations, this newly acquired operational feedback data is used to drive the dynamic updating and evolution of the entire AI optimization system. On one hand, this data is added to a historical database for periodically rerunning the causal discovery algorithm to examine and update the dynamic causal graph. If the new operational feedback data reveals previously undiscovered causal relationships or demonstrates a change in the strength of an existing causal path—for example, a decrease in catalyst activity leading to a weakening effect of temperature on yield—the structure of the dynamic causal graph or the weights of its edges are updated accordingly. On the other hand, this operational feedback data, as new empirical samples, is added to the experience replay pool of the reinforcement learning algorithm. The control policy network uses data including this new sample for incremental training or fine-tuning. In this way, the control policy network can continuously learn from the latest actual operational results and continuously optimize its decision-making logic. Figure 4 As shown, the product yield over time is compared between traditional PID control and the multi-objective deep reinforcement learning (MODRL) control of this invention. The figure visually demonstrates that MODRL control can maintain performance indicators within the target ideal range for a longer period and can quickly recover through an autonomous evolution mechanism when performance drift occurs.
[0080] For example, newly acquired runtime feedback data is added to the historical database. The causal discovery algorithm is periodically rerun. If new data reveals causal relationships previously underestimated by the model, such as finding that adjusting the flow rate has a greater impact on the yield than expected, the weights of the corresponding edges in the dynamic causal graph are updated accordingly. New runtime feedback data samples are then added. The data is added to the experience replay pool of the multi-objective deep reinforcement learning algorithm. The control policy network is then incrementally trained or fine-tuned. The network uses this high-reward sample to reinforce its tendency to perform actions such as increasing temperature and decreasing flow rate in similar states, thereby continuously optimizing its decision-making logic.
[0081] In one possible implementation, the multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals provided in this application embodiment further includes: activating an autonomous evolution mechanism when a new operating condition is detected, as detailed below:
[0082] Real-time comparison of the current causal-physicochemical fusion state descriptor with historical patterns to identify unlearned operating condition features; In some implementations, the causal-physicochemical fusion state descriptor generated at the current moment is compared in real time with a database storing historical patterns. This database contains cluster centers or distribution models of state descriptors corresponding to various typical operating conditions encountered during past stable operation. The novelty of the current operating condition is determined by calculating the distance or similarity between the current state descriptor and all known historical patterns, such as Mahalanobis distance. Once this distance exceeds a preset threshold, indicating that the current system state significantly deviates from all learned operating condition characteristics, a new operating condition is detected, and an autonomous evolution mechanism is triggered.
[0083] For example, the system generates a 32-dimensional causal-physicochemical fusion state descriptor in real time at the current moment. Comparison with a stored historical database. This database contains three typical operating condition patterns clustered during stable operation: "high activity - low flow rate," "medium activity - medium flow rate," and "low activity - high pressure." Through calculation... The Mahalanobis distance between the center of these three operating modes. Once the calculated minimum distance exceeds a preset threshold... ,For example That is, to determine the current state It deviated significantly from all learned historical patterns, indicating that a new operating condition has been detected.
[0084] When a new working condition is identified, an incremental learning process is triggered to update the definition of the causal barrier layer and the reward function; In some implementations, when a new operating condition is identified, an incremental learning process is immediately initiated, initially focusing on updating the understanding of causal relationships within the system. Process parameters and online spectral data collected under the new condition are used to re-examine and update the dynamic causal graph. This may lead to the discovery of entirely new, previously unseen negative causal chains, or significant changes in the influence strength of certain causal paths within existing causal barrier layers. For example, using a new batch of raw materials may introduce a trace impurity, which could establish a new causal pathway leading to increased byproducts. The definition of the causal barrier layer is automatically updated to include this new discovery. Simultaneously, the new operating condition may imply a change in optimization objectives. For instance, to handle new impurities, purity may need to be temporarily prioritized over yield. This triggers adjustments to the optimization objective function and safety constraint function, thereby reconstructing the causal intervention reward function to guide the agent to adapt to the new optimization trade-offs.
[0085] For example, under the new operating conditions, the newly collected data is used to rerun the causal discovery algorithm. This algorithm discovers a new negative causal chain: batch characteristics of raw material B – certain microscopic spectral characteristics. - Increased concentration of byproducts. The system automatically identifies this newly discovered causal chain and its quantified impact strength. Added to the causal barrier layer definition parameters. Due to the drastic increase in the difficulty of purity control caused by the new operating conditions, the weight coefficients in the causal intervention reward function have been adjusted: the weight of the purity reward term has been adjusted. improve To accommodate the new optimization trade-off, purity is temporarily prioritized over yield.
[0086] By retraining the control policy network online, it can adapt to new operating conditions and maintain optimized performance.
[0087] In some implementations, online retraining is used to adapt the control policy network to new operating conditions. This process utilizes new data collected under the new conditions, along with updated causal barrier layer definitions and newly constructed causal intervention reward functions, to perform additional training or fine-tuning on the existing control policy network. This online retraining allows the control policy network to quickly learn and master the optimal control logic under new conditions without completely discarding its original knowledge. Through this process, the agent learns how to respond to previously unseen system behaviors and outputs appropriate control parameter adjustments, thereby rapidly recovering and maintaining optimized performance under new operating conditions.
[0088] For example, an online retraining process is initiated. Using new data samples collected under the new operating conditions and the updated causal intervention reward function, the existing Soft Actor-Critic control policy network is incrementally fine-tuned. For instance, the network parameters are trained for an additional 10 epochs with a small learning rate. In this way, the control policy network quickly learns the decision logic under the new operating conditions, rapidly recovering and maintaining optimized performance without completely discarding existing knowledge.
[0089] It should be noted that the electrical connections between the various units described above do not necessarily represent direct or indirect connections. Any indirect connection method can be applied to the embodiments of the present invention as long as it achieves the purpose of the present invention. The above descriptions are merely exemplary embodiments of the present invention and should not be construed as limiting the scope of the present invention.
[0090] All equivalent changes and modifications made in accordance with the teachings of this invention are still within the scope of this invention. Those skilled in the art will readily conceive of other embodiments of this invention upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this invention that follow the general principles of this invention and include common knowledge or conventional techniques in the art not described herein.
Claims
1. A multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals, characterized in that, The method includes: Real-time process parameters and online spectral data of the continuous flow reactor are collected simultaneously, and causal characteristics are initially screened for the real-time process parameters and online spectral data to generate screened process parameter characteristics and spectral data characteristics. Based on the filtered process parameter features and spectral data features, an initial dynamic causal relationship graph is constructed, and a causal barrier layer is defined in the initial dynamic causal relationship graph to generate a dynamic causal relationship graph containing the causal barrier layer. By integrating the screened process parameter features, spectral data features, and topological features of the dynamic causal relationship graph, a causal-physicochemical fusion state descriptor is generated. Based on the causal-physicochemical fusion state descriptor, the future trajectory and reverse target trajectory of key performance indicators are predicted using a preset dynamic evolution prediction model, thereby generating the predicted trajectory of key performance indicators and the predicted trajectory of reverse targets. Obtain an optimized objective function to define the trade-off between product yield and purity, and construct a causal intervention reward function based on the optimized objective function and the reverse objective prediction trajectory; Using the causal-physicochemical fusion state descriptor as the state input and the causal intervention reward function as the optimization objective, a multi-objective deep reinforcement learning algorithm is used to train the control policy network and output the control parameter adjustment action. The operating parameters of the continuous flow reactor are adjusted according to the control parameters, and the dynamic causal relationship diagram and control strategy network are dynamically updated based on the operating feedback.
2. The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals according to claim 1, characterized in that, The simultaneous acquisition of real-time process parameters and online spectral data of the continuous flow reactor, and the initial screening of causal characteristics of the real-time process parameters and online spectral data, includes: Acquire real-time process parameters and online spectral data of a continuous flow reactor; The real-time process parameters are subjected to time-series feature extraction to generate a physical state feature vector; The online spectral data is subjected to dimensionality reduction processing to generate reaction kinetic feature vectors; The physical state feature vector and reaction kinetic feature vector are analyzed using a causal discovery algorithm to select a subset of features with causal driving relationships, and to generate the selected process parameter features and spectral data features.
3. The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals according to claim 1, characterized in that, Based on the filtered process parameter characteristics and spectral data characteristics, an initial dynamic causal relationship graph is constructed, and a causal barrier layer is defined in the initial dynamic causal relationship graph, including: Identify the causal paths between variables in the filtered process parameter features and spectral data features, and generate an initial dynamic causal relationship diagram; Extract the negative causal chains that cause performance drift from the initial dynamic causal relationship graph, quantify their influence strength, and generate causal barrier layer definition parameters; The parameters defining the causal barrier layer are embedded into the initial dynamic causal relationship graph, and the graph structure is dynamically updated to generate a dynamic causal relationship graph containing the causal barrier layer.
4. The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals according to claim 1, characterized in that, The topological features that integrate the filtered process parameter characteristics, spectral data characteristics, and dynamic causal relationship graph include: Extract the state information of causal barrier nodes and the strength information of causal paths from the dynamic causal relationship graph to generate a causal topological feature vector; The filtered process parameter features, spectral data features, and causal topological feature vectors are input into a multi-layer neural network for nonlinear combination to generate intermediate fusion features. The intermediate fusion features are normalized to generate a causal-physicochemical fusion state descriptor.
5. The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals according to claim 1, characterized in that, The method of using a pre-set dynamic evolution prediction model to predict the future trajectory and reverse target trajectory of key performance indicators includes: The causal-physicochemical fusion state descriptor is input into a preset dynamic evolution prediction model to calculate the values of several future time points of key performance indicators and generate the prediction trajectory of key performance indicators. Based on the causal barrier information embedded in the causal-physicochemical fusion state descriptor, the performance degradation trend under no-intervention conditions is predicted, and a reverse target prediction trajectory is generated. Verify the consistency between the predicted trajectory of the key performance indicators and the predicted trajectory of the reverse target, and output the final prediction result.
6. The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals according to claim 1, characterized in that, Based on the optimization objective function and the reverse target prediction trajectory, the causal intervention reward function is constructed as follows: The performance drift index is extracted from the predicted trajectory of the reverse target, and the deviation from the ideal interval defined by the optimization objective function is calculated. Obtain the safety constraint function that defines the safe operating range, and combine it with the deviation to construct a multi-objective composite reward function that includes positive rewards and negative penalties, thereby generating a causal intervention reward function.
7. The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals according to claim 1, characterized in that, Using the causal-physicochemical fusion state descriptor as the state input and the causal intervention reward function as the optimization objective, the control policy network is trained using a multi-objective deep reinforcement learning algorithm, including: Define a reinforcement learning environment with the causal-physicochemical fusion state descriptor as the state space and the adjustment range of the control parameters of the continuous flow reactor as the action space; Using a multi-objective deep reinforcement learning algorithm, with the causal intervention reward function as the optimization objective, the control policy network is iteratively trained in a reinforcement learning environment. When the control strategy network converges, it outputs control parameter adjustment actions based on the current state.
8. The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals according to claim 1, characterized in that, Adjusting the operating parameters of the continuous flow reactor according to the control parameters, and dynamically updating the dynamic causal relationship diagram and control strategy network based on operational feedback, includes: The control parameter adjustment action is executed to modify the operating parameters of the continuous flow reactor; Monitor changes in key performance indicators after execution and generate operational feedback data; The dynamic causal relationship graph and control strategy network are updated using the operational feedback data.
9. The multimodal AI optimization method for continuous flow reactors used in the synthesis of high-value chemicals according to claim 1, characterized in that, The method further includes: initiating an autonomous evolution mechanism when a new operating condition is detected, wherein: Real-time comparison of the current causal-physicochemical fusion state descriptor with historical patterns to identify unlearned operating condition features; When a new working condition is identified, the incremental learning process is triggered to update the definition of the causal barrier layer and the reward function; By retraining the control policy network online, it can adapt to new operating conditions and maintain optimized performance.
10. A multimodal AI optimization system for continuous flow reactors used in the synthesis of high-value chemicals, characterized in that, The system is used in the multimodal AI optimization method for continuous flow reactors for the synthesis of high-value chemicals as described in any one of claims 1-9, the system comprising: The data acquisition and causal feature screening module is used to simultaneously acquire real-time process parameters and online spectral data of the continuous flow reactor, and to perform causal feature screening on the real-time process parameters and online spectral data to generate screened process parameter features and spectral data features. The dynamic causal graph construction module is used to construct an initial dynamic causal relationship graph based on the filtered process parameter features and spectral data features, and to define a causal barrier layer in the initial dynamic causal relationship graph to generate a dynamic causal relationship graph containing the causal barrier layer. The state descriptor generation module is used to fuse the screened process parameter features, spectral data features, and topological structure features of the dynamic causal relationship graph to generate a causal-physicochemical fusion state descriptor. The trajectory prediction module is used to predict the future trajectory and reverse target trajectory of key performance indicators based on the causal-physicochemical fusion state descriptor and using a preset dynamic evolution prediction model, and to generate the predicted trajectory of key performance indicators and the predicted trajectory of reverse targets. The reward function construction module is used to obtain the optimized objective function for defining the trade-off between product yield and purity, and to construct the causal intervention reward function based on the optimized objective function and the reverse objective prediction trajectory. The reinforcement learning decision module is used to take the causal-physicochemical fusion state descriptor as the state input and the causal intervention reward function as the optimization objective, and to train the control policy network using a multi-objective deep reinforcement learning algorithm to output the control parameter adjustment action. The parameter adjustment and evolution module is used to adjust the operation parameters of the continuous flow reactor according to the control parameters, and dynamically update the dynamic causal relationship diagram and control strategy network based on the operation feedback.
Citation Information
Patent Citations
AI-driven chemical experiment simulation and result prediction system
CN120356568A
Reinforcement learning decision optimization method, system and equipment based on causal big language model
CN120911539A
Multivariable intelligent predictive control system for drying process of dried konjak rice
CN121050269A