AI-based clindamycin phosphate processing hydrolysis optimization method and system
By constructing a physical information digital twin module and a structural causal model, a multi-objective decision-making intelligent agent is generated, which solves the process bifurcation problem in the hydrolysis of clindamycin phosphate, achieves high yield and low impurity control of clindamycin phosphate, and adapts to the dynamic changes in the production process.
Patent Information
- Application Number
- CN202511164336.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Existing technologies cannot effectively address process bifurcation phenomena during the hydrolysis of clindamycin phosphate, especially dynamic disturbances in multivariable and nonlinear systems caused by fluctuations in raw material batches and changes in catalyst activity, which lead to unstable yields and impurity distributions.
By employing an AI-based approach, a multi-objective decision-making agent is generated by constructing a physical information digital twin module and a structural causal model. Combined with causal relationship analysis and dynamic safety constraints, process parameters are optimized to generate a Pareto optimal strategy set, thereby achieving proactive control of process bifurcation.
It enables precise identification and proactive control of process bifurcation in complex chemical reactions, ensuring stable product quality, improving the yield of clindamycin phosphate and impurity control capabilities, and adapting to dynamic changes in the production process.
Smart Images

Figure CN120673884B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of chemical pharmaceutical process control by specific computational models, and in particular to an AI-based clindamycin phosphate hydrolysis optimization method and system. BACKGROUND
[0002] Clindamycin phosphate is an important clinical lincosamide antibiotic, usually semi-synthesized from lincomycin. In its industrial production process, the hydrolysis step is a key link for converting intermediates into the final active pharmaceutical ingredient, and the reaction effect of this step directly determines the final yield and purity of the product. The core of quality control of pharmaceuticals lies in strict management of impurities. In the hydrolysis process of clindamycin phosphate, various process impurities will inevitably be generated, such as clindamycin-B-phosphate, 7-epi-clindamycin phosphate, etc. The content of these impurities is strictly limited by pharmacopoeias of various countries.
[0003] In the prior art, for example, Chinese patent document CN107652332B, the yield and purity are mainly improved by optimizing the chemical formula, such as using a specific hydrolysis agent. However, this kind of method provides a static, fixed process parameter window, and does not solve the problem of how to respond to dynamic disturbances in actual production, such as raw material batch fluctuations, catalyst activity decay over time, etc.
[0004] In complex chemical reaction processes, especially in such multivariate, nonlinear systems as clindamycin phosphate hydrolysis, there are some key, unstable critical regions. When small changes in process parameters, such as temperature, pH value, and reactant concentration, cause the system state to enter or cross these critical regions, the main path of the reaction will change qualitatively, resulting in significant, nonlinear differences in the properties of the final product, such as yield and impurity distribution. This phenomenon is referred to as "process bifurcation" by the present application.
[0005] Specifically, in the clindamycin phosphate hydrolysis process, at least the following several typical process bifurcation phenomena exist:
[0006] The first is yield-impurity bifurcation, i.e. in order to increase the generation rate of the main product, the reaction temperature is increased at a certain process state, which may cause the system to cross a bifurcation point, resulting in a disproportionate and sharp increase in the generation rate of a specific impurity, such as 7-epi-clindamycin phosphate;
[0007] The second is impurity spectrum bifurcation, i.e. adjusting the pH value may effectively inhibit the generation of impurity A, but at the same time may trigger another side reaction path, resulting in a significant increase in the concentration of impurity B;
[0008] The third is the kinetic-degradation bifurcation, that is, in the later stage of the reaction, in order to pursue higher conversion rate, the reaction time is prolonged or the temperature is maintained at a higher temperature, which may make the system cross a degradation bifurcation point, resulting in the generated target product clindamycin phosphate starting to decompose to generate new degradation impurities.
[0009] The positions of these process bifurcation points are not fixed, but will dynamically drift with the changes of hidden variables such as raw material batches and catalyst activity. The traditional control method such as PID (proportion-integral-derivative) controller cannot predict and manage the bifurcation phenomenon caused by the multi-variable coupling due to its single input and single output limitation. Even the advanced control strategy such as model predictive control (MPC) is also difficult to effectively deal with these dynamically drifting and nonlinear bifurcation problems because it relies on a fixed process model that usually cannot accurately describe the complex reaction kinetics. In recent years, AI technology has been tried to be applied to process optimization, but whether the prediction model based on supervised learning or the standard single-objective reinforcement learning can only learn the surface correlation in the data or can only optimize a single target, and cannot solve the multi-objective trade-off decision-making problem in front of multiple process bifurcation points. SUMMARY
[0010] To solve the above technical problems, one aspect of the present application provides a clindamycin phosphate processing hydrolysis optimization method based on AI, which is realized by a computer and includes the following steps:
[0011] Step S1, real-time and offline data of the reaction process are obtained, and the state is characterized. This step obtains data through a data acquisition layer, which includes temperature sensors, pH meters, mass flow meters and other conventional sensors, and process analytical technology (PAT) sensors for online real-time monitoring of the concentration of key components in the reaction system. At the same time, the offline analysis data such as final yield and impurity content of historical batches detected by high performance liquid chromatography (HPLC) and other methods are integrated. After the collected data is preprocessed, a high-dimensional state vector that can characterize the state of the reaction system at any time is formed.
[0012] Step S2, a reaction process prediction model capable of characterizing the process bifurcation phenomenon is constructed. The model can depict the whole picture of reaction kinetics including the process bifurcation point, and this step specifically includes:
[0013] A simulation module that combines chemical reaction mechanism and process data is constructed to generate a high-fidelity dynamic model that can predict the evolution trajectory of the system under different process conditions;
[0014] and constructing a causal relationship analysis module for identifying and quantifying causal transmission paths between process parameters and final product attributes from the data.
[0015] Step S3, based on the prediction model, training a multi-objective decision-making agent to generate a strategy set for actively controlling the process bifurcation. This step specifically includes:
[0016] Formalizing the optimization problem as a multi-objective decision-making process, where the high-dimensional state vector is defined as the process state, the adjustment amount of the adjustable process parameters is defined as the control operation, and multiple process performance indicators including yield, impurities, and energy consumption are defined as a multi-dimensional reward vector;
[0017] In the virtual environment constructed by the simulation module, the output of the causal relationship analysis module is used to guide the agent to perform exploratory training to improve learning efficiency;
[0018] During the training process, dynamic safety constraints based on known chemical reaction mechanisms are introduced to limit the range of available control operations for the agent, ensuring the safety and robustness of the generated strategy;
[0019] Through a multi-objective optimization algorithm, a strategy set covering the Pareto optimal frontier is finally generated, and each strategy in the set represents a specific trade-off between multiple conflicting objectives.
[0020] Step S4, online deployment and adaptive control. This step applies the trained model to the actual production process, specifically including:
[0021] The Pareto optimal strategy set is visualized through a human-machine interface for display to operators to select an operation strategy according to current production needs;
[0022] According to the selected strategy, real-time closed-loop control is performed in the production process, and after each production batch is completed, the reaction process prediction model is updated online using the full process data of the batch to adapt to changes in process characteristics.
[0023] Another aspect of the present application also provides an AI-based clindamycin phosphate processing hydrolysis optimization system, which includes: a data acquisition and execution unit for acquiring process data and executing control instructions; a central computing processing unit deployed with modules for executing the algorithms described in steps S2, S3 and S4 of the above method; and a visual display unit for human-machine interaction and strategy selection. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 is a structural block diagram of an AI-based clindamycin phosphate processing hydrolysis optimization system according to an embodiment of the present application.
[0025] Figure 2 is a flow chart of an AI-based clindamycin phosphate processing hydrolysis optimization method according to an embodiment of the present application.
[0026] Figure 3 is a schematic diagram for explaining the core concept of the present application, wherein, Figure 3 A schematically shows the process bifurcation phenomenon to be solved by the present application, Figure 3 B schematically shows the Pareto optimal strategy set generated by the method of the present application.
[0027] Figure 4 is a comparison chart of the prediction performance of the PI-DT model and the standard neural network model according to an embodiment of the present application.
[0028] Figure 5 is a process parameter causal relationship diagram learned by a structural causal model (SCM) according to an embodiment of the present application.
[0029] Figure 6 is a comparison chart of the exploration trajectories of the CI-MORL agent and the standard RL agent according to an embodiment of the present application.
[0030] Figure 7 is a chart of the Pareto optimal strategy set and its corresponding process control curve according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the protection scope of the present application.
[0032] Embodiment 1
[0033] The present embodiment provides an AI-based clindamycin phosphate processing hydrolysis optimization method, which is implemented by a computer, with reference to the flow chart shown in Figure 2 , and the detailed steps are as follows:
[0034] Step S1, acquire real-time and offline data of the reaction process and perform state characterization.
[0035] In the present embodiment, this step is implemented by an integrated data acquisition layer. The acquisition layer includes a plurality of sensors installed on the hydrolysis reactor, such as a Pt100 platinum resistance temperature sensor for measuring the temperature of the material in the reactor, an online pH meter for online monitoring of the pH value, and a mass flow meter for precise control of the addition speed of acid and alkali solutions. In particular, in order to obtain the concentration information of the key components in real time, the present embodiment adopts a process analysis technology (PAT) system, specifically a Raman spectrometer equipped with an immersion optical fiber probe, which has an excitation wavelength of 785 nm. The Raman spectrometer performs a spectral scan on the reaction solution every 60 seconds, and through a pre-established chemometric calibration model, the collected Raman spectrum is interpreted in real time into the concentrations of clindamycin phosphate, clindamycin-B-phosphate, 7-epi-clindamycin phosphate and other key substances. At the same time, the system database also integrates the accurate content data of the final product yield and each impurity obtained by offline analysis after the reaction is completed in the historical production batches through high performance liquid chromatography (HPLC). All collected data, including real-time sensor data and offline analysis data, are preprocessed before entering the subsequent model. The preprocessing process includes: first, using the moving average method to smooth the high-frequency noise data (such as temperature readings); second, aligning the data of different sampling frequencies (such as temperature data of 1 second and concentration data of 60 seconds) through time stamp; finally, performing Z-score normalization on all data dimensions to make the mean value 0 and the standard deviation 1, so as to eliminate the influence of different physical dimensions. After processing, at any time t, the system process state can be characterized as a high-dimensional state vector x(t), for example x(t) = [temperature(t), pH(t), acid solution dropping speed(t), alkali solution dropping speed(t), clindamycin phosphate concentration(t), impurity A concentration(t),...].
[0036] Step S2, constructing a reaction process prediction model capable of characterizing the process bifurcation phenomenon.
[0037] This step aims to establish a prediction model that can accurately depict the panorama of the clindamycin phosphate hydrolysis process. The model not only can predict the future process state of the reaction, but also can explain the internal causal mechanism, especially can identify and quantify the process bifurcation phenomenon. Referring to Figure 1 , the model is completed by the physical information digital twin module 101 and the structural causal model engine 102 working together.
[0038] Preferably, this step specifically includes:
[0039] Sub-step S201, constructing a physical information digital twin (PI-DT) module 101 for generating a reaction kinetics panorama model.
[0040] In this embodiment, the core of the module is a physics-informed neural network (PINN). The network is a deep feedforward neural network, e.g., containing 5 hidden layers with 128 neurons in each layer and a hyperbolic tangent (tanh) activation function. The training process of the PINN aims to minimize a hybrid loss function L_total = L_data + λ L_phys. Here, L_data is a data fitting loss term, calculated using mean squared error (MSE), to measure how well the network output (e.g., predicted concentrations, temperature) agrees with the historical and real-time data collected in step S1. L_phys is a physics law residual term, which embeds the core chemical reaction kinetics equations (including partial differential equations for mass conservation and energy conservation, e.g., dC / dt = f(C, T, pH) and dT / dt = g(T, Q)) of the clindamycin phosphate hydrolysis reaction into the loss function, where C represents the concentration of key substances, T represents the temperature, pH represents the acidity and alkalinity, and Q represents the heating power of the system. This term calculates to what extent the network output violates these physical laws. λ is a hyperparameter to balance the importance of data fitting and physical constraints, which can be set to 0.1, for example. By minimizing L_total using optimizers such as Adam, the trained PINN not only fits the existing data, but also follows the basic laws of chemical reactions in its internal structure. Therefore, the PI-DT module 101 can generate a high-fidelity digital twin, which can be used as a simulation environment for the reaction process to accurately simulate the dynamic evolution trajectory of the entire reaction process under the condition of any given process parameter sequence, including those regions that have not appeared in the historical data but are physically possible, thereby effectively identifying process bifurcation phenomena in the state space, such as Figure 3 A shows that the system state may evolve along different bifurcation paths after passing through a critical region. Referring to Figure 4 Compared to standard neural networks based only on data-driven, the PI-DT model shows significant prediction advantages in data sparse or extrapolation regions, and can give prediction results more consistent with physical and chemical laws, avoiding numerical unstable oscillation caused by overfitting of traditional models.
[0041] In another embodiment, the construction of the physics-informed digital twin module can be replaced by using a traditional first-principle model based on the Arrhenius formula or the like as the basis, and then training a residual neural network to learn and compensate for the deviation between the first-principle model and the actual data.
[0042] Sub-step S202, construct a structural causal model (SCM) engine 102 for identifying and quantifying bifurcation paths.
[0043] In this embodiment, the module aims to identify the causal relationships between process parameters and process outcomes from data. First, based on known chemical engineering knowledge, an initial causal graph skeleton is constructed. For example, the graph contains directed edges from "temperature" to "main reaction rate" and "impurity A formation rate." Then, using historical production data, the PC-stable causal discovery algorithm is employed to learn and refine this skeleton. This algorithm, through a series of conditional independence tests, can discover hidden causal relationships in the data that are not included in prior knowledge; for example, it may discover that "a specific property of a certain batch of raw materials" is a hidden common cause affecting "impurity B formation." Finally, a refined directed acyclic graph (DAG) is obtained, referring to... Figure 5 This diagram illustrates the causal transmission path between controllable variables, intermediate state variables, and the final goal. For each causal edge in the diagram, the system learns a structural equation. For example, for the concentration C_A of impurity A, its structural equation might be C_A = f(T,pH) + N_A, where f is a function (which can be represented by a small neural network or linear regression model), T and pH are its direct causes, and N_A is a noise term representing random perturbations. This SCM engine 102 is capable of causal inference; for example, it can quantify the direct causal effect of "raising the temperature setpoint by 0.5°C under the current conditions" on "final yield" and "7-epiclinamycin phosphate concentration".
[0044] Step S3: Based on the prediction model, train a multi-objective decision agent to generate a set of strategies for actively controlling the process bifurcation.
[0045] This step aims to train an agent capable of generating optimal timing control strategies, enabling it to learn to actively control in complex, multi-objective response environments.
[0046] Sub-step S301 formalizes the optimization problem into a multi-objective Markov decision process (MOMDP).
[0047] Specifically, the process state (State) S_t is the real-time state vector x(t) defined in step S1, and is supplemented by the PI-DT module 101, for example, supplemented with short-term prediction of future process states. The control operation (Action) A_t is defined as a discrete set of control instructions executable by the underlying programmable logic controller (PLC), for example, the set of control operations can be the Cartesian product of {temperature setpoint + 0.2℃, temperature setpoint - 0.2℃, keep unchanged} and {acid solution dropping speed + 0.5ml / min, acid solution dropping speed - 0.5ml / min, keep unchanged}. The process performance indicator (Reward) R_t is defined as a four-dimensional vector R_t = [r_yield, r_impurity1, r_impurity2, r_energy]. Among them, r_yield is positively correlated with the generation rate of clindamycin phosphate; r_impurity1 and r_impurity2 are penalty terms related to the concentration of key impurities, and a larger negative value is given when the impurity concentration exceeds the warning threshold; r_energy is a negative value related to the energy cost of heating power and stirring power.
[0048] Sub-step S302, perform causal guided exploration and training.
[0049] The training of the agent 103 is mainly offline in the PI-DT virtual environment constructed in step S201 to ensure safety and efficiency. In the exploration phase of training, the agent does not choose control operations completely randomly. On the contrary, before selecting a exploratory control operation, it will query the SCM engine 102. The SCM engine will evaluate the causal effects of each selectable control operation on multiple process performance indicator dimensions under the current process state. The agent will preferentially select those control operations that SCM predicts have strong positive causal effects on the desired target (such as yield) and weak negative causal effects on the undesired target (such as impurities). For example, if the SCM engine indicates that increasing the temperature in the current reaction stage mainly leads to the increase of yield, and has little effect on the generation of impurities, the agent will select the temperature-related control operation with a higher probability.
[0050] Further, to improve the robustness and safety of the system, dynamic taboo constraints based on reaction mechanisms are introduced when performing the trial and training of the causal guidance. In this embodiment, a dynamic taboo list generation module 104 is constructed. This module pre-encodes the dangerous operation rules based on chemical mechanisms, for example, rule 1: “when the pH value in the tank is lower than 4.5 and the temperature is higher than 55℃, the target product degradation risk is high”. During the training process, the PI-DT module 101 will evaluate the current simulation process state in real time. If it is detected that the current process state meets the conditions of rule 1, that is, the system state enters the bifurcation region that may lead to the degradation of the target product, the dynamic taboo list generation module 104 will immediately generate a taboo control operation list, for example, prohibiting all “warming” or “acid adding” related control operations in the next 15 minutes. When the agent 103 makes a decision, its set of optional control operations will be dynamically constrained by this list, so that it is forced to choose in the safe control operation subset. Referring to Figure 6 By comparing the exploration trajectories of the CI-MORL (Causal-Informed Multi-Objective Reinforcement Learning) agent and the standard reinforcement learning agent, it can be clearly seen that the agent of the present application can effectively avoid the preset dynamic taboo region, while the exploration of the standard agent is blind and may lead to entering a dangerous process state. This adds a strict safety constraint to the exploration process of the AI, ensuring that the strategy it generates is physically safe and feasible.
[0051] Alternatively, the step of introducing dynamic taboo constraints can also be achieved by: taking the degree of violation of safety rules as a negative process performance index dimension, and incorporating it into the multi-objective process performance index vector, so as to guide the agent to avoid the dangerous region in a soft constraint manner.
[0052] Sub-step S303, a set of Pareto optimal strategies is generated.
[0053] This embodiment adopts a known multi-objective reinforcement learning algorithm based on conditional networks (such as C-MORL). The strategy network trained by this algorithm not only takes the process state S as input, but also takes a preference vector w as additional input. The preference vector w represents the degree of emphasis on different objectives, for example, w = [0.7, 0.1, 0.1, 0.1] represents more emphasis on yield. By sampling various preference vectors w during training, the single strategy network finally trained can be generalized to any preference. Therefore, after training is completed, by inputting different w, a series of different optimal control strategies can be generated from the strategy network. These strategies constitute the Pareto optimal front, as shown in Figure 7the left subgraph. Each point (e.g. P1, P2, P3) in the graph represents a complete, end-to-end control strategy, corresponding to Figure 7 different forms of dynamic control curves in the right subgraph. For example, strategy P1 might represent a high-yield but slightly higher impurity solution, while strategy P3 represents a slightly lower-yield but extremely high-purity solution. No one strategy is superior to another in all objectives. This set of strategies provides multiple optimal choices for actual production.
[0054] Step S4, online deployment and adaptive control.
[0055] This step applies the trained model to the closed-loop control of the actual production process.
[0056] Sub-step S401, provide a visual decision-making interface and select a strategy.
[0057] The Pareto-optimal strategy set generated in step S303 is visualized through a human-machine interface (HMI) 105. Referring to Figure 7 , the interface displays the Pareto-optimal frontier (e.g. Figure 7 left subgraph) to the production supervisor or process operator. This interface is a two-dimensional scatter plot with the horizontal axis representing "total impurity content" and the vertical axis representing "final yield". The production supervisor or process operator can select a strategy point that best meets the needs of the current batch according to the specific business objectives (e.g. pursuing the fastest reaction speed to catch orders, or pursuing the highest purity to produce high-specification drugs), such as selecting P2 as a balance point. When a strategy point is selected, the system can further display the specific process control curve (e.g. Figure 1 right subgraph) corresponding to the strategy for the decision-maker to reference.
[0058] Sub-step S402, perform real-time closed-loop control and model adaptation.
[0059] Once the policy is selected, the system enters an automatic execution mode. Throughout the hydrolysis batch, the data acquisition layer continuously sends the real-time process state vector S_t to the central computing processing unit 100. The selected CI-MORL policy outputs the optimal control operation A_t according to the current process state S_t, which is sent to the underlying distributed control system (DCS) or programmable logic controller (PLC) through the control network to be executed accurately, for example, adjusting the steam valve opening of the heating jacket or the speed of the acid-base dropping pump. This "perception-decision-implementation" closed loop continues at a set time step (for example, every 60 seconds) until the reaction endpoint. When a production batch is completed, the full-process data of the batch (including the process state, control operation, and process performance indicators at all times) is automatically collected and used for incremental online updating of the PI-DT module 101 and the SCM engine 102. For example, the model parameters are fine-tuned using a small learning rate. This adaptive mechanism enables the system to continuously learn and effectively cope with the slow drift of process characteristics caused by factors such as catalyst activity decay over time and equipment aging, maintaining the effectiveness of the control policy.
[0060] Embodiment 2
[0061] The embodiment provides an AI-based clindamycin phosphate processing hydrolysis optimization system, and a structure thereof is shown in the accompanying Figure 1 The system is a physical carrier of the method in embodiment 1.
[0062] With reference to , the system comprises a data acquisition and execution unit 110, a central computing processing unit and a visual display unit 105.
[0063] The data acquisition and execution unit 110 is responsible for interaction with the physical production process. It comprises various sensors deployed on the reaction kettle, such as temperature sensors, pH meters, Raman spectrum probes, etc., for performing data acquisition in step S1. It also comprises various actuators, such as control valves connected to the heating system, frequency converters connected to the feeding pump, etc., for executing the control instructions issued by the central computing processing unit 100 in step S402.
[0064] The central computing processing unit 100 is the core of the system of the present application, and is usually a high-performance industrial computer or server. The unit internally deploys functional modules that implement the core algorithms of the present application, specifically including:
[0065] A physical information digital twin (PI-DT) module 101 for performing the functions described in step S201 to establish a high-fidelity process model.
[0066] A structural causal model (SCM) engine 102, which is used to perform the functions described in step S202, to conduct causal relationship mining and quantification.
[0067] A multi-objective decision-making agent 103, which integrates training and inference functions, is used to perform the strategy generation described in step S3 and the online control instruction generation in step S402.
[0068] A dynamic tabu list generation module 104, which is used to perform the safety constraint function described in step S302.
[0069] An online adaptive module, which is used to update the PI-DT module 101 and the SCM engine 102 with the new data collected after the batch ends, to realize the adaptive function described in step S402.
[0070] A visualization display unit 105, usually a touch screen or computer display, running human-machine interface (HMI) software. It is used to perform step S401 to display the Pareto optimal strategy set generated by the agent 103 to the operator in a graphical manner, and to receive the operator's strategy selection instructions.
[0071] In actual operation, the data acquisition and execution unit 110 continuously uploads process data to the central computing processing unit 100. The modules in the central computing processing unit 100 work together to calculate the optimal control instructions in real time according to the strategy selected by the operator on the visualization display unit 105, and then send them to the actuators in the data acquisition and execution unit 110 for operation, thus forming a complete and intelligent closed-loop control system, and realizing dynamic and multi-objective optimization of the clindamycin phosphate hydrolysis process.
[0072] Example 3
[0073] This example describes the implementation in a specific application scenario. The scenario is set as follows: the hydrolysis catalyst used has been operated for 45 batches; the lincomycin hydrochloride raw material used contains 0.15% of a certain precursor impurity. The control objective of this scenario is set as follows: the content of 7-epi-clindamycin phosphate in the final product is not higher than 1.0%, and the yield is maximized.
[0074] Before the batch was fed, the system first performed pre-analysis and strategy generation. The system input the off-line HPLC analysis data of the new raw material batch YL-20230815 and the record of the catalyst having been used for 45 times. Based on these new information, the PI-DT module 101 and the SCM engine 102 performed a quick online fine-tuning. Subsequently, the PI-DT module 101 simulated the standard operating procedure (SOP) in a virtual environment, and the simulation results clearly identified an offset “yield-impurity” bifurcation point: due to the presence of high content of precursor impurities and the decrease of catalyst activity, the side reaction path to generate 7-epiclinamycin phosphate was significantly activated within the temperature (52°C) and pH (4.8-5.2) control interval of the standard SOP, and the final content was predicted to reach 1.35%, which could not meet the quality requirements. At the same time, the SCM engine 102 quantified this causal effect and pointed out that under this specific condition, the positive causal effect value of temperature on the generation of this impurity increased by 40% compared to the normal case.
[0075] The system called the multi-objective decision-making agent 103 to generate a new set of Pareto optimal strategies for this specific batch, with “final 7-epiclinamycin phosphate content ≤0.9%” and “final yield maximization” as the main targets, and combined with secondary targets such as energy consumption. On the visual display unit 105, the process engineer saw that compared with the Pareto frontier of the conventional batch, the newly generated frontier was shifted to the lower left (low yield, high impurity) as a whole, which intuitively reflected the complexity of this production. The engineer selected a preset strategy point named EHP-45. Compared with the standard SOP, the yield prediction value of this strategy was slightly lower by 2.5%, but the impurity prediction value was only 0.88%.
[0076] After authorization, the system starts to execute EHP-45 to perform closed-loop control. The strategy has a high degree of dynamics: in the initial reaction (0-2 hours), the system strictly controls the reaction temperature at a lower 46°C, which is 6°C lower than the SOP, to inhibit the initial rate of epimerization, and at this time the analysis of the SCM engine shows that the negative causal effect of low temperature on the main reaction rate is not significant. In the middle of the reaction (2-5 hours), after the main product concentration reaches a critical value, the strategy starts to slowly raise the temperature at a rate of 0.1°C / 10 minutes, while precisely regulating the dropping speed of hydrochloric acid to maintain the pH value in a narrower range of 4.9±0.05 than the conventional, to seek a dynamic optimal balance point. At 4.5 hours, the Raman spectrometer detects that the concentration of an intermediate product in the system instantaneously rises, and the PI-DT module 101 predicts that the system process state is approaching a “kinetic-degradation” bifurcation point. The dynamic taboo list generation module 104 is triggered immediately, and within the next 30 minutes, all “temperature rising” and “increasing acid dropping speed” control operation options are temporarily prohibited, forcing the system to run in a safe area, effectively avoiding the degradation of the generated target product. In the late stage of the reaction, the strategy stabilizes the temperature at 50°C, which is slightly lower than the SOP, and extends the reaction time by about 90 minutes to ensure the conversion rate.
[0077] Finally, the batch reaction is completed in 9.5 hours (1.5 hours longer than the SOP). The end-point HPLC analysis shows that the final yield of clindamycin phosphate is 89.2% (the SOP is usually 91.5%-92.5%), and the content of the key impurity 7-epi-clindamycin phosphate is 0.92%, successfully controlled within the strict limit of 1.0%. The results of this example show that the present application can effectively respond to the dynamic changes of raw materials and catalyst states, accurately identify and actively control process bifurcations, generate and execute a dynamic, globally optimal, safe and reliable control strategy for specific production tasks under multiple constraints and conflicting objectives, and achieve precise control of the process.
[0078] The above merely describes preferred embodiments of the present application, but should not be used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. An artificial intelligence-based optimization method for the hydrolysis of clindamycin phosphate processing, characterized in that, This method is implemented by a computer and includes the following steps: Step S1: Acquire real-time and offline data of the reaction process and perform state characterization to construct a high-dimensional state vector that can characterize the state of the reaction system at any time. Step S2: Construct a reaction process prediction model that can characterize the process bifurcation phenomenon. This step specifically includes constructing a simulation module that integrates chemical reaction mechanism and process data, as well as a causal relationship analysis module. Step S3: Based on the prediction model, train a multi-objective decision agent to generate a set of strategies covering the Pareto optimal frontier for actively controlling the process bifurcation, wherein the training process combines: guidance using the output of the causal analysis module; and dynamic safety constraints based on known chemical reaction mechanisms. Step S4: Online deployment and adaptive control. This step includes visualizing the set of strategies covering the Pareto optimal frontier for selection, executing real-time closed-loop control according to the selected strategy, and updating the prediction model online using the batch data after the production batch is completed.
2. The method according to claim 1, characterized in that, The simulation module in step S2 is specifically a physical information digital twin module, the core of which is a physical information neural network PINN. The loss function of the physical information neural network includes a data fitting loss term and a physical law residual term based on the partial differential equation of chemical reaction kinetics.
3. The method according to claim 1, characterized in that, The causal relationship analysis module in step S2 is specifically a structural causal model (SCM) engine. Its construction steps include: constructing an initial causal graph skeleton based on prior knowledge, refining it using historical data and causal discovery algorithms, and learning structural equations for the causal edges in the refined graph.
4. The method according to claim 1, characterized in that, The step S3, which uses the output of the causal relationship analysis module to guide the agent in training, specifically involves the following steps: During the exploration phase of training, before selecting a control operation, the agent queries the causal relationship analysis module to evaluate the causal effect of each optional operation on multiple process performance indicators, and prioritizes selecting a control operation that has a stronger positive causal effect on the preset desired process performance indicators and a weaker negative causal effect on the preset undesired process performance indicators.
5. The method according to claim 1 or 4, characterized in that, The step of introducing dynamic safety constraints in step S3 is specifically implemented through at least one of the following methods: Method 1: Construct a dynamic taboo list generation module. Based on preset dangerous operation rules, this module generates a set of taboo control operations to constrain the agent's optional operation set when the simulation module evaluates that the current simulation state has entered a high-risk area during the training process. Method 2: The degree of violation of the dangerous operation rules is used as a negative process performance indicator dimension and incorporated into the multi-objective reward to guide the agent to avoid high-risk areas in a soft constraint manner.
6. The method according to claim 1, characterized in that, Before training the multi-objective decision agent in step S3, the process further includes formalizing the optimization problem into a multi-objective Markov decision process, which specifically includes: The high-dimensional state vector is defined as the process state; Define the set of instructions for adjusting adjustable process parameters as control operations; and A process performance index is defined as a multidimensional vector that includes factors related to product formation rate, critical impurity concentration, and energy consumption.
7. The method according to claim 1, characterized in that, The step of performing real-time closed-loop control in step S4 specifically includes: during the production process, the intelligent agent outputs the optimal control operation based on the high-dimensional state vector acquired in real time, and executes it through the control system to form a continuous perception-decision-execution closed loop.
8. The method according to claim 1, characterized in that, The step of updating the prediction model online in step S4 specifically includes: after a production batch is completed, using the full process data of that batch, incrementally updating the model parameters of the simulation module and the causal relationship analysis module online.
9. The method according to claim 1, characterized in that, The real-time data acquisition in step S1 is achieved through the process analysis technology PAT sensor, specifically a Raman spectrometer or a near-infrared spectrometer.
10. An artificial intelligence-based optimization system for the hydrolysis of clindamycin phosphate processing, characterized in that, include: The data acquisition and execution unit is used to acquire data on the hydrolysis process of clindamycin phosphate and execute control commands. A visualization display unit is used to run human-computer interaction interface software; as well as A central computing processing unit, connected to the data acquisition and execution unit and the visualization display unit, is configured to execute the method as described in any one of claims 1 to 9. Specifically, the central computing processing unit deploys: A simulation module; A causal relationship analysis module; A multi-objective decision-making intelligent agent; The central computing processing unit is configured to: graphically display the Pareto optimal strategy set generated by the multi-objective decision-making agent on the visualization display unit and receive strategy selection instructions from the operator; then generate real-time control instructions according to the selected strategy; and perform closed-loop control of the hydrolysis process through the data acquisition and execution unit.
Citation Information
Patent Citations
A method for preparing clindamycin phosphate
CN107652332B
Digital-twin-driven fault prognosis method and system for subsea production system of offshore oil
AU2020102863A4
Method and system for assisting production decision of methanol-to-olefin process
CN116434867A