Process control method and system for production of pharmaceutical intermediates

By establishing a quantitative correlation model and dynamically adjusting the parameters of downstream reaction units, the impact of trace impurities on subsequent reactions in the production of pharmaceutical intermediates was resolved. This achieved stability and consistency in the quality of pharmaceutical intermediate products, reduced the total amount of impurities, and improved the stability of the production process and product quality.

CN120688783BActive Publication Date: 2026-03-17XINYI DAJIANG CHEM IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing pharmaceutical intermediate production processes struggle to quantify the differences in trace impurities during upstream separation and purification processes and their impact on subsequent reactions, leading to unstable final product quality. Current control methods are unable to effectively predict and compensate for these impacts.

Method used

By acquiring trace component information of intermediate samples, a batch fingerprint spectral dataset is established, and a quantitative correlation model is constructed to predict the amount of impurities generated in downstream reaction units. Operating parameters are then dynamically adjusted to compensate for upstream fluctuations, achieving forward-looking regulation.

Benefits of technology

It significantly reduces the total amount of impurities in the final target pharmaceutical intermediate products, improves the batch qualification rate and quality stability of the products, and reduces rework and scrap, which has important economic and technical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688783B_ABST
    Figure CN120688783B_ABST
Patent Text Reader

Abstract

The application provides a medical intermediate production process regulation method and system, and relates to the technical field of medical production. The method comprises the following steps: obtaining intermediate trace component information and forming a batch fingerprint data set; establishing a quantitative correlation model; inputting the intermediate trace component information of the intermediate sample of the current batch into the quantitative correlation model to predict the generation amount of specific impurities in the downstream reaction unit; calculating the adjustment amount of the operation parameters of the downstream reaction unit according to the predicted specific impurity generation amount and the set total impurity control target; and sending the adjustment amount to the actuator of the downstream reaction unit to adjust the actual operation parameters of the downstream reaction unit. The method of the application can analyze the intermediate trace impurities of the current batch, predict the generation trend of downstream impurities, and adjust the downstream operation parameters in advance, so that the adverse effects caused by upstream fluctuations can be effectively compensated, and the stability of the quality of the final product can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pharmaceutical manufacturing technology, and more specifically, to a method and system for controlling the production process of pharmaceutical intermediates. Background Technology

[0002] The production process of pharmaceutical intermediates typically employs a multi-step reaction and multi-unit operation in series. In this model, starting materials are transformed into intermediates in a single reactor step. The intermediate mixture then enters downstream separation and purification units, such as extraction, washing, crystallization, and filtration, to effectively remove byproducts and unreacted substances generated during the reaction, thereby obtaining intermediate products that meet certain purity requirements. This intermediate product is then used as a raw material in the next reaction step, undergoing similar reaction, separation, and purification processes to finally obtain the target pharmaceutical intermediate. The ultimate goal of the entire production chain is to ensure that the total impurities in the final product are controlled within strictly defined limits to meet drug quality and safety requirements.

[0003] However, in actual industrial production, even minor fluctuations in upstream unit operating parameters, such as deviations in temperature or pressure control within the reactor, or subtle differences between batches of starting materials, can lead to slight batch-to-batch variations in the relative content and types of components, particularly various byproducts, in the one-step reaction product mixture. This batch-to-batch variability in the reaction product mixture then enters the downstream separation and purification unit. The operational performance of the separation and purification unit, such as the yield, crystal form, and separation effectiveness of the target product during crystallization, is highly sensitive to process parameters such as feed composition, supersaturation, cooling rate, and stirring intensity. Due to batch-to-batch variations in the composition of the upstream reaction products, even if the operating parameters of the downstream separation and purification unit remain constant, the actual separation process may still exhibit subtle batch-to-batch differences. For example, the presence of certain trace impurities can significantly affect the crystallization habit of the target product, making these impurities more easily trapped within the crystals, or altering the concentration distribution of impurities in the mother liquor, thereby affecting the purity of the final intermediate product. This fluctuation in separation performance caused by differences in feed composition means that the types, contents, and distribution of trace impurities in the intermediate products obtained after separation and purification vary from batch to batch, even if the contents of the main components may meet the standards.

[0004] These intermediate products, containing trace impurities, are subsequently used as feedstocks in the next reaction step. It is important to note that these trace impurities entering subsequent reactors are not inert substances. They may undergo self-transformation, participate in undesirable side reactions, or adversely affect the catalyst in the reaction system under specific temperature, pressure, and catalyst presence conditions in subsequent reactions. For example, a trace metal impurity introduced from upstream may act as a catalyst under subsequent reaction conditions, accelerating specific degradation pathways; or an organic impurity may undergo undesirable coupling reactions with subsequent reactants, generating new, structurally complex impurities; or an impurity may adsorb onto the surface of the solid catalyst used in subsequent reactions, reducing its activity and selectivity, leading to a decrease in the main reaction conversion rate and an increase in side reactions. Because the upstream separation and purification unit fails to completely and stably remove these trace impurities, batch-to-batch fluctuations in the quality of feedstocks entering subsequent reactors occur. These fluctuations are often amplified during subsequent reactions, ultimately affecting the quality of the target product. Existing process control methods typically focus on determining batch compliance through offline analysis of the final product or on feedback control based on online monitoring data (such as reaction temperature, pressure, and main product concentration) to maintain the stability of the main parameters of the current reaction step. However, these methods struggle to capture and quantify the differences in trace impurities in intermediates caused by upstream separation and purification processes, and even more so, they cannot predict the impact of these trace impurities on the formation of multiple impurities in subsequent complex reaction systems, nor can they make proactive or compensatory parameter adjustments accordingly. Therefore, although each unit operation may operate under set parameters, the implicit fluctuations in intermediate quality and their cumulative effects in subsequent reactions remain key reasons for the instability of the total impurities in the final product and the challenge to batch compliance. There is an urgent need for a technical means to quantify trace impurities in intermediates and predict their impact on subsequent reactions, thereby enabling compensatory control. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for controlling the production process of pharmaceutical intermediates. By analyzing the trace impurities in the current batch of intermediates, the generation trend of downstream impurities can be predicted, and downstream operating parameters can be adjusted in advance, thereby effectively compensating for the adverse effects of upstream fluctuations and ensuring the stability of the final product quality.

[0006] In a first aspect, the present invention provides a method for controlling the production process of pharmaceutical intermediates, comprising the following steps:

[0007] Obtain information on trace components of intermediate samples produced by the upstream separation and purification unit, and form a batch fingerprint spectral data set;

[0008] Establish a quantitative correlation model between batch fingerprint datasets and the amount of specific impurities generated in downstream reaction units;

[0009] The information on the trace components of intermediates in the current batch of intermediate samples is input into a quantitative correlation model to predict the amount of specific impurities generated in downstream reaction units.

[0010] Based on the predicted amount of specific impurities generated and the set total impurity control target, calculate the adjustment amount of the operating parameters of the downstream reaction unit;

[0011] The adjustment amount is sent to the actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit.

[0012] The pharmaceutical intermediate production process control method provided by this invention addresses batch-to-batch fluctuations in the types and contents of trace impurities in intermediate products produced by upstream separation and purification units during the tandem production of pharmaceutical intermediates. This is achieved by quantitatively analyzing these trace impurity information before the intermediates enter the downstream reaction unit, establishing a cross-unit quantitative correlation model between this model and the generation of specific impurities during the downstream reaction process, and dynamically predicting the downstream impurity generation trend based on the trace impurity analysis results of the current batch of intermediates using this correlation model. This allows for proactive adjustment of key process parameters in the downstream reaction unit to compensate for batch-to-batch differences in the quality of upstream raw materials, thereby reducing the total amount of impurities in the final target pharmaceutical intermediate product and improving batch stability.

[0013] Secondly, the present invention provides a pharmaceutical intermediate production process control system, comprising:

[0014] The acquisition module is used to acquire information on trace components of intermediate samples produced by the upstream separation and purification unit and form a batch fingerprint data set.

[0015] A module is established to create a quantitative correlation model between batch fingerprint datasets and the amount of specific impurities generated in downstream reaction units.

[0016] The prediction module is used to input the intermediate trace component information of the current batch of intermediate samples into the quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit.

[0017] The calculation module is used to calculate the adjustment amount of the operating parameters of the downstream reaction unit based on the predicted amount of specific impurity generation and the set total impurity control target.

[0018] The control module is used to send adjustment amounts to the actuators of downstream reaction units to adjust the actual operating parameters of the downstream reaction units.

[0019] As can be seen from the above, the pharmaceutical intermediate production process control method provided by this invention can quantitatively analyze the types and contents of trace impurities in intermediates produced by upstream separation and purification units, and accurately capture batch-to-batch quality differences in intermediate raw materials. By establishing a quantitative correlation model between the fingerprint spectrum of trace components in intermediates and the generation of impurities in subsequent reactions, cross-unit, intermediate quality-based impurity generation prediction has been achieved for the first time. Based on this prediction model, the system can proactively calculate and dynamically adjust key process parameters of subsequent reaction units, effectively compensating for the impact of batch fluctuations in upstream raw material quality on downstream reactions. This significantly reduces the total amount of impurities in the final target pharmaceutical intermediate product, improves the batch pass rate and quality stability of the product, reduces rework and scrap, and has significant economic and technical value.

[0020] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description

[0021] Figure 1 This is a flowchart of a pharmaceutical intermediate production process control method provided in an embodiment of the present invention.

[0022] Figure 2 This is a schematic diagram of a pharmaceutical intermediate production process control system provided in an embodiment of the present invention.

[0023] Label Explanation:

[0024] 100 Acquisition Module; 200 Establishment Module; 300 Prediction Module; 400 Calculation Module; 500 Control Module. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0026] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0027] Reference Appendix Figure 1 This invention provides a method for controlling the production process of pharmaceutical intermediates, comprising the following steps:

[0028] Obtain information on trace components of intermediate samples produced by the upstream separation and purification unit, and form a batch fingerprint spectral data set;

[0029] Establish a quantitative correlation model between batch fingerprint datasets and the amount of specific impurities generated in downstream reaction units;

[0030] The information on the trace components of intermediates in the current batch of intermediate samples is input into a quantitative correlation model to predict the amount of specific impurities generated in downstream reaction units.

[0031] Based on the predicted amount of specific impurities generated and the set total impurity control target, calculate the adjustment amount of the operating parameters of the downstream reaction unit;

[0032] The adjustment amount is sent to the actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit.

[0033] Intermediate samples from the upstream separation and purification unit are acquired and analyzed to obtain trace component information. This information is collected and organized into a batch fingerprint dataset to characterize the quality properties of different batches of intermediates. Furthermore, based on historical production data, a quantitative correlation model is established to describe the relationship between the trace component fingerprint dataset and the amount of specific impurities generated in the downstream reaction unit. Therefore, when a new batch of intermediate sample is available, its trace component information is input into the established quantitative correlation model, which outputs a predicted value for the amount of specific impurities generated in the downstream reaction unit. Based on the difference between this predicted value and a preset total impurity control target, the required adjustment amount for the operating parameters of the downstream reaction unit is calculated. Finally, this calculated adjustment amount is sent to the actuator in the downstream reaction unit, which changes the actual operating parameters of the downstream reaction unit, such as reaction temperature, reaction time, or catalyst dosage, according to the received adjustment amount.

[0034] Specifically, this method monitors the quality of intermediates entering downstream reaction units, quantifies their trace component characteristics, and establishes a quantitative relationship model between intermediate trace components and downstream impurity formation using historical data. Before a new batch of intermediates enters the downstream reaction unit, trace component analysis is performed, and the results are input into the model to predict the amount of specific impurities that this batch of intermediates may generate in the downstream reaction. If the predicted impurity generation exceeds the set control target, the required adjustments to downstream operating parameters are calculated based on the predicted values, the control target, and the influence of operating parameters on impurity generation. These adjustments are then applied to the downstream reaction unit to compensate for the impact of upstream intermediate quality fluctuations on downstream impurity generation, thereby controlling the total impurities in the final product within the target range. This method achieves proactive regulation based on upstream intermediate quality, improving the stability of the production process and the consistency of product quality.

[0035] The working principle of this invention is to construct a closed-loop quality prediction and compensation control system across units. First, by performing detailed trace component analysis on the upstream purified intermediates, batch-to-batch variations are quantified, forming a unique "quality fingerprint." Then, using historical production data, a mathematical model is established between this intermediate quality fingerprint and the amount of specific impurities generated during downstream reactions, revealing the influence of trace impurities on subsequent reactions. For the current production batch, the system acquires its quality fingerprint and inputs it into the model to predict its performance in the downstream reaction. Based on the prediction results, the system calculates the optimal adjustments needed for downstream reaction units (such as reaction temperature and time), and implements these adjustments through an automated control system. In this way, before or during the downstream reaction, proactive or compensatory intervention is provided for the quality fluctuations of upstream raw materials, thereby stabilizing and reducing the total impurities in the final product. The entire system continuously improves the accuracy of prediction and control through continuous data accumulation and model optimization.

[0036] In some specific implementations, it is assumed that the downstream reaction unit of the pharmaceutical intermediate production process is a catalytic reactor, and the specific impurity is a byproduct generated in this reaction. Intermediate samples from the upstream separation and purification unit are analyzed by high-performance liquid chromatography (HPLC) to obtain the relative content data of five trace impurities (A, B, C, D, and E), forming a batch fingerprint dataset. For example, the fingerprint data of a certain batch of intermediates is: impurity A 0.12%, impurity B 0.05%, impurity C 0.08%, impurity D 0.03%, and impurity E 0.06%. Based on the fingerprint data of the past 100 batches of intermediates and the actual generation data of the specific impurity in the corresponding downstream reaction unit, a partial least squares regression model is established. When a new batch of intermediates is analyzed, its trace component data (e.g., 0.12%, 0.05%, 0.08%, 0.03%, and 0.06% as mentioned above) are input into the model. The model predicts that this batch will generate 0.7% of the specific impurity in the downstream reaction. The target for the total amount of the specific impurity is set to not exceed 0.5%. Based on the predicted value of 0.7% and the target value of 0.5%, the required parameter adjustments are calculated. For example, according to the model or preset rules, it is determined that lowering the downstream reaction temperature by 2°C and extending the reaction time by 15 minutes can reduce the formation of a specific impurity by approximately 0.2%. The calculated adjustments (temperature -2°C, time +15 minutes) are sent to the reactor's temperature controller and timer actuator to adjust the actual operating parameters.

[0037] In some embodiments, the step of establishing a quantitative correlation model between the batch fingerprint dataset and the amount of specific impurities generated in downstream reaction units includes:

[0038] A1. Obtain historical production data; historical production data includes a set of fingerprint data of trace components of intermediates from multiple batches and the amount of specific impurities generated in the downstream reaction units corresponding to each batch of intermediates;

[0039] A2. Preprocess historical production data; specific preprocessing steps include:

[0040] A21. Principal component analysis was used to reduce the dimensionality of the fingerprint data of trace components of intermediates and extract the principal component scores that characterize the quality fluctuation of intermediates.

[0041] A22. Normalize the data on the amount of specific impurities generated in downstream reaction units to eliminate the influence of dimensions;

[0042] A3. Based on preprocessed historical production data, a quantitative correlation model is established using partial least squares regression algorithm to establish the relationship between the principal components of intermediate trace components and the amount of specific impurities generated in downstream reaction units; the quantitative correlation model includes regression coefficients to characterize the influence of different trace impurity components on the generation of downstream impurities;

[0043] A4. Using cross-validation, a quantitative correlation model is used to predict the amount of specific impurities generated in the downstream reaction units of historical batches, and the prediction error between the predicted value and the actual value is calculated to determine whether the prediction error exceeds the first preset threshold.

[0044] A5. If the prediction error exceeds the first preset threshold, return to step A3, adjust the parameters of the partial least squares regression algorithm, and rebuild and validate the quantitative correlation model; otherwise, use the current quantitative correlation model as the final model.

[0045] Specifically, this technical solution aims to improve the precision of impurity control by establishing an accurate quantitative correlation model through a data-driven approach when the reaction mechanism is unknown. First, a dataset is constructed by collecting fingerprint data of intermediate trace components from multiple historical production batches and the corresponding downstream reaction unit specific impurity generation amounts. Then, the collected data is preprocessed, including using principal component analysis to reduce the dimensionality of the high-dimensional intermediate trace component data, extracting a few principal components that represent the quality fluctuations of the intermediate, and normalizing the downstream impurity generation data to eliminate dimensional differences between different impurity generation amounts. Next, using the preprocessed principal component data and the normalized impurity generation data, a partial least squares regression algorithm is used to establish a mathematical model that quantifies the relationship between intermediate trace components (represented by principal components) and the downstream specific impurity generation amounts. This model includes regression coefficients that reflect the degree of influence of different trace components on downstream impurity generation. After establishing the model, cross-validation is used to evaluate its predictive performance. Historical data is divided into training and validation sets. The training set is used to train the model, and the validation set is used to evaluate the model's prediction error. The difference between predicted and actual values ​​(e.g., root mean square error) is calculated to determine if the model's prediction accuracy meets requirements. If the prediction error exceeds a first preset threshold, it indicates insufficient predictive ability, requiring a return to the model building steps to adjust the parameters of the partial least squares regression algorithm (e.g., number of principal components, model complexity), retrain the model, and validate it again. This process is repeated until the model's prediction error falls below the first preset threshold; at this point, the model is considered the final quantitative correlation model. This final model can predict the generation of specific downstream impurities based on interstitial trace component information, providing a basis for subsequent process parameter adjustments. Through this data-driven modeling and iterative optimization process, even when the reaction mechanism is not fully understood, an effective prediction model can be established, thereby achieving the prediction and control of downstream impurity generation.

[0046] In some specific implementations, the process of establishing a quantitative correlation model can be achieved as follows: Collect 100 interstitial samples from historical production batches, perform trace component analysis on each sample to obtain fingerprint data containing the contents of 50 different trace components. Simultaneously, record the content of a specific impurity A generated after the reaction of these 100 batches of interstitial substances in the downstream reaction unit. Use these 100 sets of 50-dimensional trace component data and the corresponding impurity A content data as historical production data. Apply principal component analysis to the trace component data to calculate the variance contribution rate of each component, determining that the top 10 principal components can explain 95% of the total variance. Project the 50-dimensional data onto these 10 principal components to obtain 100 sets of 10-dimensional principal component score data. Perform min-max normalization on the impurity A content data, scaling its numerical range to [0, 1]. Use these 100 sets of 10-dimensional principal component score data as independent variables and the normalized impurity A content as the dependent variable, and establish a model using partial least squares regression. For example, set the number of latent variables in the partial least squares model to 5. After building the model, a 5-fold cross-validation method was used to evaluate it. The 100 datasets were randomly divided into 5 parts, with one part used as the validation set and the remaining 4 parts as the training set. The training set was used to train the model, and the validation set was used to predict and calculate the prediction error (e.g., root mean square error). This process was repeated 5 times, and the average of the 5 errors was calculated. If the average root mean square error was greater than 0.05, the parameters of the partial least squares model were adjusted, for example, by increasing the number of latent variables to 6, and the model was rebuilt and cross-validated. This process was repeated until the average root mean square error was less than or equal to 0.05; at this point, the model was considered the final model. This model can predict the generation amount of downstream impurity A based on the principal component information of new intermediate batches.

[0047] In some embodiments, the specific steps in step A21 include:

[0048] A211. Calculate the variance contribution rate of each intermediate trace component in the fingerprint data of intermediate trace components, and identify at least two intermediate trace components with variance contribution rates greater than a second preset threshold as key components that have a significant impact on the quality fluctuation of the intermediate (key components are not necessarily impurities; they refer to components that have a significant impact on the quality fluctuation of the intermediate. These components can be impurities or the main components in the intermediate).

[0049] A212. Using a pre-defined principal component analysis model, calculate the covariance matrix of the key components, and obtain the eigenvectors of each principal component by solving the covariance matrix;

[0050] A213. Based on the feature vectors of each principal component, the fingerprint data of the corresponding intermediate trace components or all key components are projected from the high-dimensional space to the low-dimensional space composed of all principal components to obtain the principal component score that can characterize the quality fluctuation of the corresponding batch of intermediate samples. The principal components contain all key components (here, the principal component can be understood as a comprehensive index composed of multiple trace components (including impurities), not the necessary components for the efficacy of drugs in the traditional sense). All principal components constitute a low-dimensional representation of the quality fluctuation of the corresponding batch of intermediate samples (that is, converting the high-dimensional intermediate trace component data or all key components into a low-dimensional principal component representation). The principal component score is used as the input of the subsequent quantitative correlation model.

[0051] Calculating the variance contribution rate of each trace component in the intermediate trace component fingerprint data involves applying standard statistical methods to historical data to calculate the proportion of each trace component in the total variance. A second preset threshold is used to screen components with high variance contribution rates. Components with variance contribution rates greater than the second preset threshold are identified as key components, ensuring that subsequent analysis focuses on key factors affecting intermediate quality fluctuations and avoiding the neglect of trace components that significantly impact downstream reactions. Constructing the covariance matrix of key components involves using the values ​​of the identified key components across all historical batches to calculate the covariance between any two key components, forming a matrix. Solving the covariance matrix to obtain eigenvectors involves applying linear algebraic algorithms, such as the Jacobian iteration method, to find the eigenvalues ​​and corresponding eigenvectors of the matrix. These eigenvectors indicate the main directions of variation in the key component data. Based on the eigenvectors of each principal component, the intermediate trace component fingerprint data or key components are projected into a low-dimensional space by matrix multiplication, transforming the original data points into a new coordinate system defined by the eigenvectors. Thus, the data for each batch is represented as a set of principal component scores, which comprehensively reflect the characteristics of key components in different directions of variation, thereby reducing the data dimensionality.

[0052] Specifically, to address the issue that directly performing principal component analysis on all trace components might overlook certain crucial information, this approach first calculates the variance contribution rate of each component in the historical batches of intermediate trace component data. Based on a second preset threshold, trace components with variance contribution rates greater than this threshold are selected, identifying them as key components with a significant impact on intermediate quality fluctuations. This selection step ensures that subsequent analysis focuses on the key factors affecting intermediate quality fluctuations, improving the relevance of the analysis. Based on the selected key component data, a covariance matrix of the key components is constructed, describing the relationships and variability among them. By solving the covariance matrix, eigenvectors of each principal component are obtained; these eigenvectors represent the main directions of variation in the key component data. Finally, based on these eigenvectors, the corresponding intermediate trace component fingerprint data or all key components are projected from the original high-dimensional space to a low-dimensional space composed of the selected principal components. Thus, the quality fluctuation of each batch of intermediate samples is represented as a set of principal component scores, which serve as inputs for subsequently establishing a quantitative correlation model. This process reduces the dimensionality of the data while preserving the main variation information, improving the computational efficiency and generalization ability of the model, and providing effective input for accurately predicting the generation of specific impurities downstream.

[0053] For example:

[0054] Suppose that three key intermediate trace components (component X, component Y, and component Z) were identified through preliminary analysis. After processing historical production data, the content data of these key components in multiple batches were obtained. By calculating the covariance matrix and performing eigenvalue decomposition, two principal components (PC1 and PC2) and their corresponding eigenvectors (v1 and v2) were obtained. Eigenvectors v1 and v2 define the two main directions of data variation. For a specific batch of intermediate sample, the content data of its key components is [x, y, z]. This data point is projected onto the PC1 and PC2 axes, and its principal component score is calculated. Specifically, the score (Score1) of this batch on PC1 can be obtained by multiplying the data vector [x, y, z] by the eigenvector v1, i.e., Score1 = x * v1_x + y * v1_y + z * v1_z. Similarly, the score (Score2) of this batch on PC2 can be obtained by multiplying the data vector [x, y, z] by the feature vector v2, i.e., Score2 = x * v2_x + y * v2_y + z * v2_z. Thus, the original three-dimensional data points [x, y, z] are transformed into a two-dimensional principal component score vector [Score1, Score2]. This two-dimensional vector [Score1, Score2] represents the principal component score of the quality fluctuation of the intermediate sample in this batch. These principal component scores serve as input variables for subsequently building a partial least squares regression model to predict the generation of specific impurities in downstream reaction units. By using low-dimensional principal component scores instead of the original high-dimensional trace component content, the complexity of the model is reduced while retaining the main variation information of the data, thus improving the predictive efficiency and accuracy of the model.

[0055] In some specific implementations, intermediate sample data from 150 historical production batches are considered, with the content of 20 trace components measured in each sample. First, the variance contribution rate of these 20 trace components in the 150 batches of data is calculated. A second preset threshold of 8% is set. The analysis results show that the variance contribution rate of 6 trace components exceeds 8%, and these 6 components are identified as key components. Next, a 6x6 covariance matrix is ​​calculated using the content data of these 6 key components in the 150 batches. Eigenvalue decomposition is performed on this covariance matrix to obtain 6 eigenvalues ​​and corresponding eigenvectors. Based on the eigenvalue size, the eigenvectors corresponding to the top 3 eigenvalues ​​with a cumulative variance contribution rate of 90% are selected as the eigenvectors of the principal components. Finally, the key components (6-dimensional data) of the 150 batches are projected into a 3-dimensional space composed of these 3 eigenvectors to obtain the 3 principal component scores for each batch. These 3-dimensional principal component score vectors serve as the input for subsequent partial least squares regression models to predict downstream impurity generation. Thus, the original 20-dimensional or 6-dimensional data is effectively reduced to 3-dimensionality, simplifying the model structure while retaining key quality fluctuation information.

[0056] In some embodiments, the specific steps in step A211 include:

[0057] Stratified sampling was used to divide historical production data into multiple time periods based on the time distribution of intermediate production batches.

[0058] For the data within each time period, after calculating the variance contribution rate of each intermediate trace component, the box plot method is used to identify and remove outliers of the variance contribution rate to obtain the corrected variance contribution rate.

[0059] The corrected variance contribution rates of data from all time periods are summarized, the average variance contribution rate of each intermediate trace component is calculated, and based on the second preset threshold, the intermediate trace components with an average variance contribution rate greater than the threshold are identified as key components that have a significant impact on the quality fluctuation of the intermediate.

[0060] Historical production data is divided into multiple time periods based on the production batch time of the intermediates. This takes into account potential changes in the production process over time. For the data within each time period, the variance contribution rate of each trace component of the intermediate is calculated. Furthermore, box plots are used to identify and remove outliers in the variance contribution rate, resulting in a corrected variance contribution rate. Removing outliers reduces the impact of extreme data on the calculation results, making the variance contribution rate calculation more representative. The corrected variance contribution rates from all time periods are summarized, and the average variance contribution rate of each trace component of the intermediate is calculated. Based on a second preset threshold, trace components of the intermediate with an average variance contribution rate greater than the second preset threshold are identified as key components affecting the quality fluctuation of the intermediate. These components exhibit high fluctuation contributions across different time periods.

[0061] Specifically, to address the fluctuations in the fingerprint data of intermediate trace components and the impact of outliers on the stability of directly calculated variance contribution rates, this scheme employs stratified sampling, dividing historical production data into time periods based on the production batch time of the intermediate. This considers potential changes in the production process over time. For the data within each time period, the variance contribution rate of each intermediate trace component is calculated. Furthermore, box plots are used to identify and remove outliers in the variance contribution rate, yielding a corrected variance contribution rate. By removing outliers, the impact of extreme data on the calculation results is reduced, making the variance contribution rate calculation more representative. The corrected variance contribution rates of data from all time periods are summarized, and the average variance contribution rate of each intermediate trace component is calculated. Based on a second preset threshold, intermediate trace components with an average variance contribution rate greater than the second preset threshold are identified as key components affecting the quality fluctuations of the intermediate. These components exhibit high fluctuation contributions across different time periods. Through the above steps, this scheme takes into account the time characteristics of production data and the impact of outliers, identifies components that have a sustained or major impact on the quality fluctuations of intermediates, provides more reliable input for subsequent principal component analysis, and thus improves the accuracy of the quantitative correlation model.

[0062] In some embodiments, the specific steps in step A212 include:

[0063] Based on the identified key components, the covariance matrix of the key components is constructed by calculating the covariance between any two key components.

[0064] The covariance matrix is ​​decomposed into eigenvalues ​​using the Jacobi iteration method to obtain multiple eigenvalues ​​and corresponding eigenvectors. All eigenvectors are then sorted in descending order of eigenvalues.

[0065] Select the eigenvectors corresponding to the first N eigenvalues ​​as the eigenvectors of the principal components, and ensure that the cumulative variance contribution rate of the key components corresponding to the first N eigenvalues ​​reaches a preset ratio, where N is the minimum value when the cumulative variance contribution rate reaches the preset ratio.

[0066] Based on the identified key components, the covariance between any two key components is calculated, thus constructing a covariance matrix. This matrix is ​​a symmetric matrix, where the diagonal elements represent the variance of each key component itself, and the off-diagonal elements represent the covariance between any two key components, reflecting the degree of their common variation. The Jacobi iteration method is used to perform eigenvalue decomposition on the constructed covariance matrix. The Jacobi iteration method is a numerical algorithm used to solve for the eigenvalues ​​and eigenvectors of a symmetric matrix. This method yields multiple eigenvalues ​​and a set of orthogonal eigenvectors corresponding to the covariance matrix. The obtained eigenvectors are sorted according to the magnitude of their corresponding eigenvalues; the larger the eigenvalue, the more original data variance is explained by the principal component represented by the corresponding eigenvector. The eigenvectors corresponding to the top N sorted eigenvalues ​​are selected as the eigenvectors of the principal components. N is determined based on the cumulative variance contribution rate, i.e., the proportion of the sum of the variances corresponding to the top N eigenvalues ​​to the sum of all eigenvalues ​​reaches or exceeds a third preset threshold. N is taken as the minimum value satisfying this condition to achieve data dimensionality reduction while retaining the main information.

[0067] Specifically, this technical solution aims to extract principal components that can effectively characterize the quality fluctuations of intermediates from trace component data, addressing the issues of constructing the covariance matrix, performing eigenvalue decomposition, and determining the number of principal components. First, based on previously identified key components that significantly influence intermediate quality fluctuations, the content data of these key components in historical batches are collected. Based on this data, the covariance between any two key components is calculated, and a covariance matrix for each key component is constructed. This matrix quantifies the interrelationships and variability among the key components. Next, the Jacobi iteration method is used to perform eigenvalue decomposition on the covariance matrix. The Jacobi iteration method gradually diagonalizes the covariance matrix through a series of rotation transformations; the elements on the diagonal are the eigenvalues, and the column vectors of the rotated matrix are the corresponding eigenvectors. The eigenvalue decomposition results reveal the main directions (eigenvectors) and their extent (eigenvalues) of data variation. The obtained eigenvectors are sorted from largest to smallest according to their corresponding eigenvalues, ensuring that the direction explaining the most variance is placed first. Then, the proportion of each eigenvalue to the sum of all eigenvalues ​​is calculated, i.e., the variance contribution rate of that principal component. The variance contribution rates after sorting are accumulated until the cumulative proportion reaches a third preset threshold (e.g., 95% or 98%). The top N eigenvectors, representing the minimum number required to reach this cumulative proportion, are selected as the final principal component directions. These principal components are linear combinations of the original key components; they are mutually orthogonal and contain most of the variation information from the original data. In this way, high-dimensional key component data is effectively mapped to a low-dimensional principal component space. The extracted principal components accurately capture the main fluctuation patterns of intermediate quality, providing concise and information-rich data input for subsequent quantitative correlation model building, thus improving the model's accuracy and robustness.

[0068] In some specific implementations, it is assumed that three key components—impurity A, impurity B, and impurity C—have been identified through preliminary analysis. Content data for these three impurities from 100 historical batches are collected. First, a 3x3 covariance matrix is ​​calculated for these 100 batches. For example, the covariance between impurity A and impurity B, between impurity A and impurity C, between impurity B and impurity C, and the variances of impurities A, B, and C are calculated. These values ​​are then filled into the covariance matrix. Next, the Jacobi iterative algorithm is used to perform eigenvalue decomposition on the 3x3 covariance matrix, yielding three eigenvalues ​​λ1, λ2, and λ3 and three corresponding eigenvectors v1, v2, and v3. Assume the calculated eigenvalues ​​are λ1 = 10.5, λ2 = 2.1, and λ3 = 0.4. The eigenvectors are sorted in descending order of their eigenvalues, resulting in the sequence v1, v2, v3. Calculate the variance contribution rate of each eigenvalue: vCr1 = 10.5 / (10.5+2.1+0.4) ≈ 0.81, vcr2 = 2.1 / (10.5+2.1+0.4) ≈ 0.16, vcr3 = 0.4 / (10.5+2.1+0.4) ≈ 0.03. Calculate the cumulative variance contribution rate: cumulative vcr1 = 0.81, cumulative vcr1 + vcr2 = 0.81 + 0.16 = 0.97, cumulative vcr1 + vcr2 + vcr3 = 0.97 + 0.03 = 1.00. Set the preset cumulative variance contribution rate threshold to 95% (i.e., the third preset threshold). Since cumulative vcr1 + vcr2 = 0.97 > 0.95, and this is the minimum N value (N = 2) to reach the third preset threshold, the first two eigenvectors v1 and v2 are selected as the eigenvectors of the principal components. This means that the original three-dimensional key component data can be reduced to two dimensions, represented by these two principal components, while retaining approximately 97% of the original data variation information. These extracted principal components serve as input to subsequent quantitative correlation models, effectively reducing model complexity while ensuring accurate characterization of intermediate quality fluctuations.

[0069] In some embodiments, the specific steps in step A4 include:

[0070] A41. Divide historical production data into K non-overlapping subsets, each subset containing intermediate data from different batches;

[0071] A42. Select the i-th subset as the validation set, and the remaining K-1 subsets as the training set, where the value of i ranges from 1 to K;

[0072] A43. Based on the training set, the amount of specific impurities generated in each batch of downstream reaction units in the validation set is predicted using a quantitative correlation model, and a set of predicted values ​​is obtained.

[0073] A44. Based on the actual impurity generation amount corresponding to each batch in the predicted value set and the validation set, the root mean square error algorithm is used to calculate the prediction error of the i-th subset;

[0074] A45. Repeat steps A42 to A44 to calculate the prediction errors corresponding to the K subsets;

[0075] A46. Calculate the average of the K prediction errors as the final prediction error for cross-validation;

[0076] A47. Determine whether the final prediction error exceeds the first preset threshold.

[0077] Historical production data is divided into K non-overlapping subsets, each configured to contain intermediate data from different production batches. This ensures the representativeness of the data distribution in each subset. The system is then configured to sequentially select one subset as the validation set, while combining the remaining K-1 subsets as the training set. This process is repeated K times, selecting a different subset as the validation set each time. Based on the currently selected training set, a pre-established quantitative correlation model is used to process the intermediate data from each batch in the currently selected validation set, thereby predicting the generation amount of specific impurities in the downstream reaction unit and forming a set of predicted values. Next, based on this set of predicted values ​​and the actual impurity generation amounts corresponding to each batch in the validation set, the root mean square error (RMSE) algorithm is used to calculate the prediction error of the current validation subset. The RMSE algorithm is configured to quantify the degree of deviation between the predicted and actual values. The above steps are repeated to calculate the prediction errors for all K subsets. Finally, the average of these K prediction errors is calculated, and this average is determined as the final prediction error for cross-validation, used to comprehensively evaluate the overall prediction performance of the model. Finally, the final prediction error is compared with the first preset threshold to determine whether the model's prediction accuracy meets the requirements.

[0078] Specifically, this solution aims to address the problem that when using a quantitative correlation model built with partial least squares regression to predict the generation of specific impurities in downstream reaction units, unreasonable data subset partitioning during cross-validation leads to unrepresentative validation results and an inability to accurately assess the model's generalization ability. By dividing historical production data into K non-overlapping subsets, ensuring each subset contains intermediate data from different batches, the problem of unreasonable data partitioning is solved, guaranteeing the representativeness of the cross-validation data base. By alternately selecting different subsets as the validation set and using the remaining subsets as the training set, and repeating this process K times, a comprehensive evaluation of the model's generalization ability is achieved, avoiding situations where the model performs well only on specific data subsets. Predictions are made on the validation set based on the training set, and the root mean square error algorithm is used to calculate the prediction error, providing a quantitative assessment of the model's prediction accuracy. Repeated calculations and averaging reduce the randomness of single validation results and improve the reliability of model evaluation. Finally, by determining whether the average prediction error exceeds a first preset threshold, a clear criterion is provided for model optimization and selection, ensuring that the model's prediction accuracy meets the needs of practical applications. This improves the reliability and accuracy of quantitative correlation model evaluation, providing a solid foundation for subsequent process control based on model prediction results.

[0079] In some embodiments, the step of calculating the adjustment amount of the downstream reaction unit operating parameters based on the predicted specific impurity generation amount and the set total impurity control target includes:

[0080] Determine the type of operating parameters for the downstream reaction unit, including reaction temperature, reaction time, and catalyst dosage.

[0081] Based on the predicted amount of specific impurities generated, and using a pre-established correlation function between the amount of specific impurities generated and each type of operating parameter, the degree of influence of each type of operating parameter on impurity generation is obtained and expressed as an adjustment coefficient.

[0082] Based on the set total impurity control target, and combined with the adjustment coefficients of each type of operating parameter, calculate the adjustment amount for each type of operating parameter. The adjustment amount must ensure that the final total impurity does not exceed the control target, and that the adjustment range of each operating parameter is within the preset range, with priority given to adjusting the operating parameters that have a greater impact on impurity generation.

[0083] This method identifies the types of parameters that can be controlled in downstream reaction units, such as reaction temperature, reaction time, and catalyst dosage. Further, using predicted information on the amount of specific impurities generated, combined with a pre-established correlation function, the method quantifies the influence of each type of operating parameter on the generation of specific impurities, and expresses this quantification as an adjustment coefficient. The correlation function reflects the quantitative relationship between the amount of specific impurities generated and different operating parameters. Therefore, based on the set total impurity control target and the calculated adjustment coefficients for each type of operating parameter, the method calculates the specific adjustment amount for each operating parameter. This calculation process considers several constraints: the adjusted parameters should ensure that the final total impurity does not exceed the set control target; the adjustment range of each operating parameter must be within a preset safe or feasible range; and, provided the aforementioned conditions are met, priority is given to adjusting those operating parameters that have a greater impact on impurity generation. By determining specific parameter types, quantifying the influence of each parameter, and comprehensively considering control targets, adjustment ranges, and influence priorities to calculate the adjustment amount, this method transforms predicted impurity information into executable process adjustment instructions.

[0084] Specifically, this method identifies the controllable operating parameters related to impurity formation in downstream reaction units, thus clarifying the objects of control. For example, in a downstream reaction unit, reaction temperature, reaction time, and catalyst dosage are identified as the main operating parameters affecting the formation of specific impurities. Next, the total amount of specific impurities that may be generated in the current batch in the downstream reaction unit is predicted using trace component information from upstream intermediates and compared with the set control target for the total amount of impurities in the final product. Based on a quantitative correlation function established in advance using experimental data or historical production data, such as a multiple linear regression model or a nonlinear model, this method calculates the degree of influence of changing unit reaction temperature, unit reaction time, or unit catalyst dosage on the amount of specific impurities generated. These degrees of influence are the adjustment coefficients for each type of operating parameter. For example, the correlation function may indicate that increasing the reaction temperature significantly increases the formation of specific impurities, while extending the reaction time or increasing the catalyst dosage has a smaller effect or an inhibitory effect. Then, based on the gap between the predicted amount of impurities generated and the control target, combined with the adjustment coefficients for each type of operating parameter, this method calculates the specific adjustments required to the reaction temperature, reaction time, and catalyst dosage. This calculation process is an optimization problem. It requires ensuring that the adjustment range of each parameter does not exceed its preset maximum adjustment range, while maintaining the total final impurity amount within the control target. Furthermore, it prioritizes adjusting the operating parameter with the largest absolute value of its adjustment coefficient, achieving the maximum impurity control effect with the smallest parameter adjustment range. Therefore, this method transforms predicted impurity information into specific process parameter adjustment instructions, compensating for the impact of upstream intermediate quality fluctuations on downstream impurity generation, and achieving effective control over the total impurities in the final product.

[0085] In some specific implementations, it is assumed that a batch of intermediates is predicted to generate a specific impurity of 0.5% in the downstream reaction unit, while the set final total impurity control target requires that this specific impurity not exceed 0.3%. The operating parameters of the downstream reaction unit are determined to be reaction temperature, reaction time, and catalyst dosage. A pre-established correlation function indicates that for every 1°C increase in reaction temperature, the specific impurity increases by 0.1%; for every 10 minutes extension of reaction time, the specific impurity increases by 0.05%; and for every 1% increase in catalyst dosage, the specific impurity decreases by 0.08%. The preset adjustment ranges for each parameter are: reaction temperature ±3°C, reaction time ±30 minutes, and catalyst dosage ±5%. The specific impurity generation needs to be reduced from 0.5% to 0.3%, i.e., a reduction of 0.2%. According to the correlation function, reaction temperature has the greatest impact on impurity generation (the largest absolute value of the adjustment coefficient). Therefore, adjusting the reaction temperature is prioritized. Lowering the reaction temperature by 1°C reduces the impurity by 0.1%. A further reduction of 0.1% is needed. Next, the next most influential parameter is considered, such as the catalyst dosage. Increasing the catalyst dosage by 1% reduces the impurity by 0.08%. If the catalyst dosage is increased by another 1.25% (total increase of 2.25%), impurities can be reduced by 0.1%. At this point, the total reduction in impurities is 0.1% + 0.1% = 0.2%, achieving the control target. The calculated adjustments are: a 1°C decrease in reaction temperature, unchanged reaction time, and a 2.25% increase in catalyst dosage. These adjustments are all within the preset range (-1°C within ±3°C, 0 minutes within ±30 minutes, +2.25% within ±5%). Therefore, based on the predicted impurity generation and control target, the specific downstream operating parameter adjustments were calculated.

[0086] Reference Appendix Figure 2 This invention provides a pharmaceutical intermediate production process control system, comprising:

[0087] The acquisition module 100 is used to acquire information on trace components of intermediate samples produced by the upstream separation and purification unit and form a batch fingerprint data set.

[0088] Module 200 is established to create a quantitative correlation model between batch fingerprint data sets and the amount of specific impurities generated in downstream reaction units.

[0089] The prediction module 300 is used to input the intermediate trace component information of the current batch of intermediate samples into the quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit.

[0090] The calculation module 400 is used to calculate the adjustment amount of the operating parameters of the downstream reaction unit based on the predicted amount of specific impurity generation and the set total impurity control target.

[0091] The control module 500 is used to send adjustment amounts to the actuators of the downstream reaction units to adjust the actual operating parameters of the downstream reaction units.

[0092] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.

[0093] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method of process control for the production of pharmaceutical intermediates, characterized by, The method comprises the following steps: obtaining intermediate trace component information of an intermediate sample produced by an upstream separation and purification unit, and forming a batch fingerprint data set; establishing a quantitative correlation model between the batch fingerprint data set and the generation amount of a specific impurity in a downstream reaction unit; inputting the intermediate trace component information of the intermediate sample of the current batch into the quantitative correlation model to predict the generation amount of the specific impurity in the downstream reaction unit; calculating the adjustment amount of the operating parameters of the downstream reaction unit according to the predicted generation amount of the specific impurity and the set total amount control target of the impurity; sending the adjustment amount to an actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit; The step of establishing a quantitative correlation model between the batch fingerprint data set and the generation amount of a specific impurity in a downstream reaction unit comprises: A1. Obtain historical production data; the historical production data comprises a plurality of batch intermediate trace component fingerprint data sets and the generation amount of a specific impurity in a downstream reaction unit corresponding to each batch intermediate; A2. Preprocess the historical production data, and the specific steps comprise: A21. Perform dimensionality reduction on the intermediate trace component fingerprint data by principal component analysis, and extract principal component scores representing intermediate quality fluctuations, and the specific steps comprise: A211. Calculate the variance contribution rate of each intermediate trace component in the intermediate trace component fingerprint data, and determine at least two intermediate trace components with a variance contribution rate greater than a second preset threshold as key components that have a significant impact on intermediate quality fluctuations; A212. Calculate the covariance matrix of the key components, and obtain the eigenvectors of each principal component by solving the covariance matrix; A213. According to the eigenvectors of each principal component, project the corresponding intermediate trace component fingerprint data or all key components from a high-dimensional space into a low-dimensional space composed of all principal components to obtain principal component scores that can represent the quality fluctuations of the corresponding batch intermediate sample; the principal components include all key components; A22. Normalize the generation amount data of the specific impurity in the downstream reaction unit to eliminate the influence of dimensions; A3. Based on the preprocessed historical production data, a quantitative correlation model is established by using a partial least squares regression algorithm; the quantitative correlation model comprises regression coefficients for representing the influence degree of different trace impurity components on downstream impurity generation; A4. Using the cross-validation method, the quantitative correlation model is used to predict the generation amount of the specific impurity in the downstream reaction unit of the historical batch, and the prediction error between the predicted value and the actual value is calculated to determine whether the prediction error exceeds a first preset threshold; A5. If the prediction error exceeds the first preset threshold, return to step A3, adjust the parameters of the partial least squares regression algorithm, and re-establish the quantitative correlation model and verify it; otherwise, the current quantitative correlation model is used as the final model.

2. The process control method for pharmaceutical intermediate production according to claim 1, wherein, The specific steps in step A211 comprise: using stratified sampling method, dividing the historical production data into data of multiple time periods according to the time distribution of the intermediate production batches; For each time period, the variance contribution rate of each intermediate trace component is calculated, and the box plot method is used to identify and remove outliers of the variance contribution rate, to obtain the corrected variance contribution rate; The corrected variance contribution rates of all time periods are summarized, the average variance contribution rate of each intermediate trace component is calculated, and the intermediate trace components with an average variance contribution rate greater than a second preset threshold are identified as key components that have a significant impact on the intermediate quality fluctuation.

3. The process control method for pharmaceutical intermediate production according to claim 1, wherein, The specific steps in step A212 include: According to the determined key components, the covariance between any two key components is calculated to construct a covariance matrix of the key components; The covariance matrix is subjected to eigenvalue decomposition by the Jacobi iteration method to obtain a plurality of eigenvalues and corresponding eigenvectors, and all eigenvectors are sorted in descending order of eigenvalues; The eigenvectors corresponding to the first N eigenvalues are selected as the eigenvectors of the principal components, and the cumulative variance contribution rate of the key components corresponding to the first N eigenvalues is ensured to reach a preset proportion, N being the minimum value when the cumulative variance contribution rate reaches the preset proportion.

4. The process control method for pharmaceutical intermediate production according to claim 1, wherein, The specific steps in step A4 include: A41. The historical production data is divided into K non-overlapping subsets, each containing data of different batches of intermediates; A42. The i-th subset is selected as the validation set, and the remaining K-1 subsets are selected as the training set, where i ranges from 1 to K; A43. Based on the training set, the quantitative correlation model is used to predict the amount of specific impurities generated in the downstream reaction unit of each batch in the validation set, to obtain a set of predicted values; A44. According to the set of predicted values and the actual amount of impurities generated in each batch in the validation set, the root mean square error algorithm is used to calculate the prediction error of the i-th subset; A45. Steps A42 to A44 are repeated to calculate the prediction errors of the K subsets; A46. The average of the K prediction errors is calculated as the final prediction error of cross-validation; A47. Determine whether the final prediction error exceeds a first preset threshold.

5. The process control method for pharmaceutical intermediate production according to claim 1, wherein, The steps of calculating the adjustment amount of the operation parameters of the downstream reaction unit based on the predicted amount of specific impurities and the set total impurity control target include: Determine the type of operation parameters of the downstream reaction unit, including reaction temperature, reaction time and catalyst dosage; According to the predicted amount of specific impurities, the correlation function between the amount of specific impurities and each operation parameter type is established to obtain the influence degree of each operation parameter type on impurity generation, which is represented as an adjustment coefficient; According to the set total impurity control target, the adjustment amount of each operation parameter type is calculated based on the adjustment coefficient of each operation parameter type.

6. The method of claim 5, wherein the process is a pharmaceutical intermediate production process. The adjustment amount needs to meet the condition that the final total impurity amount does not exceed the control target, and the adjustment range of each operation parameter is within a preset range, and the operation parameter with the greatest influence on impurity generation is adjusted preferentially.

7. A pharmaceutical intermediate production process control system using the pharmaceutical intermediate production process control method according to any one of claims 1 to 6, characterized by The steps include: An acquisition module is configured to acquire intermediate trace component information of an intermediate sample produced by an upstream separation and purification unit, and form a batch fingerprint data set; A modeling module is configured to establish a quantitative correlation model between the batch fingerprint data set and the amount of specific impurities generated in a downstream reaction unit. A prediction module is configured to input the intermediate micro-component information of the intermediate sample of the current batch into a quantitative correlation model to predict the amount of a specific impurity generated in a downstream reaction unit; A calculation module is configured to calculate an adjustment amount of an operating parameter of the downstream reaction unit according to the predicted amount of the specific impurity and a set total amount control target of the impurity; A control module is configured to send the adjustment amount to an actuator of the downstream reaction unit to adjust an actual operating parameter of the downstream reaction unit.

Citation Information

Patent Citations

  • Intelligent monitoring, regulating and controlling system for medical intermediate synthesis production line

    CN119002404A