Medical intermediate production process regulation and control method and system

By establishing a quantitative correlation model between trace components of intermediates and impurity generation in downstream reactions and dynamically adjusting operating parameters, the problem of instability of product quality caused by batch differences of trace impurities in the production of pharmaceutical intermediates was solved, and the stability of product quality and the improvement of batch qualification rate were achieved.

CN120688783AActive Publication Date: 2025-09-23XINYI DAJIANG CHEM IND CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510751391.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-23
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing pharmaceutical intermediate production process makes it difficult to quantify the batch differences of trace impurities in the upstream separation and purification process and their impact on subsequent reactions, resulting in unstable quality of the final product. Existing control methods cannot effectively predict and compensate for these impacts.

Method used

By obtaining the trace component information of the intermediate samples, establishing a batch fingerprint data set, and constructing a quantitative correlation model, the amount of impurities generated in the downstream reaction unit is predicted, and the operating parameters are dynamically adjusted to compensate for upstream fluctuations, thereby achieving forward-looking regulation.

Benefits of technology

It significantly reduces the total amount of impurities in the final target pharmaceutical intermediate product, improves the product batch qualification rate and quality stability, reduces rework and scrap, and has important economic and technical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688783A_ABST
    Figure CN120688783A_ABST
Patent Text Reader

Abstract

The invention provides a medical intermediate production process regulation and control method and system, and relates to the technical field of medicine production. The method comprises the following steps: acquiring micro-component information of an intermediate, and forming a batch fingerprint spectrum data set; establishing a quantitative correlation model; inputting the intermediate trace component information of the intermediate sample of the current batch into the quantitative correlation model, and predicting the generation amount of specific impurities in a downstream reaction unit; according to the predicted specific impurity generation amount and a set impurity total amount control target, calculating the adjustment amount of the operation parameters of the downstream reaction unit; and sending the adjustment amount to an actuator of the downstream reaction unit so as to adjust actual operation parameters of the downstream reaction unit. According to the method disclosed by the invention, the generation trend of downstream impurities is predicted by analyzing the trace impurities of the intermediate in the current batch, and downstream operation parameters are adjusted in advance, so that adverse effects caused by upstream fluctuation are effectively compensated, and the stability of the quality of a final product is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pharmaceutical production, and in particular to a method and system for controlling the production process of a pharmaceutical intermediate. Background Art

[0002] The production process for pharmaceutical intermediates typically utilizes a cascade model with multiple reactions and multiple unit operations. In this model, the starting material is converted into the intermediate in a single step via a reactor. This intermediate mixture then enters downstream separation and purification units, such as extraction, washing, crystallization, and filtration, to effectively remove byproducts and unreacted products generated during the reaction, thereby obtaining an intermediate product that meets certain purity requirements. This intermediate product then serves as the feedstock for the next reaction step, where it undergoes similar reaction, separation, and purification processes to ultimately yield the target pharmaceutical intermediate. The ultimate goal of the entire production chain is to ensure that the total impurity content of the final product is within strictly specified limits to meet pharmaceutical quality and safety requirements.

[0003] However, in actual industrial production operations, slight fluctuations in the operating parameters of upstream units, such as deviations in temperature or pressure control in the reactor, or slight differences between different batches of starting materials, may lead to slight batch-to-batch variations in the relative content and types of various components, especially various by-products, in the one-step reaction product mixture. This reaction product mixture with batch differences then enters the downstream separation and purification unit. The operating performance of the separation and purification unit, such as the yield and crystal form of the target product in the crystallization operation and the separation effect of different impurities, is highly sensitive to process parameters such as feed composition, supersaturation, cooling rate, and stirring intensity. Due to batch differences in the composition of the upstream reaction products, even if the operating parameters of the downstream separation and purification unit remain constant, the actual separation process may also show slight batch-to-batch differences. For example, the presence of certain trace impurities may significantly affect the crystallization behavior of the target product, causing these trace impurities to be more easily included in the crystals, or changing the concentration distribution of impurities in the mother liquor, thereby affecting the purity of the intermediate product finally obtained. This fluctuation in separation performance caused by differences in feed composition results in batch-to-batch variations in the types, contents, and distribution of trace impurities in the intermediate products obtained through separation and purification, even though the content of the main component may meet the standards.

[0004] These intermediate products with trace impurity differences are then fed into the next reaction step as raw materials. It should be noted that the trace impurities entering the subsequent reactor are not inert substances. They may undergo self-transformation, participate in undesirable side reactions, or have an adverse effect on the catalyst in the reaction system under the specific temperature, pressure, presence of catalysts and other conditions of the subsequent reaction. For example, a trace metal impurity brought in from the upstream may act as a catalyst under the conditions of the subsequent reaction, accelerating the occurrence of a specific degradation pathway; or a certain organic impurity may undergo an undesirable coupling reaction with the subsequent reactants to generate new, structurally complex impurities; or, a certain impurity may be adsorbed on the surface of the solid catalyst used in the subsequent reaction, reducing its activity and selectivity, resulting in a decrease in the conversion rate of the main reaction and an increase in side reactions. Because the upstream separation and purification unit fails to completely and stably remove these trace impurities, the quality of the raw materials entering the subsequent reactors fluctuates from batch to batch. This fluctuation is often amplified during the subsequent reaction, ultimately affecting the quality of the target product. Existing process control methods usually focus on judging whether the batch is qualified by offline analysis of the final product, or performing feedback control based on online monitoring data (such as reaction temperature, pressure, and main product concentration) to maintain the stability of the main parameters of the current reaction step. However, these methods are difficult to capture and quantify the differences in trace impurities in intermediates caused by the upstream separation and purification process, and are even more unable to predict the impact of these trace impurities on the generation of multiple impurities in subsequent complex reaction systems, and make forward-looking or compensatory parameter adjustments accordingly. Therefore, although each unit operation may be operated under set parameters, the implicit fluctuations in the quality of the intermediates and their cumulative effects in subsequent reactions are still the key reasons for the instability of the total amount of impurities in the final product and the challenges faced by the batch qualification rate. There is an urgent need for a technical means to quantify trace impurities in intermediates and predict their impact on subsequent reactions, and then implement compensatory control. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for controlling the production process of pharmaceutical intermediates. By analyzing the trace impurities in the current batch of intermediates, the generation trend of downstream impurities is predicted, and downstream operating parameters are adjusted in advance, thereby effectively compensating for the adverse effects of upstream fluctuations and ensuring the stability of the final product quality.

[0006] In a first aspect, the present invention provides a method for controlling the production process of a pharmaceutical intermediate, comprising the following steps:

[0007] Obtaining the intermediate trace component information of the intermediate samples produced by the upstream separation and purification unit and forming a batch fingerprint data set;

[0008] Establish a quantitative correlation model between the batch fingerprint data set and the amount of specific impurities generated in the downstream reaction unit;

[0009] Input the intermediate trace component information of the current batch of intermediate samples into the quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit;

[0010] Calculate the adjustment amount of the downstream reaction unit operating parameters based on the predicted specific impurity generation amount and the set total impurity control target;

[0011] The adjustment amount is sent to the actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit.

[0012] The pharmaceutical intermediate production process control method provided by the present invention, in the process of serial production of pharmaceutical intermediates, faces batch fluctuations in the types and contents of trace impurities in the intermediate products output by the upstream separation and purification unit. By quantitatively analyzing these trace impurity information before the intermediates enter the downstream reaction unit, a cross-unit quantitative correlation model is established between the trace impurity information and the generation of specific impurities in the downstream reaction process. Based on the trace impurity analysis results of the current batch of intermediates, the correlation model is used to dynamically predict the downstream impurity generation trend, and then the key process parameters of the downstream reaction unit are proactively adjusted to compensate for the batch differences in the quality of the upstream raw materials, thereby reducing the total amount of impurities in the final target pharmaceutical intermediate product and improving batch stability.

[0013] In a second aspect, the present invention provides a pharmaceutical intermediate production process control system, comprising:

[0014] An acquisition module is used to obtain the intermediate trace component information of the intermediate sample produced by the upstream separation and purification unit and form a batch fingerprint data set;

[0015] Establishing a module for establishing a quantitative correlation model between a batch fingerprint data set and the amount of specific impurities generated in a downstream reaction unit;

[0016] The prediction module is used to input the intermediate trace component information of the current batch of intermediate samples into the quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit;

[0017] A calculation module, used to calculate the adjustment amount of the operating parameters of the downstream reaction unit based on the predicted specific impurity generation amount and the set impurity total amount control target;

[0018] The control module is used to send the adjustment amount to the actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit.

[0019] As can be seen from the above, the pharmaceutical intermediate production process control method provided by the present invention can quantitatively analyze the types and contents of trace impurities in the intermediates produced by the upstream separation and purification unit, and accurately capture the batch quality differences of the intermediate raw materials. By establishing a quantitative correlation model between the fingerprint spectrum of the intermediate trace components and the impurity generation of subsequent reactions, the cross-unit, intermediate quality-based impurity generation prediction is realized for the first time. Based on this prediction model, the system can proactively calculate and dynamically adjust the key process parameters of the subsequent reaction units, effectively compensating for the impact of batch fluctuations in the quality of upstream raw materials on downstream reactions. This significantly reduces the total amount of impurities in the final target pharmaceutical intermediate product, improves the batch qualification rate and quality stability of the product, reduces rework and scrap, and has important economic and technical value.

[0020] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the embodiments of the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flow chart of a method for controlling the production process of a pharmaceutical intermediate provided in an embodiment of the present invention.

[0022] Figure 2 A schematic structural diagram of a pharmaceutical intermediate production process control system provided in an embodiment of the present invention.

[0023] Description of labels:

[0024] 100, acquisition module; 200, establishment module; 300, prediction module; 400, calculation module; 500, control module. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0027] Reference Attachment Figure 1 The present invention provides a method for controlling the production process of a pharmaceutical intermediate, comprising the following steps:

[0028] Obtaining the intermediate trace component information of the intermediate samples produced by the upstream separation and purification unit and forming a batch fingerprint data set;

[0029] Establish a quantitative correlation model between the batch fingerprint data set and the amount of specific impurities generated in the downstream reaction unit;

[0030] Input the intermediate trace component information of the current batch of intermediate samples into the quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit;

[0031] Calculate the adjustment amount of the downstream reaction unit operating parameters based on the predicted specific impurity generation amount and the set total impurity control target;

[0032] The adjustment amount is sent to the actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit.

[0033] Intermediate samples produced by the upstream separation and purification unit are obtained and analyzed to obtain information on the intermediate trace components. This information is collected and organized into batch fingerprint data sets, which are used to characterize the quality characteristics of different batches of intermediates. Furthermore, based on historical production data, a quantitative correlation model is established that can describe the relationship between the intermediate trace component fingerprint data set and the amount of specific impurities generated in the downstream reaction unit. Thus, when a new batch of intermediate samples is available, its trace component information is input into the established quantitative correlation model, and the model outputs a predicted value for the amount of specific impurities generated in the downstream reaction unit. Based on the difference between this predicted value and the preset total impurity control target, the adjustment amount required to the operating parameters of the downstream reaction unit is calculated. Finally, the calculated adjustment amount is sent to the actuator of the downstream reaction unit, which changes the actual operating parameters of the downstream reaction unit, such as reaction temperature, reaction time, or catalyst dosage, based on the received adjustment amount.

[0034] Specifically, this method monitors the quality of the intermediates entering the downstream reaction unit, quantifies the characteristics of their trace components, and uses historical data to establish a quantitative relationship model between the trace components of the intermediates and the generation of downstream impurities. Before a new batch of intermediates enters the downstream reaction unit, they are analyzed for trace components, and the analysis results are input into the model to predict the number of specific impurities that may be produced by the batch of intermediates in the downstream reaction. If the predicted amount of impurity generation exceeds the set control target, the value of the downstream operating parameters that need to be adjusted is calculated based on the predicted value and the control target, combined with the influence of the operating parameters on impurity generation. These adjustments are then applied to the downstream reaction unit to compensate for the impact of upstream intermediate quality fluctuations on downstream impurity generation, thereby controlling the total amount of impurities in the final product within the target range. This method realizes forward-looking regulation based on the quality of upstream intermediates, improving the stability of the production process and the consistency of product quality.

[0035] The working principle of the present invention is to build a cross-unit quality prediction and compensation control closed loop. First, by conducting detailed trace component analysis on the intermediates after upstream separation and purification, their batch differences are quantified to form a unique "quality fingerprint". Then, using historical production data, a mathematical model is established between this intermediate quality fingerprint and the amount of specific impurities generated during the downstream reaction process, revealing the influence of trace impurities on subsequent reactions. For the current production batch, the system obtains its quality fingerprint and inputs it into the model to predict its performance in the downstream reaction. Based on the prediction results, the system calculates the optimal adjustments required for the downstream reaction units (such as reaction temperature, time, etc.) and implements these adjustments through the automated control system. In this way, before the start of the downstream reaction or during the process, the quality fluctuations of the upstream raw materials are proactively or compensatoryly intervened, thereby stabilizing and reducing the total amount of impurities in the final product. The entire system continuously improves the accuracy of prediction and control through continuous data accumulation and model optimization.

[0036] In some embodiments, the downstream reaction unit in the pharmaceutical intermediate production process is assumed to be a catalytic reactor, and a specific impurity is a byproduct produced in this reaction. A sample of the intermediate produced by the upstream separation and purification unit is analyzed by high-performance liquid chromatography (HPLC) to obtain data on the relative contents of five trace impurities (A, B, C, D, and E), forming a batch fingerprint data set. For example, the fingerprint data for a batch of intermediates is: impurity A 0.12%, impurity B 0.05%, impurity C 0.08%, impurity D 0.03%, and impurity E 0.06%. A partial least squares regression model is established based on the fingerprint data of the intermediates from the past 100 batches and the actual production data of the specific impurities in the corresponding downstream reaction unit. When the analysis of a new batch of intermediates is completed, its trace component data (e.g., 0.12%, 0.05%, 0.08%, 0.03%, and 0.06%) are input into the model. The model predicts that 0.7% of the specific impurity will be produced in the downstream reaction for this batch. The total amount of specific impurities set as the control target is no more than 0.5%. Based on the predicted value of 0.7% and the target value of 0.5%, the required parameter adjustments are calculated. For example, based on the model or pre-set rules, it is determined that lowering the downstream reaction temperature by 2°C and extending the reaction time by 15 minutes can reduce the formation of a specific impurity by approximately 0.2%. The calculated adjustments (temperature -2°C, time +15 minutes) are sent to the reactor's temperature controller and timer actuator, which adjust the actual operating parameters.

[0037] In certain embodiments, the step of establishing a quantitative correlation model between the batch fingerprint data set and the amount of specific impurities generated in the downstream reaction unit includes:

[0038] A1. Obtain historical production data; historical production data includes fingerprint data sets of trace components of multiple batches of intermediates and the amount of specific impurities generated in the downstream reaction units corresponding to each batch of intermediates;

[0039] A2. Preprocess historical production data. The specific steps of preprocessing include:

[0040] A21. Use principal component analysis to reduce the dimensionality of the intermediate trace component fingerprint data and extract the principal component scores that characterize the intermediate quality fluctuations.

[0041] A22. Normalize the data on specific impurity generation in downstream reaction units to eliminate dimensionality effects.

[0042] A3. Based on preprocessed historical production data, a partial least squares regression algorithm is used to establish a quantitative correlation model between the principal components of the intermediate trace components and the amount of specific impurities generated in the downstream reaction unit. The quantitative correlation model includes regression coefficients to characterize the degree to which different trace impurity components affect the formation of downstream impurities.

[0043] A4. Using a cross-validation method, use a quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit of historical batches, calculate the prediction error between the predicted value and the actual value, and determine whether the prediction error exceeds a first preset threshold;

[0044] A5. If the prediction error exceeds the first preset threshold, return to step A3, adjust the parameters of the partial least squares regression algorithm, and re-establish the quantitative association model and verify it; otherwise, use the current quantitative association model as the final model.

[0045] Specifically, this technical solution aims to solve the problem of how to establish an accurate quantitative correlation model in a data-driven manner when the reaction mechanism is unknown, thereby improving the accuracy of impurity control. First, a data set is constructed by collecting the fingerprint spectrum data of the intermediate trace components of multiple historical production batches and the corresponding specific impurity generation amounts of the downstream reaction units. Then, the collected data is preprocessed, including using the principal component analysis method to reduce the dimensionality of the high-dimensional intermediate trace component data, extracting a few principal components that can represent the quality fluctuations of the intermediate, and normalizing the downstream impurity generation amount data to eliminate the dimensional differences between different impurity generation amounts. Then, using the preprocessed principal component data and the normalized impurity generation amount data, a mathematical model is established using the partial least squares regression algorithm. The model can quantify the relationship between the intermediate trace components (represented by the principal components) and the downstream specific impurity generation amount. The model contains regression coefficients, which reflect the degree of influence of different trace components on the downstream impurity generation. After the model is established, the cross-validation method is used to evaluate the predictive performance of the model. The historical data is divided into a training set and a validation set. The training set is used to train the model, and the validation set is used to evaluate the prediction error of the model. By calculating the difference between the predicted value and the actual value (such as the root mean square error), it is determined whether the prediction accuracy of the model meets the requirements. If the prediction error exceeds the first preset threshold, it indicates that the prediction ability of the current model is insufficient, and it is necessary to return to the model building step, adjust the parameters of the partial least squares regression algorithm (such as the number of principal components, model complexity, etc.), retrain the model and verify it again. Repeat this process until the prediction error of the model is lower than the first preset threshold, and the model at this time is determined to be the final quantitative association model. The final model can predict the amount of downstream specific impurities generated based on the intermediate trace component information, providing a basis for subsequent process parameter adjustments. Through this data-driven modeling and iterative optimization process, even when the reaction mechanism is not completely clear, an effective prediction model can be established, thereby realizing the prediction and control of downstream impurity generation.

[0046] In some specific embodiments, the process of establishing a quantitative association model can be implemented as follows: collect 100 historical production batches of intermediate samples, perform trace component analysis on each sample, and obtain fingerprint data containing 50 different trace component contents. At the same time, record the content of the specific impurity A generated by these 100 batch intermediates after reaction in the downstream reaction unit. Use these 100 sets of 50-dimensional trace component data and the corresponding impurity A content data as historical production data. Apply principal component analysis to the trace component data, calculate the variance contribution rate of each component, and determine that the first 10 principal components can explain 95% of the total variance. Project the 50-dimensional data onto these 10 principal components to obtain 100 sets of 10-dimensional principal component score data. Perform minimum-maximum normalization on the impurity A content data and scale its numerical range to [0, 1]. Use these 100 sets of 10-dimensional principal component score data as independent variables and the normalized impurity A content as the dependent variable, and use the partial least squares regression algorithm to establish a model. For example, the number of potential variables in the partial least squares model is set to 5. After the model is established, the 5-fold cross-validation method is used to evaluate the model. The 100 groups of data are randomly divided into 5 parts, 1 part is taken as the validation set each time, and the remaining 4 parts are used as training sets. The training set is used to train the model, and the validation set is used to predict and calculate the prediction error (such as the root mean square error). Repeat 5 times and calculate the average of the 5 errors. If the average root mean square error is greater than 0.05, return to adjust the parameters of the partial least squares model, such as increasing the number of potential variables to 6, re-establish the model and perform cross-validation. Repeat this process until the average root mean square error is less than or equal to 0.05, and the model at this time is determined to be the final model. This model can predict the amount of downstream impurity A generated based on the principal component information of the new intermediate batch.

[0047] In some embodiments, the specific steps in step A21 include:

[0048] A211. Calculate the variance contribution of each intermediate trace component in the intermediate trace component fingerprint data, and determine at least two intermediate trace components with variance contribution rates greater than a second preset threshold as key components that have a significant impact on the quality fluctuation of the intermediate (key components are not necessarily impurities; they refer to components that have a significant impact on the quality fluctuation of the intermediate. These components can be impurities or major components in the intermediate);

[0049] A212. Calculate the covariance matrix of the key components using the preset principal component analysis model, and obtain the eigenvectors of each principal component by solving the covariance matrix;

[0050] A213. Based on the eigenvectors of each principal component, the corresponding intermediate trace component fingerprint data or all key components are projected from the high-dimensional space to the low-dimensional space composed of all principal components to obtain the principal component score that can characterize the quality fluctuation of the corresponding batch of intermediate samples; the principal component contains all key components (the principal component here can be understood as a comprehensive indicator composed of multiple trace components (including impurities), and does not refer to the necessary components for the drug to achieve therapeutic efficacy in the traditional sense); all principal components constitute a low-dimensional representation that characterizes the quality fluctuation of the corresponding batch of intermediate samples (that is, the high-dimensional intermediate trace component data or all key components are converted into a low-dimensional principal component representation); the principal component score serves as the input of the subsequent quantitative association model.

[0051] Calculating the variance contribution of each intermediate trace component in the intermediate trace component fingerprint data involves applying standard statistical methods to historical data to calculate the contribution of each trace component to the total variance. A second preset threshold is a numerical value used to screen out components with high variance contributions. Components with variance contributions greater than the second preset threshold are identified as key components, ensuring that subsequent analysis focuses on key factors influencing intermediate quality fluctuations and avoiding overlooking trace components with significant impacts on downstream reactions. Constructing the covariance matrix for key components involves using the values ​​of the identified key components across all historical batches to calculate the covariance between any two key components to form a matrix. Solving the covariance matrix to obtain eigenvectors involves applying linear algebra algorithms, such as the Jacobi iteration method, to find the matrix's eigenvalues ​​and corresponding eigenvectors. These eigenvectors indicate the primary directions of variation in the key component data. Projecting the intermediate trace component fingerprint data or key components into a low-dimensional space based on the eigenvectors of each principal component involves transforming the original data points into a new coordinate system defined by the eigenvectors through matrix multiplication. Therefore, each batch of data is represented as a set of principal component scores, which comprehensively reflect the characteristics of key components in different variation directions, thereby reducing the data dimension.

[0052] Specifically, to address the problem that directly performing principal component analysis on all trace components may overlook certain key information, this solution first calculates the variance contribution rate of each component in the trace component data of historical batches of intermediates. Based on a second preset threshold, trace components with a variance contribution rate greater than the second preset threshold are screened out, and these components are identified as key components that have a significant impact on the quality fluctuation of the intermediates. This screening step ensures that subsequent analysis focuses on the key factors affecting the quality fluctuation of the intermediates, improving the targeted nature of the analysis. Based on the screened key component data, a covariance matrix of the key components is constructed, which describes the relationship and degree of variation between the key components. By solving the covariance matrix, the eigenvectors of each principal component are obtained. These eigenvectors represent the main directions of variation in the key component data. Finally, based on these eigenvectors, the corresponding intermediate trace component fingerprint data or all key components are projected from the original high-dimensional space into a low-dimensional space composed of the selected principal components. Thus, the quality fluctuation of each batch of intermediate samples is represented as a set of principal component scores, which serve as input for the subsequent establishment of a quantitative association model. This process reduces the data dimension while retaining the main variation information of the data, improving the computational efficiency and generalization ability of the model, and providing effective input for accurately predicting the amount of specific impurities generated downstream.

[0053] For example:

[0054] Assume that three key intermediate trace components (component X, component Y, and component Z) have been identified through preliminary analysis. After processing historical production data, the content data of these key components in multiple batches are obtained. By calculating the covariance matrix and performing eigenvalue decomposition, two main principal components (PC1 and PC2) and their corresponding eigenvectors (v1 and v2) are obtained. The eigenvectors v1 and v2 define the two main directions of data variation. For a specific batch of intermediate samples, its key component content data is [x, y, z]. This data point is projected onto the PC1 and PC2 axes, and its principal component score is calculated. Specifically, the score (Score1) of this batch on PC1 can be obtained by taking the dot product of the data vector [x, y, z] and the eigenvector v1, that is, Score1 = x*v1_x+y*v1_y+z*v1_z. Similarly, the score (Score2) for this batch on PC2 can be obtained by performing a dot product of the data vector [x, y, z] and the feature vector v2, that is, Score2 = x * v2_x + y * v2_y + z * v2_z. Thus, the original three-dimensional data points [x, y, z] are converted into a two-dimensional principal component score vector [Score1, Score2]. This two-dimensional vector [Score1, Score2] is the principal component score representation of the quality fluctuation of the intermediate samples in this batch. These principal component scores serve as input variables for the subsequent establishment of a partial least squares regression model to predict the amount of specific impurities generated in the downstream reaction unit. By using low-dimensional principal component scores instead of the original high-dimensional trace component content, the complexity of the model is reduced, while retaining the main variation information of the data, improving the prediction efficiency and accuracy of the model.

[0055] In some specific embodiments, consider the intermediate sample data from 150 historical production batches, each sample measured the content of 20 trace components. First, calculate the variance contribution of these 20 trace components in the 150 batch data. Set the second preset threshold to 8%. The analysis results show that the variance contribution of 6 trace components exceeds 8%, and these 6 components are determined to be key components. Then, use the content data of these 6 key components in 150 batches to calculate a 6x6 covariance matrix. This covariance matrix is ​​subjected to eigenvalue decomposition to obtain 6 eigenvalues ​​and corresponding eigenvectors. Sort by eigenvalue size, select the eigenvectors corresponding to the first 3 eigenvalues ​​whose cumulative variance contribution reaches 90% as the eigenvectors of the principal component. Finally, the key components (6-dimensional data) of 150 batches are projected into the 3-dimensional space composed of these 3 eigenvectors to obtain the 3 principal component scores of each batch. These 3-dimensional principal component score vectors are used as the input for the subsequent partial least squares regression model to predict the amount of impurities generated downstream. As a result, the original 20-dimensional or 6-dimensional data is effectively reduced to 3 dimensions, simplifying the model structure while retaining key quality fluctuation information.

[0056] In some embodiments, the specific steps in step A211 include:

[0057] Using stratified sampling, historical production data was divided into multiple time periods based on the temporal distribution of intermediate production batches.

[0058] For the data in each time period, the variance contribution rate of each intermediate trace component is calculated, and the box plot method is used to identify and eliminate the outliers in the variance contribution rate to obtain the corrected variance contribution rate;

[0059] The corrected variance contribution rates of the data in all time periods are summarized, the average variance contribution rate of each intermediate trace component is calculated, and according to a second preset threshold, the intermediate trace components with an average variance contribution rate greater than the threshold are regarded as key components that have a significant impact on the intermediate quality fluctuation.

[0060] Historical production data is divided into data for multiple time periods based on the time of intermediate production batches. This takes into account possible changes in the production process over time. For the data in each time period, the variance contribution rate of each intermediate trace component is calculated. Furthermore, the box plot method is used to identify and eliminate outliers in the variance contribution rate to obtain a revised variance contribution rate. By eliminating outliers, the impact of extreme data on the calculation results is reduced, making the calculation of the variance contribution rate more representative. The revised variance contribution rates of the data in all time periods are summarized, and the average variance contribution rate of each intermediate trace component is calculated. According to a second preset threshold, the intermediate trace components whose average variance contribution rate is greater than the second preset threshold are determined as key components that have an impact on the quality fluctuation of the intermediate. These components all show a high fluctuation contribution in different time periods.

[0061] Specifically, in response to the fluctuations in the fingerprint data of the intermediate trace components and the impact of outliers on the stability of the direct calculation of the variance contribution rate, this solution adopts a stratified sampling method to divide the historical production data into time periods according to the time of the intermediate production batch. Thus, the changes that may occur in the production process over time are taken into account. For the data in each time period, the variance contribution rate of each intermediate trace component is calculated. Furthermore, the box plot method is used to identify and eliminate outliers in the variance contribution rate to obtain a revised variance contribution rate. By eliminating outliers, the impact of extreme data on the calculation results is reduced, making the calculation of the variance contribution rate more representative. The revised variance contribution rates of the data in all time periods are summarized, and the average variance contribution rate of each intermediate trace component is calculated. According to the second preset threshold, the intermediate trace components whose average variance contribution rate is greater than the second preset threshold are determined as key components that have an impact on the intermediate quality fluctuation. These components all show a high fluctuation contribution in different time periods. Through the above steps, this scheme takes into account the temporal characteristics of production data and the influence of outliers, identifies components that have a continuous or major impact on the quality fluctuation of intermediates, and provides more reliable input for subsequent principal component analysis, thereby improving the accuracy of the quantitative association model.

[0062] In some embodiments, the specific steps in step A212 include:

[0063] According to the determined key components, the covariance matrix of the key components is constructed by calculating the covariance between any two key components;

[0064] The Jacobi iteration method is used to perform eigenvalue decomposition on the covariance matrix to obtain multiple eigenvalues ​​and corresponding eigenvectors, and all eigenvectors are sorted in descending order of eigenvalues;

[0065] The eigenvectors corresponding to the first N eigenvalues ​​are selected as the eigenvectors of the principal components, and the variance contribution rates of the key components corresponding to the first N eigenvalues ​​are ensured to reach a preset ratio, where N is the minimum value when the cumulative variance contribution rate reaches the preset ratio.

[0066] Based on the identified key components, the covariance between any two key components is calculated to construct a covariance matrix. This matrix is ​​a symmetric matrix whose diagonal elements represent the variance of each key component itself, and the off-diagonal elements represent the covariance between any two key components, reflecting the degree of their common variation. The constructed covariance matrix is ​​subjected to eigenvalue decomposition using the Jacobi iteration method. The Jacobi iteration method is a numerical algorithm used to solve the eigenvalues ​​and eigenvectors of a symmetric matrix. This method can obtain multiple eigenvalues ​​and a set of orthogonal eigenvectors corresponding to the covariance matrix. The obtained eigenvectors are sorted according to the size of their corresponding eigenvalues. The larger the eigenvalue, the more variance in the original data is explained by the principal component represented by the corresponding eigenvector. The eigenvectors corresponding to the first N eigenvalues ​​after sorting are selected as the eigenvectors of the principal component. N is determined based on the cumulative variance contribution rate, that is, the ratio of the sum of the variances corresponding to the first N eigenvalues ​​to the sum of all eigenvalues ​​reaches or exceeds a third preset threshold. N is taken as the minimum value that meets this condition to achieve data dimensionality reduction while retaining the main information.

[0067] Specifically, this technical solution aims to extract principal components that can effectively characterize intermediate quality fluctuations from trace component data of intermediates, thereby solving the problems of constructing a covariance matrix, performing eigenvalue decomposition, and determining the number of principal components. First, based on the key components previously identified as significantly influencing intermediate quality fluctuations, historical batch content data for these key components are collected. Based on this data, the covariance between any two key components is calculated, and a covariance matrix for the key components is constructed. This matrix quantifies the intercorrelations and variability between the key components. Next, the Jacobi iteration method is used to perform eigenvalue decomposition on this covariance matrix. The Jacobi iteration method gradually diagonalizes the covariance matrix through a series of rotation transformations. The elements on the diagonal are the eigenvalues, and the column vectors of the rotation matrix are the corresponding eigenvectors. The results of the eigenvalue decomposition reveal the main directions (eigenvectors) and their degree (eigenvalues) of data variation. The resulting eigenvectors are sorted from largest to smallest according to their corresponding eigenvalues, with the directions explaining the most variance first. The proportion of each eigenvalue to the sum of all eigenvalues ​​is then calculated, representing the variance contribution of the principal component. The variance contribution rates after sorting are accumulated until the cumulative proportion reaches a third preset threshold (for example, 95% or 98%). The first N eigenvectors with the minimum number required to achieve the cumulative proportion are selected as the final principal component directions. These principal components are linear combinations of the original key components. They are mutually orthogonal and contain most of the variation information of the original data. In this way, the high-dimensional key component data are effectively mapped to the low-dimensional principal component space. The extracted principal components can accurately capture the main fluctuation patterns of the intermediate quality, provide concise and information-rich data input for the subsequent establishment of a quantitative correlation model, and improve the accuracy and robustness of the model.

[0068] In some specific embodiments, it is assumed that three key components are determined through preliminary analysis: impurity A, impurity B, and impurity C. The content data of these three impurities in 100 historical batches are collected. First, the 3x3 covariance matrix of these 100 batches of data is calculated. For example, the covariance between impurity A and impurity B, the covariance between impurity A and impurity C, the covariance between impurity B and impurity C, and the variance of impurities A, B, and C are calculated. These values ​​are filled into the covariance matrix. Then, the Jacobi iteration algorithm is used to perform eigenvalue decomposition on the 3x3 covariance matrix to obtain three eigenvalues ​​λ1, λ2, λ3 and the corresponding three eigenvectors v1, v2, and v3. Assume that the calculated eigenvalues ​​are λ1=10.5, λ2=2.1, and λ3=0.4 respectively. The eigenvectors are sorted from large to small according to the eigenvalues ​​to obtain the order v1, v2, and v3. Calculate the variance contribution rate of each eigenvalue: vCr1 = 10.5 / (10.5+2.1+0.4) ≈ 0.81, vcr2 = 2.1 / (10.5+2.1+0.4) ≈ 0.16, vcr3 = 0.4 / (10.5+2.1+0.4) ≈ 0.03. Calculate the cumulative variance contribution rate: cumulative vcr1 = 0.81, cumulative vcr1+Vcr2 = 0.81+0.16 = 0.97, cumulative vcr1+vcr2+vcr3 = 0.97+0.03 = 1.00. Set the preset cumulative variance contribution rate threshold to 95% (i.e., the third preset threshold). Since cumulative vcr1+vcr2 = 0.97 > 0.95, and this is the minimum N value (N = 2) that reaches the third preset threshold, the first two eigenvectors v1 and v2 are selected as the eigenvectors of the principal component. This means that the original three-dimensional key component data can be reduced to two dimensions, represented by these two principal components, while retaining approximately 97% of the original data variation information. These extracted principal components serve as input to the subsequent quantitative correlation model, effectively reducing the model's complexity while ensuring accurate representation of intermediate mass fluctuations.

[0069] In some embodiments, the specific steps in step A4 include:

[0070] A41. Divide historical production data into K non-overlapping subsets, each containing data from a different batch of intermediates.

[0071] A42. Select the i-th subset as the validation set and the remaining K-1 subsets as the training set, where i ranges from 1 to K.

[0072] A43. Based on the training set, use the quantitative association model to predict the amount of specific impurities generated in the downstream reaction unit of each batch in the validation set to obtain a set of predicted values;

[0073] A44. Calculate the prediction error for the i-th subset using the root mean square error algorithm based on the predicted value set and the actual impurity generation levels for each batch in the validation set.

[0074] A45. Repeat steps A42 to A44 to calculate the prediction errors corresponding to the K subsets;

[0075] A46. Calculate the average of the K prediction errors as the final prediction error of the cross-validation;

[0076] A47. Determine whether the final prediction error exceeds a first preset threshold.

[0077] Historical production data is divided into K non-overlapping subsets, each configured to contain intermediate data from different production batches. This ensures that the data distribution in each subset is representative. The system is then configured to select one of these subsets in turn as the validation set, while the remaining K-1 subsets are combined as the training set. This process is repeated K times, each time selecting a different subset as the validation set. Based on the currently selected training set, the pre-established quantitative correlation model is used to process the intermediate data from each batch in the currently selected validation set to predict the amount of specific impurities generated in the downstream reaction unit and form a set of predicted values. Next, the root mean square error (RMSE) algorithm is used to calculate the prediction error for the current validation set based on this predicted value set and the actual impurity levels for each batch in the validation set. The RMSE algorithm is configured to quantify the degree of deviation between the predicted and actual values. These steps are repeated to calculate the prediction errors for all K subsets. Finally, the average of these K prediction errors is calculated and used as the final cross-validation prediction error to comprehensively evaluate the overall prediction performance of the model. Finally, the final prediction error is compared with the first preset threshold to determine whether the prediction accuracy of the model meets the requirements.

[0078] Specifically, this solution aims to address the problem of improper cross-validation data subset partitioning when predicting specific impurity formation in downstream reaction units using a quantitative correlation model established using partial least squares regression. This can lead to unrepresentative validation results and an inability to accurately assess the model's generalization ability. By partitioning historical production data into K non-overlapping subsets, each containing data from a different batch of intermediates, this approach addresses this issue and ensures that the cross-validation data base is representative. By rotating different subsets as validation sets and the remaining subsets as training sets, and repeating this process K times, a comprehensive assessment of the model's generalization ability is achieved, preventing the model from performing well only on specific data subsets. Predictions are made on the validation set based on the training set, and the prediction error is calculated using the root mean square error (RMSE) algorithm, providing a quantitative assessment of the model's prediction accuracy. Repeated calculations and averaging reduce the variability of individual validation results and improve the reliability of the model assessment. Finally, by determining whether the average prediction error exceeds a pre-set threshold, a clear basis for model optimization and selection is provided, ensuring that the model's prediction accuracy meets the requirements of practical applications. This improves the reliability and accuracy of the quantitative correlation model evaluation, providing a solid foundation for subsequent process control based on the model prediction results.

[0079] In certain embodiments, the step of calculating the adjustment amount of the downstream reaction unit operating parameter based on the predicted specific impurity generation amount and the set impurity total amount control target includes:

[0080] Determine the operating parameter type of the downstream reaction unit, the operating parameter type including reaction temperature, reaction time, and catalyst dosage;

[0081] According to the predicted specific impurity generation amount, based on the pre-established correlation function between the specific impurity generation amount and each operating parameter type, the influence of each operating parameter type on the impurity generation is obtained and expressed as an adjustment coefficient;

[0082] Based on the set total impurity control target and combined with the adjustment coefficient of each operating parameter type, the adjustment amount of each operating parameter type is calculated; the adjustment amount must ensure that the final total impurity amount does not exceed the control target, and the adjustment range of each operating parameter is within the preset range, and priority is given to adjusting the operating parameters that have a greater impact on impurity generation.

[0083] This method identifies the types of parameters that can be controlled in downstream reaction units, such as reaction temperature, reaction time, and catalyst dosage. Furthermore, the method utilizes the predicted specific impurity generation information, combined with a pre-established correlation function, to quantify the impact of each operating parameter type on specific impurity generation and expresses this quantification as an adjustment coefficient. The correlation function reflects the quantitative relationship between specific impurity generation and different operating parameters. Based on the set total impurity control target and the calculated adjustment coefficient for each operating parameter type, the method calculates the specific adjustment amount for each operating parameter. This calculation considers multiple constraints: the adjusted parameters must ensure that the final total impurity amount does not exceed the set control target; the adjustment range for each operating parameter must be within a preset safe or feasible range; and, subject to these conditions, prioritizing adjustment of operating parameters that have a greater impact on impurity generation. By determining the specific parameter type, quantifying the impact of each parameter, and calculating the adjustment amount based on the control target, adjustment range, and impact priority, the method converts the predicted impurity information into executable process adjustment instructions.

[0084] Specifically, this method identifies the controllable operating parameters in the downstream reaction unit that are relevant to impurity formation, thereby clarifying the control targets. For example, in a particular downstream reaction unit, reaction temperature, reaction time, and catalyst dosage are identified as the primary operating parameters influencing the formation of a specific impurity. Next, the total amount of the specific impurity likely to be generated in the downstream reaction unit for the current batch is predicted using information on trace components of the upstream intermediates and compared with the set control target for the total impurity amount in the final product. Based on a quantitative correlation function established previously using experimental data or historical production data, such as a multivariate linear regression model or a nonlinear model, this method calculates the impact of changing the unit reaction temperature, unit reaction time, or unit catalyst dosage on the specific impurity formation. These impacts serve as the adjustment coefficients for each operating parameter type. For example, the correlation function may indicate that increasing the reaction temperature significantly increases the formation of a specific impurity, while extending the reaction time or increasing the catalyst dosage has a minimal or inhibitory effect. The method then calculates the required adjustments to the reaction temperature, reaction time, and catalyst dosage based on the difference between the predicted impurity formation and the control target, combined with the adjustment coefficients for each operating parameter type. This calculation process is an optimization problem. It requires ensuring that the adjustment range of each parameter does not exceed its preset maximum adjustment range, while ensuring that the final total impurity content does not exceed the control target. The operating parameters with the largest absolute adjustment coefficients are prioritized for adjustment, thereby achieving maximum impurity control with the smallest parameter adjustment range. This method thus converts predicted impurity information into specific process parameter adjustment instructions, compensating for the impact of upstream intermediate quality fluctuations on downstream impurity generation, and effectively controlling the total impurity content of the final product.

[0085] In some embodiments, suppose a batch of intermediates is predicted to generate 0.5% of a specific impurity in a downstream reaction unit, while the final total impurity control target requires that this specific impurity not exceed 0.3%. The operating parameters for the downstream reaction unit are reaction temperature, reaction time, and catalyst dosage. A pre-established correlation function indicates that for every 1°C increase in reaction temperature, the specific impurity increases by 0.1%; for every 10 minutes of reaction time, the specific impurity increases by 0.05%; and for every 1% increase in catalyst dosage, the specific impurity decreases by 0.08%. The preset adjustment ranges for each parameter are: reaction temperature ±3°C, reaction time ±30 minutes, and catalyst dosage ±5%. The specific impurity generation needs to be reduced from 0.5% to 0.3%, a reduction of 0.2%. According to the correlation function, reaction temperature has the greatest impact on impurity generation (the absolute value of the adjustment coefficient is the largest). Therefore, adjusting the reaction temperature is the priority. Lowering the reaction temperature by 1°C reduces impurities by 0.1%. A further 0.1% reduction is required. Next, consider the parameter with the next greatest impact, such as catalyst dosage. Increasing the catalyst dosage by 1% reduces impurities by 0.08%. If the catalyst dosage is increased by another 1.25% (a total increase of 2.25%), impurities can be reduced by 0.1%. At this point, the total impurity reduction is 0.1% + 0.1% = 0.2%, achieving the control target. The calculated adjustment amount is: reduce the reaction temperature by 1°C, keep the reaction time unchanged, and increase the catalyst dosage by 2.25%. These adjustments are all within the preset range (-1°C is within ±3°C, 0 minutes is within ±30 minutes, and +2.25% is within ±5%). Therefore, based on the predicted impurity generation and the control target, the specific downstream operating parameter adjustments are calculated.

[0086] Reference Attachment Figure 2 The present invention provides a pharmaceutical intermediate production process control system, comprising:

[0087] The acquisition module 100 is used to obtain the intermediate trace component information of the intermediate sample produced by the upstream separation and purification unit and form a batch fingerprint data set;

[0088] Establishing module 200, for establishing a quantitative correlation model between the batch fingerprint data set and the amount of specific impurities generated in the downstream reaction unit;

[0089] Prediction module 300, for inputting the intermediate trace component information of the current batch of intermediate samples into the quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit;

[0090] A calculation module 400 is used to calculate the adjustment amount of the operating parameters of the downstream reaction unit based on the predicted specific impurity generation amount and the set impurity total amount control target;

[0091] The control module 500 is configured to send the adjustment amount to the actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit.

[0092] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.

[0093] The foregoing description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for controlling the production process of a pharmaceutical intermediate, characterized in that: The following steps are involved: Obtaining the intermediate trace component information of the intermediate samples produced by the upstream separation and purification unit and forming a batch fingerprint data set; Establish a quantitative correlation model between the batch fingerprint data set and the amount of specific impurities generated in the downstream reaction unit; Input the intermediate trace component information of the current batch of intermediate samples into the quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit; Calculate the adjustment amount of the downstream reaction unit operating parameters based on the predicted specific impurity generation amount and the set total impurity control target; The adjustment amount is sent to the actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit.

2. The method for controlling the production process of a pharmaceutical intermediate according to claim 1 is characterized in that: The steps of establishing a quantitative correlation model between a batch fingerprint data set and the amount of specific impurities generated in a downstream reaction unit include: A1. Obtain historical production data; the historical production data includes a collection of fingerprint data of trace components of multiple batches of intermediates and the amount of specific impurities generated in the downstream reaction unit corresponding to each batch of intermediates; A2. Preprocess the historical production data; A3. Based on preprocessed historical production data, a quantitative correlation model was established using a partial least squares regression algorithm. The quantitative correlation model included regression coefficients that characterize the impact of different trace impurity components on downstream impurity formation. A4. Using a cross-validation method, use the quantitative association model to predict the amount of specific impurities generated in the downstream reaction unit of historical batches, calculate the prediction error between the predicted value and the actual value, and determine whether the prediction error exceeds a first preset threshold; A5. If the prediction error exceeds the first preset threshold, return to step A3, adjust the parameters of the partial least squares regression algorithm, and re-establish the quantitative association model and verify it; otherwise, use the current quantitative association model as the final model.

3. The method for controlling the production process of a pharmaceutical intermediate according to claim 2 is characterized in that: The specific steps in step A2 include: A21. Use principal component analysis to reduce the dimensionality of the intermediate trace component fingerprint data and extract the principal component scores that characterize the intermediate quality fluctuations. A22. Normalize the data on specific impurity generation in the downstream reaction unit to eliminate dimensional effects.

4. The method for controlling the production process of a pharmaceutical intermediate according to claim 3 is characterized in that: The specific steps in step A21 include: A211. Calculate the variance contribution rate of each intermediate trace component in the intermediate trace component fingerprint data, and determine at least two intermediate trace components with a variance contribution rate greater than a second preset threshold as key components that have a significant impact on the intermediate quality fluctuation; A212. Calculate the covariance matrix of the key components and obtain the eigenvectors of each principal component by solving the covariance matrix. A213. Based on the eigenvectors of each principal component, the corresponding intermediate trace component fingerprint data or all key components are projected from the high-dimensional space into the low-dimensional space composed of all principal components to obtain the principal component scores that can characterize the quality fluctuations of the corresponding batch of intermediate samples; the principal components include all key components.

5. The method for controlling the production process of a pharmaceutical intermediate according to claim 4 is characterized in that: The specific steps in step A211 include: Using stratified sampling, historical production data was divided into multiple time periods based on the temporal distribution of intermediate production batches. For the data in each time period, the variance contribution rate of each intermediate trace component is calculated, and the box plot method is used to identify and eliminate the outliers in the variance contribution rate to obtain the corrected variance contribution rate; The corrected variance contribution rates of the data in all time periods are summarized, the average variance contribution rate of each intermediate trace component is calculated, and according to a second preset threshold, the intermediate trace components with an average variance contribution rate greater than the threshold are regarded as key components that have a significant impact on the intermediate quality fluctuation.

6. The method for controlling the production process of a pharmaceutical intermediate according to claim 4 is characterized in that: The specific steps in step A212 include: According to the determined key components, the covariance matrix of the key components is constructed by calculating the covariance between any two key components; The Jacobi iteration method is used to perform eigenvalue decomposition on the covariance matrix to obtain multiple eigenvalues ​​and corresponding eigenvectors, and all eigenvectors are sorted in descending order of eigenvalues; The eigenvectors corresponding to the first N eigenvalues ​​are selected as the eigenvectors of the principal components, and the variance contribution rates of the key components corresponding to the first N eigenvalues ​​are ensured to reach a preset ratio, where N is the minimum value when the cumulative variance contribution rate reaches the preset ratio.

7. The method for controlling the production process of a pharmaceutical intermediate according to claim 2 is characterized in that: The specific steps in step A4 include: A41. Divide historical production data into K non-overlapping subsets, each containing data from a different batch of intermediates. A42. Select the i-th subset as the validation set and the remaining K-1 subsets as the training set, where i ranges from 1 to K. A43. Based on the training set, use the quantitative association model to predict the amount of specific impurities generated in the downstream reaction unit of each batch in the validation set to obtain a set of predicted values; A44. Calculate the prediction error for the i-th subset using the root mean square error algorithm based on the predicted value set and the actual impurity generation levels for each batch in the validation set. A45. Repeat steps A42 to A44 to calculate the prediction errors corresponding to the K subsets; A46. Calculate the average of the K prediction errors as the final prediction error of the cross-validation; A47. Determine whether the final prediction error exceeds a first preset threshold.

8. The method for controlling the production process of a pharmaceutical intermediate according to claim 1 is characterized in that: The step of calculating the adjustment amount of the downstream reaction unit operating parameter based on the predicted specific impurity generation amount and the set impurity total amount control target includes: Determine the operating parameter type of the downstream reaction unit, the operating parameter type including reaction temperature, reaction time, and catalyst dosage; According to the predicted specific impurity generation amount, based on the pre-established correlation function between the specific impurity generation amount and each operating parameter type, the influence of each operating parameter type on the impurity generation is obtained and expressed as an adjustment coefficient; According to the set total impurity control target, combined with the adjustment coefficient of each operating parameter type, the adjustment amount of each operating parameter type is calculated.

9. The pharmaceutical intermediate production process control method according to claim 8 is characterized in that the adjustment amount must be sufficient to ensure that the final total amount of impurities does not exceed the control target, and the adjustment range of each operating parameter is within a preset range, and the operating parameter with the greatest impact on impurity generation is adjusted first.

10. A pharmaceutical intermediate production process control system, characterized in that: include: An acquisition module is used to obtain the intermediate trace component information of the intermediate sample produced by the upstream separation and purification unit and form a batch fingerprint data set; Establishing a module for establishing a quantitative correlation model between a batch fingerprint data set and the amount of specific impurities generated in a downstream reaction unit; The prediction module is used to input the intermediate trace component information of the current batch of intermediate samples into the quantitative correlation model to predict the amount of specific impurities generated in the downstream reaction unit; A calculation module, used to calculate the adjustment amount of the operating parameters of the downstream reaction unit based on the predicted specific impurity generation amount and the set impurity total amount control target; The control module is used to send the adjustment amount to the actuator of the downstream reaction unit to adjust the actual operating parameters of the downstream reaction unit.

Citation Information

Patent Citations

  • Production quality prediction method based on kernel principal component analysis and multiple linear regression

    CN118761522A

  • Intelligent monitoring, regulating and controlling system for medical intermediate synthesis production line

    CN119002404A

  • Method and system for detecting impurities in glucosamine production process

    CN119619020A

  • Method for analyzing relation between process impurities and degraded impurities based on big data analysis technology

    CN119811549A

  • Oil depot whole-process comprehensive management monitoring method and system

    CN120046923A