Analysis Method for the Relationship between Process Impurities and Degradation Impurities Based on Big Data Analysis Technology

Through big data analysis technology, the correlation coefficients of process impurities and degraded impurities are calculated, and key impurities and process control points are screened out, which solves the blind problem of impurities research in drug quality control and achieves improvements in drug quality and safety.

CN119811549BActive Publication Date: 2025-08-01NAT INST FOR FOOD & DRUG CONTROL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510288211.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-08-01
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing technology lacks effective big data analysis methods to understand and control process impurities and degraded impurities in drug production, resulting in unclear key process control points and affecting drug quality and safety.

Method used

Using a method based on big data analysis technology, the correlation coefficients of process impurities and degraded impurities are calculated through typical correlation analysis methods, key impurities and process control points are selected, and the population correlation analysis method is established to reveal the intrinsic correlation between impurities and processes.

Benefits of technology

It provides a scientific basis for drug quality control, guides the control and process optimization of key drug impurities, improves drug safety and effectiveness, and promotes the development of pharmaceutical science and pharmaceutical industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811549B_ABST
    Figure CN119811549B_ABST
Patent Text Reader

Abstract

The present invention relates to an analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology. The correlation coefficients among process impurities, typical impurity variables, and degradation impurities are calculated through the canonical correlation analysis method, and then the transfer relationships among various impurities, typical impurity variables, and various degradation impurities are constructed by screening the correlation coefficients, so as to judge the correlation degree between process impurities and degradation impurities. The analysis method involved in the present invention comprehensively analyzes the impurity formation rules of drugs during the production, storage, and use processes, deeply explores the correlation relationship between process impurities and degradation impurities, reveals the internal relevance of different impurities, which not only helps to guide the direction of impurity profile research, provides a scientific basis for the control of key impurities of drugs and process optimization, but also helps to improve the safety and effectiveness of drugs, and promotes the further development of pharmaceutical science and the pharmaceutical industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data analysis technology, and in particular relates to an analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology. Background Art

[0002] Drug quality control remains a crucial component of modern drug development and production. With advances in science and technology and the development of data analysis techniques, big data-driven drug quality research has become a research hotspot and a crucial tool for drug development, safety, and efficacy assessment. However, drug quality big data research is a cutting-edge, interdisciplinary field still in its infancy. Effective big data analysis methods that can address specific issues are still lacking, hindering the expected impact of big data in practical drug research.

[0003] Pharmaceutical impurities are any substances that affect drug purity during production, storage, and use. The presence of impurities not only reduces drug purity but can also trigger adverse reactions, seriously impacting drug safety and efficacy. Impurities are primarily categorized as degradation impurities and process impurities based on their source. These impurities differ significantly in their sources, formation mechanisms, control measures, and monitoring priorities.

[0004] Process impurities are primarily introduced during the production process and may originate from raw materials, solvents, reagents, intermediates, by-products, containers, and more. Degradation impurities are primarily introduced during storage and may originate from external conditions, such as improper packaging, transportation, and storage, as well as environmental factors such as light, temperature, air, and microorganisms. While process impurities and degradation impurities differ in their sources and characteristics, they are both essential components of drug therapeutic control. Therefore, a comprehensive understanding and effective control of these two types of impurities are crucial to ensuring drug safety, efficacy, and quality control.

[0005] The currently recognized basic approach to impurity research and control is: starting with the analysis of the source of impurities, combining the actual production process and structural characteristics of the product to analyze the various potential impurities such as reaction materials, intermediates, by-products, and degradation products that may be present in the product, and through a comprehensive analysis of the impurity spectrum to grasp the impurity profile of the product, according to the risk level of each type of potential impurity, to establish appropriate analytical methods to ensure the effective detection and confirmation of each potential impurity. At the same time, through impurity spectrum analysis, the possible sources and destinations of impurities can be clarified, and targeted control measures can be set for the corresponding impurities in the preparation process. Through the study of product packaging and storage conditions, the degradation of the drug can be effectively inhibited, achieving systematic and comprehensive control of impurities in the drug from the source of impurity generation, and ensuring the quality and safety of the drug.

[0006] The study of impurity sources is a systematic project and is very complex. At present, the study of impurity profiles is basically carried out on a case-by-case basis. In most cases, it is like the blind men feeling an elephant, without a clear research direction. Often, the research results cannot obtain key impurities and key process control points. As an important pharmaceutical parameter affecting drug quality, controlling the key impurities of products and correlating them with the key process control points in the production process is an important part of quality control and also the focus of process evaluation. At present, there is still a lack of effective technical means and methods for the general study of the correlation between impurities and processes, resulting in unclear key process control points in drug quality control and drug production. Domestic drugs have always been unclear about how to improve in terms of processes; and the stability data in drug quality research has not been fully utilized, making it difficult to trace and improve the key problems and causes in the process. Summary of the Invention

[0007] To solve the problems such as the inability to obtain key process control points of drugs, the insufficient utilization of relevant stability data analysis, and the prediction of relevant safety and effectiveness risks, the present invention provides an analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology.

[0008] The technical solution adopted by the present invention is: an analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology, and the specific steps are as follows:

[0009] S1: Summarize the types of process impurities and degradation impurities according to the product synthesis process;

[0010] S2: Correlate the process impurity information with the degradation impurity information, and obtain the correlation coefficients among the process impurities, typical impurity variables, and degradation impurities through the canonical correlation analysis method;

[0011] S3: Determine the transfer relationship according to the correlation coefficients among the process impurities, typical impurity variables, and degradation impurities, and screen the correlation degrees between specific types of degradation impurities and specific types of process impurities.

[0012] Preferably, step S2 includes,

[0013] S2.1: Using the contents of process impurities and degradation impurities in drugs as the big data basis, extract multiple typical impurity variables therefrom, and calculate the correlation coefficients of each typical impurity variable; [[ID=2)]

[0014] S2.2: Calculate the conversion coefficients of each process impurity and each degradation impurity through the canonical correlation analysis method; generate multiple typical variable relationship formulas;

[0015] S2.3: Test the generated typical variable relationship formulas, and screen the typical impurity variables corresponding to the typical variables with p values less than 0.05;

[0016] S2.4: For the selected typical impurity variables, construct the correlation coefficient matrices between the typical impurity variables and process impurities and degradation impurities respectively, and obtain the correlation coefficients among process impurities, typical impurity variables, and degradation impurities.

[0017] Preferably, the canonical correlation coefficients are tested by the likelihood ratio method to screen the typical impurity variables.

[0018] Preferably, the typical impurity variables include typical variables of degradation impurities and typical variables of process impurities.

[0019] Preferably, in step S2.4, calculate the correlation coefficient matrix between degradation impurities and typical variables of degradation impurities, the correlation coefficient matrix between process impurities and typical variables of process impurities, the correlation coefficient matrix between degradation impurities and typical variables of process impurities, and the correlation coefficient matrix between process impurities and typical variables of degradation impurities respectively.

[0020] Preferably, step S3 includes

[0021] S3.1: Screen the correlation coefficients not less than 0.3 as alternative transfer relationships;

[0022] S3.2: It is determined that the degradation impurities and process impurities with correlation coefficients that are both positively correlated or both negatively correlated with the same set of typical impurity variables are correlated.

[0023] Preferably, in step S3.1, screen the correlation coefficients not less than 0.3 in the correlation coefficient matrix between degradation impurities and typical variables of process impurities and the correlation coefficient matrix between process impurities and typical variables of degradation impurities, and screen the corresponding degradation impurities or process impurities;

[0024] In step S3.2, based on the correlation coefficients in the correlation coefficient matrix between degradation impurities and typical variables of degradation impurities and the correlation coefficient matrix between process impurities and typical variables of process impurities, compare whether the degradation impurities and process impurities for the same set of typical impurity variables are both positively correlated or both negatively correlated.

[0025] Preferably, the typical impurity variables include typical variables of degradation impurities and typical variables of process impurities, and the canonical correlation coefficients of the correlated process impurities and degradation impurities are equal to the canonical correlation coefficient between the typical variables of degradation impurities and the typical variables of process impurities in the corresponding typical impurity variables.

[0026] The advantages and positive effects of the present invention are as follows: Based on big data analysis, the impurities are correlated with process characteristics, and the correlation analysis technology is applied to conduct population correlation analysis from an overall perspective, establishing an analysis method for the correlation between impurities and population processes. Information such as key process control points and key impurities obtained has guiding significance and application value for pharmaceutical process improvement, impurity profile research, and pharmaceutical quality evaluation, providing an objective basis and reference direction for how to improve product quality and process level, and promoting continuous process improvement. Description of the Drawings

[0027] Figure 1 Analysis method flow of the relationship between process impurities and degradation impurities;

[0028] Figure 2 First canonical variate relationship diagram of degradation impurities and process impurities of aztreonam for injection;

[0029] Figure 3 Second canonical variate relationship diagram of degradation impurities and process impurities of aztreonam for injection;

[0030] Figure 4 Third canonical variate relationship diagram of degradation impurities and process impurities of aztreonam for injection;

[0031] Figure 5 Transmission diagram of the correlation relationship between degradation impurity-process impurity variables and canonical variates. Detailed Embodiments

[0032] The embodiments of the present invention will be described below with reference to the drawings.

[0033] The present invention discloses an analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology. It comprehensively analyzes the impurity formation rules during the production, storage, and use of drugs, deeply explores the correlation between process impurities and degradation impurities, reveals the internal correlation and possible formation mechanisms of different impurities, which not only helps to guide the direction of impurity profile research and provides a scientific basis for the control of key impurities and process optimization of drugs, but also helps to improve the safety and effectiveness of drugs and promote the further development of pharmaceutical science and the pharmaceutical industry.

[0034] First, classify and name the process impurities and degradation impurities, calculate the correlation coefficients between process impurities, canonical impurity variables, and degradation impurities through canonical correlation analysis method, and then screen through the correlation coefficients to construct the transmission relationship between various process impurities, canonical impurity variables, and various degradation impurities, so as to judge the correlation degree between process impurities and degradation impurities. The process is as Figure 1 shown, and the specific analysis method is as follows. I

[0035] S1: Summarize the types of process impurities and degradation impurities according to the product synthesis process and name them.

[0036] First, publicly available data is used to collect drug synthesis processes and all relevant impurity information, and the impurity data is cleaned and classified. Relevant information has the characteristics of being messy, complex, and highly specialized. The sources of data currently available include impurity spectrum research data, synthesis process information, and related mechanism research information, etc. The main sources are: stability study data for new drug approvals at home and abroad, research literature related to impurities and synthesis processes, quality review data, national evaluation sampling data, CAS information, etc. In addition to multiple data sources, the key to the present invention is the need to obtain group drug data, that is, it cannot be the data of a single manufacturer or a few manufacturers, but should be from the perspective of variety, as much as possible to collect information on all products on the market for research, so as to avoid the one-sidedness of the data causing the reliability of the big data analysis results.

[0037] To ensure the authenticity and reliability of the data obtained, it is necessary to clean the acquired public data and verify the correspondence between different impurity data naming systems and structures. Impurity naming formats include "Impurity A, B, C...", "Impurity 1, 2, 3...", and "Impurity I, II, III..." However, it is worth noting that the impurity naming systems used in different studies correspond to different impurities. Therefore, the correspondence between impurities in data from different sources should be ensured based on the impurity structure information to obtain basic big data information related to impurities.

[0038] S2: Correlate the process impurity information with the degradation impurity information, and obtain the correlation coefficients among process impurities, typical impurity variables, and degradation impurities through canonical correlation analysis.

[0039] The impurity content data obtained was categorized into degradation impurities and process impurities. Process impurities not only reflect the control level of the drug production process but are also closely related to the stability of different raw material processes and are intrinsically linked to degradation impurities. Using drug process impurities as one set of variables and degradation impurities as another, canonical correlation analysis was used to explore the relationship between the two sets of variables. The correlation between different types of process impurities and degradation impurities was explored, and the inherent connection between the two was studied, thereby revealing the correlation between impurities and process.

[0040] S2.1: Based on the content of process impurities and degradation impurities in drugs, extract multiple typical impurity variables and calculate the correlation coefficient of each typical impurity variable.

[0041] S2.2: Calculate the conversion coefficients for various process impurities and various degradation impurities using canonical correlation analysis methods; generate canonical variable relationship formulas.

[0042] Taking the degradation impurity factor as the independent variable X and the process impurity factor as the dependent variable Y, according to the correlation relationship between the two sets of variables, obtain the canonical variables to replace the original variables of process impurities and degradation impurities, and then project the relationship between process impurities and degradation impurities onto the relationship of canonical variables to obtain the conversion coefficients between the original variables and the canonical variables. The calculation formulas for the canonical degradation impurity variable U and the canonical process impurity variable V are as follows:

[0043] (Equation 1);

[0044] (Equation 2);

[0045] where i = 1, 2, …, M, M is the number of common factors, and j = 1, 2, …, N, N is the number of observed samples.

[0046] S2.3: Test the generated formula for the relationship between canonical variables; specifically, test the canonical correlation coefficient by the likelihood ratio method; screen the canonical impurity variables corresponding to the canonical variables with a p-value less than 0.05. Through the correlation analysis and hypothesis testing between canonical variables, study the correlation between process impurities and degradation impurities. Assume H0, the canonical correlation coefficients corresponding to the current and subsequent canonical variables are all 0; use the likelihood ratio method to test the canonical correlation coefficient, and the obtained likelihood ratio statistic approximately follows the F distribution and distribution, and retain the canonical impurity variables with a p-value less than 0.05 for the subsequent analysis of the relationship between process impurities and degradation impurities.

[0047] S2.4 Construct and screen the correlation coefficient matrices between the canonical impurity variables and process impurities and degradation impurities respectively, and obtain the correlation coefficients among process impurities, canonical impurity variables, and degradation impurities; analyze the canonical correlation structure, explore the transfer relationship of original variable - canonical variable - original variable, and further evaluate the statistical significance of r1 and r2. Among them, the canonical impurity variables include the canonical variables of degradation impurities and the canonical variables of process impurities. The canonical correlation structure is the pairwise correlation coefficient matrix between each original variable and the canonical variable.

[0048] S3: Determine the transfer relationship according to the correlation coefficients among process impurities, canonical impurity variables, and degradation impurities, and screen the correlation degree between specific categories of degradation impurities and specific categories of process impurities.

[0049] The correlation between the original variables and the canonical variables may need to be transmitted through multiple canonical correlation coefficients (r1, r2, r3...). The correlation relationship transmission diagram of each variable and the canonical variables obtained through canonical correlation analysis includes the correlation relationship between the canonical variables of degradation impurities and process impurities, the corresponding relationship between degradation impurities and the canonical variables of degradation impurities, and the corresponding relationship between process impurities and the canonical variables of process impurities. Select the correlation coefficients with a correlation coefficient not less than 0.3 as the alternative transmission relationships. Generally, it is considered that there is a substantial correlation when the absolute value of the correlation coefficient is greater than 0.30, a significant correlation when it is greater than 0.50, and a high correlation when it is greater than 0.80. The canonical impurity variables include the canonical variables of degradation impurities and the canonical variables of process impurities; among them, the correlation transmission route with the same positive and negative correlation is the process control route that should be focused on. Pay attention to the significant correlation relationship transmission route with a canonical correlation coefficient r greater than 0.5, and analyze the corresponding relationship between degradation impurities and process impurities. The degradation impurities and process impurities with the same positive or negative correlation coefficient with the same set of canonical impurity variables are correlated. Its canonical correlation coefficient is equal to the canonical correlation coefficient between the canonical variables of degradation impurities and the canonical variables of process impurities in the corresponding canonical impurity variables.

[0050] Specifically, screen the correlation coefficients not less than 0.3 in the correlation coefficient matrix of the canonical variables of degradation impurities - process impurities and the correlation coefficient matrix of the canonical variables of process impurities - degradation impurities, and screen the corresponding degradation impurities or process impurities; based on the correlation coefficients in the correlation coefficient matrix of degradation impurities and the canonical variables of degradation impurities, and the correlation coefficient matrix of process impurities - the canonical variables of process impurities, compare whether the degradation impurities and process impurities for the same set of canonical impurity variables are both positive correlation coefficients or both negative correlation coefficients. The degradation impurities and process impurities screened for the same canonical variable are the mutually associated process impurities and degradation impurities. Judge the key points of process control and the focus of the process according to these corresponding relationships, and clarify the direction of impurity profile research.

[0051] The analysis method of the relationship between process impurities and degradation impurities based on big data analysis technology correlates the process characteristics of a group of drugs (process impurities) with the stability of individual drugs (degradation impurities), establishes a population correlation analysis method from an overall perspective, obtains key process control points through a canonical correlation analysis method based on process impurities and degradation impurities, and achieves the expected big data analysis effect. The population correlation analysis method requires that the process population has high representativeness and is less affected by the scale of individual process samples; the correlation analysis method is not affected by the limitation of data resources, only seeks internal correlations within the selected range, and the result has good reliability; using the correlation relationship to replace the causal relationship, the result reveals possible objective phenomena and gives directions for further research and evaluation. Using the thinking of big data analysis, first, utilize the canonical correlation relationship between the degradation impurity variable and the process impurity variable to mine the multi-dimensional complex correlation relationship between the process characteristics of a group of drugs (process impurities) and the stability of individual drugs (degradation impurities) in the data of a group of drugs, and then use the correlation transfer information between the population process characteristics (process impurities) and the stability of individual drugs (degradation impurities) to establish the transfer relationship between process impurities and degradation impurities, screen key process control points, and break through the technical limitations of traditional methods.

[0052] Taking aztreonam in industrial production as an example below, the specific implementation of the present invention will be described in conjunction with the accompanying drawings.

[0053] Example 1: Obtaining big data related to impurities

[0054] The key to the industrial production of aztreonam lies in the synthesis of the main ring of aztreonam. Currently, there are six relatively mature process routes, namely: sulfur trioxide cyclization method (sulfuric acid monoester), sulfamic acid process, O-benzylhydroxylamine process (ring-opening side reaction), [2+2] reaction preparation process, tert-butoxycarbonyl protection process, and chloromethylbenzyl ester protection process. The above six process routes are the main sources of process impurities. Among them, process impurities include reaction by-products (desulfonated aztreonam, aztreonam methyl ester, aztreonam ethyl ester) and reaction intermediates (aztreonam tert-butyl ester). Degradation impurities are impurities generated under strong degradation or accelerated test conditions and obtained through structure confirmation and verification, such as polymers. Including process impurities and degradation impurities, the 11 confirmed impurities in aztreonam for injection are specifically shown in Table 1; 188 samples were detected, and the content of each impurity is shown in Table 2.

[0055] Table 1 Detected impurities and structures of aztreonam for injection

[0056]

[0057]

[0058]

[0059] Table 2 Impurity content in 188 samples

[0060]

[0061]

[0062]

[0063]

[0064]

[0065] Comparative example: Simple correlation analysis

[0066] By simple correlation analysis of the data on the content of various impurities in drug samples, the data obtained are shown in Table 3. From the content of Table 3, the direct correlation relationships between various impurities can be inferred. For example, there is a correlation between process impurities IMP10, IMP11 and degradation impurity IMP9, and there is a correlation between process impurity IMP5 and degradation impurities IMP2, IMP3, etc. Coupled with the known degradation relationship between IMP1 and IMP8 reported in the literature, this information is less, and key process control points cannot be obtained, failing to achieve the expected purpose of guiding and improving the process.

[0067] Table 3 Simple correlation analysis results of aztreonam for injection impurities

[0068]

[0069] Example 2 Analysis of the relationship between process impurities and degradation mass spectrometry of aztreonam for injection based on big data analysis technology

[0070] 2.1 Big data collation of aztreonam for injection

[0071] The impurity spectrum of aztreonam for injection is divided into a process impurity variable group and a degradation impurity variable group. The process impurity variable group contains 4 variables: IMP5, IMP7, IMP10, IMP11; the degradation impurity variable group contains 7 variables: IMP1, IMP2, IMP3, IMP4, IMP6, IMP8, IMP9.

[0072] 2.2 Establishment of population correlation analysis method

[0073] Based on the data in Table 2, using the method of canonical correlation analysis, extract canonical variables to explore the internal relationship between process impurities and degradation impurities. The conversion coefficients between the original impurity variables and the canonical variables are shown in Table 4. A total of 4 pairs of canonical variables are extracted. The first and second canonical correlation coefficients are 0.7737 and 0.6980 respectively, showing significant correlation; the third canonical correlation coefficient is 0.4502, showing substantial correlation; the fourth canonical correlation coefficient is 0.1206, showing no substantial correlation. The cumulative contribution rate of the first, second, and third canonical correlation coefficients reaches 78.57%. Select the first, second, and third canonical variables for subsequent analysis.

[0074] Table 4 Results of Canonical Correlation Analysis of Process Impurities and Degradation Impurities in Aztreonam for Injection

[0075]

[0076] The relationships of the first, second, and third pairs of canonical variables are shown in Formulas 3 to 8.

[0077] (Formula 3);

[0078] (Formula 4);

[0079] (Formula 5);

[0080] (Formula 6);

[0081] (Formula 7);

[0082] (Formula 8);

[0083] The relationship diagrams of the three pairs of canonical variables are as Figures 2 - 4 shown, indicating that the correlation relationships of the three pairs of canonical variables are obvious.

[0084] 2.3 Hypothesis Testing in Big Data Analysis

[0085] Conduct hypothesis testing on the canonical correlation analysis. Assume that the current and subsequent canonical correlation coefficients are 0. In this paper, the likelihood ratio method is used to test the canonical correlation coefficients. The obtained likelihood ratio statistic approximately follows the F-distribution and distribution. The results are shown in Table 5. The F-tests of r1 to r3 and the test results are consistent, and the p-values are all 0.0000, indicating that the test and the hypothesis have extremely significant differences (generally, p < 0.05 indicates significant differences, and p < 0.01 indicates extremely significant differences), and the hypothesis is not valid; the F-test of r4 and The test results show that the p values are 0.6174 and 0.6029 respectively, both greater than 0.05, indicating that there is no significant difference between the test and the hypothesis, and the hypothesis holds. Therefore, the canonical correlation coefficients r1, r2, and r3 of the degradation impurities and process impurities are statistically significant, while r4 is not.

[0086] Table 5 Likelihood Ratio Test Statistic for Degradation Impurities and Process Impurities of Aztreonam for Injection

[0087]

[0088] 2.4 Multidimensional Correlation Results of Big Data Analysis

[0089] The canonical correlation structures of the first, second, and third pairs of canonical variables are calculated as shown in Table 6.

[0090] Table 6 Canonical Correlation Structure between Process Impurities and Degradation Impurities of Aztreonam for Injection

[0091]

[0092] From the analysis of the canonical variable correlation coefficient matrix of process impurities - degradation impurities and the canonical variable correlation coefficient matrix of degradation impurities - process impurities, there is a substantial correlation between degradation impurities IMP3, IMP4 and canonical variable 1 of process impurities; there is a substantial correlation between degradation impurities IMP1, IMP2 and canonical variable 2 of process impurities; there is a substantial correlation between degradation impurity IMP2 and canonical variable 3 of process impurities; there is a substantial correlation between process impurities IMP5, IMP10, IMP11 and canonical variable 1 of degradation impurities; there is a substantial correlation between process impurity IMP7 and degradation impurity 2; there is a substantial correlation between process impurities IMP5, IMP11 and canonical variable 3 of degradation impurities. Draw the transfer diagram of the correlation relationship between degradation impurity - process impurity variables and canonical variables as Figure 5 shown.

[0093] Data analysis of the correlation coefficient matrix of degradation impurities - typical variables of degradation impurities and process impurities - typical variables of process impurities. Among them, the correlation coefficient between degradation impurity IMP4 and typical variable 1 of degradation impurities is negative, and the correlation coefficients between process impurities IMP5, IMP10, and IMP11 and typical variable 1 of process impurities are all negative, with the same correlation direction. It can be seen that degradation impurity IMP4 is correlated with process impurities IMP5, IMP10, and IMP11, and the canonical correlation coefficient reaches 0.7737, showing a significant correlation. The correlation coefficients between degradation impurities IMP1 and IMP2 and typical variable 2 of degradation impurities are both positive, and the correlation coefficient between process impurity IMP7 and typical variable 2 of process impurities is positive, with the same correlation direction. It can be seen that degradation impurities IMP1 and IMP2 are correlated with process impurity IMP7, and the canonical correlation coefficient reaches 0.6980, showing a significant correlation. The correlation coefficient between degradation impurity IMP2 and typical variable 3 of degradation impurities is negative, and the correlation coefficient between process impurity IMP5 and the typical variable of process impurities is negative. It can be seen that degradation impurity IMP2 is correlated with process impurity IMP5, and the canonical correlation coefficient is 0.4502, showing a substantial correlation.

[0094] It should be noted that the pairwise correlation analysis results of IMP4 with IMP5 and IMP10 (Table 3) did not show a substantial correlation. Therefore, we speculate that IMP5 and IMP10 may come from the same process, and their combined action may be a necessary condition for the formation of the ring-opening degradation product IMP4. Further research and discussion can be carried out on the relevant mechanism.

[0095] Example 3: Analysis of key impurities, key process control points, and the research direction of impurity profiles

[0096] Calculate the canonical correlation determination coefficient. The results are shown in Table 7. Degradation impurities IMP3 and IMP4 are correlated with process impurities IMP5 and IMP10 through the first pair of canonical variables, with a strong correlation; degradation impurities IMP1 and IMP2 are correlated with process impurity IMP7 through the second pair of canonical variables, with a strong correlation; degradation impurities IMP2 and IMP4 are correlated with process impurity IMP11 through the third pair of canonical variables, with a weak correlation; the effects of other impurities are very small. The research direction of impurity profiles can be obtained by analyzing key impurities and key process control points.

[0097] Table 7 Canonical correlation determination coefficients of process impurities - degradation impurities of aztreonam for injection

[0098]

[0099] The above has described the embodiments of the present invention in detail, but the above content is only the preferred embodiments of the present invention and should not be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. An analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology, characterized in that: The specific steps are as follows: S1: Summarize the types of process impurities and degradation impurities according to the product synthesis process; S2: Correlate the process impurity information with the degradation impurity information, and obtain the correlation coefficients among process impurities, typical impurity variables, and degradation impurities through the canonical correlation analysis method; S2.1: Based on the content of process impurities and degradation impurities in the drug as big data, extract multiple typical impurity variables from it, and calculate the correlation coefficients of each typical impurity variable; S2.2: Calculate the conversion coefficients of various process impurities and various degradation impurities through the canonical correlation analysis method; Generate the canonical variable relationship formula; Taking the degradation impurity factor as the independent variable X and the process impurity factor as the dependent variable Y, according to the correlation relationship between the two sets of variables, obtain the canonical variables, replace the original variables of process impurities and degradation impurities, and then project the relationship between process impurities and degradation impurities onto the relationship of canonical variables to obtain the conversion coefficients between the original variables and the canonical variables; The calculation formulas for the canonical degradation impurity variable U and the canonical process impurity variable V are respectively: (Formula 1); (Formula 2); where i = 1, 2, …, M, M is the number of common factors, j = 1, 2, …, N, N is the number of observation samples; S2.3: Test the generated canonical variable relationship formula; specifically, test the canonical correlation coefficients through the likelihood ratio method; screen the typical impurity variables corresponding to the canonical variables with p-values less than 0.05; S2.4: Construct and screen the correlation coefficient matrices of the typical impurity variables with process impurities and degradation impurities respectively, and obtain the correlation coefficients among process impurities, typical impurity variables, and degradation impurities; S3: Determine the transfer relationship according to the correlation coefficients among process impurities, typical impurity variables, and degradation impurities, and screen the correlation degrees between specific categories of degradation impurities and specific categories of process impurities.

2. The analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology according to claim 1, wherein: The typical impurity variables include the canonical degradation impurity variables and the canonical process impurity variables.

3. The analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology according to claim 2, wherein: In step S2.4, calculate the correlation coefficient matrix of degradation impurity - canonical degradation impurity variables, the correlation coefficient matrix of process impurity - canonical process impurity variables, the correlation coefficient matrix of degradation impurity - canonical process impurity variables, and the correlation coefficient matrix of process impurity - canonical degradation impurity variables respectively.

4. The analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology according to claim 1, characterized in that: Step S3 includes, S3.1: Screen the correlation coefficients not less than 0.3 as alternative transfer relationships; S3.2: It is determined that the degradation impurities and process impurities with correlation coefficients that are both positively correlated or both negatively correlated with the same set of canonical impurity variables have a transfer relationship.

5. The analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology according to claim 4, wherein: In step S3.1, screen the correlation coefficients not less than 0.3 in the correlation coefficient matrix of degradation impurity - canonical process impurity variables and the correlation coefficient matrix of process impurity - canonical degradation impurity variables, and screen the corresponding degradation impurities or process impurities; In step S3.2, based on the correlation coefficients in the correlation coefficient matrix of degradation impurity - canonical degradation impurity variables and the correlation coefficient matrix of process impurity - canonical process impurity variables, compare whether the degradation impurities and process impurities for the same set of canonical impurity variables are both positively correlated or both negatively correlated.

6. The analysis method for the relationship between process impurities and degradation impurities based on big data analysis technology according to claim 4 or 5, characterized in that: Typical impurity variables include typical variables of degradation impurities and typical variables of process impurities. The canonical correlation coefficient of process impurities and degradation impurities with correlation is equivalent to the canonical correlation coefficient between the typical variable of degradation impurities and the typical variable of process impurities in the corresponding typical impurity variables.

Citation Information

Patent Citations

  • Optimization method and optimization system for preparing high-purity metal through vacuum distillation

    CN117577229A

  • Analytical method for efficiently measuring impurities in aspirin and aspirin preparation

    CN117783379A