Cytochrome P450 metabolism analysis method based on Raman spectrum and machine learning model
Through methods based on Raman spectroscopy and machine learning models, the changes in CYPs mediated drug metabolism in cells are quickly detected, solving the problems of slow detection speed and cumbersome data analysis in the prior art, and achieving efficient and rapid drug metabolism analysis.
Patent Information
- Application Number
- CN202510084703.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems such as slow detection speed, cumbersome data analysis and high cost in analyzing CYPs-mediated drug metabolism in cells, making it difficult to quickly obtain effective data.
The cytochrome P450 metabolism analysis method based on Raman spectroscopy and machine learning models was used to construct cells that stably express CYPs, collect Raman spectroscopy data, perform principal component analysis and K nearest neighbor classification to achieve rapid detection of changes in CYPs-mediated drug metabolism in cells.
It realizes rapid detection of CYPs-mediated drug metabolism in cells, with the advantages of no labeling, simple operation and wide application range, and can effectively identify the intrinsic link between drug metabolism and CYPs enzyme activity.
Smart Images

Figure CN119985435A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a cytochrome P450 metabolism analysis method based on Raman spectroscopy and a machine learning model, and belongs to the technical field of medical detection. Background Art
[0002] The cytochrome P450 (CYPs) enzyme system belongs to the superfamily of heme-thiolate proteins. More than a thousand species have been discovered so far, and they are widely distributed in various biological organisms. CYPs participate in the metabolism of a variety of exogenous and endogenous substances, and participate in the metabolism of more than 55% of drugs. They are key enzymes in the drug metabolism process. CYPs mainly participate in the phase I metabolism of drugs, catalyzing the oxidation, reduction and hydrolysis reactions of various compounds. The generated compounds are made more soluble in water, which is convenient for coupling with further phase II reactions. The enzymes involved in drug metabolism are mainly CYP1, CYP2 and CYP3 families, among which CYP3A4, CYP1A2, CYP2E1, CYP2C9, CYP2D6 and other subtypes have been studied more. It is worth noting that some compounds will produce more toxic compounds after being metabolized by CYPs, which may be the cause of various diseases.
[0003] Although in recent years, transgenic experimental animal models of humanized CYPs have provided powerful tools for studying the functions and mechanisms of CYPs by combining high-throughput sequencing technology, proteomics analysis, metabolomics and other analytical tools, these methods still have certain limitations in practical applications. In terms of models expressing CYPs, in order to solve the problem of species differences between commonly used experimental animals and human CYPs, there are high technical barriers and experimental costs to establish transgenic animal models, and the experimental cycle is long. The cost of commercial transgenic animals is too high. In addition, due to the complexity of the overall experimental animal system, it is difficult to obtain a large amount of valid data in a short period of time. In terms of in vitro cell models, most cell lines will lose CYPs expression during long-term culture. The HepaRG model expressing CYPs has problems such as long CYPs induction cycle and high culture cost. The in vitro constructed cell model that stably overexpresses CYPs has the characteristics of short experimental cycle, low cost, easy operation and wide range of applications. By analyzing the changes in different components in cell samples, it is possible to more intuitively study the functional changes of CYPs and their effects on related biological processes, providing strong experimental basis and theoretical support for its drug metabolism, disease mechanism exploration and formulation of potential treatment strategies.
[0004] In terms of analytical tools, although high-throughput sequencing and proteomics analysis can provide a large amount of data, the data analysis and verification process is often cumbersome and time-consuming. Raman spectroscopy is a spectral analysis technology based on the Raman scattering phenomenon. This technology uses laser irradiation to cause the molecules to undergo energy level transitions such as vibration and rotation, thereby generating energy scattered light. By analyzing the frequency changes of scattered light, information such as the molecular structure, chemical composition and physical state of the sample can be obtained. Label-free detection of samples by Raman spectroscopy can avoid interference from marker-related substances and simplify the operation steps. However, the disadvantage of label-free detection is that when the sample size is large and the sample components are complex, it is often difficult to identify small differences between different spectra. Therefore, it is necessary to combine machine learning methods to accurately identify key information in the sample from multiple variables, improve analysis efficiency, and thus reveal the intrinsic connection between drug metabolism and CYPs enzyme activity.
[0005] Therefore, it is necessary to establish a cytochrome P450 metabolism analysis method based on Raman spectroscopy and machine learning models to quickly analyze the changes in intracellular substances in CYPs-mediated drug metabolism in cells. Summary of the invention
[0006] Purpose of the invention: The technical problem to be solved by the present invention is to provide a cytochrome P450 metabolism analysis method based on Raman spectroscopy and machine learning model, which can realize rapid detection of CYPs-mediated drug metabolism in cells.
[0007] Technical solution: To solve the above technical problems, the present invention provides a cytochrome P450 metabolism analysis method based on Raman spectroscopy and machine learning model, comprising the following steps:
[0008] (1) Construction of cells that stably express CYPs;
[0009] (2) Collecting cells after drug treatment;
[0010] (3) collecting Raman spectroscopy data of cells;
[0011] (4) Feature extraction and dimensionality reduction of the collected Raman spectral data through principal component analysis;
[0012] (5) Use the K nearest neighbor algorithm to classify the Raman spectral features after dimensionality reduction.
[0013] Among them, the cell construction method described in step (1) includes: co-transfecting the packaging plasmids psPAX2 and pMD2.G with the target plasmid into HEK293T cells, collecting the virus-containing culture medium supernatant and then infecting the cells, and obtaining a stable cell line expressing the above-mentioned CYPs through monoclonal screening.
[0014] Wherein, the target plasmid includes pLenti-CYP2E1, pLenti-CYP3A4, pLenti-CYP1A2 or pLenti-CYP2C9.
[0015] The cells described in step (1) include human liver cancer HepG2 cell line, HLE cells, Hep3B cells and other cell models commonly used to study liver-related drug metabolism.
[0016] After the cell construction is completed, the cells in the logarithmic growth phase are cultured, the culture medium supernatant is discarded, and fresh culture medium containing drugs is added, and the cells are collected after continued cultivation.
[0017] The step (3) of collecting Raman spectral data of cells includes: dropping the cells onto a silicon wafer and drying them at low temperature, and then collecting spectra of sample points on the silicon wafer; wherein the spectral collection conditions are: using a 785nm laser, a laser power of 500mW, and a collection time of 30 seconds.
[0018] Wherein, step (4) also includes preprocessing the obtained spectral data, including baseline correction, Savitzky-Golay smoothing, and data normalization.
[0019] Among them, after the pre-processed spectral data is subjected to feature extraction and dimensionality reduction, the first 15 principal components accounting for more than 90% of the information are selected as the input features of the KNN algorithm.
[0020] Among them, CYP450, as an important liver enzyme, is mainly involved in the metabolism of various endogenous and exogenous compounds. Therefore, when studying the metabolism of drugs by CYPs in the present invention, other substrates metabolized by CYPs can be included.
[0021] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: the present invention can realize the rapid detection of cells metabolized by CYPs, and has the advantages of being label-free, simple to operate, and having a wide range of applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a flow chart of the identification and classification of CYPs metabolism of liver cancer cells using the PCA-KNN model in the present invention;
[0023] Figure 2 The Western Blot method in Example 1 of the present invention is used to detect the expression of CYP2E1 protein in HepG2 cells;
[0024] Figure 3 The cell survival rates of different concentrations of APAP in Example 2 of the present invention are shown;
[0025] Figure 4is a waterfall diagram of the Raman spectrum data set in Example 3 of the present invention;
[0026] Figure 5 is a cell average normalized spectrum diagram of CYP2E1-mediated acetaminophen metabolism in Example 3 of the present invention;
[0027] Figure 6 The principal component analysis results of the cell Raman spectrum in Example 3 of the present invention include (a) a score graph and (b) a component change graph;
[0028] Figure 7 is a loading diagram of the principal component analysis result of the cell Raman spectrum in Example 3 of the present invention;
[0029] Figure 8 The accuracy of the PCA-KNN model of the present invention;
[0030] Fig. 9 It is the confusion matrix of the PCA-KNN model of the present invention. DETAILED DESCRIPTION
[0031] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0032] Example 1: Construction of HepG2 cells stably expressing CYP2E1
[0033] The two lentiviral packaging auxiliary plasmids, psPAX2 and pMD2.G, were co-transfected into HEK293T cells with pLenti-CMV-GFP-Puro-CYP2E1 or pLenti-CMV-GFP / Puro plasmid (target plasmid purchased from Nanjing Jingmai Biotechnology), and the virus-containing culture supernatant was collected after 48 hours. HepG2 cells were infected with the virus-containing culture supernatant, and a stable cell line expressing the CYP2E1 gene (HepG2-CYP2E1) or a control cell line (HepG2-Control) was obtained by monoclonal screening.
[0034] Western blot was used to detect protein expression in HepG2 cells. Stable cell lines expressing CYP2E1 gene or control cell lines were collected, washed once with PBS, lysed on ice with Lysis buffer, and centrifuged at 4°C to obtain the supernatant. After adding 4× Loading Buffer and mixing, the samples were boiled for 5 minutes. After SDS-PAGE treatment, the samples were transferred to PVDF membranes, and the membranes were blocked with 5% skim milk powder in TBST for 1 hour at room temperature. After washing the membrane with TBST, CYP2E1-Specific Polyclonal antibody (dilution multiple of 1:2000, purchased from Wuhan Sanying Biotechnology Co., Ltd.) or β-actin monoclonal antibody (dilution multiple of 1:1000) diluted with 5% skim milk powder was added and incubated overnight at 4°C. Wash the membrane again, add horseradish enzyme-labeled goat anti-rabbit IgG (dilution factor 1:5000) or horseradish enzyme-labeled goat anti-mouse IgG (dilution factor 1:5000) corresponding to the source of the primary antibody, incubate at room temperature for 1 hour, then wash the membrane and add luminescent solution for development. Figure 2 As shown, successfully transfected HepG2 cells can significantly express CYP2E1 protein.
[0035] Example 2: Treatment of cells with acetaminophen
[0036] The drug used in the treatment of cells in the present embodiment is acetaminophen (APAP), also known as paracetamol, which is an important metabolic substrate of CYP2E1. Acetaminophen can produce toxic metabolites in the metabolic process of CYP2E1, and these products may cause damage to the liver if they accumulate too much. Therefore, studying the metabolic process of acetaminophen is of great significance for understanding its pharmacological action and potential toxicity.
[0037] The specific process of sample preparation is to continue to culture and subculture the HepG2 cells or control cells expressing CYP2E1 in Example 1. The culture medium used for cell culture is a DMEM complete culture medium prepared by adding 10% fetal bovine serum (FBS), 100U / mL penicillin and 0.1mg / mL streptomycin to DMEM high-glucose culture medium. When the cells are in the logarithmic growth phase, they are inoculated in a 6-well plate and cultured for 24h, and a drug-added group and a control group are set. After discarding the culture medium supernatant, fresh DMEM complete culture medium containing 20mM acetaminophen is added to the drug-added group, and fresh DMEM culture medium is added to the control group. After continuing to culture for 24h, the culture medium supernatant is discarded. After rinsing the cells with PBS twice, 1mL PBS solution is added and the cells are scraped off with a scraper, and the mixture is transferred to a 1.5mL centrifuge tube by blowing and mixing. Centrifuge for 5min at 4°C and 800×g, discard the PBS solution, and collect the cells.
[0038] Among them, the cell survival rate after the action of acetaminophen was determined by sulforhodamine B (SRB) staining experiment. The cells in the logarithmic growth phase were inoculated in a 96-well plate and cultured for 24 hours. DMEM complete culture medium containing APAP concentrations of 0, 10, 20, 30, and 60 mM was added and cultured for another 24 hours. Three replicate wells were set for each concentration. The culture medium was discarded, and 50% trichloroacetic acid solution was added to each well and fixed at 4°C for 1 hour. The fixative was discarded, the plate was washed 3 times with distilled water, and after drying, 0.4% SRB stain was added to each well and stained at room temperature for 20 minutes. The staining solution was discarded, the plate was washed 5 times with 1% acetic acid, and after drying, 10 mmol / L Tris solution was added to dissolve the dye. The absorbance was measured at a wavelength of 530 nm and the cell survival rate was calculated. The results are as follows Figure 3 As shown in the figure, when the APAP concentration is higher than 20mM, cytotoxicity will occur, and the cell survival rate will decrease as the concentration increases. In subsequent experimental verification, it was found that APAP showed better experimental results at 20mM than at 30mM, so the APAP dosing concentration was determined to be 20mM.
[0039] Example 3: Raman spectroscopy analysis of HepG2 cells
[0040] In this embodiment, the spectrum acquisition method is to drop the cells onto a single crystal silicon wafer and dry them at low temperature, and then collect the spectrum of the sample points on the silicon wafer. Multiple sample points are prepared for each sample, and different positions are selected at each point for multiple measurements. 50 spectra are collected for different types of samples. The spectrum acquisition conditions are: using a 785nm laser, a laser power of 500mW, a collection time of 30 seconds, and a scanning spectrum range of 200-2500cm -1 .
[0041] It should be noted that a blank silicon wafer should be used at 520 cm before spectrum acquisition. -1 Calibrate at.
[0042] In this embodiment, the spectral data is collected using an ATR3110 Raman spectrometer produced by Xiamen Aup Tiancheng Optoelectronics Co., Ltd.
[0043] Since cosmic rays (random, sharp spectral lines) are sometimes generated during the Raman spectrum acquisition process, and the fluorescence generated by biological samples will cause the Raman spectrum baseline to shift upward, pre-processing must be performed after spectrum acquisition. The Raman spectrum fingerprint spectrum of 600-1800cm is selected. -1 The area was baseline corrected, Savitzky-Golay smoothed, and data normalized. Figure 4The 3D waterfall diagram of the Raman data preprocessing of the four groups of cells (HepG2 cells expressing CYP2E1 with drug (A-CYP2E1), control cells with drug (A-Control), HepG2 cells expressing CYP2E1 with DMEM complete medium (D-CYP2E1), and control cells with DMEM complete medium (D-Control)) is shown. However, due to too much spectral data, it is impossible to obtain the characteristic data of the spectrum more intuitively. Therefore, it is necessary to calculate the average value of the spectrum of each group of cells to obtain representative spectra for comparison and analysis of characteristic peaks. Figure 5 As shown in the figure, the average normalized spectra of HepG2 cells expressing CYP2E1 with drug added (A-CYP2E1), control cells with drug added (A-Control), HepG2 cells expressing CYP2E1 with DMEM complete medium added (D-CYP2E1), and control cells with DMEM complete medium added (D-Control) are stacked, and some obvious characteristic peaks can be observed.
[0044] In addition, the Raman peak positions and tentative vibration mode assignments observed in the embodiments of the present invention are shown in Table 1. The characteristic peaks are mainly located at 618, 645, 716, 745, 780, 823, 847, 936, 999, 1029, 1089, 1123, 1169, 1319, 1340, 1440, 1539, 1615, 1705 cm -1 ,Through the division in Table 1, it can be seen that the four groups of ,cell samples collected are mainly caused by the molecular vibration of ,saccharides, proteins, nucleic acids and lipids.
[0045] Table 1 Raman peak positions and corresponding vibration mode assignments
[0046]
[0047]
[0048] In this embodiment, Origin 2024b is used to preprocess the Raman spectrum, and the PCA class and KNN classifier (KNeighborsClassifier) in the Scikit-learn library of python 3.10 are used to establish a PCA-KNN algorithm model. Among them, the steps include: taking Raman spectral data as input variables (features), the corresponding groups as output variables (labels), and using PCA class code to perform principal component analysis. The results are input into the KNN classification model created by the KNeighborsClassifier class code, cross-validated using the StratifiedKFold code, fitting the KNN model through the training set, predicting on the test set, calculating the accuracy and outputting the evaluation results. Principal Component Analysis (PCA) is an unsupervised learning method for data dimensionality reduction and feature extraction. It converts the original data into a set of new unrelated variables through linear transformation and retains the variability of the original data. K-Nearest Neighbor (KNN) is one of the most basic algorithms in machine learning and is an algorithm commonly used for classification in supervised learning. The principle of classification is to input a new sample and determine the category of the new sample based on the categories of the K sample points closest to it. The specific steps are to calculate the distance between the new data point to be classified and all the training data points. According to the calculated distance, select the K points closest to it. Count the frequency of occurrence of the categories of the first K points, and finally classify the point to be classified into the category with the highest frequency.
[0049] In this embodiment, principal component analysis is performed on the preprocessed Raman data to extract eigenvalues and reduce dimensionality, thereby avoiding the influence of irrelevant characteristic peaks on model classification. Figure 6 The PCA results shown Figure 6 The score graph of (a) projects the data from high-dimensional space to two-dimensional space. The elliptical area in the figure represents the 95% confidence interval. The first principal component (PC1) of the coordinate axis explains the largest variance in the data, which is 79.1%; the variance of the second principal component (PC2) accounts for 5.3%. The figure intuitively shows the distribution and clustering of data points of cells in the drug-treated group and the control group. Although there is a slight overlap between the CYP2E1-expressing cells and control cells in the drug-treated group and the CYP2E1-expressing cells and control cells in the control group, the overall separation results are good. Among them, the data points of the overexpressing cells in the drug-treated group are relatively concentrated and the clustering is good, indicating that the characteristics are obvious, which is consistent with our estimated results. Figure 6 (b) shows the cumulative contribution rate and variance of the first 15 principal components, indicating that the 15 principal components can explain more than 90% of the information in the Raman spectrum.
[0050] Figure 7The loading diagrams of PC1 and PC2 are shown. The main characteristic peak of PC1 is 823cm -1 (tyrosine (protein)), 936cm -1 (single bond stretching vibration of amino acid), 999cm -1 (phenylalanine ring breathing), 1089 cm -1 (single bond stretching vibration), 1123 cm -1 (lipids and proteins), 1340 cm -1 (nucleic acid, collagen), 1440cm -1 (lipid), indicating that PC1 contains more protein information. -1 and 1440cm -1 The characteristics of these two peaks are quite obvious, 999cm -1 The peak at 1440cm is derived from the vibrational breathing of the phenylalanine ring in the protein. The peak intensity of the CYP2E1 cells in the control group is higher than that in the drug-treated group, indicating that the protein content in the cells decreases after the drug is metabolized by CYP2E1. It may be that the toxic products produced by metabolism inhibit protein synthesis. -1 The peak at reflects the lipid changes in the cells. The peak of the control cells in the drug-treated group is higher than that of the control cells in the control group because HepG2 cells grow well and consume energy quickly. Lipids provide the energy required for cell growth and are difficult to accumulate in cells. The peak intensity of the drug after CYP2E1 metabolism is lower, which may be due to the activation of CYP2E1 promoting the oxidative metabolism of fatty acids in liver cancer cells and reducing lipid accumulation. -1 (CC twist (protein)) and 1615cm -1 (Tyrosine, tryptophan, C=C (protein)) are mainly derived from protein.
[0051] According to the analysis results of PCA, the first 15 principal components with a contribution of more than 90% were selected as the input features of the KNN classification model. The GridSearchCV algorithm combined with stratified five-fold cross validation was used to optimize the PCs and K parameters of the PCA-KNN model. Figure 8 As shown, the number of principal components was selected from 2 to 15, and the K parameter ranged from 2 to 8. When the PCs was 13 and K was 3, the highest classification accuracy obtained by cross-validation was 89%.
[0052] The model is trained using the optimal parameter combination, and the test set is predicted. The accuracy (Accuracy), precision (Precision), recall (Recall), and F1 score are calculated based on the real labels and prediction results. The model with the highest accuracy is the best model. Accuracy refers to the proportion of samples predicted correctly by the model to the total samples; precision refers to the proportion of samples predicted to be positive to the proportion of samples that are actually positive; recall refers to the proportion of samples that are actually positive to the proportion of samples that are correctly predicted to be positive, also known as the true rate; and F1 score refers to the harmonic mean of accuracy and recall, providing a balance between the two indicators. The accuracy, precision, recall and F1 score of the PCA-KNN model established in this embodiment are all 89%. The confusion matrix is a summary of the prediction results of the classification problem. Each row of the confusion matrix represents the true category of the data, and the total number of data in each row represents the number of data instances in that category; each column represents the predicted category of the data, and the total number of data in each column represents the number of data instances predicted to be in that category. The results are as follows. Fig. 9 As shown in the confusion matrix, the accuracy of the PCA-KNN model confusion matrix for CYP2E1 involved in acetaminophen metabolism is 89%, among which the CYP2E1 cells in the drug-added group are correctly classified, 3 data of the empty control cells in the drug-added group are incorrectly classified as control group cells, 6 data of the CYP2E1 cells in the control group are misclassified into the other three categories, and 9 data of the empty control cells in the control group are misclassified into the drug-added group, and 4 are misclassified into the CYP2E1 cells in the control group. Most of the data can be correctly classified, indicating that the PCA-KNN model can effectively identify CYP2E1-mediated acetaminophen metabolism.
Claims
1. A method for analyzing cytochrome P450 metabolism based on Raman spectroscopy and machine learning model, characterized in that: The following steps are involved: (1) Construction of cells that stably express CYPs; (2) Collecting cells after drug treatment; (3) collecting Raman spectroscopy data of cells; (4) Feature extraction and dimensionality reduction of the collected Raman spectral data through principal component analysis; (5) Use the K nearest neighbor algorithm to classify the Raman spectral features after dimensionality reduction.
2. The cytochrome P450 metabolic analysis method according to claim 1, characterized in that: The cell construction method in step (1) comprises: co-transfecting the packaging plasmid and the target plasmid into HEK293T cells, collecting the virus-containing culture medium supernatant and then infecting the cells, and obtaining a stable cell line expressing the above CYPs by monoclonal screening.
3. The cytochrome P450 metabolic analysis method according to claim 2, characterized in that: The target plasmid includes pLenti-CYP2E1, pLenti-CYP3A4, pLenti-CYP1A2 or pLenti-CYP2C9.
4. The cytochrome P450 metabolic analysis method according to claim 2, characterized in that: Infected cells include cell models used to study liver-related drug metabolism; including HepG2 cells, HLE cells, and Hep3B cells.
5. The cytochrome P450 metabolic analysis method according to claim 2, characterized in that: The packaging plasmids include psPAX2 and pMD2.G.
6. The cytochrome P450 metabolism analysis method according to claim 1, characterized in that: The drug in step (2) includes a substrate metabolized by CYPs.
7. The cytochrome P450 metabolism analysis method according to claim 1, characterized in that: The step (3) of collecting Raman spectral data of cells includes: dropping the cells onto a silicon wafer and drying them at low temperature, and then collecting spectra of sample points on the silicon wafer; wherein the spectral collection conditions are: using a 785nm laser, a laser power of 500mW, and a collection time of 30 seconds.
8. The cytochrome P450 metabolism analysis method according to claim 1, characterized in that: Step (4) also includes preprocessing the obtained spectral data, including baseline correction, Savitzky-Golay smoothing, and data normalization.
9. The cytochrome P450 metabolism analysis method according to claim 8, characterized in that: After feature extraction and dimensionality reduction of the preprocessed spectral data, the first 15 principal components accounting for more than 90% of the information were selected as the input features of the KNN algorithm.