A method for predicting drug sensitivity based on blood NMR editing spectrum
By combining blood NMR spectral editing detection technology with machine learning models, the inaccuracy of existing drug sensitivity prediction methods has been addressed, enabling efficient and non-destructive detection of individualized drug sensitivity and improving the accuracy and reliability of predictions.
Patent Information
- Application Number
- CN202510689603.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Existing methods for predicting drug sensitivity are not accurate enough in the treatment of chronic non-communicable diseases such as cancer. In particular, prediction methods based on upstream biological indicators cannot meet the needs of personalized precision treatment, and conventional methods may lead to drug resistance or excessive progression.
A blood NMR editing spectrum-based approach was adopted. Through one-dimensional proton NMR editing spectrum detection technology, NMR fingerprint data of small molecule metabolites, glycoproteins, and lipoproteins in plasma or serum before and after drug treatment were obtained. Combined with physiological and pathological indicators, unsupervised classifier models such as random forest and support vector machine were used to screen out biomarkers with high predictive ability to predict drug sensitivity.
It improves the accuracy of drug sensitivity prediction, reduces the uncertainty caused by unknown escape mechanisms, requires small sample size and is non-destructive to the body, and is fast and can be standardized.
Smart Images

Figure CN120613144B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical testing technology and relates to a method for predicting drug sensitivity based on blood NMR editing spectra. Background Technology
[0002] Chronic non-communicable diseases such as cancer and cardiovascular diseases pose a serious threat to human health. Due to their complex pathogenesis and diverse influencing factors, many drugs used to treat these diseases exhibit varying degrees of drug resistance or intolerance, making treatment challenging. Developing accurate and effective drug sensitivity prediction methods to anticipate patient sensitivity to these drugs in advance will help promote personalized and precise treatment for these diseases.
[0003] Most current methods for predicting drug sensitivity are developed for drug screening and are usually based on high-throughput drug sensitivity data and cell line genomics databases related to diseases, such as the Encyclopedia of Cancer Cell Lines (CCLE) and the Genetic Diagnostic and Statistical Study of Cancer Drug Sensitivity (GDSC). At the cellular level, machine learning methods are used to non-targetedly predict drugs that may respond to the disease.
[0004] Predicting the sensitivity to a specific drug at the fluid or tissue level typically involves selecting upstream biological indicators based on the drug's mechanism of action. Taking cancer as an example, although immunotherapy methods based on ICIs (immunoassays) such as PD-1 / PD-L1 inhibitors have been developed in recent years, significantly improving overall survival and reducing side effects in patients with various types of cancer, the efficacy of ICIs exhibits significant population variability, with only a small percentage of patients benefiting. Therefore, it is necessary to detect predictive biomarkers before medication to guide treatment. Biomarker detection is currently the main approach to assessing potential beneficiaries of ICIs and guiding rational drug use. In recent years, researchers have discovered numerous biomarkers for this therapy at the gene level and the immune microenvironment level, such as PD-L1 expression, tumor mutational burden (TMB) levels associated with cancer, cancer microenvironment-related factors, and gut microbiota. Among these, medication guidelines based on PD-L1 expression have already been applied in clinical treatment. However, even with biomarker-based efficacy predictions, the objective clinical response rate to ICI monotherapy remains below 30% for most cancer types. Some cancer patients even experience accelerated disease progression and develop hyper-progressive disease (HPD) after receiving prescribed medication. This indicates that predictive methods relying on upstream biomarkers are insufficient to meet the demands of precision medicine.
[0005] While upstream biological indicators such as gene expression and gene mutations can provide clues about potential future events, they also inevitably face numerous escape mechanisms. However, with the widespread application of metabolic phenotyping in medicine, it has been found that most chronic non-communicable diseases are closely related to metabolic disorders. For example, cardiovascular and cerebrovascular diseases are often accompanied by abnormal lipid metabolism, and cancer is closely related to energy metabolism disorders. Unlike upstream biological indicators, metabolic phenotyping information represents definite past events. Therefore, developing drug sensitivity prediction methods for chronic diseases by focusing on metabolic phenotyping analysis can avoid the uncertainties introduced by upstream biological indicator predictions and improve prediction accuracy.
[0006] In recent years, metabolic phenotyping analysis has been widely applied in medical fields such as drug metabolism research, disease diagnosis and classification, and biomarker discovery related to chronic non-communicable diseases. Besides changes in common small molecule metabolites, the metabolism of lipids and some biomolecules (such as lipoproteins and glycoproteins) is also closely related to chronic non-communicable diseases. Commonly used techniques for metabolic phenotyping analysis include mass spectrometry and nuclear magnetic resonance (NMR). NMR is particularly effective in detecting disease characteristics in bodily fluids and has been proven to be an effective diagnostic option for COVID-19 infection. For bodily fluid samples, NMR allows for non-destructive detection, meaning it can simultaneously detect small molecule metabolites and some biomolecules without separating the sample. This also means that intermolecular interactions (including between drugs and small molecules, and between biomolecules) can be preserved, thus providing additional evaluative information. Metabolic phenotypic analysis based on nuclear magnetic resonance (NMR) technology has revealed the relationship between the diversity of metabolites in human blood plasma and the interaction between plasma biomolecules and drugs (Du Y, Lan W, Ji Z, Zhang X, Jiang B, Zhou X, et al. NMR spectroscopic approach reveals metabolic diversity of human blood plasma associated with protein-drug interaction. Analytical Chemistry 85(2013)8601-8.). Since the diversity of metabolites in blood is caused by individual metabolic differences, NMR-based blood metabolic phenotypic analysis can be used to assess the relationship between individual metabolic differences and drug efficacy. Furthermore, by using machine learning methods to model blood metabolic data before and after drug treatment, the predictive accuracy of using only a single biomarker can be compensated for. Summary of the Invention
[0007] Based on the above background, this application aims to separate the upstream biological mechanisms and focus on the differences in blood metabolism of drugs in patients with chronic non-communicable diseases. It proposes a prediction method that can avoid the uncertainty caused by unknown escape mechanisms, improve prediction accuracy, and has minimal damage to the body during the sampling process, and is non-destructive and standardizable in the detection and analysis process.
[0008] To address the problems existing in the prior art, this invention first establishes two one-dimensional proton NMR editing spectroscopy detection methods (T2W and DIRE) for plasma or serum samples. These methods preserve intermolecular interactions while acquiring as much information as possible about metabolites, obtaining NMR fingerprint data of small molecule metabolites, glycoproteins, and lipoproteins in plasma or serum before and after drug treatment. Subsequently, individual differences in drug treatment efficacy are identified through physiological and pathological indicators, and the cohort is divided into a high-sensitivity group and a low-sensitivity group. Unsupervised classifiers, including random forest and support vector machine classification models, are trained based on group labels and pre-treatment blood sample fingerprint data. The optimal classification model is evaluated and selected, and high-predictive-capability biomarkers for drug sensitivity prediction are screened through ROC analysis of features. The technical route of this method is shown in the appendix to the specification. Figure 1 As shown.
[0009] Specifically, the present invention achieves the above objectives through the following methods:
[0010] A method for predicting drug sensitivity based on blood NMR editing spectra, the specific steps of which are as follows:
[0011] 1. Queue Design
[0012] A cohort of volunteers suffering from a specific chronic non-communicable disease was recruited in strict accordance with ethical requirements for the study. The research cohort was divided into two groups:
[0013] (1-1) Treatment group: Patients who received a particular drug treatment;
[0014] (1-2) Blank control group: Patients who did not receive treatment. Hierarchical cluster analysis was performed on the blood biochemistry and prognostic indicators of the treatment group after treatment and the blood biochemistry and prognostic indicators of the blank control group in the later stage of progression. The patients in the treatment group were divided into high sensitivity group and low sensitivity group.
[0015] (1-3) Blood sample retention
[0016] (1-3-1) Retain all patients’ plasma or serum samples at baseline as samples before administration;
[0017] (1-3-2) Retain plasma or serum samples from patients in the treatment group during the later stage of treatment (in this invention, the later stage of treatment refers to the period of relative stabilization of the condition) and from patients in the blank control group during the later stage of natural progression (in this invention, the later stage of progression is determined with reference to the time of the later stage of treatment in the treatment group);
[0018] 2. Acquisition and processing of NMR fingerprint data
[0019] (2-1) Obtain one-dimensional samples from each group of plasma or serum samples in step 1. 1 H NMR spectral data, the one-dimensional 1 The H NMR spectra are relaxation-weighted T2W-NMR and diffusion relaxation editing spectra DIRE;
[0020] (2-2) Obtain the two-dimensional nuclear magnetic resonance spectrum of any plasma or serum sample from step 1 to assist in the identification of T2W-NMR spectrum signals;
[0021] (2-3) Data processing:
[0022] The one-dimensional result obtained in (2-1) 1 Phase and baseline corrections were performed on the 1H NMR spectrum. The chemical shift of the T2W-NMR spectrum was calibrated using the glucose α-isomer proton signal at δ 5.23 ppm. The DIRE spectrum selected the peak with the best resolution in the spectrum as the calibration signal (preferably calibrated by SPC at the -N(CH3)3 proton signal at δ 3.20 ppm). The linewidth factor was set to 0.3 Hz, and then Fourier transform was used to generate the frequency domain spectrum.
[0023] Regarding the above 1 The chemical shift drift of the H NMR spectrum was corrected, and then segmented integration was performed with an integration width ranging from 6 to 24 Hz. The segmented integration was finally normalized based on the sum of the total integrations for each sample, and the datasets of the T2W-NMR spectrum and DIRE spectrum were obtained respectively.
[0024] (2-4) T2W-NMR spectral signal identification
[0025] Signal identification of T2W-NMR spectra of blood samples was performed based on (2-2) two-dimensional nuclear magnetic resonance spectroscopy and the human metabolome database HMDB.
[0026] (2-5) Identification of DIRE Spectrum Signals
[0027] Refer to the following literature for signal identification of blood sample DIRE spectra: Liu M, Nicholson JK, Lindon JC. High-resolution diffusion and relaxation edited one-and two-dimensional 1HNMR spectroscopy of biological fluids.Analytical chemistry 68(1996)3370-6;
[0028] The identification of two fingerprint spectrum signals is achieved through the above (2-4) and (2-5). Among them, the T2W-NMR spectrum is the fingerprint spectrum of small molecule metabolites in the blood sample, and the DIRE spectrum is the fingerprint spectrum of glycoproteins, lipoproteins and lipid macromolecule metabolites containing small molecule branched chains in the blood sample.
[0029] 3. Model Building and Evaluation
[0030] Using random forest and support vector machine algorithms, T2W-NMR and DIRE fingerprint datasets of plasma or serum samples from high-sensitivity and low-sensitivity groups before treatment were modeled, resulting in four prediction models. The random forest model consisted of 500 decision trees, and out-of-bag (OOB) error was used to evaluate its generalization ability. Leave-one-out cross-validation (LOOCV) was used to evaluate the generalization ability of the support vector machine. Based on the evaluation results, the model with the lower classification error was selected for drug sensitivity prediction.
[0031] The effectiveness of evaluating the performance of random forests by using out-of-bag error has also been proven (Mitchell MW. Bias of the Random Forest Out-of-Bag (OOB) Error for Certain Input Parameters[J]. Open Journal of Statistics,2011,01(3):205-211.).
[0032] For T2W-NMR spectra, the feature dataset used for modeling is a matrix composed of the chemical shifts of the characteristic peaks of the small molecule metabolites identified in step (2-4) and their segmented integral values; the feature dataset for DIRE spectra is a matrix composed of the chemical shifts of the glycoproteins, lipoproteins and lipid macromolecule metabolites containing small molecule branches identified in step (2-5) and their segmented integral values.
[0033] 4. Screening of biomarkers
[0034] Using the BiomarkerAnalysis module of the MetaboAnalyst analysis platform, based on the model selected in (4-1), important features were screened and ROC curve analysis was performed on all features. The ROC curves in this module are generated by Monte Carlo cross-validation (MCCV) with balanced subsampling. Through this process, important features with biomarker functions were identified. Then, the classification and prediction ability of the feature variables was evaluated based on the area under the ROC curves (AUC) of different combinations of feature variables. One or more biomarkers were selected for drug sensitivity prediction.
[0035] 5. Drug sensitivity prediction
[0036] ① Prepare the data to be predicted: Following step 2, collect and process the blood NMR data of the patients to be treated, and convert the NMR data into the input format expected by the model;
[0037] ② Input the obtained NMR data into the model using the BiomarkerAnalysis module of the MetaboAnalyst platform, and use the biomarkers or combinations thereof selected in step 4 to perform high-sensitivity and low-sensitivity group classification prediction.
[0038] Furthermore, in step (2), a transverse relaxation time-weighted T2W-NMR spectrum is obtained using a water-presaturated 1D carr-purcell-meiboomm-gill pulse sequence, with a total spin echo time of 70 ms and a cycle delay of 2.0 s; and / or
[0039] When acquiring diffusion and relaxation editing spectra (DIRE), three smooth square gradient fields were used to achieve diffusion editing with intensities of -17.13%, 13.13%, and 80%, respectively, and gradient lengths and diffusion delays of 1.53 ms and 120 ms, respectively; relaxation editing was achieved using a total spin echo time of 33 ms.
[0040] Furthermore, the method selects drugs and disease types with significant individual differences in efficacy in clinical applications. The drugs include, but are not limited to, PD-1 monoclonal antibodies: Pembrolizumab, Nivolumab, Camrelizumab; or PD-L1 monoclonal antibodies: Atezolizumab, Avelumab, Adebelimumab; and the diseases include, but are not limited to, non-small cell lung cancer, gastrointestinal cancer, and breast cancer.
[0041] The prediction model and method proposed in this invention are essentially designed to solve the binary classification problem (high-sensitivity group and low-sensitivity group) of the data to be predicted, and machine learning is one of the powerful tools for solving classification problems. To achieve the classification of the high-sensitivity group and the low-sensitivity group, a classifier needs to be trained using a machine learning algorithm. Finally, this classifier predicts drug sensitivity based on the NMR data characteristics of the blood samples of the patients to be treated.
[0042] This invention employs two classifiers: supervised Random Forest (RF) and Support Vector Machine (SVM) algorithms. Random Forest is an ensemble learning algorithm that primarily uses bagging and random feature selection to construct multiple decision trees and combine their predictions to improve model performance. Support Vector Machine is widely used in classification and regression tasks; its core principle is to find a hyperplane that maximizes the margin between two classes, thereby achieving good classification results.
[0043] Compared with the prior art, this application has the following advantages and beneficial effects:
[0044] 1) This invention establishes two one-dimensional proton spectrum detection methods: relaxation weighted spectroscopy (T2W-NMR) and diffusion relaxation editing spectroscopy (DIRE). By using T2W-NMR spectroscopy and setting an appropriate echo time, large molecules with relatively small T2 values can decay rapidly, while small molecules with larger T2 values can retain their signals due to slower signal decay. DIRE spectroscopy can enhance the signals of glycoproteins and lipoproteins, which have slow diffusion but long relaxation times.
[0045] 2) Non-targeted prediction and non-destructive blood sample testing preserve the interaction between drugs and biomolecules, thus improving prediction accuracy;
[0046] 3) The method proposed in this invention is based on the detection of trace blood samples. The sample volume is greater than 30 μL, so the sampling process causes very little damage to the body, the detection process is non-destructive to the sample, and the sample can be used for other tests after the detection is completed. Moreover, the analysis process is fast and can be standardized. Attached Figure Description
[0047] Figure 1 The flowchart of the ICIs inhibitor susceptibility prediction method proposed in this application.
[0048] Figure 2The animal experiment timeline and physiological and pathological results in the examples are as follows: A shows the experimental procedure for RMP1-14 treatment. B and C show the changes in body weight and tumor volume over time, respectively. D shows the final tumor images of 8 randomly selected mice in the treatment group and 6 mice in the control group. E compares the final tumor weight between the treatment group and the control group. F and G show the histological examination of the final tumor sections of two representative mice after treatment with PBS and RMP-14, respectively. AVT: Mean tumor volume. ATV Control : Average tumor volume in the control group.
[0049] Figure 3 The pulse sequence of the DIRE spectrum in the embodiment, wherein T e The eddy current delay time is 5ms, t s Editing time for T2 relaxation (2t) s The total time is 33ms. Diffusion editing is achieved by using a gradient length of 1.5ms δ and a diffusion delay time of 120ms Δ.
[0050] Figure 4 In the examples, the plasma T2W-NMR (A) and DIRE (B) spectra of mice in the treatment group before drug administration were used to identify the signals. Figure 4 The numbers in A represent: 1. Cholesterol; 2. Lipids; 3. Isoleucine; 4. Leucine; 5. Valine; 6. 3-Hydroxybutyric acid; 7. Lactic acid; 8. Alanine; 9. Acetic acid; 10. N-acetyl glycoprotein (NAG); 11. O-acetyl glycoprotein (OAG); 12. Methionine; 13. Acetone; 14. Pyruvic acid; 15. Succinic acid; 16. Glutamine; 17. Citric acid; 18. Lysine; 19. Creatine; 20. Phosphorylcholine; 21. Glycerylphosphorylcholine; 22. Inositol; 23. Glycerol; 24. Threonine; 25. β-glucose; 26. α-glucose; 27. Urea; 28. Tyrosine; 29. Histidine; 30. Phenylalanine; 31. Tryptophan.
[0051] Figure 5 The examples show hierarchical clustering heatmaps drawn based on post-drug administration plasma biochemical parameters (see figure for details) and tumor growth rate. In the figure, "C" represents the control group; and "T" represents the treatment group.
[0052] Figure 6 The out-of-bag error rate (OOB) of the random forest model trained on the T2W(A) and DIRE(B) spectrum datasets of pre-drug plasma in the examples.
[0053] Figure 7 The error rate of the support vector machine model trained on the T2W(A) and DIRE(B) spectrum datasets of plasma before drug administration in the examples is shown.
[0054] Figure 8 In the examples, the random forest classification features based on the DIRE spectrum of plasma samples before drug administration were screened as follows: A-1: the top 15 features with the highest frequency were selected; A-2: ROC analysis of 3-53 features in the DIRE dataset; A-3: the predictive ability of the top 5 high-frequency selected features. Detailed Implementation
[0055] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0056] Example
[0057] Using a tumor-bearing mouse model as the subject and the murine PD-1 inhibitor RMP1-14 as the treatment drug, the predictive method will be applied preclinically.
[0058] In this invention, for the preservation of plasma samples, anticoagulant tubes that have passed NMR detection should be selected, and the samples should be flash-frozen in liquid nitrogen within 30 minutes and then stored at -80°C.
[0059] Preparation of plasma samples:
[0060] For blood samples with a plasma volume of less than 100 μL, mix 30 μL of plasma with an equal volume of phosphate buffered saline (45 mM, pH = 7.4 ± 0.1). After centrifugation at 11000 g and 4 °C for 10 minutes, transfer the supernatant (50 μL) to a 1.7 mm NMR tube for analysis.
[0061] For blood samples with a plasma volume greater than 200 μL, mix 200 μL of plasma with 400 μL of phosphate-buffered saline (45 mM, pH = 7.4 ± 0.1), centrifuge at 11000 g and 4 °C for 10 minutes, and transfer the supernatant (500 μL) to a 5 mm NMR tube for analysis.
[0062] 1. Animal experiments and blood sample collection
[0063] 1.1 Materials and Methods
[0064] MC38 colorectal cancer cells (purchased from The Global Bioresource Center, cell line number ATCC#CRL-2638) were cultured according to established standard procedures. Female C57BL / 6J mice, 6-8 weeks old, were obtained from Cyagen Biosciences Inc. in Suzhou, China, and were housed under a 12-hour light / dark cycle with free access to food and water. 1*10 cells were injected into the posterior abdomen. 6 After 100 MC38 cells were collected, the tumor size was measured every 2-3 days using digital calipers. The tumor volume was calculated using the formula V = (W / W). 2*L) / 2, where W is the width of the tumor and L is the length of the tumor. The formula for calculating tumor growth is: Tumor growth (%) = Tumor volume (day n) / Tumor volume (day 0) × 100%.
[0065] The average volume reaches 90±20mm. 3 On day 0, 22 mice were randomly divided into a treatment group (n=16) and a control group (n=6). The treatment group was given 4.0 mg / kg (mpk) of anti-mouse PD-1 antibody RMP1-14 (BioXcell, BE0146), while the control group was given phosphate-buffered saline (45 mM, pH=7.4±0.1) carrier injection. Both groups were administered intraperitoneally on days 0, 3, 7, and 10.
[0066] Figure 2 As shown in Figure A, the animal experiment in this embodiment lasted for 31 days, with a total of 4 administrations. Plasma samples were collected from mice before and after treatment. Plasma samples (30–40 μL) were collected at two time points: Plasma sample ① – the day before the first treatment (day 2, with an average tumor volume of approximately 55 mm). 3 ) and plasma samples ② - before mice were sacrificed after the experiment (day 21, average tumor volume was approximately 1600 mm). 3 All plasma samples were collected in 1.5 mL EP tubes using K2EDTA solution (25 μL, 15 mg / mL) as an anticoagulant. Additionally, a portion of blood from euthanized mice was retained as serum for biochemical analysis.
[0067] Tumors from randomly selected control and experimental mice were fixed in formalin, embedded in paraffin, sectioned, and stained with hematoxylin and eosin (H&E), followed by microscopic analysis.
[0068] 1.2 Results Analysis
[0069] The results showed that RMP1-14 treatment significantly inhibited tumor growth in mice, but did not cause a significant difference in body weight. Figure 2 B). Compared with the control group, the tumor volume and weight in the treatment group were significantly reduced (P<0.01). Figure 2 (D, E) The relative standard deviation (RSD) of tumor weight in the treatment group was significantly higher than that in the control group. Histological examination of tumor sections from two representative mice showed that the tumor lesions almost completely regressed after RMP1-14 treatment. Figure 2 (F, G). These results indicate that RMP1-14 treatment was generally effective in this trial. Importantly, the tumor weight change in the treatment group was more linear (higher RSD), suggesting significant inter-individual variability in drug efficacy.
[0070] 2. Plasma samples 1 H NMR spectral acquisition
[0071] Prior to NMR analysis, plasma (30 μL) was mixed with an equal volume of phosphate-buffered saline (45 mM, pH 7.4 ± 0.1). After centrifugation at 11000 g for 10 minutes at 4 °C, the supernatant (50 μL) was transferred to a 1.7 mm NMR tube. NMR data were acquired using a Bruker AVIII 600 MHz spectrometer equipped with a cryogenic probe.
[0072] All NMR data were acquired at 25°C using an automated sample changer system (BACS, Bruker Biospin, Germany). Two types of samples were collected from each plasma sample. 1 H NMR fingerprint spectrum:
[0073] (1) Transverse relaxation time-weighted (T2W-NMR) spectra were obtained using a water-presaturated 1D carr-purcell-meiboomm-gill (CPMG) pulse sequence with a total spin echo time of 70 ms and a cycle delay of 2.0 s.
[0074] (2) When acquiring diffusion and relaxation editing spectra (DIRE) (Reference: Liu M, Nicholson JK, Lindon JC. High-resolution diffusion and relaxation edited one-and-two-dimensional) 1 HNMR spectroscopy of biological fluids. Analytical Chemistry 68(1996)3370-6), diffusion editing was achieved using three smooth square gradient fields with intensities of -17.13%, 13.13%, and 80%, and gradient lengths and diffusion delays of 1.53 ms and 120 ms, respectively; relaxation editing was achieved using a total spin echo time of 33 ms.
[0075] Two plasma samples were randomly selected and their two-dimensional nuclear magnetic resonance spectra were collected. 1 H- 1 H COSY and TOCSY, 1 H- 13 CHSQC is used for T2W-NMR spectrum signal identification.
[0076] 3. Data processing and identification of two NMR fingerprint signals
[0077] 1.1 Data Processing
[0078] The obtained NMR spectra were processed using TOPSPIN software from Bruker Biospin: phase and baseline corrections were performed on all spectra. The chemical shifts of the T2W-NMR spectra were calibrated using the glucose α-isomer proton signal at δ 5.23 ppm (a common practice in existing literature), while the DIRE spectra were calibrated using the -N(CH3)3 protons of the SPC signal at δ 3.20 ppm (selecting the peak with the best separation in the spectrum as the calibration signal). The linewidth factor was set to 0.3 Hz, and then Fourier transform was used to generate the frequency domain spectrum.
[0079] Then, NMRSpec software was used to correct and segmentally integrate the signal coverage range of plasma in the NMR 1H spectrum (in this example, the chemical shift range from δ0.50ppm to δ9.00ppm) as shown in the two fingerprint spectra, with an integration width of 6Hz. To eliminate baseline distortion caused by suppressing the water peak, the chemical shift regions of δ4.50ppm to δ5.11ppm in the T2W spectrum and δ4.50ppm to δ5.01ppm in the DIRE spectrum were excluded during the segmented integration process to ensure the accuracy of the spectral analysis results. The segmented integration was finally normalized based on the sum of the total integrations for each sample to obtain the above two fingerprints. 1 Dataset of H NMR fingerprint spectrum.
[0080] 1.2 Signal Identification
[0081] Based on two-dimensional nuclear magnetic resonance spectroscopy and related literature (Cloarec O, Dumas ME, Craig A, Barton RH, Trygg J, Hudson J, et al. Statistical Total Correlation Spectroscopy: An Exploratory Approach for Latent Biomarker Identification from Metabolic...), 1HNMR Data Sets. Anal Chem 77(2005)1282-89; Nicholson JK, Foxall PJ, Spraul M, Farrant RD, Lindon JC. 750MHz 1H and 1H-13C NMR spectroscopy of human blood plasma. Anal Chem 67(1995)793-811.) and the Human Metabolome Database (HMDB) (Wishart DS, Feunang YD, Marcu A, Guo AC, Liang K, Vázquez-Fresno R, et al. HMDB 4.0: the human metabolome database for 2018. Nucleic acids research 46(2018)D608-d17, https: / / doi.org / 10.1093 / nar / gkx1089), Thirty-one metabolites were identified in the T2W-NMR spectra of plasma from mice in the treatment group before drug administration. Figure 4 (A and Table 1). These metabolites include lipids, amino acids, carbohydrates (glucose, inositol), tricarboxylic acid (TCA) cycle intermediates (citrate, succinate, fumarate), glycoproteins (NAG, OAG), organic acids (formate, acetate), ketones (3-hydroxybutyrate, acetone), choline-containing compounds (choline, phoscholine), as well as lactic acid, glycerol, creatine, and urea.
[0082] The identification of DIRE spectral signals in the plasma of mice in the treatment group before drug administration was mainly based on the reference (Liu M, Nicholson JK, Lindon JC. High-resolution diffusion and relaxation edited one-and two-dimensional). 1 ¹H NMR spectroscopy of biological fluids. Analytical Chemistry 68(1996)3370-6). DIRE(B) spectral signal identification see [link to journal article]. Figure 4 B, three major branched macromolecules were identified: (1) acetylation signals GlycA and GlycB of various N-acetylglucosinolates, mainly α-1-acid glycoproteins; (2) -N(CH3)3 signals from supramolecular phospholipid complexes; (3) phospholipid signals of different subcomponents of high-density lipoprotein and low-density lipoprotein (SPC-A and SPC-B, respectively).
[0083] 4. Drug sensitivity grouping
[0084] Based on serum biochemical indicators and tumor growth rate after drug administration, mice in the treatment group were divided into a high-sensitivity group (HS group) and a low-sensitivity group (LS group). Fifteen serum biochemical indicators, including liver function (8 items), kidney function (3 items), and blood lipids (4 items), were detected using a Chemray 800 fully automated chemical analyzer.
[0085] The above data was imported into the MetaboAnalyst 5.0 online data analysis platform (https: / / www.metaboanalyst.ca / ) for hierarchical clustering heatmap analysis. The treated mice could be divided into two groups: some mice clustered with the control group, while the remaining mice clustered in one group. Based on this clustering result, the mice clustered with the control group were classified as the LS group, and the remaining mice were classified as the HS group. Figure 5 The grouping results serve as classification labels for subsequent classification / prediction models built based on pre-drug NMR data.
[0086] 5. Construction and evaluation of classification models
[0087] Using the Statistical Analysis [one factor] module of the MetaboAnalyst 5.0 online data analysis platform (https: / / www.metaboanalyst.ca / ), a classification model was constructed based on the T2W and DIRE spectral data of plasma before drug administration in the treatment group and the grouping labels obtained from the above grouping results, using supervised random forest and support vector machine algorithms.
[0088] For T2W spectra, the feature dataset is a matrix consisting of the chemical shifts of the characteristic peaks of 31 metabolites and their piecewise integral values. For DIRE spectra, the feature dataset is a matrix consisting of the chemical shifts (53) of multiple characteristic peaks for each identified metabolite and their piecewise integral values.
[0089] Each random forest model consists of 500 decision trees. To evaluate the predictive performance of the classifier, the out-of-bag (OOB) error rate is used to assess the classification error. Results show that the OOB error of the random forest classification model built based on T2W spectral data is 0.125 (…). Figure 6 A), while the random forest classification model based on DIRE spectral data showed better classification ability, with an OOB error of 0.0625. Figure 6 B).
[0090] The support vector machine (SVM) classifier results show that when the number of variable levels is less than 6, the error rate of the SVM classification model built based on the T2W spectral data is 0.312. Figure 7A), while the error rate of the model based on DIRE spectral data is 0.125 (A). Figure 7 B).
[0091] The above results indicate that the random forest classification model based on pre-drug administration DIRE spectrum data has the smallest classification error, and this classifier will be selected to screen biomarkers for prediction in the future.
[0092] 6. Screening and predictive capabilities of biomarkers
[0093] In the BiomarkerAnalysis module of the MetaboAnalyst platform, the pre-drug DIRE profile feature set was input. The module's random forest classification algorithm was used to filter frequently selected variable features (8A-1). ROC analysis was performed on the top 5 frequently selected variable features (including methylene signals Lipid_8 and Lipid_24 from lipids, N-acetyl signals from GlycA, and glycosyl signals from GlycA and GlycB), yielding an AUC of 0.94 and a 95% confidence interval (CI) between 0.667 and 1. Figure 8 A-2). Using these 5 variables to predict the outcomes for 8 HS samples and 8 LS samples in the treatment group, the prediction results showed that, except for one sample from the LS group which was misclassified into the HS group, all other samples were correctly grouped. Figure 8 A-3).
[0094] The above results indicate that the signal intensities of methylene, GlycA, and GlyB lipids in the DIRE profile of pre-tumor-bearing mouse plasma can accurately predict their sensitivity to the mPD-1 inhibitor (RMP1-14). Specifically, based on... Figure 8 A-1 showed a significant positive correlation between lipid signal intensity and drug low sensitivity, while the signal intensity of GlycA and GlyB showed a significant correlation with drug high sensitivity.
[0095] Table 1
[0096]
[0097]
[0098]
[0099]
Claims
1. A method for constructing a drug sensitivity prediction model based on blood NMR editing spectra, the specific steps of which are as follows: 1) Queue Design In accordance with ethical requirements for trials, a cohort of volunteers suffering from a certain chronic non-communicable disease was recruited and divided into two groups: (1-1) Treatment group: Blood samples from patients who received a certain drug treatment, including plasma or serum samples before treatment and plasma or serum samples after treatment; (1-2) Blank control group: Plasma or serum samples from patients who did not receive treatment during the later stages of disease progression; 2) Hierarchical cluster analysis was performed on the blood biochemistry and prognostic indicators of the treatment group and the blank control group after treatment to divide the patients in the treatment group into a high-sensitivity group and a low-sensitivity group. 3) Obtain modeling datasets for the high-sensitivity group and the low-sensitivity group respectively. The modeling datasets include T2W-NMR spectral data and DIRE fingerprint data, wherein: The method for obtaining the modeling dataset is as follows: (3-1) Obtain one-dimensional plasma or serum samples from patients in the treatment group before treatment in step 1. 1 H NMR spectral data, the one-dimensional 1 The H NMR spectra are relaxation-weighted T2W-NMR and diffusion relaxation editing spectra DIRE; (3-2) Obtain the two-dimensional nuclear magnetic resonance spectrum of any plasma or serum sample from the patients in the treatment group before treatment in step 1, to assist in the identification of T2W-NMR spectrum signals; (3-3) The one-dimensional result obtained from (3-1) 1 The H NMR spectral data were processed to obtain datasets of T2W-NMR and DIRE spectra: (3-4) T2W-NMR spectral signal identification Signal identification of T2W-NMR spectra of plasma or serum samples based on (3-2) two-dimensional nuclear magnetic resonance spectroscopy and the human metabolome database HMDB; (3-5) DIRE Spectrum Signal Identification Signal identification of DIRE spectra in plasma or serum samples; 4) Model construction and biomarker screening (4-1) Model Construction and Evaluation Supervised random forest and support vector machine algorithms were used to model the plasma or serum samples of high-sensitivity and low-sensitivity groups before treatment. Four models were constructed, including a random forest model consisting of 500 decision trees, with out-of-bag (OOB) error used to evaluate the model's generalization ability; and leave-one-out cross-validation (LOOCV) was used to evaluate the generalization ability of the support vector machine model. Based on the evaluation results, the model with the lowest classification error was selected for drug sensitivity prediction.
2. The method according to claim 1, characterized in that, In step (3-1), a transverse relaxation time-weighted T2W-NMR spectrum is obtained using a water-presaturated 1Dcarr-purcell-meiboomm-gill pulse sequence, with a total spin echo time of 70 ms and a cycle delay of 2.0 s; and / or When acquiring diffusion and relaxation editing spectra (DIRE), three smooth square gradient fields were used to achieve diffusion editing with intensities of -17.13%, 13.13%, and 80%, respectively, and gradient lengths and diffusion delays of 1.53 ms and 120 ms, respectively; a total spin echo time of 33 ms was used to achieve relaxation editing.
3. The method according to claim 1, characterized in that, The specific operations of the data processing described in (3-3) are as follows: The one-dimensional result obtained in (3-1) 1 Phase and baseline corrections were performed on the 1H NMR spectrum. The chemical shift of the T2W-NMR spectrum was calibrated using the glucose α-isomer proton signal at δ 5.23 ppm. The DIRE spectrum selected the peak with the best resolution in the spectrum as the calibration signal. The linewidth factor was set to 0.3 Hz, and then Fourier transform was used to generate the frequency domain spectrum. For the one-dimensional 1 The chemical shift drift of the H NMR spectrum was corrected, and then segmented integration was performed with an integration width ranging from 6 to 24 Hz. The segmented integration was finally normalized based on the sum of the total integrations for each sample, and the datasets of the T2W-NMR spectrum and DIRE spectrum were obtained respectively.
4. The method according to claim 1, characterized in that, The identification of two fingerprint spectrum signals is achieved through (3-4) and (3-5). The T2W-NMR spectrum is the fingerprint spectrum of small molecule metabolites in the sample, and the DIRE spectrum is the fingerprint spectrum of glycoproteins, lipoproteins and lipid macromolecule metabolites containing small molecule branches in the sample.
5. The method according to claim 4, characterized in that, For T2W-NMR spectra, the feature dataset used to build the model is a matrix composed of the chemical shifts of the characteristic peaks of the small molecule metabolites identified in step (3-4) and their segmented integral values; the feature dataset for DIRE spectra is a matrix composed of the chemical shifts of the glycoproteins, lipoproteins and lipid macromolecule metabolites containing small molecule branches identified in step (3-5) and their segmented integral values.
6. The method according to claim 1, characterized in that, The method selects drugs and disease types with significant individual differences in efficacy in clinical applications. The drugs are cancer immunotherapy drugs such as PD-1 inhibitors, PD-L1 inhibitors, or CTLA-4 inhibitors; the diseases are non-small cell lung cancer, gastrointestinal cancer, or breast cancer.
7. A method for predicting drug sensitivity based on blood NMR editing spectra, using the prediction model constructed by the method described in any one of claims 1-6, the specific method being as follows: (4-2) Screening of biomarkers: Based on the model constructed by the method described above, important features are identified and ROC curve analysis is performed on all features. The ROC curves are generated by Monte Carlo cross-validation (MCCV) with balanced subsampling. Important features with biomarker functions are screened out. Then, the classification and prediction ability of the feature variables is evaluated based on the area under the ROC curve (AUC) of different combinations of feature variables. One or more biomarkers are screened out for drug sensitivity prediction. 5) Drug sensitivity prediction (5-1) Prepare the data to be predicted: Collect and process the blood NMR data of the patients to be treated according to step 3) so that the NMR data is converted into the input format expected by the model; (5-2) Input the obtained NMR data into the model and use the biomarkers or combinations thereof screened in step (4-2) to perform high-sensitivity and low-sensitivity group classification prediction.
Citation Information
Patent Citations
Metabolic phenotyping
US20050074745A1
Spectographic metabolite-signature for identifying a subject's susceptibility to drugs
US20200116703A1