Application of Bone Marrow Gene Markers in Predicting the Myeloablative Sensitivity to Sulfonate Alkylating Agents
The prediction model was constructed through the expression data of Lrg1, Lilr4b, H2-Eb1 and Ak4 gene markers, which solved the problem of difficulty in accurately predicting the myeloablative sensitivity of sulfonic acid alkylating agents in the prior art, achieved guidance on individualized treatment plans, and reduced the blindness of medication and adverse reactions.
Patent Information
- Application Number
- CN202411356754.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-09-27
AI Technical Summary
The prior art is difficult to accurately predict the patient's sensitivity to sulfonic acid alkylating agents, which leads to inaccurate medication regimens, which may cause poor treatment or toxic side effects, and lack of individualized treatment regimens.
The expression data of Lrg1, Lilr4b, H2-Eb1 and Ak4 gene markers were used to construct a prediction model through machine learning algorithms to predict the subject's sensitivity to sulfonic acid alkylating agents and output the prediction results.
Drug sensitivity can be predicted without blood drawing after administration, guide individualized treatment plans, and reduce blindness and adverse reactions in medication use.
Smart Images

Figure CN119446284B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent medicine and relates to the application of bone marrow gene markers in computer prediction of the sensitivity of sulfonic acid alkylating agents for myeloablation. Background Art
[0002] Sulfonic acid alkylating agents are a class of chemotherapeutic drugs that directly damage cellular DNA to prevent the proliferation of tumor cells. They can act on all stages of the cell cycle. Because they have a damaging effect on bone marrow hematopoietic stem cells, they are later used for myeloablative pretreatment before hematopoietic stem cell transplantation. Busulfan (BU) is a clinically commonly used sulfonic acid bifunctional alkylating agent and is usually used as part of a multi-drug myeloablative pretreatment regimen. It mainly causes premature changes in hematopoietic stem cells by inducing an increase in oxidative free radicals and promoting DNA damage, clears the hematopoietic stem cells in the transplanted patients, and provides a necessary implantation space for transplantation.
[0003] The steady-state blood drug concentration of busulfan is closely related to the prognosis of hematopoietic stem cell transplantation. Insufficient dosage will lead to poor transplantation effects, such as causing disease recurrence or graft rejection reactions; overdose will lead to an increase in the incidence of treatment-related toxic and side effects, such as sinusoidal obstructive syndrome and treatment-related death. There are significant differences in the steady-state blood drug concentration of busulfan after individual medication, and there are significant differences in the sensitivity of different patients / subjects to busulfan. It is difficult for doctors to accurately judge the sensitivity of patients / subjects to busulfan based on the clinical characteristics and clinical experience of patients, and it is difficult to formulate an accurate and effective medication plan. Currently, for the research on prediction indicators of busulfan drug sensitivity, metabolomics methods are mostly used to screen indicators that can predict the in-vivo drug concentration and clearance rate of BU. However, recent studies have found that the clearance rate of busulfan cannot predict the risk of hepatic veno-occlusive disease in hematopoietic stem cell transplantation patients, and other factors affecting the efficacy and toxicity of BU, such as genetic background, may need to be considered.
[0004] With the rapid development of molecular biology and cell biology, researchers and clinicians have begun to focus on how to use biomarkers to predict the sensitivity of patients / subjects to busulfan myeloablation. Developing a brand-new biomarker, prediction model, and system that can efficiently and accurately predict the sensitivity of busulfan-induced damage to bone marrow hematopoietic stem cells can provide more effective and accurate drug sensitivity evaluation indicators for patients / subjects, thus helping clinicians formulate individualized treatment plans for patients / subjects, reduce adverse reactions caused by blind medication, and monitor drug efficacy. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention aims to provide a method for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents, a computer system for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents, a computer device for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents, and a computer-readable storage medium.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect of the present invention, there is provided a method for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents, which is completed by a computer and includes the following steps:
[0008] Obtain data, that is, obtain the expression data of gene markers of a sample to be tested; the gene markers include one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4;
[0009] Process the data by inputting the expression data of the gene markers into a constructed prediction model, and the prediction model predicts the myeloablative sensitivity of the subject to sulfonate alkylating agents based on the expression data of the gene markers;
[0010] Output the prediction result.
[0011] Further, the expression data of the gene markers includes mRNA expression level data or protein expression level data.
[0012] Further, the mRNA expression level data includes, but is not limited to, data obtained by RT-PCR, qRT-PCR, in situ hybridization, and RNA sequencing.
[0013] Further, the protein expression level data includes, but is not limited to, data obtained by immunoblotting, immunohistochemistry, enzyme-linked immunosorbent assay, and mass spectrometry analysis.
[0014] Further, the construction steps of the prediction model are as follows:
[0015] Obtain the expression data of gene markers, the gene markers include one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4; the expression data of the gene markers comes from a population sensitive to myeloablative sulfonate alkylating agents and a population insensitive to myeloablative sulfonate alkylating agents;
[0016] Divide the expression data of the gene markers into a training set and a test set, extract the expression data of the gene markers in the training set and input them into a machine learning algorithm to construct a prediction model, and verify it through the test set to evaluate the model performance.
[0017] Further, the prediction model obtains the prediction result according to the following criteria: when the expression level of any one or more of the gene markers Lrg1, Lilr4b, and H2-Eb1 is higher than the threshold, and / or the expression level of Ak4 is lower than the threshold, a classification result that the test sample is sensitive to myeloablative sulfonate alkylating agents is obtained; when the expression level of any one or more of the gene markers Lrg1, Lilr4b, and H2-Eb1 is lower than the threshold, and / or the expression level of Ak4 is higher than the threshold, a classification result that the test sample is insensitive to myeloablative sulfonate alkylating agents is obtained.
[0018] Further, the machine learning algorithm includes algorithm models developed using various development tools.
[0019] Further, the development tools include, but are not limited to, TensorFlow, Scikit Learn, PyTorch, OpenNN, RapidMiner, Azure Machine Learning, Apache Mahout, Shogun, KNIME, VertexAI, H2Oai, Anaconda, Keras, Tableau, Fast.ai, Catalyst, Amazon ML, MLJAR, Spell.
[0020] Further, the algorithm models include, but are not limited to, linear regression model, logistic regression model, Lasso regression model, Ridge regression model, linear discriminant analysis model, nearest neighbor model, decision tree model, perceptron model, neural network model, support vector machine model, naive Bayes model, AdaBoost model, GBDT model, XGBoost model, LightGBM model, CatBoost model, or random forest model.
[0021] Further, the sulfonate alkylating agents include, but are not limited to, busulfan and melphalan.
[0022] Further, the sulfonate alkylating agent is busulfan.
[0023] Further, the myeloablation includes inducing bone marrow injury.
[0024] Further, the bone marrow injury includes bone marrow hematopoietic stem cell injury.
[0025] The second aspect of the present invention provides a computer system for predicting the sensitivity of a subject to myeloablative sulfonate alkylating agents, the system comprising:
[0026] A data acquisition module for acquiring gene marker expression data of a test sample, the gene markers including one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4;
[0027] A data classification module, which is used to classify and predict the data obtained in the data acquisition module through the prediction model described in the first aspect of the present invention, so as to obtain a classification result on whether the sample to be tested is sensitive to sulfonic acid alkylating agent conditioning;
[0028] An output module, which is used to output the classification result.
[0029] The third aspect of the present invention provides a computer device for predicting the sensitivity of a subject to sulfonic acid alkylating agent conditioning, and the computer device includes:
[0030] A processor, which is suitable for implementing each instruction; and
[0031] A memory, which is suitable for storing multiple instructions, and the instructions are suitable for being loaded and executed by the processor to perform the following steps:
[0032] Obtain data, that is, obtain the expression data of gene markers of the sample to be tested; the gene markers include one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4;
[0033] Process data, input the expression data of the gene markers into the prediction model described in the first aspect of the present invention, and the prediction model predicts the sensitivity of the subject to sulfonic acid alkylating agent conditioning based on the expression data of the gene markers;
[0034] Output the prediction result.
[0035] The fourth aspect of the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded and executed by the processor to perform the following steps:
[0036] Obtain data, that is, obtain the expression data of gene markers of the sample to be tested; the gene markers include one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4;
[0037] Process data, input the expression data of the gene markers into the prediction model described in the first aspect of the present invention, and the prediction model predicts the sensitivity of the subject to sulfonic acid alkylating agent conditioning based on the expression data of the gene markers;
[0038] Output the prediction result.
[0039] Advantages and beneficial effects of the present invention:
[0040] Based on the gene marker combination of the present invention, a corresponding computer system, computer device, computer-readable storage medium, and developed prediction model and method are developed. Without taking blood to detect drug AUC after administration, the drug sensitivity of the subject to sulfonic acid alkylating agent conditioning can be predicted, so as to guide clinicians to formulate individualized treatment plans for patients or subjects, reduce the blindness of drug use, and detect the drug efficacy. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic flowchart of a method for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents;
[0042] Figure 2 It is a schematic structural diagram of a computer system for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents;
[0043] Figure 3 It is a schematic structural diagram of a computer device for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents;
[0044] Figure 4 It is a graph of the correlation results of biological replicates of mice with different sensitivities, where A1 - A3 are the C57 control group, B1 - B3 are the C57 drug - administered group, C1 - C3 are the NUK control group, D1 - D3 are the NUK drug - administered group, E1 - E3 are the LAM control group, and F1 - F3 are the LAM drug - administered group;
[0045] Figure 5 It is a graph of the GO enrichment analysis results of differentially expressed genes before and after BU administration, Figure 5 In it, A is the GO enrichment analysis graph of differentially expressed proteins before and after drug administration in LAM mice, Figure 5 In it, B is the GO enrichment analysis graph of differentially expressed proteins before and after drug administration in C57 mice, Figure 5 In it, C is the GO enrichment analysis graph of differentially expressed proteins before and after drug administration in NUK mice;
[0046] Figure 6 It is a statistical graph of the gene expression levels detected by fluorescence quantitative PCR, where Figure 6 In it, A is the statistical graph of the expression level of Lrg1 in bone marrow cells, Figure 6 In it, B is the statistical graph of the expression level of Lilr4b in bone marrow cells, Figure 6 In it, C is the statistical graph of the expression level of H2 - Eb1 in bone marrow cells, Figure 6 In it, D is the statistical graph of the expression level of Ak4 in bone marrow cells;
[0047] Figure 7 It is an ROC curve graph, where Figure 7 In it, A is the ROC curve graph for diagnosing the LAM group and the NUK group with Lrg1 in bone marrow cells, Figure 7 In it, B is the ROC curve graph for diagnosing the LAM group and the NUK group with Lilr4b in bone marrow cells, Figure 7 In it, C is the ROC curve graph for diagnosing the LAM group and the NUK group with H2 - Eb1 in bone marrow cells, Figure 7 In it, D is the ROC curve graph for diagnosing the LAM group and the NUK group with Ak4 in bone marrow cells. DETAILED DESCRIPTION OF THE INVENTION
[0048] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0049] In some processes described in the specification and claims of the present invention and the above-mentioned accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit that "first" and "second" are of different types.
[0050] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present invention.
[0051] Figure 1 It is a schematic flowchart of a method for predicting the sensitivity of a subject to myeloablation by sulfonate alkylating agents provided by the present invention. Specifically, the method includes the following:
[0052] 101: Obtain data, obtain the expression data of the gene markers of the sample to be tested; the gene markers include one or more of the following : Lrg1, Lilr4b, H2-Eb1, Ak4 ;
[0053] In some embodiments, through extensive and in-depth research, the present invention selects three different strains of GD mice with different sensitivities to the myeloablative effect of busulfan, namely high-sensitive NUK, medium-sensitive C57, and low-sensitive LAM, studies the relationship between bone marrow genes and busulfan drug sensitivity, screens gene markers indicating differences in busulfan myeloablation sensitivity, and finds that Lrg1, Lilr4b, H2-Eb1, and / or Ak4 show significant differences among the three strains of GD mice, suggesting that Lrg1, Lilr4b, H2-Eb1, and / or Ak4 can be used as good markers for predicting or promoting busulfan myeloablation sensitivity.
[0054] In the present invention, the term "gene marker" refers to a gene, gene product, which is a target for regulating one or more target phenotypes.
[0055] In some embodiments, the term also encompasses measurable entities that have been determined to indicate targets of a desired output, such as one or more diagnostic, prognostic, predictive drug sensitivity, and / or treatment outputs.
[0056] In some embodiments, the term also encompasses compositions that modulate genes or gene products, including anti-gene product antibodies and antigen-binding fragments thereof. Thus, gene markers include, but are not limited to, nucleic acids (e.g., genomic nucleic acids and / or transcribed nucleic acids), proteins, and antibodies (and antigen-binding fragments thereof).
[0057] In the present invention, the term "expression data" refers to the amount, accumulation, or rate of gene marker molecules. Expression levels can be represented by: the amount or synthesis rate of messenger RNA (mRNA) encoded by a gene, the amount or synthesis rate of a polypeptide or protein encoded by a gene, or the amount or synthesis rate of a biomolecule accumulated in a cell or biological fluid.
[0058] In a specific embodiment of the present invention, the expression data of gene markers is obtained by the following method: After lysing red blood cells of bone marrow cells from three strains of mice, LSK cells are sorted, total RNA is extracted, reverse transcribed into cDNA, and then qRT-PCR is performed. The 2 -ΔΔCt method is used for calculating the expression level to obtain the expression data of gene markers.
[0059] In a specific embodiment of the present invention, the step of obtaining the expression data of gene markers includes LSK cell sorting:
[0060] a. After lysing red blood cells of bone marrow cells from GD mice of three strains, biotin-labeled mixed primary antibodies CD4, CD8, B220, Ter119, Gr1, and cd11b are added in proportion and incubated at 4°C for 30 min;
[0061] b. Add 3 mL of PBS and wash once. The centrifugation conditions are 1500 r / min, 4°C, for 5 min, and discard the supernatant.
[0062] c. Resuspend with 1 mL of PBS and add the mixed antibodies Streptavidin (percp-labeled), Sca-1 (PE-labeled), and c-kit (APC-labeled) according to the ratio in Table 2, and incubate in the dark at 4°C for 1 h;
[0063] d. Add 3 mL of PBS and wash once. The centrifugation conditions are the same as above;
[0064] e. Sort LSK cells (Lin−c-kit+Sca-1+) by flow cytometry. The sorted LSK cells (about 1×10 for each sample 5The cells (three samples in each group) were evenly dispersed in 0.1 mL of Trizol solution for RNA-sequence detection.
[0065] In a specific embodiment of the present invention, the step of obtaining the expression data of the gene marker includes RNA extraction: total RNA was extracted, and its concentration and OD260 / OD280 were measured using a NanoDrop 2000 spectrophotometer (Thermo Scientific, USA), and the integrity of the RNA was detected by agarose gel electrophoresis.
[0066] In a specific embodiment of the present invention, the step of obtaining the expression data of the gene marker includes reverse transcription: the RNA to be detected was reverse transcribed into cDNA using the TransScript All-in-One First-Strand cDNA Synthesis SuperMIX for qPCR kit. The reverse transcription system was: 0.5 μg of total RNA; 2 μl of 5×TransScript All-in-one SuperMix for qPCR; 0.5 μl of gDNA Remover; add Nuclease-free H2O to 10 μl. The reaction program was: 42°C for 15 min, 85°C for 5 s. After reverse transcription, 90 μl of Nuclease-free H2O was added and stored in a -20°C refrigerator for later use.
[0067] In a specific embodiment of the present invention, the step of obtaining the expression data of the gene marker includes fluorescence quantitative PCR: using the PerfectStartTM Green qPCR SuperMix kit to perform the reaction on a LightCycler 480 II fluorescence quantitative PCR instrument (Roche, Swiss). The system was: 5 μl of 2×PerfectStartTM Green qPCR SuperMix; 0.2 μl of 10 μM Forward primer; 0.2 μl of 10 μM Reverse primer; 1 μl of cDNA; 3.6 μl of Nuclease-free H2O. The PCR program was: 94°C for 30 s; 94°C for 5 s, 60°C for 30 s, for 45 cycles. After the cycle ended, the specificity of the product was detected using a melting curve: slowly heating from 60°C to 97°C, and collecting fluorescence signals 5 times per °C.
[0068] In a specific embodiment of the present invention, the step of obtaining the expression data of the gene marker includes the calculation of the expression level: the expression data of the gene marker was obtained by using the -ΔΔCt 2−ΔΔCt method, and the results are as Figure 6 shown.
[0069] Lrg1 includes wild-type, mutant, or fragments thereof. The term encompasses full-length, unprocessed Lrg1, as well as any form of Lrg1 derived from processing in cells. The term encompasses naturally occurring variants of Lrg1 (e.g., splice variants or allelic variants). The term encompasses, for example, the Lrg1 gene, human Lrg1, and Lrg1 from any other vertebrate source, including mammals such as primates and rodents (e.g., mice and rats). As a preferred embodiment, in the present invention, Lrg1 is a human gene with gene ID 116844.
[0070] Lilr4b includes wild-type, mutant, or fragments thereof. The term encompasses full-length, unprocessed Lilr4b, as well as any form of Lilr4b derived from processing in cells. The term encompasses naturally occurring variants of Lilr4b (e.g., splice variants or allelic variants). The term encompasses, for example, the Lilr4b gene, human Lilr4b, and Lilr4b from any other vertebrate source, including mammals such as primates and rodents (e.g., mice and rats). As a preferred embodiment, in the present invention, Lilr4b is a human gene with gene ID 14727.
[0071] H2-Eb1 includes wild-type, mutant, or fragments thereof. The term encompasses full-length, unprocessed H2-Eb1, as well as any form of H2-Eb1 derived from processing in cells. The term encompasses naturally occurring variants of H2-Eb1 (e.g., splice variants or allelic variants). The term encompasses, for example, the H2-Eb1 gene, human H2-Eb1, and H2-Eb1 from any other vertebrate source, including mammals such as primates and rodents (e.g., mice and rats). As a preferred embodiment, in the present invention, H2-Eb1 is a human gene with gene ID 14969.
[0072] Ak4 includes wild-type, mutant, or fragments thereof. The term encompasses full-length, unprocessed Ak4, as well as any form of Ak4 derived from processing in cells. The term encompasses naturally occurring variants of Ak4 (e.g., splice variants or allelic variants). The term encompasses, for example, the Ak4 gene, human Ak4, and Ak4 from any other vertebrate source, including mammals such as primates and rodents (e.g., mice and rats). As a preferred embodiment, in the present invention, Ak4 is a human gene with gene ID 205.
[0073] In the context of the present invention, the term "sample" refers to a composition obtained from or derived from a patient / subject, which contains cells and / or other molecular entities to be characterized and / or identified according to physical, biochemical, chemical, and / or physiological characteristics. For example, a sample refers to any sample derived from a patient / subject that is expected or known to contain cells and / or molecular entities to be characterized. Samples include, but are not limited to, tissue samples, primary or cultured cells or cell lines, cell cultures, cell supernatants, cell lysates, platelets, serum, plasma, vitreous humor, lymph fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tissue culture medium, tissue extracts, homogenized tissue, cell extracts, and combinations thereof.
[0074] In a specific embodiment of the present invention, the sample is a bone marrow cell lysate extract.
[0075] The term "RT-PCR method" is also known as "reverse transcription polymerase chain reaction" or "retrotranscription polymerase chain reaction", which is a technique that combines the reverse transcription (RT) of RNA and the polymerase chain amplification (PCR) of cDNA. First, cDNA is synthesized from RNA by the action of reverse transcriptase, and then the target fragment is amplified and synthesized using cDNA as a template under the action of DNA polymerase. The RT-PCR technique is sensitive and versatile, and can be used to detect the gene expression level in cells, the content of RNA viruses in cells, and directly clone the cDNA sequence of a specific gene.
[0076] The term "qRT-PCR" is also known as "quantitative real-time polymerase chain reaction", which refers to the use of changes in fluorescence signals to detect the changes in the amount of amplification products in each cycle of the PCR amplification reaction in real time, and finally perform accurate quantitative analysis of the starting template.
[0077] The term "RNA sequencing method" is also known as "transcriptome sequencing" or "RNA-Seq", which can comprehensively and rapidly obtain the sequence information and expression information of almost all transcripts in a specific cell or tissue in a certain state, including mRNA encoding proteins and various non-coding RNAs, and the expression abundance of different transcripts generated by gene alternative splicing. While analyzing the structure and expression level of transcripts, it also discovers unknown transcripts and rare transcripts, thereby accurately analyzing important issues in life sciences such as gene expression differences, gene structure variations, and screening of molecular markers. The RNA sequencing method can detect the overall transcriptional activity of any species without pre-designing probes for known sequences. The transcriptome generally refers to the collection of all transcriptional products in a cell under a certain physiological condition, including messenger RNA, ribosomal RNA, transfer RNA, and non-coding RNA; narrowly, it refers to the collection of all mRNAs.
[0078] In a specific embodiment of the present invention, the expression data of gene markers is obtained in the following manner: Bone marrow cells of three strains of mice before and after administration are extracted for LSK cell sorting, followed by RNA extraction and library construction, and RNA sequencing, differential expression gene analysis, and gene enrichment analysis are performed on the transcriptome library.
[0079] In a specific embodiment of the present invention, LSK cell sorting is included in the steps for obtaining the expression data of gene markers:
[0080] a. Approximately 3×10 8 bone marrow cells can be isolated from each mouse, and biotin-labeled mixed primary antibodies CD4, CD8, B220, Ter119, Gr1, and cd11b are added according to the ratio in Table 1 and incubated at 4°C for 30 min;
[0081] b. Add 3 mL of PBS for one wash, centrifuge at 1500 r / min, 4°C for 5 min, discard the supernatant,
[0082] c. Resuspend with 1 mL of PBS, and add mixed antibodies Streptavidin (percp-labeled), Sca-1 (PE-labeled), and c-kit (APC-labeled) according to the ratio in Table 2 and incubate in the dark at 4°C for 1 h;
[0083] d. Add 3 mL of PBS for one wash, and centrifuge under the same conditions as above;
[0084] e. Sort LSK cells (Lin-, c-kit+, Sca-1+) by flow cytometry. The sorted LSK cells (about 1×10 5 cells per sample, three samples per group) are evenly dispersed in 0.1 mL of Trizol solution for RNA-sequence detection.
[0085] Table 1. Formulation of biotin-labeled mixed primary antibodies
[0086]
[0087] Table 2. Formulation of secondary antibodies
[0088]
[0089]
[0090] In a specific embodiment of the present invention, the step of obtaining expression data of gene markers includes RNA extraction and library construction: Total RNA is extracted using TRIzol reagent according to the instructions. The purity and quantification of RNA are identified using a NanoDrop 2000 spectrophotometer (Thermo Scientific, USA), and the integrity of RNA is evaluated using an Agilent 2100 Bioanalyzer (Agilent Technologies, Santa Clara, CA, USA). A transcriptome library is constructed using the VAHTS Universal V5 RNA-seq Library Prep kit according to the instructions. Transcriptome sequencing and analysis are performed by Shanghai OE Biotech Co., Ltd. (Shanghai, China).
[0091] In a specific embodiment of the present invention, the step of obtaining expression data of gene markers includes performing RNA sequencing on the transcriptome library: The library is sequenced using an Illumina Novaseq 6000 sequencing platform to generate 150bp paired-end reads. The fastp software is used to process the raw reads in fastq format, and clean reads are obtained after removing low-quality reads for subsequent data analysis. The HISAT2 software is used for reference genome alignment, and gene expression levels (FPKM) are calculated, and the read counts (counts) of each gene are obtained through HTSeq-count. PCA analysis and plotting of genes (counts) are performed using R (v 3.2.0) to evaluate sample biological replicates, and the results are as Figure 4 shown.
[0092] In a specific embodiment of the present invention, the step of obtaining expression data of gene markers includes differential expression gene analysis: DESeq2 software is used for differential expression gene analysis, where genes that meet the thresholds of q value < 0.05 and fold change > 2 or fold change < 0.5 are defined as differentially expressed genes (DEGs). Hierarchical clustering analysis of DEGs is performed using R (v 3.2.0) to show the expression patterns of genes in different groups and samples. The R package ggradar is used to plot a radar chart of the top 30 genes to show the expression changes of upregulated or downregulated genes.
[0093] In a specific embodiment of the present invention, the step of obtaining the expression data of gene markers includes gene enrichment analysis: performing GO, KEGG Pathway, Reactome, and WikiPathways enrichment analysis on differentially expressed genes based on the hypergeometric distribution algorithm to screen for significantly enriched functional entries. Using R (v 3.2.0) to draw bar charts, chord diagrams, or enrichment analysis circle diagrams for significantly enriched functional entries.
[0094] In a specific embodiment of the present invention, gene set enrichment analysis is performed using GSEA software. Using predefined gene sets, genes are ranked according to their differential expression levels in two types of samples, and then it is examined whether the predefined gene sets are enriched at the top or bottom of this ranking list.
[0095] As Figure 4 shown, the biological replicates within different sensitivity mouse control groups and drug administration groups have good correlations. Among them, the samples of the C57 strain control group (Group A) and the drug administration group (Group B), as well as the samples of the LAM mouse control group (Group E) and the drug administration group (Group F) have very strong correlations, indicating that the sample compositions before and after drug treatment of the C57 and LAM strains are similar; the samples of the NUK mouse control group (Group C) and the drug administration group (Group D) have relatively weak correlations, suggesting that BU has a greater impact on NUK than on the other two strains; all samples have good biological replicate correlations, ensuring the credibility of the DEGs screened by management.
[0096] The results of GO analysis are as Figure 5 shown in A of. In terms of biological processes, the differentially expressed genes of the LAM, C57, and NUK strains are all enriched in entries related to blood cell development and regulation processes, mainly related processes of red blood cells and platelets. The differentially expressed genes of LAM and NUK are also enriched in entries related to immune responses. The differences are that the differentially expressed genes of LAM are also related to apoptotic cell clearance, positive regulation of peptidyl-tyrosine phosphorylation, and positive regulation of the ERK1 and ERK2 cascades; the differentially expressed genes of C57 are also related to mitochondrial translation and regulation of transcription by RNA polymerase; the differentially expressed genes of NUK are also related to inflammatory responses, neutrophil chemotaxis processes, tumor necrosis factor regulation processes, and phagocytosis processes.
[0097] In terms of cell composition, as Figure 5 shown in B of, the differentially expressed genes of the LAM, C57, and NUK strains are all enriched in various membrane and cytoplasmic related structures. The differences are that the differentially expressed genes of LAM and NUK are also enriched in extracellular structures and cytoplasm, while C57 is enriched in nuclear and mitochondrial related structures.
[0098] In terms of molecular functions, as Figure 5As shown by C in [reference], the differential genes of the LAM, C57, and NUK strains were all enriched in the related functions of various proteins, chemokines, and enzyme binding. The difference is that the differential genes of LAM and NUK are also related to the processes of binding to various metal ions.
[0099] In the present invention, the term "myeloablation" refers to the pretreatment step before hematopoietic stem cell transplantation, which generally needs to achieve the following purposes: removing the patient's hematopoietic stem cells from the bone marrow niche to enable the implanted donor hematopoietic stem cells to obtain sufficient space to support their proliferation and differentiation. Myeloablative pretreatment regimens are generally divided into those based on total body irradiation and those based on high-dose chemotherapy drugs.
[0100] In some embodiments, the chemotherapy drug is an alkylating agent; further, the alkylating agent is a sulfonic acid alkylating agent; further, the sulfonic acid alkylating agent includes, but is not limited to, busulfan and melphalan; further, the sulfonic acid alkylating agent is busulfan.
[0101] In the specific embodiments of the present invention, busulfan was used to induce damage to hematopoietic stem cells of different strains of mice. Information on different strains of mice is as follows: 10 SPF-grade male LAM-DA (hereinafter referred to as LAM) and NUK strain genetic diversity (GD) mice, 8-10 weeks old, were provided by the Institute of Laboratory Animal Sciences, Chinese Academy of Medical Sciences
SCXK(Beijing)2014-0011
SCXK(Beijing)2014-0004
SYXK(Beijing)2015-0035
[0102] In a specific embodiment of the present invention, the experimental grouping and modeling steps for inducing hematopoietic stem cell damage in different strains of mice using busulfan are as follows: 10 mice of each of the three strains LAM, C57, and NUK are randomly divided into a control group and a busulfan administration group, with 5 mice in each group, to observe the differences in the sensitivity of different strains of GD mice to busulfan-induced bone marrow hematopoietic stem cell damage. Busulfan powder is prepared as a stock solution of 80 mg / mL with DMSO and diluted to the required concentration with PBS when in use. The administration group is intraperitoneally injected with 40 mg / kg of busulfan, administered in two days, 20 mg / kg per day, and the control group of mice is given PBS containing 5% DMSO in an equal volume as a control.
[0103] In a specific embodiment of the present invention, the effect of busulfan-induced hematopoietic stem cell damage in different strains of mice is evaluated by measuring the colony-forming ability of bone marrow cells: 14 days after busulfan administration, mouse bone marrow cells are taken under sterile conditions, and the concentration of mouse bone marrow is adjusted to 4×10 5 cells / mL with PBS. 0.2 mL of cells from each sample is added to 1 mL of pre-aliquoted M3434 medium, mixed well with an oscillator, allowed to stand for 2 - 5 min, and then added to a 24-well plate after the bubbles disappear, with a volume of 0.5 mL per well. Gently shake the culture plate to evenly distribute the medium. The 24-well plate is placed in an incubator at 37°C with 5% carbon dioxide for culture. The number of granulocyte-macrophage colony-forming units (CFU-GM) is read on the 7th day to detect the function of mouse hematopoietic stem cells. After the three strains of mice receive busulfan administration, the number of in vitro colony formations of bone marrow cells decreases, and there are significant differences (P < 0.05). This shows that busulfan has an inhibitory effect on the proliferation of hematopoietic stem and progenitor cells, and compared with C57, NUK is more sensitive to busulfan drug treatment, and LAM has a relatively lower sensitivity to busulfan.
[0104] 102: Process data, input the expression data of the gene markers into the constructed prediction model, and the prediction model predicts the sensitivity of the subject to myeloablative sulfonate alkylating agents based on the expression data of the gene markers;
[0105] In some embodiments of the present invention, the method for constructing the prediction model belongs to what is known to those skilled in the art, and the steps of associating the expression level of gene markers with a certain probability or risk can be implemented and achieved in different ways. Preferably, the measured concentrations of the markers and one or more other markers are mathematically combined, and the combined value is associated with the underlying drug sensitivity problem. The measured combinations of marker values can be combined by any suitable existing mathematical methods in the art, and a prediction model can be constructed through machine learning algorithms.
[0106] In the context of the present invention, the term "machine learning" refers to the use of a computer to simulate or implement human learning activities, and those skilled in the art usually use different development tools to construct algorithm models for machine learning. The development tools include, but are not limited to, TensorFlow, Scikit Learn, PyTorch, OpenNN, RapidMiner, Azure Machine Learning, Apache Mahout, Shogun, KNIME, Vertex AI, H2Oai, Anaconda, Keras, Tableau, Fast.ai, Catalyst, Amazon ML, MLJAR, Spell. The algorithm models include, but are not limited to, linear regression models, logistic regression models, Lasso regression models, Ridge regression models, linear discriminant analysis models, nearest neighbor models, decision tree models, perceptron models, neural network models, support vector machine models, naive Bayes models, AdaBoost models, GBDT models, XGBoost models, LightGBM models, CatBoost models, or random forest models.
[0107] After constructing the prediction model, the diagnostic efficacy of the prediction model is evaluated using ROC curve analysis.
[0108] An ROC curve is a plot of the true positive rate (sensitivity) of a test against the false positive rate (100% - specificity) of the test. It is useful for depicting the performance of a particular feature (e.g., any entry of any biomarker and / or additional biomedical information described in the present invention) when differentiating between two populations (e.g., individuals responsive and non-responsive to a therapeutic agent). Generally, feature data is selected across the entire population in ascending order based on the value of a single feature. Then, for each value of that feature, the true positive and false positive rates of the data are calculated. The true positive rate is determined by counting the number of cases above the value of that feature and dividing by the total number of cases. The false positive rate is determined by counting the number of controls above the value of that feature and dividing by the total number of controls. Although this definition refers to the situation where the feature is elevated in cases compared to controls, this definition also applies to the situation where the feature is lower in cases compared to controls (in which case, samples below the value of that feature will be counted). An ROC curve can be generated for an individual feature and can also be generated for other individual outputs. For example, combinations of two or more features can be mathematically combined (e.g., added, subtracted, multiplied, etc.) to provide a single summed value, and this single summed value can be plotted in the ROC curve. Additionally, any combination of multiple features whose combination is derived from individual output values can be plotted in the ROC curve.
[0109] In some embodiments of the present invention, the ROC curve is generated for a single feature, specifically referring to the ROC curve of a single gene marker for predicting the sensitivity of a patient / subject to busulfan myeloablation.
[0110] In a specific embodiment of the present invention, the ROC working curve of the subject is plotted using graphpad, and the diagnostic efficacy of the AUC value, sensitivity, and specificity judgment indexes is analyzed. The results show that the diagnostic efficacy of Lrg1, Lilr4b, H2-Eb1, and Ak4 is as Figure 7 shown. The results show that the AUC value of the Lrg1 expression level of LSK cells for diagnosing mice in the LAM group and the NUK group is 1, the sensitivity is 100%, and 100% - specificity% is 0%; the AUC value of the Lilr4b expression level of LSK cells for diagnosing mice in the LAM group and the NUK group is 0.8667, the sensitivity is 100%, and 100% - specificity% is 40%; the AUC value of the H2-Eb1 expression level of LSK cells for diagnosing mice in the LAM group and the NUK group is 1, the sensitivity is 100%, and 100% - specificity% is 0%; the AUC value of the Ak4 expression level of LSK cells for diagnosing mice in the LAM group and the NUK group is 1, the sensitivity is 100%, and 100% - specificity% is 0%. It shows that Lrg1, Lilr4b, H2-Eb1, and Ak4 can predict the sensitivity of busulfan-induced bone marrow injury.
[0111] In the context of the present invention, "AUC" refers to the area under the curve, which is a measure of the ability of a classifier to distinguish classes and is used as a summary of the ROC curve. It can usually be obtained by the integral method or the trapezoidal method. AUC measurement is useful for comparing the accuracy of classifiers across the entire data range. A higher AUC means that the classifier has better performance and can classify more accurately between two target groups.
[0112] In a specific embodiment of the present invention, the two target groups refer to two groups that are sensitive and insensitive to busulfan myeloablation.
[0113] 103: Output the prediction result.
[0114] Figure 2 Showing a schematic structural diagram of a computer system for predicting the myeloablation sensitivity of a subject to sulfonate alkylating agents provided by the present invention.
[0115] The system is programmed or otherwise configured to include a data acquisition module 201, a data classification module 202, and an output module 203;
[0116] The data acquisition module 201 is used to acquire the gene marker expression data of the sample to be tested, and the gene markers include one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4;
[0117] A data classification module 202, configured to perform classification prediction on the data obtained by the data acquisition module by using the prediction model obtained through the construction method described in the first aspect of the present invention, so as to obtain a classification result on whether the sample to be tested is sensitive to sulfonic acid alkylating agent conditioning;
[0118] An output module 203, configured to output the classification result.
[0119] The system may be a user's electronic device or a computer system remotely located relative to the electronic device.
[0120] Figure 3 Shown is a schematic structural diagram of a computer device for predicting the sensitivity of a subject to sulfonic acid alkylating agent conditioning provided by the present invention.
[0121] The computer device 300 includes a processor 301 and a memory 302 coupled to the processor 301. Program instructions are stored in the memory 302. When the program instructions are executed by the processor 301, the processor 301 is caused to execute the steps of predicting the sensitivity of a subject to sulfonic acid alkylating agent conditioning described in any of the above embodiments.
[0122] Among them, the processor 301 may also be referred to as a CPU (Central Processing Unit). The processor 301 may be an integrated circuit chip having signal processing capabilities. The processor 301 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0123] The computer device 300 may be a mobile electronic device.
[0124] The storage medium of the embodiments of the present invention stores program instructions capable of implementing the method for predicting the sensitivity of a subject to sulfonic acid alkylating agent conditioning. Among them, the program instructions may be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a computer device such as a computer, a server, a mobile phone, or a tablet.
[0125] It should be understood that the systems, devices, and methods described in the present invention can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.
[0126] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0127] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0128] The above is only the implementation manner of this application, and does not limit the patent scope of this application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A computer system for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents, the system comprising: A data acquisition module for acquiring gene marker expression data of a sample to be tested, the gene markers including one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4; A data classification module for classifying and predicting the data obtained in the data acquisition module through a prediction model to obtain a classification result of whether the sample to be tested is sensitive to myeloablative sulfonate alkylating agents; An output module for outputting the classification result; The sulfonate alkylating agent is busulfan; The subject is a mouse.
2. The system according to claim 1, wherein The prediction model is constructed through the following steps: Obtain the expression data of gene markers, the gene markers including one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4; the expression data of the gene markers are from mice sensitive to myeloablative sulfonate alkylating agents and mice insensitive to myeloablative sulfonate alkylating agents; Divide the expression data of the gene markers into a training set and a test set, extract the expression data of the gene markers in the training set and input them into a machine learning algorithm to construct a prediction model, and verify it through the test set to evaluate the model performance.
3. The system according to claim 1, characterized in that, The prediction model obtains a prediction result through the following criteria: When the expression level of any one or more of the gene markers Lrg1, Lilr4b, H2-Eb1 is higher than a threshold value, and / or the expression level of Ak4 is lower than the threshold value, a classification result that the sample to be tested is sensitive to myeloablative sulfonate alkylating agents is obtained; when the expression level of any one or more of the gene markers Lrg1, Lilr4b, H2-Eb1 is lower than the threshold value, and / or the expression level of Ak4 is higher than the threshold value, a classification result that the sample to be tested is insensitive to myeloablative sulfonate alkylating agents is obtained.
4. The system according to claim 2, wherein The machine learning algorithm includes algorithm models developed using various development tools.
5. The system according to claim 4, characterized in that The development tools include but are not limited to TensorFlow, Scikit Learn, PyTorch, OpenNN, RapidMiner, Azure Machine Learning, Apache Mahout, Shogun, KNIME, Vertex AI, H2Oai, Anaconda, Keras, Tableau, Fast.ai, Catalyst, Amazon ML, MLJAR, Spell.
6. The system according to claim 4, characterized in that, The algorithm models include but are not limited to linear regression models, logistic regression models, Lasso regression models, Ridge regression models, linear discriminant analysis models, nearest neighbor models, decision tree models, perceptron models, neural network models, support vector machine models, naive Bayes models, AdaBoost models, GBDT models, XGBoost models, LightGBM models, CatBoost models or random forest models.
7. A computer device for predicting the myeloablative sensitivity of a subject to sulfonate alkylating agents, the computer device comprising: A processor adapted to implement each instruction; And A memory, adapted to store a plurality of instructions, the instructions being adapted to be loaded and executed by a processor to perform the following steps: Obtain data, and obtain the expression data of gene markers of a sample to be tested; the gene markers include one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4; Process the data, and input the expression data of the gene markers into a constructed prediction model, the prediction model predicting the sensitivity of a subject to myeloablative conditioning with a sulfonic acid alkylating agent based on the expression data of the gene markers; Output a prediction result; The sulfonic acid alkylating agent is busulfan; The subject is a mouse.
8. A computer-readable storage medium, in which a plurality of instructions are stored, the instructions being adapted to be loaded and executed by a processor to perform the following steps: Obtain data, and obtain the expression data of gene markers of a sample to be tested; the gene markers include one or more of the following: Lrg1, Lilr4b, H2-Eb1, Ak4; Process the data, and input the expression data of the gene markers into a constructed prediction model, the prediction model predicting the sensitivity of a subject to myeloablative conditioning with a sulfonic acid alkylating agent based on the expression data of the gene markers; Output a prediction result; The sulfonic acid alkylating agent is busulfan; The subject is a mouse.
Citation Information
Patent Citations
Application of biomarker in predicting sensitivity of sulfonic acid alkylating agent to induced bone marrow injury
CN116397020A
Methods of characterizing immunosuppressive neutrophils
CN116710777A