Composition and method for predicting curative effect of diffuse large B-cell lymphoma chemotherapy

CN121605201APending Publication Date: 2026-03-03BOE TECHNOLOGY GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480001201.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

The existing R-CHOP treatment regimen has insufficient predictive accuracy for the efficacy of treatment in patients with diffuse large B-cell lymphoma. Approximately 20% of patients have refractory disease and poor survival outcomes, with a median overall survival (OS) of only 6.3 months.

Method used

By designing a composition comprising several different decoy oligonucleotides configured to hybridize to DNA molecules derived from multiple target genomic regions, enriching and sequencing cfDNA fragments using differential methylation features, and combining this with a trained classifier to predict chemotherapy efficacy.

Benefits of technology

It improves the predictive accuracy of chemotherapy efficacy in diffuse large B-cell lymphoma, helps identify patients who respond to and do not respond to chemotherapy, and optimizes treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121605201A_ABST
    Figure CN121605201A_ABST
Patent Text Reader

Abstract

A composition, the composition comprising: a number of different bait oligonucleotides configured to collectively hybridize to a DNA molecule derived from a plurality of genomic regions of interest; wherein each genomic region in the plurality of target genomic regions is differentially methylated in a person who is effective in diffuse large B-cell lymphoma chemotherapy, compared to a person who is ineffective in diffuse large B-cell lymphoma chemotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

Compositions and methods for predicting chemotherapy efficacy for diffuse large b-cell lymphoma TECHNICAL FIELD

[0001] The present disclosure relates to the field of biotechnology, and in particular, to a composition and methods for predicting chemotherapy efficacy for diffuse large B-cell lymphoma. BACKGROUND

[0002] Diffuse large B-cell lymphoma (DLBCL) is an aggressive tumor derived from mature B cells and is the most common type of non-Hodgkin lymphoma. Rituximab plus cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP) can effectively improve the prognosis of DLBCL patients and is considered as the standard first-line treatment, with a 10-year overall survival (OS) of 43.5%.

[0003] SUMMARY

[0004] In one aspect, a composition is provided, the composition comprising: a plurality of different decoy oligonucleotides configured to collectively hybridize to DNA molecules derived from a plurality of target genomic regions; wherein each genomic region of the plurality of target genomic regions is differentially methylated in a diffuse large B-cell lymphoma chemotherapy responder compared to a diffuse large B-cell lymphoma chemotherapy non-responder.

[0005] In some embodiments, the plurality of different decoy oligonucleotides is configured to hybridize to DNA molecules derived from at least 20%, at least 25%, or at least 50% of the plurality of target genomic regions of any one of Tables 1-33.

[0006] In some embodiments, the plurality of different decoy oligonucleotides is configured to hybridize to DNA molecules derived from at least 20%, at least 25%, or at least 50% of the plurality of target genomic regions of Tables 1-33.

[0007] In some embodiments, the plurality of different decoy oligonucleotides is configured to hybridize to DNA molecules derived from at least 20% of the plurality of target genomic regions of Tables 1-33.

[0008] In some embodiments, the plurality of DNA molecules is derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of Tables 1-33.

[0009] In another aspect, a composition is provided, the composition comprising: a plurality of different bait oligonucleotides configured to hybridize to a plurality of DNA molecules derived from at least 20% of the plurality of target genomic regions of any one of Tables 1-33.

[0010] In some embodiments, the plurality of different bait oligonucleotides are configured to hybridize to a plurality of DNA molecules derived from 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of any one of Tables 1-33.

[0011] In some embodiments, the plurality of different bait oligonucleotides are configured to hybridize to a plurality of DNA molecules derived from at least 20% of the plurality of target genomic regions of Tables 1-33.

[0012] In some embodiments, the plurality of DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of Tables 1-33.

[0013] In some embodiments, the plurality of DNA molecules are converted cfDNA fragments.

[0014] In some embodiments, the plurality of target genomic regions are hypermethylated regions, hypomethylated regions, or bimodal regions that can be hypermethylated or hypomethylated.

[0015] In some embodiments, the plurality of bait oligonucleotides are configured to hybridize to hypermethylated converted DNA molecules, hypomethylated converted DNA molecules, or both hypermethylated and hypomethylated converted DNA molecules derived from each target genomic region.

[0016] In some embodiments, each of the plurality of bait oligonucleotides is bound to an affinity moiety.

[0017] In some embodiments, each of the plurality of different bait oligonucleotides is bound to a magnetic bead surface.

[0018] In another aspect, a method for enriching converted cfDNA fragments that can provide information on diffuse large B-cell lymphoma chemotherapy effect is provided, the method comprising the steps of: contacting a composition as in any one of the above embodiments with DNA derived from a test subject, and enriching a sample of cfDNA corresponding to a plurality of genomic regions associated with diffuse large B-cell lymphoma chemotherapy effect by hybridization capture.

[0019] In yet another aspect, a method for obtaining sequence information that can provide information on the efficacy of chemotherapy for diffuse large B-cell lymphoma is provided, the method comprising the steps of: a) enriching converted DNA from a test subject by contacting the converted DNA with a composition as described in any of the above embodiments, and b) sequencing the enriched converted DNA.

[0020] In yet another aspect, a method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma is provided, comprising the steps of: a) capturing a plurality of cfDNA fragments from a test subject with a composition as described in any of the above embodiments, b) detecting the captured plurality of cfDNA fragments, and c) applying a trained classifier to the captured plurality of DNA fragments to predict the efficacy of chemotherapy for diffuse large B-cell lymphoma for the test subject.

[0021] In some embodiments, the trained classifier predicts the efficacy of chemotherapy for diffuse large B-cell lymphoma for the test subject.

[0022] In some embodiments, the trained classifier is a mixed model classifier.

[0023] In some embodiments, the classifier is trained on a plurality of converted DNA sequences derived from a target genomic region selected from any of Lists 1-33. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the present disclosure, the following will briefly introduce the drawings needed to be used in some embodiments of the present disclosure. Obviously, the drawings in the following description are only some drawings of the embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art according to these drawings. In addition, the drawings in the following description can be regarded as schematic diagrams, and are not limited to the actual size, method, etc. of the products involved in the embodiments of the present disclosure.

[0025] FIG. 1 is a technical path diagram of obtaining a target genomic region according to some embodiments;

[0026] FIG. 2 is a structure diagram of a linker sequence according to some embodiments;

[0027] FIG. 3 is a design diagram of a clinical study on patients according to some embodiments;

[0028] FIG. 4 is a receiver operating characteristic curve diagram according to some embodiments;

[0029] FIG. 5 is a confusion matrix diagram according to some embodiments;

[0030] FIG. 6 is a sensitivity and specificity of clinical sample detection in a test set according to some embodiments;

[0031] FIG. 7 is a plot of the results of the detected and theoretical values of the methylation rates of different sites of a target genomic region according to some embodiments;

[0032] FIG. 8 is a plot of the correlation analysis of the linear regression of the detected and theoretical methylation rates of different sites of a target genomic region according to some embodiments. DETAILED DESCRIPTION

[0033] Unless otherwise defined, all technical and scientific terms used in the embodiments of the disclosure have the meanings commonly understood by one of ordinary skill in the art in the field of the disclosure. As used herein, the following terms have the meanings ascribed to them below.

[0034] As used herein, any reference to "one implementation" or "an implementation" means one specific embodiment of the described implementation, feature, structure or characteristic being described, is included in at least one implementation. The phrase "in some implementations" as used throughout this description does not necessarily refer to the same implementation, although it may. Rather, the phrase "in some implementations" is used herein to allow that in some implementations there are changes to the described implementations, so as to provide a framework for various possibilities of the described implementations.

[0035] As used herein, "comprises" or "comprising" or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).

[0036] Also, the use of "a" or "an" is employed to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the scope of the description. This description should be read to include one or at least one and the singular also includes the plural, unless it is otherwise evident from the context.

[0037] As used herein, ranges and amounts can be expressed as "about" a particular value or range. About also includes the exact amount. Thus, "about 5 micrograms" means "about 5 micrograms" and also "5 micrograms." Generally, the term "about" refers to an amount that is expected to be within experimental error. In some embodiments, "about" means the indicated number or value plus or minus 20%, 10%, or 5%. Furthermore, ranges as recited in embodiments of the present disclosure are to be understood to encompass all values and subranges within the stated ranges, inclusive of the recited endpoints. For example, a range of 1 to 50 is to be understood to include any number, combination of numbers, or sub-range from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50.

[0038] The term "methylation" as used herein means the process of adding a methyl group to a DNA molecule. For example, a hydrogen atom on the pyrimidine ring of a cytosine base can be converted to a methyl group, forming a 5-methylcytosine. The term also refers to the process of adding a hydroxymethyl group to a DNA molecule, for example by oxidation of a methyl group on the pyrimidine ring of a cytosine base. Methylation and hydroxymethylation tend to occur at dinucleotides of cytosine and guanine, referred to in embodiments of the present disclosure as "CpG sites."

[0039] The term "methylation" can also mean the methylation state of a CpG site. A CpG site with a 5-methylcytosine is methylated. A CpG site with a hydrogen atom on the pyrimidine ring of a cytosine base is unmethylated.

[0040] The term "methylation site" as used herein means a region of a DNA molecule to which a methyl group can be added. CpG sites are the most common methylation sites, but methylation sites are not limited to CpG sites. For example, DNA methylation can occur at cytosines in CHG and CHH, where H is adenine, cytosine, or thymine.

[0041] The term "CpG site" is used in embodiments of the present disclosure to mean a region of a DNA molecule in which, in a linear sequence of a few bases, a cytosine nucleotide is followed by a guanine in the 5' to 3' direction along the sequence. "CpG" is a shorthand for 5'-C-phospho-G-3', a cytosine and guanine separated by only one phosphate group. The cytosine in a CpG dinucleotide can be methylated to form a 5-methylcytosine.

[0042] The term "UpG" is a shorthand for 5'-U-phospho-G-3', which is a uracil and a guanine separated by only one phosphate group. UpG can be generated, for example, by bisulfite treatment, which converts unmethylated cytosines to uracils. Cytosines can be converted to uracils by other methods known in the art, such as chemical modification, synthesis, or enzymatic conversion.

[0043] The terms "hypomethylated" or "hypermethylated", as used herein, mean a methylation state of a DNA molecule containing a plurality (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) of CpG sites, wherein a high proportion (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%) of the CpG sites can be unmethylated or methylated.

[0044] The term "training sample", as used herein, means a sample used to train a classifier and / or select one or more genomic regions that provide diffuse large B-cell lymphoma therapeutic information in embodiments of the present disclosure. The training sample can comprise genomic DNA from or derived from one or more subjects having diffuse large B-cell lymphoma. The genomic DNA can be, but is not limited to, cfDNA fragments or chromosomal DNA. The genomic DNA can be sequenced and its methylation state can be evaluated. When the genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing an individual's genome, a training sample can mean genomic DNA or cfDNA fragments having the genomic sequence.

[0045] The term "detection sample", as used herein, means a sample from a subject whose health status has been or will be detected using the classifier and / or assay detection combination described herein. The detection sample can comprise genomic DNA or its derivative. The genomic DNA can be, but is not limited to, cfDNA fragments or chromosomal DNA.

[0046] The term "target genomic region", as used herein, means a region in the genome selected for analysis in a detection sample. The assay detection combination has probes designed to hybridize to nucleic acid fragments derived from the target genomic region or a fragment of the target genomic region. A nucleic acid fragment derived from the target genomic region means a nucleic acid fragment generated by degradation, cleavage, conversion, or other processing of DNA from the target genomic region.

[0047] Various target genomic regions are described according to their location on a chromosome. Chromosomal DNA is double stranded, so a target genomic region includes two DNA strands: one strand has the sequence provided in the list, and a second strand, which is the reverse complement of the sequence in the list. Probes can be designed to hybridize to one or both sequences. Alternatively, probes hybridize to the converted sequence.

[0048] A“converted cfDNA molecule” and“modified fragment obtained from processing of the cfDNA molecule” means a DNA molecule obtained by processing DNA or cfDNA molecules in a sample to distinguish between methylated and unmethylated nucleotides in the DNA or cfDNA molecules. For example, the sample is processed to convert unmethylated cytosines (“C”) to uracils (“U”), e.g., conversion of unmethylated cytosines to uracils is accomplished using an enzymatic reaction, e.g., using a cytidine deaminase (such as APOBEC). After processing, the converted DNA or cfDNA molecule includes additional uracils that were not present in the original cfDNA sample. Amplification of the DNA strand including uracils by DNA polymerase results in adenines being added to the newly generated complementary strand, rather than the normal guanine that is the complement of cytosine or methylcytosine.

[0049] The terms“cell-free nucleic acid,”“cell-free DNA,” or“cfDNA” and the like mean nucleic acid fragments that circulate within the body (e.g., the bloodstream) of a subject and that originate from cells of one or more subjects having diffuse large B-cell lymphoma. Additionally, cfDNA can come from other sources such as a fetus.

[0050] The term“fragment” as used herein can mean a fragment of a nucleic acid molecule. For example, in one embodiment, a fragment can mean a cfDNA molecule in blood or a blood sample, or a cfDNA molecule extracted from plasma or a plasma sample. Amplification products of cfDNA molecules can also be referred to as“fragments.” In another embodiment, the term“fragment” as described herein means a sequence read, or a set of sequence reads, that have been processed (e.g., in a machine learning-based classification) for subsequent analysis. For example, as is well known in the art, raw sequence reads can be aligned to a reference genome and end sequence reads that are mated are assembled into a longer fragment for subsequent analysis.

[0051] The term“subject” means a human individual.

[0052] The term "subject" means an individual whose DNA is analyzed. A subject can be a test subject whose DNA is evaluated using a targeted test panel as described herein to assess the efficacy of treatment of diffuse large B-cell lymphoma disease in that individual.

[0053] The term "sequence read" as used herein means a nucleotide sequence read from a sample. Sequence reads can be obtained via various methods provided herein or known in the art.

[0054] The term "sequencing depth" as used herein means a count of the number of times a given target nucleic acid in a sample is sequenced (e.g., a count of sequence reads at a given target region). Increasing sequencing depth can reduce the amount of nucleic acid needed to assess disease status.

[0055] A "probe set" of a test panel or bait panel or a "set of probes containing polynucleotides" of a test panel or bait panel generally means all probes used with a particular test panel or bait panel. For example, in some embodiments, a test panel or bait panel can include (1) a number of probes having the features specified herein (e.g., a number of probes for capturing cell-free DNA fragments corresponding to or derived from genomic regions set forth in one or more lists herein) and (2) additional probes that do not contain the features specified herein. The probe set of a test panel generally means all probes used with the test panel or bait panel, including probes that do not contain the specified features.

[0056] "Diffuse large B-cell lymphoma (DLBCL)" is an aggressive tumor derived from mature B cells and is the most common type of non-Hodgkin lymphoma.

[0057] Rituximab plus cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP) is effective in improving the prognosis of patients with DLBCL and is considered the standard first-line treatment with a 10-year overall survival (OS) of 43.5%. However, 20% of patients treated with R-CHOP have refractory disease, and approximately 20% of patients who initially benefit from R-CHOP report relapse. 30-50% of patients with primary or secondary resistance to R-CHOP have significantly poorer survival outcomes, with a median OS of approximately 6.3 months.

[0058] Embodiments of the present disclosure provide a test panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma. The test panel comprises probes or probe pairs. The test panel described in embodiments of the present disclosure can be referred to as a bait panel or a composition comprising bait oligonucleotides. The probes can be polynucleotide-containing probes specifically designed to target one or more nucleic acid molecules corresponding to or derived from genomic regions that are differentially methylated between samples of responders and non-responders to chemotherapy for diffuse large B-cell lymphoma.

[0059] To design a test panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma, an analysis system can collect information on the methylation status of CpG sites of nucleic acid fragments from samples corresponding to samples of responders and non-responders to chemotherapy for diffuse large B-cell lymphoma. These samples can be processed to determine the methylation status of CpG sites, or the information can be obtained from TCGA. The analysis system can be any general computing system having a computer processor and a computer-readable storage medium having instructions for executing the computer processor to perform any or all of the operations described in embodiments of the present disclosure.

[0060] The analysis system can select target genomic regions based on the methylation patterns of nucleic acid fragments. One approach considers the pairwise discriminability between pairs of outcomes for regions (or more specifically, CpG sites within regions). Another approach considers the discriminability of regions (or more specifically, CpG sites within regions) when each outcome is considered relative to the remaining outcomes. From the selected target genomic regions with high discriminability power, the analysis system can design probes to target fragments from the selected genomic regions.

[0061] The analysis system can generate assay panels of various sizes for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma, e.g., a small assay panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma includes probes for a number of genomic regions that provide the most information, a medium assay panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma includes probes from the small assay panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma, and additional probes for a second tier of informative genomic regions, and a large assay panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma includes probes from the small and medium assay panels for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma, and more probes for a third tier of informative genomic regions.

[0062] With data obtained from such assay panels for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma (e.g., methylation states of nucleic acids derived from the assay panels for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma), the analysis system can train classifiers in various classification techniques to predict the likelihood of the efficacy of chemotherapy for diffuse large B-cell lymphoma for a sample.

[0063] In some embodiments, the composition includes a plurality of different decoy oligonucleotides configured to collectively hybridize to DNA molecules derived from a plurality of target genomic regions; wherein each genomic region of the plurality of target genomic regions is differentially methylated in a responder to chemotherapy for diffuse large B-cell lymphoma compared to a non-responder to chemotherapy for diffuse large B-cell lymphoma.

[0064] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to DNA molecules derived from the target genomic regions of any one of Tables 1-33. Wherein Tables 1-33 are selected from Table 1 shown below.

[0065] Table 1

[0066] Wherein DMR is an abbreviation for Differentially Methylated Region, which refers to a differentially methylated region that is contained in a target genomic region.

[0067] In some examples, the assay panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma includes probes, wherein each of the probes is configured to hybridize to a converted cfDNA molecule corresponding to one or more target genomic regions of any one of Tables 1-33.

[0068] In some examples, the different bait oligonucleotides are configured to hybridize to DNA molecules that are derived from at least 20%, at least 25%, or at least 50% of the target genomic regions of any one of Lists 1-33.

[0069] For example, the target genomic regions can be selected from List 1, and a method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma can comprise the step of assessing the methylation status of sequence reads derived from the target genomic regions of List 1.

[0070] For example, the target genomic regions can be selected from List 3, and a method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma can comprise the step of assessing the methylation status of sequence reads derived from the target genomic regions of List 3.

[0071] For example, the target genomic regions can be selected from List 10, and a method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma can comprise the step of assessing the methylation status of sequence reads derived from the target genomic regions of List 10.

[0072] For example, the target genomic regions can be selected from List 22, and a method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma can comprise the step of assessing the methylation status of sequence reads derived from the target genomic regions of List 22.

[0073] For example, the target genomic regions can be selected from List 33, and a method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma can comprise the step of assessing the methylation status of sequence reads derived from the target genomic regions of List 33.

[0074] For example, the target genomic regions can be selected from 2, 5, 10, 15, 20, 28, or more of Lists 1-33.

[0075] Because the probes are configured to hybridize to converted DNA or cfDNA molecules that correspond to or are derived from one or more target genomic regions, the probes can have sequences that are different from the target genomic regions.

[0076] For example, a DNA containing an unmethylated CpG site will be converted to include UpG instead of CpG because the unmethylated cytosine is converted to uracil by the conversion reaction. As a result, the probe is configured to hybridize to a sequence including UpG instead of the normally occurring unmethylated CpG. Thus, the complementary site for the unmethylated site in the probe can include CpA instead of CpG, and some probes for a methylation site in which all methylation sites are unmethylated can not have a guanine (G) base.

[0077] A test panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma can be used to predict the efficacy of chemotherapy for diffuse large B-cell lymphoma. Illustratively, a test panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma is designed to enrich for a number of genomic regions of interest that are differentially methylated based on sequencing data generated from cfDNA of individuals with diffuse large B-cell lymphoma who are responsive to chemotherapy and individuals with diffuse large B-cell lymphoma who are non-responsive to chemotherapy.

[0078] In some embodiments, the different bait oligonucleotides are configured to hybridize to DNA molecules derived from at least 20%, at least 25%, or at least 50% of the genomic regions of interest in Tables 1-33.

[0079] In some embodiments, the different bait oligonucleotides are configured to hybridize to DNA molecules derived from at least 20% of the genomic regions of interest in Tables 1-33.

[0080] In some embodiments, the DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the genomic regions of interest in Tables 1-33.

[0081] In some examples, a test panel for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma includes probes that selectively hybridize and optionally enrich for cfDNA fragments that are differentially methylated in samples from individuals with diffuse large B-cell lymphoma who are responsive to chemotherapy and individuals with diffuse large B-cell lymphoma who are non-responsive to chemotherapy. Sequencing from the enriched fragments can provide information related to the efficacy of chemotherapy for diffuse large B-cell lymphoma.

[0082] For example, the combination of probes for the assay for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma is designed to include probes for genomic regions that are determined to be differentially hypermethylated or hypomethylated in samples of patients who respond to chemotherapy for diffuse large B-cell lymphoma and patients who do not respond to chemotherapy for diffuse large B-cell lymphoma to provide additional selectivity and specificity of detection. For example, the combination of probes for the assay for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma includes probes for hypomethylated segments. For example, the combination of probes for the assay for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma includes probes for hypermethylated segments. In some embodiments, the combination of probes for the assay for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma includes probes for a first set of hypermethylated segments and probes for a second set of hypomethylated segments.

[0083] In some examples, the genomic regions can be selected when the genomic regions produce aberrantly methylated DNA molecules in samples of patients who respond to chemotherapy for diffuse large B-cell lymphoma and patients who do not respond to chemotherapy for diffuse large B-cell lymphoma.

[0084] In some examples, the genomic regions can be further filtered based on their methylation patterns so that only genomic regions that are likely to provide information are selected. For example, based on CpG sites that are differentially methylated between samples of patients who respond to chemotherapy for diffuse large B-cell lymphoma and patients who do not respond to chemotherapy for diffuse large B-cell lymphoma, a calculation can be performed for each CpG or set of CpG sites for selection.

[0085] In some examples, once the probes hybridize and capture DNA fragments corresponding to or derived from the genomic regions of interest, the hybridized probe-DNA fragment intermediates are isolated and the target DNA is amplified and the methylation status of the target DNA is determined by sequencing or hybridization to magnetic beads. The sequence reads provide information associated with the prediction of the efficacy of chemotherapy for diffuse large B-cell lymphoma. For this purpose, the combination of probes for the assay for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma is designed to include probes that can capture fragments that collectively provide information associated with the prediction of the efficacy of chemotherapy for diffuse large B-cell lymphoma.

[0086] The selected genomic regions can be located at various different locations in the genome, including but not limited to exons, introns, intergenic regions, and other portions.

[0087] In some cases, primers can be used to specifically amplify (e.g., by PCR) a number of targets / biomarkers of interest to enrich a sample for a number of targets / biomarkers. For example, forward and reverse primers can be designed for a genomic region of interest and used to amplify fragments corresponding to or derived from the desired genomic region. Thus, while embodiments of the disclosure focus on assay detection panels for diffuse large B-cell lymphoma chemotherapy response prediction and bait sets for hybrid capture, embodiments of the disclosure also encompass other methods for enrichment of cell-free DNA. Thus, one of skill in the art, with the benefit of embodiments of the disclosure, will recognize that methods similar to those of hybrid capture in embodiments of the disclosure can be accomplished instead by substituting hybrid capture with other enrichment strategies, such as PCR amplification of cell-free DNA fragments corresponding to genomic regions of interest.

[0088] In some embodiments, probes are used for enrichment (e.g., non-targeted enrichment), such as reduced representation bisulfite sequencing, methylation restriction enzyme sequencing, methylation DNA immunoprecipitation sequencing, methyl CpG binding domain protein sequencing, methyl DNA capture sequencing, or droplet PCR.

[0089] The assay detection panels for diffuse large B-cell lymphoma chemotherapy response prediction provided herein are detection panels comprising sets of hybridization probes (also referred to as "probes" in embodiments of the disclosure) designed to enrich nucleic acid fragments of interest for use in assays.

[0090] In some embodiments, probes are designed to hybridize to and enrich DNA or cfDNA molecules from samples that have been treated to convert unmethylated cytosine (C) to uracil (U). Probes can be designed to bind or hybridize to the target (complementary) strand of DNA or RNA. The target strand can be the "+" strand (e.g., the strand that is transcribed into mRNA and then translated into protein) or the complementary "-" strand.

[0091] Embodiments of the disclosure for detecting nucleic acids and determining methylation status also encompass other methods for determining the methylation status of nucleic acid sequences, as to the manner of sequencing.

[0092] In some embodiments, nucleic acid samples (DNA or RNA) are extracted from a subject. In embodiments of the disclosure, DNA and RNA can be used interchangeably unless otherwise indicated. That is, embodiments described in embodiments of the disclosure can be applicable to DNA and RNA types of nucleic acid sequences. However, examples described in embodiments of the disclosure focus on DNA for the purposes of simplicity and explanation.

[0093] In some embodiments, the sample can be any combination of the human genome, including the whole genome. The sample can include blood, plasma, serum, urine, stool, saliva, other types of bodily fluids, or any combination thereof. In some embodiments, the method for extracting a blood sample (e.g., syringe or finger prick) can be less invasive than the procedure for obtaining a tissue biopsy, which can require surgery. The extracted sample can include cfDNA and / or ctDNA. If a subject is a diffuse large B-cell lymphoma chemotherapy responder or a diffuse large B-cell lymphoma chemotherapy non-responder, the cfDNA and / or ctDNA in the extracted sample can be sufficient to predict the chemotherapy response of diffuse large B-cell lymphoma.

[0094] In some embodiments, the cfDNA fragments are treated to convert unmethylated cytosines to uracils. The conversion of unmethylated cytosines to uracils is achieved using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for converting unmethylated cytosines to uracils, such as APOBEC-Seq (NEBiolabs).

[0095] In some embodiments, the prediction of the chemotherapy response of diffuse large B-cell lymphoma requires the preparation of sequencing libraries, the details of which are described in subsequent sections and are not described here.

[0096] In some embodiments, target DNA sequences can be enriched from the sequencing libraries described above and used when the assay for predicting the chemotherapy response of diffuse large B-cell lymphoma is performed on a combination of samples. In enrichment, hybridization probes (also referred to herein as "probes") are used to target nucleic acid fragments that provide information on diffuse large B-cell lymphoma chemotherapy responders and non-responders.

[0097] In some embodiments, the hybridized nucleic acid fragments are captured and can also be amplified using PCR. Target sequences can be enriched to obtain enriched sequences, which can then be sequenced. In general, any method known in the art can be used to isolate and enrich target nucleic acids that have hybridized to the probes. For example, as is well known in the art, a biotin moiety can be added to the 5' end of the probes (i.e., biotinylated) using a streptavidin-labeled surface (e.g., streptavidin-labeled magnetic beads) to facilitate the isolation of target nucleic acids that have hybridized to the probes.

[0098] In some embodiments, a plurality of sequence reads are generated from the plurality of enriched DNA sequences, e.g., the plurality of enriched sequences. Sequence data can be obtained from the plurality of enriched DNA sequences by methods known in the art. For example, the methods can include next generation sequencing (NGS) techniques, including sequencing by synthesis (Illumina), pyrosequencing, ion semiconductor sequencing (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing.

[0099] In some embodiments, a plurality of methylation state vectors are generated from the plurality of sequence reads. The sequence reads are aligned to a reference genome, which assists in providing what location in a human genome the fragments of cfDNA originated from.

[0100] The following describes aspects related to the generation of data structures.

[0101] To create the control group data structure, the analysis system obtains information about the methylation states of a plurality of CpG sites on a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of DNA molecules or fragments of a diffuse large B-cell lymphoma chemotherapy responder. A methylation state vector is generated for each DNA molecule or fragment.

[0102] With the methylation state vector for each fragment, the analysis system subdivides the methylation state vector into strings of CpG sites. The analysis system subdivides the methylation state vector such that the resulting strings are all less than a given length. For example, a methylation state vector of length 11 can be subdivided into strings, less than or equal to a length of 3 would result in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1.

[0103] The analysis system records the plurality of strings by counting, for each possible CpG site in the vector and methylation state, the number of strings in the control group that have that particular CpG site as the first CpG site in the string and that have that possibility of methylation state.

[0104] The following describes aspects related to the validation of data structures.

[0105] In some embodiments, once a data structure has been created, the analysis system can seek to validate the data structure and / or any downstream models that use the data structure.

[0106] The first type of validation checks for consistency in the data structure of the control group. For example, if there are any outliers objects, samples, and / or fragments in the control group, the analysis system can perform various calculations to determine whether to remove any fragments from one of these categories to not affect the purity of the control group.

[0107] The second type of validation checks the probability model used to calculate p-values from the counts of the data structure itself (i.e., from the control group). Once the analysis system produces a p-value for a number of methylation state vectors in the validation group, the analysis system constructs a cumulative density function (CDF) from the p-values. The analysis system can perform various calculations on the CDF to validate the data structure of the control group.

[0108] The third type of validation uses a set of validation samples of diffuse large B-cell lymphoma chemotherapy- effective that are separate from the validation samples used to construct the data structure. The third type of validation tests whether the data structure was constructed properly and the model works. The third type of validation can quantify how well the control group generalizes the distribution of diffuse large B-cell lymphoma chemotherapy-effective samples.

[0109] The fourth type of validation tests with samples from a diffuse large B-cell lymphoma chemotherapy- ineffective validation group. The analysis system calculates p-values for the diffuse large B-cell lymphoma chemotherapy- ineffective validation group and constructs a CDF. For the diffuse large B-cell lymphoma chemotherapy- ineffective validation group, the analysis system expects to see, for at least some samples, the opposite of what was expected for the control group and the validation group in the second type of validation and the third type of validation. If the fourth type of validation fails, it indicates that the model does not properly identify the anomalies that the model was designed to identify.

[0110] In the process of validating the data structure, the analysis system performs the fourth type of validation test as described above that applies a validation group that has a composition of objects, samples, and / or fragments that are assumed to be similar to the control group. For example, if the analysis system selects diffuse large B-cell lymphoma chemotherapy-effective objects as the control group, the analysis system also uses diffuse large B-cell lymphoma chemotherapy-effective objects in the validation group.

[0111] The analysis system takes the methylation state vectors of the validation set and the analysis system performs a p-value calculation for each methylation state vector from the validation set. For each possibility of a methylation state vector, the analysis system computes a probability from the data structure of the control set. Once the probabilities for the number of possibilities of the methylation state vector are computed, the analysis system computes a p-value score for the methylation state vector based on the computed probabilities. The p-value score represents the expectation of finding that particular methylation state vector and other possible methylation state vectors with even lower probabilities in the control set. Thus, a low p-value score generally corresponds to a methylation state vector that is less expected relative to other methylation state vectors found in the control set, while a high p-value score generally corresponds to a methylation state vector that is more expected relative to other methylation state vectors found in the control set. Once the analysis system generates p-value scores for the number of methylation state vectors in the validation set, the analysis system constructs a cumulative density function (CDF) from the p-value scores from the validation set. The analysis system verifies the consistency of the CDF in a fourth type of validation test as described above.

[0112] The following introduces the relevant content of aberrant methylation fragments.

[0113] In samples of diffuse large B-cell lymphoma that respond to chemotherapy and samples of diffuse large B-cell lymphoma that do not respond to chemotherapy, aberrant methylation fragments having aberrant methylation patterns are selected as target genomic regions. The analysis system generates methylation state vectors from the cfDNA fragments of the samples. The analysis system processes each methylation state vector as follows.

[0114] For a given methylation state vector, the analysis system enumerates all possibilities of the methylation state vector having the same starting CpG site and the same length (i.e., the set of CpG sites) in the methylation state vector. Thus each methylation state possibility is either methylated or unmethylated, there are only two possible states at each CpG site, and thus the count of unique possibilities of a methylation state vector depends on a power of two, such that a methylation state vector of length n will be associated with 2npossibilities of the methylation state vector.

[0115] The analysis system computes the probability of observing each possibility of the methylation state vector for the identified starting CpG site / methylation state vector length by evaluating the control set data structure of diffuse large B-cell lymphoma that responds to chemotherapy. The probability of observing a given possibility is computed using a Markov chain probability to model the joint probability computation, different from the method of computation of Markov chain probabilities used to determine the probability of observing each possibility of the methylation state vector.

[0116] The analysis system computes a p-value score for a methylation state vector using the probabilities of each of a number of possible methylation states that could be computed. For example, this includes identifying the probabilities of possible methylation states that could be computed that are consistent with the methylation state vector under consideration. In particular, this is the probability of having the same set of CpG sites as the methylation state vector, or similarly having the same starting CpG site and length. The analysis system sums the probabilities of a number of computed methylation states to produce the p-value score. The number of computed methylation states is a number of possible computed methylation states that can have fewer or equal probabilities than the identified probabilities.

[0117] This p-value represents the probability of observing the methylation state vector of the fragment, or other even less likely methylation state vectors, in the control group. Thus, a low p-value score generally corresponds to a methylation state vector that is rare in individuals who respond to chemotherapy for diffuse large B-cell lymphoma, and causes the fragment to be labeled as aberrantly methylated relative to the control group. A high p-value score generally correlates with a methylation state vector that is expected to be present in a relative sense in individuals who respond to chemotherapy for diffuse large B-cell lymphoma. For example, if the control group is individuals who respond to chemotherapy for diffuse large B-cell lymphoma, a low p-value indicates that the fragment is aberrantly methylated relative to the group of individuals who respond to chemotherapy for diffuse large B-cell lymphoma, and thus can indicate a prediction of non-response to chemotherapy for diffuse large B-cell lymphoma in the test subject.

[0118] The analysis system computes a p-value score for each of a number of methylation state vectors, each of which represents a cfDNA fragment in the test sample. To identify which of the number of fragments are aberrantly methylated, the analysis system can filter the set of methylation state vectors based on the p-value scores of the number of methylation state vectors. For example, filtering is performed by comparing the p-value scores to a threshold and retaining only those fragments that are below the threshold.

[0119] The following describes the p-value score computation in more detail.

[0120] To compute the p-value score for a given test methylation state vector, the analysis system takes the test methylation state vector, and enumerates the number of possible methylation state vectors.

[0121] The analysis system computes the probabilities of the enumerated number of possible methylation state vectors. Because methylation is conditionally dependent on the methylation state of nearby CpG sites, one way to compute the probability of observing a given methylation state vector is to use a Markov chain model. The Markov chain model can be used to make the computation of each possible conditional probability more efficient.

[0122] To compute each Markov chain modeled probability of a possible methylation vector, the analysis system accesses the data structure of the control group, in particular the counts of the various strings of CpG sites and states.

[0123] The computation can additionally perform smoothing of the counts by applying a prior distribution. For example, the prior distribution is a uniform prior as in Laplace smoothing. As an example, algorithmic techniques, such as the Nye smoothing method, are used.

[0124] Once the computed probabilities are completed, the analysis system computes a p-value score that sums the probabilities that are less than or equal to the probability of the methylation state vector being detected.

[0125] In some embodiments, the computational burden of computing probabilities and / or p-value scores can be further reduced by caching at least some of the computations. For example, the analysis system can cache the computation of the probabilities of methylation state vectors (or windows thereof) in temporary or permanent memory. Caching the probabilities allows efficient computation of p-score values without having to recompute the underlying probabilities if other fragments have the same CpG sites. Ultimately, the analysis system can compute a p-value score for each of the probabilities of methylation state vectors associated with a set of CpG sites from a vector (or window thereof). The analysis system can cache the p-value scores for use in determining p-value scores for other fragments that include the same CpG sites. In general, the p-value scores of methylation state vectors with the same CpG sites can be used to determine p-value scores for different ones of the probabilities from the same set of CpG sites.

[0126] The following introduces the relevant content of a sliding window.

[0127] In some embodiments, the analysis system uses a sliding window to determine probabilities of methylation state vectors and compute p-values. Rather than enumerating probabilities and computing p-values for the entire methylation state vector, the analysis system enumerates probabilities and computes p-values for only a window of consecutive CpG sites, where the window is shorter in length (in CpG sites) than at least some fragments (otherwise, the window would be useless). The window length can be static, user-determined, dynamic, or otherwise selected.

[0128] In computing p-values for a methylation state vector that is larger than the window, the window identifies a consecutive set of CpG sites from the vector that is in the window, starting with the first CpG site in the vector. The analysis system computes a p-value score for the window including the first CpG site. The analysis system then "slides" the window to the second CpG site in the vector and computes another p-value score for the second window. Thus, for a window of size / and a methylation vector length m, each methylation state vector will yield m p-value scores. After completing the p-value computation for each portion of the vector, the lowest p-value score from all of the sliding windows is taken as the overall p-value score for the methylation state vector.

[0129] Using a sliding window reduces the number of possible methylation state vectors that need to be enumerated and the corresponding probability calculations that need to be performed, if not otherwise. The number of possible methylation state vectors increases exponentially with the size of the methylation state vector, as a power of two. Rather than computing 2 54 (about 1.8 x 10 16 ) possible probabilities to produce a single p-value score, the analysis system can use a window of, for example, size 5 on a fragment, resulting in 50 p-value calculations being performed for each of the 50 windows of the methylation state vector, rather than computing 2 5 (32) possible methylation state vectors for each of the 50 calculations, resulting in a total of 50 x 2 5 (1.6 x 10 3 ) probability calculations. This results in a lack of meaningful hits for accurate identification of aberrant fragments, a large reduction in the number of calculations that need to be performed. This additional step can also be applied when validating a number of methylation state vectors of the validation set against the control set.

[0130] The analysis system identifies from the filtered set of aberrant methylation fragments a number of DNA fragments that are indicative of the efficacy of chemotherapy for diffuse large B-cell lymphoma.

[0131] The following describes the identification of hypomethylated and hypermethylated fragments.

[0132] According to a first approach, the analysis system can identify from the filtered set of aberrant methylation fragments a number of DNA fragments that are considered hypomethylated or hypermethylated as fragments indicative of the efficacy of chemotherapy for diffuse large B-cell lymphoma. The number of hypomethylated or hypermethylated fragments can be defined as fragments of a particular length (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) of CpG sites that have a high percentage of methylated CpG sites (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%) or a high percentage of unmethylated CpG sites (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%).

[0133] The following describes the selection of genomic regions that are indicative of the efficacy of chemotherapy for diffuse large B-cell lymphoma.

[0134] The analysis system identifies a number of genomic regions that are indicative of the efficacy of chemotherapy for diffuse large B-cell lymphoma. To identify these informative regions, the analysis system computes for each genomic region, or more specifically for each CpG site, an information gain that describes the ability to distinguish between various outcomes.

[0135] Methods for identifying a number of genomic regions that can be used to distinguish between diffuse large B-cell lymphoma chemotherapy responders and non-responders, using a trained classification model that can be applied to a set of aberrantly methylated DNA molecules or fragments corresponding to or derived from a cohort of diffuse large B-cell lymphoma chemotherapy responders and non-responders. The trained classification model can be trained to identify any condition of interest that can be identified from a number of methylation state vectors.

[0136] In some embodiments, the trained classification model is a binary classifier that is trained based on a number of cfDNA fragments or a number of genomic sequences obtained from a population of subjects that are diffuse large B-cell lymphoma chemotherapy non-responders and a population of subjects that are diffuse large B-cell lymphoma chemotherapy responders. The binary classifier is then used to classify the probability of a test subject being a diffuse large B-cell lymphoma chemotherapy responder and a diffuse large B-cell lymphoma chemotherapy non-responder based on a number of aberrant methylation state vectors.

[0137] In some embodiments, the classifier can be trained using a number of sequence reads obtained from a sample of cells of diffuse large B-cell lymphoma chemotherapy responders. The sample is from a population of subjects known to be diffuse large B-cell lymphoma chemotherapy responders. The ability of each genomic region to distinguish between diffuse large B-cell lymphoma chemotherapy responders and non-responders in the classification model is used to rank the genomic regions from most informative to least informative. The analysis system can identify a number of genomic regions from the ranking based on the information gain in the classification between diffuse large B-cell lymphoma chemotherapy responders and non-responders.

[0138] The following describes the calculation of information gain from hypomethylated and hypermethylated fragments indicative of diffuse large B-cell lymphoma chemotherapy response.

[0139] According to some embodiments, using a number of fragments indicative of diffuse large B-cell lymphoma chemotherapy response, the analysis system can train a classifier according to a procedure. The procedure accesses two training sets of samples: one diffuse large B-cell lymphoma chemotherapy responder set and one diffuse large B-cell lymphoma chemotherapy non-responder set, obtains a diffuse large B-cell lymphoma chemotherapy responder set of methylation state vectors comprising aberrantly methylated fragments, and a diffuse large B-cell lymphoma chemotherapy non-responder set of methylation state vectors.

[0140] The analysis system determines, for each methylation state vector, whether the methylation state vector is indicative of diffuse large B-cell lymphoma chemotherapy response. Here, a number of fragments indicative of diffuse large B-cell lymphoma chemotherapy response can be defined as hypomethylated or hypermethylated fragments if a number of CpG sites have a particular state. For example, a number of cfDNA fragments can be identified as hypomethylated or hypermethylated if the number of cfDNA fragments overlap at least 4 CpG sites and at least 80%, 90%, or 100% of the CpG sites of the number of cfDNA fragments are methylated or at least 80%, 90%, or 100% of the CpG sites of the number of cfDNA fragments are unmethylated, respectively.

[0141] In other embodiments, the analysis system considers a number of portions of the methylation state vector and determines whether the portion is hypomethylated or hypermethylated and can discriminate between hypomethylated and hypermethylated portions.

[0142] In some embodiments, the procedure produces a hypomethylation score (P 低 ) and a hypermethylation score (P 过 ) for each CpG site in the genome. To produce the two scores for a given CpG site, the classifier takes four counts at the CpG site: (1) the number of (methylation state) vectors from the diffuse large B-cell lymphoma chemotherapy non-responder group that overlap the CpG site and are labeled as hypomethylated; (2) the number of vectors from the diffuse large B-cell lymphoma chemotherapy non-responder group that overlap the CpG site and are labeled as hypermethylated; (3) the number of vectors from the diffuse large B-cell lymphoma chemotherapy responder group that overlap the CpG site and are labeled as hypomethylated; and (4) the number of vectors from the diffuse large B-cell lymphoma chemotherapy responder group that overlap the CpG site and are labeled as hypermethylated. In addition, the procedure can normalize the counts for each group to account for differences in group size between the diffuse large B-cell lymphoma chemotherapy responder group and the diffuse large B-cell lymphoma chemotherapy non-responder group.

[0143] In alternative embodiments in which fragments indicative of diffuse large B-cell lymphoma chemotherapy response are used, the scores can be more broadly defined as the counts of fragments indicative of diffuse large B-cell lymphoma chemotherapy response at each genomic region and / or CpG site.

[0144] In some embodiments, to generate a low methylation fraction at a given CpG site, the procedure takes the above-described (1) divided by the sum of the above-described (1) and the above-described (3). Similarly, the hypermethylation fraction is calculated by taking the above-described (2) divided by the sum of the above-described (2) and the above-described (4). The presence of low methylation or hypermethylation in a number of fragments from the diffuse large B-cell lymphoma chemotherapy non-responding group is determined, and the low methylation fraction and the hypermethylation fraction are correlated with the probability of response to the diffuse large B-cell lymphoma chemotherapy.

[0145] The analysis system generates a total low methylation fraction and a total hypermethylation fraction for each aberrant methylation status vector. The total hypermethylation and low methylation fractions are determined based on the number of hypermethylation and low methylation fractions in the methylation status vector.

[0146] In some embodiments, the total hypermethylation and low methylation fractions are recorded as the maximum hypermethylation and low methylation fractions for the number of sites in each status vector, and the analysis system ranks each subject on two rankings according to all of the subject's methylation status vectors, one ranking by the total low methylation fraction of the number of methylation status vectors and one ranking by the total hypermethylation fraction of the number of methylation status vectors. The procedure selects a number of total low methylation fractions from the low methylation ranking and a number of total hypermethylation fractions from the hypermethylation ranking. From the selected number of fractions, the classifier generates a single feature vector for each subject.

[0147] In some embodiments, the number of fractions selected from the two rankings are selected in a fixed order, which is the same for the feature vector of each subject in each of the number of training groups. For example, the classifier selects the first, second, fourth, and eighth total hypermethylation fractions from each ranking, and the same for each total low methylation fraction, and writes these fractions in the feature vector for the subject.

[0148] The analysis system trains a binary classifier to distinguish between the feature vectors of the diffuse large B-cell lymphoma chemotherapy responding and the diffuse large B-cell lymphoma chemotherapy non-responding training groups. In general, any of a number of classification techniques can be used.

[0149] In some embodiments, the classifier is a nonlinear classifier. In a particular embodiment, the classifier is a nonlinear classifier that applies L2-regularized function kernel logistic regression with a Gaussian radial basis function kernel.

[0150] In some embodiments, the number of diffuse large B-cell lymphoma chemotherapy responding samples (n 其它 ) and the number of diffuse large B-cell lymphoma chemotherapy non-responding samples (n 阳) are counted. Next, the probability that a sample is chemotherapy-ineffective for diffuse large B-cell lymphoma is estimated by a score ("S") that is positively correlated with the number (n 阳 ) of chemotherapy-ineffective samples for diffuse large B-cell lymphoma and negatively correlated with the number (n 其它 ) of chemotherapy-effective samples for diffuse large B-cell lymphoma. The score can be calculated using the equation: (n 阳 + 1) / (n 阳 + n 其它 + 2) or (n 阳 ) / (n 阳 + n 其它 ).

[0151] The analysis system calculates the information gain for each chemotherapy-ineffective sample for diffuse large B-cell lymphoma and for each genomic region or CpG site to determine whether the genomic region or CpG site is indicative of chemotherapy efficacy for diffuse large B-cell lymphoma. The information gain is calculated for a number of training samples with a given chemotherapy-ineffective sample for diffuse large B-cell lymphoma compared to all other samples. For example, a random variable "abnormal fragment" ("AF") is used.

[0152] In some embodiments, AF, as determined by the feature vector above, is a binary variable indicating whether an abnormal fragment overlaps with a given CpG site in a given sample.

[0153] The following describes relevant content for methods of using assay detection panels for predicting chemotherapy efficacy for diffuse large B-cell lymphoma.

[0154] The method can include the steps of processing a number of DNA molecules or a number of fragments to convert unmethylated cytosines to uracils, applying an assay detection panel for predicting chemotherapy efficacy for diffuse large B-cell lymphoma to the converted DNA molecules or fragments, enriching a subset of the converted DNA molecules or fragments that hybridize to a number of probes in the detection panel, and sequencing the enriched cfDNA fragments, detecting the nucleic acid sequences and determining the methylation status of the nucleic acid sequences.

[0155] In some embodiments, a number of sequence reads can be compared to a reference genome (e.g., a human reference genome) to allow identification of methylation status at a number of CpG sites in the DNA molecules or fragments, thereby providing information about predicting chemotherapy efficacy for diffuse large B-cell lymphoma.

[0156] The following describes relevant content for analysis of sequence reads.

[0157] In some embodiments, the sequence reads can be aligned to a reference genome using methods known in the art to determine alignment position information. The alignment position information can indicate a start position and an end position in the reference genome corresponding to a starting nucleotide base and an ending nucleotide base of a given sequence read. The alignment position information can include a sequence read length, which can be determined from the start position and the end position. A region in the reference genome can be associated with a gene or a segment of a gene.

[0158] In various embodiments, a sequence read includes a pair of reads denoted as R1 and R2. For example, a first read R1 can be sequenced from a first end of a nucleic acid fragment, while a second read R2 can be sequenced from a second end of the nucleic acid fragment. Accordingly, a number of nucleotide bases of the first read R1 and the second read R2 can be aligned to nucleotide bases of a reference genome consistently (e.g., in opposite directions). Alignment position information derived from the pair of reads R1 and R2 can include a start position in the reference genome corresponding to an end of the first read (e.g., R1) and an end position in the reference genome corresponding to an end of the second read (e.g., R2). In other words, the start position and the end position in the reference genome represent possible positions of the nucleic acid fragment in the reference genome. An output file in SAM format or BAM format can be generated and output for further analysis.

[0159] From the sequence reads, the position and the methylation state of each CpG site can be determined based on alignment to a reference genome. Further, a methylation state vector for each fragment can be generated, which specifies the position of the fragment in the reference genome (e.g., by the first CpG site in each fragment or other similar metrics), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment, which is either methylated (e.g., denoted as M), unmethylated (e.g., denoted as U), or intermediate (e.g., denoted as I). The methylation state vectors can be stored in temporary or permanent computer memory for later use and processing.

[0160] The following introduces the related content of predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma.

[0161] The sequence reads obtained by the methods provided by embodiments of the disclosure can be further processed by automated algorithms. For example, an analysis system is used to receive sequence data from a sequencer and perform various aspects of the processing as described in embodiments of the disclosure. The analysis system can be one of a personal computer, a desktop computer, a laptop computer, a notebook computer, a tablet personal computer, a mobile device. The computing device can be communicatively coupled to the sequencer by wireless, wired, or a combination of wireless and wired communication technologies. Generally, the computing device is configured to have a processor and a memory storing computer instructions. When executed by the processor, the computer instructions direct the processor to perform steps as in embodiments of the disclosure. Generally, the amount of genetic data and data derived from the genetic data is large enough and the computational power required is large enough that it is not possible to perform purely on paper or by human mind.

[0162] Clinical interpretation of the methylation status of the target genomic regions is a process that includes classifying the clinical effect of each of the methylation status or combination of methylation status and reporting the results in a way that is meaningful to medical professionals. The clinical interpretation can be based on comparison of the sequence reads to a database of subjects with diffuse large B-cell lymphoma chemotherapy effective and diffuse large B-cell lymphoma chemotherapy ineffective and / or based on the number and type of cfDNA fragments identified from the sample that have a specific methylation pattern that is indicative of diffuse large B-cell lymphoma chemotherapy ineffective.

[0163] In some embodiments, the target genomic regions are ranked based on their likelihood of being differentially methylated in diffuse large B-cell lymphoma chemotherapy ineffective samples and the ranking is used in the interpretation process. The ranking can include the strength of evidence for clinical effect. Various methods of clinical analysis and genomic data interpretation can be used for analysis of the sequence reads. In some other embodiments, the clinical interpretation of the methylation status of such differentially methylated regions can be based on a machine learning approach that interprets the current sample based on a classification or regression method that uses methylation status of such differentially methylated regions from samples with known diffuse large B-cell lymphoma chemotherapy effective and diffuse large B-cell lymphoma chemotherapy ineffective to train.

[0164] The clinical significance information can include diffuse large B-cell lymphoma chemotherapy effective or ineffective.

[0165] The following introduces the classifier for predicting the efficacy of diffuse large B-cell lymphoma chemotherapy.

[0166] To train a classifier for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma, the analysis system obtains a number of training samples. The analysis system determines a feature vector for each training sample based on a set of hypomethylated and hypermethylated fragments indicative of chemotherapy ineffectiveness for diffuse large B-cell lymphoma. The analysis system computes an anomaly score for each CpG site in a number of target genomic regions.

[0167] In some embodiments, the analysis system defines the anomaly score for the feature vector as a binary score based on whether there is a hypomethylated or hypermethylated fragment from the CpG site. Once all the anomaly scores are determined for the training samples, the analysis system determines the feature vector as a vector of a number of elements, each element including one of the anomaly scores associated with one of the CpG sites. The analysis system can normalize the anomaly scores of the feature vector based on the coverage of the sample, i.e., the median or average sequencing depth of all the CpG sites.

[0168] With the feature vectors of the training samples, the analysis system can train a classifier for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma.

[0169] In some embodiments, the analysis system trains a binary classifier for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma based on the feature vectors of the training samples, distinguishing between chemotherapy effectiveness and chemotherapy ineffectiveness for diffuse large B-cell lymphoma. In this embodiment, the classifier outputs a prediction score indicating the likelihood of chemotherapy effectiveness or ineffectiveness for diffuse large B-cell lymphoma.

[0170] The analysis system trains the classifier for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma by inputting groups of training samples with their feature vectors to the classifier and adjusting the classification parameters so that the function of the classifier accurately relates the training feature vectors to their corresponding labels. The analysis system can divide the training samples into groups of one or more training samples for iterative batch training of the classifier for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma.

[0171] After inputting all the groups of training samples, including their training feature vectors, and adjusting the classification parameters, the classifier for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma is sufficiently trained to label test samples according to their feature vectors within some error range.

[0172] The analysis system can train the diffuse large B-cell lymphoma chemotherapy response prediction classifier according to any of several methods. For example, the binary diffuse large B-cell lymphoma chemotherapy response prediction classifier can be an L2-regularized function kernel logistic regression classifier trained using a log loss function. Or, the diffuse large B-cell lymphoma chemotherapy response prediction classifier can be a multinomial logistic regression. In applications, both types of diffuse large B-cell lymphoma chemotherapy response prediction classifiers can be trained using other techniques. These techniques are numerous, including the application of kernel methods, machine learning algorithms such as multilayer neural networks, and the like.

[0173] In deployment, the analysis system obtains a test sample from a subject of unknown diffuse large B-cell lymphoma chemotherapy response. The analysis system processes the test sample to obtain a set of hypomethylated and hypermethylated fragments indicative of diffuse large B-cell lymphoma chemotherapy response. The analysis system defines a test feature vector in a similar procedure as described for the training samples. The analysis system then inputs the test feature vector into the trained diffuse large B-cell lymphoma chemotherapy response prediction classifier to obtain a prediction of diffuse large B-cell lymphoma chemotherapy response, including: diffuse large B-cell lymphoma chemotherapy effective or diffuse large B-cell lymphoma chemotherapy ineffective.

[0174] Embodiments of the present disclosure also provide a method for enriching converted cfDNA fragments that can provide information of diffuse large B-cell lymphoma chemotherapy response, the method comprising the steps of: contacting the composition of any of the above embodiments with DNA derived from a test subject, and enriching the sample of cfDNA corresponding to a plurality of genomic regions associated with diffuse large B-cell lymphoma chemotherapy response by hybridization capture.

[0175] Embodiments of the present disclosure also provide a method for obtaining sequence information that can provide information of diffuse large B-cell lymphoma chemotherapy response, the method comprising the steps of: a) enriching converted DNA from a test subject by contacting the converted DNA with the composition of any of the above embodiments, and b) sequencing the enriched converted DNA.

[0176] Embodiments of the present disclosure also provide a method for diffuse large B-cell lymphoma chemotherapy response prediction, comprising the steps of: a) capturing a plurality of cfDNA fragments from a test subject with the composition of any of the above embodiments, b) detecting the captured plurality of cfDNA fragments, and c) applying a trained classifier to the captured plurality of DNA fragments to predict the diffuse large B-cell lymphoma chemotherapy response of the test subject.

[0177] In some embodiments, the trained classifier predicts the diffuse large B-cell lymphoma chemotherapy response of a test subject.

[0178] In some embodiments, the trained classifier is a hybrid model classifier.

[0179] In some embodiments, the classifier is trained on a plurality of transformed DNA sequences derived from a target genomic region selected from any one of Lists 1-33.

[0180] In some embodiments, as shown in Figure 1, a biomarker for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma is provided, and the specific scheme adopts methylation targeted capture technology. The general process includes: (1) sample collection, 1 mL-4 mL of plasma is collected before treatment of the patient; cfDNA extraction; cfDNA end repair; (2) linker ligation reaction, the cytosine (C base) contained in the linker sequence is all methylated modified, and the UMI sequence is contained; (3) oxidation reaction, the methylated cytosine in double-stranded DNA is changed to carboxyl cytosine by the action of TET2 enzyme and the like; (4) deamination reaction, the unmethylated modified cytosine (C) in the product of the previous step is deaminated to form uracil (U) under the action of APOBEC enzyme; (5) amplification enrichment, forming a methylation library; (6) denaturation, probe hybridization reaction; (7) avidin magnetic beads bind to the hybridization product; (8) magnetic bead-bound product elution; (9) target region amplification enrichment, forming the final methylation library; (10) Novaseq 6000 sequencing; (11) sequencing data analysis, data filtering, alignment, and deduplication; (12) differential methylation site analysis; (13) methylation model construction.

[0181] Based on the above, the present disclosure provides the following embodiments.

[0182] Embodiment 1:

[0183] This embodiment tests the accuracy of the standard by combining library construction with targeted capture.

[0184] (1) Standard preparation: Custom cfDNA standard from Jingliang gene, specifically including fully methylated standard (C bases in the sequence are all methylated) and fully unmethylated standard (C bases in the sequence are not methylated), mix the standards in a certain proportion to form 1 μg of 0%, 5%, 25%, 50%, 100% methylation rate standards, then take 0%, 5%, 25%, 50%, 100% methylation rate standards 1 ng, 5 ng, 10 ng, 20 ng, 50 ng into 0.2 mL PCR tubes, respectively, and label them as samples 0%-1 ng, 0%-5 ng, 0%-10 ng, 0%-20 ng, 0%-50 ng, 5%-1 ng, 5%-5 ng, 5%-10 ng, 5%-20 ng, 5%-50 ng, 25%-1 ng, 25%-5 ng, 25%-10 ng, 25%-20 ng, 25%-50 ng, 50%-1 ng, 50%-5 ng, 50%-10 ng, 50%-20 ng, 50%-50 ng, 100%-1 ng, 100%-5 ng, 100%-10 ng, 100%-20 ng, 100%-50 ng, the sample is less than 40 μL, and enzyme-free water is used to make up.

[0185] (2) cfDNA end repair: After the reagents in Table 2 are thawed, mix well and centrifuge for a short time, then add to each 0.2 mL of PCR in step (1).

[0186] Table 2

[0187] Wherein, T4 PNK refers to T4 polynucleotide kinase, ATP is the abbreviation of Adenosine triphosphate (adenosine triphosphate), PEG 4000 refers to polyethylene glycol with a molecular weight of 4000, SSB is the abbreviation of single strand binding protein, which is called single strand binding protein, and its function is to stabilize the unwound single strand of DNA. dNTP is the abbreviation of deoxy-ribonucleoside triphosphate (deoxy-ribonucleoside triphosphate), which is a general term including dATP, dGTP, dTTP and dCTP.

[0188] The pipette is gently blown and mixed and centrifuged for a short time, and then placed in a PCR instrument for cfDNA end repair, and the conditions are shown in Table 3.

[0189] Table 3

[0190] (3) Ligation reaction: After the reagents in Table 4 are thawed, mix well and centrifuge for a short time, then add to the PCR tube in the previous step in turn (make sure to operate on ice):

[0191] Table 4

[0192] In the table, the adapter sequence is shown in Table 5 and Fig. 2, the adapter sequence contains a UMI sequence, and all C bases of the adapter sequence are methylated.

[0193] Table 5

[0194] The mixture was mixed gently by pipetting and centrifuged briefly, and then placed in a PCR instrument for ligation reaction. The ligation reaction conditions are shown in Table 6.

[0195] Table 6

[0196] (4) Product purification: 80 μL of purification magnetic beads were added to the product of step (3), and after mixing well, the mixture was incubated at room temperature for 5 min, and then placed in a magnetic stand for about 5 min to make the magnetic beads completely adsorbed and the solution clear. The supernatant was removed. 200 μL of freshly prepared 80% ethanol was added for rinsing, and the mixture was incubated at room temperature for 30 s to 60 s. The supernatant was removed, and the rinsing was repeated once. After the magnetic beads were dried, 29 μL of ultrapure water was added for elution. The mixture was incubated at room temperature for 3 min, and then placed in a magnetic stand.

[0197] (5) Oxidation reaction: the reagents shown in Table 7 below were thawed and mixed by inversion, and then centrifuged briefly. Subsequently, the product obtained in step (4) was added into the PCR tube in sequence, and the operation was performed on ice. The PCR tube was mixed gently by pipetting or shaking, and then centrifuged briefly. Subsequently, 5 μL of Fe(II) solution was added, and the mixture was mixed well and centrifuged briefly. After centrifugation, the PCR tube was placed in a PCR instrument for oxidation reaction, wherein the temperature of the heat cover was 75°C, the reaction temperature of the oxidation reaction was 37°C, and the reaction time was 60 min. After the reaction, the mixture was incubated at 37°C for 30 min after 1 μL of termination solution was added. In this step, the cytosine with methylation modification in the double-stranded DNA formed a protected cytosine, and a double-stranded DNA with a protected cytosine was obtained.

[0198] Table 7

[0199] In this embodiment, the enzyme methylation conversion kit (catalog number: E7125) of NEB Company was used.

[0200] (6) Product purification: 80 μL of purification magnetic beads were added to the product of step (5), and after thorough mixing, the mixture was allowed to stand at room temperature for 5 min, and then placed in a magnetic stand for about 5 min to allow the magnetic beads to be completely absorbed and the solution to be clear, and the supernatant was removed; 200 μL of freshly prepared 80% ethanol was added for rinsing, and the mixture was incubated at room temperature for 30 s to 60 s, and the supernatant was removed, and the rinsing was repeated once; after the magnetic beads were dried, 18 μL of ultrapure water was added for elution, and the mixture was allowed to stand at room temperature for 3 min and then placed in a magnetic stand.

[0201] (7) Deamination reaction: 4 μL of formamide was added to the product of step (6), and the mixture was thoroughly mixed, centrifuged briefly, and incubated at 85°C for 10 min; immediately after the reaction was completed, the mixture was placed on ice and the reaction system shown in Table 8 below was prepared in a PCR tube; the PCR tube was gently blown to mix the reaction solution in the PCR tube, and then centrifuged briefly; after centrifugation, the PCR tube was placed in a PCR instrument for deamination reaction, wherein the temperature of the thermal cover was 75°C, the reaction temperature of the deamination reaction was 37°C, and the reaction time was 180 min. As shown in FIG. 1, in this step, the unmethylated cytosine (C) in the double-stranded DNA of the protected cytosine formed uracil (U), and the double-stranded DNA with uracil (U) was obtained.

[0202] Table 8

[0203] In this embodiment, the enzyme methylation conversion module kit (item number: E7125) of NEB Company was used in this step.

[0204] (8) Product purification: 100 μL of purification magnetic beads were added to the product of step (7), and after thorough mixing, the mixture was allowed to stand at room temperature for 5 min, and then placed in a magnetic stand for about 5 min to allow the magnetic beads to be completely absorbed and the solution to be clear, and the supernatant was removed; 200 μL of freshly prepared 80% ethanol was added for rinsing, and the mixture was incubated at room temperature for 30 s to 60 s, and the supernatant was removed, and the rinsing was repeated once; after the magnetic beads were dried, 22 μL of ultrapure water was added for elution, and the mixture was allowed to stand at room temperature for 3 min and then placed in a magnetic stand.

[0205] (9) Product amplification: After thawing the Index Primer Mix and KAPA HiFi HotStart Uracil, mix them well and prepare the reaction system as shown in Table 9 in a sterile PCR tube. Gently blow the PCR tube to mix the reaction solution in the PCR tube, and centrifuge briefly. After centrifugation, place the PCR tube in a PCR instrument for amplification reaction. The reaction conditions of the amplification reaction are shown in Table 10. As shown in FIG. 1, in this step, the double-stranded DNA with uracil undergoes amplification reaction to obtain a DNA methylation library. Among them, different adapters are needed at both ends of the sequencing library, generally called P5 and P7. Index is used to distinguish different libraries, because the sequencing instrument produces a huge amount of data, and multiple libraries are often sequenced at a time, so Index needs to be added for distinction.

[0206] Table 9

[0207] Table 10

[0208] (10) Product purification: Add 45 μL of purification magnetic beads to the product of step (9), mix well, and stand at room temperature for 5 min. Place the magnetic beads in a magnetic stand for about 5 min to make the magnetic beads completely adsorbed and the solution clear, and remove the supernatant. Add 200 μL of freshly prepared 80% ethanol for rinsing, incubate at room temperature for 30 s to 60 s, remove the supernatant, and repeat once. After the magnetic beads are dried, add 32 μL of ultrapure water for elution, and place it in a magnetic stand after standing at room temperature for 3 min.

[0209] (11) Library quality control: Use Qubit 4.0 to determine the library concentration, and dilute the library to 1 ng / μL. Take out 1 μL for Agilent 4200 Tapestation system (Agilent, USA) detection.

[0210] (12) Hybridization reagent preparation: Prepare the hybridization solution as shown in Table 11, labeled as Hyb-1, mix well after preparation, and wait for use.

[0211] Table 11

[0212] In the subsequent steps, the methylation library needs to be denatured at high temperature, and then the single-stranded DNA molecules are hybridized with the probes at a hybridization temperature of, for example, 60°C. If the single-stranded DNA reanneals, it will affect the hybridization efficiency of the probes. To solve this problem, the single-stranded binding protein SSB is used to bind to the single-stranded DNA molecules in this embodiment. SSB forms a tetramer that specifically binds to 8-16 bases, avoiding the reannealing of single-stranded DNA, which can effectively improve the hybridization efficiency.

[0213] (13) Hybridization: Take 187.5 ng of different methylation rate standard library (0%, 5%, 25%, 50%, 100%, each 10) into a 1.5 mL centrifuge tube, mix well and centrifuge, mark as Cap-1, Cap-2, Cap-3, Cap-4, Cap-5, then mix the probe Panel, Cot1 DNA and Blocker Solution after thawing, and prepare the reagent system in Table 12 in a sterile PCR tube. As shown in Figure 1, in this step, the probe is hybridized to several nucleic acid fragments (i.e. target genomic region).

[0214] Table 12

[0215] Cot1 DNA is a placental DNA, mainly in the size range of 50bp to 30bp, and rich in repetitive DNA sequences, which can effectively block the repetitive DNA sequences of the target region and reduce non-specific hybridization.

[0216] Mix the reagent system in the sterile PCR tube well and centrifuge briefly, place the sterile PCR tube in a vacuum concentrator, concentrate to dry powder, then prepare the reagent system in Table 13 in the centrifuge tube.

[0217] Table 13

[0218] Mix gently and centrifuge briefly, and place in a PCR instrument for hybridization reaction. The conditions of the hybridization reaction are shown in Table 14.

[0219] Table 14

[0220] (14) Streptavidin magnetic bead cleaning:

[0221] 1) 40 min before the completion of the previous step reaction, take the streptavidin magnetic beads from 4°C, equilibrate at room temperature for 30 min;

[0222] 2) Take 100 μL of magnetic beads into a 1.5 mL low adsorption centrifuge tube, and use Beads Binding buffer for magnetic bead cleaning;

[0223] 3) Add 200 μL of Beads Binding Buffer to the centrifuge tube, mix gently by blowing and sucking 10 times, centrifuge immediately, place in the magnetic stand for several minutes, and use the pipette to remove the supernatant after the liquid is completely clarified, and remove the centrifuge tube from the magnetic stand;

[0224] 4) Repeat step 3) twice;

[0225] 5) Add 200 μL magnetic bead suspension to the centrifuge tube, mix gently by pipetting up and down, and transfer all of the magnetic bead suspension to a new 1.5 mL low-adsorption PCR tube.

[0226] (15) Elution buffer preparation: Elution buffers WB1 and WB2 were prepared as shown in Tables 15 and 16 below.

[0227] Table 15

[0228] Table 16

[0229] Because the proportion of A and T bases in the DNA strands in the methylation library increases after methylation conversion, the tetramethylammonium chloride in the elution buffer can increase the solubility of the DNA strands rich in A and T bases (TM), increase the TM value of the regions rich in A and T bases in the DNA strands, and make it close to the TM value of the DNA strands with normal proportions of A and T bases. The gradient concentration (0.1 M, 0.5 M, 1 M, 1.5 M, and 2 M) test results of tetramethylammonium chloride show that a concentration of 1 M is preferable, and the elution buffer is incubated at 48°C during the elution step. This can effectively wash away non-specific hybridization products and retain specific hybridization products rich in A and T bases.

[0230] Formamide can reduce the TM value of DNA strands, and each increase of 1% formamide can reduce the TM value by about 0.7°C. After incubation at 48°C, the elution buffer containing 5% formamide can greatly enhance the specificity of hybridization and reduce the loss of target region hybridization products to a lesser extent.

[0231] (16) Hybridization product elution: Mix 200 μL magnetic bead suspension with the hybridization product thoroughly, incubate at room temperature for 30 min, then wash with elution buffer WB1 at 65°C and incubate at 65°C for 5 min, repeat 3 times; finally, wash with elution buffer WB2 and incubate at 48°C for 5 min, repeat 3 times. As shown in FIG. 1, in this step, several nucleic acid fragments (i.e., target genomic regions) hybridized with the probe are pulled down and enriched.

[0232] (17) Eluted product amplification: After thawing the amplification primers and Kapa Hifi hotstart ready Mix (kk2601), mix them thoroughly and invert, and prepare the reagent system as shown in Table 17 below in a sterile PCR tube.

[0233] Table 17

[0234] Mix gently by pipetting and centrifuge briefly, then place in a PCR instrument for amplification reaction. The reaction conditions for the amplification reaction are shown in Table 18.

[0235] Table 18

[0236] (18) Product purification: 90 μL of purification magnetic beads were added to step (17), and after thorough mixing, the solution was incubated at room temperature for 5 min, and then placed in a magnetic stand for about 5 min to allow the magnetic beads to be completely absorbed and the solution to be clear. The supernatant was removed. 200 μL of freshly prepared 80% ethanol was added for rinsing, and the solution was incubated at room temperature for 30 s to 60 s. The supernatant was removed, and the rinsing was repeated once. After the magnetic beads were dried, 32 μL of ultrapure water was added for elution, and the solution was placed in a magnetic stand after incubation at room temperature for 3 min.

[0237] (19) Sequencing: The methylation library obtained in step (18) was diluted to 1 ng / μL, and 1 μL was taken for detection by Agilent 4200 Tapestation system (Agilent, USA). In addition, 1 μL was taken for qPCR detection, and the concentration for sequencing was determined according to the detection results. According to the concentration obtained in the previous step, the library was diluted to the required concentration for sequencing (2 nmol), and PE150 sequencing was performed on the Novaseq sequencing platform of Illumina Company, with a data volume of 20G per sample.

[0238] (20) Data quality control: low-quality sequences and sequencing adapter sequences were filtered to obtain high-quality data, and a corresponding quality control report was generated.

[0239] (21) Genome alignment and deduplication: high-quality data were compared with the reference genome to find the location of each read on the reference genome, and repeated sequences introduced by PCR amplification were removed. The sequencing depth and sequencing coverage were also calculated.

[0240] (22) Methylation information extraction: after obtaining the deduplicated alignment results, methylation site detection was performed to obtain the methylation level under different sequence environments (CG, CHH, CHG, where H represents A, C, T).

[0241] Example 2

[0242] This example is used to obtain a test combination for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma.

[0243] (1) Clinical sample collection:

[0244] 1) Take whole blood (2 mL to 8 mL) in an 8.5 mL Roche cfDNA free nucleic acid blood collection tube (please operate at room temperature, and store the blood at room temperature (18°C to 25°C) after blood collection. The subsequent plasma extraction operation should be performed within 72 hours).

[0245] 2) First centrifugation, 1350g / min, 4℃ low temperature centrifugation for 12 min, carefully take out the light yellow supernatant (avoid contamination of leukocyte layer), transfer to 2 mL DNase free sterile centrifuge tube;

[0246] 3) Second centrifugation, 13500g / min, 4℃ low temperature centrifugation for 5 min, carefully take out the supernatant (thoroughly remove leukocytes), transfer to 2-3 tubes of 2 mL DNase free sterile centrifuge tube, freeze in -80℃ refrigerator (about 4 mL-6 mL of clean plasma should be obtained, according to the color to judge whether there is hemolysis situation to make risk control, or renotify the hospital sampling according to the standard process); the white precipitate at the bottom of the tube is leukocytes, which is marked and frozen in the -80℃ refrigerator.

[0247] 4) The obtained plasma is numbered and frozen at -80℃ refrigerator for use, and accurate electronic file plasma separation table is also recorded during the process for future reference.

[0248] (2) Use ccfDNA (Cat No. 55204) kit to extract cfDNA from plasma samples.

[0249] (3) Whole genome methylation library construction / targeted capture / library sequencing is performed according to the experimental method in Example 1, and the sequencing number is processed and analyzed to screen out the difference methylation region (DMR).

[0250] (4) 81 patients are designed for clinical research according to FIG. 3, and the training set and the test set are split according to the ratio of 7:3, wherein the training set includes 40 cases of PR+CR (treatment effective group) and 16 cases of PD+SD (treatment ineffective group), and the test set includes 18 cases of PR+CR (treatment effective group) and 7 cases of PD+SD (treatment ineffective group). In the training set, 557 DMRs are obtained by chi-square test feature selection method, and 33 final DMRs are screened out by using cross-validation feature recursive elimination algorithm RFECV for feature screening, and the specific information is shown in Table 1. The model is constructed by using the logistic regression algorithm, and the model performance is verified on the test set.

[0251] (5) In the training set, 33 final DMRs are obtained, and in the validation set, the methylation rate value is input into the model, and the model outputs the treatment response prediction probability value. The probability value 0.5 is taken as the threshold, and greater than 0.5 is the treatment effective, and less than 0.5 is the treatment ineffective.

[0252] Figure 4 is a receiver operating characteristic (ROC) curve. As can be seen from Figure 4, the test set evaluation result shows that the model AUC (area under the receiver operating characteristic curve) value is 0.81, and the closer the AUC is to 1, the better the prediction effect of the model on the chemotherapy effect of diffuse large B-cell lymphoma.

[0253] Figure 5 is a confusion matrix. As can be seen from Figure 5, the total number of PD+SD (treatment ineffective group) samples is 7, the actual detection accuracy of the model is 5, the total number of PR+CR (treatment effective group) samples is 18, the actual detection accuracy of the model is 15, and the accuracy is (15+5) / 25=80%.

[0254] Figure 6 is the sensitivity and specificity of the clinical sample detection in the test set. As can be seen from Figure 6, the sensitivity of the gene marker (target genomic region) provided by the embodiment of the present disclosure for predicting the chemotherapy effect of diffuse large B-cell lymphoma is 83%, and the specificity is 71%.

[0255] Figure 7 is a result graph of the detection value and the theoretical value of the methylation rate of different sites of the target genomic region. The abscissa represents the sample input amount, and the ordinate represents different methylation rates. As can be seen from Figure 7, the detection value and the theoretical value of the methylation rate of different sites of the target genomic region are basically consistent. Further, it is illustrated that the methylation rate detection accuracy of the gene library obtained by the library construction method and the target gene capture method provided by the embodiment of the present disclosure is high.

[0256] Figure 8 is a linear regression correlation analysis graph of the detection methylation rate and the theoretical methylation rate of different sites of the target genomic region. The abscissa represents the detection methylation rate, and the ordinate represents the theoretical methylation rate. As can be seen from Figure 8, the linear correlation coefficient R2=0.95, indicating that the detection methylation rate is basically consistent with the theoretical methylation rate; the P value is less than 0.01, wherein the P value is called the significance value, indicating that the accuracy of the gene library obtained by the library construction method and the target gene capture method provided by the embodiment of the present disclosure is high; the gene library with a sample amount of 1 ng-50 ng can effectively detect a methylation rate gradient of 0%-100%, indicating that the technical method is compatible with the detection of trace cfDNA.

[0257] To sum up, in the embodiments of the present disclosure, the methylation library construction technology is adopted and combined with target region targeted capture for high-throughput sequencing. After the accuracy of the method is verified by a standard sample, 81 clinical cohort samples are used for testing. Among them, the training set includes 40 cases of PR+CR (treatment effective group) and 16 cases of PD+SD (treatment ineffective group), the test set includes 18 cases of PR+CR (treatment effective group) and 7 cases of PD+SD (treatment ineffective group), the methylation genomic characteristics of the diffuse large B-cell lymphoma chemotherapy effective and ineffective are identified, and the model is constructed by using common classifier algorithms, so as to realize the prediction of the diffuse large B-cell lymphoma chemotherapy effect.

[0258] Moreover, the enzyme methylation conversion technology is adopted in the embodiments of the present disclosure, the cfDNA is not degraded, and the methylation data can be used for fragmentation analysis.

[0259] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art can think of changes or replacements within the technical range disclosed by the present disclosure, which should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A composition comprising: a plurality of different bait oligonucleotides configured to collectively hybridize to DNA molecules derived from a plurality of target genomic regions; wherein each of the plurality of target genomic regions is differentially methylated in diffuse large B-cell lymphoma chemotherapy responders compared to diffuse large B-cell lymphoma chemotherapy non-responders.

2. The composition of claim 1, wherein, the plurality of different bait oligonucleotides are configured to hybridize to DNA molecules derived from at least 20%, at least 25%, or at least 50% of the plurality of target genomic regions of any one of Tables 1-33.

3. The composition according to claim 1 or 2, wherein, the plurality of different bait oligonucleotides are configured to hybridize to DNA molecules derived from at least 20%, at least 25%, or at least 50% of the plurality of target genomic regions of Tables 1-33.

4. The composition of claim 1, wherein, the plurality of different bait oligonucleotides are configured to hybridize to DNA molecules derived from at least 20% of the plurality of target genomic regions of Tables 1-33.

5. The composition of claim 4, wherein, the plurality of DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of Tables 1-33.

6. A composition comprising: a plurality of different bait oligonucleotides configured to hybridize to DNA molecules derived from at least 20% of the plurality of target genomic regions of any one of Tables 1-33.

7. The composition of claim 6, wherein, the plurality of different bait oligonucleotides are configured to hybridize to DNA molecules derived from 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of any one of Tables 1-33.

8. The composition according to claim 6 or 7, wherein, the plurality of different bait oligonucleotides are configured to hybridize to DNA molecules derived from at least 20% of the plurality of target genomic regions of Tables 1-33.

9. The composition of claim 8, wherein, the plurality of DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of Tables 1-33.

10. The composition according to any one of claims 1 to 9, wherein, the plurality of DNA molecules are converted cfDNA fragments.

11. The composition of claim 10, wherein, the plurality of target genomic regions are hypermethylated regions, hypomethylated regions, or bimodal regions that can be hypermethylated or hypomethylated.

12. The composition of claim 10, wherein, the plurality of bait oligonucleotides are configured to hybridize to hypermethylated converted DNA molecules, hypomethylated converted DNA molecules, or both hypermethylated and hypomethylated converted DNA molecules derived from each target genomic region.

13. The composition according to any one of claims 1 to 12, wherein, each of the plurality of bait oligonucleotides is bound to an affinity moiety.

14. The composition of claim 13, wherein, each of the plurality of different bait oligonucleotides is bound to a magnetic bead surface.

15. A method for enriching converted cfDNA fragments that can provide information on diffuse large B-cell lymphoma chemotherapy effectiveness, the method comprising the steps of: The composition of any one of claims 1 to 14 is contacted with DNA derived from a test subject, and a sample of cfDNA corresponding to a number of genomic regions associated with diffuse large B-cell lymphoma chemotherapy effect is enriched by hybridization capture.

16. A method for obtaining sequence information that provides information on diffuse large B-cell lymphoma chemotherapy efficacy, the method comprising the steps of: a) enriching converted DNA from a test subject by contacting the converted DNA with the composition of any one of claims 1 to 14, and b) sequencing the enriched converted DNA.

17. A method for diffuse large B-cell lymphoma chemotherapy efficacy prediction, comprising the steps of: a) capturing a number of cfDNA fragments from a test subject with the composition of any one of claims 1 to 14, b) detecting the captured number of cfDNA fragments, and c) applying a trained classifier to the captured number of DNA fragments to predict diffuse large B-cell lymphoma chemotherapy efficacy of the test subject.

18. The method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma according to claim 17, wherein, The trained classifier predicts diffuse large B-cell lymphoma chemotherapy efficacy of the test subject.

19. The method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma according to claim 17 or 18, wherein, The trained classifier is a mixed model classifier.

20. The method for predicting the efficacy of chemotherapy for diffuse large B-cell lymphoma according to any one of claims 17 to 19, wherein, The classifier is trained on a number of converted DNA sequences derived from target genomic regions selected from any one of lists 1 to 33. The method further comprises the step of: d) selecting a treatment regimen for the test subject based on the prediction of diffuse large B-cell lymphoma chemotherapy efficacy of the test subject. The method further comprises the step of: e) monitoring the test subject for response to the selected treatment regimen. The method further comprises the step of: f) selecting a treatment regimen for the test subject based on the prediction of diffuse large B-cell lymphoma chemotherapy efficacy of the test subject. The method further comprises the step of: g) monitoring the test subject for response to the selected treatment regimen. The method further comprises the step of: h) selecting a treatment regimen for the test subject based on the prediction of diffuse large B-cell lymphoma chemotherapy efficacy of the test subject. The method further comprises the step of: i) monitoring the test subject for response to the selected treatment regimen. The method further comprises the step of: j) selecting a treatment regimen for the test subject based on the prediction of diffuse large B-cell lymphoma chemotherapy efficacy of the test subject. The method further comprises the step of: k) monitoring the test subject for response to the selected treatment regimen.