Composition and method for enriching transformed cfdna fragments

By designing specific hybridization to capture converted cfDNA fragments and utilizing a classifier, the sampling errors and invasiveness issues in the diagnosis of diffuse large B-cell lymphoma were resolved, achieving efficient and accurate non-invasive detection.

WO2026000126A1PCT designated stage Publication Date: 2026-01-02BOE TECHNOLOGY GROUP CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/101041
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies for the diagnosis of diffuse large B-cell lymphoma carry risks of sampling errors and invasive surgery, and rely on tumor tissue examination, making it difficult to efficiently and non-invasively detect the disease status.

Method used

A composition is provided comprising several different decoy oligonucleotides configured to hybridize to DNA molecules derived from multiple target genomic regions, thereby capturing and enriching converted cfDNA fragments through hybridization, and using a trained classifier to determine the presence or absence of a disease, combined with sequencing technology to obtain sequence information.

Benefits of technology

This technology enables non-invasive and efficient detection of diffuse large B-cell lymphoma, improving diagnostic accuracy and sensitivity while reducing testing costs and patient risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024101041_02012026_PF_FP_ABST
    Figure CN2024101041_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A composition, comprising: a plurality of different decoy oligonucleotides configured to collectively hybridize to a DNA molecule derived from a plurality of target genomic regions, wherein each genomic region of the plurality of target genomic regions is differentially methylated in diffuse large B-cell lymphomas compared to non-diffuse large B-cell lymphomas.
Need to check novelty before this filing date? Find Prior Art

Description

Composition and method for enriching converted cfDNA fragments Technical Field

[0001] This disclosure relates to the field of biotechnology, and more particularly to a composition and a method for enriching converted cfDNA fragments. Background Technology

[0002] Diffuse large B-cell lymphoma (DLBCL) is an aggressive tumor originating from mature B cells and is the most common type of non-Hodgkin's lymphoma, accounting for 45.8% in China. Approximately 30% of DLBCL patients present with limited-stage (i.e., early-stage) disease, which can also be defined as stage I or II, while 70% of patients present with advanced-stage disease. The disease control rate for limited-stage DLBCL is as high as 90%, and clinical treatment studies show that the overall survival (OS) rate of about 4 years for limited-stage DLBCL is approximately 92%. However, according to data from the past 20 years, the cure rate for advanced-stage DLBCL is only 60%.

[0003] Summary of the Invention

[0004] On one hand, a composition is provided comprising: several different decoy oligonucleotides configured to collectively hybridize to a DNA molecule derived from a plurality of target genomic regions; wherein each of the plurality of target genomic regions is differentially methylated in diffuse large B-cell lymphoma compared to non-diffuse large B-cell lymphoma.

[0005] In some embodiments, the plurality of different bait oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20%, at least 25%, or at least 50% of the plurality of target genomic regions of any of Lists 1 to 31.

[0006] In some embodiments, the plurality of different bait oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20%, at least 25%, or at least 50% of the plurality of target genomic regions listed in 1 to 31.

[0007] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions listed in 1 to 31.

[0008] In some embodiments, the plurality of DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions listed in 1 to 31.

[0009] On the other hand, a composition is provided comprising several different decoy oligonucleotides configured to hybridize to several DNA molecules, said DNA molecules being at least 20% derived from any of the several target genomic regions listed in 1 to 31.

[0010] In some embodiments, the plurality of different bait oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions derived from any of Lists 1 to 31.

[0011] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions listed in 1 to 31.

[0012] In some embodiments, the plurality of DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions listed in 1 to 31.

[0013] In some embodiments, the plurality of DNA molecules are converted cfDNA fragments.

[0014] In some embodiments, the plurality of target genomic regions are hypermethylated regions, hypomethylated regions, or binary regions that are either hypermethylated or hypomethylated.

[0015] In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a hypermethylated converted DNA molecule, a hypomethylated converted DNA molecule, or both hypermethylated and hypomethylated converted DNA molecules derived from each target genomic region.

[0016] In some embodiments, each of the plurality of decoy oligonucleotides is conjugated to an affinity moiety.

[0017] In some embodiments, each of the several different bait oligonucleotides is bound to the surface of the magnetic bead.

[0018] In another aspect, a method for enriching converted cfDNA fragments, which can provide information on diffuse large B-cell lymphoma, is provided. The method includes the steps of: contacting a composition as described in any of the above embodiments with DNA derived from a test subject, and enriching a sample of cfDNA corresponding to several genomic regions associated with diffuse large B-cell lymphoma by hybridization capture.

[0019] In another aspect, a method for obtaining sequence information that can provide information on the presence or absence of diffuse large B-cell lymphoma is provided, the method comprising the steps of: a) enriching the converted DNA by contacting the converted DNA from the test subject with the composition as described in any of the above embodiments, and b) sequencing the enriched converted DNA.

[0020] In another aspect, a method for determining whether a subject has diffuse large B-cell lymphoma is provided, the method comprising the steps of: a) capturing several cfDNA fragments from the subject using the composition described in any of the above embodiments, b) detecting the captured several cfDNA fragments, and c) applying a trained classifier to the captured several DNA fragments to determine whether the subject has diffuse large B-cell lymphoma.

[0021] In some embodiments, the trained classifier determines the presence or absence of diffuse large B-cell lymphoma.

[0022] In some embodiments, the trained classifier is a hybrid model classifier.

[0023] In some embodiments, the classifier is trained on a plurality of converted DNA sequences derived from a target genomic region selected from any of List 1 to 31. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in this disclosure, the accompanying drawings used in some embodiments of this disclosure will be briefly described below. Obviously, the drawings described below are merely drawings of some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings. Furthermore, the drawings described below can be considered as schematic diagrams and are not intended to limit the products or methods involved in the embodiments of this disclosure.

[0025] Figure 1 is a technical path diagram for obtaining target genome regions according to some embodiments;

[0026] Figure 2 is a structural diagram of a first connector sequence according to some embodiments;

[0027] Figure 3 is a structural diagram of a second connector sequence according to some embodiments;

[0028] Figure 4 shows the distribution of cfDNA fragment lengths in samples according to some embodiments;

[0029] Figure 5 shows the DNA length distribution in a gene library according to some embodiments;

[0030] Figure 6 is a graph of subject operating characteristics according to some embodiments;

[0031] Figure 7 shows the sensitivity and specificity of the clinical samples detected in the test set according to some embodiments;

[0032] Figure 8 is a comparison of the methylation rates of bases 1-150 (5′-3′ orientation) in Read2 sequences from libraries of Examples 1 and 3.

[0033] Figure 9 shows the results of the detected and theoretical values ​​of methylation rate at different sites in the target genome region according to some embodiments;

[0034] Figure 10 is a correlation analysis diagram of linear regression between the detected methylation rate and the theoretical methylation rate at different sites in the target genome region according to some embodiments. Detailed Implementation

[0035] Unless otherwise defined, all technical and scientific terms used in the embodiments of this disclosure have the meanings commonly understood by one of ordinary skill in the art to which this description pertains. As used herein, the following terms have the meanings attributed to them hereinafter.

[0036] As used herein, any reference to "an embodiment" or "an embodiment" means a particular embodiment, feature, structure, or characteristic described in association with the described embodiment, which is included in at least one embodiment. The appearance of the term "in some embodiments" throughout the specification does not necessarily refer to the same embodiment, thereby providing a framework for the various possibilities of several described embodiments operating together.

[0037] As used herein, “comprising” or any other variation thereof is intended to cover a non-exclusive inclusion. For example, a procedure, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent in such a procedure, method, article, or apparatus. Furthermore, unless expressly defined to the contrary, “or” means inclusive or rather, not exclusive. For example, a condition A or B is satisfied by any of the following: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); and both A and B are true (or exist).

[0038] Furthermore, the use of the word "a" is applied to describe elements and components of several embodiments herein. This is merely for convenience and to give the general meaning of this description. This description should be read as including one or at least one, and the singular includes the plural, unless clearly otherwise implied.

[0039] As used herein, ranges and amounts can be expressed as “about” for a specific value or range. “About” also includes the precise amount. Therefore, “about 5 micrograms” means “about 5 micrograms” and also “5 micrograms”. Generally, the term “about” includes an amount expected to be within experimental error. In some embodiments, “about” means a indicated number or value that is “+” or “-” 20%, 10%, or 5%. Furthermore, ranges referenced in embodiments of this disclosure are understood to be all values ​​within the range, including the referenced endpoints. For example, a range of 1 to 50 is understood to include any number, combination of several numbers, or subrange of numbers that comes from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50.

[0040] The term “methylation,” as used herein, refers to the process by which a methyl group is added to a DNA molecule. For example, a hydrogen atom on the pyrimidine ring of a cytosine base can be converted to a methyl group, forming 5-methylcytosine. The term also refers to the process by which a hydroxymethyl group is added to a DNA molecule, for example, through the oxidation of the methyl group on the pyrimidine ring of a cytosine base. Methylation and hydroxymethylation tend to occur at the dinucleotides of cytosine and guanine, referred to as “CpG sites” in embodiments of this disclosure.

[0041] The term "methylation" can also refer to the methylation state of a CpG site. A CpG site with 5-methylcytosine is methylated. A CpG site with a hydrogen atom on the pyrimidine ring of the cytosine base is unmethylated.

[0042] The term “methylation site,” as used herein, refers to a region of a DNA molecule to which a methyl group can be added. “CpG” sites are the most common methylation sites, but methylation sites are not limited to CpG sites. For example, DNA methylation can occur in cytosine at CHG and CHH, where H is adenine, cytosine, or thymine.

[0043] The term “CpG site” is used in embodiments of this disclosure to refer to a region of a DNA molecule in which a cytosine nucleotide is followed by a guanine nucleotide in a linear sequence of several bases along the 5′ to 3′ direction. “CpG” is short for 5′-C-phosphate-G-3′, which consists of cytosine and guanine separated by only one phosphate group. The cytosine in the CpG dinucleotide can be methylated to form 5-methylcytosine.

[0044] The term "UpG" is short for 5′-U-phosphate-G-3′, which refers to uracil and guanine separated by only one phosphate group. UpG can be produced, for example, by bisulfite treatment, which converts unmethylated cytosine into uracil. Cytosine can also be converted into uracil by other methods known in the art, such as chemical modification, synthesis, or enzymatic conversion.

[0045] Terms such as “hypomethylation” or “hypermethylation”, as used herein, refer to the methylation state of a DNA molecule containing multiple (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) CpG sites, where a high proportion (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%) of the CpG sites may be unmethylated or methylated.

[0046] The term “healthy sample,” as used herein, means a sample comprising genomic DNA from an individual not diagnosed with diffuse large B-cell lymphoma. The genomic DNA may be, but is not limited to, a fragment of cfDNA or chromosomal DNA from an individual without diffuse large B-cell lymphoma (e.g., a healthy individual), which may be sequenced and whose methylation status may be assessed. When the genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing the genome of an individual without diffuse large B-cell lymphoma, a healthy sample may mean genomic DNA or a fragment of cfDNA containing the genomic sequence. The term “several healthy samples,” as a plural, means several samples comprising genomic DNA from multiple individuals, each diagnosed without diffuse large B-cell lymphoma. In various embodiments, healthy samples from more than 10, 20, 50, 100, 200, 300, 500, 1000, 2000, 5000, 10000, 20000, 40000, 50000 or more individuals diagnosed without diffuse large B-cell lymphoma were used.

[0047] The term “training sample,” as used herein, means a sample used to train the classifier described in embodiments of this disclosure and / or select one or more genomic regions for the detection of diffuse large B-cell lymphoma. The training sample may include genomic DNA or derivatives thereof from one or more healthy subjects or from one or more subjects with diffuse large B-cell lymphoma. The genomic DNA may be, but is not limited to, several cfDNA fragments or chromosomal DNA. The genomic DNA may be sequenced and its methylation status may be assessed. When the genomic sequence is obtained from a public database (e.g., Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing the genome of an individual, a training sample may mean genomic DNA or cfDNA fragments having said genomic sequence.

[0048] The term “test sample,” as used herein, means a sample from an object whose health condition has been or will be tested using the classifier and / or laboratory assay combination described herein. The test sample may include genomic DNA or derivatives thereof. The genomic DNA may be, but is not limited to, several cfDNA fragments or chromosomal DNA.

[0049] The term "target genomic region," as used herein, refers to a region of the genome selected for analysis in a test sample. An assay assay is a combination of several probes designed to hybridize to several nucleic acid fragments derived from or a segment of the target genomic region. Nucleic acid fragments derived from the target genomic region refer to nucleic acid fragments produced through degradation, cleavage, transformation, or other treatments of DNA from the target genomic region.

[0050] Various target genomic regions are described according to their chromosomal location. Chromosomal DNA is double-stranded, therefore a target genomic region comprises two DNA strands: one strand having the sequence provided in the list, and a second strand that is the opposite complementary strand of the sequence in the list. Probes can be designed to hybridize to one or both sequences. Optionally, probes hybridize to the converted sequence.

[0051] "Converted cfDNA molecule" and "processed modified fragment obtained from said cfDNA molecule" refer to DNA molecules obtained by treating a sample of DNA or cfDNA molecules to distinguish between methylated and unmethylated nucleotides in DNA or cfDNA molecules. For example, the sample is treated to convert unmethylated cytosine ("C") to uracil ("U"). This conversion, for instance, is accomplished using an enzymatically catalyzed conversion reaction, for example, using a cytidine deaminase (such as APOBEC). After treatment, the converted DNA molecule or cfDNA molecule includes additional uracil not present in the original cfDNA sample. The DNA strand containing uracil is amplified by DNA polymerase, resulting in the addition of adenine to the newly generated complementary strand, instead of guanine, which normally complements cytosine or methylcytosine.

[0052] The terms "cell-free nucleic acid," "cell-free DNA," or "cfDNA" refer to nucleic acid fragments that circulate within an individual's body (e.g., in the bloodstream) and originate from one or more healthy cells and / or from one or more cells belonging to an individual with diffuse large B-cell lymphoma. Furthermore, cfDNA can originate from other sources such as viruses or fetuses.

[0053] The term “fragment,” as used herein, can mean a segment of a nucleic acid molecule. For example, in one embodiment, a fragment can mean a cfDNA molecule in blood or a blood sample, or a cfDNA molecule extracted from a plasma or a plasma sample. An amplified product of a cfDNA molecule can also be referred to as a “fragment.” In another embodiment, the term “fragment,” as described herein, means a sequence read, or a set of sequence reads, that has been processed for (e.g., in machine learning-based classification) subsequent analysis. For example, as is well known in the art, raw sequence reads can be aligned to a reference genome, and matched end sequence reads can be assembled into a longer fragment for subsequent analysis.

[0054] The term “individual” refers to a human being. The term “healthy individual” refers to an individual presumed not to have diffuse large B-cell lymphoma.

[0055] The term “subject” refers to an individual whose DNA is analyzed. A subject can be a test subject whose DNA is evaluated using a targeted assay combination as described herein to assess whether the individual has diffuse large B-cell lymphoma (DBL). A subject can also be an individual in the control group who is known not to have DBL. A subject can also be an individual with DBL, who is known to have DBL.

[0056] The term “sequence readout” as used herein refers to a nucleotide sequence readout from a sample. Sequence readouts can be obtained by various methods provided herein or known in the art.

[0057] The term “sequencing depth,” as used in this article, refers to the count of the number of times a given target nucleic acid in a sample is sequenced (e.g., the count of sequence reads at a given target region). Increasing sequencing depth can reduce the amount of nucleic acid required to assess disease status.

[0058] The term "probe set" or "polynucleotide-containing probe set" of a detection combination or decoy set generally means all probes used with a particular detection combination or decoy set. For example, in some embodiments, a detection combination or decoy set may include (1) several probes having the characteristics specified herein (e.g., several probes for capturing cell-free DNA fragments corresponding to or derived from genomic regions presented herein in one or more lists) and (2) additional probes that do not contain the characteristics specified herein. The probe set of a detection combination generally means all probes used with the detection combination or decoy set, including probes that do not contain the specified characteristics.

[0059] Diffuse large B-cell lymphoma (DLBCL) is an aggressive tumor that originates from mature B cells and is the most common type of non-Hodgkin lymphoma.

[0060] Currently, according to the lymphoma classification principles established by the World Health Organization (WHO), the diagnosis of DLBCL mainly relies on detailed examination of tumor tissue, primarily through biopsy specimens evaluated by hematologists. However, due to the high heterogeneity of tumor tissue, sampling errors or false negatives may occur, and invasive surgery also poses risks to patients.

[0061] Embodiments of this disclosure provide a diagnostic assay for diffuse large B-cell lymphoma, comprising a plurality of probes or probe pairs. The plurality of diagnostic assays described in the embodiments of this disclosure may alternatively be referred to as a plurality of decoy sets, or as a plurality of compositions comprising a plurality of decoy oligonucleotides. The plurality of probes may be a plurality of polynucleotide-containing probes specifically designed to target one or more nucleic acid molecules corresponding to, or derived from, a plurality of genomic regions differentially methylated between diffuse large B-cell lymphoma and non-diffuse large B-cell lymphoma samples. In some embodiments, the plurality of target genomic regions (or nucleic acids derived from the plurality of target genomic regions) are selected to maximize classification accuracy, subject to a size budget (determined by sequencing budget and desired sequencing depth).

[0062] To design a diagnostic assay for diffuse large B-cell lymphoma, the analysis system can collect information on the methylation status of several CpG sites from several nucleic acid fragments, said fragments being derived from several samples, such as samples known to have diffuse large B-cell lymphoma, samples considered healthy, etc. These samples can be processed to determine the methylation status of the CpG sites, or the information can be obtained from TCGA. The analysis system can be any general-purpose computing system having a computer processor and a computer-readable storage medium having several instructions for executing the computer processor to perform any or all of the operations described in the embodiments of this disclosure.

[0063] The analysis system can select target genomic regions based on the methylation patterns of several nucleic acid fragments. One approach considers the pairwise resolvability between several regions (or more specifically, CpG sites within several regions). Another approach considers the resolvability of several regions (or more specifically, CpG sites within several regions) when each result is considered relative to the remaining several results. From the selected target genomic regions with high resolvability, the analysis system can design probes to target several fragments from the selected genomic regions.

[0064] The analysis system can generate diffuse large B-cell lymphoma assay kits of varying sizes. For example, a small-sized diffuse large B-cell lymphoma assay kit includes several probes targeting several of the most informative genomic regions; a medium-sized diffuse large B-cell lymphoma assay kit includes several probes from the small-sized kit, plus additional probes targeting informative genomic regions in the second layer; and a large-sized kit includes several probes from the small and medium-sized kits, plus more probes targeting informative genomic regions in the third layer.

[0065] With data obtained from such a diffuse large B-cell lymphoma assay suite (e.g., the methylation status of nucleic acids derived from the diffuse large B-cell lymphoma assay suite), the analysis system can train a classifier using various classification techniques to predict the likelihood of a sample having diffuse large B-cell lymphoma.

[0066] In some embodiments, the composition includes several different decoy oligonucleotides configured to collectively hybridize to a DNA molecule derived from multiple target genomic regions. Each of the multiple target genomic regions is differentially methylated in diffuse large B-cell lymphoma compared to non-diffuse large B-cell lymphoma.

[0067] For example, each of several genomic regions includes at least four methylation sites, and at least four of these methylation sites have an anomalous methylation pattern in diffuse large B-cell lymphoma samples, or have a differential methylation status between different diffuse large B-cell lymphoma samples. For instance, in one embodiment, at least four methylation sites are differentially methylated between diffuse large B-cell lymphoma and non-diffuse large B-cell lymphoma samples.

[0068] In some embodiments, several different decoy oligonucleotides are configured to hybridize to several DNA molecules, which are derived from several target genomic regions of any of Lists 1 to 31. Lists 1 to 31 are selected from Table 1 below.

[0069] Table 1

[0070] DMR stands for Differentially Methylated Region, which refers to differentially methylated regions that are included in the target genome region.

[0071] In some examples, the diffuse large B-cell lymphoma assay kit includes several probes, each of which is configured to hybridize to a converted cfDNA molecule that corresponds to one or more target genomic regions of any of the probes listed in Lists 1 to 31.

[0072] In some examples, several different decoy oligonucleotides are configured to hybridize to several DNA molecules, which are derived from at least 20%, at least 25%, or at least 50% of several target genomic regions of any of Lists 1 to 31.

[0073] For example, several target genomic regions may be selected from List 1. A method for detecting diffuse large B-cell lymphoma includes the step of: assessing the methylation status of several sequence readings derived from several target genomic regions of List 1.

[0074] For example, several target genomic regions may be selected from List 2. A method for detecting diffuse large B-cell lymphoma includes the step of: assessing the methylation status of several sequence readings derived from several target genomic regions of List 2.

[0075] For example, several target genomic regions may be selected from List 8. A method for detecting diffuse large B-cell lymphoma includes the step of: assessing the methylation status of several sequence readings derived from several target genomic regions of List 8.

[0076] For example, several target genomic regions may be selected from List 20. One method for detecting diffuse large B-cell lymphoma includes the step of: assessing the methylation status of several sequence readings derived from several target genomic regions of List 20.

[0077] For example, several target genomic regions may be selected from 2, 5, 10, 15, 20, 28 or more of those listed in 1 to 31.

[0078] Because several probes are configured to hybridize to converted DNA or cfDNA molecules that correspond to or are derived from one or more target genomic regions, the probes may have sequences different from the target genomic regions.

[0079] For example, a DNA sequence containing an unmethylated CpG site will be converted to include UpG instead of CpG because the unmethylated cytosine is converted to uracil via a conversion reaction. As a result, the probe is configured to hybridize to a sequence that includes UpG, rather than the normally present unmethylated CpG. Therefore, the complementary site to the unmethylated site in the probe can include CpA instead of CpG, and some probes targeting a monomethylated site where all methylated sites are unmethylated may not contain a guanine (G) base.

[0080] A Diffuse Large B-Cell Lymphoma (DBL) assay kit can be used to detect the presence or absence of DBL. For example, the DBL assay kit is designed to enrich several target genomic regions based on differentially methylated sequencing data derived from cfDNA of individuals with DBL and non-DBL.

[0081] In some embodiments, several different decoy oligonucleotides are configured to hybridize to several DNA molecules, which are derived from at least 20%, at least 25%, or at least 50% of several target genomic regions listed in 1 to 31.

[0082] In some embodiments, several different decoy oligonucleotides are configured to hybridize to several DNA molecules, the DNA molecules being at least 20% derived from the several target genomic regions listed in Lists 1 to 31.

[0083] In some embodiments, several DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the several target genomic regions listed in 1 to 31.

[0084] In some examples, the diffuse large B-cell lymphoma (DFBLC) assay suite includes several probes that can selectively hybridize and optionally enrich several cfDNA fragments that are differently methylated in both DFBLC and non-DFBLC samples. Sequences from these enriched fragments can provide information relevant to the detection of DFBLC.

[0085] For example, several probes are designed to target several genomic regions identified as having an aberrant methylation pattern in a diffuse large B-cell lymphoma sample. Several probes are designed to target several genomic regions identified as differentially hypermethylated or hypomethylated in a specific diffuse large B-cell lymphoma to provide additional selectivity and specificity for detection. For example, a diffuse large B-cell lymphoma diagnostic kit includes several probes targeting several hypomethylated fragments. For example, a diffuse large B-cell lymphoma diagnostic kit includes several probes targeting several hypermethylated fragments. In some embodiments, a diffuse large B-cell lymphoma diagnostic kit includes several probes targeting a first group of several hypermethylated fragments and several probes targeting a second group of several hypomethylated fragments.

[0086] In some examples, several genomic regions may be selected when those regions produce anomalously methylated DNA molecules in samples with diffuse large B-cell lymphoma.

[0087] In some examples, several genomic regions can be further filtered based on their methylation patterns to select only a few genomic regions that may provide information. For example, based on the differentially methylated CpG sites between diffuse large B-cell lymphoma and non-diffuse large B-cell lymphoma samples, calculations can be performed for each CpG or several CpG sites for selection.

[0088] In some examples, once several probes hybridize and capture several DNA fragments corresponding to or derived from the target genomic region, the hybridized probe-DNA fragment intermediate is separated, the target DNA is amplified, and the methylation status of the target DNA is determined by sequencing or hybridization to magnetic beads. Sequence readouts provide information associated with the detection of diffuse large B-cell lymphoma. For this purpose, diffuse large B-cell lymphoma detection assemblies are designed to include several probes that can capture several fragments that collectively provide information associated with the detection of diffuse large B-cell lymphoma.

[0089] Several selected genomic regions can be located in various different positions within the genome, including but not limited to exons, introns, intergenic regions, and other parts.

[0090] In some cases, primers can be used (e.g., via PCR) to specifically amplify several target / biomarkers of interest, thereby enriching the sample with the desired target / biomarkers. For example, forward and reverse primers can be designed for each genomic region of interest and used to amplify several fragments corresponding to or derived from the desired genomic region. Therefore, while embodiments of this disclosure focus on assay suites for diffuse large B-cell lymphoma and decoy sets for hybridization capture, embodiments of this disclosure also encompass other methods for enriching cell-free DNA. Therefore, those skilled in the art, with the help of embodiments of this disclosure, will recognize that several methods similar to the hybridization capture methods in the embodiments of this disclosure can alternatively be performed by replacing hybridization capture with other enrichment strategies, such as PCR amplification of cell-free DNA fragments corresponding to several genomic regions of interest.

[0091] In some embodiments, probes are used for enrichment (e.g., non-targeted enrichment), such as simplified bisulfite sequencing, methylation restriction enzyme sequencing, methylated DNA immunoprecipitation sequencing, methyl CpG binding domain protein sequencing, methyl DNA capture sequencing, or droplet PCR.

[0092] The diffuse large B-cell lymphoma assay kit provided herein is an assay kit that includes a genome hybridization probe (also referred to as a "probe" in embodiments of this disclosure), which is designed to enrich several nucleic acid fragments of interest for assay purposes.

[0093] In some embodiments, several probes are designed to hybridize and enrich DNA or cfDNA molecules from several samples that have been treated to convert unmethylated cytosine (C) to uracil (U). The probes may be designed to ligate to or hybridize to a target (complementary) strand of DNA or RNA. The target strand may be a “positive” strand (e.g., the strand transcribed into mRNA and subsequently translated into protein) or a complementary “negative” strand.

[0094] The embodiments of this disclosure are used to detect nucleic acids and determine methylation status. As for sequencing methods, the embodiments of this disclosure also include other methods for determining the methylation status of several nucleic acid sequences.

[0095] In some embodiments, a nucleic acid sample (DNA or RNA) is extracted from a subject. In the embodiments of this disclosure, DNA and RNA may be used interchangeably unless otherwise indicated. That is, the several embodiments described in the embodiments of this disclosure can be applied to both DNA and RNA-type nucleic acid sequences. However, the several examples described in the embodiments of this disclosure focus on DNA for the purposes of brevity and explanation.

[0096] In some embodiments, the sample can be any combination of human genomes, including the whole genome. The sample may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. In some embodiments, methods for obtaining a blood sample (e.g., syringe or finger prick) may be less invasive than procedures for obtaining a tissue biopsy, which may require surgery. The extracted sample may include cfDNA and / or ctDNA. In healthy individuals, the body naturally removes cfDNA and other cellular debris. If a subject has diffuse large B-cell lymphoma, the cfDNA and / or ctDNA in an extracted sample may be sufficient to detect the presence of diffuse large B-cell lymphoma.

[0097] In some embodiments, several cfDNA fragments are processed to convert unmethylated cytosine to uracil. The conversion of unmethylated cytosine to uracil is achieved using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for converting unmethylated cytosine to uracil, such as APOBEC-Seq (NEBiolabs).

[0098] In some embodiments, the detection of diffuse large B-cell lymphoma requires the preparation of a sequencing library, the specific methods of which are described in subsequent sections and will not be elaborated here.

[0099] In some embodiments, several target DNA sequences may be enriched from the sequencing library described above and used when a diffuse large B-cell lymphoma detection combo assay is performed on several samples. During enrichment, several hybridization probes (also referred to herein as “probes”) are used to identify several nucleic acid fragments that provide information about the presence or absence of diffuse large B-cell lymphoma.

[0100] In some embodiments, the hybridized nucleic acid fragments are captured and may also be amplified using PCR. The target sequences can be enriched to obtain enriched sequences, which can then be sequenced. Generally, any method known in the art can be used to isolate and enrich the target nucleic acids hybridized to the probes. For example, as is well known in the art, the biotinylate portion can be added to the 5′ end of several probes using a streptavidin-labeled surface (e.g., streptavidin-labeled magnetic beads) (i.e., biotinylation) to facilitate the isolation of several target nucleic acids hybridized to the probes.

[0101] In some embodiments, several sequence reads are generated from several enriched DNA sequences, for example, several enriched sequences. Sequence data can be obtained from several enriched DNA sequences using methods known in the art. For example, the method may include next-generation sequencing (NGS) technologies, including sequencing-while-synthesis (Illumina), pyrosequencing, Ion Torrent sequencing, single-molecule real-time sequencing (Pacific Biosciences), ligation-based sequencing (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing.

[0102] In some embodiments, several methylation state vectors are generated from several sequence reads. The sequence reads are aligned to a reference genome, which helps to indicate where the fragment of cfDNA originates from a human genome.

[0103] The following section introduces the relevant content regarding the generation of data structures.

[0104] To create a health control group data structure, the analysis system obtains information about the methylation status of several CpG sites on several sequence reads, which are derived from several DNA molecules or fragments from several healthy subjects. A methylation status vector is generated for each DNA molecule or fragment.

[0105] The analysis system subdivides the methylation state vector of each fragment into several strings representing several CpG sites. The system subdivides the methylation state vector so that the resulting strings are all shorter than a given length. For example, a methylation state vector of length 11 can be subdivided into several strings, with lengths less than or equal to 3 resulting in nine strings of length 3, ten strings of length 2, and eleven strings of length 1.

[0106] The analysis system counts the number of possible strings in the control group that have a specific CpG site as the first CpG site in the string and have a methylation state for each possible CpG site and methylation state in the vector, and records several strings.

[0107] The following section introduces the relevant content on data structure validation.

[0108] In some embodiments, once the data structure has been created, the analysis system may seek to validate the data structure and / or any downstream models that use the data structure.

[0109] The first type of validation checks the consistency of the data structure in the control group. For example, if there are any outlier objects, samples, and / or fragments in the control group, the analysis system can perform various calculations to determine whether to remove any fragments from one of these categories without affecting the purity of the control group.

[0110] The second type of validation involves examining several counts from the data structure itself (i.e., from the health control group) to check the probabilistic model used to calculate the p-values. Once the analysis system generates a p-value for several methylated state vectors in the validation group, it constructs a cumulative density function (CDF) using these p-values. The analysis system can then perform various calculations on the CDF to validate the control group's data structure.

[0111] Type III validation uses a healthy control group, separated from the control groups used to construct the data structure. Type III validation tests whether the data structure is properly constructed and whether the model functions. It can quantify the distribution of the healthy control group across the healthy control group.

[0112] The fourth type of validation is tested with several samples from the non-healthy validation group. The analysis system calculates several p-values ​​and constructs a CDF for the non-healthy validation group. For the non-healthy validation group, the analysis system expects to see, for at least some samples, the opposite of what was expected for the healthy control group and the healthy validation group in the second and third types of validation. If the fourth type of validation fails, it indicates that the model does not properly identify the anomalies that the model was designed to identify.

[0113] During the validation of the data structure, the analysis system performs the fourth type of validation test as described above. This fourth type of validation test applies a validation group, which has a composition of objects, samples, and / or fragments that are assumed to be similar to the control group. For example, if the analysis system selects several healthy subjects without diffuse large B-cell lymphoma as the control group, the analysis system will also use several healthy subjects without diffuse large B-cell lymphoma in the validation group.

[0114] The analysis system takes methylated state vectors from the validation set and performs p-value calculations for each methylated state vector from the validation set. For each possibility of a methylated state vector, the analysis system calculates probabilities from the data structure of the control set. Once several probabilities for several methylated state vectors have been calculated, the analysis system calculates a p-value score for that methylated state vector based on the calculated probabilities. The p-value score represents the predictability of finding that particular methylated state vector in the control set, as well as other possible methylated state vectors with even lower probabilities. Thus, a low p-value score generally corresponds to a methylated state vector that is less expected compared to other methylated state vectors in the control set, while a high p-value score generally corresponds to a methylated state vector that is more expected compared to other methylated state vectors found in the control set. Once the analysis system produces p-value scores for several methylated state vectors in the validation set, it constructs a cumulative density function (CDF) using the p-value scores from the validation set. The analysis system verifies the consistency of the CDF in a Type IV validation test as described above.

[0115] The following section introduces information about aberrantly methylated fragments.

[0116] In subjects with diffuse large B-cell lymphoma, several aberrant methylation fragments exhibiting abnormal methylation patterns were selected as target genomic regions. The analysis system generated several methylation state vectors from several cfDNA fragments of the sample. The analysis system processed each methylation state vector as follows.

[0117] For a given methylation state vector, the analysis system enumerates all possibilities of methylation state vectors that have the same starting CpG site and the same length (i.e., the set of CpG sites). Therefore, each methylation state can be methylated or unmethylated, with only two possible states at each CpG site, and thus the unique count of possible methylation state vectors depends on powers of 2, such that a methylation state vector of length n will be associated with 2^n possibilities of methylation state vectors.

[0118] The analysis system evaluates the health control group data structure and calculates the probability of observing each possible methylation state vector for a given identified initiation CpG site / methylation state vector length. The probability of observing a given possible vector is calculated using a Markov chain probability model to model joint probability calculations; a different method than the Markov chain probability calculation is used to determine each possible probability of observing a methylation state vector.

[0119] The analysis system uses several probabilities that can be calculated for each methylation state vector to calculate a p-value score. For example, this involves identifying probabilities that might correspond to the considered methylation state vector. Specifically, this is a probability that has the same set of CpG sites as the methylation state vector, or similarly has the same starting CpG sites and length. The analysis system sums the calculated probabilities to produce the p-value score. The calculated probabilities are several possible calculated probabilities, and several may have fewer than or equal to the identified probabilities.

[0120] This p-value represents the probability of observing a fragment's methylation state vector or other, even less likely, methylation state vector in the healthy control group. Therefore, a low p-value score roughly corresponds to a methylation state vector that is rare in healthy individuals, and causes the fragment to be labeled as aberrantly methylated relative to the healthy control group. A high p-value score is generally associated with a methylation state vector that is conceptually expected to be present in healthy subjects. For example, if the healthy control group is a non-diffuse large B-cell lymphoma group, a low p-value indicates that the fragment is aberrantly methylated relative to the non-diffuse large B-cell lymphoma group, and therefore may indicate the presence of diffuse large B-cell lymphoma in the tested subjects.

[0121] The analysis system calculates a p-score for each of several methylation state vectors, each representing a cfDNA fragment in the sample being tested. To identify which fragment is aberrantly methylated, the system can filter a set of methylation state vectors based on their p-scores. For example, filtering is performed by comparing the p-scores to a threshold and retaining only those fragments below the threshold.

[0122] The following section introduces the relevant content on calculating P-values.

[0123] To calculate the p-value score for a given detected methylation state vector, the analysis system takes the detected methylation state vector and lists several possible methylation state vectors.

[0124] The analysis system calculates several enumerated possible probabilities for a number of methylation state vectors. Because methylation conditionally depends on the methylation state of nearby CpG sites, one approach to calculating the possible probabilities of observing a given methylation state vector is to use a Markov chain model. The Markov chain model can be used to make the calculation of each possible conditional probability more efficient.

[0125] To calculate the probability of modeling each possible Markov chain for the methylation vector, the data structure of the system access control group was analyzed, particularly the counts of various strings of several CpG sites and states.

[0126] The computation can perform additional smoothing of the counts by applying a prior distribution. For example, the prior distribution could be a uniform prior, as in Laplace smoothing. Algorithmic techniques, such as Nie's smoothing, are used as an example.

[0127] Once the calculated probabilities are complete, the analysis system calculates the p-value score, adds the p-value score to the total number of probabilities, and the number of probabilities is less than or equal to the probability of matching the detected methylation state vector.

[0128] In some embodiments, the computational burden of computational rates and / or p-scores can be further reduced by caching at least some computations. For example, the analysis system can cache the computation of the probabilities of several methylation state vectors (or windows thereof) in temporary or permanent memory. If other fragments have the same CpG sites, caching the probabilities allows for the efficient computation of p-score values ​​without needing to recalculate the potential probabilities. Ultimately, the analysis system can compute a p-score for each of the probabilities of several methylation state vectors associated with a set of CpG sites from a vector (or a window thereof). The analysis system can cache p-scores for use in determining the p-scores of other fragments that include the same CpG sites. Generally, the possible p-scores of methylation state vectors with the same CpG sites can be used to determine the p-score of a possibly different one from the same set of CpG sites.

[0129] The following section introduces the relevant content of sliding windows.

[0130] In some embodiments, the analysis system uses a sliding window to determine the possibilities of the methylation state vector and calculate p-values. The analysis system lists possibilities and calculates p-values ​​only for a window of a few consecutive CpG sites, rather than for the entire methylation state vector, wherein the window is shorter in length (of CpG sites) than at least some segments (otherwise, the window would be useless). The window length can be static, user-determined, dynamic, or otherwise selected.

[0131] When calculating the p-value for a methylation state vector larger than a window, the window is started from the first CpG site in the vector, and a consecutive set of CpG sites from the vector is identified within the window. The analysis system calculates the p-value score for the window including the first CpG site. The analysis system then "slides" the window to the second CpG site in the vector and calculates another p-value score for the second window. Therefore, for a window of size l and a methylation vector length m, each methylation state vector will generate m p-value scores. After completing the p-value calculation for each part of the vector, the lowest p-value score from all sliding windows is taken as the overall p-value score for the methylation state vector.

[0132] Using a sliding window helps reduce the number of possible methylation state vectors that can be enumerated, and the corresponding probability calculations that would otherwise be required. The number of possible methylation state vectors increases exponentially with the size of the methylation state vector. An analysis system could use a window of size 5 for a fragment, resulting in 50 p-value calculations being performed for each of the 50 windows containing the methylation state vector, instead of calculating 2 p-values. 54 (approximately 1.8 × 10) 16 ) possible probabilities to produce a single p-value score. 2^50 probabilities for each enumerated methylation state vector in the calculation. 5 (32) possibilities, resulting in a total of 50 × 2 5 (1.6×10 3 This results in a significant reduction in the computations required to accurately identify anomalous fragments due to a lack of meaningful hits. This additional step can also be applied when validating control groups with several methylation state vectors in the validation group.

[0133] The analysis system identified several DNA fragments indicating diffuse large B-cell lymphoma from a filtered group of aberrant methylated fragments.

[0134] The following section introduces information about hypomethylated and hypermethylated fragments.

[0135] According to a first method, the analysis system can identify several DNA fragments considered hypomethylated or hypermethylated from a filtered set of aberrantly methylated fragments as indicators of diffuse large B-cell lymphoma. The several hypomethylated or hypermethylated fragments can be defined as several fragments of a specific length (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) with a high percentage of methylated CpG sites (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%) or a high percentage of unmethylated CpG sites (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%).

[0136] The following section introduces the selection of genomic regions that indicate diffuse large B-cell lymphoma.

[0137] The analysis system identifies several genomic regions that indicate diffuse large B-cell lymphoma. To identify these informational regions, the system calculates information gain for each genomic region, or more specifically, for each CpG site. Information gain describes the ability to distinguish between various outcomes.

[0138] A method for identifying several genomic regions distinguishable between diffuse large B-cell lymphoma and non-diffuse large B-cell lymphoma employs a trained classification model. This model can be applied to groups of aberrantly methylated DNA molecules or fragments corresponding to or derived from the diffuse and non-diffuse large B-cell lymphoma groups. The trained classification model can be trained to identify any situation of interest that can be identified from several methylation state vectors.

[0139] In some embodiments, the trained classification model is a binary classifier trained based on several cfDNA fragments or several genomic sequences obtained from a population of subjects with diffuse large B-cell lymphoma and a population of healthy subjects without diffuse large B-cell lymphoma. The binary classifier is then used to classify the probability of a subject having or not having diffuse large B-cell lymphoma based on several aberrant methylation state vectors.

[0140] In some embodiments, the classifier can be trained using several sequence reads obtained from samples of cells rich in diffuse large B-cell lymphoma. The samples come from a population of subjects known to have diffuse large B-cell lymphoma. The ability of each genomic region to distinguish between diffuse and non-diffuse large B-cell lymphoma types in the classification model is used to rank several genomic regions by classification performance, from most informative to least informative. The analysis system can identify several genomic regions from the ranking, which is based on the information gain of classification between non-diffuse and diffuse large B-cell lymphoma types.

[0141] The following section describes how to calculate the information gain of hypomethylated and hypermethylated fragments from diffuse large B-cell lymphoma.

[0142] According to some embodiments, using several fragments indicating diffuse large B-cell lymphoma, the analysis system can programmatically train a classifier. The program accesses two training groups of several samples: a non-diffuse large B-cell lymphoma group and a diffuse large B-cell lymphoma group, resulting in a non-diffuse large B-cell lymphoma group containing several aberrant methylation fragments and several methylation state vectors, and a diffuse large B-cell lymphoma group containing several methylation state vectors.

[0143] The analysis system determines whether each methylation state vector indicates diffuse large B-cell lymphoma. Here, several fragments with a certain number of CpG sites exhibiting specific states indicative of diffuse large B-cell lymphoma can be classified as hypermethylated or hypomethylated fragments. For example, if several cfDNA fragments overlap with at least four CpG sites, and at least 80%, 90%, or 100% of the CpG sites in the several cfDNA fragments are methylated, or at least 80%, 90%, or 100% of the CpG sites in the several cfDNA fragments are unmethylated, then the several cfDNA fragments are identified as hypomethylated or hypermethylated, respectively.

[0144] In other embodiments, the analysis system considers several parts of the methylation state vector and determines whether the part is hypomethylated or hypermethylated, and can distinguish whether the part is hypomethylated or hypermethylated.

[0145] In some embodiments, the procedure generates a hypomethylation fraction (P) for each CpG site in the genome. 低 ) and permethylation fraction (P 过 To generate two scores at a given CpG site, the classifier takes four counts at that CpG site: (1) counts of several vectors (methylation status) of the diffuse large B-cell lymphoma group that overlaps with the CpG site and is labeled as hypomethylated; (2) counts of several vectors of the diffuse large B-cell lymphoma group that overlaps with the CpG site and is labeled as hypermethylated; (3) counts of several vectors of the non-diffuse large B-cell lymphoma group that overlaps with the CpG site and is labeled as hypomethylated; and (4) counts of several vectors of the non-diffuse large B-cell lymphoma group that overlaps with the CpG site and is labeled as hypermethylated. Furthermore, the procedure can normalize the counts for each group to account for group size differences between the non-diffuse large B-cell lymphoma group and the diffuse large B-cell lymphoma group.

[0146] In several alternative embodiments where several segments are used to indicate diffuse large B-cell lymphoma, the several fractions can be more broadly defined as the count of several segments indicating diffuse large B-cell lymphoma at each genomic region and / or CpG site.

[0147] In one embodiment, to generate the hypomethylation fraction at a given CpG site, the procedure takes (1) above and divides it by the sum of (1) and (3). Similarly, the hypermethylation fraction is calculated by taking (2) above and dividing it by the sum of (2) and (4). The presence of hypomethylation or hypermethylation in several fragments from the diffuse large B-cell lymphoma group is determined, and the hypomethylation fraction and hypermethylation fraction are correlated with an estimate of the probability of diffuse large B-cell lymphoma.

[0148] The analysis system generates a total hypomethylation score and a total hypermethylation score for each anomalous methylation state vector. The total hypermethylation and hypomethylation scores are determined based on several hypermethylation and hypomethylation scores in the methylation state vector.

[0149] In some embodiments, the total overmethylation and low methylation scores are denoted as the maximum overmethylation and low methylation scores at several sites in each state vector, respectively. The analysis system ranks each object twice based on all methylation state vectors, with the rankings being the total low methylation score and the total overmethylation score based on several methylation state vectors. This process selects several total low methylation scores from the low methylation rankings and several total overmethylation scores from the overmethylation rankings. Based on the selected scores, the classifier generates a single feature vector for each object.

[0150] In some embodiments, several scores selected from two rankings are chosen in a fixed order, which is the same for the feature vectors of each object in several training groups. For example, the classifier selects the first, second, fourth, and eighth total overmethylated scores from each ranking, and the same for each total undermethylated score, and writes these scores into the feature vector of the object.

[0151] The analysis system trains a binary classifier to distinguish the feature vectors of the diffuse large B-cell lymphoma and non-diffuse large B-cell lymphoma training groups. Generally, any of several classification techniques can be used.

[0152] In some embodiments, the classifier is a nonlinear classifier. In a particular embodiment, the classifier is a nonlinear classifier that applies L2-normalized kernel logistic regression with a Gaussian radial basis function kernel.

[0153] In some embodiments, the number of non-diffuse large B-cell lymphoma samples (n) 其它 The number of diffuse large B-cell lymphoma samples (n) with aberrant methylation fragments overlapping with CpG sites. 阳The number of samples was counted. Next, the probability that a sample was diffuse large B-cell lymphoma was estimated by a score (“S”), which was correlated with the number of diffuse large B-cell lymphoma samples (n). 阳 ) is positively correlated with n 其它 They are negatively correlated. This score can be expressed using the equation: (n 阳 +1) / (n 阳 +n 其它 +2) or (n 阳 ) / (n 阳 +n 其它 ) is calculated.

[0154] The analysis system calculates information gain for each diffuse large B-cell lymphoma and for each genomic region or CpG site to determine whether the genomic region or CpG site indicates diffuse large B-cell lymphoma. Information gain is calculated against several training samples with a given diffuse large B-cell lymphoma, compared to all other samples. For example, a random variable “anomalous fragment” (“AF”) is used.

[0155] In some embodiments, AF, as determined by the feature vector above, is a binary variable indicating whether there are anomalous fragments in a given sample that overlap with a given CpG site.

[0156] The following describes the methods for using a combination of laboratory tests for diffuse large B-cell lymphoma.

[0157] The method may include the steps of: processing several DNA molecules or fragments to convert unmethylated cytosine to uracil; applying a diffuse large B-cell lymphoma assay to the converted DNA molecules or fragments; enriching a sub-combination of several probes hybridized (or bound to) in the assay; sequencing the enriched cfDNA fragments; detecting the nucleic acid sequence; and determining the methylation status of the nucleic acid sequence.

[0158] In some embodiments, several sequence reads can be compared with a reference genome (e.g., the human reference genome), allowing identification of the methylation status at several CpG sites in a DNA molecule or fragment, thereby providing information about the detection of diffuse large B-cell lymphoma.

[0159] The following section introduces the relevant content of sequence reading analysis.

[0160] In some embodiments, several sequence reads may be aligned to a reference genome using methods known in the art to determine alignment position information. This alignment position information may indicate the start and end positions of the start and end nucleotide bases corresponding to a given sequence read in the reference genome. The alignment position information may include the sequence read length, which may be determined from the start and end positions. Regions in the reference genome may be associated with genes or segments of genes.

[0161] In various embodiments, the sequence reads comprise a read pair denoted as R1 and R2. For example, the first read R1 may be sequenced from the first end of the nucleic acid fragment, and the second read R2 may be sequenced from the second end of the nucleic acid fragment. Therefore, several nucleotide base pairs of the first read R1 and the second read R2 may be aligned consistently (e.g., in opposite directions) with the nucleotide bases of a reference genome. Alignment information derived from the read pair R1 and R2 may include a start position in the reference genome corresponding to the end of the first read (e.g., R1) and an end position in the reference genome corresponding to the end of the second read (e.g., R2). In other words, the start and end positions in the reference genome represent possible positions of the nucleic acid fragment within the reference genome. Output files in SAM or BAM format can be generated and exported for further analysis.

[0162] From these sequence reads, the location and methylation status of each CpG site can be determined based on alignment with a reference genome. Furthermore, a methylation status vector for each fragment can be generated specifying the fragment's location in the reference genome (e.g., by the first CpG site in each fragment or other similar metrics), the number of CpG sites in the fragment, and the methylation status of each CpG site in the fragment—either methylated (e.g., denoted as M), unmethylated (e.g., denoted as U), or intermediate (e.g., denoted as I). These methylation status vectors can be stored in temporary or permanent computer memory for later use and processing.

[0163] The following section introduces relevant information on the detection of diffuse large B-cell lymphoma.

[0164] The sequence reads obtained by the methods provided in the embodiments of this disclosure can be further processed by automated algorithms. For example, an analysis system is used to receive sequence data from the sequencer and perform various aspects of the processing as described in the embodiments of this disclosure. The analysis system can be a personal computer, desktop computer, laptop computer, notebook computer, tablet PC, or mobile device. The computing device can be communicatively coupled to the sequencer via wireless, wired, or a combination of wireless and wired communication technologies. Generally, the computing device is configured to have a processor and memory storing several computer instructions. When executed by the processor, the several computer instructions instruct the processor to perform several steps as described in the embodiments of this disclosure. Generally, the amount of genetic data and data derived from genetic data is large enough, and the required computing power is large enough, that it is impossible to perform this solely on paper or by human thought.

[0165] Clinical interpretation of several methylation states across several target genomic regions is a procedure that includes classifying the clinical effects of individual or combinations of several methylation states and reporting the results in a manner meaningful to healthcare professionals. Clinical interpretation can be based on comparisons of several sequence reads with databases of specific diffuse large B-cell lymphoma or non-diffuse large B-cell lymphoma subjects, and / or on the number and type of cfDNA fragments with diffuse large B-cell lymphoma-specific methylation patterns identified from the samples.

[0166] In some embodiments, several target genomic regions are ranked based on their likelihood of differential methylation in several diffuse large B-cell lymphoma samples, and this ranking is used in the interpretation process. The ranking may include the strength of evidence for clinical efficacy. Various clinical analysis and genomic data interpretation methods can be used for the analysis of several sequence reads. In some other embodiments, the clinical interpretation of several methylation states of such differentially methylated regions can be based on machine learning, using a classification or regression method to interpret the current sample. This classification or regression method is trained using several methylation states of such differentially methylated regions from samples with known diffuse large B-cell lymphoma and non-diffuse large B-cell lymphoma.

[0167] Clinically significant information may include the presence or absence of diffuse large B-cell lymphoma.

[0168] The following section introduces the relevant content of the diffuse large B-cell lymphoma classifier.

[0169] To train a diffuse large B-cell lymphoma classifier, the analysis system obtained several training samples. For each training sample, the system determined a feature vector based on a set of hypomethylated and hypermethylated fragments indicative of diffuse large B-cell lymphoma. The system also calculated an anomaly score for each CpG site in several target genomic regions.

[0170] In some embodiments, the analysis system defines anomaly scores as a binary score for the feature vector based on the presence or absence of hypomethylated or hypermethylated fragments from CpG sites. Once all anomaly scores have been determined for the training samples, the analysis system determines the feature vector as a vector of several elements, each element including one of several anomaly scores associated with one of several CpG sites. The analysis system may normalize the anomaly scores of the feature vector based on the coverage of the samples, i.e., the median or average sequencing depth of all CpG sites.

[0171] Using several feature vectors from several training samples, the analysis system can train a diffuse large B-cell lymphoma classifier.

[0172] In some embodiments, the analysis system trains a binary diffuse large B-cell lymphoma classifier with several labels based on several feature vectors from several training samples, distinguishing between diffuse large B-cell lymphoma and non-diffuse large B-cell lymphoma. In this embodiment, the classifier outputs a prediction score, which indicates the probability of the presence or absence of diffuse large B-cell lymphoma.

[0173] The analysis system trains the diffuse large B-cell lymphoma classifier by inputting several groups of training samples, each with its own feature vector, into the classifier and adjusting several classification parameters. This allows the classifier's function to accurately associate the training feature vectors with their corresponding labels. The system can also divide several training samples into one or more groups for iterative batch training of the diffuse large B-cell lymphoma classifier.

[0174] After inputting training samples from all groups, including their training feature vectors, and adjusting several classification parameters, the diffuse large B-cell lymphoma classifier is sufficiently trained to label several detection samples based on the feature vectors of several detection samples within a certain error range.

[0175] The analysis system can train a diffuse large B-cell lymphoma classifier using any of several methods. For example, a binary diffuse large B-cell lymphoma classifier could be an L2-normalized kernel logistic regression classifier trained using a logarithmic loss function. Alternatively, a diffuse large B-cell lymphoma classifier could be a multinomial logistic regression classifier. In applications, both types of diffuse large B-cell lymphoma classifiers can be trained using other techniques. These techniques are numerous and include applications of kernel methods and machine learning algorithms such as multilayer neural networks.

[0176] During deployment, the analysis system acquires detection samples from subjects with unknown diffuse large B-cell lymphomas. The system processes these samples to obtain a set of hypomethylated and hypermethylated fragments indicative of diffuse large B-cell lymphoma. The system defines a test feature vector using a procedure similar to that described for several training samples. The system then inputs the test feature vector into a trained diffuse large B-cell lymphoma classifier to obtain a prediction of diffuse large B-cell lymphoma, including: diffuse large B-cell lymphoma or non-diffuse large B-cell lymphoma.

[0177] Embodiments of this disclosure also provide a method for enriching converted cfDNA fragments, which can provide information about diffuse large B-cell lymphoma. The method includes the steps of: contacting a composition of any of the above embodiments with DNA derived from a test subject, and enriching a sample of cfDNA corresponding to several genomic regions associated with diffuse large B-cell lymphoma by hybridization capture.

[0178] Embodiments of this disclosure also provide a method for obtaining sequence information that can provide information on the presence or absence of diffuse large B-cell lymphoma, the method comprising the steps of: a) enriching the converted DNA by contacting the converted DNA from the test subject with the composition as described in any of the above embodiments, and b) sequencing the enriched converted DNA.

[0179] Embodiments of this disclosure also provide a method for determining whether a subject has diffuse large B-cell lymphoma, the method comprising the steps of: a) capturing several cfDNA fragments from the subject using the composition as described in any of the above embodiments, b) detecting the captured several cfDNA fragments, and c) applying a trained classifier to the captured several DNA fragments to determine whether the subject has diffuse large B-cell lymphoma.

[0180] In some embodiments, the trained classifier is a hybrid model classifier. The classifier is trained on several transformed DNA sequences derived from target genomic regions selected from List 1 to 31.

[0181] In some embodiments, a trained classifier is used to determine the presence or absence of diffuse large B-cell lymphoma.

[0182] In some embodiments, as shown in Figure 1, a biomarker for early screening of diffuse large B-cell lymphoma is provided, specifically employing methylation-targeted capture technology. The general process includes: (1) clinical sample collection, where 1 mL to 5 mL of plasma is collected from the patient before treatment; (2) cfDNA (cell-free DNA) extraction; (3) denaturation, where cfDNA is denatured into a single strand, where M represents C base methylation modification; (4) double adapter ligation reaction, where the cytosine (C base) contained in the adapter sequence is methylated; (5) extension reaction, where dUTP is added to dNTP, one of the raw materials, in a certain proportion to extend and form double-stranded DNA, the extended strand of which contains random (6) Oxidation reaction: Methylated cytosine in double-stranded DNA is converted into carboxycytosine (C-Ca / g) by enzymes such as TET2; (7) UDG enzyme digestion: UDG enzyme degrades the extended chain formed by the extension reaction; (8) Deamination reaction: The unmethylated cytosine in the previous step is deaminated to form uracil by APOBEC enzyme; (9) Amplification and enrichment: A methylated library is formed; (10) Denaturation and probe hybridization reaction; (11) Avidin magnetic beads bind hybridization products; (12) Eluting magnetic bead binding products; (13) Amplification and enrichment of target regions: A final methylated library is formed; (14) Novaseq 6000 sequencing; (15) Sequencing data analysis: Data filtering, alignment, and deduplication; (16) Methylation information extraction: Differential methylation region analysis; (17) Methylation model construction.

[0183] Based on the above, this disclosure provides the following embodiments.

[0184] Example 1:

[0185] This embodiment provides a method for constructing a cfDNA methylation library and a cfDNA methylation sequencing method.

[0186] (1) Preparation of standard products: cfDNA standard products were customized from Jingliang Gene, specifically including fully methylated standard products (i.e., all C bases in the sequence of the standard products are methylated) and fully unmethylated standard products (i.e., all C bases in the sequence of the standard products are not methylated). The standard products were mixed in a certain proportion to prepare 1 μg of standard products with a 50% methylation rate. Then, the 50% methylated standard products were added to a 0.6 mL PCR tube.

[0187] (2) Denaturation: Take 1 ng, 10 ng and 50 ng of 50% methylation standard into 0.2 mL PCR tubes and label them as Sample 1, Sample 2 and Sample 3. Place Sample 1, Sample 2 and Sample 3 in the PCR instrument and incubate at 95℃ for 5 min. Then immediately place the PCR tubes on ice to cool and let stand for 2 min to allow the DNA in the standard to fully denature into single-stranded DNA.

[0188] The embodiments of this disclosure use single-stranded DNA for the ligation reaction. Both DNA double strands can be used as templates for ligation, ensuring that the single strands formed after denaturation of the cfDNA strand gap are effectively utilized, avoiding the strand loss problem caused by traditional double-strand ligation. Furthermore, end repair is not performed. End repair introduces unmethylated cytosine, leading to a lower calculated methylation rate. Therefore, the embodiments of this disclosure do not employ end repair to address the problem of 3′ end methylation distortion in DNA.

[0189] (3) Ligation reaction: Thaw the reagents shown in Table 2 below, invert and mix well, and then briefly centrifuge. Add the mixture sequentially to a PCR tube containing 16 μL of the product obtained in step (2). This step is performed on ice. Gently pipette or shake the PCR tube to mix the reaction solution, and then briefly centrifuge to bring the reaction solution to the bottom of the tube. After centrifugation, place the PCR tube in a PCR instrument and perform the ligation reaction at 37°C for 45 min. As shown in Figure 1, in this step, single-stranded DNA is ligated with the adapter sequence to obtain single-stranded DNA with a double-linked adapter sequence.

[0190] Table 2

[0191] In this context, T4 PNK refers to T4 polynucleotide kinase, PEG 8000 refers to polyethylene glycol with a molecular weight of 8000, and DTT refers to dithiothreitol with the molecular formula C4H. 10 O2S2 has a molecular weight of 154.25. ATP is an abbreviation for Adenosine triphosphate.

[0192] The adapter sequence is formed by annealing four oligonucleotides, and the oligonucleotide sequences are shown in Table 3.

[0193] Table 3

[0194] Connector sequence annealing steps:

[0195] Step 1: Resuspend the amino-modified oligonucleotides 1 to 4 in 100 μM solutions of 100 μM each.

[0196] Step 2: Prepare 100 μL of buffer solution. The solution consists of: 10 mM (Tris(hydroxymethyl)aminomethane)-HCl buffer, pH 7.5, 2 mM EDTA, and 50 mM NaCl.

[0197] Step 3: Place 10 μL of oligonucleotide 1 and oligonucleotide 2 into the first PCR tube, and 10 μL of oligonucleotide 3 and oligonucleotide 4 into the second PCR tube. Add 80 μL of buffer reagent to each of the PCR tubes, mix thoroughly, and centrifuge for 10 seconds.

[0198] Step 4: Place the PCR tubes in the PCR instrument and incubate at 95°C for 10 minutes.

[0199] Step 5: After the reaction is complete, turn off the PCR instrument and wait for the temperature to drop to room temperature before removing the PCR tube.

[0200] The product in the first PCR tube shows the first adapter sequence as shown in Figure 2, and the product in the second PCR tube shows the second adapter sequence as shown in Figure 3. Here, Pho indicates phosphate group modification, which facilitates DNA ligation; N represents a degenerate base; AmMC6:AmM indicates amino modification, and C6 represents the number of carbon atoms in the methylene straight chain between the amino group and the DNA, mainly serving to prevent adapter self-ligation; AmMC12:AmM indicates amino modification, and C6 represents the number of carbon atoms in the methylene straight chain between the amino group and the DNA, mainly serving to prevent adapter self-ligation.

[0201] (4) Extension reaction: After the previous step is completed, immediately prepare the following reaction system in the PCR tube according to Table 4 below. Gently pipette or shake the PCR tube to mix the reaction solution in the PCR tube, and then centrifuge briefly to bring the reaction solution to the bottom of the tube. After centrifugation, place the PCR tube in the PCR instrument for the extension reaction. The temperature of the hot cap is 100℃, the reaction temperature of the extension reaction is 20℃, the reaction time is 15min, and the temperature of the PCR tube is maintained at 4℃ after the reaction. As shown in Figure 1, this step extends the single-stranded DNA with the double-linker sequence to form an extension fragment. The extension fragment is then ligated to the double-linker sequence of the single-stranded DNA with the double-linker sequence to form a double-stranded DNA with the double-linker sequence.

[0202] Table 4

[0203] Here, dNTP is an abbreviation for deoxyribonucleoside triphosphate, a collective term including dATP, dGTP, dTTP, and dCTP. dUTP is an abbreviation for 2′-deoxyuridine 5′-triphosphate. Tween 20 refers to polyoxyethylene sorbitan monolaurate, the Klenow fragment, also known as the Klenow fragment, which is a large fragment of E. coli DNA polymerase I.

[0204] (5) Product purification: Add 80 μL of purified magnetic beads to the product from step (4), mix thoroughly, let stand at room temperature for 5 min, place on a magnetic rack for about 5 min to allow the magnetic beads to be completely adsorbed and the solution to become clear, and remove the supernatant; add 200 μL of freshly prepared 80% ethanol for rinsing, incubate at room temperature for 30 to 60 s, remove the supernatant, and repeat once; after the magnetic beads are dry, add 22 μL of ultrapure water for elution, let stand at room temperature for 3 min, and then place on a magnetic rack.

[0205] (6) Oxidation reaction: After thawing the reagents shown in Table 5 below, invert and mix them, and then briefly centrifuge. Then add them sequentially to a PCR tube containing 20 μL of the product obtained in step (5). This step is performed on ice. Gently pipette or shake the PCR tube to mix the reaction solution in the PCR tube, and then briefly centrifuge. Then add 5 μL of Fe(II) solution, mix thoroughly again, and briefly centrifuge. After centrifugation, place the PCR tube in a PCR instrument for oxidation reaction. The reaction temperature of the oxidation reaction is 37℃, and the reaction time is 60 min. After the reaction, place it on ice, add 1 μL of stop solution, and incubate at 37℃ for 30 min. As shown in Figure 1, in this step, the methylated cytosine in the double-stranded DNA with the double-linker sequence forms protected cytosine, resulting in double-stranded DNA with protected cytosine.

[0206] Table 5

[0207] In this embodiment, this step uses the NEB enzymatic methylation conversion module kit (catalog number: E7125).

[0208] (7) UDG Digestion Reaction: After the previous step is completed, immediately prepare the following reaction system in the PCR tube according to Table 6 below. Gently pipette or shake the PCR tube to mix the reaction solution in the PCR tube, and then centrifuge briefly to bring the reaction solution to the bottom of the tube. After centrifugation, place the PCR tube in the PCR instrument for digestion reaction; the reaction temperature of the first digestion stage is 37℃ and the reaction time is 20min; the reaction temperature of the second digestion stage is 50℃ and the reaction time is 5min. As shown in Figure 1, in this step, the extended fragment in the protected cytosine double-stranded DNA is removed, and protected cytosine single-stranded DNA is obtained.

[0209] Table 6

[0210] (8) Product purification: Add 80 μL of purified magnetic beads to the product of step (7), mix thoroughly, let stand at room temperature for 5 min, place on a magnetic rack for about 5 min to allow the magnetic beads to be completely adsorbed and the solution to become clear, remove the supernatant; add 200 μL of freshly prepared 80% ethanol for rinsing, incubate at room temperature for 30 to 60 s, remove the supernatant, and repeat once; after the magnetic beads are dry, add 18 μL of ultrapure water for elution, let stand at room temperature for 3 min, and then place on a magnetic rack.

[0211] (9) Deamination reaction: Add 4 μL of formamide to the product of step (8), mix thoroughly, centrifuge briefly, incubate at 85°C for 10 min, and immediately place on ice after the reaction is complete. Prepare the reaction system shown in Table 7 in the PCR tube; gently pipette the PCR tube to mix the reaction solution in the PCR tube, and centrifuge briefly. After centrifugation, place the PCR tube in the PCR instrument for deamination reaction. The temperature of the hot cap is 75°C, the reaction temperature of the deamination reaction is 37°C, and the reaction time is 180 min. As shown in Figure 1, in this step, the unmethylated cytosine in the protected cytosine single-stranded DNA forms uracil, resulting in single-stranded DNA with uracil.

[0212] Table 7

[0213] In this embodiment, this step uses the NEB enzymatic methylation conversion module kit (catalog number: E7125).

[0214] (10) Product purification: Add 100 μL of purified magnetic beads to the product of step (9), mix thoroughly, let stand at room temperature for 5 min, place on a magnetic rack for about 5 min to allow the magnetic beads to be completely adsorbed and the solution to become clear, remove the supernatant; add 200 μL of freshly prepared 80% ethanol for rinsing, incubate at room temperature for 30 to 60 s, remove the supernatant, and repeat once; after the magnetic beads are dry, add 22 μL of ultrapure water for elution, let stand at room temperature for 3 min, and then place on a magnetic rack.

[0215] (11) Product amplification: After thawing the Index Primer Mix and KAPA HiFi HotStart Uracil, invert and mix them thoroughly. Prepare the reaction system as shown in Table 8 in a sterile PCR tube. Gently pipette the PCR tube to mix the reaction solution in the PCR tube, and briefly centrifuge. After centrifugation, place the PCR tube in a PCR instrument for amplification reaction. The reaction conditions for the amplification reaction are shown in Table 9. As shown in Figure 1, in this step, single-stranded DNA containing uracil undergoes amplification reaction to obtain a DNA methylated library.

[0216] Table 8

[0217] Table 9

[0218] Note: The cycle number is based on the following reference: 14 cycles when the DNA input is 1 ng, 12 cycles when the input is 5 ng, and 10 cycles when the input is 10 ng.

[0219] (12) Product purification: Add 45 μL of purified magnetic beads to the product of step (11), mix thoroughly, let stand at room temperature for 5 min, place on a magnetic rack for about 5 min to allow the magnetic beads to be completely adsorbed and the solution to become clear, remove the supernatant; add 200 μL of freshly prepared 80% ethanol for rinsing, incubate at room temperature for 30 to 60 s, remove the supernatant, and repeat once; after the magnetic beads are dry, add 32 μL of ultrapure water for elution, let stand at room temperature for 3 min, and then place on a magnetic rack.

[0220] (13) Sequencing: Dilute the methylated library obtained in step (12) to 1 ng / μL, take 1 μL for detection using an Agilent 4200 Tapestation system (Agilent Technologies, USA); take another 1 μL for qPCR detection, and determine the sequencing concentration based on the detection results. Based on the concentration obtained in the previous step, dilute the library to the required sequencing concentration (2 nmol) and perform PE150 sequencing on the Illumina Novaseq sequencing platform, with each sample generating 1.5 G of data.

[0221] Example 2

[0222] Compared with Example 1, in step (4) of the extension reaction process, the ratio of dNTP to dUTP used in Example 1 is 1:1. In this example, the ratio of dNTP to dUTP is 0.25, 0.5, 2 and 4. The amount of test standard input is 10ng. The test samples are labeled as sample 4, sample 5, sample 6 and sample 7. The sequencing data of each sample is 20G.

[0223] Example 3

[0224] This embodiment employs an enzymatic double-stranded library construction method to construct a free DNA methylated library. 1 ng, 10 ng, and 50 ng of 50% methylation standard were respectively added to 0.2 mL PCR tubes and labeled as Sample 8, Sample 9, and Sample 10. The sequencing data for each sample was 20 G. The specific procedure involves directly oxidizing and protecting the methylated C bases of cfDNA using an enzyme such as TET2, followed by deamination to deaminate the unmethylated C bases into U bases. Single-end adapter ligation and primer extension were then performed to form double strands, followed by double-end adapter ligation, and subsequent amplification and enrichment to form a methylated library.

[0225] Example 4

[0226] This embodiment combines targeted capture for standard accuracy testing.

[0227] (1) Preparation of standards: cfDNA standards were custom-made from Jingliang Gene, specifically including fully methylated standards (all C bases in the sequence are methylated) and fully unmethylated standards (all C bases in the sequence are not methylated). The standards were mixed in a certain proportion to prepare 1 μg of standards with 0%, 5%, 25%, 50%, and 100% methylation rates. Subsequently, 1 ng, 5 ng, 10 ng, 20 ng, and 50 ng of the 0%, 5%, 25%, 50%, and 100% methylation rate standards were respectively added to 0.2 mL PCR tubes and labeled as sample 0%-1 ng, 0%-5 ng, 0%-1 ng, and 0%-1 ng. 0ng, 0%-20ng, 0%-50ng, 5%-1ng, 5%-5ng, 5%-10ng, 5%-20ng, 5%-50ng, 25%-1ng, 25%-5ng, 25%-10ng, 25%-20ng, 25%-50ng, 50%-1ng, 50%-5ng, 50%-10ng, 50%-20ng, 50%-50ng, 100%-1ng, 100%-5ng, 100%-10ng, 100%-20ng, 100%-50ng. If the sample is less than 40μL, make up the difference with enzyme-free water.

[0228] (2) The above standards were used to construct methylated libraries according to steps (2)-(11) of Example 1. When the amount of DNA input was 20ng, the number of amplification cycles in step (11) was 9, and when the amount of DNA input was 50ng, the number of amplification cycles was 8.

[0229] (3) Library quality control: The library concentration was determined using Qubit4.0 and the library was diluted to 1 ng / μL. 1 μL was taken out for testing using an Agilent 4200 Tapestation system (Agilent Technologies, Inc.).

[0230] (4) Preparation of hybridization reagent: Prepare the hybridization solution according to Table 10 below, labeled as Hyb-1. After preparation, mix thoroughly and set aside for use.

[0231] Table 10

[0232] The final concentration of ethylenediaminetetraacetic acid (EDTA) in SSPE is 10 mmol, and additional EDTA needs to be added to bring the final concentration to 20 mmol.

[0233] In subsequent steps, the methylated library needs to undergo high-temperature denaturation, after which single-stranded DNA molecules and probes hybridize at, for example, 60°C. If the single-stranded DNA renatures, it will affect the probe hybridization efficiency. To address this issue, this embodiment uses a single-stranded binding protein (SSB) to bind to single-stranded DNA molecules. SSB forms a tetramer that specifically binds 8–16 bases, preventing single-stranded DNA renaturation and effectively improving hybridization efficiency.

[0234] (5) Hybridization reaction: Take 187.5 ng of standard libraries with different methylation rates (0%, 5%, 25%, 50%, and 100%, 10 of each) into 1.5 mL centrifuge tubes, mix thoroughly, and centrifuge. Label them as Cap-1, Cap-2, Cap-3, Cap-4, and Cap-5. Then, thaw the probe panel, Cot1 DNA, and Blocker Solution, mix them by inversion, and prepare the reagent system as shown in Table 11 in a sterile PCR tube. As shown in Figure 1, in this step, the probe hybridizes to several nucleic acid fragments (i.e., target genomic regions).

[0235] Table 11

[0236] Cot1 DNA is a type of placental DNA, mainly ranging in size from 50bp to 30bp, and is rich in repetitive DNA sequences. It can effectively block repetitive DNA sequences in the target region and reduce non-specific hybridization.

[0237] Mix the reagent system in the sterile PCR tube thoroughly and centrifuge briefly. Place the sterile PCR tube in a vacuum concentrator and concentrate it to dry powder. Then prepare the reagent system as shown in Table 12 in this centrifuge tube.

[0238] Table 12

[0239] Gently pipette to mix and briefly centrifuge, then place in a PCR instrument for hybridization reaction. The hybridization reaction conditions are shown in Table 13.

[0240] Table 13

[0241] (6) Streptavidin magnetic bead cleaning:

[0242] 1) 40 minutes before the previous reaction is complete, remove the streptavidin magnetic beads from 4°C and allow them to equilibrate at room temperature for 30 minutes.

[0243] 2) Pipette 100 μL of magnetic beads into a 1.5 mL low-adsorption centrifuge tube and clean the magnetic beads with Beads Binding buffer;

[0244] 3) Add 200 μL of Beads Binding Buffer to the centrifuge tube, gently pipette to mix 10 times, centrifuge briefly, place on a magnetic rack for several minutes until the liquid is completely clear, discard the supernatant with a pipette, and remove the centrifuge tube from the magnetic rack.

[0245] 4) Repeat step 3) twice;

[0246] 5) Add 200 μL of magnetic bead suspension to the centrifuge tube, gently mix by blowing and aspirating, and transfer all the magnetic bead suspension to a new 1.5 mL low-adsorption PCR tube.

[0247] (7) Preparation of eluent: Prepare eluents WB1 and WB2 according to Tables 14 and 15 below.

[0248] Table 14

[0249] Table 15

[0250] After methylation, the proportion of A and T bases in the DNA strands of the methylated library increases. Tetramethylammonium chloride in the elution buffer can improve the solubility (TM) of A and T-rich DNA strands, increasing the TM value of A and T-rich regions within the DNA strand, making it closer to the TM value of DNA strands with normal A and T base proportions. Gradient concentration tests of tetramethylammonium chloride (0.1M, 0.5M, 1M, 1.5M, and 2M) showed that a concentration of 1M was optimal. During the elution step, the elution buffer was incubated at 48°C. This effectively washed away non-specific hybridization products while retaining specific hybridization products rich in A and T bases.

[0251] Formamide can reduce the TM value of DNA strands. Each 1% increase in formamide can reduce the TM value by about 0.7°C. An elution buffer containing 5% formamide, after incubation at 48°C, can maximize the specificity of hybridization and minimize the loss of hybridization products in the target region.

[0252] (8) Hybridization product elution: 200 μL of magnetic bead suspension was thoroughly mixed with the hybridization product and incubated at room temperature for 30 min. Then, the mixture was washed with elution buffer WB1 at 65 °C and incubated at 65 °C for 5 min, repeated 3 times. Finally, the mixture was washed with elution buffer WB2 and incubated at 48 °C for 5 min, repeated 3 times. As shown in Figure 1, in this step, several nucleic acid fragments (i.e., target genomic regions) that hybridized with the probe were pulled down and enriched.

[0253] (9) Amplification of elution products: After thawing the amplification primers and Kapa Hifi hotstart ready Mix (kk2601), mix them by inversion and prepare the reagent system as shown in Table 16 in a sterile PCR tube.

[0254] Table 16

[0255] Gently pipette to mix and briefly centrifuge, then place in a PCR instrument for amplification. The reaction conditions for the amplification reaction are shown in Table 17.

[0256] Table 17

[0257] (10) Product purification: Add 90 μL of purified magnetic beads to step (9), mix thoroughly, let stand at room temperature for 5 min, place on a magnetic rack for about 5 min to allow the magnetic beads to be completely adsorbed and the solution to become clear, remove the supernatant; add 200 μL of freshly prepared 80% ethanol for rinsing, incubate at room temperature for 30 s to 60 s, remove the supernatant, and repeat once; after the magnetic beads are dry, add 32 μL of ultrapure water for elution, let stand at room temperature for 3 min, and then place on a magnetic rack.

[0258] (11) Sequencing: Dilute the methylated library obtained in step (10) to 1 ng / μL, and take 1 μL for detection using an Agilent 4200 Tapestation system (Agilent Technologies, USA); take another 1 μL for qPCR detection, and determine the concentration for sequencing based on the detection results. Based on the concentration obtained in the previous step, dilute the library to the required level (2 nmol) and perform PE150 sequencing on the Illumina Novaseq sequencing platform, with each sample generating 20 G of data.

[0259] (12) Data quality control: Filter low-quality sequences and sequencing adapter sequences to obtain high-quality data and generate corresponding quality control reports.

[0260] (13) Genome alignment and deduplication: High-quality data are compared with the reference genome to find the position of each read on the reference genome, and repetitive sequences introduced by PCR amplification are removed. Sequencing depth and sequencing coverage are statistically analyzed.

[0261] (14) Methylation information extraction: After obtaining the deduplication comparison results, methylation site detection is performed to obtain the methylation level under different sequence environments (CG, CHH, CHG, where H represents A, C, T).

[0262] Example 5

[0263] This embodiment is used to obtain a detection kit for diffuse large B-cell lymphoma.

[0264] (1) Clinical sample collection:

[0265] 1) Collect whole blood (2mL~8mL) into an 8.5mL Roche cfDNA free nucleic acid collection tube (please operate at room temperature, store at room temperature (18-25℃) after blood collection, and perform subsequent plasma extraction operations within a maximum of 72 hours);

[0266] 2) Centrifuge for the first time at 1350g / min at 4℃ for 12min. Carefully remove the pale yellow supernatant (avoid contamination of the white blood cell layer) and transfer it to a 2mL DNase-free sterile centrifuge tube.

[0267] 3) Centrifuge a second time at 13500g / min at 4℃ for 5min. Carefully remove the supernatant (to completely remove white blood cells) and transfer it to 2-3 2mL DNase-free sterile centrifuge tubes. Store at -80℃ (approximately 4mL-6mL of clean plasma should be obtained. The color should be used to determine if hemolysis has occurred and to implement risk control measures, or the hospital should be notified to collect samples again according to the standard procedure). The white precipitate at the bottom of the tube is white blood cells. Mark the tube and store it at -80℃.

[0268] 4) After writing the obtained plasma number, freeze it at -80℃ for later use. During this process, also enter the accurate electronic plasma separation table for future reference.

[0269] (2)Use The ccfDNA (catalog number 55204) kit extracts cfDNA from plasma samples.

[0270] (3) cfDNA fragment analysis: 1 μL was taken for detection using an Agilent 4200 Tapestry system (Agilent Technologies, USA). Figure 4 shows the cfDNA fragment length distribution in the sample. The horizontal axis represents the length of the cfDNA fragment in bp, and the vertical axis represents the normalized fluorescence intensity. "Maximum" and "minimum" are reference values ​​for cfDNA fragment length; "maximum" refers to a fragment length of 1000 bp, and "minimum" refers to a fragment length of 15 bp. When the cfDNA fragment length is within the range of 100 bp to 700 bp, the cfDNA fragment length distribution is considered normal. As shown in Figure 4, the main peaks of the normalized fluorescence intensity include a 183 bp peak and a 371 bp peak, indicating that the cfDNA fragment length distribution meets the requirements for subsequent library preparation.

[0271] (4) The whole-genome methylation library was constructed / targeted capture / sequencing was performed according to the experimental method in Example 4. Figure 5 shows the DNA length distribution in the gene library, where "maximum" and "minimum" are reference values ​​for cfDNA fragment length. "Maximum" means the cfDNA fragment length is 1000 bp, and "minimum" means the cfDNA fragment length is 15 bp. When the DNA length is within the range of 300 bp to 350 bp and 450 bp to 530 bp, it indicates that the DNA length distribution is normal. As shown in Figure 5, the normalized fluorescence intensity includes the main peaks at 339 bp and 514 bp, indicating that the length distribution of the clinical sample DNA library meets the requirements.

[0272] (5) As shown in Figure 1, the dataset was split into a training set and a test set in an 8:2 ratio. DMR identification was performed in the training set. The minimum number of CpG sites covered by each DMR region was set to 4, and the significance q-value (q value) and the average methylation rate difference were set to 20%, respectively. A total of 238 DMRs were obtained, which were associated with 191 protein-coding genes. The methylation data of the 238 DMR regions were used as features, and the feature recursive elimination algorithm RFECV was used for feature selection to obtain the optimal candidate features DMRs, a total of 31, as shown in Table 1 (lists 1 to 31). Common classifier algorithms were used to build the model. Under the condition that the model meets certain performance indicators (AUC, specificity, sensitivity), the model candidate features were used as potential biomarkers in the early screening scenario of diffuse large B-cell lymphoma.

[0273] The 191 protein-coding genes are listed below: TP73, SPSB1, KAZN, PADI1, IFFO2, MAN1C1, TRNP1, RRAGC, BTBD19, PLPP3, FGGY, SLC44A3, MAGI3, MAB21L3, PDE4DIP, SMCP, KIRREL1, NAV1, PTPN7, PPFIA4, NFASC, SLC26A9, RHEX, PLXNA2, FAM89A, CHRM3, KLHL29, MRPL33, CAPN13, EHD3, SPRED2, TGFA, SPR, STAMBP, DCTN1, TRABD2A, SEM A4C, VWA3B, NPAS2, MAP4K4, MFSD9, ANAPC1, HS6ST1, NRP2, TNP1, TNS1, SP140, HDAC4, NUP210, DCLK3, EXOG, CTNNB1, POMGNT2, CACNA2D2, LRTM1, KALRN, CFAP100, KY, DZIP1L, PXYLP1, XXYLT1, MUC4, SLC2A9, ZNF518B, C1QTNF7, RBPJ, APBB2, LIMCH1, STOX2, FAT1, SLC1A3, PTCD2, CEP120, ZNF608, POU4F3, AD RA1B, CDYL, ATXN1, FGD2, ZFAND3, PAQR8, TRAM2, RPS6KA2, TARP, DDC, GRB10, GSAP, CLDN15, CDHR3, AKR1B10, CNTNAP2, VIPR2, ERICH1, CHD7, ARFGEF1, K LF10, TSNARE1, PLEC, TRPM3, KLF4, GARNL3, AKR1C4, NPY4R, TBATA, OIT3, PRXL2A, MYOF, ENTPD1, ​​PIK3AP1, INSYN2A, TUBGCP2, DENND2B, SOX6, KCNJ11, CB LIF, PPFIA1, P2RY2, KCTD21, HTR3B, NNMT, DSCAML1, GRIK4, SORL1, ETS1, BARX2, ADAMTS8, WNT5B, RHNO1, NTF3, ANO2, APOBEC1, PRH1, PLEKHA5, NCKAP5L, KRT74, OR10P1, R3HDM2, RASSF3, CPSF6, PLXNC1, UBE3B, ANAPC7, DHRS12, MYO16, RNASE13, DAD1, ZFHX2, SPTSSA, CNIH1, TTC9, TTC7B, MOK, CDCA4, ITPKA,MYO5C, TPM1, FURIN, SRL, SHISA9, SLC6A2, MT1G, WWOX, ABR, TVP23C, CENPV, LLGL1, RAB11FIP4, LHX1, ETV4, MPP2, ABCC3, MRPS23, TBCD, SMAD 7. CTDP1, SBNO2, LRRC8E, SIN3B, CARD8, PPP1R15A, EIF2S2, ATP9A, PMEPA1, SAMD10, PDXK, DNMT3L, TSPEAR, PCBP3, PRAME, MYO18B, KLHDC7B. ,

[0274] (7) Input the detected methylation rate value into the logistic regression model in the validation set. The model outputs the predicted probability value of treatment response. Use 0.5 as the threshold. If the probability value is greater than 0.5, the treatment is effective. If it is less than 0.5, the treatment is ineffective.

[0275] Figure 6 shows the receiver operating characteristic (ROC) curve. As can be seen from Figure 6, the test set evaluation results show that the model's AUC (area under the receiver operating characteristic curve) value is 0.942. The closer the AUC is to 1, the better the model's evaluation effect on diffuse large B-cell lymphoma.

[0276] Figure 7 shows the sensitivity and specificity of the clinical samples in the test set. As can be seen from Figure 7, the gene marker (target genomic region) provided in this embodiment has a sensitivity of 85% and a specificity of 94% for the detection of diffuse large B-cell lymphoma.

[0277] In the embodiments of this disclosure, methylation library construction technology was used in combination with target region targeted capture for high-throughput sequencing. After the methodology was validated for accuracy using standard samples, it was tested using 214 clinical cohort samples, including 117 healthy control samples and 97 DLBCL patients. Methylomics characteristics of healthy individuals and patients were identified, and a model was constructed using common classifier algorithms to achieve accurate diagnosis of early diffuse large B-cell lymphoma.

[0278] The test data of the average methylation rate of samples 1, 2, 3, 4, 5, 6, 7, 8, 9 and 10 provided in the above embodiments are shown in Table 18.

[0279] Table 18

[0280] As shown in Table 18, the methylation rates of different fragments in samples 1, 2, and 3 are all close to the theoretical value of 50%. Compared with the traditional double-stranded enzyme library construction method (Example 3), this demonstrates that the cfDNA methylation library construction method provided in this disclosure can maintain the original methylation status of cfDNA and ensure the accuracy of methylation rate detection. Moreover, when the molar ratio of dNTPs to dUTPs in the extension reaction is in the range of 0.5 to 2 (samples 5 and 6), the accuracy of methylation rate detection is high.

[0281] Figure 8 compares the methylation rates of bases 1-150 (5′-3′ orientation) in Read2 sequences of libraries from Examples 1 and 3. The horizontal axis represents the base sequence number, i.e., the base read position, and the vertical axis represents the methylation rate at that position. As can be seen from Figure 8, the methylation rate of Read2 in Library 3 from positions 80-150 is significantly lower than the theoretical value of 50%, indicating a decrease in methylation rate. In contrast, the methylation rates of bases 1-150 in Read2 of Library 1 from Example 1 are close to the theoretical value of 50%. This demonstrates that the methylated library construction method provided in this disclosure can effectively solve the problem of reduced cfDNA end methylation detection and improve the accuracy of methylation rate detection.

[0282] Figure 9 shows the results of the detected and theoretical values ​​of methylation rates at different sites in the target genome region. The horizontal axis represents the sample input amount, and the vertical axis represents different methylation rates. As can be seen from Figure 9, the detected and theoretical values ​​of methylation rates at different sites in the target genome region are basically consistent. This further illustrates that the methylation rate detection of the gene library obtained by the library construction method and target gene capture method provided in the embodiments of this disclosure has high accuracy.

[0283] Figure 10 shows the correlation analysis of linear regression between the detected methylation rate and the theoretical methylation rate at different sites in the target genome region. The horizontal axis represents the detected methylation rate, and the vertical axis represents the theoretical methylation rate. As can be seen from Figure 10, the linear correlation coefficient R0... 2 =0.95, indicating that the detected methylation rate is basically consistent with the theoretical methylation rate; P value <0.01, where P value is called significance value, indicating that the gene library obtained by the library construction method and target gene capture method provided in the embodiments of this disclosure has high accuracy; gene libraries with sample amounts of 1ng to 50ng can effectively detect methylation rates of 0% to 100%, indicating that the technical method is compatible with the detection of trace amounts of cfDNA.

[0284] In summary, the embodiments of this disclosure employ methylation-targeted capture technology. By comparing the differentially methylated regions between the healthy control group and the patient group, a set of gene markers (target genomic regions) are screened out. Through validation with clinical cohort samples, the AUC is 0.942, the sensitivity is 85%, and the specificity is 94% in the test set, achieving accurate diagnosis of early diffuse large B-cell lymphoma.

[0285] Furthermore, the embodiments of this disclosure do not use bisulfite to methylate cfDNA, and the reaction conditions are mild, thus solving the problems of excessive degradation, damage and loss of cfDNA.

[0286] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A composition comprising: Several different bait oligonucleotides are configured to collectively hybridize to DNA molecules derived from multiple target genomic regions; In this context, each of the multiple target genomic regions is differentially methylated in diffuse large B-cell lymphoma compared to non-diffuse large B-cell lymphoma.

2. The composition according to claim 1, wherein, The several different bait oligonucleotides are configured to hybridize to several DNA molecules, which are derived from at least 20%, at least 25%, or at least 50% of the several target genomic regions of any of Lists 1 to 31.

3. The composition according to claim 1 or 2, wherein, The several different bait oligonucleotides are configured to hybridize to several DNA molecules, which are derived from at least 20%, at least 25%, or at least 50% of the several target genomic regions listed in Lists 1 to 31.

4. The composition according to claim 1, wherein, The several different bait oligonucleotides are configured to hybridize to several DNA molecules, which are derived from at least 20% of the several target genomic regions listed in Lists 1 to 31.

5. The composition according to claim 4, wherein, The plurality of DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions listed in 1 to 31.

6. A composition comprising: Several different decoy oligonucleotides are configured to hybridize to several DNA molecules, said DNA molecules being at least 20% derived from any of the target genomic regions listed in 1 to 31.

7. The composition according to claim 6, wherein, The several different bait oligonucleotides are configured to hybridize to several DNA molecules, which are derived from 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the several target genomic regions of any of Lists 1 to 31.

8. The composition according to claim 6 or 7, wherein, The several different bait oligonucleotides are configured to hybridize to several DNA molecules, which are derived from at least 20% of the several target genomic regions listed in Lists 1 to 31.

9. The composition according to claim 8, wherein, The plurality of DNA molecules are derived from at least 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions listed in 1 to 31.

10. The composition according to any one of claims 1 to 9, wherein, The aforementioned DNA molecules are converted cfDNA fragments.

11. The composition according to claim 10, wherein, The target genomic regions are hypermethylated regions, hypomethylated regions, or binary regions that are either hypermethylated or hypomethylated.

12. The composition according to claim 10, wherein, The decoy oligonucleotides are configured to hybridize to overmethylated converted DNA molecules, hypomethylated converted DNA molecules, or both overmethylated and hypomethylated converted DNA molecules derived from each target genomic region.

13. The composition according to any one of claims 1 to 12, wherein, Each of the several bait oligonucleotides is conjugated to an affinity moiety.

14. The composition according to claim 13, wherein, Each of the several different bait oligonucleotides is bound to the surface of the magnetic bead.

15. A method for enriching converted cfDNA fragments, said converted cfDNA fragments providing information on diffuse large B-cell lymphoma, said method comprising the steps of: The composition of any one of claims 1 to 14 is contacted with DNA derived from the test subject, and samples of cfDNA corresponding to several genomic regions associated with diffuse large B-cell lymphoma are enriched by hybridization capture.

16. A method for obtaining sequence information, said sequence information providing information on the presence or absence of diffuse large B-cell lymphoma, said method comprising the steps of: a) Enriching the converted DNA by contacting it with the composition according to any one of claims 1 to 14, and b) Sequencing the enriched converted DNA.

17. A method for determining whether a subject has diffuse large B-cell lymphoma, comprising the steps of: a) Capture several cfDNA fragments from the target subject using the composition according to any one of claims 1 to 14. b) Detect several captured cfDNA fragments, and c) The trained classifier is applied to several captured DNA fragments to determine whether the subject has diffuse large B-cell lymphoma.

18. The method for determining whether a test subject has diffuse large B-cell lymphoma according to claim 17, wherein, The trained classifier determines the presence or absence of diffuse large B-cell lymphoma.

19. The method for determining whether a test subject has diffuse large B-cell lymphoma according to claim 17 or 18, wherein, The trained classifier is a hybrid model classifier.

20. The method for determining whether a test subject has diffuse large B-cell lymphoma according to any one of claims 17 to 19, wherein, The classifier is trained on a plurality of converted DNA sequences derived from target genomic regions selected from any of Lists 1 to 31.

Citation Information

Patent Citations

  • Detecting cancer, cancer tissue of origin, and / or cancer cell type

    CN113728115A

  • Detecting cancer, cancer tissue of origin, and / or a cancer cell type

    CN114026254A

  • Detection of non-hodgkin lymphoma

    CN116670299A

  • Prognostic methods for diffuse large b-cell lymphoma

    WO2015154018A1

  • Systems and methods for cell-free nucleic acids methylation assessment

    WO2024124207A2