A methylation sequencing method and apparatus
The MERMAID methylation detection method, which utilizes machine learning to identify comethylated blocks, solves the sensitivity and impairment problems of existing blood DNA methylation analysis techniques, enabling reliable detection of early-stage cancer.
Patent Information
- Application Number
- CN202210629330.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-03
- Filing Date
- 2022-06-02
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-06-02
AI Technical Summary
Existing blood-based DNA methylation analysis techniques suffer from limited sensitivity and DNA damage, limiting their application in early cancer detection.
A MERMAID methylation detection method was developed, which utilizes a machine learning classifier to achieve efficient detection of cfDNA by identifying comethylated blocks and analyzing the correlation coefficient, methylation level, and location information of CpG sites.
It improves the detection capability of low-frequency ctDNA, reduces sequencing artifacts, and can provide reliable early cancer diagnosis without biopsy. It is also suitable for methylation analysis of blood samples.
Smart Images

Figure QLYQS_1 
Figure QLYQS_3 
Figure QLYQS_5
Abstract
Description
Technical Field
[0001] This application relates to the biomedical field, specifically to a methylation sequencing method and apparatus. Background Technology
[0002] In 2018, cancer caused 9.6 million deaths worldwide, the majority of which were diagnosed at an advanced stage. To date, intervention before distant metastasis offers the greatest opportunity to improve prognosis, making the development of sensitive, reliable, and minimally invasive tests to detect cancer before symptoms appear highly desirable. Unfortunately, serum-based biomarkers are often limited to cancer surveillance (e.g., carbohydrate antigen 19-9) and are not sensitive and / or specific enough for screening purposes.
[0003] Cell-free DNA (cfDNA) refers to degraded DNA fragments in the bloodstream, most of which originate from normal white blood cells (WBCs). In cancer patients, a portion of cfDNA is tumor-derived (ctDNA), providing a real-time snapshot of the cancer genome. Mutation characterization of ctDNA has achieved promising results for cancer diagnosis, prognosis, and monitoring. However, these methods are largely limited to advanced-stage patients due to limited sensitivity or the need for biopsies to guide downstream analysis.
[0004] With the rapid development of next-generation sequencing (NGS) technology, epigenetic analysis has gained significant attention. Compared to searching for rare somatic mutations in the blood, methylome analysis has shown several advantages: (i) aberrant methylation often occurs extensively during the onset of cancer, thus reflecting early changes in tumors; (ii) methylation modifications are frequently found in specific genomic regions such as “CpG islands,” providing a great opportunity to analyze large numbers of alterations through targeted sequencing; and (iii) methylation status is cell-type specific, and therefore can be used to infer the tissue origin of ctDNA.
[0005] Currently, NGS-based DNA methylation analysis techniques can be categorized into two types: bisulfite conversion-based methods (WGBS, RRBS) and enrichment-based methods (MeDIP, MBD-seq). Among these, bisulfite sequencing (BS-seq) is considered the gold standard for DNA methylation analysis because it provides quantification at single-base resolution. Unfortunately, the harsh conditions of bisulfite conversion (BC) can cause significant damage to DNA, limiting its use in blood-based applications. Furthermore, the converted DNA typically exhibits poor sequence diversity, leading to issues such as biased target enrichment and high sequencing errors, further complicating the analysis.
[0006] To overcome these problems, this application develops the MERMAID methylation detection method, which is based on a sequencing method that maximizes the use of cfDNA and significantly reduces artifacts in methylation sequencing. With the assistance of a robust machine learning classifier, MERMAID significantly outperformed other non-biopsy methods in a proof-of-concept study of low-frequency ctDNA detection using ELSA-seq (see international patent publications WO2019 / 191900A1 and WO2019 / 192489A1), providing new opportunities for advancing clinical applications. Summary of the Invention
[0007] To overcome these challenges, this application developed the MERMAID sequencing method. With the assistance of a robust machine learning classifier, MERMAID significantly outperformed other biopsy-free methods in a proof-of-concept study for low-frequency ctDNA detection, providing new opportunities for advancing clinical applications.
[0008] On the one hand, this application provides a method for detecting the methylation modification of a target nucleic acid, the method comprising the following steps: step (a-1) determining a comethylation block based on the correlation coefficient of CpG sites in the target nucleic acid, the methylation level of the CpG sites, and the location information of the CpG sites; and / or step (a-2) determining the comethylation block based on the correlation coefficient of CpG sites in the target nucleic acid, the information content of candidate comethylation blocks, and the degree of balance in the division of the candidate comethylation blocks; and step (b) determining the presence and / or content of the target nucleic acid in the sample to be tested based on the methylation degree of the comethylation block.
[0009] On the other hand, this application provides an analytical device for detecting methylation modification of a target nucleic acid, the device comprising: a block partitioning module (a-1) for determining comethylated blocks based on the correlation coefficient of CpG sites in the target nucleic acid, the methylation level of the CpG sites, and the position information of the CpG sites; and / or a block partitioning module (a-2) for determining the comethylated blocks based on the corrected correlation coefficient of CpG sites in the target nucleic acid, the information content of candidate comethylated blocks, and the partitioning balance of the candidate comethylated blocks; and a determination module (b) for determining the presence and / or content of the target nucleic acid in the sample to be tested based on the methylation degree of the comethylated blocks.
[0010] Methylation sequencing has attracted considerable interest due to its significant potential to improve current ctDNA detection. This application presents MERMAID as a novel epigenetic analysis method characterized by well-preserved molecular diversity, strong noise suppression, and robust high-dimensional modeling. In addition to these properties, MERMAID may be particularly useful for blood-based applications: (i) a subset of cfDNA can be in single-stranded form, thus this ssDNA-compatible approach maximizes the use of limited starting materials, increasing the chances of detecting rare ctDNA. (ii) the capture group is designed with an excess of long RNA probes (>100 nucleotides long) complementary to various methylation patterns. This strategy is more tolerant of sequence-related biases and polymorphisms compared to amplicon-based target methods (~20 nucleotides long). (iii) MERMAID does not require prior knowledge of the analysis (e.g., tissue examined via biopsy), thus providing a solution for patients without surgically removed samples. While this method has only been validated on LC, it can be customized for other types of cancer (e.g., CRC) or bodily fluids (e.g., urine). It can be extended to answer basic questions, such as tumor heterogeneity, or applied to other clinical scenarios, such as evaluating treatment effectiveness.
[0011] Other aspects and advantages of this application will readily be apparent to those skilled in the art from the detailed description below. Only exemplary embodiments of this application are shown and described in the following detailed description. As will be appreciated by those skilled in the art, the content of this application enables them to make modifications to the disclosed specific embodiments without departing from the spirit and scope of the invention to which this application pertains. Accordingly, the descriptions in the accompanying drawings and specification of this application are merely exemplary and not restrictive. Attached Figure Description
[0012] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention can be better understood by referring to the exemplary embodiments and drawings described in detail below. A brief description of the drawings is as follows:
[0013] Figure 1 The diagram illustrates an overview of the MERMAID method of this application, including (D) to (E) in (A) to (E), wherein,
[0014] (A) Schematic diagram of the sample preparation procedure.
[0015] A schematic diagram of the (BC) library construction and targeted sequencing workflow. First, the cfDNA fragment is denatured and transformed with sodium bisulfite (lightning). Then, a "tailing" (Tail & Tag-1) step is performed via TdT (terminal deoxynucleotidyl transferase) to add an additional nucleotide (primarily dC, dark purple, right end directly below Tail & Tag-1) to the 3' end of the template strand (gray, between BC and Tail & Tag-1). Splint adapter 1 (pink and yellow, right end complementary double strand of the small fragment below Tail & Tag-1) facilitates the ligation process in the presence of E. coli ligase (dotted arc) via a 5' protruding "arm" (primarily dG, light purple, left end protrusion of the small fragment below Tail & Tag-1). A copy of the original template was then generated from the common anchoring site (yellow, the right end of the two strands below Pre-amp and Tail & Tag-2) by uracil-resistant polymerase, followed by ligation mediated by adapter 2 (green and blue, the left end of the complementary double strand of the small fragment below Tail & Tag-2). A fully methylated genome sequencing library was generated by PCR amplification (PCR-1), and a dual-indexing (barcoding) system was performed to distinguish up to 384 samples. A library of biotinylated RNA probes (purple, the dark gray short fragment directly below the strand of Hyb & Cap) and streptoacidin beads (solid gray, dendritic spheres above PCR-2) were used to pull down both strands of the DNA target, and the captured molecules were then subjected to PCR amplification (PCR-2) and NGS sequencing.
[0016] (D) Schematic diagram of noise suppression driven by deep sequencing and single-molecule methylation pattern recognition.
[0017] (E) Schematic diagram of machine learning-assisted classification.
[0018] Figure 2 The illustration shows a comparison of different WGBS experimental protocols using an exemplary WGBS library construction method of this application. Partially methylated *E. coli* (DH5α) DNA was cut to ~200 bp to mimic cfDNA. A template-free control (NTC) was performed to evaluate background readings (e.g., residual adapter dimers) for each experimental protocol, which were then subtracted from the library yield measured in each experiment. ELSA: ELSA-seq, an exemplary sequencing method of this application; SWT: Accel-NGS Methyl-SEQ; NEB: NEBNext ultra II. Unless otherwise noted, experiments were performed with two technical replicates and two biological replicates, and error bars represent one SD. Figure 2 Including (A) to (F), where,
[0019] (A) Final library yields obtained for ELSA, SWT, and NEB using 12, 12, and 16 PCR cycles, respectively. Note: 12 NEB cycles did not provide sufficient DNA for quantification.
[0020] (B) Unique read depths (excluding PCR duplicates) were observed with different input amounts (0.5, 1, 5 ng) when sequencing reached a median depth of ~2,000X.
[0021] (C) Unique read depths observed by constructing shallow sequencing libraries at different depths using 500 pg of E. coli DNA. The DH5α genome size is approximately 4.5 Mbp, and 0.5 ng refers to approximately 100,000 haploid genome copies prior to BC. Normalization was performed by subtracting the unique read depths observed in NTC.
[0022] (D) A graph showing the overlap (depth > 0) of detected methylation sites between 500 pg and 5 ng DNA inputs.
[0023] (EF) Mutation spectra and frequencies observed without (E) or with (F) depth sequencing-induced error suppression. The same libraries were sequenced on Illumina Hiseq 2500, Novaseq 6000, and MiniSeq, respectively.
[0024] Figure 3 The figures illustrate the design and performance of an exemplary target group of this application, including (A) through (E), wherein,
[0025] (A) Heatmap of methylation levels in 2,765 probe regions (Illumina 450K) of WBC (n=656), cancerous tissue (n=4,539), and normal tissue (n=521) based on GSE and TCGA databases.
[0026] (BD) Performance of an exemplary target group of this application using 10 ng cfDNA from a representative donor, wherein,
[0027] (B) The only depth across all genomic regions in the target group.
[0028] (C) Fragment size obtained using an exemplary library construction method of this application or the Trussq library construction method.
[0029] (D) Methylation level of each CpG site (iAF) calculated using the "+" or "-" chain.
[0030] (E) Panel-wide technical background for assessment using CHH methylation levels.
[0031] Figure 4 The diagram illustrates the definition and pattern recognition of methylated blocks, including (A) through (D), where,
[0032] (A) Representative genomic regions (CHR14: 26674253-26674534) showing comethylation status as measured by Poisson correlation (left panel) or “block index” (right panel).
[0033] (B) Box plot showing the AUC statistical distribution for classifying malignant (n=48) and normal lung tissue (n=20) using methylation status (mAF or iAF) of SHOX2 (CHR3: 157813800-157822158). All samples were sequenced at unique read depths of 500X (left panel) and 100X (right panel). Unless otherwise noted, plots are marked with one asterisk (*), two asterisks (**) for P<0.05, three asterisks (***) for P<0.01, and three asterisks (***) for P<0.001. In this plot, P-values were calculated using the Wilcoxon Sign test.
[0034] (C) The formula for MBS and examples of genomic loci with various methylation patterns.
[0035] (D) Density curves of mAF (left panel) and MBS (right panel) observed in a λ-spiking-in experiment treated with methyltransferase. Colored lines represent samples with different dilution ratios (technically repeated with identical coloring), while the X-axis shows the logarithmically transformed mAF or MBS values.
[0036] Figure 5 The diagram illustrates the analytical validation of MERMAID, including (A) through (F), where,
[0037] (A) Poisson correlation (ρ) matrix of MBS values among different replicates, individuals, and diseases. Data were generated from 4 different healthy donors (H1–4) and 3 LC patients (C1–3), with three replicates for H1 (H1R1–R3). Computation was performed across the entire group using 8,312 blocks.
[0038] (B) Detection of methylation signals using serial dilutions (0.0001–0.05) of CRC cell DNA. Error bars depict the 95% CI of the estimated tumor fraction, while dotted lines represent y = x.
[0039] (C) FDR measured by sequencing normal leukocytes at different depths and with different input volumes (X-axis). The Y-axis represents the percentage of false positive calls.
[0040] (D) Bioinformatics (in sillico) sensitivity measured by mixing sequencing reads from CRC into normal cfDNA data at different proportions (0.00001–0.05). The X-axis represents the expected tumor score. The Y-axis represents the percentage of positive calls. P-values were calculated using a t-test.
[0041] (EF) Sensitivity of MERMAID (E), ddPCR (F, left panel), and HS-UMI (right panel) assays measured by sequentially diluting cancer cell (LC, CRC) DNA with normal leukocytes at ratios of 0.0001–0.1. The Y-axis represents the percentage of observed positive markers (E) or allele frequency (F). The dashed line represents the detection threshold (95% CI) for ddPCR or HS-UMI for the indicated mutation.
[0042] Figure 6 This application illustrates an exemplary sequencing method and the adapter ligation efficiency of TruSeq, including (A) through (E), wherein,
[0043] (A) A diagrammatic illustration of droplet digital PCR (ddPCR) used to estimate ligation efficiency. A synthesized DNA template is ligated to adapter 1 or adapter 2 using an exemplary sequencing method of this application. Primers are designed to amplify either the “total” (For.1-Rev.1) or the “ligated” fragment (For.1-Rev.2). Reactions are arranged in separate wells with the same reporter fluorescent probe.
[0044] (B) Representative fluorescence curves of ddPCR reactions. Each dot represents a single PCR reaction (droplet) using one DNA molecule. Gray dots indicate background fluorescence, and green dots represent single droplets with amplification signals.
[0045] (CE) Copy number curves (left panel) and ligation efficiency table (right panel) for ddPCR using double-stranded TruSeq adapters (C), adapter 1 (D), and adapter 2 (E). For the copy number curves, the Y-axis represents copies / μL, while the X-axis represents three technical replicates (T1, 2, 3) and three biological replicates (B1, 2, 3). Notably, the curves show the count per well, while the table reflects the count normalized to the reaction volume (Tail-Tag.1 vs. Tail-Tag.2).
[0046] Figure 7 The diagrams illustrate the principles of the library preparation method, including (A) to (B), where,
[0047] (A) A schematic diagram of the library construction workflows used by NEB (NEBNext Ultra II), SWT (Accel-NGS Methyl-Seq), TELP, ELSA (an exemplary sequencing method of this application), SPLAT, SALP (SALP-seq), and Padlock.
[0048] (B) Summary of library preparation methods for whole-genome bisulfite sequencing (WGBS). Data are derived from current studies (NEB, SWT, ELSA) or previous studies (TELP, SPLAT, SALP). ds: dual-linker; ss: single-linker; ds-TA: TA cloning-mediated linker attachment; ss-Tailing: synthetic tail-mediated linker attachment; ss-guided: random-guided linker attachment; chimera: a combination of read pairs from two or more different sequences. Note: Padlock can only be used for targeted bisulfite sequencing.
[0049] Figure 8 The figures illustrate the free capture performance of an exemplary sequencing method of this application, including (A) to (G), wherein,
[0050] (A) Fragment size density curves using λDNA with different input amounts.
[0051] (B) Genomic range Poisson correlation of methylation levels of E. coli (DH5α) DNA using 0.5, 1 and 5 ng (sliding window: 20, 400 bp).
[0052] (C) Representative libraries were constructed using ELSA with 500 pg (red, dark red, lower peak), 1 ng (green, blue-purple, middle peak), and 2 ng cfDNA (yellow, aquamarine, higher peak) from two patients. Each sample was prepared with a technical replicate and included a template-free control (NTC, blue line, flatter bottom).
[0053] (D) Representative libraries constructed using 500 pg of E. coli DNA via NEB, SWT, and ELSA. Each sample was prepared with technical replicates and included a template-free control (NTC, dark red line, line with the highest peak).
[0054] (E) Mutation spectrum and frequency of libraries constructed using NEB.
[0055] (F) Coverage distribution across the entire λ genome when sequenced at different input levels at ~2000X. Gray shading indicates unique depth across the entire genome, while black bars highlight coverage at CpG sites.
[0056] (G) Methylation signal for each sequencing cycle estimated on Illumina Novaseq 6000 with different amounts of λDNA using C / C+T (R1) and G / G+A (R2) (the bars in each group are from left to right: Hiseq, Miniseq, Novaseq).
[0057] Figure 9 The figures illustrate the targeting performance of an exemplary sequencing method of this application, including (A) to (G), wherein,
[0058] (A) The selected DMLs (n = 80, 672) overlap with known human genome regions.
[0059] (B) Distribution of CpG target sites relative to known gene characteristic categories.
[0060] (C) Coverage uniformity curves show the fraction (y-axis) of target bases with a unique depth equal to or greater than the normalized coverage (x-axis). Here, normalized coverage is calculated by dividing the observed unique read depth for each base by the average unique read depth for all target bases.
[0061] (D) Displays the density curve of GC content against normalized read depth. The Y-axis represents the unique depth in the small subregion, while the X-axis indicates the GC content of the reference genome.
[0062] (E) Sequencing quality (Phred score) of read 1 and read 2 measured in each sequencing cycle on an Illumina Novaseq 6000.
[0063] (F) Size distribution of original cfDNA (blue, dashed line indicating peak to the left), amplified cfDNA without capture (Pre-Lib, red, dashed line indicating peak to the center), and amplified cfDNA after capture (post-Lib, dark red, dashed line indicating peak to the right) from two representative donors. The shift of peaks (dotted lines) towards higher molecular weights reflects the attachment of adapter / ligand sequences to the template during PCR-1 (80 bp, dashed line on the left) and PCR-2 (69 bp, dashed line in the middle).
[0064] (G) "+" or "-" indicates the unique depth of each CpG site on the chain.
[0065] Figure 10 The illustrations depict descriptions of methylated blocks and patterns, including (A) through (E), where,
[0066] (AB) Determination of the penalty coefficients α2(A) and α1(B) for block separation.
[0067] (C) Distribution of block lengths (base pairs) on the capture group.
[0068] (D) Distribution of the number of CpG sites in each block of the capture group.
[0069] (E) Methylation status (gray) and unmethylation status (black) illustrated by each unique sequencing read (vertical line) within the representative SHOX2 region (chr3: 157821291-157821596).
[0070] Figure 11 The diagram illustrates the reproducibility of MERMAID, including (A) through (C), where,
[0071] (A) A scatter plot showing intra-batch consistency of iAF (left panel) and MBS (right panel) using the same NA12878 DNA sample (10 ng).
[0072] (B) Scatter plot of MBS values observed in two intra-batch repeats (left panel) and inter-batch repeat sequences (right panel) using the same cfDNA sample (10 ng).
[0073] (C) Density curves were obtained by using different NA12878 DNA input amounts (2, 5, 10, 30 ng) or MBS values at different sequencing depths (1000-5000X).
[0074] Figure 12 The diagram illustrates the accuracy of MERMAID, including (A) through (D), where,
[0075] (A) IGV (Integrated Genomics Viewer) observation of MERMAID data from H2122, H2228, and NA12878 cells at the SHOX2 and SEPT9 loci. For forward reads, red "C" indicates methylated "C" (unconverted), while blue "T" indicates unmethylated C (C->T conversion). For reverse reads, the interpretation is exactly the opposite.
[0076] (B) Methylation of SHOX2 and SEPT9 was verified by qMSP (quantitative methylation-specific PCR). The Y-axis represents the relative fold change in methylation level (msp_PCR on the left and NGS on the right for each group of bars).
[0077] (C) Density curves of iAF differences measured in NA12878 cells using MERMAID and Illumina Epic TruSeq (public data). Only overlapping CpG sites (n = 32,383) were used. The red dashed line indicates the median of plateau-related iAF differences.
[0078] (D) Scatter plot of the logarithmic transformation iAF generated using Illumina Epic TruSeq Methyl (Y-axis) and MERMAID (X-axis).
[0079] Figure 13 The diagram illustrates the FDR and LoD of MERMAID, including (A) through (F), where,
[0080] (A) Box plot showing the dependence of FDR (Y-axis) and sequencing depth (Y-axis) by using 10 ng cfDNA.
[0081] (B) Numerical simulation of sampling noise using a Poisson distribution (unique segment = 500, noise = 0.002). The X-axis represents the number of observations (depth), and the Y-axis represents the variance.
[0082] (C) Detection of tumor-derived methylation counts varying with tumor burden (0.000001–0.001), unique depth (left panel, marker = 1000), and number of markers (right panel, depth = 500). The X-axis represents the predefined tumor fraction θ in each simulated “tumor” sample. i The Y-axis represents the average detection rate of the sample in 10,000 repetitions. The probability of a successful observation is defined as P < 0.05, and the P-value is tested using a likelihood ratio test. i The null hypothesis is used to determine the value of 0.
[0083] (D) Evaluation of batch effects in bioinformatics LoD assays. Samples A (left) and B (right) were from two different healthy donors and sequenced in two separate rounds. Reads were spiked and analyzed as described in the Methods section of Example 1.
[0084] (EF)ddPCR fluorescence curves show the detection of EGFR p.G719S mutation (A) and EML4-ALK fusion (B) in spiking experiments on SW48 and NCI-H2228 cell lines, respectively. In each plate, the gate thresholds are indicated by horizontal solid lines (e.g., 6000, 2400, 3000, and 2760).
[0085] Figure 14 The diagram illustrates the selection and classification of tissue-specific LC biomarkers, including (A) through (D), where,
[0086] (A) Volcano plots of MBS measured in LC tissue (n=48) versus normal lung tissue (n=20, left panel) and LC tissue versus healthy cfDNA (n=30, right panel). Dark dots indicate important markers while light gray dots indicate unimportant markers.
[0087] (B) Violin plot showing MBS on representative specifiers in LC, normal lung, and control cfDNA samples. Each dot represents one sample.
[0088] (C) Supervised classification of tumor and normal tissue samples with different numbers of specifiers. The Y-axis represents the prediction accuracy in 100 replicates, and the X-axis represents the number of specifiers.
[0089] (D) PCA (principal component analysis) clustering of tumor and normal tissue samples with different numbers of specifiers.
[0090] Figure 15 The illustration shows the clinical characteristics of the validation cohort (plasma). The table shows 308 LC patients and 261 non-cancer controls stratified by clinical characteristics. UNK: Unknown. LUAD: Lung adenocarcinoma; LUSC: Lung squamous cell carcinoma.
[0091] Figure 16 The figures illustrate parallel comparisons of MERMAID, HS-UMI, and patient-specific ddPCR, including (A) through (F), where,
[0092] (A) Research design comparing “double”, “triple” and “quadruple”.
[0093] (BD) shows the ddPCR fluorescence curves for the detection of EGFR P.S746_I750DEL (B) for patient P65, KRAS p.G12D (C) for patient P64, and EGFR P.L858R (D) for patient P66. The DNA input for each ddPCR reaction was 38, 62, and 54 ng, respectively. NTC: template-free control; NC: normal WBC; PC: positive control with the expected mutation having 0.1% or 0.5% AF (multiplex I cfDNA reference standard group, Horizon Discovery). Notably, 0.5% PC was generated by mixing WT and 1% reference DNA in a 1:1 ratio. Details are provided in Table 7.
[0094] (E) maxAF (n=115) of all participants in Group 1 and Group 2 categorized by disease status.
[0095] (F) Predicted scores of all participants in groups 1 and 2 classified by maxAF (n = 115). Dotted lines represent the threshold (positive or negative) with 96% training specificity.
[0096] Figure 17 The figure illustrates the performance results of detecting cancer-related changes based on the methylation level iAF at a single site and the average methylation level mAF (mean methylation allele frequency) per block.
[0097] Figure 18 The diagram illustrates the influence of "regional median length" and "regional length variation coefficient" on the values of α2 and α1, respectively.
[0098] Figure 19 The figure illustrates the density curves of MBS and mAF statistics in samples with different blending ratios. Detailed Implementation
[0099] The following specific embodiments illustrate the implementation of the invention. Those skilled in the art can easily understand other advantages and effects of the invention from the content disclosed in this specification.
[0100] Terminology Definition
[0101] In this application, the terms "next-generation sequencing (NGS)," "high-throughput sequencing," or "next-generation sequencing" generally refer to second-generation high-throughput sequencing technology and subsequent higher-throughput sequencing methods. Next-generation sequencing platforms include, but are not limited to, existing sequencing platforms such as Illumina. With the continuous development of sequencing technology, those skilled in the art will understand that other sequencing methods and devices can also be used in this method. For example, next-generation sequencing can have advantages such as high sensitivity, high throughput, high sequencing depth, or low cost. Based on differences in development history, influence, sequencing principles, and technologies, the main types include: Massively Parallel Signature Sequencing (MPSS), Polony Sequencing, 454 pyro sequencing, Illumina (Solexa) sequencing, Ion semi-conductor sequencing, DNA nano-ball sequencing, and Complete sequencing. Genomics' DNA nanoarrays and combined probe anchoring ligation sequencing methods, among others. The aforementioned next-generation sequencing makes it possible to perform detailed and comprehensive analysis of the transcriptome and genome of a species, hence it is also known as deep sequencing. For example, the method described in this application can also be applied to first-generation sequencing, second-generation sequencing, third-generation sequencing, or single-molecule sequencing (SMS).
[0102] In this application, the term "sample to be tested" generally refers to a sample that needs to be tested. For example, it can be used to detect whether one or more gene regions on the sample to be tested are modified.
[0103] In this application, the term "complementary region" generally refers to a region that is complementary to a reference nucleotide sequence. For example, complementary nucleic acids can be nucleic acid molecules that optionally have opposite orientations. For example, the complementarity can refer to having the following complementary associations: guanine and cytosine; adenine and thymine; adenine and uracil.
[0104] In this application, the term "hybridization" generally refers to a reaction in which one or more polynucleotides react to form a complex stable by hydrogen bonds between the bases of nucleotide residues. Hydrogen bonding can occur through Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner based on base complementarity. The complex may comprise two strands forming a double helix, three or more strands forming a multi-stranded complex, a self-hybridized single strand, or any combination thereof. Hybridization reactions can constitute steps in broader methods, such as the initiation of PCR or the enzymatic cleavage of polynucleotides by endonucleases. A second sequence that is completely complementary to the first sequence, or that is polymerized by a polymerase using the first sequence as a template, is referred to as "complementary" to the first sequence. The term "hybridizable," as applied to polynucleotides, refers to the ability of a polynucleotide to form a complex stable by hydrogen bonds between the bases of nucleotide residues in a hybridization reaction. In some embodiments, the hybridizable nucleotide sequence is at least about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to the sequence it hybridizes to.
[0105] In this application, the terms "polynucleotide," "nucleotide," "nucleic acid," and "oligonucleotide" are used interchangeably. They refer to polymeric forms of nucleotides (deoxyribonucleotides or ribonucleotides) of any length, or analogs thereof. Polynucleotides can have any stereostructure and can perform any function, whether known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene segments, loci defined by linkage analysis (locuses), exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribonucleases, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, primers, and adapters. Polynucleotides may include one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs.
[0106] In this application, the term "modification state" generally refers to the modification state of a gene fragment, nucleotide, or base thereof in this application. For example, a modification state in this application may refer to a modification state of cytosine. For example, a gene fragment with a modified state in this application may have altered gene expression activity. For example, a modification state in this application may refer to methylation modification of a base. For example, a modification state in this application may refer to the covalent binding of a methyl group at the 5' carbon position of cytosine in the CpG region of genomic DNA, for example, it may be 5-methylcytosine (5mC). For example, a modification state may refer to the presence or absence of 5-methylcytosine ("5-mCyt") within the DNA sequence.
[0107] In this application, the term "methylation" generally refers to the methylation state of a gene fragment, nucleotide, or base thereof in this application. For example, the DNA fragment containing the gene in this application may be methylated on one or more strands. For example, the DNA fragment containing the gene in this application may be methylated at one or more sites.
[0108] In this application, the term "transformation" generally refers to the conversion of one or more structures into another. For example, the transformations in this application can be specific. For example, cytosine without methylation modification can be transformed into other structures (e.g., uracil), while cytosine with methylation modification can remain substantially unchanged after transformation. For example, cytosine without methylation modification can be cleaved after transformation, while cytosine with methylation modification can remain substantially unchanged after transformation.
[0109] In this application, the term "bisulfite," or "bisulfite," generally refers to a reagent capable of distinguishing between modified and unmodified DNA regions. For example, a bisulfite may include bisulfite, its analogues, or combinations thereof. For example, a bisulfite can deamination of the amino group of unmodified cytosine to distinguish it from modified cytosine. In this application, the term "analyte" generally refers to a substance having a similar structure and / or function. For example, an analogue of a bisulfite may have a similar structure to a bisulfite. For example, an analogue of a bisulfite may refer to a reagent that can also distinguish between modified and unmodified DNA regions.
[0110] In this application, the term "comprising" generally means including the explicitly specified features, but does not exclude other elements.
[0111] In this application, the term "about" generally refers to a variation within a range of 0.5% to 10% above or below a specified value, such as a variation within a range of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% above or below a specified value. Invention Details
[0113] On the one hand, this application provides a method for detecting methylation modification of a target nucleic acid, the method comprising the following steps: step (a-1) determining a comethylation block based on the correlation coefficient of CpG sites in the target nucleic acid, the methylation level of the CpG sites, and the location information of the CpG sites; and / or step (a-2) determining the comethylation block based on the correlation coefficient of CpG sites in the target nucleic acid, the information content of candidate comethylation blocks, and the degree of balance in the division of the candidate comethylation blocks; and step (b) determining the presence and / or content of the target nucleic acid in the sample to be tested based on the methylation degree of the comethylation block.
[0114] For example, the method detects the presence and / or quantity of the methylated target nucleic acid in the test sample. As used herein, the “object” from which the test sample originates can be a mammal, such as a non-primate (e.g., cow, pig, horse, cat, dog, rat, etc.) or a primate (e.g., monkey or human). In some embodiments, the object is a human. In some embodiments, the object is a mammal (e.g., a human) that has or is potentially suffering from a disease, condition, or ailment described herein as an example thereof. In some embodiments, the object is a mammal (e.g., a human) at risk of developing a disease, condition, or ailment described herein as an example thereof.
[0115] For example, the correlation coefficient of the CpG sites includes the Pearson correlation coefficient between two or more CpG sites in the target nucleic acid.
[0116] For example, the methylation level of the CpG site includes the difference in average methylated allele frequency (mAF) between two or more CpG sites in the target nucleic acid. For example, the methylation level of the CpG site includes the ratio of the difference in mAF between two or more CpG sites in the target nucleic acid to the sum of the mAFs.
[0117] For example, the CpG site location information includes the difference in genomic location between two or more CpG sites in the target nucleic acid. For example, the CpG site location information includes the ratio of the genomic location distance between two or more CpG sites in the target nucleic acid to the length of the target nucleic acid.
[0118] For example, step (a-1) includes: determining the corrected correlation coefficient between every two CpG sites of the target nucleic acid, and the corrected correlation coefficient d between site i and site j. ij Calculated using the following formula: Where, ρ ij E(y) represents the Pearson correlation coefficient. i E(y) represents the average methylated allele frequency (mAF) of all samples at site i. j ) represents the average methylated allele frequency (mAF) of all samples at position j, pos i pos represents the genomic location of site i. j The genomic location of site j is represented by L, and the length of the target nucleic acid region is represented by λ1 and λ2, which are independently selected from 0 or larger numbers. For example, the value of λ1 ranges from 0 to 1. For example, the value of λ2 ranges from 0 to 1. For example, in this application, λ1 and λ2 are independently selected from 0. For example, in this application, λ1 is selected from 0; for example, in this application, λ2 is selected from 0.
[0119] For example, the correlation coefficient based on CpG sites in the target nucleic acid in step (a-2) includes the corrected correlation coefficient of CpG sites in the target nucleic acid, which includes the Pearson correlation coefficient between two or more CpG sites in the target nucleic acid after correction based on the methylation level of the CpG sites and / or the position information of the CpG sites.
[0120] For example, the corrected correlation coefficient of the CpG site in the target nucleic acid includes the corrected correlation coefficient in the method of step (a-1) above.
[0121] For example, the information content of the candidate comethylation block includes the number of CpG sites in the candidate comethylation block.
[0122] For example, the degree of balance in the partitioning of the candidate comethylation blocks includes the difference in the number of CpG sites in the different candidate comethylation blocks. For example, the degree of balance in the partitioning of the candidate comethylation blocks includes the coefficient of variation of the number of CpG sites in the different candidate comethylation blocks.
[0123] For example, step (a-2) includes: maximizing the block index of the candidate comethylation block, determining the comethylation block, and the candidate comethylation block in the target nucleic acid. The block index Calculated using the following formula:
[0124] in, B i The number of CpG sites in the i-th candidate comethylation block is represented by α1 and α2, which are independently selected from 0 or greater. For example, the block breakpoints that make up the comethylation block are determined by an iterative method using unique breakpoints. For example, the value of α1 ranges from 0 to 10. For example, the value of α2 ranges from 0 to 10. For example, in this application, the values of α1 and α2 are independently selected from rational numbers from 0 to 10. For example, in this application, α1 and α2 are independently selected from 0. For example, in this application, α1 is selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10; for example, in this application, α2 is selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0125] For example, step (b) includes determining the presence and / or content of the target nucleic acid based on the length of consecutive CpGs in each sequencing read of the comethylated block, the number of CpGs on the read, and the total number of reads in the comethylated block.
[0126] For example, step (b) includes: determining the methylation block score (MBS) of the comethylated block, the MBS being calculated using the following formula:
[0127]
[0128] Wherein, n is the total number of reads covering all CpG sites in the comethylation block, L i It is the number of CpG sites contained in the i-th read, l ij is the length of consecutive methylated CpG sites on the i-th read, and m is the sequencing depth on the i-th read.
[0129] For example, UMI correction is applied to the sequencing data of the sample to be tested.
[0130] For example, the method further includes: extracting the feature values of the MBS of the comethylated blocks in tumor samples and healthy samples using a machine learning model, and determining the presence and / or content of the target nucleic acid based on the MBS of the comethylated blocks in the sample to be tested.
[0131] An analytical device for detecting methylation modification of a target nucleic acid, the device comprising: a block partitioning module (a-1) for determining comethylated blocks based on the correlation coefficient of CpG sites in the target nucleic acid, the methylation level of the CpG sites, and the positional information of the CpG sites; and / or a block partitioning module (a-2) for determining the comethylated blocks based on the corrected correlation coefficient of CpG sites in the target nucleic acid, the information content of candidate comethylated blocks, and the partitioning balance of the candidate comethylated blocks; and a determination module (b) for determining the presence and / or content of the target nucleic acid in the sample to be tested based on the methylation degree of the comethylated blocks.
[0132] For example, the analytical device for detecting methylation modification of the target nucleic acid of this application may include the steps of the method for detecting methylation modification of the target nucleic acid of this application.
[0133] As one aspect of this disclosure, this application provides a method for detecting ctDNA using methylation sequencing, comprising: selecting differentially methylated CpG sites; dividing the CpG sites into multiple comethylated blocks based on the similarity of the methylation status of the CpG sites; sequencing the sample to obtain methylation sequencing reads; and detecting the average methylation level of each comethylated block of the sample for further DNA methylation analysis.
[0134] In a preferred embodiment of this disclosure, differentially methylated CpG sites are selected from a TCGA database generated by an Infinium HumanMethylation 450K array.
[0135] In another preferred embodiment of this disclosure, the average methylation level is the average methylation allele frequency.
[0136] In another preferred embodiment of this disclosure, the number of comethylated blocks is between 1 / 30 and 1 / 5 of the number of CpG sites.
[0137] In another preferred embodiment of this disclosure, comethylation blocks are further restricted by comparing specific tumor samples with normal tissue samples, such as primary lung tumors and normal lung tissue samples.
[0138] In another preferred embodiment of this disclosure, comethylated blocks of insufficient depth (<100) on most samples (>80%) are excluded from downstream analysis.
[0139] In another preferred embodiment of this disclosure, the comethylated blocks are separated based on an improved correlation matrix referred to as the “block index”.
[0140] In another preferred embodiment of this disclosure, the method further includes normalizing the depth difference of each methylated block using a methylated block score (MBS), which is used to distinguish very small tumor signals, such as 0.1%, 0.2%, 0.5%, and 1%.
[0141] In another preferred embodiment of this disclosure, the sequencing reads to be analyzed are trimmed by any one or more of the following criteria: (i) the number of G bases longer than a fixed value m is retrieved; (ii) the number of non-G bases is less than a fixed number p; and the next base is a high-quality A / T / C (Phred score > 30).
[0142] In another preferred embodiment of this disclosure, repeated removal is applied at the starting points of both R1 and R2 with a tolerance of + / -3bp to minimize artifacts associated with improperly assigned fragment end positions.
[0143] In another preferred embodiment of this disclosure, UMI is applied during correction.
[0144] In another preferred embodiment of this disclosure, the average methylation level of each comethylated block of unmethylated phage λDNA is detected to measure genome-wide “technical noise” (C / C+T in read 1 and G / G+A in read 2).
[0145] In another preferred embodiment of this disclosure, a machine learning classifier of methylation patterns is applied to assess tumor levels, preferably in early tumor screening.
[0146] As another aspect of this disclosure, it provides the use of the above-described method in assessing general tumor levels during early tumor screening. In a preferred embodiment of this disclosure, the above-described method is used in assessing tumor levels during early tumor screening, wherein the tumors are homogeneous tumors, heterogeneous tumors, hematologic malignancies, and / or solid tumors; preferably, the tumors are from one or more of the following groups of cancers: brain cancer, lung cancer, skin cancer, nasopharyngeal carcinoma, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, colorectal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, gastric cancer, solid tumors, ovarian cancer, esophageal cancer, gallbladder cancer, bile duct cancer, breast cancer, cervical cancer, uterine cancer, prostate cancer, head and neck cancer, sarcoma, thoracic malignancies (excluding lung), melanoma, and testicular cancer. In a preferred embodiment of this disclosure, the above-described method is used in assessing lung tumor levels during early tumor screening.
[0147] As another aspect of this disclosure, it provides a kit for detecting ctDNA using methylation sequencing, wherein the kit can be used to capture at least 50, 100, 150, 200, 300, 500, 800, 1000, 1500 or 2000 comethylated blocks as shown in Table 5-2; preferably, it can be used to capture at least all of the comethylated blocks as shown in Table 5-1.
[0148] As another aspect of this disclosure, it provides an apparatus for performing the above-described methods in assessing general tumor levels during early tumor screening.
[0149] As another aspect of this disclosure, it provides a non-volatile memory for storing programs that can be used to perform the above-described methods in assessing general tumor levels during early tumor screening.
[0150] On one hand, this application provides a method for detecting the level of base modification, comprising providing the nucleic acid molecule combination of this application and / or the kit of this application. For example, the base modification includes methylation modification.
[0151] On one hand, this application provides a storage medium containing a program capable of running the methods of this application. For example, the non-volatile computer-readable storage medium may include floppy disks, flexible disks, hard disks, solid-state storage (SSS) (e.g., solid-state drives (SSDs)), solid-state cards (SSCs), solid-state modules (SSMs)), enterprise-grade flash drives, magnetic tapes, or any other non-transitory magnetic media. The non-volatile computer-readable storage medium may also include punched cards, paper tape, cursor sheets (or any other physical medium with perforated patterns or other optically identifiable markings), compact disc read-only memory (CD-ROM), rewritable optical disc (CD-RW), digital versatile optical disc (DVD), Blu-ray disc (BD), and / or any other non-transitory optical media.
[0152] On one hand, this application provides an apparatus that includes the storage medium of this application. For example, the apparatus further includes a processor coupled to the storage medium, the processor being configured to execute, based on a program stored in the storage medium, the method of this application.
[0153] The embodiments described below are not intended to be limited by any theory, but are merely for illustrating the methods and uses of this application and are not intended to limit the scope of the invention.
[0154] Example
[0155] Example 1: Detection method of this application
[0156] Methods section
[0157] Biomarker discovery and verification
[0158] Initially, differentially methylated sites were screened from the TCGA database generated by the Infinium HumanMethylation 450K array. A total of 4539 tumor samples and 521 normal tissue samples were analyzed. Data from the GEO dataset (GSE40279) of 656 normal WBC samples were used to remove hypermethylated CpG sites (>0.1) in the hematopoietic lineage. CpG sites located on the X or Y chromosome were also excluded. DML selection was performed using the "limma (V2.0)" software, with the cutoff value set to a BH-corrected FDR <0.05. In addition, CpG sites associated with common cancers in previous studies were also included. This resulted in a total of 80,672 CpG sites in the biomarker discovery phase.
[0159] CpG sites were then segmented into 8,312 blocks (described in “Comethylation Block Segmentation”) and validated by comparing internal sequencing data from 48 primary lung tumor samples and 20 normal lung tissue samples. Blocks with insufficient depth (<100) in most samples (>80%) were excluded from downstream analysis. Linear regression was used to select differentially methylated blocks, with cutoff values set to log(fold change) > 0.05 and BH-corrected FDR < 0.05. A total of 2,473 blocks were selected as taxonomic features, and genomic coordinates are listed in Table 5.
[0160] Data processing framework
[0161] FASTQ files were generated from the raw BCL data using bcl2fastq (V2.19.1). Illumina-specific adapters and low-quality sequences were trimmed using trimmomatic (V0.36) (SLIDINGWINDOW:4:15TRAILING:20). For Accel-NGS Methyl-Seq (Swift Biosciences), additional trimming was performed as described in previous work. For ELSA-seq, an exemplary sequencing method of this application, parameters were tested at different levels of rigor to remove low-complexity tail sequences. The final trimming criteria were set as follows: (i) the number of G base segments longer than a fixed value m; (ii) the score of non-G bases less than a fixed number p; and (iii) the next base was a high-quality A / T / C (Phred score > 30). Based on simulation and experimental analysis, M = 10 and P = 0.1 were used in this study. Then use FastQC (V0.11.5) to check the quality of the trimmed reads and require a minimum read length of 50 bases.
[0162] Secondly, reads were aligned with the bioinformatically transformed hg 19 human genome (C->T for R1, G->A for R2) using the bwa-meth software (V0.2.0) with default parameters. Repeated reads were flagged using Samblaster (V0.1.24), and BAM files were sorted using Picard (V1.138). Reads with alignment scores <20, mismatches >5, improper pairings, or multiple mapping positions were excluded from secondary analysis. To minimize artifacts associated with improper fragment end placement, duplicate removal was applied at the start points (“fuzzy windows”) of both R1 and R2 with a tolerance of + / -3 bp. It is important to note that while this strategy correctly identified the original molecule more than 90% of the time, accuracy decreased when the library sequencing was severely over- or under-sequential.
[0163] For context-based error correction and methylation metric calculations, an internal module was built to fold forward and reverse reads into a single, consistent sequence relating to base call quality, cigar string, and mutation direction (e.g., C->A) relative to a reference genome. Base calls in low-fidelity regions (e.g., near the ends of reads) are underestimated, and missing bases are recovered based on supporting reads using the nearest Hamming distance.
[0164] Comethylated block separation
[0165] This application is first based on a concept called "block index". The improved correlation matrix divides the design region of the capture group into comethylation blocks.
[0166] (i) Represent the iAF correlation matrix of each CpG site in region r as follows: The correlation coefficient between site i and site j is calculated as follows:
[0167] ρ ij Represents the Poisson correlation coefficient.
[0168] E(y i ) represents the mAF on all samples at position i.
[0169] pos i Indicates genomic location,
[0170] L represents the length of the original region.
[0171] λ1 and λ2 are parameters estimated using prior information.
[0172] (ii) For the new split within the original region r The "block index" is defined by the following function:
[0173]
[0174]
[0175] Where B represents i This represents the number of sites in the i-th block.
[0176] α1 and α2 are the penalty coefficients for unbalanced splitting and over-splitting, respectively. In this study, based on the desired block length and uniformity, both α1 and α2 are set to 1.0.
[0177] This application defines two indicators: "regional median length" and "regional length coefficient of variation." "Regional median length" measures the size of the defined region; a smaller value indicates less information contained within the region, while a larger value indicates more information. "Regional length coefficient of variation" (standard deviation / mean) measures the difference in size between different regions; a larger value indicates a more unbalanced region division (more independent points, more information omissions), while a smaller value indicates a more balanced region division. The value of α1 primarily affects the "regional length coefficient of variation," and the value of α2 primarily affects the "regional median length," with the relationship as follows: Figure 18 As shown.
[0178] (iii) Repeat the partitioning process using iterative division with unique breakpoints. Assuming k breakpoints, then by making... The maximum value is reached to select the (k+1)th new breakpoint.
[0179] (iv) Terminate the interactive algorithm when the (k+1)th breakpoint matches the kth breakpoint.
[0180] Definition of Methylated Block Score (MBS)
[0181] This application defines the measurement of MBS as follows:
[0182]
[0183] For a given block,
[0184] n is the total number of reads covering multiple CpG sites.
[0185] Li is the number of CpG sites covered on the i-th read.
[0186] lij represents the length of consecutive methylated CpG sites (>1).
[0187] And m represents the total count over the i-th read length. The number of read lengths in each block is used to normalize the depth difference.
[0188] Simulation of fragment recognition based on UMI correction or shift correction
[0189] The following is a simulation of the current pipeline and UMI-assisted method used in this application:
[0190] (i) Template generation: Generate a library of 2kb DNA segments and “cut” them into various lengths by calculation (average = 170, sd = 30);
[0191] (ii) Tail addition: Based on empirical observation, tail sequences (90% G) were added to the 5'-end of R2 at different lengths;
[0192] (iii) UMI incorporation: A random 6-base UMI tag (corresponding to the modified linker 1) is added to the 5'-end of R2 to mark the original template;
[0193] (iv) Technical errors: Three datasets were generated with low / medium / high substitution error rates typical for Illumina sequencers (0.005, 0.01, 0.02); all errors (e.g., A->T; T->A) were set with equal probability; indels (+ / - 1 base) were introduced in the tail region to account for homopolymer properties; the accumulation of PCR errors was simulated by 10 rounds of PCR reactions with a repetition probability of 0.9 (1 means complete doubling) per round; incomplete extension was simulated by random position fluctuations (<10 bases) at the 5'-end of R1.
[0194] (v) UMI and non-UMI solutions:
[0195] a) UMI: Apply UMI extraction (5′-end of R2) and shift correction (5′-end of R1, fw=3) to each end separately; to avoid over-counting of UMIs due to amplification or sequencing errors, the minimum edit distance is allowed to be 2.
[0196] b) Non-UMI: Perform shift correction on both ends (5' end of R1 and 5' end of R2, fw=3);
[0197] (vi) Template Count: Set the original (post-BC) fragment depth to 250-500X, and set the original fragment depth to 500-1000X.
[0198] The entire process was repeated 9 times, and the average performance of the simulation is summarized in Table 2.
[0199] UMI-facilitated template counting experiment
[0200] A DNA mixture (10 ng) of λ DNA and filler DNA with ~5,000 haploid copies (before BC) was prepared. To incorporate UMIs into ELSA-seq, an exemplary sequencing method of this application, the 3' adapter was modified with a 6-base random UMI (Adapter 1-UMI) inserted adjacent to the dangling sequence, and thus sequenced in the first six cycles of R2 (Table 1). To aid in UMI localization, two nucleotides (e.g., 5'-NNDDNN-3', D:A / T / G) were predefined to act as anchors. To minimize the risk of false UMI recognition due to polymerase or base calling errors, a minimum edit distance of 2 was allowed for error correction. ~2 kb regions in the λ genome (4,500–6,539) were targeted and analyzed, yielding ~50,000 raw reads. Unique fragments measured by the current pipeline (non-UMI) of this application were compared with the UMI-enhanced counting strategy by random read downsampling. The results are summarized in Table 2.
[0201] Supplementary methods section
[0202] The adapter ligation efficiency of ELSA-seq and TruSeq, two exemplary sequencing methods of this application, was determined by ddPCR.
[0203] To evaluate ligation efficiency, the ddPCR reaction was set up as follows:
[0204] 1. ELSA-seq ligation reaction (20 μl)
[0205] 10X CutSmart buffer 2 <![CDATA[CoCl2(2.5mM)]]> 2 NAD (0.5mM) 1 Template DNA 8.6 <![CDATA[H2O]]> 2.1 dCTP:dNTP (2.5mM) 0.8 Adapter (connector) (10uM) 2 TdT 0.5 E. coli ligase 1
[0206] 2. ddPCR reaction (20 μl):
[0207]
[0208]
[0209] Substrate:
[0210] ELSA-SEQ: Synthetic ssDNA (KRAS-10N)
[0211] TruseQ: dsDNA amplified using primers KRASF and KRASR (KRAS-177)
[0212] ELSA Connection-1:
[0213] Total copies were detected using primers LEF and KRASR and probe KRAS-G13D.
[0214] The ligation copy was detected using primers LEF and LER and probe KRAS-G13D.
[0215] ELSA Connection-2:
[0216] Total copies were detected using primers LEF and KRASR and probe KRAS-G13D.
[0217] The copy of the ligation was detected using primers LEF and LER-ATNR1 and probe KRAS-G13D.
[0218] TruSeq connection:
[0219] Total copies were detected using primers LEF and KRASR and probe KRAS-G13D.
[0220] The copies of the ligation were detected using primers LEF and LER-ATNR1 and probe KRAS-G13D.
[0221] The design using the 'N' in the KRAS-10N template was employed to avoid sequence-dependent bias. All oligonucleotide sequences are listed in Table 1 - Summary of Oligonucleotide Sequences.
[0222] Construction and analysis of whole-genome bisulfite sequencing (WGBS) libraries
[0223] WGBS libraries were constructed using ELSA-seq, NEBNext Ultra II (New England Biolabs), and Accel-NGS Methyl-Seq (Swift Biosciences) as described in the methods or according to the manufacturer's instructions. Genomic DNA was sonicated to ~200 bp (peak). Library quality was evaluated using a LabChip GXII touch 24 (Perkin Elmer). Paired-end sequencing (2 × 150 bp) was then performed on an Illumina NovaSeq 6000 system.
[0224] Simulation of fragment recognition based on UMI correction or shift correction
[0225] The following is a simulation of the current pipeline and UMI-assisted method used in this application:
[0226] (vii) Template generation: Generate a library of 2kb DNA segments and "cut" them into various lengths by calculation (average = 170, sd = 30);
[0227] (viii) Tail addition: Based on empirical observation, tail sequences (90% G) were added to the 5'-end of R2 at different lengths;
[0228] (ix) UMI incorporation: A random 6-base UMI tag (corresponding to the modified linker 1) is added to the 5'-end of R2 to mark the original template;
[0229] (x) Technical errors: Three datasets were generated with low / medium / high substitution error rates typical for Illumina sequencers (0.005, 0.01, 0.02); all errors (e.g., A->T; T->A) were set with equal probability; indels (+ / - 1 base) were introduced in the tail region to account for homopolymer properties; the accumulation of PCR errors was simulated by 10 rounds of PCR reactions with a repetition probability of 0.9 (1 means complete doubling) per round; incomplete extension was simulated by random position fluctuations (<10 bases) at the 5'-end of R1.
[0230] (xi) UMI and non-UMI solutions:
[0231] a) UMI: Apply UMI extraction (5′-end of R2) and shift correction (5′-end of R1, fw=3) to each end separately;
[0232] b) Non-UMI: Perform shift correction on both ends (5' end of R1 and 5' end of R2, fw=3);
[0233] (xii) Template count: In order to reflect the ELSA-seq library in a clinical setting, the original (post-BC) fragment depth was set to 250-500X, while the original fragment depth was set to 500-1000X.
[0234] The entire process was repeated 9 times, and the average performance of the simulation is summarized in Table 2.
[0235] UMI-facilitated template counting experiment
[0236] A DNA mixture (10 ng) of λ DNA and filler DNA with ~5,000 haploid copies (before BC) was prepared. To incorporate UMIs into ELSA-seq, an exemplary sequencing method of this application, the 3' adapter was modified with a 6-base random UMI (Adapter 1-UMI) inserted adjacent to the dangling sequence, and thus sequenced in the first six cycles of R2 (Table 1). To aid in UMI localization, two nucleotides (e.g., 5'-NNDDNN-3', D:A / T / G) were predefined to act as anchors. To minimize the risk of false UMI recognition due to polymerase or base calling errors, a minimum edit distance of 2 was allowed for error correction. ~2 kb regions in the λ genome (4,500–6,539) were targeted and analyzed, yielding ~50,000 raw reads. Unique fragments measured by the current pipeline (non-UMI) of this application were compared with the UMI-enhanced counting strategy by random read downsampling. The results are summarized in Table 2.
[0237] Methyltransferase-mediated in vitro methylation
[0238] Set the reaction as follows:
[0239]
[0240] The in vitro methylation process was carried out at 37°C for 1 hour and terminated by heating at 65°C for 20 minutes.
[0241] Calculate the sample size with the required error range (<0.05).
[0242]
[0243] The sample size (n) is calculated as follows:
[0244]
[0245] in and It is the inverse function of the standard cumulative normal distribution.
[0246] Analysis of group composition
[0247] To assess potential confounding effects (such as age and sex) in the cohort design, a chi-square test was used, with the null hypothesis that categorical variables are comparable in the sample or control classes (P > 0.05). Additional data on patients and tumors are available in [link to relevant documentation]. Figure 15 And in Table 6.
[0248] Multivariate analysis of predictors
[0249] Univariate and multivariate analyses were performed using logistic regression to identify important clinical factors for the MERMAID trial outcome. Logit linearity (predictive scores) was tested for all independent variables (clinical factors) with the dependent variable. The results are presented in Table 6.
[0250] Calculation of potential clinical benefit in the hypothetical screening population
[0251] MERMAID was evaluated for diagnostic yield in LC in a hypothetical screening population of 10,000 participants. Based on the results of this study, sensitivity was set at 63.0% (194 / 308) and specificity at 96.2% (251 / 261). The hypothetical prevalence of LC in older adults at average risk was 0.53%, according to the Surveillance, Epidemiology, and Final Outcomes (SEER) program (SEER.cancer.gov / data / access, 2020). The number of individuals with true positive (TP), false positive (FP), true negative (TN), and false negative (FN) results was predicted in the 10,000 participants. Positive predictive value (PPV) and negative predictive value (NPV) were calculated as follows:
[0252]
[0253]
[0254]
[0255] The predicted FP for each TP indicates the number of FP subjects observed when TP subjects were tested. A caveat is that while all patients in this application had early (stage I-III) LC and 65% (199 / 308) were diagnosed with stage I (Ia and Ib) at study entry, the stage distribution in the actual screening context remains unknown, so the PPV may be less than reported here.
[0256]
[0257] For tests where a positive result leads to action, performance can be evaluated as follows:
[0258] (1)
[0259] (2) Here, R refers to the risk of the lung nodule being malignant, for which this application can choose surveillance or active investigation without bias. Based on the report from the National Lung Screening Trial (NLST) study group, R is set at 1.1%.
[0260] In the hypothetical average risk population, MERMAID's performance meets the criteria based on formulas (1) and (2) (MERMAID: 16.6 > 2.1).
[0261] Example 2: Results of the methylation detection method of this application
[0262] This application relates to the design of an ultrasensitive BS-seq method.
[0263] The methylation detection method of this application can be based on methylation sequencing data known in the art. For example, the methylation detection method of this application can be based on an ultrasensitive BS-seq (bisulfite sequencing). For example, a usable BS-seq can be ELSA-Seq as described in WO2019192489A1. For example, the design principle of the sequencing method for methylation sequencing data in this application involves both molecular and computational aspects. Figure 1 To improve detection capabilities, this application first focuses on increasing the number of templates that can be efficiently sequenced: (i) DNA molecules require specific adapters at both the 5' and 3' ends to be "read" by the high-throughput sequencer. Breakage of the tagged molecule will lead to failure of "seeding" on the surface of flowing cells. To minimize adapter loss caused by BC, this application first arranges the denaturation step, followed by a single-stranded DNA (ssDNA) compatible step, to maximize the retrieval of the original template. (ii) Adapter ligation is another common limiting factor, therefore this application devises a novel strategy called "tail and tag" to improve efficiency. In short, the bisulfite-treated DNA is denatured, dephosphorylated, and extended to a cytosine-rich nucleotide tail by TdT (terminal deoxynucleotidyl transferase). Then, in the presence of E. coli ligase, the clip adapter and tail are annealed to promote a highly efficient ligation step (Tail-Tag. 1). (iii) Secondly, a copy strand is generated from the common anchor site by a uracil-resistant polymerase, thereby providing high molecular redundancy for the single-tag intermediate to minimize template loss during the next round of linker attachment (Tail-Tag.2). Figure 1 ).
[0264] To estimate the template recovery rate of the sequencing method used in this application for methylation sequencing data, this application first compared its ligation efficiency with that of the conventional method (TruSeq) using ddPCR. Figure 6 The ratio of ligated to total DNA copies was 82% (Tail-Tag.1) and 86% (Tail-Tag.2) for the sequencing method used in this application for methylation sequencing data, and 64% (Tail-Tag.2) for TruSeq. Figure 6(CE). Considering that both methods require two rounds of connection, the recovery rate is almost doubled by applying step (ii) alone (0.64*0.64 vs. 0.82*0.86). Previous work has shown that the concepts employed in steps (i) and (iii) can increase template retention, so it is expected that the combination of all these steps can significantly improve library complexity. Figure 7 The paper summarizes the sequencing method of the methylation sequencing data in this application and compares it with other library preparation methods.
[0265] To further evaluate its performance, this application compared the sequencing method of the methylation sequencing data presented herein with two commercial kits: Accel Methyl-Seq (SWT) and NEBNext Ultra (NEB). It was found that the whole-genome bisulfite sequencing (WGBS) library constructed using the sequencing method of the methylation sequencing data presented herein showed a 10-fold increase in yield and exhibited the highest number of unique molecules regardless of input volume or sequencing depth. Figure 2 Furthermore, for inputs as low as 500 pg, the sequencing method for methylation sequencing data presented in this application demonstrated the highest methylome coverage, very small amplification bias, and highly reproducible methylation levels. Figure 2 D, Figure 8 (AE). It is worth noting that although the read lengths undergo rigorous adapter and tail removal, residual synthetic sequences and incomplete extension of the copy strand can still alter the start or end positions of DNA fragments, leading to incorrect estimates of library complexity. Therefore, this application employs "conditional pruning" and "shift correction" in data processing to overcome these challenges (Method section of Example 1). Through simulation- and experimental analysis, this application found that >90% of fragments could be correctly identified by the current informatics strategy of this application, approaching the level of UMI-assisted measurements (Supplementary Method section of Example 1, Table 2).
[0266] Traditional BS-seq introduces a significant amount of technical artifacts, hindering its use for rare allele detection. To address this issue, this application utilizes deep sequencing to suppress errors within PCR repeat families. This application first applies the sequencing method described in this application to unmethylated phage λDNA using methylated sequencing data to measure genome-wide "technical noise" (C / C+T in read 1, G / G+A in read 2). Compared to previous studies, the error rate of the sequencing method described in this application using methylated sequencing data (maximum 0.0025, average 0.0017) is reduced by almost 10-fold, regardless of the number of sequencing cycles (…). Figure 8 F). The non-reference allele frequencies for each substitution type (0.00003–0.00135) appear to be comparable to those described for the Illumina platform. Figure 2 EF, Figure 8 (G). Recently, a method called EM-seq, characterized by a mild enzymatic conversion, has been reported, allowing for library preparation from a small input. Analysis of human methylome data revealed that both EM-seq and the sequencing method for methylation sequencing data presented in this application exhibit superior performance in terms of library complexity and coverage uniformity, with EM-seq showing less AT bias and the sequencing method for methylation sequencing data presented in this application demonstrating lower graphical noise (two unconverted Cs in tandem, Table 2). The full capabilities of this new technique remain to be tested.
[0267] Team design and target capture performance
[0268] Because deep sequencing of the human full methylome would be prohibitively expensive, this application focuses on epigenetic changes associated with common cancers by analyzing data from previous studies. Figure 3 A). Therefore, 80,672 CpG sites were selected, spanning approximately 1.05 MB of genomic region. The group demonstrated significant enrichment in CpG islands, consistent with the reported role of DNA methylation in transcriptional regulation during tumorigenesis. Figure 9 AB).
[0269] The performance of targeted sequencing was evaluated using human lymphocyte DNA (NA12878) and plasma samples. Quite uniform capture of the amplified DNA fragments was achieved with as little as 2 ng of cfDNA, with 60–80% of reads aligned only to the bait region (target ratio) and >90% of the bait region covered by >200% of the reads (uniformity). Figure 3 B, Figure 9 C, Table 2). The read coverage of the GC content plot reveals a typical unimodal distribution ( Figure 9 D) and the estimated base call accuracy is above 99.9% for most (>0.94) bases (Phred score >30). Figure 9 E, Table 2). The sequenced cfDNA fragments had a single nucleosome peak length of ~160 bp, which is consistent with the results of the traditional method (TruSeq). Figure 3 C). Furthermore, negligible amplification or capture bias was observed for fragments associated with single and binucleate bodies. Figure 9 F). To assess common capture biases, this application calculated the percentage of methylated cytosine at each individual CpG site (iAF, frequency of a single methylated allele) and found highly correlated values on both the positive and negative strands (ρ = 0.90, F). Figure 3 D). This is confirmed by the almost balanced read depth of the positive chain to the negative chain ( Figure 9G). The typical technical error rate after target enrichment is 0.0012, similar to the error rate observed using the capture-free method. Figure 3 E).
[0270] The detection method MERMAID in this application identifies signals through single-molecule derivatization patterns.
[0271] Traditional DNA methylation analysis is largely based on iAF, which is sensitive to sampling variance and technical noise. In response, the methylation detection method MERMAID of this application designs a metric, the “block index” (BI), to separate CpG sites with similar methylation states into distinct blocks. Figure 4 A). A total of 8312 blocks were defined, with a median block size of ~143 bp and an average of ~13 CpG sites / block. Figure 10 This application defines the average methylation level in each block as mAF (mean methylated allele frequency) and compares its performance with iAF in detecting cancer-related changes. By examining SHOX2, a gene frequently methylated in LC, this application found that mAF showed a significantly higher AUC value than iAF, indicating that "block" is a more distinguishable unit than "site". Figure 4 B).
[0272] The advantage of MERMAID is the significantly improved separation of biological signals and technical noise. Analysis of each DNA fragment reveals distinct patterns for cancer cells and normal cells, whereas chemical or sequencing errors are typically sporadic. Figure 10 E). To highlight the contrast, this application developed a new unit of measurement, “methylated block score” (MBS), which is defined as the ratio of the weighted occurrence rate of consecutive methylation patterns to the total number of CpG sites per read (E). Figure 4 C). To evaluate the performance of MBS, this application mixed DNA treated with in vitro methyltransferases into λDNA in different proportions. For example, in Figure 4 As shown in D. With MBS, even samples with tiny spikes (0.001) can be clearly distinguished from negative controls, while with mAF, significant signal overlap is observed, thus supporting the success of pattern recognition in improving the signal-to-noise ratio.
[0273] like Figure 17Traditional DNA methylation analysis primarily relies on the methylation level (iAF) at individual sites, which is highly sensitive to sequencing depth and technical noise. Therefore, this application designs an index called the "block index" to separate CpG sites into different blocks based on their methylation status and genomic location. This application defines the average methylation level of each block as the mean methylated allele frequency (mAF) and compares it with the performance of iAF in detecting cancer-related changes. By detecting SHOX2, a gene frequently exhibiting methylation variations in lung cancer patients, this application found that the area under the receiver operating characteristic (AUROC) curve (mAF) was significantly higher than that of iAF. This indicates that "blocks" are more discriminative units than "sites," and that mAF performs better than iAF at lower sequencing depths.
[0274] like Figure 19 As shown, this application also compares the advantages of the MBS statistic compared to the traditional mAF statistic. Using a set of blending data between negative and positive standards, density curves of the MBS and mAF statistics were plotted for samples with different blending ratios. As can be seen from the figure, the MBS statistic shows a greater difference than the mAF statistic between the "negative standard" and the "positive standard with a blending ratio of 0.1%", indicating that the MBS statistic is better able to distinguish differential methylation signals.
[0275] Evaluation of MERMAID for tumor-origin detection signals
[0276] The analytical performance of MERMAID was examined at both the assay and bioinformatics levels. Highly correlated MBS values (ρ = 0.94–0.97) were observed in healthy individuals (including repetitive sequences), suggesting excellent reproducibility. Figure 5 A). Similar results were also observed with different input volumes (2-30 ng), sequencing depths (1000-5000X), or DNA sources (from cells or plasma). Figure 11 However, the differences between the healthy group and the cancer group, as well as among different cancer patients, were highly significant (ρ = 0.60–0.83), suggesting that the methylation profile was greatly influenced by disease status. Figure 5 A).
[0277] The quantitative accuracy of MERMAID was evaluated using tumor cell spiking assays. In a one-dilution series of colorectal cancer (CRC) DNA compared with normal WBC DNA, a near-perfect correlation was found between the observed and expected tumor proportions at a dilution of 0.0005. Figure 5 B,r 2=0.99), implying precise quantitative accuracy. Highly consistent results were also observed when this application compared MERMAID with quantitative methylation-specific PCR (qMSP) or different target methods. Figure 12 ).
[0278] The false discovery rate (FDR) of this assay was empirically evaluated by repeatedly sequencing normal white blood cells (WBCs). As shown in Figure 5C and Table 2, the FDR steadily decreased with increasing DNA input or sequencing depth. This is likely because markers with low methylation counts generated most of the false calls, and these markers benefited more from the reduction in Poisson noise. Figure 13 AB).
[0279] This application evaluated the detection limit (LoD) of MERMAID at three different levels: (i) numerical simulation: by simulating methylation counts from a binomial distribution, this application found that increasing the number of biomarkers or sequencing depth improved the success rate (sensitivity). Figure 13 C). Notably, the marker size expanded dramatically when the ctDNA fraction decreased from 1 / 10,000 to 1 / 100,000, revealing that tumor burden is a fundamental limiting factor for detection sensitivity. (ii) Bioinformatics simulation: This theory was further tested by calculating the mixing of sequencing reads from cancer cells with those from healthy cfDNA at different ratios. For example, in Figure 5 As shown in D, the bioinformatics sensitivity of MERMAID reached 1 / 100,000 in both simulated data from LC and CRC (two-tailed t-test, P < 0.001). This observation appears to be batch-independent, as the pooled control data from different rounds of sequencing did not produce significant detection (two-tailed t-test, P > 0.05). Figure 13 (D) (iii) Experimental evaluation: LC and CRC DNA were spiked into normal WBC DNA at a series of dilutions. Cancer signals quantified by MBS were detected in both cell lines at dilutions as low as 1 / 10000 (detection rate) (two-tailed t-test, P < 0.01). Figure 5 E).
[0280] Finally, this application compares the LoD of MERMAID with ddPCR and ultra-deep mutation sequencing with unique molecular specifiers (HS-UMI), two methods that are exceptions when detecting variants at extremely low frequencies. For each diluted sample, approximately 9000 copies of the human haploid genome were used to mimic the average amount of cfDNA in 10 ml of blood. For predefined hotspot mutations, EGFR p.G719S was identified in all 1 / 1,000 dilutions by ddPCR and HS-UMI at 0.03–0.11% AF. Figure 5 F, Figure 13 E, Table 4). Detection of the carcinogenic EML4-ALK fusion in a 1 / 1,000 dilution was validated by ddPCR at AF of 0.03%, but failed with HS-UMI ( Figure 5 F, Figure 13 F (Table 4). Depending on read depth and background noise, including mutations with uncertain functional significance, the maximum dilution ratio was not increased to more than 1 / 1000 (Table 4). In summary, under the same conditions, MERMAID showed at least 10-fold greater capability than mutation analysis.
[0281] The effectiveness of the MERMAID methylation sequencing method in this application
[0282] Methylation sequencing has attracted considerable interest due to its significant potential to improve current ctDNA detection. This application presents MERMAID as a novel epigenetic analysis method characterized by well-preserved molecular diversity, strong noise suppression, and robust high-dimensional modeling. In addition to these properties, MERMAID is particularly useful for blood-based applications: (i) a portion of cfDNA may be in single-stranded form, thus this ssDNA-compatible method maximizes the use of limited starting materials, increasing the chances of detecting rare ctDNA. (ii) capture groups are designed with excess long RNA probes (>100 nucleotides long) complementary to various methylation patterns. This strategy is more tolerant of sequence-related biases and polymorphisms compared to amplicon-based target methods (~20 nucleotides long). (iii) MERMAID does not require prior knowledge of the analysis (e.g., tissue examined via biopsy), thus providing a solution for patients without surgically removed samples. While this method has only been validated on LC, it can be customized for other types of cancer (e.g., CRC) or bodily fluids (e.g., urine). It can be extended to answer basic questions, such as tumor heterogeneity, or applied to other clinical scenarios, such as evaluating treatment effectiveness.
[0283] A drawback of introducing synthetic sequences to improve adapter labeling efficiency is the failure to protect the natural ends of the DNA template, which may lead to incomplete duplication removal. The computational strategy employed by MERMAID can, at least partially, overcome this problem of incomplete duplication removal. Simultaneously, MERMAID can employ bias handling methods commonly used in the art to address the associated risk of C->T / G->A artifacts arising from the conversion of cytosine to uracil after oxidative stress. The MERMAID of this application can also incorporate tissue-specific biomarkers into target groups for multiple cancer classification.
[0284] Supplementary Table
[0285] Table 1-1 Oligonucleotide sequences used in the methods of this application
[0286]
[0287]
[0288] Table 1-2 Oligonucleotide sequences used in the method + UMI correction of this application
[0289]
[0290] Table 1-3 Oligonucleotide sequences used for ligation efficiency measurement (ddPCR)
[0291]
[0292] Table 2-1 Full methylome sequencing metrics (depth titration) of ELSA, SWT, and NEB using 500 pg of E. coli DNA.
[0293]
[0294] Table 2-2 Method of this application and Method of this application + UMI correction (data simulation)
[0295]
[0296]
[0297] Table 2-3 Method of this application and Method of this application + UMI correction (experiment)
[0298]
[0299] Table 2-4 EM-seq and the methods of this application
[0300]
[0301] Note 1: Sequencing data were downloaded from the NCBI Sequence Reading Archive (SRC) containing PRJNA591788 and PRJNA534206. Note 2: To objectively compare these methods, all data were downsampled to ~100M read length.
[0302] Table 2-5 Quality control indicators for the methods of this application using 2 to 30 ng human cfDNA input.
[0303]
[0304]
[0305] Table 2-6 Quality control indicators for the methods of this application using 2 to 30 ng of human WBC input.
[0306]
[0307] Table 4-1 shows the detection of hotspot mutations using ddPCR with titrated SW48 and H2228 cancer cells (biological replicates).
[0308]
[0309]
[0310]
[0311]
[0312]
[0313] Table 4-2 Hotspot mutation detection of SW48 and H2228 cancer cells (biological replicas) using HS-UMI titration.
[0314]
[0315]
[0316]
[0317]
[0318] Table 4-3 HS-UMI Detection of Non-Hotspot Mutations Using Titrated SW48 and H2228 Cancer Cells (Bioreplicates)
[0319]
[0320]
[0321]
[0322]
[0323] Table 4-4 QC indicators for HS-UMI using spike-in experiments with cancer cell lines (NCI-H2228, SW48)
[0324]
[0325]
[0326]
[0327] Table 5-1 Organizational Classification of "Specifiers" (n=157)
[0328]
[0329]
[0330]
[0331] Table 5-2 "Specifier" Plasma Classification (n = 2,473)
[0332]
[0333]
[0334]
[0335]
[0336]
[0337]
[0338]
[0339]
[0340]
[0341]
[0342]
[0343]
[0344]
[0345]
[0346]
[0347]
[0348]
[0349]
[0350]
[0351]
[0352]
[0353]
[0354]
[0355]
[0356]
[0357]
[0358]
[0359]
[0360]
[0361]
[0362]
[0363]
[0364]
[0365] Table 6 Multivariate analysis of ctDNA detection using the method described in this application
[0366]
[0367] Table 7, Group 2, Summary of "Fourfold Comparison"
[0368] Table 7-1 Summary of HS-UMI and ddPCR detection of ctDNA in “Group 2 - Quadruple Comparison”
[0369]
[0370]
[0371]
[0372]
[0373]
[0374] Table 7-2 Summary of ctDNA mutations detected by ddPCR in “Group 2-Quadruple Comparison”
[0375]
[0376]
[0377]
[0378]
[0379]
[0380] Table 7-3 shows the ddPCR detection used in "Group 2 - Quadruple Comparison".
[0381]
[0382] The foregoing detailed description is provided by way of explanation and example and is not intended to limit the scope of the appended claims. Various variations of the embodiments listed herein will be apparent to those skilled in the art and are reserved within the scope of the appended claims and their equivalents.
Claims
1. A method for detecting methylation modification of a target nucleic acid, the method comprising the following steps: step (a-1) determining a comethylation block based on the correlation coefficient of CpG sites in the target nucleic acid, the methylation level of the CpG sites, and the position information of the CpG sites; and / or step (a-2) determining the comethylation block based on the correlation coefficient of CpG sites in the target nucleic acid, the information content of candidate comethylation blocks, and the degree of balance in the division of the candidate comethylation blocks; and step (b) determining the presence and / or content of the target nucleic acid in the sample to be tested based on the methylation degree of the comethylation block; Step (b) includes: determining the methylation block score (MBS) of the comethylated block, wherein the MBS is calculated using the following formula: in, The n is the total number of reads covering all CpG sites in the comethylation block, L i It is the number of CpG sites contained in the i-th read, l ij is the length of consecutive methylated CpG sites on the i-th read, and m is the sequencing depth on the i-th read; The presence and / or content of the target nucleic acid are determined based on MBS of the comethylated blocks in the sample to be tested.
2. The method of claim 1, wherein the method detects the presence and / or content of the methylated target nucleic acid in the sample to be tested.
3. The method according to any one of claims 1-2, wherein the correlation coefficient of the CpG sites comprises the Pearson correlation coefficient between two or more CpG sites in the target nucleic acid.
4. The method of any one of claims 1-2, wherein the methylation level of the CpG site comprises the difference in average methylation allele frequency (mAF) between two or more CpG sites in the target nucleic acid.
5. The method of any one of claims 1-2, wherein the methylation level of the CpG site comprises the ratio of the difference in mAF between two or more CpG sites in the target nucleic acid to the sum of mAF.
6. The method of any one of claims 1-2, wherein the location information of the CpG sites includes differences in genomic location between two or more CpG sites in the target nucleic acid.
7. The method of any one of claims 1-2, wherein the location information of the CpG sites comprises the ratio of the genomic location distance between two or more CpG sites in the target nucleic acid to the length of the target nucleic acid.
8. The method of any one of claims 1-2, wherein step (a-1) comprises: determining the corrected correlation coefficient between every two CpG sites of the target nucleic acid, wherein the corrected correlation coefficient between site i and site j Calculated using the following formula: ,in, This represents the Pearson correlation coefficient. The average methylated allele frequency (mAF) represents the frequency of all samples at site i. The average methylated allele frequency (mAF) represents the frequency of all samples at position j. Indicates the genomic location of site i. λ1 represents the genomic location of site j, L represents the length of the target nucleic acid region, and λ1 and λ2 are independently selected from numbers 0 or greater.
9. The method as described in claim 8, wherein the value of λ1 ranges from 0 to 1.
10. The method as described in claim 8, wherein the value of λ2 ranges from 0 to 1.
11. The method of any one of claims 1-2, wherein the correlation coefficient based on CpG sites in the target nucleic acid in step (a-2) comprises a corrected correlation coefficient of CpG sites in the target nucleic acid, and the corrected correlation coefficient of CpG sites in the target nucleic acid comprises a Pearson correlation coefficient between two or more CpG sites in the target nucleic acid corrected based on the methylation level of the CpG sites and / or the position information of the CpG sites.
12. The method of claim 11, wherein the corrected correlation coefficient of the CpG site in the target nucleic acid comprises the corrected correlation coefficient of the method of any one of claims 8-10.
13. The method of any one of claims 1-2, wherein the information content of the candidate comethylation block includes information on the number of CpG sites in the candidate comethylation block.
14. The method of any one of claims 1-2, wherein the degree of balance in the partitioning of the candidate comethylation blocks includes the difference in the number of CpG sites in the different candidate comethylation blocks.
15. The method of any one of claims 1-2, wherein the degree of balance in the partitioning of the candidate comethylation blocks comprises the coefficient of variation of the number of CpG sites in the different candidate comethylation blocks.
16. The method of any one of claims 1-2, wherein step (a-2) comprises: maximizing the block index of the candidate comethylation block, determining the comethylation block, and the candidate comethylation block in the target nucleic acid. The block index Calculated using the following formula: in, , The number of CpG sites in the i-th candidate comethylation block is represented by α1 and α2, which are independently selected from numbers of 0 or greater.
17. The method of claim 16, wherein the block breakpoint of the comethylated block is determined by an iterative method with a unique breakpoint.
18. The method of claim 16, wherein the value of α1 ranges from 0 to 10.
19. The method of claim 16, wherein the value of α2 ranges from 0 to 10.
20. The method of any one of claims 1-2, wherein UMI correction is applied to the sequencing data aggregate of the sample to be tested.
21. The method according to any one of claims 1-2, further comprising: extracting the feature values of the MBS of the comethylated blocks in the tumor sample and the healthy sample using a machine learning model, and determining the presence and / or content of the target nucleic acid based on the MBS of the comethylated blocks in the sample to be tested.
22. An analytical device for detecting methylation modification of a target nucleic acid, the device comprising: a block partitioning module (a-1) for determining comethylated blocks based on the correlation coefficient of CpG sites in the target nucleic acid, the methylation level of the CpG sites, and the positional information of the CpG sites; and / or a block partitioning module (a-2) for determining the comethylated blocks based on the corrected correlation coefficient of CpG sites in the target nucleic acid, the information content of candidate comethylated blocks, and the partitioning balance of the candidate comethylated blocks; and a determination module (b) for determining the presence and / or content of the target nucleic acid in the test sample based on the methylation degree of the comethylated blocks; Step (b) includes: determining the methylation block score (MBS) of the comethylated block, wherein the MBS is calculated using the following formula: in, The n is the total number of reads covering all CpG sites in the comethylation block, L i It is the number of CpG sites contained in the i-th read, l ij is the length of consecutive methylated CpG sites on the i-th read, and m is the sequencing depth on the i-th read; The presence and / or content of the target nucleic acid are determined based on MBS of the comethylated blocks in the sample to be tested.
23. A method for detecting ctDNA using methylation sequencing, comprising: Select differentially methylated CpG sites; Based on the similarity of the methylation states of the CpG sites, the CpG sites are divided into multiple comethylation blocks; The sample was sequenced to obtain methylation sequencing reads; The average methylation level of each comethylated block in the sample was detected for further DNA methylation analysis. The method further includes using a methylated block score (MBS) to normalize the depth difference of each methylated block, which is used to distinguish very small tumor signals with tumor fractions of 0.1%, 0.2%, 0.5%, or 1%. The MBS mentioned therein is: For a given block, n is the total number of reads covering multiple CpG sites; L i It is the number of CpG sites covering the i-th read; l ij represents the length of consecutive methylated CpG sites > 1, and m is the total count on the i-th read.
24. The method of claim 23, wherein the differentially methylated CpG sites are selected from the TCGA database generated by the Infinium HumanMethylation 450K array.
25. The method according to any one of claims 23-24, wherein the average methylation level is the average methylation allele frequency.
26. The method according to any one of claims 23-24, wherein the number of comethylated blocks is between 1 / 30 and 1 / 5 of the number of CpG sites.
27. The method of any one of claims 23-24, wherein the comethylated blocks are further restricted by comparing a specific tumor sample with a normal tissue sample.
28. The method according to claim 27, wherein, The comparison of specific tumor samples includes primary lung tumor samples, and the comparison of normal tissue samples includes normal lung tissue samples.
29. The method of claim 27, wherein in the majority of samples (>80%), comethylated blocks with sequencing depth <100X are excluded from downstream analysis.
30. The method of any one of claims 23-24, wherein comethylated blocks are separated based on an improved correlation matrix referred to as a "block index".
31. The method of claim 30, wherein (1) The iAF correlation matrix of each CpG site in region r is expressed as follows: The correlation coefficient between site i and site j is calculated as follows: ; Represents the Poisson correlation coefficient. This represents the mAF of all samples at position i. λ1 and λ2 represent the genomic location, L represents the length of the original region, and λ1 and λ2 are parameters estimated using prior information. (2) For the new split within the original region r The block index is defined as the following function: , Wherein represents Let α1 and α2 represent the number of sites in the i-th block, and let α1 and α2 be the penalty coefficients for unbalanced splitting and over-splitting, respectively. In this study, based on the desired block length and uniformity, both α1 and α2 are set to 1.
0. (3) Repeat the partitioning process by iterative division with a unique breakpoint; assuming k breakpoints, then by making To reach the maximum, select the (k+1)th new breakpoint; The interactive algorithm terminates when the (k+1)th breakpoint matches the kth breakpoint.
32. The method of any one of claims 23-24, wherein the sequencing reads to be analyzed are trimmed by any one or more of the following criteria: (i) Query the number of segments of G bases longer than a fixed value m; (ii) The fraction of non-G-bases is less than the fixed number p; (iii) The next base is a high-quality A / T / C, in which the Phred score of the high-quality A / T / C is >30.
33. The method of any one of claims 23-24, wherein repeated removal is applied at the starting points of both R1 and R2 with a tolerance of + / -3bp to minimize artifacts associated with the improperly assigned end positions of segments.
34. The method of any one of claims 23-24, wherein UMI is applied during correction.
35. The method of claim 34, wherein the UMI-assisted correction comprises: (1) Template generation: A library of 2kb DNA segments is generated and "cut" into various lengths by calculation; the average of each length is 170 base pairs and the standard deviation of each length is 30 base pairs; (2) Tail addition: Based on empirical observation, tail sequences of different lengths were added to the 5'-end of R2, and the proportion of guanine (G) bases in the tail sequences was 90%; (3) UMI incorporation: A random 6-base UMI tag is added to the 5'-end of R2 to mark the original template; (4) Technical errors: Three datasets were generated with low / medium / high substitution error rates of 0.005, 0.01 and 0.02, which are typical for Illumina sequencers; all base error directions were set with equal probability; insertion / deletion errors of 1 base were added or removed in the tail region to take into account homopolymer properties; the accumulation of PCR errors was simulated by 10 rounds of PCR reactions with a repetition probability of 0.9 per round; incomplete extension was simulated by random position fluctuations of less than 10 bases at the 5' end of R1. (5) UMI and non-UMI solutions: a) UMI: Apply UMI extraction to the 5' end of each R2, while accepting shift corrections within 3 bases; b) Non-UMI: Perform a shift correction of up to 3 bases on the 5' ends of both R1 and R2; (6) Template counting: The depth of the original bisulfite-converted fragment is set to 250-500X, while the depth of the original fragment is set to 500-1000X.
36. The method according to any one of claims 23-24, wherein the average methylation level of each comethylated block of unmethylated phage λ DNA is detected to measure genome-wide "technical noise"; wherein, For read length 1, calculate the proportion of cytosine (C) to the sum of cytosine and thymine (C+T). For read length 2, calculate the proportion of guanine (G) to the sum of guanine and adenine (G+A).
37. The method of any one of claims 23-24, wherein a machine learning classifier of methylation patterns is applied to assess tumor levels, said assessment of tumor levels being used in early tumor screening.
38. A kit for detecting ctDNA using methylation sequencing, wherein the kit is used to capture at least 50, 100, 150, 200, 300, 500, 800, 1000, 1500 or 2000 comethylated blocks as shown in Table 5-2.
39. The kit according to claim 38, wherein, Use the kit to capture at least all comethylated blocks as shown in Table 5-1.
40. Use of the method as described in any one of claims 1-21 and 23-37 in assessing tumor levels.
41. The use as described in claim 40, wherein the use includes assessing the tumor level in early tumor screening.
42. The use as described in claim 40, wherein the tumor is derived from homogeneous tumors, heterogeneous tumors, hematologic malignancies, and / or solid tumors.
43. The use according to claim 40, wherein the tumor is derived from one or more cancers from the group consisting of: brain cancer, lung cancer, skin cancer, nasopharyngeal carcinoma, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, ovarian cancer, esophageal cancer, gallbladder cancer, bile duct cancer, breast cancer, cervical cancer, uterine cancer, prostate cancer, head and neck cancer, sarcoma, thoracic malignant tumors other than lung cancer, melanoma, and testicular cancer.
44. A storage medium containing a program for performing the method according to any one of claims 1-21 and 23-37.
45. An apparatus comprising the storage medium of claim 44, and optionally comprising a processor coupled to the storage medium, the processor being configured to execute, based on a program stored in the storage medium, the method of any one of claims 1-21 and 23-37.
Citation Information
Patent Citations
Compositions and methods for preparing nucleic acid libraries
WO2019191900A1
Compositions and methods for preparing nucleic acid libraries
WO2019192489A1
Interpretation of genetic and genomic variants via an integrated computational and experimental deep mutational learning framework
CN111095422A
Methylated biomarker for detecting breast cancer and application thereof
CN112375822A