Nucleic acid sequence analysis methods

By spatially separating and ligating 5' and 3' terminal ends of DNA reads with a common linker, the method addresses sequencing errors in non-overlapping reads, ensuring accurate high-throughput analysis of clonal populations.

JP2026062846APending Publication Date: 2026-04-10INVIVOSCRIBE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INVIVOSCRIBE INC
Filing Date
2025-12-25
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Current high-throughput bidirectional sequencing methods face limitations due to non-overlapping sequence reads, leading to sequencing errors and misclassification, especially in analyzing clonal populations like neoplastic states, as they rely on overlapping 3' ends which are prone to errors and length variations, distorting test results.

Method used

The method involves spatially separating template DNA molecules on a solid support, amplifying them to form amplicon clusters, and ligating 5' and 3' terminal ends of forward and reverse reads with a common linker, ensuring the target sequence is within 80% of the read length, allowing accurate alignment and analysis without requiring full-length overlap.

Benefits of technology

This approach enhances the accuracy and reliability of sequencing results by mitigating sequencing errors and misclassification, enabling high-throughput analysis of non-overlapping reads, particularly in detecting clonal populations like rearranged immunoglobulin and T cell receptor gene segments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062846000010
    Figure 2026062846000010
  • Figure 2026062846000011
    Figure 2026062846000011
  • Figure 2026062846000012
    Figure 2026062846000012
Patent Text Reader

Abstract

The problem that this invention aims to solve is to provide a method for analyzing the nucleotide read sequence of a target nucleic acid sample using high-throughput bidirectional sequencing. [Solution] The present invention relates to a method that works even when bidirectional sequencing produces forward and reverse reads that are not long enough to be paired via complementary hybridization of overlapping sequences at the 3' end of the sequence reads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority from U.S. Provisional Application No. 62 / 953,270, filed on December 24, 2019, the entire content of which is incorporated herein by reference.

[0002] Field of the Invention The present invention generally relates to methods for analyzing the nucleotide sequence of a nucleic acid sample of interest, more specifically, to methods for analyzing the nucleotide sequence of a nucleic acid sample of interest using high - throughput bidirectional sequencing. The methods of the present invention are based on the determination that accurate alignment and analysis of sequencing results can be facilitated when, even when bidirectional sequencing yields forward and reverse reads that are not of sufficient read length for complementary hybridization of overlapping sequences at the 3' ends of the sequence reads to pair, the 3' terminal ends of the sequence reads are removed and defined portions of the 5' ends of co - localized forward and reverse sequence reads are ligated via a nucleic acid linker common to all ligated reads. The development of the methods of the present invention is useful in a variety of applications including, but not limited to, the diagnosis of conditions characterized by the presence of a clonal population of cells (such as neoplastic states, etc.) or microorganisms, the monitoring of the progression of such conditions, the prediction of the likelihood of recurrence of a subject from a remission state to a diseased state, the evaluation of the effectiveness of existing therapeutic agents and / or new therapeutic agents, or immune monitoring.

[0003] Incorporation by Reference of a Sequence Listing The sequence listing, an ASCII text file named 38093WO.P41235PCUS.SeqListing.txt, created on December 16, 2020 and submitted to the United States Patent and Trademark Office via EFS - Web, is incorporated herein by reference.

Background Art

[0004] No reference in this Specified Publication to any prior publication (or information derived therefrom) or any publicly known matter shall be construed, and should not be construed, as an acknowledgment, recognition, or suggestion in any form that the prior publication (or information derived therefrom) or publicly known matter constitutes part of the common general knowledge in the field addressed herein.

[0005] The bibliographic details of the publications mentioned by the authors in this specification are listed alphabetically at the end of the description.

[0006] Clones are generally understood as a group of cells that share a common progenitor cell lineage. Diagnosing and / or detecting the presence of clonal populations of cells or organisms in a subject has generally constituted a relatively problematic procedure. Specifically, clonal populations can constitute only a small component within a larger population of cells or organisms. For example, in mammalian organisms, one more common situation requiring the detection of clonal populations of cells occurs in relation to the diagnosis and / or detection of neoplasms such as cancer. However, the detection of one or more clonal populations can also be important in the diagnosis of conditions such as spinal dysplasia or polycythemia vera, and in the detection of antigen-derived clones produced by the immune system in situations of infection, autoimmune diseases, allergies, or transplantation.

[0007] When clone members are characterized by molecular markers such as modified DNA sequences, the detection problem can be rephrased as detecting a population of molecules with all identical molecular sequences within a larger population of molecules with different sequences. The level of detection of marker molecules that can be achieved depends heavily on the sensitivity and specificity of the detection method, but almost always, as the proportion of the target molecule in a larger population of molecules decreases, it becomes difficult to detect the signal from the target molecule due to signal noise from the larger population.

[0008] A special class of molecular markers, highly specific but exhibiting unique complexities in their detection, arise from genetic recombination events. Recombination of genetic material in somatic cells involves the merging of two or more initially separate regions of the genome. This can occur as a random process, or as part of the developmental process in normal lymphocytes.

[0009] In relation to cancer, recombination can be simple or complex. Simple recombination can be considered as the juxtaposition of two unrelated genes or regions. Complex recombination can be considered as the recombination of more than two genes or gene segments. A classic example of complex recombination is the rearrangement of immunoglobulin and T cell receptor variable genes, which occurs during the normal development of lymphocytes and involves recombination of the V, D, and J gene segments. The loci for these gene segments are widely separated in the germline, but recombination during lymphocyte development results in the juxtaposition of the V, D, and J gene segments, or the V and J gene segments, and the junctions between these gene segments are characterized by small regions of nucleotide insertions and deletions (N1 and N2 regions). Because this process occurs randomly, each normal lymphocyte will end up with a unique V(D)J rearrangement, which can be a complete VDJ rearrangement or a VJ or DJ rearrangement, depending on both the genes being rearranged and the nature of the rearrangement. Lymphoid cancers such as acute lymphoblastic leukemia, chronic lymphocytic leukemia, lymphoma, or myeloma arise as a result of neoplastic changes in a single normal cell; therefore, all cancer cells have, at least initially, the V(D)J rearrangement at the junction originally present in the founder cell. Subclones can arise during the expansion of the neoplastic population, and further V(D)J rearrangements can occur in them.

[0010] Unique DNA sequences arising from recombination and present in cancer clones or subclones provide unique genetic markers that can be used to monitor the response to treatment and make treatment decisions. Clonal monitoring can be performed by various techniques, including PCR, flow cytometry, or next-generation sequencing, each of which exhibits different advantages and disadvantages.

[0011] PCR revolutionized DNA analysis thanks to its ability to exponentially amplify target DNA, especially DNA present at low start copy numbers, but conventional sequencing methods such as Sanger sequencing remained time-consuming. Therefore, large-scale sequencing-based analysis of PCR-amplified patient DNA was virtually impossible. The emergence of next-generation sequencing revolutionized sequencing-based analysis by providing a high-throughput approach to DNA sequencing. This reduced the turnaround time and cost associated with conventional sequencing, making nucleic acid sequencing available on a large scale. Coupled with the evolution from PCR to solid-phase bridge amplification-based colony generation, the significantly more sophisticated, informative, and far more accurate information provided by nucleic acid sequencing analysis became routinely available.

[0012] There are various DNA library amplification methods and next-generation sequencing methods under development. For example, three of the more common PCR-based amplification methods are emulsion PCR, rolling circle amplification, and solid-phase amplification.

[0013] In emulsion PCR, a DNA library is first generated. Single-stranded DNA fragments are attached to the surface of beads using adapters or linkers, and one bead is attached to a single DNA fragment from the DNA library. The surface of the bead contains oligonucleotide probes having a sequence complementary to the adapter that binds to the DNA fragment. The beads are then compartmentalized within water-oil emulsion droplets. In the aqueous water-oil emulsion, each droplet capturing a single bead is a PCR microreactor that generates an amplified copy of the single DNA template.

[0014] Gridded Rolling Circle Nanoballs describe the amplification of a population of single DNA molecules by rolling circle amplification in solution, followed by the capture of spots smaller than the immobilized DNA on a grid.

[0015] DNA colony generation (bridge amplification) uses forward and reverse primers densely covalently bonded to a flow cell slide. The ratio of primers to template on the support defines the surface density of the amplified clusters. The flow cell is exposed to reagents for polymerase-based extension, and priming occurs when the free / distal ends of the ligated fragments "bridge" to complementary oligonucleotides on the surface. Repeated denaturation and extension result in localized amplification of DNA fragments at millions of separate locations across the flow cell surface. Solid-phase amplification generates 100-200 million spatially separated template clusters, providing free ends to which universal sequencing primers then hybridize to initiate the sequencing reaction.

[0016] Regarding next-generation sequencing approaches, four well-known techniques include pyrosequencing, reversible terminator chemistry sequencing, ligase-mediated ligation sequencing, and phospholinked fluorescent nucleotide sequencing.

[0017] Pyrosequencing is a non-electrophoretic bioluminescence method that measures the release of inorganic pyrophosphate by proportionally converting it to visible light using a series of enzymatic reactions. Unlike other sequencing approaches that use modified nucleotides to terminate DNA synthesis, pyrosequencing manipulates DNA polymerase by adding a limited amount of dNTPs in a single step. Upon incorporating complementary dNTPs, the DNA polymerase extends the primer and stops. DNA synthesis resumes after the addition of the next complementary dNTP in the dispensing cycle. The order and intensity of the light peaks are recorded as a flowgram, which reveals the underlying DNA sequence.

[0018] Reversible terminator chemistry sequencing uses reversible terminator-bound dNTPs in a periodic method involving nucleotide incorporation, fluorescence imaging, and cleavage. Fluorescently labeled terminators are imaged as each dNTP is added and then cleaved to allow for the incorporation of the next base. Each incorporation is a unique event because these nucleotides are chemically blocked. The imaging step follows each base incorporation step, after which the blocked group is chemically removed, preparing each strand for subsequent incorporation by DNA polymerase. This sequence continues for a specific number of cycles, determined by user-defined instrument settings. The 3' blocking group was initially considered to be achieved through enzymatic or chemical reversal. This method forms the basis of instruments from Solexa and Illumina. Reversible terminator chemistry sequencing can be performed as a four-color cycle, as used by Illumina / Solexa, or as a single-color cycle, as used by Helicos BioSciences. Helicos BioSciences uses "virtual terminators," which are unblocked terminators containing a second nucleoside analog that acts as an inhibitor. These terminators incorporate appropriate modifications to terminate or inhibit the base so that DNA synthesis terminates after a single base addition. Reversible terminator sequencing can be designed as bidirectional (paired-end) sequencing or single-read sequencing.

[0019] Ligation-mediated sequencing uses a sequence extension reaction performed by a DNA ligase and a single-nucleotide or double-nucleotide coding probe, rather than a polymerase. In its simplest form, a fluorescently labeled probe hybridizes with its complementary sequence adjacent to a primed template. DNA ligase is then added to ligate the dye-labeled probe to the primer. The unligated probe is washed away, and the identity of the ligated probe is determined by subsequent fluorescence imaging. This cycle can be repeated by removing the fluorescent dye and using a cleavable probe to regenerate the 5'-PO4 group for a subsequent ligation cycle (linked ligation), or by removing a new primer and hybridizing to the template (unlinked ligation).

[0020] Phosphoconjugated fluorescence nucleotide sequencing is a real-time sequencing method that involves imaging the continuous incorporation of dye-labeled nucleotides during DNA synthesis. A single DNA polymerase molecule is attached to the bottom surface of individual zero-mode waveguide detectors, which can obtain sequence information while the phosphoconjugated nucleotide is incorporated into the growing primer chain. For example, Pacific Biosciences uses a proprietary DNA polymerase that effectively incorporates phosphoconjugated nucleotides and enables rearrangement of closed circular templates.

[0021] These technologies are available on various commercial platforms, including those summarized in Table 1 below.

[0022] [Table 1A]

[0023] [Table 1B]

[0024] The combination of solid-phase bridge amplification of a target DNA followed by reversible dye terminator bidirectional sequencing has proven to be a particularly effective means of achieving high-throughput amplification and sequencing. However, one limitation to the usefulness of bidirectional sequencing is the maximum number of cycles that can be performed, which in turn limits the maximum sequence read length that can be generated. For example, the Illumina HiSeq instrument can generate 2×250 base bidirectional reads, while the MiSeq instrument can generate 2×300 base bidirectional reads. Both the NextSeq and NovaSeq instruments generate 2×150 base bidirectional reads. In the context of long DNA targets such as long sections of chromosomes or other genomes, the generation of relatively short reads, nevertheless, is useful because those reads can be paired (also referred to as "taped" or "stitched") based on the complementarity of sequences that overlap at their 3' ends, thereby generating double-stranded DNA sequence sections. Each of these taped sequences can then be further aligned based on sequence overlaps with other taped reads to assemble longer stretches of the genomic sequence. This alignment is often performed against a reference sequence. In this regard, where sequence reads do not overlap, the use of a reference sequence to align these reads can provide a means of analyzing the reads against the reference sequence. However, in the absence of sequence reads against which analysis can be performed, non-overlapping reads are of little utility outside of the context of any information that can currently be provided as individual independent sequencing results.

[0025] In the context of a specific DNA target region, such as a rearranged immunoglobulin (hereinafter referred to as "Ig") or T cell receptor (hereinafter referred to as "TCR") molecule, when each individual amplicon is analyzed to determine whether it represents one member of a population of clonal sequences in the biological sample of interest, or alternatively, whether it represents a residual or regenerated clonal sequence, it is usually necessary for bidirectional sequence reads to provide sufficient forward and reverse read lengths so that the 3' ends of the reads overlap and can be tapeted based on their complementarity, thereby providing an entire target sequence region, such as a rearranged VJ gene segment of a T or B cell, or a span of genomic DNA that may contain a mutation, chromosomal translocation, DNA break or inversion, or indel region. If the DNA region required to be amplified to detect this nucleotide feature is longer than what can be sequenced by the chemistry of the selected instrument, the bidirectional forward and reverse reads generated from the 5' and 3' terminal ends of such a template may not be long enough to overlap and therefore cannot be tapeed together. Therefore, the high-throughput instrumentation and methodologies currently available limit the types and scope of sequencing analyses that can be performed in the context of screening specific sequences or investigating the diversity of a target DNA population.

[0026] In the studies leading to the present invention, it has unexpectedly been found that even when bidirectional sequencing chemistry is insufficient to generate overlapping forward and reverse reads, it is possible to screen a DNA sample of interest to express one or more target nucleotide sequences by generating a template DNA library from an initial biological sample, and that regardless of the length of each individual template DNA molecule, the target nucleotide sequences are located at the 5' and 3' termini of the template DNA, specifically, the template is designed such that it is within a 5' or 3' terminal nucleotide stretch corresponding to approximately 80% of the length of the bidirectional sequence read length selected for use. Thus, the bidirectional sequencing step effectively sequences the target nucleotide sequence because the target nucleotide sequence is located in a region known to be within the range of the read length. These sequence reads do not contain a read length sufficient for the forward and reverse read lengths to overlap, but if they are generated from amplicons generated on a solid phase via cluster amplification of individual template DNA molecules themselves, the spatial co-localization of the reads provides a means of identifying potential bidirectional sequence read pairs.

[0027] However, due to the increased likelihood of sequencing errors as bidirectional sequencing reads progress in the 3' direction, these reads cannot be reliably aligned and analyzed using currently available analytical tools. This is because these tools rely on the hybridization of the overlapping 3' ends of paired reads to help distinguish between random sequencing errors and the presence of SNPs or point mutations. Furthermore, due to the fact that there is variability in the final sequence length between reads (not all amplicons are sequenced to the maximum theoretical read length for the selected instrument), these reads are unexpectedly misclassified as separate and distinct sequences, even when their actual sequences are otherwise identical across the generated sequence lengths. Thus, the combination of naturally occurring sequencing errors at the 3' ends of sequence reads, along with the misclassification of reads that are otherwise identical but of different lengths, significantly distorts the test results.

[0028] When conventional overlapping bidirectional sequencing reads are generated, both of the above problems are mitigated. Forward and reverse reads overlap and can hybridize based on the complementarity of the overlapping sequences, generating double-stranded molecules. 3' sequencing errors are easily identified and discarded (rather than being classified as unique sequences) by complementary paired terminal reads expressing the correct complementary nucleotides, making the issue of sequence length variation practically meaningless. Therefore, without the generation of overlapping sequence reads, analysis of non-overlapping reads in their original form has proven to yield substantially erroneous results, which can prove to be a significant problem in clinical settings.

[0029] With respect to the present invention, surprisingly, in addition to the specific template designs described herein, forward and reverse sequence reads are cleaved to remove the 3' sequence read to the point where the remaining read is at least 80% of the maximum bidirectional sequence read length selected for use, and the cleaved and colocalized forward and reverse bidirectional reads are linked to sequences complementary to the reverse and forward reads, respectively, to form a linear molecule via a linear linker sequence common to all paired colocalized reads, and when the resulting “tape” sequence reads are aligned with other reads and / or analyzed by other means, it has been found that highly accurate results are obtained regarding the presence, nature and / or diversity of target nucleotide sequences in the DNA sample of interest. Furthermore, in the context of immunoglobulin and TCR gene rearrangements, it has been found that even when 5' and 3' reads from two or more clusters are identical, there remains the possibility that these reads may be generated from two different template molecules, even if the target sequences are the same between these molecules but the intervening (unamplified) sequences are different. In this situation, these reads are classified as originating from a common clone. However, in the context of rearranged VDJ gene segments, it has now been found that the incidence of this sequencing anomaly does not actually adversely affect the sensitivity or specificity of the test results. By designing and generating template DNA libraries to ensure that the target sequence is localized to the 5' and 3' ends of the template molecule, it is now possible to perform high-throughput next-generation sequencing without necessarily ensuring that the template DNA library fragments are of a size that can be sequenced to their full length using selected bidirectional sequencing instruments. Thus, this development has greatly expanded the applications of current next-generation bidirectional sequencing chemistry and instrumentation, so that the length of the target DNA template is no longer limited by the maximum read length of a given instrument, provided that the appropriate instrumentation is selected. If the target sequence can be expressed within the 5' and 3' terminal DNA regions described herein, the total length of the DNA template into which the amplicon cluster is generated and sequenced becomes irrelevant and no longer limited.Furthermore, this method also enables the matching and analysis of non-overlapping sequence reads without requiring this process to be performed on the reference sequence to which each read is aligned. [Overview of the Initiative] [Means for solving the problem]

[0030] Throughout this specification and the following claims, unless the context otherwise indicates, the term “comprise,” and variations such as “comprises” and “comprising,” are understood to mean that they encompass the integer or process, or group or set of integers, but do not exclude any other integer or process, or group or set of integers.

[0031] The present invention is intended for illustrative purposes only and is not limited to the scope of the specific embodiments described herein. Functionally equivalent products, compositions, and methods are clearly within the scope of the present invention as described herein.

[0032] As used herein, the term “derived from” shall be interpreted as indicating that a particular integer or group of integers derives from a specified species, but is not necessarily directly obtained from a specified source. Furthermore, as used herein, the singular forms of “one,” “and,” and “it” shall include multiple referents unless the context explicitly indicates otherwise.

[0033] The specification of this subject includes nucleotide sequence information created using the program PatentIn version 3.1, which is presented herein after the bibliography. Each nucleotide sequence is represented numerically in the sequence listing. <210> and the subsequent array identifier (for example, <210> 1. <210> Identified by (2nd class). For each nucleotide sequence, the length, type, and source organism of the sequence (DNA, etc.) are, respectively, represented by a numerical field. <211> , <212> and <213> The information provided herein indicates the nucleotide sequences referred to herein are identified by the sequence number designation followed by a sequence identifier (e.g., Sequence Number 1, Sequence Number 2, etc.). Sequence identifiers referred to herein are shown in the sequence listings in the numerical field. <400> and the subsequent array identifier (for example, <400> 1. <400> It correlates with the information provided in (2) below. That is, Sequence ID No. 1, detailed herein, corresponds to the sequence listing. <400> It correlates with the sequence shown as 1.

[0034] One aspect of the present invention is a method for screening a target nucleic acid sample to express one or more target nucleotide sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the nucleic acid sample on a solid support, wherein the template DNA molecules are generated such that the target nucleotide sequence is localized in the adjacent nucleotide region at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and This applies to methods that include [specific methods].

[0035] In another embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a solid support, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0036] In yet another embodiment, a method for screening a DNA sample containing B and / or T cell DNA for expressing one or more rearranged V, D, or J gene segments, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a solid support, wherein the template DNA molecules are generated such that the rearranged V, D, or J gene segments are localized to adjacent nucleotide regions at the 5' and / or 3' terminal ends of the template, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequencing results; (v) Steps to analyze the sequence results and A method is provided that includes this.

[0037] In another embodiment, the adjacent nucleotide region in step (i) corresponds to approximately 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0038] In other embodiments and in the context of V(D)J rearrangements, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ. In another embodiment, the target nucleotide sequence is a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0039] In yet another embodiment, the target nucleotide sequence is a V gene segment region and / or a J gene segment region encoding a portion of CDR3, such as a hypermutation-prone region.

[0040] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a V region susceptible to somatic hypermutation, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0041] In yet another embodiment, the target nucleotide sequence is a BCL1 / JH translocation or a BCL2 / JH t(14:18).

[0042] In a further embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, and the adjacent nucleotide region corresponds to approximately 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0043] Preferably, the glass surface is a glass slide or a flow cell.

[0044] In yet another embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, the adjacent nucleotide region corresponds to about 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an index, barcode, unique molecular identifier, sequencing primer hybridization site and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0045] In another further embodiment, a method for screening a DNA sample of interest to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0046] In one embodiment, the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where up to 20 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0047] In another embodiment, the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, but up to 30 nucleotide terminals in the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0048] In a further embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an index, barcode, unique molecular identifier, sequencing primer hybridization site and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules by bridge amplification, wherein each cluster is generated from individual spatially separated template DNA molecules. (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0049] In yet another embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an index, barcode, unique molecular identifier, sequencing primer hybridization site and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules by bridge amplification, wherein each cluster is generated from individual spatially separated template DNA molecules. (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads over the entire length of the amplicon, and the bidirectional sequencing is a synthetic sequencing using reversibly terminated labeled nucleotides. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (b) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (c) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (d) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b); (v) Steps to analyze the sequence results and A method is provided that includes this.

[0050] According to the above embodiment, in one example, the glass surface is a glass slide or a flow cell.

[0051] In another embodiment, the adjacent nucleotide region in step (i) corresponds to approximately 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0052] In another embodiment, the nucleic acid sample of the object of interest comprises B and / or T cell DNA, wherein the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0053] In yet another embodiment, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0054] In yet another embodiment, the target nucleotide sequence is a V gene segment region, such as a region susceptible to hypermutation, and / or a J gene segment region encoding a portion of CDR3.

[0055] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a somatic hypermutable V region, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0056] In a further embodiment, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the forward and reverse read portions are 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0057] In yet another embodiment, the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where 20 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0058] In yet another embodiment, the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where up to 30 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0059] In another further embodiment, the linker is 5 to 30 nucleotides long, preferably 5 to 25, more preferably 5 to 20 nucleotides long. In yet another embodiment, the linker is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides long.

[0060] In another further embodiment, the analysis includes aligning the nucleic acid sequence results generated in step (iv) and determining the expression of a target nucleic acid sequence of interest.

[0061] In related embodiments, a method for diagnosing, monitoring, or otherwise screening a patient's condition, wherein the condition is characterized by the expression of one or more target nucleotide sequences. (i) A step of spatially separating a library of individual template DNA molecules derived from a nucleic acid sample on a solid support, wherein the template DNA molecules are generated such that the target nucleotide sequence is localized in the adjacent nucleotide region at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0062] In one embodiment, the state is characterized by a clonal population of cells or microorganisms.

[0063] In another embodiment, the clonal cells are a population of clonal lymphocyte cells.

[0064] In another embodiment, the state is characterized by one or more target nucleotide sequences expressed by immune cells.

[0065] In yet another embodiment, the adjacent nucleotide region in step (i) corresponds to approximately 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0066] In yet another embodiment, the state is characterized by the expression of features of one or more rearranged V, D, or J gene segment sequences.

[0067] In another embodiment, the DNA sample of interest comprises B and / or T cell DNA, and the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0068] In yet another embodiment, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0069] In yet another embodiment, the target nucleotide sequence is a V gene segment region, such as a region susceptible to hypermutation, and / or a J gene segment region encoding a portion of CDR3.

[0070] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a somatic hypermutable V region, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0071] In a further embodiment, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the forward and reverse read portions are 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0072] In yet another embodiment, the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where 20 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0073] In yet another embodiment, the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where up to 30 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0074] In another embodiment, the linker is 5 to 25 nucleotides long. In yet another embodiment, the linker is 5 to 20 nucleotides long. In a further embodiment, the length of the linker is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides, most preferably 9, 10, 11, or 12 nucleotides long.

[0075] In another embodiment, the analysis includes aligning the nucleic acid sequence results generated in step (iv) and determining the expression of the target nucleic acid sequence.

[0076] In yet another embodiment, the state characterized by the expression of one or more rearranged V, D, or J gene segment features is any other state characterized by infection, transplantation, autoimmunity, immunodeficiency, allergy, neoplasm, or T or B cell clonal proliferation.

[0077] The aforementioned method is useful in situations of diagnosis, prognosis, classification, prediction of disease risk, detection of disease recurrence, immune surveillance, or monitoring of prophylactic or therapeutic effects.

[0078] Disease conditions suitable for analysis in the context of lymphoid neoplasms include acute lymphoblastic leukemia, acute lymphoblastic leukemia, acute myeloid leukemia, acute promyelocytic leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, myeloproliferative neoplasms such as myeloma, systemic mastocytosis, lymphoma, and hairy cell leukemia.

[0079] In one particular embodiment, the method of the present invention is used to detect minimal residual lesions in the context of lymphoid neoplasms.

[0080] In another embodiment, non-neoplastic diseases characterized by clonal lymphocyte proliferation include infection, allergy, autoimmunity, graft rejection, immunotherapy, polycythemia vera, myelodysplasia and leukocytosis, such as lymphocytosis.

[0081] Another aspect of this disclosure relates to a computer implementation method for producing nucleic acid sequence results for analysis from non-overlapping sequence reads. The method comprises the steps of identifying forward sequence reads and reverse sequence reads from sequence reads of an amplicon cluster, wherein the cluster is generated from individual spatially separated template DNA molecules, each sequence read is generated by a selected bidirectional sequencing technique, the forward sequence reads and reverse sequence reads are non-overlapping, and no adjacent reads are provided across the entire length of any amplicon; and obtaining a plurality of first nucleic acid sequence results by ligating forward sequence reads with reverse sequence reads such that each forward sequence read is ligated to a reverse sequence read, and each reverse sequence read is ligated to a forward sequence read via a first nucleic acid linker sequence, wherein each ligation ligates the first nucleic acid linker sequence between the 3' end of a 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of a 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read, and Therefore, the method includes the steps of obtaining a first nucleic acid sequence result comprising, in that order, a portion of a forward sequence read, a first nucleic acid linker sequence, and a reverse complement of a portion of a reverse sequence read, wherein (1) the length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) the length of the portion from the reverse sequence read is the same for all reverse sequence reads analyzed, (3) the length of the portion from the forward sequence read is the same for all forward sequence reads analyzed, but may be the same as or different from the length of the portion from the reverse sequence read, and (4) the first nucleic acid linker sequence is the same for all first nucleic acid sequence results.

[0082] In some embodiments, the computer implementation method comprises the steps of obtaining a plurality of second nucleic acid sequence results by linking forward sequence reads with reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a second nucleic acid linker sequence, wherein each linkage is achieved by linking a second nucleic acid linker sequence between the 3' end of a portion of the 5' adjacent nucleic acid sequence at the end of the reverse sequence read and the reverse complement of a portion of the 5' adjacent nucleic acid sequence at the end of the forward sequence read, thereby obtaining a second nucleic acid sequence result comprising, in that order, a portion from the reverse sequence read, a second nucleic acid linker sequence, and the reverse complement of a portion from the forward sequence read. (1) The length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique; (2) The length of the portion from the reverse sequence read linked to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read linked to the first nucleic acid linker; (3) The length of the portion from the forward sequence read linked to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read linked to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read linked to the second nucleic acid linker; (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

[0083] Another aspect of the present disclosure is a non-temporary computer-readable storage medium having embodied program instructions, wherein the program instructions, executable by a processing element of the device, include the steps of identifying forward sequence reads and reverse sequence reads from sequence reads of amplicon clusters, wherein the clusters are generated from individual spatially separated template DNA molecules, each sequence read is generated by a selected bidirectional sequencing technique, the forward sequence reads and reverse sequence reads are non-overlapping, and no adjacent reads are provided across the entire length of any amplicon; and the steps of ligating forward sequence reads with reverse sequence reads such that each forward sequence read is ligated to a reverse sequence read, and each reverse sequence read is ligated to a forward sequence read via a first nucleic acid linker sequence, wherein each ligation connects the first nucleic acid linker sequence between the 3' end of a 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of a 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read. A method for creating nucleic acid sequence results for analysis from non-overlapping sequence reads is implemented in a device, which is achieved by obtaining a first nucleic acid sequence result comprising, in that order, a portion of a forward sequence read, a first nucleic acid linker sequence, and a reverse complement of a portion of a reverse sequence read, wherein (1) the length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) the length of the portion from the reverse sequence read is the same for all reverse sequence reads to be analyzed, (3) the length of the portion from the forward sequence read is the same for all forward sequence reads to be analyzed, but may be the same as or different from the length of the portion from the reverse sequence read, and (4) the first nucleic acid linker sequence is the same for all first nucleic acid sequence results.

[0084] In some embodiments, a non-temporary computer-readable storage medium comprises the steps of: (1) the length of the portion from the forward sequence read being determined by a selected bidirectional sequencing technique (1) The length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) The length of the portion from the reverse sequence read linked to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read linked to the first nucleic acid linker, (3) The length of the portion from the forward sequence read linked to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read linked to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read linked to the second nucleic acid linker, and (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

[0085] Another aspect of the present disclosure relates to a device for producing nucleic acid sequence results for analysis from non-overlapping sequence reads. The device includes a hardware processor configured to identify forward and reverse sequence reads from sequence reads of amplicon clusters, where the clusters are generated from individual spatially separated template DNA molecules, each sequence read is generated by a selected bidirectional sequencing technique, the forward and reverse sequence reads are non-overlapping, and no adjacent reads are provided across the entire length of any amplicon, and further configured to link forward and reverse sequence reads to a number of first nucleic acid sequence results, where each linkage is between the 3' end of a 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of a 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read. This is achieved by linking a first nucleic acid linker sequence between the forward sequence reads, the first nucleic acid linker sequence, and the reverse complement of the reverse sequence reads in that order, wherein (1) the length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) the length of the portion from the reverse sequence read is the same for all reverse sequence reads analyzed, (3) the length of the portion from the forward sequence read is the same for all forward sequence reads analyzed, but may be the same as or different from the length of the portion from the reverse sequence read, and (4) the first nucleic acid linker sequence is the same for all first nucleic acid sequence results.

[0086] In some embodiments, the hardware processor is configured to link forward sequence reads with reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a second nucleic acid linker sequence, thereby obtaining a plurality of second nucleic acid sequence results, each linkage being achieved by connecting the second nucleic acid linker sequence between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read, thereby obtaining a second nucleic acid sequence result comprising, in that order, the portion from the reverse sequence read, the second nucleic acid linker sequence, and the reverse complement of the portion from the forward sequence read, wherein (1) the length of the portion from the forward sequence read is provided by the selected bidirectional sequencing technique (1) The length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) The length of the portion from the reverse sequence read linked to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read linked to the first nucleic acid linker, (3) The length of the portion from the forward sequence read linked to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read linked to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read linked to the second nucleic acid linker, and (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

[0087] In some embodiments, the first nucleic acid linker sequence and the second nucleic acid linker sequence are at least 11 nucleotides long.

[0088] In some embodiments, the length of the forward sequence read portion is the same as the length of the reverse sequence read portion.

[0089] In some embodiments, the forward sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the forward sequence read, and the reverse sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the reverse sequence read. In some embodiments, the specified number of adjacent nucleotides includes between approximately 80 and approximately 180 nucleotides.

[0090] In some embodiments, the forward and reverse sequence reads are DNA sequence reads. In some embodiments, the amplicon clusters are amplified from B and / or T cell DNA.

[0091] In some embodiments, the amplicon cluster includes at least one rearranged V, D, or J gene segment. [Brief explanation of the drawing]

[0092] [Figure 1] This is a block diagram of a system according to the aspects of this disclosure. CPU: Central Processing Unit ("Processor"). [Figure 2] This is a flowchart of an embodiment for creating nucleic acid sequence results for analysis from non-overlapping sequence reads. [Figure 3] This is a flowchart of an embodiment for creating nucleic acid sequence results for analysis from non-overlapping sequence reads. [Modes for carrying out the invention]

[0093] The present invention is partly based on the development of a means of using non-overlapping bidirectional sequencing reads to screen one or more target nucleotide sequences. Specifically, the colocalization of bidirectional sequence read results to amplicon clusters, which are generated from a single template DNA immobilized on a solid platform and are therefore clones, makes it possible to identify that the sequencing information of these reads originates from a common template DNA. Conventional methods have relied on the use of overlapping forward and reverse read sequences to enable assembly of the entire template DNA sequence from bidirectional sequence reads, or reference sequences on which the reads are aligned to determine their orientation and position relative to each other. This also has the advantage that sequencing errors are known to occur more frequently at the 3' terminal end of sequence reads, but the overlapping complementary sequences of the paired reads make it possible to identify the presence of a single base error on a single strand (as opposed to a mutation), which can then be confidently discarded, thus facilitating relatively accurate alignment and analysis of the tapered reads. However, if bidirectional sequence reads do not overlap, their pairing and assembly by overlapping complementary 3' sequences is impossible. Furthermore, even when bidirectional sequence reads are analyzed individually, it has been found that, aside from the problem of arbitrary sequencing errors that can occur at the 3' end of the reads and result in a single read being classified as a different (e.g., mutated) sequence compared to a comparison read that does not show errors, the generation of different sequence read lengths alone can lead to inaccurate classification of these reads as different sequences, even if their actual sequences are otherwise identical, thereby distorting the sequencing results for the target DNA sample.

[0094] However, it was unexpectedly found that this unintended phenomenon is corrected if the sequence reads are modified to be sufficiently cleaved from the 3' bidirectional sequence read ends so that all sequence reads of the forward and reverse reads are the same length. Furthermore, if the forward and reverse reads are thus adjusted, and the 3' ends of the forward and reverse reads, which are then identified as colocalizing to a single amplicon cluster on a solid support, are linked using nucleic acid linkers attached to the 5' ends of sequences complementary to the reverse and forward reads, respectively, to generate linear sequence reads, and the linkers are the same for all assembled reads for a given biological sample, then accurate alignment and comparative analysis of the assembled sequence results can be achieved. By designing an initiation DNA template library such that the target nucleotide sequences are located at the 5' and 3' ends of the template and are therefore sequenced by selected bidirectional sequencing techniques, a means is provided for analyzing target nucleotide sequences that may be located quite far apart, such as VDJ gene segments rearranged into immunoglobulins or TCR genes, even if the entire template is not completely sequenced. By choosing sequencing instrument usage based on the generated read length rather than other functional characteristics of instrument use, and thus no longer being limited to designing template DNA libraries so that the template molecules are short enough to allow the generation of overlapping bidirectional sequence reads, a wide range of applications for high-throughput next-generation sequencing analysis has now become possible.

[0095] Accordingly, one aspect of the present invention is a method for screening a target nucleic acid sample to express one or more target nucleotide sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the nucleic acid sample on a solid support, wherein the template DNA molecules are generated such that the target nucleotide sequence is localized in the adjacent nucleotide region at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and This applies to methods that include [specific methods].

[0096] In one embodiment, the non-adjacent sequence reads are not analyzed against the reference sequence in order to pair the forward and reverse reads.

[0097] References to “nucleic acid,” “nucleotide,” “base,” or “nucleic acid base” should be understood as references to both deoxyribonucleic acid or nucleotides and ribonucleic acid or nucleotides or purine or pyrimidine bases or their derivatives or analogs. In this regard, it should be understood that, in particular, the term encompasses phosphate esters of ribonucleotides and / or deoxyribonucleotides, including DNA (cDNA or genomic DNA), RNA, or mRNA. Nucleic acid molecules of the present invention may be of any origin, including naturally occurring (e.g., derived from biological samples), recombinantly produced, or synthetically produced. Nucleic acids may also be non-standard nucleotides such as inosine.

[0098] References to “derivatives” should be understood to include references to fragments, parts, portions, homologs, and mimics of the nucleic acid molecule from natural, synthetic, or recombinant sources. “Functional derivatives” should be understood to be derivatives exhibiting any one or more functional activities of purine or pyrimidine bases, nucleotides, or nucleic acid molecules. Derivatives of nucleotides or nucleic acid sequences include fragments having specific regions of a nucleotide or nucleic acid molecule fused to other proteinaceous or nonproteinaceous molecules. Biotinylation of nucleotides or nucleic acid molecules is an example of a “functional derivative” as defined herein. Derivatives of nucleic acid molecules may result from one or more nucleotide substitutions, deletions, and / or additions. The term “functional derivative” should also be understood to include nucleotides or nucleic acids exhibiting any one or more functional activities of nucleotides or nucleic acid sequences, such as products obtained after natural product screening.

[0099] The “analogs” as used herein include, but are not limited to, modifications to nucleotides or nucleic acid molecules, such as modifications to their entire chemical composition or three-dimensional structure, or any other type of modification to nucleotides that do not exist naturally. This includes, for example, modifications to the way in which nucleotides or nucleic acid molecules interact with other nucleotides or nucleic acid molecules, such as at the level of skeletal formation or complementary base-pair hybridization. Without limiting the present invention to any theory or mode of action, nucleic acids consist of three parts: a phosphate backbone, a pentose sugar, ribose, or deoxyribose, and one of four bases. Analogs may have any of these modified. Typically, analog bases confer different base-pairing and base-stacking properties, among other things. Examples include universal bases that can pair with all four standard bases, and phosphate-sugar backbone analogs such as PNA that affect the properties of the chain. Nucleic acid analogs are also called senonucleic acids. Nucleic acids that do not exist in nature include peptide nucleic acids (PNA), morpholino and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). Each of these is distinguished from naturally occurring DNA or RNA by modifications to the molecular backbone.

[0100] The target nucleic acid sample and / or target nucleotide sequence may be DNA, RNA, or derivatives or analogs thereof. The nucleic acid sample may take the form of genomic DNA, cDNA generated from mRNA transcripts, DNA generated by nucleic acid amplification, synthetic DNA, or DNA generated by recombination. If the target nucleic acid sample is RNA, it is understood that it is necessary to first reverse transcribe the RNA into DNA using RT-PCR or the like. The target RNA may be RNA in any form, such as mRNA, primary RNA transcript, ribosomal RNA, transfer RNA, or microRNA. Preferably, the nucleic acid sample and the target nucleotide sequence are DNA.

[0101] According to this embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a solid support, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0102] In one embodiment, the adjacent nucleotide region in step (i) corresponds to approximately 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0103] References to “target nucleotide sequence” should be understood as references to any DNA or RNA sequence that is to be analyzed. This could be a gene, a portion of a gene, such as a gene segment or region, or an intergenetic region. For this purpose, references to “gene” should be understood as references to a DNA molecule encoding a protein product, whether it be a full-length protein or a protein fragment. With respect to chromosomal DNA, a gene contains both intron and exon regions. However, as long as the nucleic acid sample is cDNA, intron regions may not be present, as can occur when the target nucleotide sequence is vector DNA or reverse-transcribed mRNA. Nevertheless, such DNA may contain 5' or 3' untranslated regions. Therefore, references to “gene” in this specification should be understood to include, for example, genomic DNA and any form of DNA encoding a protein or protein fragment, including cDNA. The target nucleotide sequence in question may also correspond to a non-coding portion of genomic DNA that is not known to be associated with any particular gene (e.g., commonly referred to as “junk” DNA region). This can correspond to any region of genomic DNA produced by recombination, between two regions of genomic DNA or between a region of genomic DNA and a region of foreign DNA such as a virus or introduced sequence. It can also correspond to a region that may contain breakpoints such as SNPs, chromosomal translocations, insertions, deletions, or chromosomal breakpoints. The target sequence can also correspond to a region of a nucleic acid molecule produced partially or entirely by synthesis or recombination. The target sequence may also be a region of DNA previously amplified by any nucleic acid amplification method, including polymerase chain reaction (PCR) (i.e., it was produced by the amplification method).

[0104] The method of the present invention is designed to screen for the “expression” of one or more target nucleotide sequences. “Expression” means the presence of the sequence in the nucleic acid sample being tested. It should be understood that the sequence in question may or may not correspond to a nucleic acid sequence that is transcribed and / or translated.

[0105] The fact that the method of the present invention may be designed to screen "one or more" target nucleotide sequences of interest should be understood to mean that one or more distinct target sequences can be screened. Examples of distinct target sequences include SNPs, point mutations, hypermutations, DNA insertions, DNA deletions, chromosome breakpoints, specific gene segments, specific regions, parts or sections of genes, intergenetic regions, etc. In the context of a single analysis, one or more of these target sequences can be screened. These target sequences may be located at separate and distinct positions in the nucleic acid of the sample, or they may be located consecutively along the nucleic acid chain. It should be understood that they may even occur at the same position along the nucleic acid chain, for example, when a mutation is found within a gene segment and both the mutation and the gene segment itself are the target sequence of interest. In one embodiment, the nucleic acid sample of interest comprises B and / or T cell DNA, and the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0106] According to this embodiment, a method for screening a DNA sample containing B and / or T cell DNA for expressing one or more rearranged V, D, or J gene segments, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a solid support, wherein the template DNA molecules are generated such that the rearranged V, D, or J gene segments are localized to adjacent nucleotide regions at the 5' and / or 3' terminal ends of the template, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0107] It should be understood that the reference to “B and / or T cell DNA” refers to DNA derived from any lymphocyte that has rearranged at least one germline set of immunoglobulin or TCR variable region gene segments. Immunoglobulin variable regions encoding rearrangeable genomic DNA include variable regions associated with the heavy chain or κ or λ light chain, while TCR chain variable regions encoding rearrangeable genomic DNA include the α, β, γ, and δ chains. In this regard, it should be understood that a cell is within the scope of “lymphocyte” if it has rearranged the variable region encoding DNA in at least one immunoglobulin or TCR gene segment region. The cell does not also need to transcribe and translate the rearranged DNA. In this regard, “lymphocyte” includes, but is not limited to, immature T and B cells that have rearranged the TCR or immunoglobulin variable region gene segment but do not express further rearranged chains (e.g., TCR thymocytes), or have not further rearranged both chains of those TCR or immunoglobulin variable region gene segments. This definition further extends to lymphoid cells that have undergone rearrangement of at least some TCRs or immunoglobulin variable regions, but which may not otherwise exhibit all of the phenotypic or functional characteristics conventionally associated with mature T cells or B cells.

[0108] Furthermore, it should be understood that in one embodiment, the target rearrangement is a complete rearrangement, such as a complete rearrangement of at least one variable region gene region, while in another embodiment, the target rearrangement is a partial rearrangement. For example, B cells that have undergone only a DJ recombination event are cells that have undergone only a partial rearrangement. A complete rearrangement is not achieved until the DJ recombination segment is further recombined with a V segment. Therefore, the method of the present invention may be designed to screen for partial or complete variable region rearrangements of TCRs or immunoglobulin chains.

[0109] While the present invention is not limited to any theory or mode of action, V(D)J recombination in organisms with adaptive immune systems is an example of a type of site-directed recombination that helps rapidly diversify immune cells to recognize and adapt to new pathogens. Each lymphocyte cell is approximately 10 16 To generate the total antigenic diversity of individual distinct variable region structures, somatic recombination occurs in the germline variable region gene segment (either V and J, D and J, or V, D and J segments) depending on the specific gene segment being rearranged. In any given lymphocyte cell, such as a T cell or B cell, at least two distinct variable region gene segment rearrangements can occur due to the rearrangement of two or more of the two chains containing the α, β, γ, or δ heavy and light chains of the TCR or immunoglobulin molecule, specifically the TCR and / or immunoglobulin molecule. In addition to the rearrangement of the VJ, DJ, or VDJ segments of any given immunoglobulin or TCR gene, nucleotides are randomly removed and / or inserted at the junctions between segments. This leads to the generation of enormous diversity.

[0110] The loci for these gene segments are widely separated in the germline, but recombination during lymphocyte development results in the juxtaposition of the V, (D), and J genes, and the junctions between these genes are characterized by small regions of nucleotide insertions and deletions. Because this process occurs randomly, each normal lymphocyte will have its own unique V(D)J rearrangement. Lymphoid cancers such as acute lymphoblastic leukemia, chronic lymphoblastic leukemia, lymphoma, or myeloma arise as a result of neoplastic changes in a single normal cell, so all cancer cells have the V(D)J rearrangement at the junction that was originally present in the founder cell, at least initially. Subclones can arise during the expansion of the neoplastic population, and further V(D)J rearrangements can occur in them.

[0111] References to “gene segments” should be understood as references to the V, D, and J regions of immunoglobulin and T cell receptor genes. The V, D, and J gene segments are clustered into families. For example, for the κ immunoglobulin light chain, there are 52 distinct functional V gene segments and 5 J gene segments. For the immunoglobulin heavy chain, there are 55 functional V gene segments, 23 functional D gene segments, and 6 J gene segments. Across the entire immunoglobulin and T cell receptor V, D, and J gene segment families, there are numerous individual gene segments, thereby allowing for a vast diversity in terms of unique combinations of V(D)J rearrangements that may be affected. For the sake of clarity, a rearranged immunoglobulin or T cell receptor [V(D)J] variable nucleic acid region is referred to herein as a rearranged “gene,” and individual V, D, or J nucleic acid regions are referred to as “gene segments.” Thus, the technical term “gene segment” is not limited to references to segments of a gene. Rather, in the context of Ig and TCR gene rearrangements, this is itself a reference to genes in which these gene segments are clustered into a family. A “rearranged” immunoglobulin or T cell receptor variable region gene should be understood herein as a gene in which two or more of one V segment, one J segment, and one D segment (if the D segment is incorporated into the particular rearranged variable gene in question) are spliced ​​together to form a single rearranged “gene.” In fact, this rearranged “gene” is actually a stretch of genomic DNA containing one V gene segment, one J gene segment, and one D gene segment spliced ​​together. Thus, it is sometimes referred to as a “gene region” because it actually consists of two or three distinct V, D, or J genes (referred to herein as gene segments) that are spliced ​​together. Accordingly, individual “gene segments” of a rearranged immunoglobulin or T cell receptor gene are defined as individual V, D, and J genes.These genes are described in detail in the IMGT database. The term “gene” is used herein to refer to rearranged immunoglobulin or T cell receptor variable genes. The term “gene segment” is used herein to refer to V, D, and J segments. However, it should be noted that there is a significant inconsistency in the use of the terms “gene” / “gene segment” with respect to immunoglobulin and T cell receptor rearrangements. For example, IMGT refers to the individual V, D, and J “genes,” while some scientific publications refer to them as “gene segments.” Some sources refer to rearranged variable immunoglobulins or T cell receptors as “gene regions,” while others refer to them as “genes.” The nomenclature used herein is as previously defined.

[0112] Furthermore, while the present invention is not limited to any theory or mode of action, the nature of genetic recombination events is such that the junctions between recombinant genes or gene segments (as defined herein) can be characterized by random nucleotide deletions and insertions resulting in the formation of “N regions.” These N regions are also unique and, therefore, are sometimes useful targets in the context of targeted sequence analysis. Thus, it is generally understood that V(D)J rearrangements provide diversity of combinations, while the addition of N nucleotides or palindromic (P) nucleotides provides diversity of junctions.

[0113] Furthermore, in the context of V(D)J rearrangements, it should be understood that the secondary structures of the translated protein molecules themselves, while relating to the DNA sequence regions within the V(D)J rearrangements that encode the properties of these secondary structures, also contain inherent properties that are often the subject of analysis. For example, the translated variable regions of IgH (immunoglobulin heavy chain) or TCRβ or δ chains typically take the form of three loop-shaped hypervariable regions, referred to as complementarity-determining regions (CDRs) 1, 2, and 3. These CDR regions are adjacent to four framework regions (FRs) 1, 2, 3, and 4. While the present invention is not limited to any theory or mode of action, it is understood that the V gene segment encodes CDR1, CDR2, the leader sequence, FR1, FR2, and FR3. The CDR3 region is encoded by part of the V gene segment, all of the D gene segment, and part of the J gene segment. The remainder of the J gene segment generally encodes FR4.

[0114] Therefore, in one embodiment and in the context of V(D)J rearrangements, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ. In another embodiment, the target nucleotide sequence is a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0115] In yet another embodiment, the target nucleotide sequence is a V gene segment region and / or a J gene segment region encoding a portion of CDR3, such as a hypermutation-prone region.

[0116] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a V region susceptible to somatic hypermutation, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0117] In yet another embodiment, the target nucleotide sequence is a BCL1 / JH or BCL2 / JH t(14:18) translocation.

[0118] In yet another embodiment, the target nucleotide sequence is an internal tandem duplication or other mutation related to the FLT3 or TP53 gene.

[0119] Regarding the properties of target nucleotide sequences, the method of the present invention facilitates screening for the presence of specific nucleotide sequences, such as specific V, D, or J gene segment sequences, or for screening target nucleotide sequence regions to determine the diversity of sequences expressed by DNA molecules in that region. In this example, the target nucleotide sequence may be a V, D, or J gene segment family rather than a specific V, D, or J gene segment, so that the properties and diversity of gene segments within the family expressed by the DNA sample of interest can be determined.

[0120] The method of the present invention provides a significant improvement over conventional solid-phase next-generation sequencing techniques based on the use of cluster amplification of individual template sequences followed by bidirectional sequencing. While the present invention is not limited to any theory or mode of operation, in one embodiment of this type of technique, following the preparation of a library of DNA templates for analysis, these templates are immobilized on a solid support via adapter sequences. Once attached, cluster generation can be initiated. The objective is to produce hundreds of identical strands of the template DNA. Some correspond to forward strands, and others to complementary reverse strands. Clusters are then generated by bridge amplification. Polymerase moves along the DNA strand, generating its complementary strand. The original strand is washed away, leaving only the reverse strand. Another adapter sequence is present on top of the reverse strand. The DNA strand bends and attaches to an immobilized oligonucleotide complementary to this adapter sequence. Polymerase then attaches to the reverse strand, generating its complementary strand (which is identical to the original strand). Here, the double-stranded DNA is denatured so that each strand can separately attach to other unoccupied, fixed oligonucleotide sequences that are complementary to the adapters present at each end of the amplicon. This bridge amplification proceeds to simultaneously generate thousands of clusters corresponding to individual templates across a solid support (often referred to as a "flow cell"). Thus, since each cluster is generated from a single start template DNA, the amplification is cloned in the context of the individual cluster.

[0121] Following clonal amplification, the reverse strand is washed away from the flow cell, leaving only the forward strand. Synthetic sequencing is then initiated using reversibly terminated fluorescently labeled oligonucleotides. Primers attach to the forward strand, and polymerase adds the fluorescently tagged nucleotide to the DNA strand. Only one base is added per round. Reversible terminators present on all nucleotides prevent multiple additions in a single round. Each of the four bases produces a unique emission, and after each round, the instrument used records which base was added based on the emitted fluorescence. Once the forward DNA strand is read and the sequence reads are washed away, the reverse strand is generated by bridge amplification in another round. The forward strand is then washed away, and the synthetic sequencing process is repeated for the reverse strand. In this way, bidirectional sequencing is achieved.

[0122] The present invention improves this method by designing means for generating non-overlapping bidirectional sequence reads of a DNA template longer than a selected bidirectional sequence read length, and for precisely pairing and assembling them. This is achieved in part by the intrinsic design of a library of template DNA molecules derived from a nucleic acid sample. The reference to “template” DNA molecules in this regard should be understood as a reference to a DNA molecule that is immobilized on a solid support (spatially separated) and subsequently amplified to generate clusters of clonal amplicons. That is, this molecule includes both a target nucleic acid region and any further nucleic acids or non-nucleic acid regions described in more detail below herein (e.g., nucleic acid adapter sequences, sequencing primer hybridization regions, index regions, unique molecular identifiers, etc.). In this regard, it should be understood that the template DNA molecule undergoing cluster amplification and sequencing is a single-stranded molecule, but upon immobilization on a solid support, the DNA template may be in single-stranded form or may form part of a molecular complex such as a double-stranded DNA molecule, or a complex with non-nucleic acid components. For example, it may be desirable to enrich the template population before fixation, which can be achieved by coupling beads or a chemical compound (e.g., biotin) to the specific template DNA molecule of interest to enable their isolation and subsequent enrichment before fixation. However, as long as double-stranded or other molecular complexes are fixed, those skilled in the art will understand that the complexes need to be made single-stranded before cluster amplification so that only the fixed template DNA is amplified. In this regard, it is assumed that, as long as the template DNA is coupled with a non-nucleic acid molecule that does not interfere with amplification, such as biotin, this non-nucleic acid molecule does not necessarily need to be cleaved. Thus, references to “template” DNA molecules are intended to refer to the DNA molecules that actually undergo amplification. A “library” of template DNA means a collection of template DNA molecules (in single-stranded, double-stranded, or some other complex forms) that are initially applied to and fixed on a solid support. It should be understood that the template DNA may consist of naturally occurring or non-naturally occurring nucleotides, as described herein above.

[0123] The template DNA molecule applied to the solid support "derives" from the nucleic acid sample of interest. "Derived" means that the template DNA is either directly isolated from the sample, as is done when the sample's DNA is simply fragmented before application to the solid support, or it takes the form of an amplification product generated from the DNA sample of interest. In this regard, the template DNA library can be prepared using any suitable method. The library can be generated by fragmentation of the nucleic acid sample of interest, such as using an endonuclease, particularly a restriction enzyme, exonuclease, exo-endonuclease, or any other means of site-directed DNA cleavage. Depending on the nature and location of the target nucleotide sequence, this method may be sufficient to generate the library. Alternatively, to facilitate enrichment of the target nucleotide sequence, it may be chosen to amplify the sample of interest using primers that specifically target and amplify the nucleotide sequence of interest, such as primers induced to amplify specific immunoglobulin or TCR gene segment rearrangements, primers to amplify gene regions that may have generated SNPs, or primers to amplify across specific indels, breakpoints, or other chromosomal translocations or mutations. The template DNA molecule can be of any suitable length, for example, 250-1000, 250-900, 300-700, or 300-600 nucleotides. Since the template DNA may also incorporate adapter regions to facilitate solid-phase amplification and sequencing, it will be understood by those skilled in the art that the portion of the template DNA molecule corresponding to the target nucleic acid region is generally shorter than the length of the template DNA. In this regard, these further non-target regions may consist of 15-75 nucleotides, preferably 20-40, more preferably 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides at each end of the template DNA molecule.

[0124] Regardless of whether the template DNA molecule takes the form of fragmented DNA or is amplified from all or part of the DNA sample of interest, the template DNA may also be further modified to introduce additional nucleic acid or non-nucleic acid components, which are necessary or desirable to enhance the effectiveness of the high-throughput amplification and sequencing platform technology used in the context of the present invention. Such additional sequences may include, for example, restriction enzyme sites or certain nucleic acid tags to enable identification of the amplified product of a given nucleic acid template sequence. Other desirable sequences may include “control” DNA sequences that direct protein / DNA interactions, such as foldback DNA sequences (which form hairpin loops or other secondary structures when single-stranded), promoter DNA sequences recognized by nucleic acid polymerases, or operator DNA sequences recognized by DNA-binding proteins. In another example, to enable the immobilization of the template DNA to a solid support, the means for attaching the template DNA to the solid support requires coupling to the template DNA. In this regard, as used herein, “means for attaching template DNA to a solid support” refers to any chemical or non-chemical attachment method that includes chemically modifiable functional groups. "Adhesion" refers to the immobilization of template DNA on a solid support by covalent or non-covalent bonding, including by irreversible passive adsorption or by intermolecular affinity (e.g., immobilization on an avidin-coated surface by biotinylated molecules), or hybridization (e.g., between short complementary nucleic acid fragments). The adhesion must be strong enough that it cannot be removed by washing with water or an aqueous buffer under DNA denaturation conditions. As used herein, "chemically modifiable functional group" refers to groups such as phosphate groups, carboxyl or aldehyde moieties, thiols, or amino groups. For this purpose, reference to "solid support" should be understood as referring to any solid surface to which nucleic acids can be covalently bonded, such as latex beads, dextran beads, polystyrene, polypropylene surfaces, polyacrylamide gels, gold surfaces, glass surfaces, and silicon wafers. The selection of a suitable solid support and the means for attaching template DNA are well known to those skilled in the art.In one embodiment, the solid support is a solid matrix capable of confirming a two-dimensional position. In another embodiment, the solid support is a glass surface (such as a glass slide or flow cell), and the means for fixing the mold to the glass surface is a nucleic acid anchor.

[0125] According to this embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized in the adjacent nucleotide region at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0126] Preferably, the glass surface is a glass slide or a flow cell.

[0127] In another embodiment, the nucleic acid sample of the object of interest comprises B and / or T cell DNA, wherein the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0128] In yet another embodiment, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0129] In yet another embodiment, the target nucleotide sequence is a V gene segment region, such as a region susceptible to hypermutation, and / or a J gene segment region encoding a portion of CDR3.

[0130] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a somatic hypermutable V region, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0131] A typical example of a nucleic acid immobilization system is a short linear nucleic acid sequence (referred to herein as a “nucleic acid adapter”) attached to the 5' and / or 3' ends of a template DNA molecule. The anchor takes the form of a complementary nucleic acid sequence covalently bound to a solid support. When the template DNA is applied to the solid support, any nucleic acid adapter sequence complementary to the covalently bound nucleic acid anchor results in hybridization of the two sequences, thereby immobilizing the template DNA to the solid support. In this regard, the 5' nucleic acid adapter sequence attached to the template DNA can be designed to express the same sequence as that of the corresponding anchor sequence, so that only sequences complementary to the 5' adapter hybridize with the anchor, while the 3' nucleic acid adapter sequence is complementary to its corresponding anchor. Thus, when the entire length of the template DNA sequence undergoes cluster amplification, the hybridization of the adapter sequence on the 3' end of the DNA template with the corresponding anchor, and the amplification of the amplicon generated from the DNA template are always facilitated, thereby enabling bridge amplification and cluster formation to always occur. As will be understood by those skilled in the art, this is the principle that operates using, for example, Illumina MiSeq, HiSeq, NovaSeq, and NextSeq instruments.

[0132] Therefore, the reference to “spatially separating” individual template DNA molecules on a solid support should be understood as referring to immobilizing these molecules on the solid support in order to enable cluster amplification of the template. For this purpose, if the concentration of molecules applied to the solid support is such that the distribution and immobilization of these molecules across the solid support leaves anchor molecules that are not sufficiently occupied proximal to each immobilized template DNA molecule, then the template molecules are “spatially” separated so that localized clonal cluster amplification can occur without any amplicon of any one clonal cluster fusing into substantially another cluster, thereby enabling the pairing of bidirectional sequencing data from a single template with high accuracy based on colocalization data. That is, the amplicon of a single cluster is maintained within a separate region on the solid support, and the cluster density is optimized so that the data can be spatially allocated. In this regard, determining the optimal cluster density for the instrument usage selected for use is well within the realm of the art. As will be understood by those skilled in the art, each cluster may contain both a forward strand and a complementary reverse strand for each start template DNA molecule.

[0133] In addition to adapter molecules that can be incorporated into the template DNA molecule to facilitate immobilization of the template DNA to a solid support, the template DNA molecule can also be modified to incorporate further properties useful in clinical or research settings, such as indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, and index sequencing primer hybridization sites. For example, in addition to localizing the target nucleotide sequence of interest to the 5' and 3' ends of the template as described herein, the template DNA molecule can be designed to be modified to incorporate an additional nucleic acid sequence region, which (a) is adjacent to the target nucleotide sequence region and (b) is located at either or both of the 5' and 3' ends of the template DNA molecule together with the adapter. Thus, this additional nucleic acid sequence region expresses one or more of the adapter sequence and inverse multiplexing index (commonly also referred to as a barcode), allowing for the simultaneous analysis of multiple different nucleic acid samples, and the unique molecular identifier enables the identification of individual amplicons, sequencing primer hybridization sites, and index sequencing primer hybridization sites. The combination of properties selected to be incorporated into the 5' end of the template DNA does not need to be the same as that incorporated into the 3' end. For example, a reverse multiplexing index can be incorporated into only one end of the template DNA strand. Designing such additional properties into the template DNA to facilitate optimal experimental design is well within the realm of the art. Means for incorporating such additional nucleic acid components are well known and include blunt-end ligation of nucleic acid fragments containing these properties to the 5' and / or 3' ends of the template DNA molecule. Alternatively, if the template library is prepared by amplifying the DNA of the sample of interest, for example by PCR, amplification primers can be designed to include these additional properties at their 5' ends. Thus, primers designed to amplify the target nucleotide sequence of interest can be designed to incorporate these additional nucleic acid sequences simultaneously, thereby generating the library in a single amplification step.Alternatively, one could choose to use a two-step amplification procedure to prepare the library, where the first round of amplification uses primers targeting the generation of a template DNA amplicon expressing the target nucleotide sequence, followed by the use of primers (e.g., consensus primers) targeting all amplicons generated from the first round, which achieve the incorporation of exogenous DNA such as the previously described index.

[0134] In one embodiment, the template DNA molecule further expresses one or more nucleic acid sequences corresponding to an index, a barcode, a unique molecular identifier, a sequencing primer hybridization site, and an index sequencing primer hybridization site at the 5' and / or 3' positions of the terminal.

[0135] According to this embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an index, barcode, unique molecular identifier, sequencing primer hybridization site and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0136] Preferably, the glass surface is a glass slide or a flow cell.

[0137] In another embodiment, the nucleic acid sample of the object of interest comprises B and / or T cell DNA, wherein the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0138] In yet another embodiment, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0139] In yet another embodiment, the target nucleotide sequence is a V gene segment region, such as a region susceptible to hypermutation, and / or a J gene segment region encoding a portion of CDR3.

[0140] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a somatic hypermutable V region, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0141] As detailed above in this specification, the present invention has facilitated the routine use of high-throughput bidirectional sequencing, even when the template DNA is longer than what bidirectional sequencing chemistry can read. However, this development is partly based on the design of a template DNA molecule such that the target nucleotide sequence is located within a contiguous nucleotide region at the 5' and / or 3' terminal ends of the template. More specifically, the target sequence should be located within a stretch of nucleotides at the 5' and / or 3' terminals, corresponding to about 80% of the maximum read length that can be obtained by the bidirectional sequencing technique selected for use. In this regard, reference to “bidirectional sequencing” (also commonly referred to as paired-end sequencing) should be understood as a reference to obtaining sequence information related to the template DNA molecule from both its 5' and 3' ends. In practice, this is achieved by sequencing the template DNA amplified by cluster formation on a solid support. Sequence of a strand complementary to the target strand (also known as the “template strand” or “template amplicon”) from its 3' end generates a “reverse read.” The sequence of this read is complementary to the target strand. The sequencing of the complementary strand from the 3' end of this complementary strand to the target strand generates a "forward read." The sequence of this read corresponds to the template strand. Thus, the two reads are the reverse complement of approximately 100 (depending on the sequencing chemistry used) of the most 3' nucleotides of the template strand, and its complementary strand.

[0142] If the template strand is shorter than the combined length of the forward and reverse bidirectional sequence reads, the forward and reverse reads overlap and exhibit complementarity in the overlapping region. Based on these reads, the full-length sequences of the template strand and its complement can be estimated. However, this is not possible if the template strand is longer than the combined length of the bidirectional forward and reverse reads, because the central region of the template strand is not sequenced by either of the reads. As discussed herein, the method of the present invention provides an improved means for performing high-throughput bidirectional sequencing so that its application can be extended to any template DNA molecule (and thus its template strand amplicon) regardless of its length.

[0143] The sample of the present invention includes both the strand expressing the target nucleotide sequence and the reverse strand of the target nucleotide sequence of interest. DNA includes two complementary strands of DNA that hybridize together to form a molecule. The target nucleotide sequence of interest is defined in the context of the present invention as the “forward strand” (also known as the “template strand” or “target strand”), while the complementary strand is referred to as the “reverse strand”. Those skilled in the art will also understand that the two strands of a DNA double helix are often referred to as the “sense” strand, the “coding” strand, the “plus (+)” strand, the “top” strand, or the “upper” strand. These latter three terms are most commonly used when the DNA region of interest does not produce a protein expression product. The corresponding complementary strand is often referred to as the “antisense” strand, the “non-coding” strand, the “minus (-)” strand, the “lower” strand, or the “bottom” strand. This should be understood to mean the strand that, in the context of a chromosome locus, is complementary to the top / + / upper strand and, in its native state, hybridizes with the top strand to form a characteristic double helix structure. As those skilled in the art will understand, this nomenclature has become increasingly inaccurate as it has been discovered that there are many gene regions that do not code for proteins (and are therefore not precisely described as being found on the sense or coding strand), and furthermore, these genes may be found on either the + / upper strand or the - / lower strand, depending on how those skilled in the art define these strands. Now, even protein-coding genes are known to be found on what was traditionally considered the - / bottom / antisense strand. Therefore, referring to this terminology alone, without mentioning a specific chromosomal location, or referring to a specific + / - strand nomenclature used in annotated human genome databases, can be inaccurate. In this regard, in the context of the present invention, a reference to the "forward strand" refers to the DNA strand containing the target nucleotide sequence, whichever of the two strands it may be, while a "reverse strand" refers to the complementary strand. Thus, the target strand may correspond to either the + / - (top / bottom, upper / lower) strand in the original DNA biological sample, depending on where the gene is located in the chromosomal double helix.The terms "forward lead" and "reverse lead" should be distinguished from the definitions of "forward lead" and "reverse lead" used herein.

[0144] As detailed above in this specification, a DNA template derived from a nucleic acid sample is designed such that one or more target nucleotide sequences of interest are localized at the 5' and / or 3' terminal ends of the template. In this regard, the reference to the “terminal ends” of the DNA template refers to the region of nucleic acid sequence that extends adjacently from the most terminal 5' nucleotide in the 3' direction along the template strand and from the most terminal 3' nucleotide in the 5' direction along the template strand. More specifically, for a number of consecutive nucleotides corresponding to approximately 80% of the maximum forward or reverse read length obtained by the bidirectional sequencing technique selected for use, the target nucleotide sequence is located within the adjacent stretches of nucleotides extending from the terminal 5' and / or 3' nucleotides in the 3' and 5' directions, respectively. The reference to “forward and reverse read lengths” should be understood as a reference to the read length of a single read, not the combined length of both reads. For example, using an Illumina NovaSeq 6000 instrument allows for a maximum of 300 cycles, which corresponds to a bidirectional sequencing read length of 150 nucleotides for forward reads and 150 nucleotides for reverse reads, 80% of which is 105 nucleotides per read. Therefore, the reference to "maximum read length" refers to the maximum read length (e.g., 150 for the NovaSeq 6000) for either forward or reverse reads that can be achieved under optimal conditions using the selected instrument or chemistry, and this information is widely and routinely available to those skilled in the art. In this regard, it should be understood that not all reads generated in a single sequencing run necessarily produce the maximum possible read length. Furthermore, the lengths of millions of forward reads and millions of reverse reads generated in a high-throughput bidirectional sequencing process are not equal. Typically, variation between sequence read lengths is observed; that is, forward read lengths, like reverse read lengths, can vary by up to 5%.As detailed above in this specification, when aligning a series of unpaired forward or unpaired reverse reads, all originating from the same template molecule and therefore expressing the same sequence, it has been unexpectedly found that currently available alignment software and algorithms sometimes classify these sequences as different sequences, solely due to the generation of reads with slightly different lengths. Such analytical errors can negatively impact the specificity and / or sensitivity of results in clinical applications such as screening for minimal residual disease, clonal evolution, or the presence or appearance of a small number of clones.

[0145] As detailed above in this specification, the target nucleotide sequence is located within a 5' and / or 3' adjacent stretch of the nucleotide terminus, the length of which corresponds to approximately 80% of the maximum forward and reverse bidirectional read length. In one embodiment, the percentage of the maximum read length is 70%–85%, in another embodiment, 75%–85%, and in yet another embodiment, 75%–80%. In yet another embodiment, the percentage of the maximum read length is 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83%. The reference to a target nucleotide sequence “localized” to a defined adjacent nucleotide region should be understood to mean that the target sequence is located within that region, but does not necessarily have to span the entire length of that region. That is, a stretch of the sequence may exist within a defined region that does not express the target sequence. This is more likely to occur when the target nucleotide sequence is small. As long as two target nucleotide sequences can exist, they can be located distal to the 5' and 3' ends of the template, for example, when a portion of a particular V gene segment is located at the 5' end of the template and part or all of the CDR3 region is located at the 3' end of the template. It should be understood that if only one target nucleotide sequence of interest exists, neither the 5' nor the 3' terminal of the template will express the target nucleotide sequence. It should also be understood that more than one target nucleotide sequence can exist located within a single defined 5' or 3' region. For example, both V gene segment-specific sequences and the occurrence of somatic hypermutations within a particular V gene segment sequence can be screened. In this case, there are two target nucleotide sequences to be analyzed, both located within defined adjacent nucleotide regions at the ends of the template DNA.

[0146] According to this embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0147] As detailed above in this specification, the target nucleotide sequence must be located within a defined 5' or 3' terminal facile nucleotide region of the template DNA, corresponding to approximately 80% of the maximum theoretical read length of the selected bidirectional sequencing technique. It should be understood that the reference to this region of the template is to a defined region, regardless of whether it is functionally available for expressing the target nucleotide sequence. Therefore, the facile nucleotide region in which the target sequence can actually be located may be less than the equivalent of the maximum read length. For example, insofar as the template DNA may be designed to incorporate further nucleic acid properties such as adapters, indices, barcodes, primer hybridization sites, etc. (referred to herein as “adapter regions”), all or part of this stretch of terminal nucleotides will be unavailable to the target sequence, depending on where the sequencing primer hybridization site is located within the adapter region, because this further adapter region inevitably forms part of the bidirectional sequence read. Specifically, the section of the adapter region sequence located at 3' relative to the sequencing primer hybridization site, rather than the section of the adapter sequence located at 5' relative to the primer hybridization site, forms part of the sequence read. Those skilled in the art will understand that such non-target nucleic acid characteristics may include, for example, adjacent nucleotide lengths of 10-30 nucleotides located at the terminal 5' and 3' positions. As long as the bidirectional sequence read is 2 × 100-150 nucleotides, the 10-30 nucleotide region that is not available for the target sequence corresponds to a larger proportion of read length that cannot be used to maximize the target sequence read length than if the selected sequence read length were 2 × 200-300 nucleotides. However, as those skilled in the art will understand, bidirectional read length is not only a consideration when selecting a particular instrument or chemistry for use. For example, the Illumina MiSeq instrument provides a bidirectional read length of 2 × 300 nucleotides, but offers a read depth more than an order of magnitude less than the NovaSeq instrument, which provides only a read length of 2 × 150.For example, when attempting to apply this method to MRD analysis, sequence depth becomes a crucial factor. Therefore, the ability to select the instrumentation and chemistry of any high-throughput bidirectional sequencing system for use has significantly expanded the applicability of this class of technology, regardless of whether overlapping bidirectional reads may be generated.

[0148] In one embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, and the 20 nucleotide terminal ends of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, using sequencing chemistry that produces a maximum forward read length of 150 nucleotides and a maximum reverse read length of 150 nucleotides; (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, The process involves the aforementioned portion being 120 nucleotides each of the forward and reverse read lengths, the linker sequence being the same for all nucleic acid sequence results in (a), and the linker sequence being the same for all nucleic acid sequence results in (b), (v) Steps to analyze the sequence results and A method is provided that includes this.

[0149] In another embodiment, the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, but up to 30 nucleotide terminals in the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0150] It will be understood that, as defined above in this specification, generating a DNA template localized to the 5' and / or 3' ends of a template with one or more target nucleotide sequences is well within the scope of the art. Since the overall length of the DNA template is now of little importance, those skilled in the art only need to identify the target sequences and then determine how to incorporate them into the DNA template at their precise locations. If only one target sequence of interest exists, it may be possible to generate a template by, for example, simply cutting the DNA of a biological sample near the target sequence using a suitable restriction enzyme and then ligating any required adapter region into the fragment, or by amplifying the fragment using a consensus primer that includes the adapter region sequence at the template end of the primer as a non-hybridize tail region, thereby incorporating the adapter region into the amplified product to generate a template library. Alternatively, amplification of a DNA sample can be carried out using primers in which either the forward or reverse primer is adjacent to the target sequence, thereby enabling its amplification, while the other primer is bound to any suitable region of the DNA to allow PCR to proceed. These primers can incorporate an adapter region sequence at the terminal end of the primer as a non-hybridized region, thereby enabling the incorporation of the adapter region into the amplification product in a single step, or allowing a second round of amplification to be performed using a consensus primer targeting the amplification product of the first round to introduce the adapter region. When attempting to analyze more than one target sequence, those skilled in the art can design amplification primers adjacent to the 5' end of the upstream target nucleotide sequence and the 3' end of the downstream target nucleotide sequence. The length of the intervening sequence is irrelevant, as long as the target nucleotide sequence selected for analysis can localize to the terminal 5' and 3' regions as defined herein above. Designing primers adjacent to and amplifying one or more target nucleotide sequences is a routine and straightforward procedure.Those skilled in the art will understand that, depending on the relative positions of the target sequences and the orientation of the primers in question, the length of the target nucleotide sequence that can be localized to the defined 5' and / or 3' ends of the DNA template can be maximized, thereby enabling sequencing, by positioning the amplification primers to be adjacent to the target sequence as close as possible to where the target nucleotide sequence begins or ends, depending on the relative positions of the target sequences and the orientation of the primers in question. In this regard, primers can be designed so that they hybridize within the target sequence itself, thereby forming a portion of the amplified target sequence nucleotide sequence, in which case the length of the primer sequence forms a portion of the 5' and / or 3' DNA template region to be sequenced. If the primer hybridizes outside the target region, one can choose to design a primer sequence having a cleavage site at its 3' end, which can cleave the primer sequence from the amplicon in a site-specific manner. In any of these examples, the adapter region can be introduced in either a single or two-step procedure as described above. In yet another example, it may be desirable to generate template DNA using non-PCR-based methods, such as splicing a region of DNA expressing the target nucleotide sequence within a vector and amplifying the vector via host cell replication. The DNA templates thus generated require excision from the vector before they can be facilitated to adhere to a solid support.

[0151] As detailed above in this specification, the methods of the present invention relate to means of applying high-throughput bidirectional sequencing to screen nucleic acid samples even when it is not possible to obtain overlapping bidirectional reads for template DNA whose read lengths are longer than those of sequencing chemistry. This is achieved in part by spatially separating individual template DNA molecules on a solid support so that amplification can be carried out by any suitable method for generating clusters of amplicons. In this regard, the reference to “amplicon” refers to an amplified copy of the template DNA and / or its complementary sequence. Accordingly, the reference to “cluster” is intended to refer to a colony of amplicons that is generated and fixed proximal to the template DNA so that colonies of the clonal target sequence and the clonal complementary sequence are generated around a single template DNA. Methods for carrying out cluster DNA are well known to those skilled in the art and can be carried out as a standard procedure. An exemplary method for achieving such cluster amplification is bridge amplification. In this method, when template DNA containing adapter sequences at both the 5' and 3' ends is immobilized on a solid support at an appropriate density, nucleic acid clusters can be generated by performing an appropriate number of amplification cycles on the immobilized template DNA such that each colony contains multiple copies of the original immobilized template DNA and its complementary sequence. One amplification cycle consists of hybridization, extension, and denaturation steps, which are generally carried out using reagents and conditions well known in the art for PCR. A typical amplification reaction involves subjecting the solid support and attached template DNA to conditions that induce primer hybridization and extension in the presence of nucleic acid polymerase, along with a supply of nucleoside triphosphate molecules or any other nucleotide precursor, e.g., modified nucleoside triphosphate molecules. The primers are extended by the addition of nucleotides complementary to the template DNA.Examples of nucleic acid polymerases that can be used in the present invention include DNA polymerases (Klenow fragments, T4 DNA polymerase), heat-stable DNA polymerases derived from various heat-stable bacteria (Taq, VENT, Pfu, Tfl DNA polymerase, etc.), and their genetically modified derivatives (TaqGold, VENTexo, Pfu exo). A combination of RNA polymerase and reverse transcriptase can also be used to generate amplification of DNA colonies. Preferably, the nucleoside triphosphate molecules used are deoxyribonucleotide triphosphates, such as dATP, dTTP, dCTP, and dGTP. The nucleoside triphosphate molecules may or may not be naturally occurring.

[0152] Following the hybridization and extension steps, two immobilized nucleic acids are present, the first being a template strand and the second a complementary nucleic acid strand. Both of these nucleic acid molecules can then initiate further rounds of amplification by the formation of a bridge and hybridization of the unimmobilized ends of the amplicon with its complementary immobilized anchor. Such further rounds of amplification produce nucleic acid clusters containing multiple immobilized clonal copies of the template strand and its complementary sequence. The initial immobilization of the template DNA means that the template DNA can hybridize with adapter anchors located at a distance within the length of the template DNA, forming only a bridge. Thus, the cluster boundaries are confined to the relatively localized region where the initial template DNA is immobilized. Clearly, when copies of the template strand and its complement are synthesized again by performing further rounds of amplification, the cluster boundaries formed are still confined to the relatively localized region where the initial template DNA is immobilized, but the resulting clusters can be further extended. The amplification of the subject can be carried out qualitatively or quantitatively.

[0153] In one embodiment, the amplification is bridge amplification.

[0154] According to this embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an index, barcode, unique molecular identifier, sequencing primer hybridization site and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules by bridge amplification, wherein each cluster is generated from individual spatially separated template DNA molecules. (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0155] Preferably, the glass surface is a glass slide or a flow cell.

[0156] In another embodiment, the nucleic acid sample of the object of interest comprises B and / or T cell DNA, wherein the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0157] In yet another embodiment, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0158] In yet another embodiment, the target nucleotide sequence is a V gene segment region, such as a region susceptible to hypermutation, and / or a J gene segment region encoding a portion of CDR3.

[0159] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a somatic hypermutable V region, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0160] In another embodiment, the adjacent nucleotide region in step (i) corresponds to approximately 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0161] In a further embodiment, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the forward and reverse read portions are 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0162] In yet another embodiment, the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where 20 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0163] In yet another embodiment, the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where up to 30 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0164] Following cluster formation, bidirectional sequencing is performed on one or more amplicons of one or more clusters. However, in most situations, parallel bidirectional sequencing of all clusters and all amplicons within these clusters is expected. Any high-throughput method for bidirectional sequencing of nucleic acids can be used in the method of the present invention. In one example, sequencing by synthesis using reversibly terminated labeled nucleotides is applied. As detailed above herein, the present invention is not limited to any theory or mode of operation, but in one embodiment of bidirectional sequencing using reversibly terminated labeled nucleotides, following clonal amplification, the reverse strand is washed away from the solid support, leaving only the forward (template) strand. Sequencing is then initiated. A primer attaches to the forward strand, and a polymerase adds a fluorescently tagged nucleotide to the DNA strand. Only one base is added per round. Reversible terminators present in all nucleotides prevent multiple additions in a single round. Each of the four bases produces a unique emission, and after each round, the instrument used records which base was added based on the emitted fluorescence. Once the forward DNA strand is read and the sequence reads are washed away, the reverse strand is generated by bridge amplification in another round. The forward strand is then washed away, and the synthesis sequencing process is repeated for the reverse strand. In this way, bidirectional sequencing is achieved.

[0165] In one embodiment, the method involves sequencing by synthesis using reversibly terminated labeled nucleotides.

[0166] According to this embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an index, barcode, unique molecular identifier, sequencing primer hybridization site and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules by bridge amplification, wherein each cluster is generated from individual spatially separated template DNA molecules. (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads over the entire length of the amplicon, and the bidirectional sequencing is a synthetic sequencing using reversibly terminated labeled nucleotides. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0167] Preferably, the glass surface is a glass slide or a flow cell.

[0168] In another embodiment, the nucleic acid sample of the object of interest comprises B and / or T cell DNA, wherein the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0169] In yet another embodiment, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0170] In yet another embodiment, the target nucleotide sequence is a V gene segment region, such as a region susceptible to hypermutation, and / or a J gene segment region encoding a portion of CDR3.

[0171] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a somatic hypermutable V region, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0172] In another embodiment, the adjacent nucleotide region in step (i) corresponds to approximately 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0173] In a further embodiment, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the forward and reverse read portions are 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0174] In yet another embodiment, the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where 20 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0175] In yet another embodiment, the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where up to 30 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0176] As detailed above in this specification, the method of the present invention is based on the development of means for analyzing non-overlapping bidirectional sequence reads that provide accurate and reproducible results. This development is partly based on the unexpected finding that current analysis software classifies reads differently based on any difference in read length, despite the fact that one or more clusters of forward or reverse reads originate from the same template sequence and therefore produce the same sequence read results, and despite the fact that most of the sequences of the reads are identical among these reads. Further complicating the analysis of results is the additional complexity that sequencing errors become more frequent with respect to the 3' ends of the sequencing reads. When bidirectional sequence reads contain overlapping and complementary 3' ends, the issue of individual read lengths becomes practically meaningless because the reads are tapeped together before alignment and further analysis. Furthermore, the problem of sequencing errors is mitigated because information from the complementary strand to the strand exhibiting the sequencing anomaly helps determine whether any such sequence difference is real or not. This is impossible when analyzing reads for which overlapping complementary strand reads are not available. For this reason, current teachings relating to high-throughput bidirectional sequencing are that the template DNA must always be designed so that its length matches the read length of the instrument used. Furthermore, as those skilled in the art know, while the instrument use of bidirectional sequencing provides a theoretical maximum sequence read length, the actual reads obtained do not necessarily accurately reflect that read length, and the actual read lengths obtained can vary by as much as 5% between reads.

[0177] According to this method, forward and reverse reads are identified for one or more of the sequenced clusters. “Identified” means that sequence information is determined for forward and reverse reads co-localized to a single cluster. In this regard, when multiple high-throughput screening is performed, those skilled in the art may choose to initially identify forward and reverse read sequence information for some, but not all, clusters. For example, when multiple reactions are performed to analyze multiple patient samples, the results may be demultiplexed so that information for one patient, rather than others, is analyzed first. This demultiplexing step is performed by using a patient-specific index or barcode. Alternatively, if more than one target sequence is screened for the use of separate primer pairs (which themselves may be designed to be identifiable by an index or other suitable means well known to those skilled in the art), it may be chosen to initially analyze only one of these target nucleotide sequences. In one embodiment, all clusters for which bidirectional sequencing information is generated are analyzed. In this regard, the analysis of sequence reads and the generation and analysis of sequence results may be carried out in any convenient manner, as will be described in more detail below. For example, the sequence data can be examined manually, or an appropriate algorithm can be used to efficiently automate one or more of the analytical steps described in step (iv). Alternatively, a combination of methods and algorithms can be used to carry out the steps described in step (iv). It should be understood that this analysis, including the generation of sequence results, is most conveniently performed in silico.

[0178] As detailed above in this specification, the forward and reverse reads for individual template DNA molecules subjected to cluster amplification and bidirectional sequencing according to this method are identifiable based on the colocalization of these reads to the location of a single cluster on a solid support. However, these reads overlap at their 3' ends and do not exhibit complementary sequence regions. Once these “paired” reads are identified, nucleic acid sequencing results can be generated. “Sequencing results” means sequences assembled from forward and reverse reads and then aligned to each of the cluster sequencing results for evaluating the clonality or diversity of the target DNA sample; alignment of the sequencing results to a reference sequence for further classification of the sequences (e.g., to determine the specific identity of V, D, or J gene segments when the template DNA is amplified using gene families or consensus primers); identification of the occurrence and nature of hypermutations, indels, DNA breaks, SNPs, etc.; evaluation of clonal evolution; or determination of the emergence of new clones. In another example, in the context of MRD monitoring, it may be desired to identify patient-specific sequences. This is because this may indicate a recurrence of the disease. It should be understood that the sequencing result may include the positions of the 5' and 3' adapter regions, depending on where the sequencing primer hybridization site is positioned. In this regard, those skilled in the art may choose to cleave this additional sequence so that the sequencing result contains only the sequence corresponding to the DNA sample of interest, along with the intervening linker region. However, those skilled in the art may also decide that this is unnecessary and that the sequencing result retains this additional sequence at its 5' and 3' ends, as it is identifiable.

[0179] The nucleic acid sequence results are typically generated by assembling in silico portions of the 5' adjacent nucleic acid sequences of the forward and reverse reads, which may or may not include any terminal nucleotides corresponding to the adapter region. The reference to “portion” should be understood as a reference to a portion, though not necessarily the entire, of the forward and reverse read sequence lengths, relating to shorter reads, and the entire sequence may be used. The portion to be used is determined by those skilled in the art, but it is at least 80% of the maximum reads obtained by the selected bidirectional sequencing technique, and the selected portion is the same for all forward and reverse reads analyzed for a given DNA sample of interest. The reference to “maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique” should be understood to have the same meaning as previously detailed. By selecting portions within these parameters, it has been found that this provides sufficient target nucleotide sequence data to achieve sequence accuracy with respect to specificity with respect to the target sequence information of interest and sufficient removal of 3' sequence data that is likely to contain sequence errors, thereby enabling both highly sensitive and specific screening results for the DNA sample of interest. With regard to the determination of the portion used for screening a DNA sample, if considered in light of the teachings provided herein, determining this is well within the realm of those skilled in the art. As long as multiple assays are performed using samples from multiple patients, multiple different tissues, and / or target different target sequences, for example, those skilled in the art can determine different portion lengths between categories of results. However, in the context of a single DNA sample source, the portion is the same for all forward sequence reads and the same for all reverse sequence reads. In this regard, the length of the portion selected for use in the forward read does not need to be the same as the length of the portion selected for the reverse read.By ensuring that the lengths of the nucleic acids in the forward and reverse portions are the same as the lengths between all forward read portions and all reverse read portions, we prevent the unexpected occurrence of misclassification of clone sequences as different sequences simply because one sequence is longer than another.

[0180] The forward and reverse read portions are assembled to generate a sequence read result by linking the 3' end of the forward read with sequence information derived from the reverse read via a nucleic acid linker. In this regard, those skilled in the art will understand that the sequences of the forward and reverse reads correspond to the sequences of the 5' end of the template / forward strand and the 5' end of the complementary / reverse strand, respectively. Thus, if these reads are extended along the entire length of the sequence to be hybridized, the two reads are complementary. Therefore, in the context of the present invention, which concerns taping the 5' and 3' ends of the template DNA, and the 5' and 3' ends of the strand complementary to the template strand, it is necessary to easily and quickly achieve in silico the determination of sequences complementary to each of the forward and reverse read sequences, and to tape the forward read sequence with the complement of the reverse read sequence. Similarly, the complement of the forward read sequence is tapered with the reverse read sequence. This then generates a template sequence result, although only for the 5' and 3' end sequences, and the corresponding sequence result for the strand complementary to the template strand.

[0181] References to “nucleic acid linker” should be understood as references to nucleic acid sequences, preferably linear sequences, attached to the 3' ends of the forward and reverse read portions, and to the 5' ends of the sequences complementary to the forward and reverse read portions, such that the 3' end of the forward read sequence is linked to a sequence complementary to the reverse read sequence, and the 3' end of the reverse read sequence is linked to a complement to the forward read sequence, forming a single linear adjacent nucleic acid sequence. The nucleotides of the linker may be any naturally occurring or non-naturally occurring nucleotides, but as long as this aspect of the invention is carried out in silico, the actual chemical structure of the nucleotides in the assembled sequence result is less important than the in silico functional information related to these nucleotides, such as showing accurate complementary base pairings, where relevant. References to “naturally occurring and non-naturally occurring” nucleotides should have the same meaning as provided above herein. In one embodiment, the nucleic acid linker is N xHere, N represents a natural or non-natural nucleotide, and x represents the number of adjacent nucleotides in the linker. Regarding the nature of the linker sequence itself, it can be a random sequence, but if a randomly generated sequence is used, it must be the same for all sequencing results. This is because differences in the linker sequences used for assembled forward and reverse read pairs, which are otherwise derived from clones and therefore identical, will result in these sequences being classified as different due to the variability of the linker sequence. Furthermore, this means that comparisons between sequencing results of a single DNA sample in the context of immune receptor diversity, etc., are meaningless. Preferably, if the sequence of interest is linked in silico, the N nucleotide is simply designated as N, thereby different from and identifiable from the naturally occurring nucleotides A, T, G, and C. The length of the linker sequence can be any appropriate length determined by those skilled in the art. In this regard, it has been found that the number of nucleotides in the linker should not be too small. This is because a "linker" consisting of only one or two N nucleotides is interpreted as a random nucleotide insertion, and thereby is not interpreted as a linker, leading to misalignment of the sequence. In one embodiment, the linker is 5 to 30 nucleotides long, preferably 5 to 25, more preferably 5 to 20 nucleotides long. In another embodiment, the length of the linker is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides.

[0182] According to this embodiment, a method for screening a target DNA sample to express one or more target DNA sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the DNA sample on a glass surface, wherein the template DNA molecules are generated such that the target DNA sequence is localized to an adjacent nucleotide region at the 5' and / or 3' terminal ends of the template, and the terminal ends of the adjacent nucleotide region express one or more nucleic acid sequences corresponding to an index, barcode, unique molecular identifier, sequencing primer hybridization site and index sequencing primer hybridization site, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules by bridge amplification, wherein each cluster is generated from individual spatially separated template DNA molecules. (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads over the entire length of the amplicon, and the bidirectional sequencing is a synthetic sequencing using reversibly terminated labeled nucleotides. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is 5 to 30 nucleotides long and is the same for all nucleic acid sequence results in (a); and the linker sequence is the same for all nucleic acid sequence results in (b); (v) Steps to analyze the sequence results and A method is provided that includes this.

[0183] Preferably, the glass surface is a glass slide or a flow cell.

[0184] In another embodiment, the nucleic acid sample of the object of interest comprises B and / or T cell DNA, wherein the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0185] In yet another embodiment, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0186] In yet another embodiment, the target nucleotide sequence is a V gene segment region, such as a region susceptible to hypermutation, and / or a J gene segment region encoding a portion of CDR3.

[0187] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a somatic hypermutable V region, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0188] In another embodiment, the adjacent nucleotide region in step (i) corresponds to approximately 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0189] In a further embodiment, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the forward and reverse read portions are 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0190] In yet another embodiment, the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where 20 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0191] In yet another embodiment, the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where up to 30 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0192] In another embodiment, the linker is 5 to 25 nucleotides long. In yet another embodiment, the linker is 5 to 20 nucleotides long. In a further embodiment, the length of the linker is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides, most preferably 9, 10, 11, or 12 nucleotides long.

[0193] Once the sequence results are assembled, the assembled sequence can be analyzed. The type of analysis performed is determined by those skilled in the art and depends on the nature of the information sought. For example, these results can be mined to identify the presence or absence of specific mutations or other sequence characteristics such as specific V(D)J immunoglobulins or TCR rearrangements. This can be useful for diagnostic or MRD purposes, or for determining the relative effectiveness of treatment. Some diseases are identified by the presence of specific mutations (e.g., Flt3 or NPM1), hypermutations, indels, gene breakpoints (e.g., BCR-ABL), etc. Alternatively, instead of screening for the presence of previously known target sequences, one may request an investigation of sequence diversity in the gene region of interest, and this sequence information can then be used to track the progression and / or development of the disease. For example, a leukocyte neoplasm resulting from the neoplastic transformation of a single leukocyte is itself useful for identification and tracking based on the identification of the unique V, D, and / or J rearrangements of the neoplastic cell. This can be particularly useful for assessing minimal residual disease. Due to the vast diversity of the immune cell repertoire, virtually all leukocytes exhibit unique immunoglobulin or TCR rearrangements. Specific cells can be tracked by identifying one or more specific gene segments rearranged in a neoplasm population. In relation to the application of the present invention, DNA from biological samples can also be screened to assess the diversity of specific rearrangements, such as IgH VJ rearrangements. If all rearranged IgH VJ sequences from blood or bone marrow samples are screened, the alignment of the sequence results provides a qualitative or quantitative readout of the diversity of IgH VJ gene segment rearrangements. This can be extremely useful in the context of investigating the immune system to determine the circumstances or progression of any other events that may be beneficial in assessing whether (desirable or undesirable) immunotherapy, infection, transplantation, autoimmunity, allergy, immunodeficiency, or T or B cell clonal proliferation are occurring as indicators of immune activity.If a clone exhibits an expansion of the clonal population (e.g., due to an acute immune response to a pathogen or autoantigen), an increase in the number of sequence reads corresponding to a single specific rearrangement will be evident in the rearrangement at the IgH VJ locus, compared to a heterogeneous background array otherwise. Identifying the presence of this clone makes it possible to identify a specific gene segment rearrangement and track that clone. This can be particularly important in the context of autoimmunity. If multiple clones are proliferating, this may indicate a broad immune response, such as a response to multiple antigens in the context of infection, transplantation, or allergy.

[0194] With respect to the sequence analysis performed herein, multiple identical sequence results for a single cluster are aligned, and identical sequences are fused into a single sequence result. Non-identical sequences within a cluster are discarded on the basis that they may contain sequencing errors if they differ from the sequences of other amplicons from the same cluster. Complementary sequences may be paired to produce a DNA double-stranded result. Single-stranded or double-stranded sequences between clusters are then aligned. In one example, a threshold is set for the tolerance of 2 or 3 nucleotide differences between sequences of different clusters; below this threshold, those sequences may be classified as originating from the clonal population present in the start DNA sample of interest. Relative or actual proportions (depending on whether amplification was performed quantitatively) are then evaluated to determine, for example, whether there is evidence of clonal proliferation or whether specific sequences (such as those relevant to MRD assessment) are present.

[0195] According to this embodiment, the analysis includes the step of aligning the nucleic acid sequence results generated in step (iv) and determining the expression of the target nucleic acid sequence of interest.

[0196] Therefore, this method can be used for diagnosis, prognosis, classification, prediction of disease risk, detection of disease recurrence, immune surveillance, or monitoring of preventive or therapeutic effects in contexts characterized by the expression of one or more target nucleotide sequences, or in any disease or non-disease state. Furthermore, this method is applicable to any other contexts where the analysis of sequences in a particular target DNA and RNA region or the screening for the presence of a particular target DNA and RNA sequence is required, such as in the context of research and development. For example, the present invention provides solutions to current and emerging needs that scientists and the biotechnology industry are seeking to address in the fields of genomics, pharmacogenomics, drug discovery, food characterization, and genotyping.

[0197] Using lymphoid neoplasms as a non-limiting example, the present invention provides a method for determining whether a mammal (e.g., human) has a neoplasm, or whether a biological sample taken from a mammal contains neoplastic cells or DNA derived from neoplastic cells, for estimating the risk or likelihood of a mammal developing a neoplasm, monitoring the effectiveness of anti-cancer treatments, or selecting appropriate treatments in a mammal with cancer. Such a method is based on the determination that lymphoid neoplasms are characterized by clonal proliferation of cells expressing a specific V(D)J rearrangement.

[0198] The method of the present invention can be used to assess individuals known to have or suspected to have neoplasms, or as a routine clinical trial in individuals not necessarily suspected to have neoplasms. Furthermore, the method can be used to evaluate the effectiveness of a treatment process. For example, the effectiveness of an anti-cancer treatment can be evaluated by monitoring DNA methylation over time in mammals with lymphoid cancer. For instance, a decrease or absence of a clonal population characterized by a specific target nucleotide sequence in a biological sample taken from a mammal after treatment indicates an effective treatment.

[0199] Accordingly, the method of the present invention is useful as a one-time test or as continuous monitoring of an individual, whether in the context of lymphoid neoplasms or in the context of any other application described herein. In these situations, screening for target sequences is a useful indicator of the individual's condition, for example, the condition of their immune system.

[0200] Therefore, in another embodiment, a method for diagnosing, monitoring, or otherwise screening a condition in a patient, wherein the condition is characterized by the expression of one or more target nucleotide sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from a nucleic acid sample on a solid support, wherein the template DNA molecules are generated such that the target nucleotide sequence is localized in the adjacent nucleotide region at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion thereof is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion thereof of the adjacent sequences of the reverse read is the same for all reverse reads to be analyzed; (3) The preceding portion of the adjacent sequences of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0201] The term "nucleic acid sample" should be understood as referring to any sample of DNA derived from any organism, including but not limited to plants, animals, or microorganisms, such as cellular material, blood, mucus, feces, urine, tissue biopsy specimens, or liquids introduced into and subsequently removed from the body of an animal (e.g., saline solution or solution recovered from enema lavage extracted from the lungs after lung lavage), microorganisms (e.g., bacteria, viruses, parasites), tissue cultures, or any recombinant, synthetic, or artificial source such as recombinant DNA processes. Biological samples tested according to the method of the present invention may be tested directly or may require some form of processing before testing. For example, biopsy specimens may require homogenization before testing. Furthermore, unless the biological sample is in liquid form, the addition of reagents such as buffers may be required to mobilize the sample.

[0202] As long as the target DNA is present in the sample, the sample may be tested directly, or all or part of the nucleic acid material present in the sample may be isolated before testing. Pre-treatment of the target nucleic acid molecule before testing, such as inactivation of live viruses or electrophoresis on a gel, is within the scope of this invention. It should also be understood that the sample may be newly collected, stored before testing (e.g., by freezing), or otherwise processed before testing (e.g., undergoing culture). The sample may also be subjected to in vitro culture or manipulation (e.g., immortalization or recombination) to generate a cell line or cell culture.

[0203] The selection of the most suitable sample type for testing according to the methods disclosed herein depends on the nature of the situation, such as the nature of the condition being monitored. For example, in a preferred embodiment, a neoplasm is the subject of analysis. If the neoplasm is lymphocytic leukemia, blood samples, lymph samples, or bone marrow aspirates may be suitable test samples. If the neoplasm is lymphoma, lymph node biopsies or blood or bone marrow samples may be suitable tissue sources for testing. It is also necessary to consider whether to monitor the original source of neoplasm cells, or whether to monitor the presence of metastasis of the neoplasm from its origin or the spread of other forms. In this regard, it may be desirable to collect and test a number of different samples from any one mammal. In another example, in the case of infection, either or both of cell proliferation and microbial clonal proliferation, such as viral proliferation, can be tested. Selecting an appropriate sample for any given detection scenario is within the scope of the art of the art.

[0204] As used herein, the term “mammal” includes humans, primates, domestic animals (e.g., horses, cattle, sheep, pigs, donkeys), laboratory test animals (e.g., mice, rats, rabbits, guinea pigs), companion animals (e.g., dogs, cats), and captured wild animals (e.g., kangaroos, deer, foxes). Preferably, the mammal is a human or a laboratory test animal. More preferably, the mammal is a human.

[0205] The nucleic acid sample being tested may be cell-free DNA, such as that found in circulation in the context of certain disease conditions, or it may originate from cells.

[0206] References to “cells or more cells” should be understood as references to all forms of cells from any species, and their variants or offshoots. In one embodiment, the cells are lymphocytes, but the methods of the present invention can be carried out on any type of cell that can undergo partial or complete immunoglobulin or TCR rearrangement. Without limiting the present invention to any one theory or mode of action, cells can constitute an organism (in the case of a single-celled organism), or they can be subunits of a multicellular organism in which individual cells can be more or less specialized (differentiated) for a particular function. All living organisms are composed of one or more cells. The cells of interest may form part of a biological sample being tested in a syngeneic, homogeneous, or heterogeneous context. The syngeneic context means that a population of cloned cells and the biological sample in which the cloned population resides share the same MHC genotype. This is most likely to occur, for example, when screening for the presence of neoplasms in an individual. The “homogeneous” context is when the clonal population of interest actually expresses a different MHC than that of the individual from which the biological sample was taken. This can occur, for example, when screening the proliferation of donor cell populations transplanted in the context of conditions such as graft-versus-host disease (e.g., immunocompetent bone marrow transplants). The "xenogeneic" context refers to cases where the target cloned cells belong to a completely different species from the subject from which the biological sample originates. This can occur, for example, when a potential neoplasm donor population originates from xenotransplantation.

[0207] The term "variant" of the target cells includes, but is not limited to, cells that exhibit some, but not all, of the morphological or phenotypic characteristics or functional activity of the variant cell. The term "mutant" includes, but is not limited to, naturally modified or unnaturally modified cells, such as genetically modified cells.

[0208] In one embodiment, the state is characterized by a clonal population of cells or microorganisms.

[0209] "Clone" means that a target population of cells or microorganisms originates from a common cellular origin. For example, a population of neoplastic cells originates from a single cell that underwent transformation at a particular stage of differentiation. In this regard, neoplastic cells that undergo further genomic rearrangement or mutation to produce a genetically distinct population of neoplastic cells are also a "clone" population of cells, although they are distinct clonal populations of cells. In another example, T or B lymphocytes that proliferate in response to acute or chronic infection or immune stimulation are also a "clone" population of cells within the definition provided herein. In yet another example, a clonal population of cells is a clonal microbial population or viral clone, such as drug-resistant clones that arise within a larger microbial population. Preferably, the target clonal population of cells is a neoplastic population of cells or a clonal immune cell population.

[0210] In one embodiment, the clonal cells are a population of clonal lymphocyte cells.

[0211] It should be understood that the term "lymphocyte" refers to any cell that has rearranged at least one germline set of immunoglobulin or TCR variable region gene segments. Immunoglobulin variable regions encoding rearrangeable genomic DNA include variable regions associated with the heavy chain or the κ or λ light chain, while TCR chain variable regions encoding rearrangeable genomic DNA include the α, β, γ, and δ chains. In this regard, it should be understood that a cell falls within the definition of a "lymphocyte" if it has rearranged the variable region encoding DNA in at least one immunoglobulin or TCR gene segment region. The cell does not also need to transcribe and translate the rearranged DNA. In this regard, "lymphocyte" includes, but is not limited to, immature T and B cells that have rearranged a TCR or immunoglobulin variable region gene segment but have not yet expressed the rearranged chain (e.g., TCR-thymocytes), or have not yet rearranged both chains of those TCR or immunoglobulin variable region gene segments. This definition further extends to lymphoid cells that have undergone rearrangement of at least some TCR or immunoglobulin variable region, but which may not otherwise exhibit all of the phenotypic or functional features conventionally associated with mature T cells or B cells. Therefore, the method of the present invention can be used to monitor cellular neoplasms, including lymphocytes, activated lymphocytes, or non-lymphoid / lymphoid cells at any stage of developmental differentiation, but not limited to, if rearrangement of at least some of a variable region gene region has occurred. It can also be used to monitor clonal proliferation that occurs in response to specific antigens.

[0212] In another embodiment, the state is characterized by one or more target nucleotide sequences expressed by immune cells. In yet another embodiment, the state is characterized by the expression of one or more rearranged V, D, or J gene segment features.

[0213] According to this embodiment, a method for diagnosing, monitoring, or otherwise screening a patient's condition, wherein the condition is characterized by the expression of features of one or more rearranged V, D, or J gene segment sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from a DNA sample containing B and / or T cell DNA on a solid support, wherein the template DNA molecules are generated such that the rearranged V, D, or J gene segments are localized to adjacent nucleotide regions at the 5' and / or 3' terminal ends of the template, (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying forward and reverse sequence reads for each of the one or more clusters sequenced according to step (iii), and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) A portion of the 5' adjacent nucleic acid sequence at the end of a reverse read, where the 3' end of one of the terminal ends of the nucleic acid linker sequence is ligated to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read, with the other terminal end of the linker sequence being ligated to a portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a), and the linker sequence is the same for all nucleic acid sequence results of (b), and (v) Steps to analyze the sequence results and A method is provided that includes this.

[0214] In another embodiment, the DNA sample of interest comprises B and / or T cell DNA, and the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

[0215] In yet another embodiment, the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In yet another embodiment, the rearrangement is a kappa deletion element rearrangement.

[0216] In yet another embodiment, the target nucleotide sequence is a V gene segment region, such as a region susceptible to hypermutation, and / or a J gene segment region encoding a portion of CDR3.

[0217] In yet another embodiment, the target nucleotide sequence is a V leader sequence, a somatic hypermutable V region, or a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0218] In another embodiment, the adjacent nucleotide region of step (i) corresponds to about 80% of the maximum forward and reverse read lengths provided by the bidirectional sequencing technology selected for use in step (iii).

[0219] In a further embodiment, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82% or 83% of the maximum forward and reverse read lengths provided by the bidirectional sequencing technology selected for use in step (iii), and the forward and reverse read portions are 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82% or 83% or more of the maximum forward and reverse read lengths provided by the bidirectional sequencing technology selected for use in step (iii).

[0220] In yet another embodiment, the target DNA sequence is located in 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, but 20 nucleotide terminal ends of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site or index sequencing primer hybridization site.

[0221] In still yet another embodiment, the target DNA sequence is located in 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, but up to 30 nucleotide terminal ends of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site or index sequencing primer hybridization site.

[0222] In another embodiment, the linker is 5 to 25 nucleotides long. In yet another embodiment, the linker is 5 to 20 nucleotides long. In a further embodiment, the length of the linker is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides, most preferably 9, 10, 11, or 12 nucleotides long.

[0223] According to this embodiment, the analysis includes the step of aligning the nucleic acid sequence results generated in step (iv) and determining the expression of the target nucleic acid sequence of interest.

[0224] In yet another embodiment, the state characterized by the expression of one or more rearranged V, D, or J gene segment features is any other state characterized by infection, transplantation, autoimmunity, immunodeficiency, neoplasm, or T or B cell clonal proliferation.

[0225] The aforementioned method is useful in situations of diagnosis, prognosis, classification, prediction of disease risk, detection of disease recurrence, immune surveillance, or monitoring of prophylactic or therapeutic effects.

[0226] In this aspect of the present invention, the reference to “monitoring” should be understood as a reference to testing the subject for the presence or level of a target clonal population of cells after the initial diagnosis of the presence of the population. “Monitoring” includes references to conducting both a single, one-time test or a series of tests over several days, weeks, months, or years. Tests may be conducted for any number of reasons, including, but are not limited to, to help arrive at a decision regarding appropriate treatment or to test a new form of treatment, including predicting the likelihood of relapse in a mammal in remission, screening for minimal residual disease, monitoring the effectiveness of a treatment protocol, confirming the status of a patient in remission, or monitoring the progression of the condition before or after the application of a treatment regimen. Accordingly, the methods of the present invention are useful as both clinical and research means.

[0227] References to “neoplastic cells” should be understood as references to cells exhibiting abnormal “growth.” The term “growth” should be understood in its broadest sense, including references to proliferation. In this regard, an example of abnormal cell growth is uncontrolled proliferation of cells. Uncontrolled proliferation of lymphocytes may result in a population of cells that take the form of either a solid tumor or a single-cell suspension (e.g., as observed in the blood of a leukemia patient). Neoplastic cells can be benign or malignant. In a preferred embodiment, neoplastic cells are malignant. In this regard, references to “neoplastic conditions” refer to the presence of neoplastic cells in the mammal of interest. The term "neoplastic lymphocyte state" includes references to disease states characterized by the presence of an abnormally large number of neoplastic cells, such as those occurring in leukemia, lymphoma, and myeloma. However, this phrase should also be understood to include references to events in which the number of neoplastic cells found in a mammal falls below the threshold usually considered to define the transition of the mammal from an apparent disease state to remission, or vice versa (the number of cells present during remission is often referred to as "minimal residual disease"). Furthermore, even if the number of neoplastic cells present in a mammal falls below a threshold detectable by screening methods used before the advent of this invention, the mammal is nevertheless considered to be in a "neoplastic state."

[0228] Disease conditions suitable for analysis in the context of this embodiment include acute lymphoblastic leukemia, acute lymphoblastic leukemia, acute myeloid leukemia, acute promyelocytic leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, myeloproliferative neoplasms, and any lymphoid neoplasm such as myeloma, systemic mastocytosis, lymphoma, and hairy cell leukemia.

[0229] In one particular embodiment, the method of the present invention is used to detect minimal residual lesions in the context of lymphoid neoplasms.

[0230] In another embodiment, non-neoplastic diseases characterized by clonal lymphocyte proliferation include infection, allergy, autoimmunity, graft rejection, immunotherapy, polycythemia vera, myelodysplasia and leukocytosis, such as lymphocytosis.

[0231] According to all of the embodiments described above, in one embodiment, the glass surface is a glass slide or a flow cell.

[0232] In another embodiment, the terminal end of the adjacent nucleotide region expresses one or more nucleic acid sequences corresponding to an index, a barcode, a unique molecular identifier, a sequencing primer hybridization site, and an index sequencing primer hybridization site.

[0233] In yet another embodiment, the amplification is bridge amplification.

[0234] In a further embodiment, the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the forward and reverse read portions are 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

[0235] In yet another embodiment, the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where 20 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0236] In yet another embodiment, the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, where up to 30 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

[0237] Computer implementation method, computer-readable storage medium and device Some aspects of this disclosure relate to computer implementations, computer-readable storage media and devices, that implement a method for producing nucleic acid sequence results for analysis from non-overlapping sequence reads for screening a nucleic acid sample of interest to express one or more target nucleotide sequences.

[0238] The computer implementation methods, computer-readable storage media, and devices described herein offer advantages over prior art methods by enabling the analysis of non-overlapping sequence reads without the use of a reference sequence. The method includes the steps of: identifying forward and reverse sequence reads from a colocalized non-overlapping read sequence; trimming the identified forward and reverse sequence reads (i.e., obtaining a predetermined length from the 5' portion of the forward sequence read and a predetermined length from the 5' portion of the reverse sequence read); and then tapeing them together with a nucleic acid linker containing a predetermined number of N (where N refers to any nucleotide, e.g., A, G, T, or C) in between (always maintaining one set of sequence reads (forward or reverse) and obtaining the reverse complement of the other set). In some embodiments, the computer implementation methods, computer-readable storage media, and devices described herein process millions to billions of sequence reads. In some embodiments, the computer implementation methods, computer-readable storage media, and devices described herein process at least 1 million, 5 million, 10 million, 20 million, 30 million, 40 million, 50 million, 100 million, 250 million, 500 million, 1 billion, 5 billion, 10 billion or more sequence reads.

[0239] As used herein, the term “memory” includes program memory and working memory. Program memory may contain one or more programs or software modules. Working memory stores data or information used by the CPU when performing the functionalities described herein.

[0240] The term “processor” can include a single-core processor, a multi-core processor, multiple processors located in a single device, or multiple processors distributed via a network of devices, the Internet, or the Cloud, either wired or wirelessly. Therefore, as used herein, a function, characteristic, or instruction executed or configured to be executed by a “processor” can include the execution of a function, characteristic, or instruction by a single-core processor, the collective or coordinated execution of a function, characteristic, or instruction by multiple cores of a multi-core processor, or the collective or coordinated execution of a function, characteristic, or instruction by multiple processors, and each processor or core does not need to execute all functions, characteristics, or instructions individually. A processor may be a CPU (Central Processing Unit). A processor can also include other types of processors, such as a GPU (Graphics Processing Unit). In other aspects of this disclosure, instead of, or in addition to, CPU execution instructions programmed into program memory, a processor may be an ASIC (Application-Specific Integrated Circuit), an analog circuit, or other functional logic such as an FPGA (Field-Programmable Gate Array), PAL (Phase Alternating Line), or PLA (Programmable Logic Array).

[0241] The CPU is configured to execute programs (also referred to herein as modules or instructions) stored in program memory in order to perform the functionalities described herein. Memory may be, but is not limited to, RAM (random access memory), ROM (read-only memory), and persistent storage. Memory is any part of hardware that can temporarily and / or permanently store information such as, for example, but not limited to, data, programs, instructions, program code, and / or other appropriate information.

[0242] Various aspects of the present disclosure may be embodied or stored as a program, software, or computer instructions in a group of media that are computer or machine usable or readable, or that cause steps of a method to be executed on a computer or machine. A machine-readable program storage device, e.g., a computer-readable medium, that tangibly embodies a program of machine-executable instructions for performing the various functionalities and methods described herein is also provided.

[0243] In some embodiments, the present disclosure includes a system including a CPU, a display, a network interface, a user interface, a memory, a program memory, and a working memory (FIG. 1), the system being programmed to execute a program, software, or computer instructions directed to the method or processor of the present disclosure. Exemplary and non-limiting embodiments are shown in FIGS. 2 and 3.

[0244] Computer-implemented method Aspects of the present disclosure are directed to a computer-implemented method for creating nucleic acid sequence results for analysis from non-redundant sequence reads from a cluster of amplicons.

[0245] In some embodiments, the computer-implemented method includes identifying forward and reverse sequence reads from sequence reads of a cluster of amplicons. In some embodiments, the forward and reverse sequence reads are DNA sequence reads.

[0246] In some embodiments, the cluster of amplicons is generated from individual spatially separated template DNA molecules, and each sequence read is generated by a selected bidirectional sequencing technique. In some embodiments, the bidirectional sequencing technique is selected from the techniques listed in Table 1. In some embodiments, the forward and reverse sequence reads do not overlap and do not provide adjacent reads over the full length of any amplicon.

[0247] In some embodiments, the amplicon cluster is amplified from B and / or T cell DNA. In some embodiments, the amplicon cluster includes at least one rearranged V, D, or J gene segment. In some embodiments, the amplicon cluster includes a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In specific embodiments, the VJ rearrangement is a kappa deletion element rearrangement. In some embodiments, the amplicon cluster includes a V gene segment region such as a hypermutagenic region, and / or a J gene segment region encoding a portion of CDR3. In some embodiments, the amplicon cluster includes a V leader sequence, a somatic hypermutagenic V region, and a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0248] In some embodiments, the computer implementation method includes the step of linking forward sequence reads with reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a first nucleic acid linker sequence, thereby obtaining a plurality of first nucleic acid sequence results.

[0249] In some embodiments, each linkage is achieved by connecting a first nucleic acid linker sequence between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read, thereby obtaining a first nucleic acid sequence result comprising, in that order, the portion of the forward sequence read, the first nucleic acid linker sequence, and the reverse complement of the portion of the reverse sequence read.

[0250] In some embodiments, the identification step is achieved by one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the forward and reverse sequence reads, and the one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the forward sequence reads are different from the one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the reverse sequence reads.

[0251] In some embodiments, the computer implementation method comprises the steps of obtaining a plurality of second nucleic acid sequence results by linking forward sequence reads with reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a second nucleic acid linker sequence, wherein each linkage is achieved by linking a second nucleic acid linker sequence between the 3' end of a portion of the 5' adjacent nucleic acid sequence at the end of the reverse sequence read and the reverse complement of the portion of the 5' adjacent nucleic acid sequence at the end of the forward sequence read, thereby obtaining a second nucleic acid sequence result comprising, in that order, a portion from the reverse sequence read, a second nucleic acid linker sequence, and the reverse complement of the portion from the forward sequence read, wherein (1) the length of the portion from the forward sequence read is determined by a selected bidirectional sequencing technique (1) The length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) The length of the portion from the reverse sequence read linked to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read linked to the first nucleic acid linker, (3) The length of the portion from the forward sequence read linked to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read linked to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read linked to the second nucleic acid linker, and (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

[0252] In some embodiments, the length of the portion from the forward sequence read is approximately 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is approximately 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum read length obtained by the selected bidirectional sequencing technique. In some embodiments, the length of the portion from the reverse sequence read is the same for all reverse sequence reads being analyzed. In some embodiments, the length of the portion from the forward sequence read is the same for all forward sequence reads being analyzed, but may be the same as or different from the length of the portion from the reverse sequence read. In some embodiments, the length of the portion from the forward sequence read is the same as the length of the portion from the reverse sequence read.

[0253] In some embodiments, the forward sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the forward sequence read, and the reverse sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the reverse sequence read. In some embodiments, the specified number of adjacent nucleotides includes between approximately 80 and approximately 180 nucleotides. As used in this disclosure, the term “approximately” refers to ±10% of a given value. In some embodiments, the specified number of adjacent nucleotides includes approximately 80, approximately 90, approximately 100, approximately 110, approximately 120, approximately 130, approximately 140, approximately 150, approximately 160, approximately 170, or approximately 180 nucleotides.

[0254] In some embodiments, the first nucleic acid linker sequence is the same for all first nucleic acid sequence results. In some embodiments, the first nucleic acid linker sequence has a nucleotide length between 5 and 30, a nucleotide length between 5 and 25, or a nucleotide length between 5 and 20. In some embodiments, the length of the first nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides.

[0255] In some embodiments, the first nucleic acid linker sequence and the second nucleic acid linker sequence are at least 11 nucleotides long. In some embodiments, the first nucleic acid linker sequence and the second nucleic acid linker sequence are between 5 and 30 nucleotides long, between 5 and 25 nucleotides long, or between 5 and 20 nucleotides long. In some embodiments, the length of the first nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides long. In some embodiments, the length of the second nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides long.

[0256] Computer-readable storage media Aspects of this disclosure relate to a non-temporary computer-readable storage medium having embodied program instructions, which are executable by a processing element of the device causing the device to implement a method for creating nucleic acid sequence results for analysis from non-overlapping sequence reads from amplicon clusters.

[0257] In some embodiments, a non-temporary computer-readable storage medium includes instructions for identifying forward and reverse sequence reads from sequence reads of an amplicon cluster. In some embodiments, the forward and reverse sequence reads are DNA sequence reads.

[0258] In some embodiments, amplicon clusters are generated from individual, spatially separated template DNA molecules, and each sequence read is generated by a selected bidirectional sequencing technique. In some embodiments, the bidirectional sequencing technique is selected from the techniques listed in Table 1. In some embodiments, the forward and reverse sequence reads do not overlap, and no adjacent reads are provided across the entire length of any amplicon.

[0259] In some embodiments, the amplicon cluster is amplified from B and / or T cell DNA. In some embodiments, the amplicon cluster includes at least one rearranged V, D, or J gene segment. In some embodiments, the amplicon cluster includes a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In specific embodiments, the VJ rearrangement is a kappa deletion element rearrangement. In some embodiments, the amplicon cluster includes a V gene segment region such as a hypermutagenic region, and / or a J gene segment region encoding a portion of CDR3. In some embodiments, the amplicon cluster includes a V leader sequence, a somatic hypermutagenic V region, and a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0260] In some embodiments, a non-temporary computer-readable storage medium includes instructions for concatenating forward sequence reads with reverse sequence reads such that each forward sequence read is concatenated with a reverse sequence read, and each reverse sequence read is concatenated with a forward sequence read via a first nucleic acid linker sequence, thereby obtaining a plurality of first nucleic acid sequence results.

[0261] In some embodiments, each linkage is achieved by connecting a first nucleic acid linker sequence between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read, thereby obtaining a first nucleic acid sequence result comprising, in that order, the portion of the forward sequence read, the first nucleic acid linker sequence, and the reverse complement of the portion of the reverse sequence read.

[0262] In some embodiments, the non-temporary computer-readable storage medium includes further instructions for linking forward sequence reads with reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a second nucleic acid linker sequence, wherein each linkage is achieved by linking the second nucleic acid linker sequence between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read, thereby obtaining a second nucleic acid sequence result comprising, in that order, the portion from the reverse sequence read, the second nucleic acid linker sequence, and the reverse complement of the portion from the forward sequence read, and (1) the length of the portion from the forward sequence read is also determined by a selected bidirectional sequencing technique. (1) The length of the portion from the reverse sequence read is 75% or more of the maximum read length that can be obtained, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique; (2) The length of the portion from the reverse sequence read linked to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read linked to the first nucleic acid linker; (3) The length of the portion from the forward sequence read linked to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read linked to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read linked to the second nucleic acid linker; (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

[0263] In some embodiments, the identification step is achieved by one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the forward and reverse sequence reads, and the one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the forward sequence reads are different from the one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the reverse sequence reads.

[0264] In some embodiments, the identification step is achieved by one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the forward and reverse sequence reads, and the one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the forward sequence reads are different from the one or more indices, barcodes, unique molecular identifiers, sequencing primer hybridization sites, or index sequencing primer hybridization sites found in the reverse sequence reads.

[0265] In some embodiments, the length of the portion from the forward sequence read is approximately 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is approximately 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum read length obtained by the selected bidirectional sequencing technique. In some embodiments, the length of the portion from the reverse sequence read is the same for all reverse sequence reads being analyzed. In some embodiments, the length of the portion from the forward sequence read is the same for all forward sequence reads being analyzed, but may be the same as or different from the length of the portion from the reverse sequence read. In some embodiments, the length of the portion from the forward sequence read is the same as the length of the portion from the reverse sequence read.

[0266] In some embodiments, the forward sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the forward sequence read, and the reverse sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the reverse sequence read. In some embodiments, the specified number of adjacent nucleotides includes between approximately 80 and approximately 180 nucleotides. As used in this disclosure, the term “approximately” refers to ±10% of a given value. In some embodiments, the specified number of adjacent nucleotides includes approximately 80, approximately 90, approximately 100, approximately 110, approximately 120, approximately 130, approximately 140, approximately 150, approximately 160, approximately 170, or approximately 180 nucleotides.

[0267] In some embodiments, the first nucleic acid linker sequence is the same for all first nucleic acid sequence results. In some embodiments, the first nucleic acid linker sequence has a nucleotide length between 5 and 30, a nucleotide length between 5 and 25, or a nucleotide length between 5 and 20. In some embodiments, the length of the first nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides.

[0268] In some embodiments, the first nucleic acid linker sequence and the second nucleic acid linker sequence are at least 11 nucleotides long. In some embodiments, the first nucleic acid linker sequence and the second nucleic acid linker sequence are between 5 and 30 nucleotides long, between 5 and 25 nucleotides long, or between 5 and 20 nucleotides long. In some embodiments, the length of the first nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides long. In some embodiments, the length of the second nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides long.

[0269] device Another aspect of this disclosure relates to a device for generating nucleic acid sequence results for analysis from non-overlapping sequence reads. The device includes a hardware processor configured to identify forward and reverse sequence reads from sequence reads of amplicon clusters.

[0270] In some embodiments, the hardware processor is configured to identify forward and reverse sequence reads from the sequence reads of the amplicon cluster. In some embodiments, the forward and reverse sequence reads are DNA sequence reads.

[0271] In some embodiments, the hardware processor is configured to obtain a plurality of first nucleic acid sequence results by linking forward sequence reads with reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a first nucleic acid linker sequence.

[0272] In some embodiments, each linkage is achieved by connecting a first nucleic acid linker sequence between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read, thereby obtaining a first nucleic acid sequence result comprising, in that order, the portion of the forward sequence read, the first nucleic acid linker sequence, and the reverse complement of the portion of the reverse sequence read.

[0273] In some embodiments, amplicon clusters are generated from individual, spatially separated template DNA molecules, and each sequence read is generated by a selected bidirectional sequencing technique. In some embodiments, the bidirectional sequencing technique is selected from the techniques listed in Table 1. In some embodiments, the forward and reverse sequence reads do not overlap, and no adjacent reads are provided across the entire length of any amplicon.

[0274] In some embodiments, the amplicon cluster is amplified from B and / or T cell DNA. In some embodiments, the amplicon cluster includes at least one rearranged V, D, or J gene segment. In some embodiments, the amplicon cluster includes a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ. In specific embodiments, the VJ rearrangement is a kappa deletion element rearrangement. In some embodiments, the amplicon cluster includes a V gene segment region such as a hypermutagenic region, and / or a J gene segment region encoding a portion of CDR3. In some embodiments, the amplicon cluster includes a V leader sequence, a somatic hypermutagenic V region, and a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

[0275] In some embodiments, the non-temporary computer-readable storage medium includes further instructions for linking forward sequence reads with reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a second nucleic acid linker sequence, wherein each linkage is achieved by linking the second nucleic acid linker sequence between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read, thereby obtaining a second nucleic acid sequence result comprising, in that order, the portion from the reverse sequence read, the second nucleic acid linker sequence, and the reverse complement of the portion from the forward sequence read, and (1) the length of the portion from the forward sequence read is also determined by a selected bidirectional sequencing technique. (1) The length of the portion from the reverse sequence read is 75% or more of the maximum read length that can be obtained, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique; (2) The length of the portion from the reverse sequence read linked to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read linked to the first nucleic acid linker; (3) The length of the portion from the forward sequence read linked to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read linked to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read linked to the second nucleic acid linker; (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

[0276] In some embodiments, the length of the portion from the forward sequence read is approximately 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is approximately 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum read length obtained by the selected bidirectional sequencing technique. In some embodiments, the length of the portion from the reverse sequence read is the same for all reverse sequence reads being analyzed. In some embodiments, the length of the portion from the forward sequence read is the same for all forward sequence reads being analyzed, but may be the same as or different from the length of the portion from the reverse sequence read. In some embodiments, the length of the portion from the forward sequence read is the same as the length of the portion from the reverse sequence read.

[0277] In some embodiments, the forward sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the forward sequence read, and the reverse sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the reverse sequence read. In some embodiments, the specified number of adjacent nucleotides includes between approximately 80 and approximately 180 nucleotides. As used in this disclosure, the term “approximately” refers to ±10% of a given value. In some embodiments, the specified number of adjacent nucleotides includes approximately 80, approximately 90, approximately 100, approximately 110, approximately 120, approximately 130, approximately 140, approximately 150, approximately 160, approximately 170, or approximately 180 nucleotides.

[0278] In some embodiments, the first nucleic acid linker sequence is the same for all first nucleic acid sequence results. In some embodiments, the first nucleic acid linker sequence has a nucleotide length between 5 and 30, a nucleotide length between 5 and 25, or a nucleotide length between 5 and 20. In some embodiments, the length of the first nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides.

[0279] In some embodiments, the first nucleic acid linker sequence and the second nucleic acid linker sequence are at least 11 nucleotides long. In some embodiments, the first nucleic acid linker sequence and the second nucleic acid linker sequence are between 5 and 30 nucleotides long, between 5 and 25 nucleotides long, or between 5 and 20 nucleotides long. In some embodiments, the length of the first nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides long. In some embodiments, the length of the second nucleic acid linker sequence is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides long.

[0280] Further characteristics of the present invention are fully described by the following non-limiting embodiments. [Examples]

[0281] method Paired-end sequencing is a standard method for analyzing B-cell or T-cell clonality. When the sequencing length is sufficient, the entire rearrangement can be sequenced by utilizing the overlap between the two paired reads. This "complete" sequencing allows for a simpler analysis without the need for any further formatting steps. However, when the sequencing length is sufficient (e.g., due to platform limitations or assay design reasons), the analysis used in the "complete" sequencing scenario is prone to errors. Methods for analyzing non-overlapping sequencing data for the purpose of evaluating clonality are described herein.

[0282] The analytical method for "complete" sequencing (where paired reads overlap and the entire amplicon sequence can be identified) begins by identifying the overlap and generating a concatenated sequence containing the non-overlapping sequence of the unique read 1 (R1), followed by the overlapping sequence between read 1 and read 2 (R1 and R2), and finally the non-overlapping sequence of the unique read 2 (R2). If the sequencing platform / assay does not support generating overlapping sequences, the following modifications allow for downstream analysis.

[0283] Simple Taping: The simplest method is to "tape" the read pair (R1 and R2) together with the intrinsic sequence between them. Since downstream analysis involves alignment with a reference, it is important to use sequences that cannot be involved in this alignment process. A sequence of 11 "N"s is selected (11-Nmer) because such sequences generally do not align with the implementation of standard alignment algorithms (they are considered unknown nucleotides and therefore do not attempt to align the "N"s). First, the R2 read is reverse complementary (rcR2) so that it is sense-directed relative to R1. Next, the 11-Nmer is tethered to the end of R1. Finally, the R2 read is tethered to the end of the R1+11-Nmer sequence to produce the R1+11-Nmer+rcR2 read. This tethered read is now ready for downstream analysis.

[0284] Smart Taping: "Smart taping" is similar to simple taping, except that the read pairs are modified before being ligated to the 11-Nmer. R1 and R2 reads are first identified by gene-specific primers that amplified these reads, which is done by examining the first 20-25 nt of the sequence and matching it to a known primer sequence. A further 100 nt is saved from the end (i.e., the anchor point) of the primer sequence, and the remaining sequence is removed (for both R1 and R2 reads) to obtain "trimmed" R1 and R2 reads. At this point, the trimmed reads are processed in the same way as in simple taping: the trimmed R2 is reverse complementarized, the 11-Nmer is ligated to the trimmed R1, and the trimmed rcR2 is ligated to the trimmed R1+11-Nmer. These ligated trimmed reads are now ready for downstream analysis.

[0285] Downstream analysis: In short, identical reads are folded into a single entry with a counter attached to their header to annotate how many copies existed in the dataset. The folded reads are then aligned with references and assigned to the V and J genes based on optimal alignment, outputting quantitative information on the total count and relative frequency of each read. [Examples]

[0286] MISEQ paired-end array determination Dataset: A MiSeq sequencing run (2 x 251 cycles) consisting of 10% artificial cell line DNA diluted in tonsil background DNA was used to demonstrate the efficiency of the taping method. While a 2 x 251 cycle run allows for "complete" sequencing analysis of the selected target (LymphoTrack IGH FR1 assay), the data included in this run was truncated to mimic a 2 x 151 cycle by removing the last 100nt of all reads contained within the R1 and R2 pair files. The 2 x 251 cycle data is called the "control" dataset, while the truncated 2.151 cycle data is called the "tape test" dataset.

[0287] Furthermore, a Nextseq sequencing run (2 × 151 cycles) consisting of 100% cell line DNA was used to demonstrate a real-world application of the taping method's efficiency.

[0288] result Results for the MiSeq control dataset using complete sequencing: The control dataset was analyzed using a "complete" analysis consisting of paired read duplication before performing downstream analysis. The results are shown in Table 2.

[0289] [Table 2]

[0290] This is an expected result for this 10% artificial dataset using a "complete" sequencing platform / assay, where V3-J4 rearrangements are found at around 10% frequency (9.45% here).

[0291] Results from the MiSeq tape test dataset using simple taping: The MiSeq tape test dataset was analyzed using a "simple tape" analysis, which consists of adding an 11-Nmer sequence between the R1 and R2 reads. The results are shown in Table 3.

[0292] [Table 3]

[0293] This result shows that a simple taping method produces 10% of the cloned sequences, which are split into multiple sequences of different lengths. This appears to be due to the choice of location for placing the 11-Nmer during the taping process. Below are the upstream and downstream region alignments of the 11-Nmer for these top five reads, with dashed lines representing alignment gaps where sequences are absent in the read. Read ranks 2 and 5 have a single gap, while read rank 3 has a 4nt gap.

[0294] [ka]

[0295] During the simple taping process, the 11-Nmer is directly attached to the end of the R1 lead. A detailed examination of the taped area reveals that the end of the R1 lead is not a matched end at the same position as a lead presumed to have the same sequence. This phenomenon has clearly negative consequences, particularly because the lead sequences are no longer identical and do not fold during downstream analysis, thus reducing the upper lead signal.

[0296] Results of MiSeq tape test dataset using smart taping: The MiSeq tape test dataset was then analyzed using the smart taping method, which involves trimming sequences from R1 and R2 reads located more than 100 nt away from the primer site. The results are found in Table 4.

[0297] [Table 4]

[0298] These results demonstrate that reducing sequence length by using anchor points to trim the "ambiguous" ends of reads can restore the expected ratio measured by a full sequencing approach. [Examples]

[0299] NEXTSEQ paired-end sequencing determination Results from the NextSeq tape test dataset using simple taping: The NextSeq tape test dataset was analyzed using a "simple tape" analysis, which consists of adding an 11-Nmer sequence between the R1 and R2 reads. The results are shown in Table 5.

[0300] [Table 5]

[0301] This result shows that a simple taping method produces 100% clone sequences that are divided into multiple sequences of different lengths. This appears to stem from the choice of where to place the 11-Nmer during the taping process. Below are the upstream and downstream region alignments of the 11-Nmer for these top five reads, with dashed lines representing alignment gaps where sequences are absent in the read. Read rank 1 has a single gap, ranks 2 and 5 have three gaps, rank 3 has no gaps, and rank 4 has two gaps.

[0302] [ka]

[0303] During a simple taping process, 11-Nmer is directly connected to the end of the R1 lead and the start of rcR2. A detailed examination of the taped area reveals that the start of the rcR2 lead (which is also the end of the R2 lead) is not a coincident start at the same position for leads presumed to have the same sequence. This phenomenon has clearly negative consequences, particularly because the lead sequences are no longer identical and do not fold during downstream analysis, thus reducing the upper lead signal.

[0304] Results of NextSeq tape test dataset using smart taping: NextSeq tape test dataset was analyzed using the smart taping method, which involves trimming sequences from R1 and R2 reads located more than 100 nt away from the primer site. The results are found in Table 6.

[0305] [Table 6]

[0306] These results demonstrate that reducing the sequence length by using anchor points to trim the "ambiguous" ends of reads can greatly improve the captured signal.

[0307] Those skilled in the art will understand that the inventions described herein are susceptible to modifications and alterations other than those specifically described. It should be understood that the inventions include all such modifications and alterations. The inventions also include all of the processes, properties, compositions and compounds referred to or shown herein individually or collectively, as well as any and all combinations of any two or more of the aforementioned processes or properties.

Claims

1. A method for screening a target nucleic acid sample to express one or more target nucleotide sequences, (i) A step of spatially separating a library of individual template DNA molecules derived from the nucleic acid sample on a solid support, wherein the template DNA molecules are generated such that the target nucleotide sequence is localized in adjacent nucleotide regions at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying the forward and reverse sequence reads for the one or more clusters sequenced in step (iii) and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of a forward read, which is ligated at one of the terminal ends of a nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) The portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, which is linked at one end of the 3' end of the nucleic acid linker sequence, and whose other end is linked to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a); and the linker sequence is the same for all nucleic acid sequence results of (b); (v) A step of analyzing the sequence results and Methods that include...

2. A method for diagnosing, monitoring, or otherwise screening a patient's condition, wherein the condition is characterized by the expression of one or more target nucleotide sequences. (i) A step of spatially separating a library of individual template DNA molecules derived from a nucleic acid sample on a solid support, wherein the template DNA molecules are generated such that the target nucleotide sequence is localized in adjacent nucleotide regions at the 5' and / or 3' terminal ends of the template; (ii) A step of generating amplicon clusters by amplifying the spatially separated template DNA molecules, wherein each cluster is generated from individual spatially separated template DNA molecules, (iii) A step of sequencing one or more amplicons of one or more clusters in a bidirectional manner, wherein the forward and reverse sequence reads of the amplicon do not provide adjacent reads along the entire length of the amplicon. (iv) A step of identifying the forward and reverse sequence reads for the one or more clusters sequenced in step (iii) and generating a nucleic acid sequence result, wherein the nucleic acid sequence result is (a) a portion of the 5' adjacent nucleic acid sequence at the end of the forward read, which is ligated at one of the terminal ends of the nucleic acid linker sequence, with the other terminal end of the linker sequence being ligated at a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, and / or (b) The portion of the 5' adjacent nucleic acid sequence at the end of the reverse read, which is linked at one of the terminal ends of the nucleic acid linker sequence at its 3' end, and whose other terminal end is linked to a sequence complementary to the portion of the 5' adjacent nucleic acid sequence at the end of the forward read. Includes, (1) The portion is 75% or more of the maximum forward and reverse read lengths obtained by the selected bidirectional sequencing technique; (2) The portion of the adjacent sequence of the reverse read is the same for all reverse reads to be analyzed; (3) The portion of the adjacent sequence of the forward read is the same for all forward reads to be analyzed, but may be the same or different for the portion of the reverse read; (4) The linker sequence is the same for all nucleic acid sequence results of (a); and the linker sequence is the same for all nucleic acid sequence results of (b); (v) A step of analyzing the sequence results and Methods that include...

3. The method according to claim 1 or 2, wherein the nucleic acid region is DNA.

4. The method according to claim 2, wherein the nucleic acid sample for the purpose comprises B and / or T cell DNA, and the one or more target nucleotide sequences are one or more rearranged V, D, or J gene segments.

5. The method according to claim 3, wherein the target nucleotide sequence is a DJ or VDJ rearrangement of IgH, TCRβ, or TCRδ, or a kappa deletion element rearrangement.

6. The method according to claim 3, wherein the target nucleotide sequence is a VJ rearrangement of Igκ, Igλ, TCRα, or TCRγ.

7. The method according to claim 3, wherein the target nucleotide sequence is a J gene segment region encoding a V gene segment region such as a hypermutation-prone region and / or a portion of CDR3.

8. The method according to claim 3, wherein the target nucleotide sequence is a V leader sequence, a V region susceptible to somatic hypermutation, and a gene segment region encoding all or part of IgH FR1, IgH FR2, or IgH FR3.

9. The method according to claim 3, wherein the target nucleotide sequence is a BCL1 / JH or BCL2 / JH translocation or an internal tandem duplication or other mutation related to the FLT3 or TP53 gene.

10. The method according to any one of claims 1 to 3, wherein the solid support is a glass surface.

11. The method according to claim 10, wherein the glass surface is a glass slide or a flow cell.

12. The method according to any one of claims 1 to 11, wherein the template DNA molecule expresses one or more nucleic acid sequences corresponding to an index, a barcode, a unique molecular identifier, a sequencing primer hybridization site, and an index sequencing primer hybridization site at the 5' and / or 3' positions of the terminal.

13. The method according to any one of claims 1 to 12, wherein the adjacent nucleotide region in step (i) corresponds to about 80% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

14. The method according to any one of claims 1 to 13, wherein the adjacent nucleotide region corresponds to 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii), and the forward and reverse read portions are 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, or 83% or more of the maximum forward and reverse read lengths obtained by the bidirectional sequencing technique selected for use in step (iii).

15. The method according to claim 14, wherein the target DNA sequence is localized to 120 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, and the 20 nucleotide terminal ends of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

16. The method according to claim 14, wherein the target DNA sequence is localized to 125 adjacent nucleotides at the 5' and / or 3' terminal ends of the template, and up to 30 nucleotide terminals of the adjacent nucleotide region express one or more nucleotide sequences corresponding to an adapter, index, barcode, unique molecular identifier, sequencing primer hybridization site, or index sequencing primer hybridization site.

17. The method according to any one of claims 1 to 15, wherein the amplification is a bridge amplifier.

18. The method according to any one of claims 1 to 16, wherein the sequence is determined by synthesis using reversibly terminated labeled nucleotides.

19. The method according to any one of claims 1 to 18, wherein the nucleic acid linker is 5 to 30 nucleotides long, preferably 5 to 25, and more preferably 5 to 20 nucleotides long.

20. The linker has a length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides. The method according to claim 19.

21. The method according to any one of claims 1 to 20, wherein the analysis includes the step of aligning the nucleic acid sequence results generated in step (iv) and determining the expression of the target nucleic acid sequence of interest.

22. The method according to claim 2, wherein the state is characterized by a clonal population of cells or microorganisms.

23. The method according to claim 22, wherein the cloned cells are a population of cloned lymphocyte cells.

24. The method according to claim 2, wherein the state is characterized by one or more target nucleotide sequences expressed by immune cells.

25. The method according to claim 24, wherein the target nucleotide sequence is characterized by one or more rearranged V, D, or J gene segment sequences.

26. The method according to claim 25, wherein the condition characterized by the expression of features of one or more rearranged V, D, or J gene segment sequences is any other condition characterized by infection, transplantation, autoimmunity, immunodeficiency, allergic neoplasm, or T or B cell clonal proliferation.

27. The method according to claim 26, wherein the neoplasm is a lymphoid or myeloid neoplasm.

28. The method according to claim 27, wherein the lymphoid or myeloid neoplasm is acute lymphoblastic leukemia, acute lymphoblastic leukemia, acute myeloid leukemia, acute promyelocytic leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, myeloproliferative neoplasm, such as myeloma, systemic mastocytosis, lymphoma, or hairy cell leukemia.

29. The method according to claim 27 or 28, used for detecting minimal residual lesions.

30. The method according to claim 26, wherein the condition is graft rejection, immunotherapy, polycythemia vera, myelodysplasia and leukocytosis.

31. The method according to claim 30, wherein the leukocytosis is lymphocytosis.

32. The method according to claim 2, applicable to diagnosis, prognosis, prediction of disease risk, detection of disease recurrence, immune surveillance, or monitoring of prophylactic or therapeutic effects.

33. A computer implementation method for generating nucleic acid sequence results for analysis from non-duplicate sequence reads, A step of identifying forward sequence reads and reverse sequence reads from sequence reads of an amplicon cluster, wherein the cluster is generated from individual spatially separated template DNA molecules, each sequence read is generated by a selected bidirectional sequencing technique, the forward sequence read and the reverse sequence read do not overlap, and no adjacent reads are provided across the entire length of any amplicon. A step of obtaining a plurality of first nucleic acid sequence results by linking the forward sequence read to the reverse sequence read such that each forward sequence read is linked to the reverse sequence read, and each reverse sequence read is linked to the forward sequence read via a first nucleic acid linker sequence, wherein each linkage is The first nucleic acid linker sequence is linked between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read, thereby obtaining a first nucleic acid sequence result that includes the forward sequence read portion, the first nucleic acid linker sequence, and the reverse complement of the reverse sequence read portion in that order. The process and Includes, A computer implementation method wherein (1) the length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) the length of the portion from the reverse sequence read is the same for all reverse sequence reads to be analyzed, (3) the length of the portion from the forward sequence read is the same for all forward sequence reads to be analyzed, but may be the same as or different from the length of the portion from the reverse sequence read, and (4) the first nucleic acid linker sequence is the same for all first nucleic acid sequencing results.

34. A step of obtaining a plurality of second nucleic acid sequence results by linking the forward sequence reads to the reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to the forward sequence read via a second nucleic acid linker sequence, wherein each linkage is The second nucleic acid linker sequence is linked between the 3' end of the portion of the 5' adjacent nucleic acid sequence at the end of the reverse sequence read and the reverse complement of the portion of the 5' adjacent nucleic acid sequence at the end of the forward sequence read, thereby obtaining a second nucleic acid sequence result that includes, in that order, the portion from the reverse sequence read, the second nucleic acid linker sequence, and the reverse complement of the portion from the forward sequence read. The process further includes, (1) The length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) The length of the portion from the reverse sequence read connected to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read connected to the first nucleic acid linker, (3) The length of the portion from the forward sequence read connected to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read connected to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read connected to the second nucleic acid linker, and (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

35. The computer implementation method according to claim 34, wherein the first nucleic acid linker sequence and the second nucleic acid linker sequence are at least 11 nucleotides long.

36. The computer implementation method according to claim 33, wherein the length of the forward array read portion is the same as the length of the reverse array read portion.

37. The computer implementation method according to claim 33, wherein the portion of the forward sequence read includes a specified number of adjacent nucleotides at the 5' end of the forward sequence read, and the portion of the reverse sequence read includes a specified number of adjacent nucleotides at the 5' end of the reverse sequence read.

38. The computer implementation method according to claim 37, wherein the specified number of adjacent nucleotides includes between approximately 80 and approximately 180 nucleotides.

39. The computer implementation method according to any one of claims 33 to 38, wherein the forward and reverse sequence reads are DNA sequence reads.

40. The computer implementation method according to any one of claims 33 to 39, wherein the cluster of the amplicon is amplified from B and / or T cell DNA.

41. The computer implementation method according to claim 40, wherein the cluster of the amplicon comprises at least one rearranged V, D, or J gene segment.

42. A non-temporary computer-readable storage medium having embodied program instructions, wherein the program instructions, which are executable by the device's processing elements, A step of identifying forward sequence reads and reverse sequence reads from sequence reads of an amplicon cluster, wherein the cluster is generated from individual spatially separated template DNA molecules, each sequence read is generated by a selected bidirectional sequencing technique, the forward sequence read and the reverse sequence read do not overlap, and no adjacent reads are provided across the entire length of any amplicon. A step of obtaining a plurality of first nucleic acid sequence results by linking the forward sequence read to the reverse sequence read such that each forward sequence read is linked to the reverse sequence read, and each reverse sequence read is linked to the forward sequence read via a first nucleic acid linker sequence, wherein each linkage is The first nucleic acid linker sequence is linked between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read, thereby obtaining a first nucleic acid sequence result that includes the forward sequence read portion, the first nucleic acid linker sequence, and the reverse complement of the reverse sequence read portion in that order. The process and The device implements a method for creating nucleic acid sequence results for analysis from non-overlapping sequence reads. (1) The length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) The length of the portion from the reverse sequence read is the same for all reverse sequence reads to be analyzed, (3) The length of the portion from the forward sequence read is the same for all forward sequence reads to be analyzed, but may be the same as or different from the length of the portion from the reverse sequence read, and (4) The first nucleic acid linker sequence is the same for all first nucleic acid sequencing results, a non-temporary computer-readable storage medium.

43. A step of obtaining a plurality of second nucleic acid sequence results by linking the forward sequence reads to the reverse sequence reads such that each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to the forward sequence read via a second nucleic acid linker sequence, wherein each linkage is The second nucleic acid linker sequence is linked between the 3' end of the portion of the 5' adjacent nucleic acid sequence at the end of the reverse sequence read and the reverse complement of the portion of the 5' adjacent nucleic acid sequence at the end of the forward sequence read, thereby obtaining a second nucleic acid sequence result that includes, in that order, the portion from the reverse sequence read, the second nucleic acid linker sequence, and the reverse complement of the portion from the forward sequence read. The process further includes, (1) The length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) The length of the portion from the reverse sequence read connected to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read connected to the first nucleic acid linker, (3) The length of the portion from the forward sequence read connected to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read connected to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read connected to the second nucleic acid linker, and (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

44. The non-temporary computer-readable storage medium according to claim 42, wherein the first nucleic acid linker sequence and the second nucleic acid linker sequence are at least 11 nucleotides long.

45. The non-temporary computer-readable storage medium according to claim 42, wherein the length of the forward array read portion is the same as the length of the reverse array read portion.

46. The non-temporary computer-readable storage medium according to claim 42, wherein the portion of the forward sequence read includes a specified number of adjacent nucleotides at the 5' end of the forward sequence read, and the portion of the reverse sequence read includes a specified number of adjacent nucleotides at the 5' end of the reverse sequence read.

47. The non-temporary computer-readable storage medium according to claim 46, wherein the specified number of adjacent nucleotides includes between approximately 80 and approximately 180 nucleotides.

48. The non-temporary computer-readable storage medium according to any one of claims 42 to 47, wherein the forward and reverse sequence reads are DNA sequence reads.

49. The non-temporary computer-readable storage medium according to any one of claims 42 to 48, wherein the cluster of the amplicon is amplified from B and / or T cell DNA.

50. The non-temporary computer-readable storage medium according to claim 49, wherein the cluster of the amplicon comprises at least one rearranged V, D, or J gene segment.

51. A device including a hardware processor for generating nucleic acid sequence results for analysis from non-duplicate sequence reads, The aforementioned hardware processor is It is configured to identify forward and reverse sequence reads from sequence reads of an amplicon cluster, where the cluster is generated from individual spatially separated template DNA molecules, each sequence read is generated by a selected bidirectional sequencing technique, the forward and reverse sequence reads do not overlap, and no adjacent reads are provided across the entire length of any amplicon, and furthermore, Each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a first nucleic acid linker sequence, thereby obtaining a plurality of first nucleic acid sequence results, where each linkage is The first nucleic acid linker sequence is linked between the 3' end of the 5' adjacent nucleic acid sequence portion at the end of the forward sequence read and the reverse complement of the 5' adjacent nucleic acid sequence portion at the end of the reverse sequence read, thereby obtaining a first nucleic acid sequence result that includes the forward sequence read portion, the first nucleic acid linker sequence, and the reverse complement of the reverse sequence read portion in that order. Achieved by, (1) The length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) The length of the portion from the reverse sequence read is the same for all reverse sequence reads to be analyzed, (3) The length of the portion from the forward sequence read is the same for all forward sequence reads to be analyzed, but may be the same as or different from the length of the portion from the reverse sequence read, and (4) The first nucleic acid linker sequence is the same for all first nucleic acid sequencing results.

52. The aforementioned hardware processor Each forward sequence read is linked to a reverse sequence read, and each reverse sequence read is linked to a forward sequence read via a second nucleic acid linker sequence, thereby obtaining a plurality of second nucleic acid sequence results, and each linkage is, The second nucleic acid linker sequence is linked between the 3' end of the portion of the 5' adjacent nucleic acid sequence at the end of the reverse sequence read and the reverse complement of the portion of the 5' adjacent nucleic acid sequence at the end of the forward sequence read, thereby obtaining a second nucleic acid sequence result that includes, in that order, the portion from the reverse sequence read, the second nucleic acid linker sequence, and the reverse complement of the portion from the forward sequence read. Further configured to be achieved by, (1) The length of the portion from the forward sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, and the length of the portion from the reverse sequence read is 75% or more of the maximum read length obtained by the selected bidirectional sequencing technique, (2) The length of the portion from the reverse sequence read connected to the second nucleic acid linker is the same for all reverse sequence reads and is the same as the length of the portion from the reverse sequence read connected to the first nucleic acid linker, (3) The length of the portion from the forward sequence read connected to the second nucleic acid linker is the same for all forward sequence reads and is the same as the length of the portion from the forward sequence read connected to the first nucleic acid linker, but may be the same as or different from the length of the portion from the reverse sequence read connected to the second nucleic acid linker, and (4) The second nucleic acid linker sequence is the same for all second nucleic acid sequencing results.

53. The device according to claim 52, wherein the first nucleic acid linker sequence and the second nucleic acid linker sequence are at least 11 nucleotides long.

54. The device according to claim 51, wherein the length of the forward array read portion is the same as the length of the reverse array read portion.

55. The device according to claim 51, wherein the forward sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the forward sequence read, and the reverse sequence read portion includes a specified number of adjacent nucleotides at the 5' end of the reverse sequence read.

56. The device according to claim 55, wherein the specified number of adjacent nucleotides includes between approximately 80 and approximately 180 nucleotides.

57. The device according to any one of claims 51 to 56, wherein the forward and reverse sequence reads are DNA sequence reads.

58. The device according to any one of claims 51 to 57, wherein the cluster of the amplicon is amplified from B and / or T cell DNA.

59. The device according to claim 58, wherein the cluster of the amplicon comprises at least one rearranged V, D, or J gene segment.