Intermolecular consensus of partially read sequences

The method addresses incomplete reads in nanopore sequencing by merging full and partial length reads, improving sequencing accuracy and efficiency through clustering and tiering techniques, thus optimizing bioinformatics workflows.

WO2025235888A1PCT designated stage Publication Date: 2025-11-13ROCHE SEQUENCING SOLUTIONS INC

Patent Information

Application Number
PCT/US2025/028649
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-10
Filing Date
2025-05-09
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Existing nanopore sequencing techniques often result in incomplete reads due to polymerase enzyme detachment from template DNA molecules, leading to inaccurate and inefficient nucleotide sequence identification, with incomplete reads being discarded and wasting processing resources.

Method used

A computer-implemented method for identifying consensus molecular sequences by merging full length and partial length reads based on positional distance and other criteria, utilizing clustering and tiering techniques to enhance sequencing accuracy and efficiency.

Benefits of technology

Improves sequencing accuracy and reduces the number of reads required, enhancing the quality of consensus sequences and optimizing downstream bioinformatics workflows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025028649_13112025_PF_FP_ABST
    Figure US2025028649_13112025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for identifying a consensus molecular sequence from full length and partial length read sequencing data of a sample. The method includes sorting the sequencing data into full length and partial length reads with 5' and / or 3' binding sites, clustering the full length reads and the partial length reads, the clustering based on a position of binding site(s), and comparing a partial length read cluster to a full length read cluster to determine a positional distance between the compared clusters. In response to determining that the positional distance between the compared clusters falls within a redetermined threshold, merging the compared clusters, and identifying a consensus molecular sequence of the sample based on the merged clusters.
Need to check novelty before this filing date? Find Prior Art

Description

INTERMOLECULAR CONSENSUS OF PARTIALLY READ SEQUENCESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This Application claims the benefit of United States Provisional Patent Application No. 63 / 645,433, titled “INTERMOLECULAR CONSENSUS OF PARTIALLY READ SEQUENCES”, filed on May 10, 2024, which is hereby incorporated by reference in its entirety.BACKGROUND

[0002] Biological samples can be used to identify a nucleotide sequence of nucleic acid material that indicates ordered sets of nucleic acid bases in the nucleic acid material. Rather recently, single-molecule sequencing has been rapidly advancing in techniques and expanding in applications. These techniques can sequence individual molecules of nucleic acid material (and analogs or molecules derived therefrom) and can be performed in real-time without PCR amplification. Traditional PCR amplification techniques generally involve significant time to produce enough molecules to facilitate reads in the tens of millions of base pairs and may be subject to high error rates.

[0003] Some newer techniques involve chemistry that translates a sequence of DNA into a measurable surrogate molecule called an Xpandomer. Xpandomer synthesis is based on the natural function of DNA replication uses expandable nucleotide triphosphates (X-NTPs) that act as substrates for template-dependent, polymerase-based replication of nucleic acid material. As the Xpandomer molecule transits through nanopore-sized openings in an electrode-resistant membrane, each nanopore corresponding to a selective channel, a distinct electrical signal is generated for each base reporter and identifiable to enable highly accurate and high throughput nanopore-based nucleic acid sequencing (also referred to generally as nanopore sequencing).

[0004] The techniques hold incredible promise with regard to a number of different applications, including, but not limited to: building of comprehensive libraries relating genes to diseases; identifying and characterizing new diseases; characterizing rare diseases; and identifying therapies to treat various diseases. Although rapid and accurate, such sequencing techniques (e.g., third-generation or next-generation techniques) can result in incompletereads which, by themselves, fail to adequately identify a nucleotide sequence and are thus discarded, wasting processing resources. For example, during Xpandomer synthesis, the polymerase enzyme can detach from the template DNA molecules leading to incomplete synthesis of the surrogate Xpandomer molecules and, consequently, incomplete reads during sequencing operations. Thus, improved techniques for utilization and processing of incomplete reads are desired.SUMMARY[00051 The embodiments described herein relate to systems and methods for sequencing nucleic acid material, or surrogate molecules derived therefrom. More particularly, the embodiments described herein relate to techniques for generating consensus molecular sequences based on full length and partial length reads produced by sequencer instruments.

[0006] Computer computer-implemented methods and systems are provided for identifying consensus molecular sequences in which the methods may utilize sequencing data including both full length reads and partial length reads. The method may be used to improve the accuracy and speed of sequencing where partial reads may be frequently encountered, for example. In some embodiments, the method is applied to “nanopore” sequencing in which the sequencing data is based on a template-directed synthesis targeting a nucleic acid structured with a segment that produces a synchronization signal upon passage through a nanopore.

[0007] In accordance with a first aspect of the present disclosure, a computer- implemented method for identifying a consensus molecular sequence includes obtaining sequencing data from a physical sample comprising a plurality of full length reads and partial length reads; sorting the sequencing data into clusters of full length reads and partial length reads; comparing a partial length read cluster to a full length read cluster to determine a positional distance between the compared clusters; in response to determining that the positional distance between the compared clusters falls within a predetermined threshold, merging the compared partial length read cluster and full length read cluster; and identifying and reporting a consensus molecular sequence of the sample based on the merged partial length read cluster and full length read cluster. Each read in the sequencing data comprises atleast one 5’ and / or 3’ binding sites. The clusters are based on position information of the 5’ and / or 3’ binding sites included in the reads.

[0008] In some embodiments of the first aspect, merging the compared partial length read cluster and full length read cluster is further in response to determining that a count of reads in the full length read cluster is at least two.

[0009] In some embodiments of the first aspect, the merged full length read cluster has fewer than two reads.

[0010] In some embodiments of the first aspect, merging the compared partial length read cluster and full length read cluster is further in response to determining that a count of reads in the full length read cluster with a positional distance from the respective partial length read cluster is not more than one.

[0011] In some embodiments of the first aspect, merging the compared partial length read cluster and full length read cluster is further in response to determining that none of the reads in the compared partial length read cluster has a length greater than a length of the read(s) of the compared full length read cluster.

[0012] In some embodiments of the first aspect, the method further including merging two full length read clusters having two or fewer reads and a positional distance within a predetermined threshold with each other.

[0013] In some embodiments of the first aspect, sorting the sequence data into clusters comprises sorting and aligning the reads in relation to bases and sequence positions prior to the sorting into full length reads and partial length reads.

[0014] In some embodiments of the first aspect, merging the compared partial length read cluster and full length read cluster is further in response to determining that a total number of the reads in the partial length read cluster comprises thirty percent or more of a total count of reads in the combined full length read cluster and partial length read cluster.

[0015] In some embodiments of the first aspect, the sequencing data from the physical sample is based on a template-directed synthesis for targeting a nucleic acid structured with a segment that produces a synchronization signal upon passage through a nanopore.

[0016] In accordance with a second aspect of the present disclosure, a computer- implemented method for identifying a consensus molecular sequence includes obtaining sequencing data for a physical sample comprising a plurality of full length reads and partial length reads; sorting the sequencing data into clusters of full length and partial length reads; sorting the clusters into tiers of clusters, a first tier comprising only full length reads, a second tier comprising both full length reads and partial length reads, and a third tier comprising only partial length reads; discarding the third tier of clusters; determining a consensus sequence among each of the first tier clusters; determining a consensus sequence among each of the second tier clusters by calculating a likelihood of partial length reads or full length reads within the second tier clusters of supporting the consensus sequence and determining that the calculated likelihood exceeds a particular threshold; and identifying and reporting one or more consensus molecular sequences of the physical sample based on the determined consensus sequences for the first tier and / or second tier of clusters.

[0017] In some embodiments of the second aspect, determining the consensus sequence within the respective first tier cluster comprises determining a sequence with a greatest number of occurrences included in the reads of the first tier cluster.

[0018] In some embodiments of the second aspect, the calculating the likelihood of partial length reads or full length reads supporting the consensus sequence is based on a modeling of a probability of error in a candidate consensus sequence to follow a binomial distribution with a probability number representing a base coverage and a count of reads within the respective cluster supporting a position of the base of the candidate consensus sequence.

[0019] In some embodiments of the second aspect, the calculating the likelihood of partial length reads or full length reads supporting the consensus sequence is based on a deep machine learning inference model comprising a gap-aware transformer-encoder configured to determine a sequence correction utilizing an alignment-based loss function.

[0020] In accordance with a third aspect of the present disclosure, system for identifying a consensus molecular sequence includes one or more processors associated with a sequencer instrument. The one or more processors are configured to: obtain, from the sequencer instrument, sequencing data from a physical sample comprising a plurality of full lengthreads and partial length reads; sort the sequencing data into clusters of full length reads and partial length reads; compare a partial length read cluster to a full length read cluster to determine a positional distance between the compared clusters; in response to determining that the positional distance between the compared clusters falls within a predetermined threshold, merge the compared partial length read cluster and full length read clusters; and identify and report a consensus molecular sequence of the sample based on the merged partial length read cluster and full length read cluster.

[0021] In accordance with some embodiments of the third aspect, the sequencer instrument is in communication with the one or more processors via a network.

[0022] In accordance with some embodiments of the third aspect, the sequencer instrument is configured to generate the sequencing data from a template-directed synthesis for targeting a nucleic acid structured with a segment that produces a synchronization signal upon passage through a nanopore.

[0023] In accordance with some embodiments of the third aspect, merging the compared partial length read cluster and full length read cluster is further in response to determining that a count of reads in the full length read cluster is at least two.

[0024] In accordance with some embodiments of the third aspect, the one or more processors are further configured to merge a candidate full length read cluster having fewer than two reads with the merged full length read cluster and partial length read cluster in response to determining that the candidate full length read cluster has a positional distance within a predetermined threshold from respective partial length reads of the merged full length read cluster and partial length read cluster.BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The detailed description is set forth with reference to the accompanying figures.

[0026] FIG. 1 is a process flow for identifying a consensus molecular sequence, in accordance with some embodiments.

[0027] FIG. 2 is an illustrative mapping and clustering of molecular sequence reads, in accordance with some embodiments.

[0028] FIG. 3 is an illustrative clustering and tiering of molecular sequence reads, in accordance with some embodiments.

[0029] FIG. 4 illustrates a process flow for identifying a consensus molecular sequence, in accordance with some embodiments.

[0030] FIG. 5 illustrates separate methods of merging clusters utilizing a tiering system, in accordance with some embodiments.

[0031] FIG. 6A is a chart of clustering results of reads obtained from a sample, in accordance with some embodiments.

[0032] FIG. 6B is a chart representing separate clustering results based on minimum read count of first tier (unmatched full read) clusters, in accordance with some embodiments.

[0033] FIG. 7 illustrates an example computer system that may be utilized to implement techniques disclosed herein.DETAILED DESCRIPTION

[0034] The embodiments described herein can be implemented in numerous ways, including as a process; an apparatus; a system; a computer program product embodied on a computer readable storage medium; and / or in conjunction with a processor, such as a processor configured to execute instructions stored on and / or provided by a memory coupled to the processor. As used herein, the term "processor" refers to one or more electronic devices, circuits, and / or processing cores configured to process data, such as computer program instructions. In this specification, these implementations, or any other form that the embodiments may take, may be referred to as techniques. In general, the order of the steps of one or more disclosed processes may be altered without departing from the scope of the disclosed embodiments.

[0035] A detailed description of one or more embodiments shown in the accompanying figures is provided below to illustrate the principles of the techniques disclosed herein. These techniques are described in connection with such embodiments, but the invention is not limited to any particular embodiment. The scope of the invention is limited only by the claims and the invention encompasses numerous alternatives, modifications, and / orequivalents. Numerous specific details are set forth in the following description in order to provide a thorough understanding of the invention to one of skill in the art. These details are provided for the purpose of example and the invention may be practiced according to the claims without some of the specific details or with additional details not described herein.

[0036] Disclosed herein are systems and methods for processing sequencing data received from a sequencer instrument. The sequencing data may include both full length reads and partial length reads, which are used to determine consensus sequences. Many sequencer instruments obtain a high rate of partial length reads and / or a low number of full length reads from a sample. Embodiments described herein utilize both full length and partial length reads to determine one or more “consensus” sequence(s), thereby reducing the total number of reads, and increasing the quality of consensus sequences, analyzed in downstream bioinformatics workflows.Definition of terms

[0037] As used herein throughout this disclosure, the term “Xpandomer” refers to a class of surrogate molecules derived from nucleic acids (such as DNA or RNA), which may in turn be used in a nanopore-based sequencing technique known as sequencing by expansion (SBX) to identify an original nucleic acid sequence. Xpandomer molecules are derived or synthesized based on the natural function of DNA replication where expandable nucleotide triphosphates (X-NTPs) act as substrates for template-dependent, polymerase-based replication. Specifically, four differentiable X-NTPs are used during Xpandomer synthesis, one for each DNA base, and engineered polymerases incorporate the X-NTPs into a resulting Xpandomer, which then serves as a surrogate for the complement of a nucleic acid template. When nanopore-based sequencing is conducted using the synthesized Xpandomers, accordingly, each incorporated X-NTP within the Xpandomer will function as a high signal- to-noise reporter of the original DNA base to which the incorporated X-NTP corresponds.Thus, as the Xpandomer molecule transits through a nanopore, the distinct electrical signal of each base reporter is easily identifiable to enable highly accurate and high throughput nanopore-based nucleic acid sequencing.

[0038] As used herein throughout this disclosure, the term “about” when used in connection with a referenced numeric indication means the referenced numeric indicationplus or minus up to 10% of that referenced numeric indication. For example, the term about 50 covers the range of 45 to 55, inclusive.

[0039] As used herein throughout this disclosure, the specific words chosen to describe one or more embodiments and optional elements or features are not intended to limit the invention. For example, spatially relative terms - such as “beneath,” “below,” “lower,” “above,” “upper,” “proximate,” and the like - may be used to describe the relationship of one element or feature to another element or feature as illustrated in the figures. These spatially relative terms are intended to encompass different positions (i.e. , translational placements) and orientations (i.e., rotational placements) of a device, component, element, or feature in use or operation in addition to the position and orientation shown in the figures. For example, if a device in the figures is turned over, elements described as below or beneath other elements or features would then be above or over the other elements or features in the new orientation of the device. Thus, the term below can encompass both positions and orientations of above and below. A device component, element, or feature may be otherwise oriented (e.g., rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly. Likewise, descriptions of movement along (i.e., translation) or around (i.e., rotation) various axes include various spatial device positions and orientation.

[0040] Similarly, as used herein throughout this disclosure, geometric terms such as but not limited to “parallel,” “perpendicular,” “round,” or “square” are not intended to require absolute geometric precision, unless clearly contradicted by context. Instead, such geometric terms encompass tolerances that allow for variations due to manufacturing or equivalent functions. For example, if an element is described as “round” or “generally round,” a component that is not precisely circular (e.g., one that is elliptical or polygonal, given the shape approximates a circle) is still encompassed by this description.

[0041] In addition, the singular forms “a,” “an,” and “the” are intended to include the corresponding plural forms as well, unless clearly contradicted by the context. The terms “comprises,” “includes,” “has,” and the like specify the presence of the stated features, steps, operations, elements, components, etc. but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, or groups. The term “at least one” followed by a list of one or more items (for example, “at least one of A and B”) is to be construed to mean one item selected from the listed items (i.e., A or B) or any combination oftwo or more of the listed items (i.e., A and B), unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illustrate the disclosed subject matter and is not intended to pose a limitation on the scope of the claimed invention.Intermolecular Consensus Workflow

[0042] FIG. 1 is a flowchart of a method 100 for identifying a consensus molecular sequence, in accordance with some embodiments. The method 100 may be implemented by a computing device including at least one processor, configured to execute programmable instructions stored on a non-transitory computer readable medium. In some embodiments, the method 100 may be implemented by a control system of a sequencer instrument configured to generate sequencing data during a biological assay of a sample. The control system may include logic, such as hardware, firmware, software, or any combination thereof, for implementing the steps of the method 100. In yet other embodiments, the method 100 may be implemented by a remote device (e.g., a server device) in communication with the sequencer instrument, and may receive sequencing data over a network. Any system capable of implementing one or more steps of the method 100 is within the scope of the present disclosure.

[0043] As shown in FIG. 1, at block 100, sequencing data is obtained from a biological sample. The sequencing data can include a number of individual reads. Each read can represent a chain of bonded elements of a molecule in the sample, such as that of a nucleic acid (or surrogate molecule), methyl group, or protein molecule, for example. In some embodiments, the sequencing data is positionally aligned based on the identification of 5 ’ and / or 3’ binding sites, which refer to positions on either end of the target fragment being sequenced where adapters that include unique molecular identifiers (UMI) and / or other random or fixed indexes are ligated to the target fragments before PCR amplification. The sequencing data may be obtained (e.g., received) from a sequencer instrument such as one that utilizes a template-directed synthesis for targeting a nucleic acid or other molecule structured with a segment that produces a signal upon passage through a nanopore. Exemplary embodiments of such sequencer instruments include, but are not limited to, sequencer instruments from Illumina, Inc. (MiniSeq™, NextSeq™ 550, etc.), OxfordNanopore Technologies, pic (MinlON™, GridlON™, etc.). Pacific Biosciences of California, Inc. (Revio™, Vega™, etc.), or the like.

[0044] At block 110, the reads are sorted into full length and partial length reads. Full length reads may be based on reads of a target molecular sequence having both 5’ and 3’ binding sites, and partial length reads may be based on reads of the target molecular sequence having only one of the 5’ or 3’ binding sites included in the read. At block 115, partial reads not identified as having 5’ binding sites are omitted from clustering.[00451 At block 120, full length reads matched by UMI position and / or binding site positional information are sorted into respective full length clusters. At block 130, partial length reads matched by UMI position and / or binding site positional information of the 5’ end are sorted into respectively matched partial read clusters. In some embodiments, full and partial reads are classified differently and / or apart from the binding sites of DNA strands (e.g., for protein or methyl group sequences).

[0046] At block 140, based on the clustering, matches (or families of reads) are determined between and among the partial and full read clusters. Matching may include determining that a full length read cluster has binding site (5’) positional distances within a predetermined threshold from a respective partial length read cluster or other full length read cluster. In some embodiments, the predetermined threshold is a maximal positional distance of one or two positions between the clusters. At block 135, partial read clusters that are not matched with full read clusters are discarded and no longer used for determining a consensus.

[0047] In some embodiments, collision avoidance measures are performed such as to avoid merging reads belonging to different molecules (e.g., of different DNA). Collisions can occur such as from incorrectly associating UIDs or binding site information among different molecules. At block 145, a determination is made that partial read clusters are discarded if they are matched with multiple full read clusters each exceeding a certain threshold of reads. The threshold can be based on determining that multiple (matched) full read clusters of a certain size provide sufficient data for establishing a consensus sequence, and that the corresponding (matched) partial read clusters are only likely to provide unneeded cumulative information or may have been incorrectly matched. In some embodiments, a threshold ofgreater than two full read sequences in multiple clusters provides a basis for excluding / discarding a corresponding (matched) partial read cluster.

[0048] At block 155, matched full and partial read clusters (not otherwise excluded at block 145) are further analyzed to determine possible irregularities (or inconsistencies) among the matched clusters. An irregularity may include, for example, a partial read sequence in a cluster that is longer than a full read sequence of the matched cluster. Based on identifying such irregularities, the associated partial cluster may be discarded.

[0049] In some embodiments, further measures are taken to optimize merging of the partial read clusters and full read clusters. At block 150, a determination is made as to whether a partial length read cluster is matched with full length read clusters having sizes of 1, 2, and greater than 2. If a partial length read cluster is matched with separate full length read clusters of the three different size categories, then the partial length read cluster is merged with the full length read cluster having a size greater than 2 and the remaining full length read clusters are merged with each other. Otherwise, the partial length read cluster and the full read clusters with which it is matched are merged together.

[0050] FIG. 2 is an illustrative mapping and clustering of molecular sequence reads, in accordance with some embodiments, A first cluster 210 of full length reads includes those identified as having both 5’ and 3’ binding sites positioned within a predetermined threshold distance. A second cluster 215 of partial length reads includes those identified as having 5’ binding sites and positioned within a predetermined threshold distance of each other (e.g., one or two positions). In some embodiments, the first and second clusters 210 and 215 are matched with each other based on the proximity of the position(s) of the 5’ binding sites between the clusters (e.g., within a predetermined threshold distance of one or two positions). In some embodiments, other identifiers may be used for matching (e.g., UMI).

[0051] Clusters 210 and 220 each represent full length reads (e.g., reads matched based on both their 5’ and 3’ binding sites). Clusters 225 and 235 each include partial length reads matched with each other based on the respective 5’ binding sites (and lack of 3’ binding site information). Reads within a cluster can be matched based on differences in binding site positions and / or UMI with respect to each other. Partial length read cluster 225 is matched with full length read cluster 220 based on a proximate range / difference of identified 5’binding sites. Partial length read cluster 235 is similarly matched with both full length read clusters 220 and 230. When a partial length read cluster is matched with multiple full length read clusters, a further optimization can be performed to select which (if any) full length read clusters to merge the partial length read cluster with such as described above with respect to blocks 145 and 150 of FIG. 1.

[0052] A partial length read cluster 240 is determined to be matched with full length read cluster 230 such as further described herein. Based on determining irregularities / inconsistencies between partial length read cluster 240 and full length read cluster 230 (apart from their matching characteristics), partial length read cluster 240 may be discarded such as in order to avoid introducing invalid data and / or read collisions. In some embodiments, an irregularity includes a cluster with partial length reads identified as exceeding the length of corresponding full length reads in a matching cluster, reflecting a significant probability that the reads are not associated with the same molecule. Unmatched clusters 245, 250, and 255 of partial length reads can be omitted from further consensus determinations as further described herein.

[0053] FIG. 3 is an illustrative clustering and tiering of molecular sequence reads, in accordance with some embodiments. Tiering can be used to rank preferences for utilizing respective clusters for consensus calling such as further described herein. A first tier 310 includes full length read clusters 312 and 317 that have been matched with each other while not matched with partial length read clusters. A second tier 320 of clusters includes a mixture of full length read clusters 325 and 327 and partial length read clusters 330 and 335 matched with each other. In some embodiments, exceeding a predetermined threshold of full length read clusters matched with each other is used apart from partial read clusters to form a consensus.

[0054] In some embodiments, a consensus for a cluster or selection of matched full length reads is determined by identifying the most common sequence among them and / or other methods known to one of ordinary skill in the art (e.g., a neural network). However, if a cluster of full length reads contains only two reads or one read, a determination (e.g,, a consensus Q score based on a probability of error) may be made that insufficient information is gleaned to establish an accurate consensus and the read(s) are discarded from consideration.

[0055] In some embodiments, second tier clusters of full length and partial length reads are merged such as further described herein (e.g., with respect to FIG. 1). A third tier 340 of clusters includes only partial length read clusters 345, 347, and 350. As further described herein, clusters including and only matched with other partial length reads are discarded from further consideration (e.g., block 135 in FIG. 1). In some embodiments, at least one of the partial length reads is considered as a full length read and the analysis proceeds accordingly. However, because third tier clusters cannot form a duplex molecule (family that contains reads on both forward and reverse strands), these non-duplex clusters cannot mitigate an error due to DNA damage since those tend to be present in the majority of reads on only one strand. In this case, duplex molecule reads can be helpful.

[0056] FIG. 4 illustrates a process flow for identifying a consensus molecular sequence, in accordance with some embodiments. At block 410, sequencing data is obtained for a sample, such as from a nanopore sequencer instrument. At block 415, the sequencing data can be sorted (e.g., clustering of reads) and aligned by base positions and / or UMI. At block 420, the clusters are further compared and matched with each other (e.g., as further described herein), and sorted into tiers of clusters. A first tier 440 includes full length read clusters not matched with partial length read clusters. A second tier 445 includes full length read clusters matched with partial length read clusters and a third tier 450 includes only partial length read clusters that were matched with each other (and not matched with any full length read clusters).

[0057] At block 425, consensus sequence(s) are calculated using first tier clusters. A consensus can be calculated such as by determining a commonality among the clusters with an acceptable level of probability (or error). Numerous techniques such as those known to one of ordinary' skill in the art can be employed to determine commonality among clustered full length sequence reads.

[0058] At block 430, consensus sequence(s) are calculated among second tier clusters. In some embodiments, at each base, a consensus call is evaluated. Since the per-base support varies in a family of (merged) reads with partial length reads, our confidence in a consensus call varies. A likelihood may be calculated for the second tier clusters of supporting a consensus sequence and determining that the calculated likelihood exceeds a particular threshold. In some embodiments, we introduce a consensus Q score by modeling theprobability of error in the consensus to follow' a binomial distribution with N draw' (per-base coverage) with M success (number of reads supporting the consensus base at this position). Therefore, for instance, if only a single full length read supports a particular base call, and no partial length reads support that base call, the consensus base call will be similar to the base call identified in the full length read (family size 1), but a very low7consensus Q score will be assigned.

[0059] In some embodiments, a deep machine learning inference model comprising a gap-aware transformer-encoder configured is used to determine a sequence correction utilizing an alignment-based loss function. In some embodiments, the model is trained or produces an inference using a series of subreads combined with a consensus read. The combination is divided into partitions and transformed into a tensor used as input to the model. The tensor provides pulse widths and interpulse duration. Additionally, signal-to- noise and strand information may be used for each nucleotide. A loss function can be utilized to consider the alignment between labeled positions and subreads. Deep learning techniques are described, for example, in U.S. Patent Application Publication No. 2023 / 0298701 Al , which is incorporated by reference in its entirety'.

[0060] In some embodiments, third tier clusters (partial length read clusters not matched with full length read clusters) are ignored / discarded at block 455 from consensus consideration. This may reduce chances of an erroneous consensus and / or collision, for example. In some embodiments, a partial length read may be considered as a full length read. For example, a partial length read may be considered as a full length read if it is determined to have a minimum number or cross section of base reads. Calculating a consensus read may then proceed based on the full length read assumption and matching the reclassified partial length read with other partial length reads to calculate a consensus. In some embodiments, a consensus / quality score and threshold such as identified above is used to accept or discard a consensus determination from only partial length reads.

[0061] At block 435, based on the consensus calculations made above, one or more consensus sequences are identified. The sequences may represent nucleic acid molecules (or molecules derived therefrom) identified from a particular biological sample, for example, and / or sequences from other ty pes of molecules. The sequences may be generated in a report such as within a user interface or file format (e.g., BAM, FASTA file format).

[0062] FIG. 5 illustrates separate methods of merging clusters utilizing a tiering system according to some embodiments herein. A first exemplary set 510 of full length reads of a certain size (e.g., 3 or more reads in the cluster) may generally provide sufficient data where partial length reads are not used to contribute to a respective sequence family or consensus sequence. The first set 510 may be part of a first tier (Tier 1) of clusters. A second exemplary' set 520 of merged mixed full length and partial length reads has only two full length reads. Additional partial length reads provide information that is used to generate a consensus sequence 500, where the small number of full length reads might otherwise have had insufficient information to generate a consensus call. A third exemplary set 530 of merged mixed full length and partial length reads has two full length reads where additional partial length reads increase a sequence family size. The second and third sets 520 & 530 may be part of a second tier (Tier 2) of clusters, which combine full length and partial length reads. A fourth exemplary set 540 of merged mixed full length and partial length reads has only one full length read while the partial length reads provide sufficient information to complete a portion of a family for determining a consensus sequence. In this case, at leas t a portion of the consensus sequence 510 for the fourth set 540 may be included based on the base calls in the single full length read, but also associated with a low' quality score to indicate that the consensus calls were not made based on two or more reads.

[0063] FIG, 6A is a table of clustering results of reads obtained from a sample, in accordance with some embodiments. The sample contained cell-free DNA (cfDNA) and the sequencing was performed using a nanopore sequencer utilizing Xpandomer synthesis (sometimes referred to as Sequencing By expansion or SBX). The reads were sorted into clusters of full length reads and partial length reads such as further described herein. Tiering of the clusters (i.e., into the three tiers described herein) was also performed based on matching full length read clusters with partial length read clusters where the merging of full length read clusters may occur with sizes of a) one (singleton) read, and b) two or more reads. An analysis was performed in relation to determining consensus calculations based on the separate criteria for full length read only clusters and matching full length read clusters with partial length read clusters. Although collision rates were higher when permitting partial length read matches with singleton full length read clusters, the overall number of discarded reads was significantly lower.

[0064] FIG. 6B is a chart representing the separate clustering results of FIG. 6 A based on minimum read count for matching full length read clusters. Using the looser criteria for matching full length read clusters (two or more reads between clusters), a greater overall number of clusters are utilized for generating consensus reads, including the total number of full length read clusters exclusively matched with other full length read clusters (first tier) and matched with partial length read clusters (second tier).

[0065] FIG. 7 illustrates an example computer system 700 that may be utilized to implement techniques disclosed herein. Any of the computer systems mentioned herein, such as for hosting the systems and implementing the processes described for calculating consensus sequences, may utilize any suitable number of subsystems. Examples of such subsystems are shown in FIG. 7 in computer system 700. In some embodiments, a computer system 700 includes a single computer apparatus, where the subsystems can be the components of the computer apparatus. In other embodiments, a computer system 700 can include multiple computer apparatuses, each being a subsystem, with internal components. A computer system 700 can include desktop and laptop computers, tablets, mobile phones, telecommunication devices or other mobile devices. The computer system 700 may include one or more processors, including, for example, a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), or the like. In some embodiments, a cloud infrastructure (e.g., Amazon Web Services) can be used to implement the disclosed techniques.

[0066] The subsystems shown in FIG. 7 are interconnected via a system bus 775.Additional subsystems such as a printer 774, keyboard 778, storage device(s) 779, monitor 776, which is coupled to display adapter 782, and others are shown. Peripherals and input / output (170) devices, which couple to I / O controller 771, can be connected to the computer system 700 by any number of means known in the art such as input / output (I / O) port 777 (e.g., USB, FireWire®). For example, I / O port 777 or external interface 781 (e.g. Ethernet, Wi-Fi, etc. ) can be used to connect computer system 700 to a wide area network such as the Internet, a mouse input device, or a scanner.

[0067] The interconnection via system bus 775 allows the central processor 773 to communicate with each subsystem and to control the execution of a plurality of instructions from system memory 772 or the storage device(s) 779 (e.g., a fixed disk, such as a hard drive.or optical disk), as well as the exchange of information between subsystems. The system memory 772 anchor the storage device(s) 779 may embody a computer readable medium. Another subsystem is a data collection device 785, such as a camera, microphone, accel erometer, and the like. Any of the data mentioned herein can be output from one component to another component and can be output to the user.

[0068] A sequencer instrument 790 (e.g., a nanopore sequencer) is connected through external interface 781 for providing sequencing data to a data collection device 785 and / or storage devices 779, In another embodiment, the computer system 700 may be included as a subsystem of the sequencer instrument 790.

[0069] A computer system 700 can include a plurality’ of the same components or subsystems, e.g., connected together by external interface 781 or by an internal interface (e.g., a system bus). In some embodiments, computer systems, subsystems, or apparatuses can communicate over a network. In such instances, one computer can be considered a client and another computer a server, where each can be part of the same computer system. A client and a server can each include multiple systems, subsystems, or components.

[0070] Aspects of embodiments can be implemented in the form of control logic using hardware (e.g. an application specific integrated circuit or field programmable gate array) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein, a processor includes a single-core processor, multi-core processor on a same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, a person of ordinary’ skill in the art will know and appreciate other ways and / or methods to implement embodiments of the present invention using hardware and a combination of hardware and software.

[0071] Machine learning models utilized herein may include one or more of a Naive Bayes (NB) model, a logistic regression (LR) model, a random forest (RF) model, a support vector machine (SVM) model, an artificial neural network model, a multilayer perceptron (MLP) model, a convolutional neural network (CNN), a Large Language model (LLM), and / or other machine learning or deep learning models, etc. The machine learning models can be updated / trained using a supervised learning technique, an unsupervised learning technique, etc. as known to those of ordinary’ skill in the art.

[0072] Any of the software components or functions described in this applicati on may be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C, C++, C#, Objective-C, Swift, or scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code may be stored as a senes of instructions or commands on a computer readable medium for storage and / or transmission. A suitable non-transitory computer readable medium can include random access memory (RAM), a read only memory' (ROM), a magnetic medium such as a hard-driv e or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk), flash memory, and the like. The computer readable medium may be any combination of such storage or transmission devices.

[0073] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium may be created using a data signal encoded with such programs. Computer readable media encoded with the program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer readable medium may reside on or within a single computer product (e.g. a hard drive, a CD, or an entire computer system), and may be present on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results mentioned herein to a user.

[0074] Although presented as numbered steps, steps of methods or process workflows presented herein can be performed at the same time or in a different order, unless clearly contradicted by context. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of a step may be optional. Additionally, any of the steps of any of the methods can be performed with modules, units, circuits, or other means for performing these steps.

[0075] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of embodiments of the invention. However, other embodiments of the invention may be directed to specific embodiments relating to each individual aspect, or specific combinations of these individual aspects.

[0076] The above description of example embodiments of the invention has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form described, and many modifications and variations are possible in ligh t of the teaching above.

Claims

CLAIMS1. A computer-implemented method for identifying a consensus molecular sequence, the method comprising: obtaining sequencing data from a physical sample comprising a plurality of full length reads and partial length reads, wherein each read in the sequencing data comprises at least one 5’ and / or 3’ binding sites; sorting the sequencing data into clusters of full length reads and partial length reads, wherein the clusters are based on position information of the 5’ and / or 3’ binding sites included in the reads; comparing a partial length read cluster to a full length read cluster to determine a positional distance between the compared clusters; in response to determining that the positional distance between the compared clusters fails within a predetermined threshold, merging the compared partial length read cluster and full length read cluster; and identifying and reporting a consensus molecular sequence of the sample based on the merged partial length read cluster and full length read cluster.

2. The method of claim 1 , wherein merging the compared partial length read cluster and full length read cluster is further in response to determining that a count of reads in the full length read cluster is at least two.

3. Tire method of claim 1, wherein the merged full length read cluster has fewer than two reads.

4. The method of claim 1 , wherein merging the compared partial length read cluster and full length read cluster is further in response to determining that a count of reads in the full length read cluster with a positional distance from the respective partial length read cluster is not more than one.

5. 'The method of claim 1, wherein merging the compared partial length read cluster and full length read cluster is further in response to determining that none of the reads in thecompared partial length read cluster has a length greater than a length of the read(s) of the compared full length read cluster.

6. The method of claim 1 , further comprising merging two full length read clusters having two or fewer reads and a positional distance within a predetermined threshold with each other.

7. The method of claim 1, wherein sorting the sequence data into clusters comprises sorting and aligning the reads in relation to bases and sequence positions prior to the sorting into full length reads and partial length reads.

8. The method of claim 1, wherein merging the compared partial length read cluster and full length read cluster is further in response to determining that a total number of the reads in the partial length read cluster comprises thirty percent or more of a total count of reads in the combined full length read cluster and partial length read cluster.

9. The method of claim 1 , wherein the sequencing data from the physical sample is based on a template-directed synthesis for targeting a nucleic acid structured with a segment that produces a synchronization signal upon passage through a nanopore.

10. A computer-implemented method for identifying a consensus molecular sequence, the method comprising: obtaining sequencing data for a physical sample comprising a plurality' of full length reads and partial length reads, wherein each read in the sequencing data comprises at least one 5’ and / or 3’ binding sites; sorting the sequencing data into clusters of full length and partial length reads, wherein the clusters are based on position information of the 5’ and / or 3' binding sites included in the read; sorting the clusters into tiers of clusters, a first tier comprising only full length reads, a second tier comprising both full length reads and partial length reads, and a third tier comprising only partial length reads; discarding the third tier of clusters;determining a consensus sequence among each of the first tier clusters; determining a consensus sequence among each of the second tier clusters by calculating a likelihood of partial length reads or full length reads within the second tier clusters of supporting the consensus sequence and determining that the calculated likelihood exceeds a particular threshold; and identifying and reporting one or more consensus molecular sequences of the physical sample based on the determined consensus sequences for the first tier and / or second tier of clusters.11 , The method of claim 10, wherein determining the consensus sequence within the respective first tier cluster comprises determining a sequence with a greatest number of occurrences included in the reads of the first tier cluster.

12. The method of claim 11, wherein the calculating the likelihood of partial length reads or full length reads supporting the consensus sequence is based on a modeling of a probability of error in a candidate consensus sequence to follow a binomial distribution with a probability number representing a base coverage and a count of reads within the respective cluster supporting a position of the base of the candidate consensus sequence.

13. The method of claim 12, wherein the calculating the likelihood of partial length reads or full length reads supporting the consensus sequence is based on a deep machine learning inference model comprising a gap-aware transformer-encoder configured to determine a sequence correction utilizing an alignment-based loss function.

14. A system for identifying a consensus molecular sequence, the system comprising: one or more processors associated with a sequencer instrument, the one or more processors configured to: obtain, from the sequencer instrument, sequencing data from a physical sample comprising a plurality of full length reads and partial length reads, wherein each read in the sequencing data comprises at least one 5’ and / or 3’ binding sites;sort the sequencing data into clusters of full length reads and partial length reads, wherein the clusters are based on position information of the 5’ and / or 3’ binding sites included in the reads; compare a partial length read cluster to a full length read cluster to determine a positional distance between the compared clusters; in response to determining that the positional distance between the compared clusters falls wdthin a predetermined threshold, merge the compared partial length read cluster and full length read clusters; and identify and report a consensus molecular sequence of the sample based on the merged partial length read cluster and full length read cluster.

15. The system of claim 14, wherein the sequencer instrument is in communication with the one or more processors via a network.

16. The system of claim 15, wherein the sequencer instrument is configured to generate the sequencing data from a template-directed synthesis for targeting a nucleic acid structured with a segment that produces a synchronization signal upon passage through a nanopore.

17. The system of claim 14, wherein merging the compared partial length read cluster and full length read cluster is further in response to determining that a count of reads in the full length read cluster is at least two.

18. The system of claim 17, where the one or more processors are further configured to merge a candidate full length read cluster having fewer than two reads with the merged full length read cluster and partial length read cluster in response to determining that the candidate full length read cluster has a positional distance within a predetermined threshold from respective partial length reads of the merged full length read cluster and partial length read cluster.

19. The system of claim 14, wherein merging the compared partial length read cluster and full length read cluster is further in response to determining that a count of full length reads with a positional distance from the respective partial length read cluster is not more than one.

20. The system of claim 14, wherein merging the compared partial length read cluster and full length read cluster is further in response to determining that none of the reads in the compared partial length read cluster has a length greater than a minimum read length of the compared full length read cluster.

Citation Information

Patent Citations

  • Deep-learning-based techniques for generating a consensus sequence from multiple noisy sequences

    US20230298701A1

Cited By

  • Systems and methods for identifying somatic structural variants

    WO2026044179A1