Machine learning methods and systems for detecting structural variants

By aligning sequence reads to a reference genome and using machine learning models to analyze colocation plots, the method addresses the challenge of mapping fragmented sequences, enhancing the detection of structural variants in genomic sequences with improved accuracy and precision.

WO2026080196A1PCT designated stage Publication Date: 2026-04-16ILLUMINA INC
View PDF 36 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/046615
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-01
Filing Date
2025-09-16
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Traditional nucleic acid sequencing methods, including next-generation sequencing by synthesis (SBS), struggle to accurately map sequence fragments back to their original positions in the genome due to the loss of connectivity and proximity information during fragmentation, making it difficult to detect structural variants in genomic sequences.

Method used

The method involves obtaining flow cell data with nucleic acid sequence reads and flow cell locations, aligning these reads to a reference genome, generating colocation data based on linking information, and processing this data with a machine learning model to identify structural variants, using techniques such as deep learning and neural networks to analyze colocation plots for visual signatures of structural variants.

Benefits of technology

This approach enhances the detection of structural variants, particularly in repetitive and hard-to-map genomic regions, achieving breakpoint precision and improving the accuracy of identifying complex structural variants like F8 inversions and ring chromosomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025046615_16042026_PF_FP_ABST
    Figure US2025046615_16042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and systems for detecting a structural variant from a DNA sample placed on a flow cell. In some embodiments, the methods and systems obtain nucleic acid sequence reads and flow cell locations of sequence read clusters on the flow cell; align the sequence reads to one or more reference genomes to obtain a genomic location of the sequence reads; obtain linking information between pairs of sequence reads on the flow cell based on the flow cell location and genomic location of each sequence read in the pairs of sequence reads; generate colocation data based on the linking information; provide the colocation data to a machine learning model; and process the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acids and determine a confidence score. Further disclosed are methods and systems for training a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

ILLINC.851WO / IP-2870-PCT PATENTMACHINE LEARNING METHODS AND SYSTEMS FOR DETECTINGSTRUCTURAL VARIANTSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 705,439, filed October 9, 2024, and U.S. Provisional Application No. 63 / 715,462, filed November 1, 2024, the content of each of which is incorporated by reference in its entirety.BACKGROUNDField

[0002] The present disclosure relates to DNA sequencing systems and methods. In particular, this disclosure relates to systems and methods for detecting structural variants in DNA sequences using machine learning methods and related systems.Description

[0003] Structural variants (SVs) of a genome are large variants in the genome with respect to a known reference genome. For example, a structural variant may be an insertion of hundreds of nucleotides into a new position of the genome as compared to the reference genome. One widely-used threshold of variant size for classifying variants as “small” or “large” has been 50 base pairs (bp), where variations of less than 50bp are considered small and variations of more than 50bp are considered as large. However, some methods may use 35 bp as the threshold to denote the difference between a small or large variant. SVs include DNA losses, gains, and rearrangements relative to the reference genome. Simple SVs are typically classified as deletions (DNA loss); insertions and duplications (DNA gain); inversions and translocations (DNA rearrangements). Simple SVs can combine to cause more complex SVs at a genomic locus. For example, one genomic region may have a segment of DNA deleted, and in addition, a new segment of DNA is inserted at the same position so that the single position is missing nucleotides, and also includes new nucleotides as compared to a reference genome. The importance of SVs has already been well established due to their role in gene regulation, various diseases, and ethnic diversity. Others have commented on the importance of SVs and their applications in medicine and molecular biology. Mahmoud, M. etal. (2019) ‘Structural variant calling: the long and the short of it’, Genome Biology, 20(1), p. 246. Available at: doi.org / 10.1086 / sl3059-019-1828-7.

[0004] Traditional nucleic acid sequencing methods, and several types of nextgeneration sequencing methods including Sequencing by Synthesis (SBS), use a shotgun approach to sequence large genomic DNA fragments, sometimes called template genomic sequences. During SBS sequencing the template genomic sequences are first fragmented into smaller pieces that are amenable to next-generation sequencing methods on a flow cell. One of the difficulties of this approach is that by the time the smaller sequence fragments from the template genomic sequences have been sequenced, knowledge of their original position in the genome, and their connectivity and proximity to each other in the original template genomic sequence is lost.SUMMARY

[0005] The methods and systems disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, some prominent features will now be discussed briefly. Numerous other embodiments are also contemplated, including embodiments that have fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled “Detailed Description”, one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.

[0006] Disclosed herein are methods and systems for detecting a structural variant in DNA from a sample placed on a flow cell. In some embodiments, the methods and systems obtain flow cell data from the DNA, wherein the flow cell data comprises nucleic acid sequence reads and flow cell locations of sequence read clusters on the flow cell. The methods and systems can then align the sequence reads to one or more reference genomes to obtain a genomic location of the sequence reads. The methods and systems can obtain linking information between pairs of sequence reads on the flow cell based on the flow cell location and the genomic location of each sequence read in the pairs of sequence reads. The methods and systems can generate colocation data based on the linking information. As used herein,“colocation data” refers to data which represents or includes a set of linking information between sequence reads. The methods and systems can provide the colocation data to a machine learning model. Then, the methods and systems can process the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acids and determine a confidence score, thereby detecting at least one structural variant for the sample.

[0007] Further disclosed herein are methods and systems for training a machine learning model to detect structural variants in a genomic sequence. In some embodiments, the methods and systems obtain a plurality of training datasets, wherein each training dataset comprises i) a structural variant classification and ii) linking information between pairs of sequence reads based on their geographic location on the flow cell. The methods and systems can generate colocation data for each training dataset based on the linking information. The methods and systems can provide each training dataset to a machine learning model. Then, the methods and systems can train the machine learning model to detect structural variants based on each colocation plot data and its corresponding structural variant classification.

[0008] According to an aspect, disclosed herein are methods for detecting a structural variant in DNA from a sample placed on a flow cell. In some embodiments, the method includes obtaining flow cell data from the DNA, wherein the flow cell data comprises nucleic acid sequence reads and flow cell locations of sequence read clusters on the flow cell; aligning the sequence reads to one or more reference genomes to obtain a genomic location of the sequence reads; obtaining linking information between pairs of sequence reads on the flow cell based on the flow cell location and genomic location of each sequence read in the pairs of sequence reads; generating colocation data based on the linking information; providing the colocation data to a machine learning model; and processing the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acids and determining a confidence score, thereby detecting at least one structural variant for the sample.

[0009] In some embodiments, generating colocation data based on the linking information comprises: dividing the alignment of sequence reads to the reference genome into a plurality of bins; and counting an estimated number of links between pairs of sequence reads by bin. In some embodiments, the bins are between 500 bp to 100 kbp in size.

[0010] In some embodiments, the colocation data comprises a colocation plot image or a graph representation of colocation data. In some embodiments, the machine learning model comprises an object detection model, image classification model, an anomaly detection model, a deep learning image processing model, a deep convolutional neural network (CNN), graph neural network (GNN), or a transformer-based model. In some embodiments, the method comprises applying non-max suppression to a colocation plot image. In some embodiments, the method comprises applying non-max suppression across a plurality of combined colocation plot images.

[0011] In some embodiments, processing the image comprises identifying a putative structural variant region based on a first colocation data; generating a second colocation data based on the putative structural variant, wherein the second colocation data is at a higher or lower resolution than the first colocation data; providing the second colocation data to the machine learning model; and processing the second colocation data with the machine learning model to identify a structural variant within the genomic nucleic acids. In some embodiments, the second colocation data is based on a larger or smaller bin size as compared to the first colocation data. In some embodiments, the first colocation data is based on bins of 10 kbp to 100 kbp in size, and wherein the second colocation data is based on bins of 500 bp to 10 kbp in size. In some embodiments, the first colocation data is based on bins of 500 bp to 10 kbp in size, and wherein the second colocation data is based on bins of 10 kbp to 100 kbp in size. In some embodiments, the method comprises displaying the putative structural variant region with a user interface, and wherein the putative structural variant region is selected via the user interface.

[0012] In some embodiments, the method comprises flowing nucleic acid molecules greater than 500 bp in length across a flow cell and fragmenting the nucleic acid molecules on the flow cell. In some embodiments, the method comprises providing a sequence read depth, a mapping quality measurement, GC content bias measurement, an initial CNV / structural variant prediction, abnormal paired read information, a link quality score, a genomic distance from alignment, split read information, phasing information, linking information statistical parameters, or a comparison to one or more normal samples, to the machine learning model.

[0013] In some embodiments, identifying a predicted structural variant region within the genomic nucleic acids comprises determining one or more of: an area with a suspected SV, a structural variant event class, a structural variant position, a structural variant anomaly region corresponding to a complex or overlapping structural variant, and a structural variant size. In some embodiments, the structural variant event class comprises an insertion, a deletion, an inversion, or a translocation.

[0014] In some embodiments, the method includes storing structural variant information in an electronic file. In some embodiments, the method comprises storing a set of detected structural variants in a VCF file and a set of anomaly regions corresponding to complex or overlapping structural variants in a BED file. In some embodiments, the method further comprises displaying, with a graphical user interface, a first colocation plot showing a predicted structural variant regions or a bounding box for a predicted structural variant region. In some embodiments, displaying the first colocation plot comprises displaying graphical indicia indicating a position of the predicted structural variant region or anomalous region on the display. In some embodiments, wherein the graphical indicia include a highlight or change in color. In some embodiments, the method further comprises, in response to receiving a user selection, displaying a magnified view of the predicted structural variant region. In some embodiments, the method further comprises updating a mapping location of a sequence read based on linking information.

[0015] According to another aspect, disclosed herein are methods for training a machine learning model to detect structural variants in a genomic sequence. In some embodiments, the method includes obtaining a plurality of training datasets, wherein each training dataset comprises i) a structural variant classification and ii) linking information between pairs of sequence reads based on their geographic location on the flow cell; generating colocation data for each training dataset based on the linking information; providing each training dataset to a machine learning model; and training the machine learning model to detect structural variants based on each colocation plot data and its corresponding structural variant classification.

[0016] In some embodiments, the method comprises generating modified genomic sequence flow cell data by introducing simulated structural variants within reference genome data and generating simulated sequence reads. In some embodiments, the method comprisesgenerating modified genomic sequence flow cell data by introducing synthetic structural variants into a diploid genome sample, and generating simulated sequence reads from the diploid genome sample. In some embodiments, the method comprises providing a sequence read depth, a mapping quality measurement, GC content bias measurement, an initial CNV / structural variant prediction, abnormal paired read information, a link quality score, a genomic distance from alignment, split read information, phasing information, linking information statistical parameters, or a comparison to one or more normal samples, to the machine learning model.

[0017] According to a further aspect, disclosed herein are systems for classifying a structural variant condition for a subject. In some embodiments, the system includes one or more processors that are programmed to execute a method comprising: obtaining flow cell data from sample DNA, wherein the flow cell data comprises a nucleotide sequence read, and flow cell location of sequence read clusters on the flow cell; aligning sequence reads to one or more reference genomes to obtain the genomic location of each sequence read using the flow cell location of the sequence reads, thereby obtaining linking information between pairs of sequence reads based on flow cell location and genomic location; generating colocation data based on the linking information; providing the colocation data to a machine learning model; and processing the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acids and determining a confidence score, thereby detecting at least one structural variant for the sample.

[0018] According to a further aspect, disclosed herein are systems for training a machine learning model to detect structural variants in a genomic sequence. In some embodiments, the system includes one or more processors that are programmed to execute a method comprising: obtaining a plurality of training datasets, wherein each training dataset comprises i) structural variant classification and ii) linking information between pairs of sequence reads based on their geographic location on the flow cell; generating colocation data for each training dataset based on the linking information; providing each colocation data and its corresponding structural variant condition to a machine learning model; and training the machine learning model to detect structural variants based on each colocation plot data and its corresponding structural variant classification.

[0019] According to a further aspect, disclosed herein are systems for detecting a structural variant. In some embodiments, the system includes a user interface configured to display one or more first colocation plots based on linking information from genomic nucleic acids, and a detection refinement option; and a processor configured to, in response to receiving a selection of the detection refinement option, generate one or more second colocation plots wherein the one or more second colocation plot images are at a different magnification or resolution than the one or more first colocation plot images. In some embodiments, the processor is further configured to provide the one or more second colocation plot images to a trained machine learning model and process the one or more second colocation plot images with the machine learning model to identify a structural variant within the genomic nucleic acids.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. In addition to the features described herein, additional features and variations will be readily apparent from the following descriptions of the drawings and exemplary embodiments. It is to be understood that these drawings depict typical embodiments, and are not intended to be limiting in scope.

[0021] FIG. 1 A is a block diagram of an exemplary sequencing system that may be used to perform the disclosed methods.

[0022] FIG. IB is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 1A.

[0023] FIG. 2 is a flow diagram that schematically illustrates an exemplary method for detecting a structural variant in DNA from a sample placed on a flow cell.

[0024] FIG. 3 is a flow diagram that schematically illustrates an exemplary method for training a machine learning model to detect structural variants in a genomic sequence.

[0025] FIG. 4 depicts an exemplary colocation plot for control sample HG002.

[0026] FIG. 5 depicts a colocation plot from a sample with an F8 gene inversion in intron 22.

[0027] FIG. 6A schematically depicts the flow of an object detection network.

[0028] FIG. 6B depicts an example of object detection with a machine learning model where there are overlapping boundary boxes.

[0029] FIG. 7A depicts a colocation plot used in the visualizations of FIGS. 7B- 7E.

[0030] FIG. 7B depicts a visualization of a colocation plot with a chromosome overlaid on the diagonal of the colocation plot.

[0031] FIG. 7C depicts a further visualization of a colocation plot in a graphical user interface.

[0032] FIG. 7D depicts a further visualization of a colocation plot in a graphical user interface.

[0033] FIG. 7E depicts a further visualization of a colocation plot in a graphical user interface.

[0034] FIG. 8A depicts an exemplary colocation plot with a homozygous insertion event.

[0035] FIG. 8B depicts an exemplary colocation plot with a homozygous deletion event.

[0036] FIG. 9 depicts an exemplary colocation plot and a close-up view of features of the colocation plot.

[0037] FIG. 10 has panels a) - h) depicting exemplary detections of a structural variant with a machine learning model.

[0038] FIG. 11 has panels a) - d) depicting exemplary detections of a structural variant with a machine learning model.DETAILED DESCRIPTION

[0039] The foregoing and other aspects of the present disclosure will now be described in more detail with respect to the description and methodologies provided herein. This description is not intended to be a detailed catalogue of all the ways in which the embodiments of the present disclosure may be implemented, or of all the features that may beadded to the present disclosure. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein, which do not depart from the instant disclosure, will be apparent to those skilled in the art in light of the instant detailed description, figures and claims. Hence, the following specification is intended to illustrate some particular embodiments, and not to exhaustively specify all permutations, combinations and variations thereof.

[0040] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.

[0041] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.Overview

[0042] Whole genome sequencing (WGS) technologies using short sequence reads have been widely adopted. However, comprehensively characterizing the entire structural profile of individual genomes has been difficult due to challenges with characterizing genomic regions which include long repetitive regions and segmental duplications. As used herein,“segmental duplications” refer to duplications of DNA sequences in the genome where the duplicated region is generally greater than 1000 base pairs in length. For example, segmental duplications may include duplicated regions of 1 kbp to 400 kbp. In many cases, the duplicated regions may be identical to one another. In other cases, the DNA sequences that are repeated in the genome share a high level, e.g., more than 90%, 95%, or 98% of sequence identity with one another. In these regions SBS methods may not generate sequences which are long enough to include a single sequence read which covers the repetitive region or segmental duplication such that the sequence read can be correctly mapped to its correct position. For example, a relatively short sequence read may be located within a segmental duplication while the reference sequence may map the sequence read to a first position. However, its actual location in the template genomic sequence being analyzed may be found in a wholly different second position of the genome due to the duplication of that sequence in the template genomic sequence.

[0043] To improve the accuracy of mapping sequence reads to their correct position on the template genomic sequence, embodiments herein relate to linking sequence reads together based on the location of their respective clusters on a flow cell and the genomic location of the sequence reads when mapped to a reference genome. Various sequencing library preparation methods and sequencing techniques may be used to capture flow cell and genomic proximity information. For example, in some embodiments, sequence reads are generated from fragments of a genomic DNA sample which are bound to the flow cell. For example, relatively long (e.g., greater than 500 bp) DNA fragments may be flowed across a sequencing flow cell and may attach to transposome complexes that are embedded on the surface of the flow cell. The transposome complexes can fragment the long DNA fragments into shorter fragments (e.g., of less than 500 bp), and attach sequencing adapters. Each shorter fragment becomes bound to the flow cell and can be amplified to form clusters. The sequence of each shorter fragment can be determined by performing SBS on the clusters of fragments. During this process, the physical location of each cluster on the flow cell can be recorded in addition to sequence reads. The “link” or “linkage information” as discussed herein refers, in some embodiments, to the probability that two pairs of sequence reads on a sequencing flow cell are derived from the same original nucleic acid molecule. Sequences which are nearer to each other on the flow cell were found to be more likely to be derived from the same originalnucleic acid molecule, and so linkage measurements can be determined to provide a probability that two sequence reads were derived from the same original nucleic acid molecule based on their flow cell location and position on the genome.

[0044] Such on-flowcell library preparation methods supplement standard sequencing reads with long-range flow cell proximity information, which aids in detection of rearrangements in repetitive and hard-to-map regions that have previously been challenging for standard WGS. On-flow cell proximity technology physically associates reads that originated from the same input DNA molecule, providing long range information between two distant regions of the genome.

[0045] In some embodiments, there is an expected relationship between the proximity between clusters on a flow cell and the genomic distance between corresponding sequence reads after aligning to a reference genome. For example, clusters which are proximate to one another on the flow cell are expected to correspond to sequence reads which are genomically proximate to one another, and vice-versa.

[0046] However, structural variants can perturb the relationship between genomic proximity and flow cell proximity. For example, if a genomic DNA sample has a structural variant as compared to the reference genome, clusters may be in close proximity to one another on the flow cell because they are derived from the same original long DNA fragment, but when the sequence reads are mapped to the reference genome, their genomic distance is greater or less than would be expected based on the flow cell proximity / linking information because they are being mapped closer or farther apart for the sample than they are on the reference genome, due to the structural variant in the genomic DNA sample. Furthermore, linkage information can be used to determine a primary mapping location for sequence reads that have a plurality of candidate mapping locations, for example due to a structural variant or due to sequence homology between two locations.

[0047] Thus, perturbations to the relationship between genomic proximity and flow cell proximity can be used to identify structural variants. For example, the number of links between sequence reads mapped to one area of the genome and sequence reads mapped to another area of the genome can be counted. It is expected that sequence reads coming from clusters located near each other on a flow cell will be linked to each other and found genomically close to each other in an alignment to a reference genome. An absence ofexpected links in a region of the reference genome can suggest a deletion event, which would reduce the links found on the flowcell between the sequence reads. The presence of ample links across breakpoints of the putative deletion event can help confirm there was a deletion event, because the breakpoints would have been physically proximate to one another in the sample genomic DNA and thus on the flow cell.

[0048] In some embodiments, linkage information between sequence reads is visualized using a colocation plot, which shows counts of the number of links between sequence reads mapped across regions of the reference genome. As used herein, a “colocation plot” refers to a graphical representation that visualizes the number of links between sequence reads mapped to different regions along the genome. In some embodiments, the colocation plot has x and y coordinates that represent spans along an alignment of the sequence reads to a reference genome. In some embodiments, the number of links between sequence reads in one area of the reference genome and sequence reads in another area of the reference genome can be illustrated on the colocation plot. For example, in some embodiments, the number of links is represented using visual features such as color intensity on the colocation plot.

[0049] In some embodiments, because neighboring regions of the genome are expected to have high counts of links between each other, a colocation plot is expected to show a diagonal line feature representing links between neighboring regions of a reference genome. In other colocation plots with different shapes, the feature representing links between neighboring regions may have a different shape. In some embodiments, a colocation plot visually depicts the unexpected presence or absence of links between two areas of the genome as distortions in the diagonal line feature. In some embodiments, structural variants leave a unique visual signature on a colocation plot image that depends on the structural variant type (for example, insertions, deletions, inversion, etc.) and the size of the feature or variant, which can be interpreted visually.

[0050] In one embodiment, machine learning systems and processes are used to train a neural network to recognize the visual signatures found on colocation plots. For example, in one embodiment, specific structural variants are introduced via simulation into a reference genome such as the HG002 complete diploid human genome sample. Relatively large DNA fragments are isolated from each haplotype of the genomic sample and placed onto a flow cell where the large fragments are cleaved into smaller fragments by transposomes. Thesmaller fragments are used to create clusters on the flow cell and a set of links can be determined between the smaller fragments on the flow cell based on their physical proximity to each other on the flow cell and genomic distance on the HG002 genome. A colocation map can be generated based on the linked sequence reads, and the colocation map may be used to train the neural network to recognize the specific visual patterns caused by the specific structural variants introduced by simulation into the HG002 genome. This process can be repeated using a variety of different specific structural variants until the neural network is trained to recognize the visualizations within a colocation map corresponding to each specific structural variant. The machine learning model can also be trained based on synthetic link count data, which does not represent a real sample. For example, synthetic SVs can be introduced to a reference sample by adjusting link counts, to create colocation plots with known signatures. The simulated colocation plots can then be used for training the machine learning model. The machine learning model can also be trained by simulating reads from a simulated reference genome, and training the machine learning model on those simulated reads.

[0051] Newly generated colocation plots can then be run through the trained system to process linkage information and detect structural variants based on the detected visual signatures in the colocation plots. For example, machine learning (ML)-based image analysis methods and systems can be trained to detect and label deletions, insertions, duplications, translocations, inversions and complex SVs based on the detectable visual signature that each of these types of structural variants create on a colocation plot. Breakpoint precision down to base pair resolution can be achieved through short read analysis. For example, in some embodiments, structural variant events are detected from the colocation plot, but breakpoint precision and sequence (e.g. for inserted bases) can be resolved by analysis of short reads after analyzing their linking information to map them precisely to a particular SV.

[0052] Colocation plots have visual features that can be associated with a variety of structural variants. However, the realities of “noise” and heterozygous structural variants can complicate this association. As discussed herein, using machine learning models which are trained to recognize these features associated with structural variants may overcome these technical problems. Converting linkage information to a visual format can enable the use of image-based deep networks to detect structural variants. For example, a machine learningnetwork can analyze images and associate each image with a structural variant classification. In alternative embodiments, machine learning networks can analyze linkage information directly and associate the linkage information for a sample with a structural variant classification.

[0053] The ability to accurately call large rearrangements, including clinically important F8 inversions and ring chromosomes in real samples, has been demonstrated. These balanced events in regions of the genome that are usually difficult to map and overlapping segmental duplications have traditionally been the most difficult to call accurately with WGS and other sequencing technologies. As used herein, a “balanced event” refers to a structural variant wherein there is the same amount of genomic material, but the order has been rearranged. In some embodiments, balanced events are more difficult to detect than events where the amount of genomic material has changed (e.g. deletions, where sequences have been removed so reads do not map to those locations, or duplications where sequences have been added which generates extra reads). In some embodiments, a balanced event does not change the number of reads, and detection is based on accurate mapping in the breakpoint regions. In regions that are repetitive / long, it can be difficult to map reads unambiguously to their correct location, therefore the breakpoints can be difficult to characterize with confidence. The novel methods and systems described herein can visualize and highlight complex and difficult-to- detect structural variation, providing new insights into genomic structure.Definitions

[0054] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0055] Although the following terms are believed to be well understood by one of skill in the art, the following definitions are set forth to facilitate understanding of the presently disclosed subject matter.

[0056] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art.

[0057] As used herein, the terms “a” or “an” or “the” may refer to one or more than one. For example, “a” marker can mean one marker or a plurality of markers.

[0058] As used herein, the term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).

[0059] Throughout this specification, unless the context requires otherwise, the words “comprise,” “comprises,” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.

[0060] As used herein, the term “consists essentially of’ (and grammatical variants thereof), as applied to the compositions and methods of the present disclosure, means that the compositions / methods may contain additional components so long as the additional components do not materially alter the composition / method.

[0061] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, diTP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.

[0062] As used herein, the term "fragment," when used in reference to a first nucleic acid, is intended to mean a second nucleic acid having a part or portion of the sequence of the first nucleic acid. Generally, the fragment and the first nucleic acid are separate molecules. The fragment can be derived, for example, by physical removal from the larger nucleic acid, by replication or amplification of a region of the larger nucleic acid, by degradation of other portions of the larger nucleic acid, a combination thereof or the like. Theterm can be used analogously to describe sequence data or other representations of nucleic acids. As used herein, the term "haplotype" refers to a set of alleles at more than one locus inherited by an individual from one of its parents. A haplotype can include two or more loci from all or part of a chromosome. Alleles include, for example, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), gene sequences, chromosomal insertions, chromosomal deletions etc. The term "phased alleles" refers to the distribution of the particular alleles from a particular chromosome, or portion thereof. Accordingly, the "phase" of two alleles can refer to a characterization or representation of the relative location of two or more alleles on one or more chromosomes.

[0063] “Fragmentation” as described herein refers to the shearing or fragmenting of nucleic acid into shorter lengths. Fragmentation methods include enzymatic (including fragmentase), physical (including sonication, nebulization, needle shearing, microwave, etc.), and chemical (including depurination, hydrolysis, oxidation, etc.). The term “fragmentase” as used herein refers to enzymes that fragment nucleic acids. A fragmentase can be a single enzyme or two or more enzymes that work together to fragment the nucleic acid. Some fragmentases work on single-stranded nucleic acids whereas other fragmentases work on double-stranded nucleic acids. Yet other fragmentases work on one strand of a doublestranded nucleic acid. Fragmentases can cut randomly or they can cut specifically at a particular sequence of nucletotides. Non-limiting examples of fragmentases include transposases, restriction enzymes, Argonaute proteins, CRISPR -associated nucleases (Cas), endonucleases, exonucleases, and topoisomerases. Preferred fragmentation embodiments include methods that fragment a nucleic acid while also retaining proximity information of the fragments.

[0064] As used herein, the term "nucleotide sequence" or simply “sequence” is intended to refer to the order and type of nucleotide monomers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of formats including, for example, a depiction, image, electronic medium, series of symbols, series of numbers, series of letters, series of colors, etc. The information can be represented, for example, at single nucleotide resolution, at higher resolution (e.g. indicating molecular structure for nucleotide subunits) or at lower resolution (e.g. indicating chromosomal regions, such as haplotype blocks). A series of "A," "T," "G," and "C" letters isa well-known sequence representation for DNA that can be correlated, at single nucleotide resolution, with the actual sequence of a DNA molecule. A similar representation is used for RNA except that "T" is replaced with "U" in the series.

[0065] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 900 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. Other examples of reference sequences include genomes of other species, such as of control organisms as disclosed herein, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual. Examples of reference genomes include GRCh38 from the Genome Reference Consortium.

[0066] The term “nucleic acid sample” as used herein may refer to a sample, typically derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid molecule. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. Alternatively, the sample may be microbial such as bacteria, viral, or fungal. The nucleic acid sample may be taken from any organism, including but not limited to animals, plants, fungi and microbes. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparingplasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, the pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein. A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.

[0067] The sample can include high molecular weight material, such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some implementations, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some implementations, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another implementation, the sample can include nucleic acid molecules obtained from a non-mammalian source, such as a plant, bacteria, virus, or fungus. In some implementations, the source of the nucleic acid molecules may be an archived or extinct sample or species.

[0068] As further used herein, the term “sequencing run” refers to an iterative process on a sequencing device to determine a primary structure of nucleotide sequences from a sample (e.g., genomic sample). In particular, a sequencing run includes cycles of sequencing chemistry and imaging performed by a sequencing device (including an imaging device, such as a CCD or CMOS) that incorporate nucleobases into growing oligonucleotides to determinenucleotide reads from nucleotide sequences extracted from a sample (or other sequences within a library fragment) and seeded throughout a flow cell or other nucleotide-sample slide. In some cases, a sequencing run includes replicating oligonucleotides derived or extracted from one or more genomic samples seeded in clusters throughout a flow cell. Upon completing a sequencing run, a sequencing device can generate base-call data in a file, such as a binary base call (BCL) sequence file or a fast-all quality (FASTQ) file.

[0069] Relatedly, the term “sequencing cycle” (or “cycle”) refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to a sample’s sequence (e.g., a genomic or transcriptomic sequence from a sample) or a corresponding adapter sequence. In some cases, a sequencing cycle includes an iteration of both incorporating nucleobases into clusters of oligonucleotides using sequencing chemistry and capturing images of such clusters attached to a nucleotide-sample slide (e.g., a flow cell). Accordingly, cycles can be repeated as part of sequencing a nucleic-acid polymer (e.g., a sample genomic sequence). For example, in one or more embodiments, each sequencing cycle involves incorporating nucleobases into either a single nucleotide read in which DNA or RNA strands are read in only a single direction or paired-end reads in which DNA or RNA strands are read from both ends but in different cycles. Further, in certain cases, each sequencing cycle involves a camera taking an image of the nucleotide-sample slide or multiple sections of the nucleotide-sample slide to generate image data for determining a particular nucleobase added or incorporated into particular oligonucleotides. Following the image capture stage, a sequencing system can remove certain fluorescent labels from incorporated nucleobases and perform another sequencing cycle until the nucleic-acid polymer has been completely sequenced. In one or more embodiments, a sequencing cycle includes a cycle within an SBS run. A sequencing cycle can include one or both of an indexing cycle and a genomic sequencing cycle. For instance, one cluster of oligonucleotides or a set of clusters of oligonucleotides may be undergoing a genomic sequencing cycle in which nucleobases corresponding to a sample genomic sequence are incorporated and another cluster of oligonucleotides or another set of clusters of oligonucleotides may be concurrently undergoing an indexing cycle in which nucleobases corresponding to an indexing sequence for a nucleotide read are incorporated.

[0070] Further, as used herein, the term “nucleotide-sample slide” (or “nucleotide- sample substrate”) refers to a plate or substrate, such as a flow cell, comprising oligonucleotides for sequencing nucleotide sequences from genomic samples or other sample nucleic-acid polymers. In particular, a nucleotide-sample slide can refer to a substrate containing fluidic channels through which reagents and buffers can travel as part of sequencing. For example, in one or more embodiments, a flow cell (e.g., a patterned flow cell or non-patterned flow cell) may comprise small fluidic channels and oligonucleotide samples that can be bound to adapter sequences on the substrate. In other implementations, a nucleotide-sample slide can be an open substrate with one or more regions for oligonucleotide samples to be analyzed and the oligonucleotide samples may be positioned using charged pads or other means. In yet another implementation, the nucleotide-sample slide can be a membrane having a nanopore through which one or more oligonucleotide samples may pass.

[0071] Relatedly, as used herein, the term “region of a nucleotide-sample slide” (or “nucleotide-sample slide region”) refers to an area that is part of a nucleotide-sample slide. In particular, a region of a nucleotide-sample slide can refer to a discrete portion of a nucleotide- sample slide that differs from other portions of the nucleotide-sample slide. For instance, a region of a nucleotide-sample slide can include a subsection of a patterned flow cell comprising one or more wells (e.g., nano-wells) or a discrete subsection of a non-patterned flow cell (e.g., a subsection corresponding to one or more clusters). In some cases, a region (e.g., section) of a nucleotide-sample slide includes a tile or a sub-tile of a flow cell having clusters of oligonucleotides growing in parallel.

[0072] The term “read” or “sequence read” (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, thatcan be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained by any method known in the art. For example, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques.

[0073] Embodiments described herein can be used with any suitable sequencing chemistry, such as sequencing by synthesis (SBS), sequencing by avidity (also referred to as sequencing by binding), or sequencing by ligation.

[0074] SBS can be with or without the use of reversible terminators. For example, SBS can be initiated by contacting the target nucleic acids with one or more nucleotides (e.g., labelled, synthetic, modified, or a combination thereof), DNA polymerase, etc. Those features where a primer is extended using the target nucleic acid as a template will incorporate a labeled nucleotide that can be detected. The incorporation time used in a sequencing run can be significantly reduced using the altered polymerases described herein. Optionally, the labeled nucleotides can further include a reversible termination property that terminates further primer extension once a nucleotide has been added to a primer. For example, a nucleotide analog having a reversible terminator moiety can be added to a primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washes can be carried out between the various delivery steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, fluidic systems, and detection platforms that can be readily adapted for use with an array produced by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008); WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U.S. Pat. Nos. 7,057,026 B2, 7,329,492 B2, 7,211,414 B2, 7,315,019 B2, 7,405,281 B2, and 8,343,746 B2. Sequence reads can be generated using instruments such as MiniSeqTM, MiSeqTM, NextSeqTM, HiSeqTM, and NovaSeqTM sequencing instruments from Illumina, Inc. (San Diego, CA).

[0075] One example of SBS is termed sequencing by avidity. In sequencing by avidity, fluorescent dye-labeled cores are termed avidities. Sequencing by avidity is described in Arslan, S., Garcia, F.J., Guo, M. et al. Sequencing by avidity enables high accuracy with low reagent consumption. Nat Biotechnol 42, 132-138 (2024). doi.org / 10.1038 / s41587-023- 01750-7, which is incorporated by reference in its entirety.

[0076] One example of SBS using an open flow cell and without using reversible terminators is disclosed in Almogy, G.(2022) “Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform” doi.org / 10.1101 / 2022.05.29.493900, which is incorporated by reference in its entirety.

[0077] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on the detection of released protons can use an electrical detector and associated techniques that are described in U.S. Pat. Nos. 8,262,900 B2, 7,948,015 B2, 8,349,167 B2, and U.S. Pat. Pub. 2010 / 0137143 Al, which are incorporated by reference in its entirety.

[0078] Some embodiments can use methods involving the real-time monitoring of DNA polymerase activity. For example, nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate- labeled nucleotides, or with zeromode waveguides. Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33, 1026-1028 (2008); Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), which are incorporated by reference in its entirety. Techniques sequencing using zeromode waveguides is described in U.S. Pat. No. 6,917,726 B2, which is incorporated by reference in its entirety.

[0079] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood that the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. For example, the alignment of a read to the reference sequence for human chromosome 13 will tell the likelihood that the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to in the reference sequence. Forexample, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference genome (e.g., chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence.

[0080] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).

[0081] Alignment may be performed by modifications and / or combinations of methods such as Burrows-Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.

[0082] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, e.g., a reference genome, by alignment.

[0083] As used herein, the term “paired-end reads” or “paired end reads” refers to paired reads generated from sequencing the forward and reverse ends of a larger nucleic acid fragment. In some examples, the forward and reverse ends of a larger nucleic acid fragment may share the same name. The paired-end reads may be generated from paired end sequencing that obtains one read from each end of a nucleic acid fragment.

[0084] Moreover, as used herein, the term “alignment score” refers to a numeric score, metric, or other quantitative measurement evaluating an accuracy of an alignment between one or more nucleotide reads or a fragment of a nucleotide read and another nucleotidesequence from a reference genome. In particular, an alignment score includes a metric indicating a degree to which the nucleobases of one or more nucleotide reads (or a fragment thereof) match or are similar to a reference sequence or an alternate contiguous sequence from a reference genome. In certain implementations, an alignment score takes the form of a Smith- Waterman score or a variation or version of a Smith- Waterman score for local alignment, such as various settings or configurations used by DRAGEN by Illumina, Inc. for Smith-Waterman scoring.

[0085] As used herein, a “short sequence read” refers to a sequence read of between 50-500 bp, for example, about 50 - 100 bp, and includes paired end sequence reads.

[0086] As used herein, a “long sequence read” refers to a sequence read of more than about 500 bp, for example 500 - 250,000 bp or more. A long sequence read may be obtained from a long-read sequencing technology, or may be synthetically constructed by assembling multiple short sequence reads.

[0087] As used herein, a “file” includes electronic files. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.

[0088] The terms “solid support,” “solid surface,” and other grammatical equivalents herein refer to any substrate that is appropriate for or can be modified to be appropriate for the attachment of enzymes, nucleic acids, and complexes thereof. As will be appreciated by those in the art, the number of possible substrates is very large. Possible substrates include, but are not limited to, glass and modified or functionalized glass, polymers (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, plastics, optical fiber bundles, quartz, metal oxides, inorganic oxides, other suitable transparent materials, other suitable non-transparent materials, other suitable translucent materials, and combinations thereof. The composition and geometry of the solid support can vary with its use.

[0089] In some embodiments, the solid support or solid surface is a planar structure, such as a flowcell, slide, chip, microchip, array, microarray, wafer, panel, chargepad, and / or web. The planar structure can be a single surface structure having a single surface of sample / reaction sites. The planar structure can be a dual surface structure. One example of a dual surface structure includes a top substrate having a top surface of sample / reactions sites, a bottom substrate having a bottom surface of sample / reactions sites, and a spacer layer separating the top substrate and the bottom substrate. The solid support or solid surface can be open to direct application of a fluid. One example of an open solid support or open solid surface is an open flow cell having a single surface structure without an inlet port.

[0090] In some embodiments, the solid support comprises one or more surfaces of a flowcell or flow cell. The term “flowcell” or “flow cell” as used herein refers to a solid surface across which one or more fluid reagents can be flowed. Examples of flowcells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U.S. 7,057,026 B2; WO 91 / 06678; WO 07 / 123744; U.S. 7,329,492 B2; U.S. 7,211,414 B2; U.S. 7,315,019 B2; U.S. 7,405,281 B2, and U.S. Pat. Pub. 2008 / 0108082 Al, each of which is incorporated herein by reference in its entirety. In some embodiments, the flowcells can be one or more flow lanes. For flow cells having a plurality of flow lanes, each of the flow lanes can be independently accessed or two or more flow lanes can be accessed as a group.

[0091] In some embodiments, the solid support or solid surface is a non-planar structure, such beads, microspheres, and / or inner and / or outer surface of a tube or vessel. The terms “beads”, “microspheres,” or “particles” or grammatical equivalents herein is refer to small discrete particles. Suitable bead compositions include, but are not limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, thoria sol, carbon graphite, titanium dioxide, latex, polysaccharide (e.g. Dextran™ , Sepharose™, cellulose, nylon, cross-linked micelles, Teflon™, as well as any other materials outlined herein for solid supports may all be used. “Microsphere Detection Guide” from Bangs Laboratories, Fishers Ind. is a helpful guide. In certain embodiments, the microspheres are magnetic microspheres or beads. The beads need not be spherical; irregular particles may be used. Alternatively or additionally, the beads may be porous. The bead sizes range from nanometers, e.g. 100 nm, to millimeters, e.g. 1 mm, with beads from about 0.2 micron to about 200 micronsbeing preferred, and from about 0.5 to about 5 micron being particularly preferred, although in some embodiments smaller or larger beads may be used.

[0092] In some embodiments, the solid support comprises a patterned surface suitable for immobilization of molecules, such as enzymes, nucleic acids, and complexes thereof, in an ordered pattern. A “patterned surface” refers to an arrangement of different regions in or on an exposed layer of a solid support. The features can be separated by interstitial regions that contribute to the pattern. In some embodiments, the interstitial regions can be a different height, creating wells or raised platform patterns. In other embodiments, the interstitial regions can have different surface charges. In yet other embodiments, the interstitial regions can have different attachment moieties. In some embodiments, the pattern can be any suitable pattern, such as a grid pattern, radial patterns, and combinations thereof. Examples of grid patterns include rectangular patterns, hexagonal patterns, triangular, and other suitable grid patterns. The regions for immobilization of molecules may be depressed regions, elevated regions, or planar regions relative to the interstitial regions. The regions may be fabricated as is generally known in the art using a variety of techniques, including, but not limited to, photolithography, stamping techniques, molding techniques, microetching techniques, and combinations thereof. As will be appreciated by those in the art, the technique used will depend on the composition and shape of the regions. For example, the regions for immobilization of molecules of a patterned surface may be wells, pits, channels, posts, pillars, ridges, stripes, swirls, lines, and other suitable topographies. For example, the wells may have any opening in any shape, such as circular, oval, polygonal (e.g., hexagonal, octagonal, square, rectangular, elliptical, etc.). Exemplary patterned surfaces that can be used in the methods and compositions set forth herein are described in U.S. Pat. No. 8,778,849 B2, which is incorporated herein by reference in its entirety.

[0093] In some embodiments, the solid support comprises a surface suitable for immobilization of molecules, such as enzymes, nucleic acids, and complexes thereof, in a random distribution over the solid support. Exemplary random distribution over a solid support is described in U.S. Pat. No. 8,241,573 B2, which is incorporated herein by reference in its entirety.

[0094] As used herein, the term "flow cell" is intended to mean a chamber having a surface across which one or more fluid reagents can be flowed. Generally, a flow cell willhave an ingress opening and an egress opening to facilitate flow of fluid. A flow cell can have multiple surfaces. Examples of flow cells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al, Nature 456:53-59 (2008), WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,129,492; US 7,211,414; US 7,115,019; US 7,405,281, and US 2008 / 0108082, each of which is incorporated herein by reference.

[0095] In many embodiments, a solid support to which nucleic acids are attached in a method set forth herein will have a continuous or monolithic surface. Thus, fragments can attach at spatially random locations wherein the distance between nearest neighbor fragments (or nearest neighbor clusters derived from the fragments) will be variable. The resulting arrays will have a variable or random spatial pattern of features. Alternatively, a solid support used in a method set forth herein can include an array of features that are present in a repeating pattern. In such embodiments, the features provide the locations to which modified nucleic acid polymers, or fragments thereof, can attach. Particularly useful repeating patterns are hexagonal patterns, rectilinear patterns, grid patterns, patterns having reflective symmetry, patterns having rotational symmetry, or the like. The features to which a modified nucleic acid polymer, or fragment thereof, attach can each have an area that is smaller than about 1mm2, 500 pm2, 100 pm2, 25 pm2, 10 pm2, 5 pm2, 1 pm2, 500 nm2, or 100 nm2. Alternatively, or additionally, each feature can have an area that is larger than about 100 nm2, 250 nm2, 500 nm2, 1 pm2, 2.5 pm2, 5 pm2, 10 pm2, 100 pm2, or 500 pm2. A cluster or colony of nucleic acids that result from amplification of fragments on an array (whether patterned or spatially random) can similarly have an area that is in a range above or between an upper and lower limit selected from those exemplified above.

[0096] As used herein, the term "surface," when used in reference to a material, is intended to mean an external part or external layer of the material. The surface can be in contact with another material such as a gas, liquid, gel, polymer, organic polymer, second surface of a similar or different material, metal, or coat. The surface, or regions thereof, can be substantially flat. The surface can have surface features such as wells, pits, channels, ridges, raised regions, pegs, posts or the like. The material can be, for example, a solid support, gel, or the like.

[0097] As used herein, the term "target," when used in reference to a nucleic acid polymer, is intended to linguistically distinguish the nucleic acid, for example, from othernucleic acids, modified forms of the nucleic acid, fragments of the nucleic acid, and the like. Any of a variety of nucleic acids set forth herein can be identified as target nucleic acids, examples of which include genomic DNA (gDNA), messenger RNA (mRNA), copy or complimentary DNA (cDNA), and derivatives or analogs of these nucleic acids.

[0098] As used herein, the term "transposase" is intended to mean an enzyme that is capable of forming a functional complex with a transposon element- containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposon element-containing composition into a target DNA with which it is incubated, for example, in an in vitro transposition reaction. The term can also include integrases from retrotransposons and retroviruses. Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of U.S. Pat. App. Pub. 2010 / 0120098, which is incorporated herein by reference in its entirety. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon element with sufficient efficiency to tag a target nucleic acid can be used. In particular embodiments, a preferred transposition system is capable of inserting the transposon element in a random or almost random manner to tag the target nucleic acid. As used herein, the term "transposome" is intended to mean a transposase enzyme bound to a nucleic acid. Typically, the nucleic acid is double-stranded. For example, the complex can be the product of incubating a transposase enzyme with double-stranded transposon DNA under conditions that support non-covalent complex formation. Transposon DNA can include, without limitation, Tn5 DNA, a portion of Tn5 DNA, a transposon element composition, a mixture of transposon element compositions or other nucleic acids capable of interacting with a transposase such as the hyperactive Tn5 transposase.

[0099] As used herein, the term "transposon element" is intended to mean a nucleic acid molecule, or portion thereof, that includes the nucleotide sequences that form a transposome with a transposase or integrase enzyme. Typically, the nucleic acid molecule is a double-stranded DNA molecule. In some embodiments, a transposon element is capable of forming a functional complex with the transposase in a transposition reaction. As non-limiting examples, transposon elements can include the 19-bp outer end ("OE") transposon end, inner end ("IE") transposon end, or "mosaic end" ("ME") transposon end recognized by a wild-typeor mutant Tn5 transposase, or the R1 and R2 transposon end as set forth in the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Transposon elements can comprise any nucleic acid or nucleic acid analogue suitable for forming a functional complex with the transposase or integrase enzyme in an in vitro transposition reaction. For example, the transposon end can comprise DNA, RNA, modified bases, nonnatural bases, modified backbone, and can comprise nicks in one or both strands.

[0100] A standard NGS sequencing run yields millions of short sequences that are eventually mapped on a reference genome. A percentage of good-quality reads (1-5%) are discarded because of ambiguous genomic location. Increasing read length (2x500 or long-read sequencing), designing a specialized algorithm to map reads on specific regions of the genome (targeted callers), using expensive and time-consuming library preparation (Illumina CLR), or a combination thereof may be implemented to address the need for disambiguating such reads that would normally be discarded. However, such approaches are costly, laborious, and time intensive. Spatial information (X and Y coordinates) obtained from a solid support surface) can be leveraged to identify fragments that are generated from a single long input fragment and subsequentially be used to improve mapping reads in ambiguous positions.

[0101] In one or more embodiments, the system identifies and / or stores sequencing metrics within one or more sequencing data files. As used herein, the term “sequencing data file” refers to a digital file that includes genetic sequencing information concerning genotype calls or nucleotide reads generated by one or more genomic sequencing procedures. Such sequencing information may include, for example, nucleotide reads, alignment and mapping information, nucleotide reads at one or more genomic coordinates, and so forth.

[0102] Moreover, in one or more embodiments, one or more sequencing data files in which the system identifies or stores sequencing metrics include an alignment data file containing information from a read processing and mapping procedure. As used herein, the term “alignment data file” refers to a digital file that indicates mapping and alignment information for nucleotide reads of a sample nucleotide sequence. For example, an alignment data file can include a binary alignment map (BAM) file, a compressed reference-oriented alignment map (CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence.

[0103] Moreover, as used herein, the term “cluster of oligonucleotides” (or “cluster” or “oligonucleotide cluster”) refers to a localized group or collection of DNA or RNA on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster may include tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, the cluster can comprise one or more concatemers, such as, for example, a nanoball. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-patterned flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell. In some embodiments, a cluster can be monoclonal or polyclonal. Clusters may also be imaged or detected with localized changes in pH or changes in electrical conductivity.

[0104] The term “immobilized”, “affixed” and “attached” are used interchangeably herein and both terms are intended to encompass direct or indirect, covalent or non-covalent attachment unless indicated otherwise, either explicitly or by context.

[0105] Exemplary covalent attachment includes, for example, those that result from the use of click chemistry techniques. Exemplary non-covalent attachment includes, but are not limited to, non-specific interactions (e.g. hydrogen bonding, ionic bonding, van der Waals interactions etc.) or specific interactions (e.g. affinity interactions, receptor-ligand interactions, antibody-epitope interactions, avidin-biotin interactions, streptavidin-biotin interactions, lectin-carbohydrate interactions, etc.). Exemplary attachments are set forth in U. S. Pat. Nos. 6,737,236 Bl; 7,259,258 B2; 7,375,234 B2 and 7,427,678 B2; and U.S. Pat. Pub. 2011 / 0059865 Al, each of which is incorporated herein by reference in its entirety.

[0106] In certain embodiments, the molecules (e.g. nucleic acids, enzymes) remain immobilized or attached to the solid support under the conditions in which it is intended to use the solid support, for example in applications requiring nucleic acid amplification and / or sequencing. In other embodiments, the molecules are reversibly immobilized and can be removed from the solid support through the use of cleavable sites, linkers, and the like.

[0107] Some embodiments further comprise amplifying and / or replicating one or more nucleic acid templates, including fragments thereof. The amplifying and / or replicating comprises use of one or more of a bridge amplification reaction, an isothermal bridge amplification reaction, a rolling circle amplification (RCA) reaction, a modified rolling circle multiple displacement amplification, a helicase-dependent amplification reaction, a recombinase-dependent amplification reaction, a single-stranded DNA binding (SSB) protein mediated Isothermal amplification, a PCR reaction, a strand-displacement reaction, a ligase chain reaction, a transcription-mediated reaction, a loop-mediated amplification reaction, other suitable reactions, and combinations thereof. Some embodiments further include generating “polonies,” which refers to polymerase colonies created by RCA on a flow cell.

[0108] Some embodiments further comprise rolling circle amplification / replication used to form nucleic acid nanoballs. The term “nucleic acid nanoball” may be a concatemer comprising multiple copies of a target nucleic acid molecule. These nucleic acid copies may be arranged one after another in a continuous linear strand of nucleotides. These nucleic acid copies may result in a nanoball folding configuration. The multiple copies of a target nucleic acid molecule in a nucleic acid nanoball may each contain an adaptor sequence of known sequence to facilitate amplification or sequencing. The adaptor sequence of each target nucleic acid molecule may be the same or different. The nucleic acid nanoball can be loaded on the surface of a solid support. The nanoball can be attached to the surface of a solid support by any suitable method. Non-limiting examples of such methods include nucleic acid hybridization, biotin-streptavidin binding, thiol binding, photoactive binding, covalent binding, antibodyantigen, physical constraints via hydrogels or other porous polymers, etc., or combinations thereof. In some cases, the nanoball can be digested with an enzyme (nuclease, etc.) to produce a smaller nanoball or a fragment from the nanoball.

[0109] Embodiments of the present disclosure relate to methods and systems which use “links” or “linkage information” between sequence reads. The “link” or “linkage information” as discussed herein refers, in some embodiments, to the probability that two pairs of reads on a sequencing flow cell are derived from the same original nucleic acid molecule. In some next generation sequencing (NGS) systems, fragments of long nucleic acids, such as genomic DNA, from a sample are sheared to create shorter fragments which can be sequencedin a single read. The shearing process can create these shorter fragments which land on the flow cell and the flow cell proximity of each fragment may be related to the original nucleic acid molecule from which the fragment was derived. For example, fragments which come from the same nucleic acid molecule have been found to bind closer together on the flow cell as compared to fragments which come from different original nucleic acid molecules. Accordingly, if two clusters of reads on a flow cell are in close proximity and also close together on the genome, the clusters are more likely to have come from the same nucleic acid molecule. However, unrelated fragments may also bind to the flow cell near one another, which leads to an uncertainty in the probability that adjacent clusters originate from the same molecule. A number of factors could affect the probability that unrelated clusters would land in a similar area, and these factors may change based on a variety of experimental conditions. Embodiments of the disclosure provide a statistical method for calculating the probability that two reads are linked, such that on a flow cell the two reads were derived from the same nucleic acid molecule.

[0110] Embodiments of the disclosure relate to systems and methods for sequencing target nucleic acids by fragmenting the target nucleic acid and distributing the fragments onto a flow cell. As the fragments are distributed along the flow cell, they bind capture primers and are then used to create clusters by well-known technologies, such as those provided by Illumina Inc. (San Diego, CA). As described above, according to the methods of this disclosure, fragments which were derived from the same template genomic sequence are more likely to bind to the flow cell in proximity to one another as compared to fragments that are from different template genomic sequences, particularly when the fragmentation is performed directly on the flow cell using immobilized transposome complexes on the surface of the flow cell. In some library preps with fragmentation happening prior to loading, fragments can land anywhere in the flow cell independently of whether they came from the same molecule. However, in some embodiments when fragmentation is performed directly on the flow cell, flow cell proximity information is retained. This flow cell proximity information can be used to help guide assembly and variant calling of the original template genomic sequence, as will be described in more detail below.

[0111] For example, transposome complexes may be provided as part of the sequencing process. In some embodiments, the transposome complexes include a transposaseand a first polynucleotide having end sequences which can be used to fragment the target polynucleotides and insert into each fragment an end sequence or tag which can be used to bind to capture probes located on the substrate. The method can include contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides and add capture sequences to the ends of each fragment. In some embodiments, the capture sequences include P5 or P7 sequences as provided by Illumina®, Inc. (San Diego, CA). In some embodiments, the complexed strand and transposome are in solution, and are then brought towards a substrate and immobilized thereon. In some embodiments, prior to immobilization of the transposome complexes on the substrate, one or more of the transposome complexes bind the target polynucleotides in solution. In this embodiment, the transposome complexes in solution become immobilized to the substrate.

[0112] Once the fragments have been bound to the substrate, the bound fragments can be amplified to form a plurality of nucleic acid clusters on the substrate. The location of each cluster on the flow cell can then be determined before, during or after performing sequencing by synthesis reactions (SBS) to obtain the nucleotide sequence of each fragment located in each cluster. Once the nucleotide sequence of each cluster has been determined, the method can start to map those reads to determine the original target polynucleotide from which the read originated. In some embodiments, the mapping process takes into account the flow cell proximity of each cluster, such that clusters which are closer to each other on the flow cell are more likely to have originated from the same target polynucleotide. In some embodiments, the library preparation steps are performed on the flow cell, which may reduce the complexity and the amount of equipment associated with the systems. Furthermore, by mapping the sequenced fragments to target polynucleotides using the flow cell proximity information accompanying each cluster, the method performs more accurate mapping operations as compared to methods that do not take the flow cell proximity of each cluster into account during the mapping process. Therefore, flow cell proximity that includes relative distances between various clusters on a flow cell may be leveraged to adjust mapping information, thereby increasing the read quality of previously identified multi-mapped reads. In the past, identified multi-mapped reads may have been discarded. Increasing the read quality of these previously discarded reads, by improving the confidence of the read pair’s alignment based on linking information with a high link quality score, may improve the alignment information andquality of information used in certain genomic analysis applications including, but not limited to, variant calling. Processing DNA samples suitable for high-throughput sequencing that retain information on the original configuration of the DNA samples provides useful information on co-located fragments.

[0113] In some embodiments, linking information is determined by analyzing, for example, statistically analyzing with a model, the genomic distance between two reads and flow cell proximity between the two clusters on the flow cell. In some embodiments, the methods and systems determine whether the genomic distance and / or flow cell proximity is below a threshold. In some embodiments, the methods and systems determine the presence or absence of a link between the two sequence reads. In some embodiments, the methods and systems determine a linking quality score between the two sequence reads. In some embodiments, the methods and systems analyze genomic distance and flow cell proximity for a plurality of pairs of two sequence reads (for example, each possible pair) of two sequence reads in a dataset. Further details regarding sequencing conditions that result in links or downstream analyses utilizing linking information can be found in International Patent Application Nos. PCT / US2024 / 035447 and PCT / US2024 / 045996, International Patent Application Publication Nos. WO2015 / 189636, WO2015 / 095226 and WO2023 / 122755, and U.S. Provisional Patent Application Nos. 63 / 600460, 63 / 614066, 63 / 800,049, and 63 / 800,262, the disclosure of each of which is incorporated herein by reference in its entirety.

[0114] As used herein, the term “machine learning model” refers to a computer- implemented model that can be tuned (e.g., trained) based on inputs to approximate unknown functions. In particular, in some embodiments, a machine-learning model includes a model that utilizes computer processes or procedures to learn from, and make predictions on, known data by analyzing the known data to learn to generate outputs that reflect patterns and attributes of the known data. For instance, in some cases, a machine-learning model includes, but is not limited to, a neural network (e.g., a convolutional neural network, recurrent neural network or other deep learning network), a decision tree (e.g., a gradient boosted decision tree, such as XGBoost), association rule learning, inductive logic programming, support vector learning, a Bayesian network, a regression-based model (e.g., censored regression), principal component analysis, or a combination thereof.

[0115] Further, as used herein, the term “neural network” refers to a model of interconnected artificial neurons (e.g., organized in layers) that communicate and learn to approximate complex functions and generate outputs based on inputs provided to the model. In some instances, a neural network includes one or more machine learning processes or algorithms. Further, in some cases, a neural network includes a process (or set of processes) that implement deep learning techniques that utilize a set of procedures to model high-level abstractions in data. To illustrate, in some embodiments, a neural network includes a convolutional neural network, a recurrent neural network (e.g., a long short-term memory neural network), a generative adversarial network, a graph neural network, a multi-layer perceptron, or a diffusion neural network. In some embodiments, a neural network includes a combination of neural networks or neural network components.Systems for Detecting a Structural Variant

[0116] Further disclosed herein are electronic systems for detecting a structural variant in DNA from a sample placed on a flow cell. In some embodiments, the system includes a processor configured to perform a method comprising: obtaining flow cell data from the DNA, wherein the flow cell data comprises nucleic acid sequence reads and flow cell locations of sequence read clusters on the flow cell; aligning the sequence reads to one or more reference genomes to obtain a genomic location of the sequence reads; obtaining linking information between pairs of sequence reads on the flow cell based on the flow cell location and genomic location of each sequence read in the pairs of sequence reads; generating colocation data based on the linking information; providing the colocation data to a machine learning model; and processing the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acids and determining a confidence score, thereby detecting at least one structural variant for the sample.

[0117] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: obtain flow cell data from the DNA, wherein the flow cell data comprises nucleic acid sequence reads and flow cell locations of sequence read clusters on the flow cell; align the sequence reads to one or more reference genomes to obtain a genomic location of the sequencereads; obtain linking information between pairs of sequence reads on the flow cell based on the flow cell location and genomic location of each sequence read in the pairs of sequence reads; generate colocation data based on the linking information; provide the colocation data to a machine learning model; and process the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acids and determine a confidence score, thereby detecting at least one structural variant for the sample.

[0118] FIG. 1A illustrates a diagram of an environment in which a system for detecting a structural variant can operate in accordance with one or more implementations. The following paragraphs describe one embodiment of a structural variant detection system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG. 1A illustrates a schematic diagram of a computing system 1000 in which a structural variant detection application 1106 operates in accordance with one or more implementations. As illustrated, the computing system 1000 includes one or more server device(s) 1102 connected to a user client device 1108, a local device 1118, and a sequencing device 1114 via a network 1112. The network 1112 can comprise any suitable network over which computing devices can communicate.

[0119] As shown in FIG. 1A, the computing system 1000 includes the server device(s) 1102. In various implementations, the server device(s) 1102 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls or sequenced nucleic- acid polymers. In some implementations, the server device(s) 1102 receive various data from the sequencing device 1114, such as data from a sample genome and / or sequence reads. The server device(s) 1102 may also communicate with the user client device 1108. In particular, the server device(s) 1102 can send data for sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 1108.

[0120] As shown, the server device(s) 1102 includes a sequencing application 1110. In general, the sequencing application 1110 analyzes the data (such as call data) received from the sequencing device 1114 or elsewhere to determine nucleobase sequences for nucleic- acid polymers. For example, the sequencing application 1110 can receive raw data from the sequencing device 1114 and determine a nucleobase sequence for a sample genome or a nucleic-acid segment. In some implementations, the sequencing application 1110 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.

[0121] As also shown, the sequencing application 1110 includes the structural variant detection application 1106. As described below, in some embodiments, the structural variant detection application 1106 can receive input and data from the sequencing application as described in more detail below to generate one or more colocation plots. The colocation plots can be generated and stored within the server device 1102, the sequencing device 1114, a local device 1118 or a user client device 1108, or some combination of these devices. For example, in some embodiments, the structural variant detection application 1106 may include links to a trained neural network configured to input colocation plots and determine the positions of structural variants within a genome. The trained neural network may be maintained within the service device 1102 or located on one or more external systems. For example, the structural variant detection application 1106 can include the methods shown in Figures 2 or 3 and described herein.

[0122] While the sequencing application 1110 has been described as including the structural variant detection application 1106, other systems or methods may be included within the sequencing application 1110, such as an application to determine single nucleotide polymorphisms or to assemble sequence reads (not illustrated).

[0123] Moreover, while the structural variant detection application 1106 is described being implemented on the server device(s) 1102, as part of the sequencing application 1110, in some implementations, the structural variant detection application 1106 is implemented by (such as located entirely or in part) on the user client device 1108, the sequencing device 1114, and / or the local device 1118. As mentioned, in some implementations, structural variant detection application 1106 is implemented by one or more other components of the computing system 1000, such as the sequencing device 1114. In particular, the structural variant detection application 1106 can be implemented in a variety of different ways across the server device(s) 1102, the network 1112, the user client device 1108, the local device 1118, and the sequencing device 1114.

[0124] As further shown in FIG. 1A, the computing system 1000 includes the user client device 1108. In various implementations, the user client device 1108 can generate, store, receive, and send digital data. In particular, the user client device 1108 can receive the data from the sequencing device 1114. As further illustrated, the user client device 1108 includes a sequencing application 1110. The sequencing application 1110 may be a web application or anative application stored and executed on the user client device 1108 (for example, a mobile application, desktop application, or web application). The sequencing application 1110 can receive data from the sequencing application 1110 and / or structural variant detection application 1106. For example, the user client device 1108 can receive variant call files and / or alignment files from the sequencing application 1110.

[0125] The sequencing application 1110 can also include instructions that (when executed) cause the user client device 1108 to receive data from the structural variant detection application 1106 and present data from the sequencing device 1114 and / or the server device(s) 1102. Furthermore, the sequencing application 1110 can instruct the user client device 1108 to display data for variant calls, such as nucleobase calls or an indication of a structural variant. Indeed, the user client device 1108 can display nucleobase call results for a genome sample and / or an indication of a predicted structural variant.

[0126] As further shown in FIG. 1A, the computing system 1000 includes the sequencing device 1114. In various implementations, the sequencing device 1114 can sequence a genomic sample or other nucleic-acid polymer. For example, the sequencing device 1114 analyzes nucleic-acid segments or oligonucleotides extracted from genomic samples to generate data either directly or indirectly on the sequencing device 1114. More particularly, the sequencing device 1114 receives and analyzes, within nucleotide-sample slides (such as flow cells), nucleic-acid sequences extracted from genomic samples. In one or more implementations, the sequencing device 1114 utilizes sequencing by synthesis (SBS) to sequence a genomic sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 1112, in some implementations, the sequencing device 1114 bypasses the network 1112 and communicates directly with the user client device 1108.

[0127] As further depicted in FIG. 1A, in some implementations, the server device(s) 1102 includes a distributed collection of servers, where the server device(s) 1102 include several server devices distributed across the network 1112 and located in the same or different physical locations. For instance, the server device(s) 1102 can be implemented, in whole or in part, on the local device 1118. To illustrate, the local device 1118 may implement the sequencing application 1110 and / or the structural variant detection application 1106. Further, the server device(s) 1102 and / or the local device 1118 can include a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0128] The user client device 1108 illustrated in FIG. 1 A can include various types of client devices. For example, in some implementations, the user client device 1108 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In various implementations, the user client device 1108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.

[0129] Though FIG. 1A illustrates the components of the computing system 1000 communicating via the network 1112, in certain implementations, the components of the computing system 1000 can also communicate directly with each other, bypassing the network 1112. For instance, in some implementations, the user client device 1108 communicates directly with the sequencing device 1114. Additionally, in some implementations, the user client device 1108 communicates directly with the structural variant detection application 1106 and / or the server device(s) 1102. In some implementations, the user client device 1108 communicates directly with the local device 1118. Moreover, the structural variant detection application 1106 can access one or more databases housed on or accessed by the server device(s) 1102 or elsewhere in the computing system 1000.

[0130] FIG. IB is a block diagram of an exemplary server device 1102 that may be used in connection with the computing system 1000 of FIG. 1A. The server device 1102 may be configured to receive flow cell data from the DNA, wherein the flow cell data comprises nucleic acid sequence reads and flow cell locations of sequence read clusters on the flow cell; align the sequence reads to one or more reference genomes to obtain a genomic location of the sequence reads; obtain linking information between pairs of sequence reads on the flow cell based on the flow cell location and genomic location of each sequence read in the pairs of sequence reads; generate colocation data based on the linking information; provide the colocation data to a machine learning model; and process the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acids and determine a confidence score, thereby detecting at least one structural variant for the sample. The general architecture of the server device 1102 depicted in FIG. IB includes an arrangement of computer hardware and software components. The server device 1102 may include many more (or fewer) elements than those shown in FIG. IB. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 1102 includes a processing unit 110, anetwork interface 120, a computer readable medium drive 130, an input / output device interface 140, a display 150, and an input device 160, all of which may communicate with one another by way of a communication bus. The network interface 120 may provide connectivity to one or more networks or computing systems. The processing unit 110 may thus receive information and instructions from other computing systems or services via a network. The processing unit 110 may also communicate to and from memory 170 and further provide output information for an optional display 150 via the input / output device interface 140. The input / output device interface 140 may also accept input from the optional input device 160, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0131] The memory 170 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 110 executes in order to implement one or more embodiments. The memory 170 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 170 may store an operating system 172 that provides computer program instructions for use by the processing unit 110 in the general administration and operation of the server device 1102. The memory 170 may store a reference genome 173, such as for use by the sequencing application 1110. The memory 170 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0132] For example, in one embodiment, the memory 170 includes a sequencing application 1110, which may include a structural variant detection application 1106. The structural variant detection application 1106 can perform the methods disclosed herein. In addition, memory 170 may include or communicate with the data store 190 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of aligning sequence reads, and / or one or more reference genomes.

[0133] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments,the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.

[0134] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.

[0135] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.

[0136] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein is installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.

[0137] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for the implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktopcomputer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (such as PDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.

[0138] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.

[0139] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computerreadable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.

[0140] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.

[0141] In some embodiments, a hardware platform for providing a computational environment comprises a processor (such as CPU) wherein processor time and memory layout such as random access memory (such as RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.

[0142] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (such as grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perlinterfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.Methods for Detecting a Structural Variant

[0143] FIG. 2 is a flow diagram that schematically illustrates an exemplary method 200 for detecting one or more structural variants in a DNA sample, particularly a genomic DNA sample, placed on a flow cell. For example, a DNA sample 201 may be placed on flow cell 205. In some embodiments, the DNA sample 201 comprises DNA fragments that are greater than 500 bp in length, greater than 1 kbp in length, greater than 5 kbp in length, greater than 10 kbp in length, greater than 100 kbp in length, greater than 200 kbp in length, greater than 250 kbp in length, greater than 300 kbp in length, greater than 500 kbp in length, or more. In some embodiments the DNA fragments are high molecular weight DNA of about 250 kbp in length or more.Obtaining Sequence Reads

[0144] Sequence reads can be obtained from sample nucleic acids. For example, in some embodiments, the flow cell comprises transposome complexes bound to the flow cell. The transposome complexes include a transposase and a first polynucleotide comprising an end sequence and a first tag. A library preparation of the genomic sequences is then contacted with the flow cell and transposomes in order to contact the transposome complexes with the target genomic DNA sample under conditions to cause the transposome to fragment the genomic DNA sample. Because the transposomes are bound to the flow cell, following cleavage of the genomic DNA sample, the resulting fragments become bound to the flow cell. The process may then include amplifying the fragmented genomic DNA to form a plurality of nucleic acid clusters on a flow cell. A sequencing by synthesis process may then be started to sequence the nucleic acids in each cluster on the flow cell to generate sequence reads. In some embodiments, the sequence reads comprise paired end sequence reads where each nucleic acid is sequenced by two primers bound in opposite directions to one another on the bound fragment.

[0145] Embodiments of the disclosure relate to systems and methods for sequencing target nucleic acids by fragmenting the target nucleic acid and distributing thefragments onto a flow cell. As the fragments are distributed along the flow cell, they bind capture primers and are then used to create clusters by well-known technologies, such as those provided by Illumina Inc. (San Diego, CA). According to the methods of this disclosure, fragments which were derived from the same template genomic sequence are more likely to bind to the flow cell in close physical proximity as compared to fragments that are from different template genomic sequences, particularly when the fragmentation is performed directly on the flow cell using immobilized transposome complexes on the surface of the flow cell. In some embodiments, the library preparation steps are performed on the flow cell, which may reduce the complexity and the amount of equipment associated with the systems. In some library preparations with fragmentation happening prior to loading, fragments can land anywhere in the flow cell independently of whether they came from the same molecule. However, when fragmentation is performed directly on the flow cell proximity information is retained. This flow cell proximity information can be used to help guide assembly and variant calling of the original template genomic sequence, as will be described in more detail below.

[0146] For example, transposome complexes may be provided as part of the sequencing process. In some embodiments, the transposome complexes include a transposase and a first polynucleotide having end sequences which can be used to fragment the target polynucleotides and insert into each fragment an end sequence or tag which can be used to bind to capture probes located on the substrate. The method can include contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides and add capture sequences to the ends of each fragment. In some embodiments, the capture sequences include P5 or P7 sequences as provided by Illumina, Inc. In some embodiments, the complexed strand and transposome is in solution, and is then brought towards a substrate and immobilized thereon. In some embodiments, prior to immobilization of the transposome complexes on the substrate, one or more of the transposome complexes bind the target polynucleotides in solution. In this embodiment, the transposome complexes in solution become immobilized to the substrate.

[0147] While embodiments involving transposome complexes have been described above, a variety of sequencing library preparation methods and sequencing techniques may be used to capture flow cell and genomic proximity information for sequence reads. For example, in some embodiments, the methods herein include rolling circle amplification or DNA nanoballsequencing. For example, in some embodiments, any library preparation method which facilitates the recording and storing of flow cell proximity and genomic proximity information. For example, the methods can include library preparation by a multitude of different methods to prepare the nucleic acid fragments on the flow cell.

[0148] In some embodiments, the method 200 includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell. The method 200 may proceed to block 210, wherein sequence reads are obtained by SBS or other methods from each of the clusters generated on the flow cell.Obtaining Geographic Location Information

[0149] The method 200 may proceed to block 220, wherein geographic location information for each of the sequence reads is obtained. Geographic location information can include locations of sequence read clusters on the flow cell 205, where each cluster corresponds to a nucleic acid fragment that is sequenced at block 210. For example, flow cell locations of the sequence read clusters on the flow cell can be obtained. For example, in some embodiments, the geographic location information comprises spatial coordinates in a cartesian coordinate system. For example, the dimensions of the flow cell may be mapped with a cartesian coordinate system, e.g., with x and y dimensions, and locations of clusters on a flow cell may be assigned to a spatial coordinate using this system. The spatial coordinates can be used to determine the flow cell proximity between two or more clusters on the flow cell.

[0150] For example, once the fragments have been bound to substrate, the bound fragments can be amplified to form a plurality of nucleic acid clusters on the substrate. While block 220 is described as taking place after block 210 in the example of method 200, the location of each cluster on the flow cell can then be determined before, during or after performing sequencing by synthesis reactions (SBS) to obtain the nucleotide sequence of each fragment located in each cluster.Comparing to a Reference Genome

[0151] The method 200 may proceed to block 230, wherein sequence reads are compared to a reference genome. For example, the sequence reads can be aligned to one or more reference genomes to obtain a genomic location of the sequence reads. The alignmentcan be accomplished by a local alignment method or by any other method. For example, a set of reference genomes, a multireference genome, or a reference genome comprising alternative contigs including common structural variants, may be used for the alignment.

[0152] In some embodiments, at least a subset of sequence reads is aligned to a reference genome without regard to flow cell proximity of those sequence reads. In some embodiments, “linkage information” as further described below may be used to facilitate or improve alignments of sequence reads to a reference genome.Identify Read Pairs with Links

[0153] The method 200 may proceed to block 240, wherein “links” between read pairs located on the flow cell are identified. Embodiments of the present disclosure relate to methods and systems which use “links,” or “linkage information” between sequence reads. The “link,” “linking,” or “linkage information” as discussed herein refers, in some embodiments, to the probability that two read pairs on a sequencing flow cell are derived from the same original nucleic acid molecule.

[0154] In some embodiments, at block 240, linking information is obtained between pairs of sequence reads on the flow cell based on the flow cell location and genomic location of each sequence read in the pairs of sequence reads. For example, the linking information can include a probability that two read pairs are derived from the same long (over 500 bp) fragment from the DNA sample 201 prepared on the flow cell 205. If the probability is over a predetermined threshold, the two read pairs can be determined to be linked.. Methods for quantifying sequence read connectivity based on flow cell proximity are further described in U.S. Patent Application No. 63 / 700,049, which is hereby incorporated by reference in its entirety. Links can also be determined by counting the number of times fragments from each bin are in close proximity (e.g. below a predetermined threshold), or by accumulating the probability of the fragments in each bin being from the same template molecule given flow cell proximity (given an estimated proximity model).

[0155] For example, as described above, in some next generation sequencing (NGS) systems, fragments of long DNA, such as genomic DNA, from a biological source, for example the DNA sample 201, are sheared to create shorter fragments which can be sequenced in a single read, for example, at block 210. The shearing process can create these shorterfragments which land on the flow cell 205 and the flow cell proximity of each fragment may be related to the original nucleic acid molecule from which the fragment was derived. For example, fragments that come from the same nucleic acid molecule have been found to bind geographically closer to one another on the flow cell as compared to fragments which come from different original nucleic acid molecules. Accordingly, if two clusters of reads on a flow cell are in close physical proximity and also close together on the genome, the clusters have a higher probability of being derived from the same nucleic acid molecule. However, in some cases, unrelated fragments may also bind to the flow cell near one another, which leads to an uncertainty in the probability that adjacent clusters originate from the same molecule. A number of factors can affect the probability that unrelated clusters would land in a similar area, and these factors may change based on a variety of experimental conditions. Embodiments of the disclosure provide a statistical method for calculating the probability that two reads are linked, such that on a flow cell the two reads were derived from the same nucleic acid molecule.

[0156] In some embodiments, the methods and systems shown in Fig. 2 align an initial set of reads to the reference genome (e.g., a subset of all of the sequence reads), for example at block 240. In some embodiments, the methods and systems use an initial alignment in order to define linking model parameters based on the flow cell proximity and mapping location. For example, based on the geographic and genomic locations of the initial set of reads, a linking model can be constructed. In some embodiments, the linking model can determine a probability that a first read pair is derived from the same original long DNA fragment as a second read pair from a different cluster on the flow cell.

[0157] In some embodiments, once linking model parameters for the current sample are defined, the full link-informed alignment process begins. For example, when reads are being aligned to a reference sequence, various alignment candidates are evaluated for each read, and potential links for each candidate alignment position of a read are evaluated and a boost to the mapping quality is calculated from the linking information in each possible candidate position. For example, in some embodiments, all candidate alignments are obtained and a pairwise comparison between potentially linked reads is performed. In some embodiments, a link bonus for each candidate alignment is determined based on the linking model parameters.

[0158] Furthermore, by mapping the sequenced fragments to target polynucleotides using the flow cell proximity information accompanying each cluster, the method performs more accurate mapping operations as compared to methods that do not take the flow cell proximity of each cluster into account during the mapping process. Therefore, flow cell proximity information that includes relative distances between various clusters on a flow cell may be leveraged to adjust mapping information, thereby increasing the read quality and alignment accuracy of previously identified multi-mapped reads. In the past, identified multi-mapped reads may have been discarded. Increasing the read quality of these previously discarded reads, by improving the confidence of read pair’s alignment based on linking information with a high link quality score, may improve the alignment information and quality of information used in certain genomic analysis applications including, but not limited to, variant calling. Processing DNA samples suitable for high-throughput sequencing that retain information on the original configuration of the DNA samples provides useful information on co-located fragments. Furthermore, in the past, reads mapped with equal probability to more than one location would map with a Phred scale mapping score mapqO, indicating no confidence that the sequence reads are mapped to the right location and that it is plausible they originated from a first position or at least one other position. With the linking information, in some embodiments, this ambiguity can be resolved and the reads mapped in the correct position.

[0159] In some embodiments, once a linking model is obtained, linking information is used to align sequence reads to the reference genome. In some embodiments, linking information is updated based on the alignments. In some embodiments, the alignments can be updated based on the updated linking information. In some embodiments, the process continues to iteratively update alignments and linking information.

[0160] Further details regarding sequencing conditions that result in links or downstream analyses utilizing linking information can be found in International Patent Application Nos. PCT / US2024 / 035447 and PCT / US2024 / 045996, International Patent Application Publication Nos. WO2015 / 189636, WO2015 / 095226 and WO2023 / 122755, and U.S. Provisional Patent Application Nos. 63 / 600,460, 63 / 614,066, 63 / 800,049, and 63 / 800,262, the disclosure of each of which is incorporated herein by reference in its entirety.Plot Linking Information on Colocation Plot

[0161] The method 200 can then proceed to block 250, wherein linking information is plotted on a colocation plot, for example the exemplary colocation plot 255.

[0162] As used mentioned above, a colocation plot refers to a graphical representation to visualize the number of links between sequence reads. In some embodiments, the reference genome is divided into a plurality of subsections called “bins.” In some embodiments, the systems and methods can count the number of links between pairs of sequence reads by bin. In some embodiments, the systems and methods can visualize the count of links using the colocation plot. For example, in some embodiments, the colocation plot has x and y coordinates corresponding to bins along reference genome. In some embodiments, the colocation plot provides a visualization of the number of links between bins on the reference genome. Bins can be sized as needed in order to visualize structural variants based on the links. Various sizes of bins are contemplated to accommodate different implementations. Various considerations can be taken into account when determining the size of bin. For example, if the bins are too small then they may include too few sequence reads, which can cause increased noise per bin (more alignments increase SNR). If the bins are too large, it can be difficult to detect SVs, e.g. events of interest may be a similar size to a bin, so identifying breakpoints and event types may become difficult. In some embodiments, bin size is selected such that the SV events are at least a few bins wide, e.g. 5-10 bins or more. If SV events are so big such that they larger than a colocation plot image, event breakpoints can be concatenated (e.g. split the image so that it shows 2 different positions) or zoom out wherein the bin size is made larger. One of skill in the art can determine a bin size based on these and other considerations.

[0163] Bins may be about 500 bp, about 1 kbp, about 2 kbp, about 5 kbp, about 10 kbp, about 20 kbp, about 30 kbp, about 40 kbp, about 50 kbp, about 60 kbp, about 70 kbp, about 80 kbp, about 90 kbp, about 100 kbp, or larger, or a range constructed from any of the aforementioned values. For example, bins can be large (such as lOkbp - 100 kbp, or any range or value therebetween) or smaller, (such as 500 bp - 10 kbp, or any range or value therebetween). Larger bins can allow for visualization of links between sequence reads that map farther away from each other on the reference genome. Smaller bins can allow for visualization of links with greater specificity and resolution of a smaller section of the reference genome. The number of bins used to generate a colocation plot can be optimized to facilitatecomputational efficiency, thereby increasing the speed by which the colocation plots are generated. The person of ordinary skill in the art may determine a bin size based on these and other considerations.

[0164] In some embodiments, generating colocation data based on the linking information comprises: dividing the alignment of sequence reads to the reference genome into a plurality of bins. In some embodiments, the methods and systems count an estimated number of links between pairs of sequence reads by bin. In some embodiments, the methods and systems accumulate the probability of links in each bin based on a probabilistic model. In some embodiments, links depend on read alignments. In some embodiments, read alignment depends on which other reads sequence reads are linked to. For example, in some embodiments, mapqO ambiguities are resolved by identifying nearby reference segments to which there is a link. Thus, in some embodiments, links and read alignments can iteratively update each other.

[0165] While the exemplary method 200 includes a step of plotting linking information on a colocation plot, in other embodiments, colocation data can be generated that is based on linking information in other ways. For example, the colocation data can comprise a graph representing linking relationships between different regions of the reference genome. In some graph embodiments, each node (also referred to as a vertice) is a genomic bin and each edge connecting each node represents linking information between two bins, for example, the number of links between two bins. For example, if a sequence read in bin A is linked to a sequence read in bin D, the edge can represent the link between bin A and bin D. Similarly, the graph may not include a node between bin A and other bins if bin A is not linked to bin B, bin C, or other bins. Thus, in some embodiments, involving a graph can advantageously provide a more efficient representation of linking information compared to images, because there may be less data to process. For example, graphs may generally be useful when connections are sparse. For example, in the case of linking information, in some embodiments, there are dense connections along the diagonal (e.g., between neighboring bins) and sparse connections off- diagonal (e.g., between non-neighboring bins). There are numerous implementations of neural networks and other machine learning models that can operate on such graph representations. Graphs may be preferable in some situations over image representations, because image representations may be constrained to Cartesian grids, while graphs may allow depiction of more complex data. In some embodiments, the colocation data comprises counts of linksbetween sequence reads in different subsections (e.g., bins) of the reference genome. For example, these counts can be formatted on a matrix.Feed Colocation Plot Data into Trained Model

[0166] The method 200 can proceed to block 260, wherein colocation plot data is fed into a trained machine learning model or neural network. For example, the colocation data can be provided to a machine learning model which has been trained to recognize various structural variant features found in a colocation plot.

[0167] In some embodiments, the machine learning model comprises an object detection model, for example, the object detection network 600 described further below.

[0168] In some embodiments, the machine learning model comprises an anomaly detection model. Alternatively, an anomaly detection model can be a statistical model, not just a machine learning model. For example, a statistical model could be a distribution analysis to determine if the distribution of links in a sample region is different from known homozygous reference regions, or if there is an edge event in the sample region.

[0169] In some embodiments, the machine learning model comprises a deep learning image processing model. In some embodiments, the machine learning model comprises a deep convolutional neural network (CNN). In some embodiments, the machine learning model comprises a transformer-based machine learning model.

[0170] In some embodiments, the machine learning model comprises a graph neural network (GNN). In some embodiments, the machine learning model comprises an image classification model. For example, an image classification model could move a window over the diagonal of a colocation plot image and classify each position.

[0171] In some embodiments, additional information beyond the colocation data is provided to the machine learning model to increase the accuracy of detecting structural variants from the colocation plots. The machine learning model can use this additional information in further processing steps. For example, the additional information can include a mapping quality measurement, such as a count of MAPQ = 0 sequence reads, which are reads which cannot be assigned to a genomic location on the reference genome, and / or the average mapping quality for each bin. In some embodiments, the colocation data is segmented based on sequence read mapping quality. For example, in some embodiments, the colocation plot data includesone input for reads with MAPQO, another for reads with MAPQ between 0 and 20, another for reads with MAPQ between 20 and 40, and another for reads with MAPQ above 40. For example, a separate colocation plot can be generated for each input.

[0172] The additional information can also include a GC content bias measurement, such as a measure of GC content in a region of the reference genome and / or in sequence reads. The additional information can also include an initial CNV / structural variant prediction, for example one that has been determined from another method.

[0173] The additional information can also include abnormal paired read information, such as an indication that read pairs are abnormal in a region. The additional information can also include split read information, such as sequence reads that are partially mapped to one bin and partially to a different one. In some embodiments, the split reads partially map to each of two target bins.

[0174] The additional information can also include a link quality score, such as a measurement of confidence in links between one sequence read and another, and / or a measurement of average / median link quality scores in a region of the reference genome. The additional information can also include the genomic distance between sequence reads from the alignment to the reference genome. The additional information can also include a comparison to a normal sample or a set of normal samples. The additional information can also include a set of linking information statistical parameters that characterize the sample under test, e.g. link length distribution, and coverage distribution (mean, variance).

[0175] The additional information can also include phasing of sequence reads to parental haplotypes. In some embodiments, where sequence reads are assigned to a haplotype (e.g. phased), the machine learning model processes each haplotype separately. This can improve accuracy and efficiency because in the case of heterozygous variants, it may be less complicated to recognize a visual feature when haplotypes are separated rather than processed together.Identify Structural Variants Based on Colocation Data

[0176] After the colocation plot data has been fed into the trained model at the step 260, the method 200 can proceed to block 270, wherein structural variants are identified based on the output of the trained model from the colocation data. In addition, the system may outputa visual representation of the colocation data, such as colocation plot 275 which can be viewed and compared with the structural variant data output by the trained model. The user may compare the structural variant data and calls output by the trained model with the data illustrated on the colocation plot 275 to ensure that the colocation plot 275 appears to indicate structural variations at the positions identified from the trained model analysis.

[0177] In some embodiments, the methods and systems process the colocation data with the trained machine learning model to identify one or more predicted structural variant regions within the genomic DNA. For example, in some embodiments the methods and systems determine one or more of: an area with a suspected SV, a structural variant event class, a structural variant position, and a structural variant size.

[0178] In some embodiments, the structural variant event class comprises an insertion, a deletion, an inversion, or a translocation. In some embodiments, the structural variant event classes include different types of insertion variant, for example, duplication, novel insertion, translocations, or a mix of multiple types. In some embodiments, the structural variant event classes include an indication of whether the variant is a homozygous or heterozygous variant. In some embodiments, homozygous vs heterozygous structural variants are separate categories, e.g., homozygous deletion is classified separately from heterozygous depletion.

[0179] In some embodiments, each colocation plot is processed, and the methods and systems output multiple SV classifications. For example, some regions of the sample genome may have more than 1 SV present. Therefore, in some embodiments, an object detection network can detect multiple objects (e.g., SV type, heterozygous or homozygous), each classified independently, and a determined size and position for each.

[0180] The methods and systems can process the colocation data at block 270 using a variety of techniques, including techniques described further herein: active region detection, an object detection network, one-stage object detection, two-stage object detection, and nonmax suppression, or any suitable method known to those of skill in the art.

[0181] In some embodiments where linking information is represented with a graph, the graph is used as input to a Graph Neural Network (GNN). In some embodiments, to identify a SV, the trained GNN classifies new edges between nodes, for example, edges that are new compared to a graph from a reference sequence. In some embodiments, these newedges represent SV classes, such as insertions, deletions, inversions, etc. To identify breakpoints, the methods and systems can identify nodes connected by SV edges and extract their genomic coordinates.

[0182] In some embodiments, the machine learning model also determines a confidence score. In some embodiments, the confidence score is a statistical representation of the probability that a detected structural variant exists. In some embodiments, the methods and systems determine whether the confidence score is above a predetermined threshold. In some embodiments, the methods and systems detect at least one structural variant for the sample.

[0183] In some embodiments, the machine learning model comprises a statistical model that detects anomalous links, for example, links between genomically distant regions of the reference genome or an absence of links between regions that are expected to be closer in the reference genome but appear to be further away in a sample, for example due to an insertion between the regions.

[0184] In some embodiments, the machine learning model detects a structural variant anomaly region corresponding to a complex or overlapping structural variant. For example, in some regions, there may be multiple SVs, and / or the SVs can be complex or overlapping and therefore hard for the machine learning model to classify correctly as a specific structural variant. In such embodiments, the machine learning model may output a detection of an ‘anomaly.’ For example, in some embodiments, the machine learning outputs a detection of an anomaly if it detects a disruption in a standard diagonal feature of a colocation plot, where a standard diagonal would represent a region in the sample that matches the reference genome, but the machine learning model was unable to classify a specific or single structural variant.

[0185] In some embodiments, the methods and systems store structural variant information in an electronic file. In some embodiments, the file is a BED file or a VCF file. For example, a structural variant call based on a colocation plot can be converted and stored in a VCF file.Active Region Detection

[0186] In some embodiments, the methods and systems process colocation data by detecting an active region containing a potential structural variant based on a first set ofcolocation data, and then refine the colocation data and process the active region to identify a structural variant.

[0187] For example, at block 270, the method 200 can include identifying a putative structural variant region based on a first colocation data. As used herein, a “putative structural variant” refers to an initial classification of a potential structural variant based on an initial analysis of linking information or other sequence read information. The first colocation data can include linking information represented in various ways, such as a graph, matrix, or an image. For example, in some embodiments, the first colocation data includes a first colocation plot using bins of a first size. In some embodiments, first colocation data includes a graph based on linking information.

[0188] Next, the methods and systems can generate second colocation data based on the putative structural variant. For example, in some embodiments, the second colocation data is at a higher resolution than the first colocation data. The higher resolution data may be generated by, for example using bins of a smaller size than the first colocation data. By example and not by way of limitation, in some embodiments, the first colocation data is based on bins of 10 kbp to 100 kbp in size, and wherein the second colocation data is based on bins of 500 bp to 10 kbp in size. Alternatively, in some embodiments, the second colocation data is at a lower resolution than the first colocation data, for example is generated using bins of a larger size than the first colocation data.

[0189] Next, the methods and systems can provide the second colocation data to a machine learning model. In some embodiments, the second colocation data comprises a colocation plot or other image or graph representation of linkage information, and the machine learning model is an object detection network.

[0190] Next, the methods and systems can process the second colocation data with the machine learning model to identify a structural variant within the genomic nucleic acids. The machine learning model can process the second colocation data using the methods described herein. For example, an object detection network as described herein can process a visual representation of colocation data, such as a colocation plot, to identify a structural variant and determine a confidence score.

[0191] In some embodiments, instead of the active region detection as a separate function, the methods and systems use the image network to process the entire diagonal (e.g.every segment of the genome is analyzed by the image network). The methods and systems can process the diagonal feature of a colocation plot image to identify significant events, e.g. where the diagonal has a visual feature that differs from being homozygous for matching the reference sequence. Then the methods and systems can process the off-diagonal regions of a colocation plot image to detect an SV.Methods for training a machine learning model to detect structural variants in a genomic sequence

[0192] Further disclosed herein are methods for training a machine learning model to detect structural variants in a genomic sequence. FIG. 3 is a flow diagram that schematically illustrates an exemplary method 300 for training a machine learning model to detect structural variants in a genomic sequence. The method 300 can begin from start block 310.Obtaining a plurality of training datasets

[0193] For example, the method 300 can proceed to block 320, wherein sequence reads are determined from clusters on a flowcell. This can be accomplished as described above, for example at block 210 of the method 200. For example, sequence reads can be obtained by SBS or other methods from each of the clusters generated on the flow cell

[0194] Next, the method 300 can proceed to block 330, wherein locations of clusters on the flow cell are determined. This can be accomplished as described above, for example at block 220 of the method 200. For example, geographic location information for each of the sequence reads may be obtained. Geographic location information can include locations of sequence read clusters on the flow cell, where each cluster corresponds to a nucleic acid fragment that is sequenced. For example, flow cell locations of the sequence read clusters on the flow cell can be obtained. For example, in some embodiments, the geographic location information comprises spatial coordinates in a cartesian coordinate system. For example, the dimensions of the flow cell may be mapped with a cartesian coordinate system, e.g., with x and y dimensions, and locations of clusters on a flow cell may be assigned to a spatial coordinate using this system. The geographic location information can be used to determine relative proximity between each cluster on the flow cell.

[0195] Next, the method 300 can proceed to block 340, wherein sequence reads are mapped to a reference genome. This can be accomplished as described above, for example atblock 230 of the method 200. The sequence reads can be aligned to one or more reference genomes to obtain a genomic location of the sequence reads. The alignment can be accomplished by a local alignment method or by any other method.

[0196] Next, the method 300 can proceed to block 350, wherein links between sequence reads are determined based on mapping and flow cell locations. This can be accomplished as described above, for example at block 240 of the method 200. For example, the linking information can include a probability that two read pairs are derived from the same long (over 500 bp) fragment from the DNA sample that was flowed across the flow cell. If the probability is over a predetermined threshold, the two read pairs can be determined to be linked. As a further example, linking probabilities of all sequence reads within a bin can be summed to give the expected number of link counts given the probabilistic model. The probabilistic model can be generated using a statistical model approach, e.g. maximum likelihood estimation, or using a ML method e.g. neural networks or boosting trees.

[0197] Next, the method 300 can proceed to block 360, wherein linking information is plotted on a colocation plot. This can be accomplished as described above, for example at block 250 of the method 200. For example, in some embodiments, the colocation plot has x and y coordinates corresponding to bins along reference genome. In some embodiments, the colocation plot provides a visualization of the number of links between bins on the reference genome. Bins can be sized as needed in order to visualize structural variants based on the links.

[0198] In some embodiments, multiple colocation plots are generated and can be used in a plurality of training datasets. In some embodiments, each training dataset further includes a structural variant classification. For example, a training dataset can include a colocation plot and a structural variant classification. A training dataset can include colocation plots representative of the different types of structural variants described herein, as well as those types of structural variants known to those skilled in the art, including but not limited to deletions, duplications, inversions, insertions, translocations, or other chromosomal rearrangements, or a combination of multiple structural variants in a region. A dataset may comprise curated types of colocation plots representative of the different types of structural variants to improve efficiency of the training of the machine learning model as well as speedand accuracy of detection of structural variants from data received from sequencing systems employing the methods described herein (or similar).

[0199] While the method 300 includes the steps at blocks 320, 330, 340, 350, and 360 to generate colocation plots, in other embodiments, training datasets can be obtained in other ways. For example, training datasets can include simulated data. For example, in some embodiments, modified genomic sequence flow cell data by is generated from reference genome data by introducing simulated structural variants within reference genome data and generating simulated sequence reads. Linking information can be generated based on the simulated sequence reads. Colocation data, including colocation plots, can be generated based on the simulated linking information. Simulations can also be done simply at the colocation plot level, e.g, by generating simulated colocation plot images with simulated features corresponding to a structural variant. Such an embodiment would not require structural variants to first be simulated into a reference sample.

[0200] In some embodiments, homozygous and / or heterozygous simulated structural variants are simulated. For example, in some embodiments, modified genomic sequence flow cell data is generated by introducing synthetic structural variants into a diploid genome sample, and simulated sequence reads are generated from the diploid genome sample.

[0201] Furthermore, while the exemplary method 300 relates to colocation plots, colocation data can be used in other formats and can be used to train a machine learning model. For example, the colocation data may be represented with a graph as further described herein. For example, each bin can be represented as a node, with edges connecting the bins representing linking information.Training the model with the plurality of training datasets

[0202] The method 300 can then proceed to block 370, wherein a model is trained on one or more colocation plots to identify patterns corresponding to structural variants. In some embodiments, the machine learning model is trained to detect structural variants based on each colocation plot and its corresponding structural variant classification.

[0203] For example, in some embodiments, training datasets with colocation data and corresponding structural variant classifications are provided to a machine learning model. For example, the machine learning model may be trained to recognize the visual signaturesfound on colocation plots from reference genomes having known structural variants. For example, in one embodiment, specific structural variants are introduced into a reference genome such as the HG002 complete diploid human genome sample. As discussed above, a colocation map can be generated based on the linked sequence reads from the modified reference genome, and the colocation map may be used to train the neural network to recognize the specific visual patterns caused by the specific structural variants introduced into the HG002 genome. This process can be repeated using a variety of different specific structural variants until the neural network is trained to recognize the visualizations within a colocation map corresponding to each specific structural variant.

[0204] In some embodiments, the machine learning model is any of the machine learning models discussed above with reference to block 260 of the method 200.

[0205] In some embodiments, additional information is also provided to the machine learning model for training. For example, additional sequence read information beyond colocation data can be provided alongside the colocation data. In some embodiments, the additional information is as described with respect to block 260 of the method 200.

[0206] The method 300 can then proceed to decision state 380, wherein the system queries whether there are additional colocation plots available. If yes, the method 300 can return to block 370 and proceed as previously described. Thus, in some embodiments, the machine learning model can iteratively update itself as it processes additional colocation plots, for example by updating the patterns it associates as corresponding to structural variants. If there are no further colocation plots available, the method 300 can return to proceed to end block 390.Colocation Plots

[0207] In some embodiments, colocation data includes a colocation plot. For example, in some embodiments, linking information is represented as an image, for example, as a colocation plot. In some embodiments, colocation plots visualize the number of links between genomic regions of a sample.

[0208] FIG. 4 depicts how linking information can be visualized with a colocation plot. The x and y axes are successive genomic bins. Intensity on the heatmap corresponds to the number of sequence reads that are proximal to each other on the flowcell from the tworeference bins in the colocation matrix. The length 401 of the bright perpendicular across the diagonal is representative of and proportional to link strength. In FIG. 4, the length 401 of the diagonal representing the contiguous base sequence shows a strong pigment contrast, because sequence reads in neighboring bins have both flow cell proximity and genomic proximity. There is a reduction in connections to more distant genomic regions, which are not proximate on the flow cell to one another. There are no large SV events in this region. Although there appears to be a structural variant deletion at 402, this is not the case. Due to mapping difficulties and undetermined signals (e.g., noise) in linking information, this would be a false positive. Further processing may be needed to distinguish a true structural variant from a false positive caused by undetermined signals.

[0209] FIG. 5 depicts a colocation plot showing an F8 gene inversion (>500 kb SV) from a real sample. There are interruptions 510 and 511 in the diagonal and bowtie-shaped features 520 and 521 off the diagonal due to the gene inversion. The gene inversion causes links between regions that are genomically distant on the reference genome, but in the sample, the regions are near one another due to the gene inversion and thus have both genomic proximity and flow cell proximity. This difference leads to the visual artifacts shown in FIG. 5. FIG. 5 also includes a schematic representation of the F8 gene inversion.

[0210] In some embodiments, the machine learning model processes linking information between neighboring bins of the reference genome. Because neighboring bins are shown as a diagonal in colocation plots, this is referred to herein as detecting structural variants “on-diagonal.” In some embodiments, on-diagonal processing deletions in the DNA sample. In some embodiments, the machine learning model processes linking information between non-neighboring bins of the reference genome, referred to herein as “off-diagonal.” In some embodiments, off-diagonal processing enables matching of paired breakpoints for insertions, deletions, and other structural variants.Object Detection Networks and Non-Max Suppression

[0211] Embodiments herein relate to object detection networks. Various types of object detection networks can be used in the methods and systems of the present disclosure. An object detection network may be neural or non-neural. Examples of non-neural approaches include a Viola-Jones object detection framework based on Haar features, a scale-invariantfeature transform (SIFT), a histogram of oriented gradients (HOG) features. Examples of neural network approaches include OverFeat, Region Proposals (R-CNN, Fast R-CNN, [Faster R-CNN, cascade R-CNN), You Only Look Once (YOLO), Single Shot MultiBox Detector (SSD), Single-Shot Refinement Neural Network for Object Detection (RefineDet), Retina-Net, and deformable convolutional networks.

[0212] Components of an exemplary object detection network are illustrated in FIG. 6A. Briefly, in the exemplary object detection network of FIG. 6A, the flow can include the following steps: an input 610 where an image is fed into the network; a backbone 620 which extracts and transforms image features; a neck 630 which refines and combines these features to improve detection quality; and a head 640 which processes the refined features to output object classes and bounding box coordinates.

[0213] The input 610 is the initial stage where the image or video frame is fed into the network. This is the raw data that the network will process to detect objects (e.g., structural variants). The input 610 stage can involve preprocessing steps such as resizing, normalization, and in some cases data augmentation to prepare the image for analysis.

[0214] The backbone 620 is responsible for extracting features from the input image. A deep convolutional neural network (CNN), for example, ResNet, VGG, or EfficientNet, may be used to transform the raw pixel values into a set of high-level features that represent various patterns and structures in the image. The backbone 620 acts as the feature extractor, converting the image into a more abstract representation that is easier for subsequent layers to analyze.

[0215] The neck 630 connects the backbone to the head of the network and helps aggregate and refine the features extracted by the backbone. It can include additional layers or modules such as Feature Pyramid Networks (FPNs) or Attention Mechanisms. The goal of the neck 630 is to enhance the feature representations by combining features from different layers of the backbone, which helps in detecting objects at various scales and improving the overall performance of the network.

[0216] The head 640 is the final component responsible for generating the actual object detection outputs. The head 640 can include several sub-components, including a classification head 641, which predicts the class of each detected object (e.g., insertion, deletion, inversion), a localization head 642 (e.g., bounding box regression head), whichrefines the coordinates of the bounding boxes around detected objects to accurately localize them, and other components for additional tasks such as object segmentation or keypoint detection.

[0217] In some embodiments, the machine learning model comprises a two-stage object detection network. Two-stage object detection network can involve a two-step process to identify and locate objects within an image. For example, in some embodiments, the backbone 620 is used to extract features from an image, which are input to the two-stage object detection network.

[0218] In some embodiments, the two-stage object detection network includes two phases: proposal generation and refinement. During the first stage, a proposal generator generates candidate regions (proposals) where structural variants might be located. This can be accomplished using a Region Proposal Network (RPN) or a similar module. The RPN scans the feature map produced by the backbone network to identify regions of interest (ROIs) that potentially contain structural variants. The RPN can use anchor boxes (predefined bounding boxes of various sizes and aspect ratios) to propose potential structural variant locations. The RPN can predict the likelihood of each anchor box containing a structural variant and refine the box coordinates to better fit the structural variant.

[0219] The second stage includes object classification and bounding box refinement. In some embodiments, a box classifier extracts the proposed regions from the RPN from the feature map and processes the proposed regions by an ROI pooling or ROI align layer. This step converts the proposals into a fixed-size feature map suitable for further processing. The extracted features are fed into the classification head 641 to determine the object class (e.g., insertion, deletion, inversion) and a localization head 642 to refine the coordinates of the bounding boxes. The classification head 641 assigns class probabilities to each proposed region, while the localization head 642 adjusts the box coordinates to better fit the detected object.

[0220] Advantages of a two-stage object detection network include accuracy, as two-stage networks generally achieve high accuracy in detecting objects because they separate the tasks of proposal generation and object classification / regression, allowing for more precise detections. The separation into two stages allows for better refinement of object locations and class labels compared to single-stage approaches.

[0221] In some embodiments, non-max suppression may be used to remove redundant overlapping bounding boxes. Non-Maximum Suppression (NMS) refers to a postprocessing step in object detection networks that helps refine the final detection results. NMS eliminates redundant bounding boxes that overlap significantly, ensuring that each object is represented by only one bounding box. In some embodiments, NMS is applied after the object detection network has made predictions. In some embodiments, NMS occurs after the localization head has predicted multiple bounding boxes and their corresponding confidence scores, and after the classification head has determined the class probabilities for those bounding boxes.

[0222] FIG. 6B depicts an example of a colocation plot showing various object detection regions generated from a machine learning model. As indicated, there are overlapping boundary boxes around putative structural variants found on the colocation plot. In the example of FIG. 6B the machine learning model has detected many duplicate SVs overlapping the true structural variant(s).

[0223] For example, in some embodiments, the network outputs a set of bounding boxes along with associated class scores and confidence scores for each predicted box. In some embodiments, the predicted bounding boxes are first sorted by their confidence scores or probability scores in descending order, so that boxes with the highest confidence are processed first. Starting with the highest confidence bounding box, NMS compares it with all other bounding boxes. It calculates the Intersection over Union (loU) between the highest- confidence box and other boxes. If the loU between the highest-confidence box and another box exceeds a predefined threshold (e.g., 0.5), the lower-confidence box is discarded because it is considered redundant. This step is repeated iteratively for each remaining box. The NMS process is repeated until all bounding boxes have been processed, leaving only one bounding box per object. The remaining bounding boxes, which are the ones with the highest confidence and least overlap, are considered the final detections and are used for reporting the results.

[0224] In some embodiments, bounding boxes are clustered based off anchor locations and / or event type (e.g., insertion, deletion, inversion). In some embodiments, the machine learning model selects the highest confidence SV box per cluster.

[0225] In some embodiments, non-max suppression is applied across multiple colocation plot images. For example, because variant calling concerns variant calls across thegenome and not just a single image, a single variant could be called from multiple images. Thus, in some embodiments, NMS is applied across multiple images. For example, the methods and systems can use a global genomic coordinate system to identify objects that are represented on overlapping images and reconcile objects detected across multiple images to avoid extra reports, e.g. by processing overlapping images together. In some embodiments, detected events on different images are the same event, so the methods and systems can use NMS to suppress extra SV calls corresponding to the same event.Structural Variant Visualizations and User Interface

[0226] In some embodiments, the methods and systems display a visualization of a colocation plot on a user interface. For example, in some embodiments, the systems generate a colocation plot according to the methods described herein and display the colocation plot via a user interface. In some embodiments, the methods and systems display an indication of a detected structural variant region on a user interface. In some embodiments, the system is configured to receive a selection from a user, and to adjust the parameters of visualization window for the colocation plot. In some embodiments, the methods and systems are configured to update a colocation plot visualization in response to the selection.

[0227] In some embodiments, the user interface includes further visualizations, for example, a structural variant call plot, a short read plot, or other visualizations related to sequence read data. In some embodiments, the colocation plot and the structural variant call plot, or the short read plot, or any other plot, may interact with each other in response to user input. For example, in response to receiving a selection from a user in a structural variant call plot, the system can display an indication of a genomic region of a structural variant in the colocation plot. As another example, in response to receiving a selection from a user in a colocation plot, the system can display an indication of a genomic region of a structural variant in the structural variant plot. As another example, in response to user input in a colocation plot, the system can display an indication of a genomic region in the short read plot.

[0228] In some embodiments, genomic regions of the colocation plot are aligned with genomic regions of other plots such as the read visualization, variant call visualization and other visualizations. In some embodiments, variant call information is overlaid on thecolocation plot. In some embodiments, a visualization of a chromosome is overlaid on the colocation plot.

[0229] In some embodiments, the user interface includes an option where in response to a user selection, the system can show / hide additional information such as read visualization, variant call visualization, structural variant visualization, or other visualizations, overlaying the colocation plot visualization.

[0230] In some embodiments, the user interface can display a filtering option to filter a visualization by variant type. For example, the system can add, remove, or highlight a specified type of structural variant based on a user selection.

[0231] In some embodiments, the methods disclosed herein include displaying, with a graphical user interface, a first colocation plot showing the predicted structural variant regions. In some embodiments, displaying the first colocation plot comprises displaying graphical indicia indicating the positions of the predicted structural variant regions on the display. In some embodiments, the method further comprises in response to receiving a user selection, displaying a magnified view of the predicted structural variant region. In some embodiments, after receiving a user selection to increase or decrease a magnification of the colocation plot zooming, the methods and systems generate and display a visualization with an increased or decreased bin size. In some embodiments, the visualization includes a bin level indicator based upon magnification level. For example, upon receiving a selection from a user to increase the magnification, the methods and systems can generate a colocation plot visualization with a decreased bin size, and display a bin level indicator that shows that the bin size has decreased.

[0232] Further disclosed herein are systems for detecting a structural variant which include a user interface configured to display one or more first colocation plots based on linking information from genomic nucleic acids, and a detection refinement option; and a processor configured to, in response to receiving a selection of the detection refinement option, generate one or more second colocation plots wherein the one or more second colocation plot images are at a different magnification or resolution than the one or more first colocation plot images. In some embodiments, the processor is further configured to provide the one or more second colocation plot images to a trained machine learning model and process the one or moresecond colocation plot images with the machine learning model to identify a structural variant within the genomic sample.

[0233] FIG. 7A depicts an exemplary colocation plot 700. The colocation plot 700 has x and y axes corresponding to genomic bins along the X chromosome. Count of sequence reads in one bin along the x axis with sequence reads in the bin along the y axis is shown in intensity.

[0234] FIG. 7B depicts an exemplary visualization 705 of a colocation plot 700 in a user interface. In FIG. 7B, the visualization 705 was generated by rotating the colocation plot 700 to the left by 45 degrees and cropping along the diagonal line. The visualization 705 is displayed next to a visualization of a chromosome 710, to show colocation intensity along the chromosome. In the example of FIG. 7B, the chromosome 710 is an X chromosome.

[0235] In some embodiments, a user can select, via the user interface, a window of the visualization 705 for generating a further visualization. For example, the user can select a smaller window 715 within the colocation plot 700 for further visualization. In some embodiments, the methods and systems can receive the selection. In some embodiments, in response to receiving the selection, the methods and systems can display a further visualization of the colocation plot 700.

[0236] FIG. 7C depicts an exemplary visualization based on receiving a selection of the window 715. Visualization 720 is a zoomed-in view of the window 715. Legend 722 includes symbols that depict what different structural variants such as a translocation, inversion, deletion, insertion, or tandem duplication, may look like on a colocation plot. In the example of FIG. 7C, the user interface also includes other visualizations of other genomic information such as genes visualization 723, protein domains visualization 725, structural variant visualization 727, and sequence reads and depth of coverage visualization 729. The structural variant 727 includes indication of a tandem duplication event that was detected by the methods and systems based on the colocation plot 700, for example by machine learning as described herein.

[0237] FIG. 7D depicts a further exemplary visualization. Visualization 730 is a visualization of another section of the colocation plot 700. In the example of FIG. 7D, the methods and systems receive a selection from a user via the display 731, where a user can click and drag the window 735 to cover an area of the colocation plot 705 corresponding to updatedgenomic coordinates. In response to receiving the selection from user, the methods and systems display the visualization 730. The visualization 730 corresponds to the updated genomic coordinates, shown in coordinate display 732. The methods and systems can similarly update the genes visualization 723, the protein domains visualization 725, the structural variant visualization 727, and the sequence reads and depth of coverage visualization 729 in response to receiving the selection from the user via the display 731.

[0238] FIG. 7E depicts a further exemplary visualization. In the example of FIG. 7E, the methods and systems receive a selection from a user via the zoom option 745. In response to receiving the selection from user, the methods and systems display the visualization 740, which is a zoomed-in view of the colocation plot 700 corresponding the user’s selection. The visualization 740 corresponds to the updated genomic coordinates, shown in coordinate display 732. The display 731 indicates the window of the colocation plot 700 that is being shown in visualization 740. The methods and systems can similarly update the gene visualization 723, the protein domain visualization 725, the structural variant visualization 727, and the sequence reads and depth of coverage visualization 729 in response to receiving the selection from the user via the display 731.Electronic Files

[0239] The methods and systems described herein can include a step of storing information generated by the methods and systems, in an electronic file, or retrieving information from an electronic file. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, BED, FASTQ, SAM, CRAM, JSON, CIGAR, HDF5, or VCF file.

[0240] In some embodiments, the electronic file includes information related to a co-proximity property, e.g., links as described herein, where reads that are nearby within an individual’s genome are also nearby on the flow cell surface. This property enables improved mapping, variant calling, and phasing, among other enhancements. The following is a list of terms that are used to describe data in Tables 1 and 2 below.

[0241] Template molecule: A template molecule is a long DNA molecule (e.g., at least 500 bp) from either a standard or high molecular weight extraction that is loaded into the sequencing cartridge for on-flow cell tagmentation.

[0242] Proximal reads: Proximal reads are reads from the same template molecule that are from clusters located near one another on the flow cell surface.

[0243] Proximity group: A proximity group is a set of reads that are likely to have been derived from the same template molecule.

[0244] Flow cell proximity: The flow cell proximity is the distance between DNA nanowells of reads within a proximity group, measured either in common units of measurement including nanometers, millimeters, inches, or in flow cell units. Genomic proximity: Genomic proximity is the genomic distance between reads within a proximity group.

[0245] Template length: Genomic distance (e.g., distance on a reference genome, measured in bp) spanned by a proximity group from the same original template.

[0246] Co-proximity rate: Percentage of all reads that are co-proximal with at least one other read.

[0247] Co-proximity quality: A Phred-scaled quality score that estimates the probability that two reads derive from the same original template molecule.

[0248] In some embodiments, the method includes a step of storing sequence reads or other sequence information in a BAM file or a VCF file. In some embodiments, the BAM file has tags and / or VCF file has fields as described below. In some embodiments, metrics for these tags are stored in a separate CSV file with the definitions provided.

[0249] Table 1 describes new tags for BAM files. In some embodiments, the BAM file also includes other tags, for example debug tags (xn, xp, xf, xa, xb, Id, ad, bd, ws, wm, and pq)-Table 1 : New BAM tags

[0250] In some embodiments, the VCF file includes fields related to variant calling and haplotype information (for example, Whatshap phase).

[0251] In some embodiments, the method includes creating an electronic file that includes one or more of the following metrics:Table 2EXAMPLES

[0252] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims.Example 1

[0253] FIG. 8A depicts an exemplary colocation plot with a homozygous insertion event. Genomic positions are listed by bin along the x and y axes, and count of links between sequence reads in each bin are shown. FIG. 8B depicts an exemplary colocation plot with a homozygous deletion event.Example 2

[0254] FIG. 9 depicts active region detection.. A first colocation plot 910 is generated at coarse resolution with bins that are over 10,000 kbp in size. The first colocation plot 910 is processed by a machine learning model and an active region 930 is detected in area 920. The methods and systems generate a second colocation plot (not shown) for processing with a downstream object detection model. The second colocation plot covers the active region 930 at higher resolution using smaller bins.Example 3

[0255] In the following example, an object detection machine learning model was trained and was used to detect structural variants based on a colocation plot. The machine learning model also determined a confidence score. The machine learning model’s predictions of structural variants was compared to ground truth knowledge of structural variants.

[0256] The structural variant detections are shown in FIG. 10 panels a) - h) and FIG. 11 panels a) - d). The machine learning model detected insertions (INS), deletions (DEL), and inversions (INV).Additional Notes

[0257] Various embodiments of the present disclosure may be a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or mediums) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0258] For example, the functionality described herein may be performed as software instructions are executed by, and / or in response to software instructions being executed by, one or more hardware processors and / or any other suitable computing devices. The software instructions and / or other executable code may be read from a computer readable storage medium (or mediums). Computer readable storage media may also be referred to herein as computer readable storage or computer readable storage devices.

[0259] The computer readable storage medium can be a tangible device that can retain and store data and / or instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device (including any volatile and / or non-volatile electronic storage devices), a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a solid state drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0260] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storagemedium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0261] Computer readable program instructions (as also referred to herein as, for example, “code,” “instructions,” “module,” “application,” “software application,” and / or the like) for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language, such as Smalltalk, C++, or the like, and procedural programming languages, such as the "C" programming language or similar programming languages. Computer readable program instructions may be callable from other instructions or from itself, and / or may be invoked in response to detected events or interrupts. Computer readable program instructions configured for execution on computing devices may be provided on a computer readable storage medium, and / or as a digital download (and may be originally stored in a compressed or installable format that requires installation, decompression or decryption prior to execution) that may then be stored on a computer readable storage medium. Such computer readable program instructions may be stored, partially or fully, on a memory device (e.g., a computer readable storage medium) of the executing computing device, for execution by the computing device. The computer readable program instructions may execute entirely on a user's computer (e.g., the executing computing device), partly on the user’ s computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, throughthe Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0262] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0263] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart(s) and / or block diagram(s) block or blocks.

[0264] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer may load the instructions and / or modules into its dynamic memory and send the instructions over a telephone, cable, or optical line using a modem. A modem local to aserver computing system may receive the data on the telephone / cable / optical line and use a converter device including the appropriate circuitry to place the data on a bus. The bus may carry the data to a memory, from which a processor may retrieve and execute the instructions. The instructions received by the memory may optionally be stored on a storage device (e.g., a solid-state drive) either before or after execution by the computer processor.

[0265] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a service, module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. In addition, certain blocks may be omitted in some implementations. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states relating thereto can be performed in other sequences that are appropriate.

[0266] It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions. For example, any of the processes, methods, algorithms, elements, blocks, applications, or other functionality (or portions of functionality) described in the preceding sections may be embodied in, and / or fully or partially automated via, electronic hardware such application-specific processors (e.g., application-specific integrated circuits (ASICs)), programmable processors (e.g., field programmable gate arrays (FPGAs)), application-specific circuitry, and / or the like (any of which may also combine custom hard-wired logic, logic circuits, ASICs, FPGAs, etc. with custom programming / execution of software instructions to accomplish the techniques).

[0267] Any of the above-mentioned processors, and / or devices incorporating any of the above-mentioned processors, may be referred to herein as, for example, “computers,”“computer devices,” “computing devices,” “hardware computing devices,” “hardware processors,” “processing units,” and / or the like. Computing devices of the above-embodiments may generally (but not necessarily) be controlled and / or coordinated by operating system software, such as Mac OS, iOS, Android, Chrome OS, Windows OS (e.g., Windows XP, Windows Vista, Windows 7, Windows 8, Windows 10, Windows 11, Windows Server, etc.), Windows CE, Unix, Linux, SunOS, Solaris, Blackberry OS, VxWorks, or other suitable operating systems. In other embodiments, the computing devices may be controlled by a proprietary operating system. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file system, networking, I / O services, and provide a user interface functionality, such as a graphical user interface (“GUI”), among other things.

[0268] Reference throughout the specification to “one example”, “another example”, “an example”, and so forth, means that a particular element (e.g., feature, structure, and / or characteristic) described in connection with the example is included in at least one example described herein, and may or may not be present in other examples. In addition, it is to be understood that the described elements for any example may be combined in any suitable manner in the various examples unless the context clearly dictates otherwise.

[0269] It is to be understood that the ranges provided herein include the stated range and any value or sub-range within the stated range, as if such value or sub-range were explicitly recited. For example, a range from about 2 kbp to about 20 kbp should be interpreted to include not only the explicitly recited limits of from about 2 kbp to about 20 kbp, but also to include individual values, such as about 3.5 kbp, about 8 kbp, about 18.2 kbp, etc., and sub-ranges, such as from about 5 kbp to about 10 kbp, etc.

[0270] While several examples have been described in detail, it is to be understood that the disclosed examples may be modified. Therefore, the foregoing description is to be considered non-limiting.

[0271] While certain examples have been described, these examples have been presented by way of example only, and are not intended to limit the scope of the disclosure. Indeed, the novel methods described herein may be embodied in a variety of other forms. Furthermore, various omissions, substitutions and changes in the methods described herein may be made without departing from the spirit of the disclosure. The accompanying claimsand their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the disclosure.

[0272] Features, materials, characteristics, or groups described in conjunction with a particular aspect, or example are to be understood to be applicable to any other aspect or example described in this section or elsewhere in this specification unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The protection is not restricted to the details of any foregoing examples. The protection extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.

[0273] Furthermore, certain features that are described in this disclosure in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations, one or more features from a claimed combination can, in some cases, be excised from the combination, and the combination may be claimed as a sub-combination or variation of a sub-combination.

[0274] Moreover, while operations may be depicted in the drawings or described in the specification in a particular order, such operations need not be performed in the particular order shown or in sequential order, or that all operations be performed, to achieve desirable results. Other operations that are not depicted or described can be incorporated in the example methods and processes. For example, one or more additional operations can be performed before, after, simultaneously, or between any of the described operations. Further, the operations may be rearranged or reordered in other implementations. Those skilled in the art will appreciate that in some examples, the actual steps taken in the processes illustrated and / or disclosed may differ from those shown in the figures. Depending on the example, certain of the steps described above may be removed or others may be added. Furthermore, the featuresand attributes of the specific examples disclosed above may be combined in different ways to form additional examples, all of which fall within the scope of the present disclosure.

[0275] For purposes of this disclosure, certain aspects, advantages, and novel features are described herein. Not necessarily all such advantages may be achieved in accordance with any particular example. Thus, for example, those skilled in the art will recognize that the disclosure may be embodied or carried out in a manner that achieves one advantage or a group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein.

[0276] Conditional language, such as “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain examples include, while other examples do not include, certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required for one or more examples or that one or more examples necessarily include logic for deciding, with or without user input or prompting, whether these features, elements, and / or steps are included or are to be performed in any particular example.

[0277] Conjunctive language such as the phrase “at least one of X, Y, and Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to convey that an item, term, etc. may be either X, Y, or Z. Thus, such conjunctive language is not generally intended to imply that certain examples require the presence of at least one of X, at least one of Y, and at least one of Z.

[0278] Language of degree used herein, such as the terms “approximately,” “about,” “generally,” and “substantially” represent a value, amount, or characteristic close to the stated value, amount, or characteristic that still performs a desired function or achieves a desired result. In some embodiments these terms encompass minor variations (for example, up to + / - 10%) from the stated value or amount.

[0279] The scope of the present disclosure is not intended to be limited by the specific disclosures of preferred examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the presentspecification or during the prosecution of the application, which examples are to be construed as non-exclusive.

Claims

WHAT IS CLAIMED IS:

1. A method for detecting a structural variant in DNA from a sample placed on a flow cell, comprising: obtaining flow cell data from the DNA, wherein the flow cell data comprises nucleic acid sequence reads and flow cell locations of sequence read clusters on the flow cell; aligning the sequence reads to one or more reference genomes to obtain a genomic location of the sequence reads; obtaining linking information between pairs of sequence reads on the flow cell based on the flow cell location and genomic location of each sequence read in the pairs of sequence reads; generating colocation data based on the linking information; providing the colocation data to a machine learning model; and processing the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acids and determining a confidence score, thereby detecting at least one structural variant for the sample.

2. The method of claim 1 , wherein generating colocation data based on the linking information comprises: dividing the alignment of sequence reads to the reference genome into a plurality of bins; and counting an estimated number of links between pairs of sequence reads by bin.

3. The method of claim 2, wherein the bins are between 500 bp to 100 kbp in size.

4. The method of any one of claims 1-3, wherein the colocation data comprises a colocation plot image or a graph representation of colocation data.

5. The method of any one of claims 1-4, wherein the machine learning model comprises an object detection model, image classification model, an anomaly detection model, a deep learning image processing model, a deep convolutional neural network (CNN), graph neural network (GNN), or a transformer-based model.

6. The method of any one of claims 1-5, wherein the method comprises applying non-max suppression to a colocation plot image.

7. The method of claim 6, wherein the method comprises applying non-max suppression across a plurality of combined colocation plot images.

8. The method of any one of claims 1 -7, wherein processing the image comprises: identifying a putative structural variant region based on a first colocation data; generating a second colocation data based on the putative structural variant, wherein the second colocation data is at a higher or lower resolution than the first colocation data; providing the second colocation data to the machine learning model; and processing the second colocation data with the machine learning model to identify a structural variant within the genomic nucleic acids.

9. The method of claim 8, wherein the second colocation data is based on a larger or smaller bin size as compared to the first colocation data.

10. The method of claim 8, wherein the first colocation data is based on bins of 10 kbp to 100 kbp in size, and wherein the second colocation data is based on bins of 500 bp to 10 kbp in size.

11. The method of claim 8, wherein the first colocation data is based on bins of 500 bp to 10 kbp in size, and wherein the second colocation data is based on bins of 10 kbp to 100 kbp in size.

12. The method of claim 8, wherein the method comprises displaying the putative structural variant region with a user interface, and wherein the putative structural variant region is selected via the user interface.

13. The method of any one of claims 1-12, wherein the method comprises flowing nucleic acid molecules greater than 500 bp in length across a flow cell and fragmenting the nucleic acid molecules on the flow cell.

14. The method of any one of claims 1-13, wherein the method comprises providing a sequence read depth, a mapping quality measurement, GC content bias measurement, an initial CNV / structural variant prediction, abnormal paired read information, a link quality score, a genomic distance from alignment, split read information, phasing information, linking information statistical parameters, or a comparison to one or more normal samples, to the machine learning model.

15. The method of any one of claims 1-14, wherein identifying a predicted structural variant region within the genomic nucleic acids comprises determining one or more of: an area with a suspected SV, a structural variant event class, a structural variant position, a structural variant anomaly region corresponding to a complex or overlapping structural variant, and a structural variant size.

16. The method of claim 15, wherein the structural variant event class comprises an insertion, a deletion, an inversion, or a translocation..

17. The method of claim any one of claims 1-16, further comprising storing structural variant information in an electronic file.

18. The method of claim 17, wherein the method comprises storing a set of detected structural variants in a VCF file and a set of anomaly regions corresponding to complex or overlapping structural variants in a BED file.

19. The method of claim any one of claims 1-18, further comprising: displaying, with a graphical user interface, a first colocation plot showing a predicted structural variant region or a bounding box for a predicted structural variant region.

20. The method of claim 19, wherein displaying the first colocation plot comprises displaying graphical indicia indicating a position of a predicted structural variant region or anomalous region on the display.

21. The method of claim 20, wherein the graphical indicia includes a highlight or change in color.

22. The method of claim 20, further comprising in response to receiving a user selection, displaying a magnified view of the predicted structural variant region.

23. The method of any one of claims 1-22, wherein the method further comprises updating a mapping location of a sequence read based on linking information.

24. A method for training a machine learning model to detect structural variants in a genomic sequence, comprising: obtaining a plurality of training datasets, wherein each training dataset comprises i) a structural variant classification and ii) linking information between pairs of sequence reads based on their geographic location on the flow cell;generating colocation data for each training dataset based on the linking information; providing each training dataset to a machine learning model; and training the machine learning model to detect structural variants based on each colocation plot data and its corresponding structural variant classification.

25. The method of claim 24, wherein the method comprises generating modified genomic sequence flow cell data by introducing simulated structural variants within reference genome data and generating simulated sequence reads.

26. The method of claim 24 or claim 25, wherein the method comprises generating modified genomic sequence flow cell data by introducing synthetic structural variants into a diploid genome sample, and generating simulated sequence reads from the diploid genome sample.

27. The method of any one of claims 24 - 26, wherein the method comprises providing a sequence read depth, a mapping quality measurement, GC content bias measurement, an initial CNV / structural variant prediction, abnormal paired read information, a link quality score, a genomic distance from alignment, split read information, phasing information, linking information statistical parameters, or a comparison to one or more normal samples, to the machine learning model.

28. A system for classifying a structural variant condition for a subject, comprising: one or more processors that are programmed to execute a method comprising: obtaining flow cell data from sample DNA, wherein the flow cell data comprises a nucleotide sequence read, and flow cell location of sequence read clusters on the flow cell; aligning sequence reads to one or more reference genomes to obtain the genomic location of each sequence read using the flow cell location of the sequence reads, thereby obtaining linking information between pairs of sequence reads based on flow cell location and genomic location; generating colocation data based on the linking information; providing the colocation data to a machine learning model; and processing the colocation data with the machine learning model to identify a predicted structural variant region within the genomic nucleic acidsand determining a confidence score, thereby detecting at least one structural variant for the sample.

29. A system for training a machine learning model to detect structural variants in a genomic sequence, comprising: one or more processors that are programmed to execute a method comprising: obtaining a plurality of training datasets, wherein each training dataset comprises i) structural variant classification and ii) linking information between pairs of sequence reads based on their geographic location on the flow cell; generating colocation data for each training dataset based on the linking information; providing each colocation data and its corresponding structural variant condition to a machine learning model; and training the machine learning model to detect structural variants based on each colocation plot data and its corresponding structural variant classification.

30. A system for detecting a structural variant, comprising: a user interface configured to display one or more first colocation plots based on linking information from genomic nucleic acids, and a detection refinement option; and a processor configured to, in response to receiving a selection of the detection refinement option, generate one or more second colocation plots wherein the one or more second colocation plot images are at a different magnification or resolution than the one or more first colocation plot images.

31. The system of claim 30, wherein the processor is further configured to provide the one or more second colocation plot images to a trained machine learning model and process the one or more second colocation plot images with the machine learning model to identify a structural variant within the genomic nucleic acids.

Citation Information

Patent Citations

  • Polymerase enzymes and reagents for enhanced nucleic acid sequencing

    US20080108082A1

  • Transposon end compositions and methods for modifying nucleic acids

    US20100120098A1

  • Methods and apparatus for measuring analytes

    US20100137143A1

  • Modified Molecular Arrays

    US20110059865A1

  • Determining a mudweight of drilling fluids for drilling through naturally fractured formations

    US62636004P0