Systems and methods for identifying somatic structural variants
The method of extracting and grouping soft-clipped reads with k-mer matching in sequencing data effectively addresses the challenge of identifying somatic structural variants, enhancing diagnostic and treatment capabilities for diseases like cancer.
Patent Information
- Application Number
- PCT/US2025/043101
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-23
- Filing Date
- 2025-08-22
- Publication Date
- 2026-02-26
AI Technical Summary
Existing systems face challenges in accurately identifying somatic structural variants in sequencing data due to computational intensity and error-prone processes, particularly in sequencing by expansion (SBX) data, which lacks discordant reads and may not provide sufficient data for somatic structural variant detection.
A method involving soft-clipped read extraction, grouping by breakpoints, and k-mer matching is employed to identify somatic structural variants, utilizing sequencing data from processes like targeted, whole genome, or whole transcriptome sequencing, with error correction and secondary alignment information to enhance accuracy.
This approach improves the identification of somatic structural variants, particularly those associated with cancer, by reducing errors and enhancing computational efficiency, enabling better diagnosis, prognosis, and treatment strategies.
Smart Images

Figure US2025043101_26022026_PF_FP_ABST
Abstract
Description
PATENT Client Reference No.: P39512-WO-1 SYSTEMS AND METHODS FOR IDENTIFYING SOMATIC STRUCTURAL VARIANTS CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This Application claims priority to United States Provisional Patent Application No.63 / 686,484, filed on August 23, 2024, which is herein incorporated by reference in its entirety. BACKGROUND
[0002] Identifying genetic mutations using sequencing data is a crucial task in bioinformatics that has implications in diagnosis, prognosis, and treatment of multiple diseases including cancer. Specifically, molecular sequencing of nucleic acid material (e.g., deoxyribonucleic acid (DNA) molecules or ribonucleic acid (RNA) molecules), or analogues derived therefrom, can be used to identify variations in the sequence from a reference genome. These genetic variants can include single nucleotide polymorphisms (SNPs), insertions and deletions (InDels), and structural variations (SVs). SVs are larger scale variations than the other types of variants, typically involving sequences of nucleotides of 50 base pairs or more, and can include insertions, deletions, duplications, inversions, and translocations.
[0003] Genetic variants may also come in two forms: germline variants and somatic variants. Germline variants occur in reproductive cells (e.g., sperm and eggs) and can be passed down from parents to offspring. Germline variants are found in every cell of an organism’s tissue and can be inherited across generations. In contrast, somatic structural variants are variations in regions of DNA that can occur during cell development and aging and, therefore, are not passed down from parents to offspring. Somatic variants may be present in only some tissues of an organism.
[0004] Somatic structural variants are associated with many types of cancer and other diseases, and identification of somatic structural variants remains a computationally intensive process, which in previous systems was prone to error. Thus, there is a need in the art for improved systems, devices, and methods for identifying somatic structural variants in sequencing data.PATENT Client Reference No.: P39512-WO-1 SUMMARY
[0005] The embodiments described herein relate to systems and methods for identifying somatic structural variants from sequencing data. More particularly, the embodiments described herein related to computer-implemented methods and systems for identifying somatic structural variants based on soft-clipped reads included in the sequencing data.
[0006] In accordance with a first aspect of the present disclosure, a computer- implemented method is provided for identifying somatic structural variants. The method includes: receiving, at a processor, a dataset comprising base pair data output from a sequencing by expansion process; extracting, at the processor, soft-clipped reads from the received dataset; grouping, at the processor, the extracted soft-clipped reads based at least on their respective breakpoints; determining, at the processor, a presence or an absence of a somatic structural variant by k-mer matching the grouped and extracted soft-clipped reads; and providing, at the processor, at least one determined somatic structural variant responsive to determining the presence of the somatic structural variant.
[0007] In at least one embodiment of the first aspect, the base pair data is derived from at least one of targeted sequencing, whole genome sequencing, whole transcriptome sequencing, and whole exome sequencing.
[0008] In at least one embodiment of the first aspect, the method further includes generating the dataset comprising the base pair data by at least one of clustering, consensus, and realignment of data obtained from the output of the sequencing by expansion process.
[0009] In at least one embodiment of the first aspect, the dataset comprising the base pair data is generated based on at least one of the following: mapping the data obtained from the output of the sequencing by expansion process to a reference sequence using unique molecular identifiers (UMIs), clustering the mapped data based on at least one of a start position, end position, or index indicated by the UMIs, and correcting errors based on the clustered mapped data.
[0010] In at least one embodiment of the first aspect, the somatic structural variant comprises at least one of a fusion, translocation, inversion, inter-gene deletion, intra-gene deletion, or a duplication.PATENT Client Reference No.: P39512-WO-1
[0011] In at least one embodiment of the first aspect, the somatic structural variant comprises at least ten (10) base pairs. In other embodiments of the first aspect, the somatic structural variant comprises at least twenty-five (25) base pairs. In yet other embodiments of the first aspect, the somatic structural variant comprises at least fifty (50) base pairs.
[0012] In at least one embodiment of the first aspect, the dataset comprising the base pair data characterizes soft-clipped genomic data.
[0013] In at least one embodiment of the first aspect, the somatic structural variant has an allele frequency that is less than 10 percent.
[0014] In at least one embodiment of the first aspect, the somatic structural variant has an allele frequency that is about 0.1 to about 10 percent.
[0015] In at least one embodiment of the first aspect, providing the at least one determined somatic structural variant comprises generating a dataset indicative of the at least somatic structural variant or displaying a representation of the at least one somatic structural variant. In some embodiments, the representation of the at least one somatic structural variant is displayed in a graphical user interface.
[0016] In at least one embodiment of the first aspect, providing the at least one determined somatic structural variant comprises generating an output file in a variant calling format indicative of a location of one or more of the at least one determined somatic structural variant.
[0017] In at least one embodiment of the first aspect, the method further includes receiving, at the processor, secondary alignment information. Determining the presence or absence of the somatic structural variant is based on the received secondary alignment information. For example, the secondary alignment information is used to identify candidate gene pairs, the candidate gene pairs applied to filter consensus breakpoints input to a k-mer matching process.
[0018] In at least one embodiment of the first aspect, the method further includes generating a prediction of at least one of cancer risk or cancer treatment using the determined at least one somatic structural variant.PATENT Client Reference No.: P39512-WO-1
[0019] In at least one embodiment of the first aspect, determining the presence of the somatic structural variant further comprises comparing the received dataset to a dataset for a reference genome.
[0020] In accordance with a second aspect of the present disclosure, a system is provided for identifying somatic structural variants. The system includes at least one processor configured to: receive a dataset comprising base pair data output from a sequencing by expansion process; extract soft-clipped reads from the received dataset; group the extracted soft-clipped reads based at least on their respective breakpoints; determine a presence or an absence of a somatic structural variant by k-mer matching the grouped and extracted soft- clipped reads; and provide at least one determined somatic structural variant responsive to determining the presence of the somatic structural variant.
[0021] In at least one embodiment of the second aspect, the system may further includes a sequencing instrument. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The detailed description below is set forth with reference to the accompanying figures, which include the following.
[0023] FIG.1 depicts an exemplary diagram for a computer system configured to identify somatic structural variants (SVs), in accordance with at least some embodiments of the present disclosure.
[0024] FIG.2 depicts an exemplary method for identifying somatic SVs , in accordance with at least some embodiments of the present disclosure.
[0025] FIG.3 depicts an exemplary method for identifying somatic SVs, in accordance with at least some embodiments of the present disclosure.
[0026] FIG.4 depicts a block diagram for a computer system for identifying somatic SVs, in accordance with at least some embodiments of the present disclosure. DETAILED DESCRIPTIONPATENT Client Reference No.: P39512-WO-1
[0027] Systems and methods for identifying somatic SVs are provided, such as those that can be useful in the diagnosis, prognosis, and / or treatment of multiple diseases such as cancer based on sequencing data generated by a sequencer instrument. In particular, the sequencing data can be generated by a next generation sequencing (NGS) instrument, such as a nanopore-based sequencer instrument optimized for sequencing-by-expansion (SBX) chemistry. Definitions
[0028] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0029] All of the U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patent applications and non-patent publications referred to in this specification and / or listed in the Application Data Sheet are incorporated herein by reference, in their entirety. Aspects of the embodiments can be modified, if necessary to employ concepts of the various patents, applications and publications to provide yet further embodiments.
[0030] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0031] Although various features of the disclosure can be described in the context of a single embodiment, the features can also be provided separately or in any suitable combination. Conversely, although the present disclosure can be described herein in the context of separate embodiments for clarity, different aspects of the various embodiments described in the present disclosure can also be implemented in a single embodiment. It is to be understood that the present disclosure is not limited to the particular embodiments described herein and as such can vary. Those of skill in the art will recognize that there arePATENT Client Reference No.: P39512-WO-1 variations and modifications of the present disclosure, which are encompassed within its scope.
[0032] All terms are intended to be understood as they would be understood by a person skilled in the art. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure pertains.
[0033] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed.
[0034] When a feature or element is herein referred to as being “on” another feature or element, it can be directly on the other feature or element or intervening features and / or elements may also be present. In contrast, when a feature or element is referred to as being “directly on” another feature or element, there are no intervening features or elements present. It will also be understood that, when a feature or element is referred to as being “connected," “attached” or “coupled” to another feature or element, it can be directly connected, attached or coupled to the other feature or element or intervening features or elements may be present. In contrast, when a feature or element is referred to as being “directly connected," “directly attached” or “directly coupled” to another feature or element, there are no intervening features or elements present. Although described or shown with respect to one embodiment, the features and elements so described or shown can apply to other embodiments. It will also be appreciated by those of skill in the art that references to a structure or feature that is disposed “adjacent” another feature may have portions that overlap or underlie the adjacent feature.
[0035] Terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. For example, as used herein, the singular forms “a," “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features,PATENT Client Reference No.: P39512-WO-1 steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items and may be abbreviated as
[0036] Spatially relative terms, such as “under," “below," “lower," “over," “upper” and the like, may be used herein for ease of description to describe one element or feature’s relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if a device in the figures is inverted, elements described as “under” or “beneath” other elements or features would then be oriented “over” the other elements or features. Thus, the exemplary term “under” can encompass both an orientation of over and under. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly. Similarly, the terms “upwardly," “downwardly," “vertical," “horizontal” and the like are used herein for the purpose of explanation only unless specifically indicated otherwise.
[0037] Although the terms “first” and “second” may be used herein to describe various features / elements (including steps), these features / elements should not be limited by these terms, unless the context indicates otherwise. These terms may be used to distinguish one feature / element from another feature / element. Thus, a first feature / element discussed below could be termed a second feature / element, and similarly, a second feature / element discussed below could be termed a first feature / element without departing from the teachings of the present invention. Overview
[0038] The identification of structural variants (SVs) or genomic alterations by identifying regions of nucleic acid molecular (e.g., DNA or RNA) sequences that are variations from a reference sequence is a critical task in bioinformatics. SVs in DNA can be identified by detection of breakpoints or the chromosomal position at which a double stranded break occurs, such as during meiosis, as a result of environmental factors, or errors in DNA repair mechanisms. The identification of SVs can be useful in the diagnosis, prognosis, and / or treatment of multiple diseases including cancer.PATENT Client Reference No.: P39512-WO-1
[0039] Identification of SVs can be difficult due to the varying outputs from sequencing platforms, the variability of SV size (e.g., dozens to millions of base pairs), the complexity of the sample that is sequenced, as well as the low frequency of SVs being present in a dataset of sequencing data. For example, a sequencing by expansion (SBX) process may generate single end reads without discordant reads, thus preventing easy identification of at least some somatic SVs in the SBX sequencing data. Additionally, a process that uses data from a targeted sequencing process (i.e., that includes sequencing data from a select number of targeted genes or chromosomal regions) may not have enough data to determine some somatic SVs using conventional methods.
[0040] Somatic SVs are variations in regions of DNA that can occur during cell development and aging. SVs, including somatic SVs, are associated with cancer and other diseases, and identification of SVs remains a computationally intensive process, which in previous systems was prone to error.
[0041] The techniques described herein provide an improved method for identifying somatic SVs, such as those that can be useful in the diagnosis, prognosis and / or treatment of multiple diseases such as cancerous conditions, and / or discriminate between different forms of cancer. The techniques disclosed herein may be particularly suited to sequencing data collected using SBX chemistry, such as with a nanopore-based sequencer instrument. Systems and Devices for Identifying Somatic SVs
[0042] FIG.1 sets forth an illustrative system 100 including a sequencing device 110 (which may be alternately referred to as a sequencer instrument) communicatively coupled to a computing system 102. Sequencing device 110 can be coupled to computing system 102 either directly (e.g., through one or more communication cables) or through network 130, which may be the Internet or any other combination of wide-area, local area, wired, and / or wireless networks. In some embodiments, computing system 102 may be included in or integrated with the sequencing device 110. In some embodiments, sequencing device 110 may sequence (e.g., perform a biochemical assay) a sample containing genetic material and produce resulting sequencing data. The sequencing data can be sent to computing system 102 (e.g., through network 130) or stored on a storage device and at a later stage transferred to computing system 102 (e.g., through network 130). In some embodiments, computing system 102 may or may not include a display 108 and one or more input devices (not illustrated) forPATENT Client Reference No.: P39512-WO-1 receiving commands from a user or operator (e.g. a technician or a geneticist). In some embodiments, computing system 102 and / or sequencing device 110 can be accessed by users or other devices remotely through network 130. Thus, in some embodiments various methods discussed herein may be run remotely on computing system 102.
[0043] Computing system 102 may include one computing device or a combination of a number of computing devices of any type, such as personal computers, laptops, network servers (e.g., local servers or servers included on a public / private / hybrid cloud), mobile devices, etc., where some or all of the devices can be interconnected. Computing system 102 may include one or more processors (not illustrated), each of which can have one or more logic cores. In some embodiments, computing system 102 can include one or more general- purpose processors (e.g., CPUs), special-purpose processors such as graphics processors (GPUs), digital signal processors, or any combination of these and other types of processors. In some embodiments, some or all processors in computing system can be implemented using customized or customizable circuitry, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). Computing system 102 can also in some embodiments retrieve and execute non-transitory computer-readable instructions stored in one or more memories or storage devices (not illustrated) integrated into or otherwise communicatively coupled to computing system 102. The memory / storage devices can include any combination of non-transitory computer readable storage media including semiconductor memory chips of various types (e.g., dynamic random access memory (DRAM), static random access memory (SRAM), synchronous dynamic random access memory (SDRAM), flash memory, programmable read-only memory, etc.) and so on. Magnetic and / or optical disks can also be used. The memories / storage devices can also include removable storage media that can be readable and / or writeable; examples of such media include compact disc (CD), read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), read-only and recordable Blu-ray® disks, ultra-density optical disks, flash memory cards (e.g., SD cards, mini-SD cards, micro-SD cards, etc.), and so on. In some embodiments, data and other information (e.g. sequencing data) can be stored in one or more remote locations, e.g., cloud storage, and synchronized with other the components of system 100.
[0044] In some embodiments, the sequencing device 110 can generate sequencing data by a sequencing by expansion (SBX) process. Examples of the SBX process include those described in U.S. Patent Application No.17 / 456,342 (U.S. Publication No.PATENT Client Reference No.: P39512-WO-1 US20220411458A1), entitled “Translocation control elements, reporter codes, and further means for translocation control for use in nanopore sequencing,” filed November 23, 2021, which is herein incorporated by reference in its entirety. During library preparation in the SBX process, a number of surrogate molecules are derived from and characterize nucleic acid material provided in a sample.
[0045] More particularly, the SBX process may translate a sequence of DNA into a measurable surrogate molecule called an Xpandomer. Xpandomer synthesis based on the natural function of DNA replication uses expandable nucleotide triphosphates (X-NTPs) that act as substrates for template-dependent, polymerase-based replication. These Xpandomer molecules are then processed by a sequencer instrument (e.g., the sequencing device 110) to measure the sequence in the original DNA template. As the Xpandomer molecule transits through nanometer-sized openings in an electrode-resistant membrane (a “nanopore”), each nanopore corresponding to a selective channel, a distinct electrical signal is generated for each base reporter and identifiable to enable highly accurate and high throughput nanopore- based nucleic acid sequencing (also referred to generally as nanopore sequencing).
[0046] Sequencing device 110 can generate a plurality of sequence reads corresponding to a genetic sample (e.g., a sample comprising a patient’s DNA or RNA material). For example, a sequence may be identified by processing a sample, which may include (for example) a blood, saliva, or tissue biopsy collected from a subject. Sequence reads can be obtained either directly from sequencing device 110, or from one or more local or remote volatile or non-volatile memories, storage devices, or databases communicatively coupled to computing system 102. Sequence reads can be pre-processed (e.g., pre-aligned) or they can be “raw,” in which case a downstream method may include a preprocessing (e.g., pre- aligning) step. Also, while in some embodiments, entire sequence reads (as generated by sequencing device 110) can be obtained, in other embodiments only sections of the sequence reads can be obtained. Thus, “obtaining a sequence read,” as used herein, refers generally to obtaining one or more sections of one or more (e.g., adjacent) sequence reads. Method for Identifying Somatic Structural Variants
[0047] FIG.2 illustrates a flow chart of a method 200 for identifying somatic structural variants, in accordance with at least some embodiments. The method 200 may be implemented, for example, in the form of software (i.e., a set of instructions stored in one orPATENT Client Reference No.: P39512-WO-1 more non-transitory computer-readable media accessible and executable by one or more processors, e.g., of computing system 102), or in the form of firmware, hardware, or any combination thereof. A. Receiving Base Pair Data
[0048] In step 201, a dataset including base pair data output from a sequencing process is received. In some embodiments, the sequencing process can be implemented by a sequencer instrument, such as sequencing device 110, using a SBX process. The dataset can include base pair data such as whole genome sequencing (WGS) data, whole exome sequencing (WES) data, targeted sequencing data, whole transcriptome sequencing data, and the like. The dataset including base pair data can be generated by taking base pair data output by a SBX process and performing processing steps on the sequencing data. Examples of processing steps include clustering, consensus, alignment, mapping, and / or realignment of the base pair data.
[0049] For example, a sequencing process may output a dataset in one or more files (e.g., FASTQ files) that includes base pair data for a plurality of sequence reads. The base pair data can include a sequence of bases (e.g., nucleotides) and their respective position in a corresponding sample.
[0050] Prior to being provided to the process illustrated in FIG.2, in some embodiments, the raw sequencing data output by a sequencing process can be pre-processed using one or more algorithms configured to perform clustering and / or consensus. In some embodiments, the pre-processing of the raw sequencing data can also include alignment of the sequence reads to a reference sequence. The pre-processing of the raw sequencing data output by the sequencing process can result in an alignment dataset provided in one or more binary alignment map (.bam) files that can be used by the process in FIG.2.
[0051] For example, the sequencing process described above can result in the generation of incomplete read data (e.g., full and partial reads) due, at least in part, to polymerase detaching from DNA molecules leading to incomplete synthesis of Xpandomer molecules. Thus, pre-processing of raw sequencing data output by a sequencing process can be used to generate file(s) composed of a re-aligned consensus of reads from a number of reads corresponding to different molecules.PATENT Client Reference No.: P39512-WO-1
[0052] As another example, the raw sequencing base pair data, which can include both full and partial reads obtained from a sequencing process, can be clustered using one or more algorithms configured to sort full and partial length reads based on their binding sites or ends, clustering the sorted reads based on their unique molecular identifiers (UMIs) and / or base positions, merging reads that are identified to be the same, and determining a consensus sequence based on likelihoods that that the clustered reads supports a consensus sequence. The resulting re-aligned consensus reads can be used to generate a .bam file that can be used in a process for identifying somatic structural variants.
[0053] Examples of a clustering and consensus process include those described in International Patent Application No. PCT / US2025 / 028649, entitled “Intermolecular Consensus of Partially Read Sequences,” filed May 9, 2025 and having a priority date of May 10, 2024, which is herein incorporated by reference in its entirety.
[0054] For example, in some embodiments, a consensus sequence may be generated using a set of sequences. Consensus sequences can include a sequence of nucleotides that are the most common at a specific position in the sequencing data. In some instances, at least some or all of the set of sequences may have been generated using a same sample, a same molecule, a same sequencing technique and / or a same sequencing device. Alternatively or additionally, at least some or all of the set of sequences may have been generated using different samples, different molecules, different sequencing techniques and / or different sequencing devices. As one illustrative example, the set of sequences may have been generated by using a same sequencing system and technique (e.g., an SBX process) and using a same sample. A consensus sequence may be determined locally within the sequencing system and output to a user, e.g., as one or more files.
[0055] The set of sequence representations output by the sequencing device can be aligned to each other to form the consensus sequence. The alignment can be performed using a cost or loss function. Alternatively or additionally, the sequences can be aligned using a multiple sequence alignment technique, iterative methods, or progressive techniques.
[0056] In some embodiments, the dataset can be processed by mapping the data obtained from the output of the SBX process to a reference sequence using unique molecular identifiers (UMIs), e.g., by clustering the mapped data based on at least one of a startPATENT Client Reference No.: P39512-WO-1 position, end position, or index indicated by the UMIs, and correcting errors based on the clustered data.
[0057] The received dataset in step 201 can be in any suitable format such as in a FASTQ file, a Binary Alignment Map (.bam) file, Sequence Alignment Map (.sam) file, and the like. The dataset received in step 201 can include pre-processed raw sequencing data that has been processed using any of the above described techniques. In some cases, the raw sequencing data or pre-processed sequencing data can be further processed by any subset of the techniques described above after being received to ensure the data is in a standard format. For example, raw sequencing data received in one or more FASTQ files can be processed using the above described techniques to generate corresponding BAM / SAM files. As another example, pre-processed sequencing data received in a SAM file format can be converted to a BAM file format, or vice versa. B. Extracting and Grouping Soft-Clipped Reads
[0058] In step 203, soft-clipped reads can be extracted from the received dataset and, at step 205, the extracted soft-clipped reads can be grouped based on their respective breakpoint(s). Soft-clipped reads can be sequencing reads with bases that do not align with a reference sequence removed from the ends of reads. The extracted soft-clipped reads can be grouped by identifying their respective breakpoints. Breakpoints can be identified based on a comparison with a reference sequence. Soft-clipped reads with identical breakpoints can be determined to form the same group.
[0059] For example, in some embodiments, the method 200 receives a BAM file which includes the realigned consensus reads derived from processing the raw sequencing data output by the SBX process. The soft-clipped reads can be identified in the received aligned file by locating the reads indicated as being soft-clipped in the BAM file. For example, soft- clipped reads can be associated with a tag such as a ‘S’ in the .BAM file. Soft-clipped reads can be tagged with information in a Concise Idiosyncratic Gapped Alignment Report (CIGAR) string associated with the received BAM file. Breakpoints associated with the soft- clipped reads can be determined by comparing the soft-clipped reads to a reference sequence. The extracted soft-clipped reads can be grouped by their respective breakpoints. The soft- clipped reads with the same breakpoint can be used to determine a consensus sequence by a majority vote.PATENT Client Reference No.: P39512-WO-1
[0060] In some embodiments, additional filters can be applied to the extracted and grouped soft-clipped reads. For example, a high guanine-cytosine (GC)- content filter can be applied to reduce false positives. In some embodiments the GC-filter can be configured to filter out values that are less than 30% or greater than 70%. In some embodiments, long homopolymer representatives can be removed from the extracted and grouped soft-clipped reads.
[0061] The reference sequence may be obtained, for example, from one or more private or public repositories such as the Reference Sequence (RefSeq), an open-access, annotated and curated collection of nucleotide base sequences built by the National Center for Biotechnology Information (NCBI) or the NCBI Genomes FTP site storing a set of complete genomes of different organisms. In some embodiments, a specific copy of the reference sequence may be stored locally (e.g., in a memory of computing system 102), while in other embodiments the reference sequence may be obtained from a remote server, e.g., through network 130. Furthermore, in some embodiments, the entire reference sequence may be obtained while in other embodiments one or more sections of the reference sequence may be obtained – e.g., only the section(s) that is / are associated with a particular assay. Thus, a “reference sequence” as used herein refers generally to one or more sections of the reference sequence, which may or may not include the entire reference sequence. C. Identification of Somatic SVs
[0062] In step 207, the presence or absence of a somatic SV can be determined by k-mer matching the grouped and extracted soft-clipped reads. In an embodiment, nucleotide sequences can be analyzed in groups to determine matches with the reference sequence. For example, the soft-clipped reads with identical breakpoints will include repetitive nucleotides. Using the k-mer matching process, the groups of soft-clipped reads can be analyzed to determine the length and sequence of the repetitive DNA elements. For example, a consensus sequence determined by extracting and grouping the soft-clipped reads in step 203 can be used for k-mer matching.
[0063] The k-mer matching process starts with the consensus soft-clipped reads grouped in step 205. For each group, a k-mer hash table is generated in which the k-mers of each consensus read are listed. A k-mer is a short sequence of bases that is k bases long and each successive k-mer overlaps the previous k-mer by k-1 bases. The k-mer matching process isPATENT Client Reference No.: P39512-WO-1 done by counting how many k-mers are the same between the k-mer hash table of the soft- clipped portion of consensus read 1 and the k-mer hash table of the mapped portion of consensus read 2, and vice versa. A matching score and threshold can be applied to the k-mer process in order to identify a sequence corresponding to the original sample.
[0064] More specifically, the k-mer matching process is used to identify consensus sequences aligned to different locations of the reference sequence that may be related to a same somatic SV at a particular location in the sample that varies from the reference sequence due to, e.g., a translocation or deletion event. In other words, the soft-clipped portion of one breakpoint may match the mapped portion of another breakpoint, and vice versa, which could indicate the presence of a somatic SV rather than separate and distinct reads associated with different locations of the sample genome.
[0065] In some embodiments, the k-mer matching process can apply a coarse-grained algorithm to improve the running time. The coarse-grained algorithm selects the k-mers in certain positions of the consensus reads to match instead of using all k-mers. For example, if k-mers from the start, middle and end positions of the consensus reads are used instead of all possible k-mers, the k-mer matching process can process the data in less than about 10 minutes for 60GB data.
[0066] As discussed above, the k-mer matching process can be used to identify whether the dataset includes SVs. More specifically, the k-mer matching process can be used to output a sequence of base pairs, which can then be analyzed in order to identify SVs. For example, overlapping segments of k-mers can be used to reconstruct a sequence of base pairs corresponding to the sample, which can then be used to determine if the sample includes SVs. Again, SVs include variations from a reference sequence such as deletions, duplications, insertions, inversions, translocations, fusions, inter-gene deletions, and the like. SVs can be indicative of the deletion, amplification, or reordering of genomic sequences. Deletions are indicative of the point at which one or more contiguous nucleotides were excised in the sequence. Duplications are indicative of the point at which one or more nucleotides were repeated in the sequence, when compared to a reference sequence. Insertions are indicative of the introduction of one or more new nucleotides to the sequence. Inversions are indicative of when a continuous nucleotide sequence is inverted in the same position of the sequence. Translocations are indicative of when a segment of nucleotides are moved in the sequence. Fusions are indicative of two sequences corresponding to genes fusing together or joining toPATENT Client Reference No.: P39512-WO-1 form a new gene with altered phenotype. Inter-gene deletions are indicative of deletions that occur between coding regions.
[0067] Somatic SVs can include a plurality of base pairs. For example, somatic SVs can have a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, or more nucleotides. For example, in some embodiments, the somatic SVs can include ten (10) base pairs or more. In some embodiments, somatic SVs can include twenty-five (25) base pairs or more. In some embodiments, the somatic SVs can include fifty (50) base pairs or more.
[0068] In some embodiments, the somatic SVs have an allele frequency that is less than a threshold percentage. In some embodiments, the somatic SVs have an allele frequency that is less than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more percent of alleles represented in the sequencing data. In some embodiments, the somatic SVs have an allele frequency that is less than about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, or 0.9 percent of alleles represented in the sequencing data. In some embodiments, the somatic SVs can have an allele frequency between about 0.1-10 percent of alleles represented in the sequencing data. In other embodiments, the somatic SVs can have any of the lower limit in combination with an upper limit of the numbers disclosed above as a range.
[0069] In an optional step, secondary alignment information can be used to determine the presence or absence of the somatic SVs. For example, secondary alignment information can indicate a chromosome that the soft-clipped read aligned with and / or positional information regarding the soft-clipped read alignment. For example, information regarding the chromosome corresponding to the grouping of soft-clipped reads can be used in grouping the soft-clipped reads and / or reducing the consensus sequences that undergo the k-mer matching process. In some embodiments, the optional step including secondary alignment information can reduce the number of possible combinations of fusion genes, thereby reducing the search space and boosting efficiency of the processes described herein. D. Providing Data Indicative of Presence or Absence of Somatic SVs
[0070] In step 209, the data indicative of the determined presence or absence of the somatic SVs can be provided. For example, the k-mer matching process can result in a list ofPATENT Client Reference No.: P39512-WO-1 SVs associated with the sample. Each SV is associated with a chromosome, position, supporting counts, and the gene symbol information. In some embodiments, the output file can include orientation data (e.g., + strand, - strand) and location data of the genes present. SVs can be determined based at least on the orientation data and the location data.
[0071] Providing the determined presence or absence of the somatic SV can include generating a dataset indicative of at least one somatic SV determined to be present. For example, an output file can be generated in a variant calling format (VCF) indicative of the location of at least one somatic SV determined to be present. Additionally, or alternatively, providing the determined presence or absence of the somatic SV can include displaying a representation of the at least one somatic SV, when it is determined that a somatic SV is present. Additionally, or alternatively, the absence of a somatic SV can be reported to a user. In some embodiments, the somatic SVs can be displayed in a graphical user interface. For example, a nucleotides listing can be provided to a user along with one or more graphical icons indicating the location of a somatic SV. Alternatively, a graphical icon can indicate the type of SV detected.
[0072] The output file can be a variant calling format file (VCF file), binary variant calling format file (BCF) or the like. The output file can include the location of detected SVs. The output file can report detected breakpoints and the like. The output file can also include a categorization of the type of SV detected. In some embodiments, a classification algorithm can be applied to generate categorizations for the detected SVs. The classification algorithm can be applied to the orientation and location data included herein. In an embodiment, a classification algorithm would be able to classify the following SVs based on breakpoint information. For example, if two breakpoints are not in the same chromosome, they can be classified as a translocation. In another example, if two breakpoints are in the same chromosome, but have different orientations, then the breakpoints can be classified as an inversion. In another example, if two breakpoints are in the same chromosome and their orientations are bothand if the position of one breakpoint is smaller than the position of the second breakpoint they can be classified as a deletion. In another example, if two breakpoints are in the same chromosome and their orientations are both "+" and if the position of one breakpoint is not smaller than the position of the second breakpoint they can be classified as a duplication.PATENT Client Reference No.: P39512-WO-1
[0073] The output file can be produced after a validation and filtering process is applied to the results of the k-mer matching process. Example filters include high GC-content filters, long homopolymer filters, low allele frequency filters and the like.
[0074] In an embodiment, the candidate list of SVs produced by the k-mer matching process can undergo validation and filtration. For example, multiple SV candidates can be reviewed to determine if they can be merged into one SV candidate. The resulting SV candidate can be associated with a re-calculated position which is output to the final output file. In some embodiments an allele frequency cut off can be applied. In some embodiments, a homopolymer cut off can be applied.
[0075] In some embodiments, the allele frequency cutoff can be adjusted based on whether the underlying sample corresponds to a plasma sample or is a solid tumor tissue sample. For example, an allele frequency threshold between about 0.5 % to 15 % can be applied. For example, for a plasma sample, a range between 0.1% to 15% can be applied.
[0076] As discussed herein, data generated by identifying somatic SVs, which may be in the form of an output file, can be generated in a variant calling format indicative of the location of at least one somatic SV. The indication that at least one somatic SV is determined to be present can be useful in the diagnosis, prognosis and / or treatment of multiple diseases such as cancer. For example, a prediction of at least one of cancer risk or cancer treatment can be generated using the information regarding the presence or absence of a somatic SV in a sequence, and / or the location information for one or more somatic SVs. Examples
[0077] FIG.3 illustrates a schematic diagram for a process 300 for identifying somatic SVs, analogous to that illustrated in FIG.2, in accordance with at least some embodiments of the present disclosure. As illustrated in FIG.3, a dataset including base pair data output by a sequencing device 110 using, e.g., an SBX process can be read 301 by a computing system 102. In some embodiments, the dataset can be in a BAM format file 301. Soft-clipped reads can be extracted from the base pair data 313 to produce a collection of soft-clipped reads 303. The resulting soft-clipped reads 303 can be grouped by their corresponding breakpoints. In the example illustrated in FIG.3, two breakpoints (i.e., Breakpoint 1305, Breakpoint 2307) are identified, and the corresponding soft-clipped reads are aligned accordingly. The soft-PATENT Client Reference No.: P39512-WO-1 clipped reads associated with each breakpoint are input into a consensus algorithm 317 in order to generate consensus sequences corresponding to each breakpoint 309, 311. In some embodiments, the consensus algorithm 317 can be configured to output one consensus read per breakpoint group, and include one or more filters. For example, the filters can remove breakpoints associated with high GC-content and / or long homopolymer representatives. A k- mer matching algorithm 319 is applied to each consensus sequence corresponding to a breakpoint to identify mapped genes and / or soft-clipped sequences. This process can include an additional alignment step.
[0078] The resulting identified mapped genes and / or soft-clipped sequences can undergo a validation and filtering process 325. The validation and filtering process can be used to screen and remove artifacts and outputs a file including a listing of somatic SVs 327, including their positional information. Alternatively, or additionally, validation and filtering can be performed prior to the application of a k-mer matching process (e.g., at step 203 in FIG.2).
[0079] Artifacts may result from one or more of the following: misalignments (causing false positives, especially near the start or end of reads, or within repetitive regions of the genome), low base quality, DNA damage (which can cause artifactual single-base substitutions during sample handling and library preparation), polymerase chain reaction (PCR) errors, cross-individual contamination (which can lead to false positives), germline mutations (low variant allele frequency due to read sampling bias can be misclassified as somatic), and / or tumor-normal cross contaminations, tumor ploidy, sub-clonality or intra- tumor heterogeneity, and local copy number variation (each of these can cause low variant allele frequencies in tumor samples).
[0080] In one embodiment, the variant caller excludes variants with allele frequencies below the allele frequency call threshold or variants with an allele frequency below the allele frequency filter threshold and tagged with a low allele frequency filter tag. The default allele frequency call threshold is 1% and the default allele frequency (AF) filter threshold is 5%. Alternative filtering methods include, but are not limited to, using counts of somatic variants supporting reads as a hard filter, where the threshold is variants that fall below 1.0% SV caller is a simple allele frequency cutoff filter, identifying a blacklist of variants frequent within a normal population to remove variants having a negative predictive value (NPVs),PATENT Client Reference No.: P39512-WO-1 and / or using a machine learning-based algorithm to associate low-frequency alleles with the disease of interest.
[0081] In some embodiments, optionally, supplementary reads can be determined from the dataset provided in 301. The supplementary reads can include information regarding candidate gene pairs 331. The candidate gene pairs can be used to remove breakpoints that are not in the candidate gene pairs 333, thus aiding in the identification of mapped genes by reducing the search space for the k-mer matching algorithm.
[0082] More specifically, supplementary read information can be included in the SAM / BAM files of the sequencing dataset provided in 301. Supplementary read information refers to an alignment that is not the primary alignment for a given read. For example, supplementary read information can be provided in the CIGAR string of the BAM file, which can indicate that a particular read could be mapped (i.e., aligned) to a secondary location. In the case of a somatic SV, the sequence information can be the result of a fusion of two genes. Thus, the sequence may contain base pairs for two different genes located in different locations in the reference sequence. The mapped portion of the sequence may be aligned to a first gene, while the soft-clipped portion of the sequence does not match the first gene in the reference sequence. However, the supplementary read information may show that the soft- clipped portion of the sequence could be aligned to a second gene in the reference sequence. Thus, a given read might be associated with two genes located in different positions of the reference sequence, which could be referred to as a candidate gene pair that may be indicative of a fusion event that resulted in the somatic SV.
[0083] As shown in FIG.3, the supplementary read information in the sequencing dataset provided in 301 can be used to identify candidate gene pairs 331 used to filter breakpoints to reduce complexity in the k-mer matching process 319. At 329, supplementary reads are extracted from the sequencing dataset (e.g., based on SA tags in the BAM file). The supplementary read information is used to identify a set of candidate gene pairs 331 which may indicate the presence of a somatic SV resulting from a fusion event (e.g., translocation, deletion, etc.) at the location of the read. At 333, the set of candidate gene pairs 331 can then be used to filter breakpoints that go through the k-mer matching process at 319 in order to improve speed and accuracy of the somatic SV calling.PATENT Client Reference No.: P39512-WO-1
[0084] The resulting output file 327 composed of somatic SV information can be used for diagnostics, prognostic, and treatment of disease. Computer Programming
[0085] FIG.4 illustrates a functional block diagram of a machine in the example form of computer system 400, within which a set of instructions for causing the machine to perform any one or more of the methodologies, processes or functions discussed herein may be executed. In some examples, the machine may be connected (e.g., networked) to other machines as described above. The machine may operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may be any special-purpose machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine for performing the functions describe herein. Further, while only a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein. In some examples, the computing system 102 of FIG.1 may be implemented by the example machine shown in FIG.4 (or a combination of two or more of such machines).
[0086] Example computer system 400 may include processing device 403, memory 407, data storage device 409 and communication interface 415, which may communicate with each other via data and control bus 401. In some examples, computer system 400 may also include display device 413 and / or user interface 411. In some embodiments, the user interface 411 may include a graphical user interface.
[0087] Processing device 403 may include, without being limited to, a microprocessor, a central processing unit, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP) and / or a network processor. Processing device 301 may be configured to execute processing logic 405 for performing the operations described herein. In general, processing device 403 may include any suitable special-purpose processing device specially programmed with processing logic 405 to perform the operations described herein.PATENT Client Reference No.: P39512-WO-1
[0088] Memory 407 may include, for example, without being limited to, at least one of a read-only memory (ROM), a random access memory (RAM), a flash memory, a dynamic RAM (DRAM) and a static RAM (SRAM), storing computer-readable instructions 417 executable by processing device 403. In general, memory 407 may include any suitable non- transitory computer readable storage medium storing computer-readable instructions 417 executable by processing device 403 for performing the operations described herein. Although one memory device 407 is illustrated in FIG.4, in some examples, computer system 400 may include two or more memory devices (e.g., dynamic memory and static memory).
[0089] Computer system 400 may include communication interface device 415, for direct communication with other computers (including wired and / or wireless communication), and / or for communication with a network. In some examples, computer system 400 may include display device 413 (e.g., a liquid crystal display (LCD), a touch sensitive display, etc.). In some examples, computer system 400 may include user interface 411 (e.g., an alphanumeric input device, a cursor control device, etc.).
[0090] In some examples, computer system 400 may include data storage device 409 storing instructions (e.g., software) for performing any one or more of the functions described herein. Data storage device 409 may include any suitable non-transitory computer-readable storage medium, including, without being limited to, solid-state memories, optical media and magnetic media.
[0091] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, can also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, can also be provided separately or in any suitable subcombination. All combinations of the embodiments pertaining to the disclosure are specifically embraced by the present disclosure and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub- combinations of the various embodiments and elements thereof are also specifically embraced by the present disclosure and are disclosed herein just as if each and every such sub combination was individually and explicitly disclosed herein.PATENT Client Reference No.: P39512-WO-1
[0092] It is intended that every maximum numerical limitation given throughout this specification includes every lower numerical limitation, as if such lower numerical limitations were expressly written herein. Every minimum numerical limitation given throughout this specification will include every higher numerical limitation, as if such higher numerical limitations were expressly written herein. Every numerical range given throughout this specification will include every narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were all expressly written herein.
[0093] Although various illustrative embodiments are described above, any of a number of changes may be made to various embodiments without departing from the scope of the invention as described by the claims. For example, the order in which various described method steps are performed may often be changed in alternative embodiments, and in other alternative embodiments one or more method steps may be skipped altogether. Optional features of various device and system embodiments may be included in some embodiments and not in others. Therefore, the foregoing description is provided primarily for exemplary purposes and should not be interpreted to limit the scope of the invention as it is set forth in the claims.
[0094] The examples and illustrations included herein show, by way of illustration and not of limitation, specific embodiments in which the subject matter may be practiced. As mentioned, other embodiments may be utilized and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Such embodiments of the inventive subject matter may be referred to herein individually or collectively by the term “invention” merely for convenience and without intending to voluntarily limit the scope of this application to any single invention or inventive concept, if more than one is, in fact, disclosed. Thus, although specific embodiments have been illustrated and described herein, any arrangement calculated to achieve the same purpose may be substituted for the specific embodiments shown. This disclosure is intended to cover any and all adaptations or variations of various embodiments. Combinations of the above embodiments, and other embodiments not specifically described herein, will be apparent to those of skill in the art upon reviewing the above description.
Claims
PATENT Client Reference No.: P39512-WO-1 CLAIMS 1. A computer-implemented method of identifying somatic structural variants, the method comprising: receiving, at a processor, a dataset comprising base pair data output from a sequencing by expansion process; extracting, at the processor, soft-clipped reads from the received dataset; grouping, at the processor, the extracted soft-clipped reads based at least on their respective breakpoints; determining, at the processor, a presence or an absence of a somatic structural variant by k-mer matching the grouped and extracted soft-clipped reads; and providing, at the processor, at least one determined somatic structural variant responsive to determining the presence of the somatic structural variant.
2. The computer-implemented method of claim 1, wherein the base pair data is derived from at least one of targeted sequencing, whole genome sequencing, whole transcriptome sequencing, and whole exome sequencing.
3. The computer-implemented method of claim 1 or 2, further comprising: generating the dataset comprising the base pair data by at least one of clustering, consensus, and realignment of data obtained from the output of the sequencing by expansion process.
4. The computer-implemented method of any one of claims 1-3, wherein the dataset comprising the base pair data is generated based on at least one of the following: mapping the data obtained from the output of the sequencing by expansion process to a reference sequence using unique molecular identifiers (UMIs), clustering the mapped data based on at least one of a start position, end position, or index indicated by the UMIs, and correcting errors based on the clustered mapped data.PATENT Client Reference No.: P39512-WO-1 5. The computer-implemented method of any one of claims 1-4, wherein the somatic structural variant comprises at least one of a fusion, translocation, inversion, inter- gene deletion, intra-gene deletion, or a duplication.
6. The computer-implemented method of any one of claims 1-5, wherein the somatic structural variant comprises at least ten (10) base pairs.
7. The computer-implemented method of any one of claims 1-5, wherein the somatic structural variant comprises at least twenty-five (25) base pairs.
8. The computer-implemented method of any one of claims 1-5, wherein the somatic structural variant comprises at least fifty (50) base pairs.
9. The computer-implemented method of any of claims 1-8, wherein the dataset comprising the base pair data characterizes soft-clipped genomic data.
10. The computer-implemented method of any of claims 1-9, wherein the somatic structural variant has an allele frequency that is less than 10 percent.
11. The computer-implemented method of any of claims 1-10, wherein the somatic structural variant has an allele frequency that is about 0.1 to about 10 percent.
12. The computer-implemented method of any of claim 1-11, wherein providing the at least one determined somatic structural variant comprises generating a dataset indicative of the at least somatic structural variant or displaying a representation of the at least one somatic structural variant.
13. The computer-implemented method of claim 12, wherein the representation of the at least one somatic structural variant is displayed in a graphical user interface.
14. The computer-implemented method of any of claims 1-11, wherein providing the at least one determined somatic structural variant comprises generating an output file in aPATENT Client Reference No.: P39512-WO-1 variant calling format indicative of a location of one or more of the at least one determined somatic structural variant.
15. The computer-implemented method of any of claims 1-14, further comprising: receiving, at the processor, secondary alignment information; and wherein determining the presence or absence of the somatic structural variant further comprises the received secondary alignment information.
16. The computer-implemented method of any of claims 1-14, further comprising: generating a prediction of at least one of cancer risk or cancer treatment using the determined at least one somatic structural variant.
17. The computer-implemented method of any of claims 1-14, wherein determining the presence of the somatic structural variant further comprises: comparing the received dataset to a dataset for a reference genome.
18. A system for identifying somatic structural variants, the system comprising: at least one processor configured to: receive a dataset comprising base pair data output from a sequencing by expansion process; extract soft-clipped reads from the received dataset; group the extracted soft-clipped reads based at least on their respective breakpoints; determine a presence or an absence of a somatic structural variant by k-mer matching the grouped and extracted soft-clipped reads; and provide at least one determined somatic structural variant responsive to determining the presence of the somatic structural variant.
Citation Information
Patent Citations
Translocation control elements, reporter codes, and further means for translocation control for use in nanopore sequencing
US20220411458A1
Intermolecular consensus of partially read sequences
WO2025235888A1