Systems and methods for marking duplicate reads of nucleic acid sequences
The two-pass approach with smaller clip-padding and specialized indices resolves shard conflicts in long RNA sequences, optimizing memory usage and enabling efficient parallel processing for accurate duplicate marking.
Patent Information
- Application Number
- PCT/US2025/029483
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-20
- Filing Date
- 2025-05-15
- Publication Date
- 2025-11-27
AI Technical Summary
Traditional duplicate marking tools struggle with the computational and memory challenges posed by long RNA sequences due to large genomic distances and intron-spanning alignments, leading to inefficient handling of distant mate pairs and increased resource consumption.
A two-pass approach is employed, utilizing smaller clip-padding in shards and specialized indices to identify distant mate pairs and friends, resolving conflicts by injecting relevant reads into appropriate shards, thereby optimizing memory usage and enabling parallel processing.
This method efficiently addresses shard conflicts and reduces memory exhaustion, allowing for time-effective analysis of large RNA datasets with improved accuracy in duplicate marking.
Smart Images

Figure US2025029483_27112025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR MARKING DUPLICATE READS OF NUCLEICACID SEQUENCESCROSS REFERENCE TO RELATED APPLICATION
[0001] The application claims priority to U.S. Provisional Application No. 63 / 649,742, filed on May 20, 2024, which is incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to the field of bioinformatics and nucleic acid sequencing and, more specifically, to systems and methods for improving duplicate marking tools to address the challenges posed by long reads in nucleic acid sequences, including RNA sequences.BACKGROUND
[0003] Traditional duplicate marking tools have sufficiently handled short reads common in DNA sequencing. However, their efficiency diminishes when confronted, for example, with the intricacies of long RNA sequences. Nucleic acid sequences in RNA may traverse extensive genomic distances due to intron-exon junctions, as opposed to the more contiguous nature of DNA fragments. Technological challenges abound when attempting to process these nucleic acid sequences. For example, substantial memory resources may be needed to effectively process the greater volume of data associated with these longer RNA reads, which may impact the performance and feasibility of conventional duplicate marking tools. One or more aspects of this disclosure may address one or more of the issues described above.
[0004] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art, or suggestions of the prior art, by inclusion in this section.SUMMARY OF THE DISCLOSURE
[0005] According to certain aspects of the disclosure, systems and methods are described for performing duplicate marking on long nucleic acid, e.g., RNA, sequence reads.
[0006] In one aspect, a computer-implemented is provided. The computer- implemented method may include operations including: accessing, using a processor of a computing device, a data file containing aligned sequencing data; initiating, using the processor, a duplicates marking process on the aligned sequencing data, wherein the initiating comprises segmenting the aligned sequencing data into a plurality of shards; identifying, using the processor and in a first scan of the data file, a first set of data associated with a distant mate pair in the aligned sequencing data; storing, using the processor, the first set of data into a first data structure; identifying, using the processor and in a second scan of the data file, a second set of data corresponding to a distant mate friend associated with the distant mate pair; storing, using the processor, the second set of data into a second data structure; identifying, using the processor and by accessing the first set of data in the first data structure, a conflict between two shards of the plurality of shards; and resolving the conflict by accessing the second data structure and introducing sequence information contained in the second set of data into one of the two shards needing the sequence information.
[0007] In another aspect, a system is provided. The system may include: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations comprising: accessing, using the one or more processors of a computing device associated with the system, a data file containing aligned sequencing data; initiating, using the one or more processors, a duplicates marking process on the aligned sequencing data, wherein the initiating comprises segmenting the aligned sequencing data into a plurality of shards; identifying, using the one or more processors and in a first scan of the data file, a first set of data associated with a distant mate pair in the aligned sequencing data; storing, using the one or more processors, the first set of data into a first data structure; identifying, using the one or more processors and in a second scan of the data file, a second set of data corresponding to a distant mate friend associated with the distant mate pair; storing, using the one or more processors, the second set of data into a second data structure; identifying, using the one or more processors and by accessing the first set of data in the first data structure, a conflict between two shards of the plurality of shards; and resolving the conflict by accessing the second data structure and introducing sequence information contained in the second set of data into one of the two shards needing the sequence information.
[0008] In yet another aspect, a non-transitory computer-readable medium storing computer-executable instructions is provided. The computer-executable instructions, when executed by a system, cause the system to perform operations comprising: accessing, using a processor of a computing device, a data file containing aligned sequencing data; initiating, using the processor, a duplicates marking process on the aligned sequencing data, wherein the initiating comprises segmenting the aligned sequencing data into a plurality of shards; identifying, usingthe processor and in a first scan of the data file, a first set of data associated with a distant mate pair in the aligned sequencing data; storing, using the processor, the first set of data into a first data structure; identifying, using the processor and in a second scan of the data file, a second set of data corresponding to a distant mate friend associated with the distant mate pair; storing, using the processor, the second set of data into a second data structure; identifying, using the processor and by accessing the first set of data in the first data structure, a conflict between two shards of the plurality of shards; and resolving the conflict by accessing the second data structure and introducing sequence information contained in the second set of data into one of the two shards needing the sequence information.
[0009] Additional objects and advantages of the disclosed embodiments will be set forth in part in the description that follows, and in part will be apparent from the description, or may be learned by practice of the disclosed embodiments. The objects and advantages of the disclosed embodiments will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.
[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and together with the description, serve to explain the principles of the disclosure.
[0012] FIG. 1 depicts a diagram illustrating an exemplary situation in a duplicate marking sharding process.
[0013] FIG. 2 depicts a diagram illustrating another exemplary situation in a duplicate marking sharding process.
[0014] FIG. 3 depicts a diagram illustrating another exemplary situation in a duplicate marking sharding process.
[0015] FIG. 4A depicts an exemplary computer system for executing the methods described herein.
[0016] FIG. 4B depicts an exemplary software platform for executing the methods described herein.
[0017] FIG. 5 depicts an exemplary pre-filtering process for removing sequence reads, in accordance with an embodiment.
[0018] FIG. 6 depicts an exemplary workflow for resolving conflicts between shards during a duplicate marking process, according to one or more embodiments of the present disclosure.
[0019] FIG. 7 depicts an exemplary diagram that illustrates the performance improvements resulting from the processes described herein, according to one or more embodiments of the present disclosure.
[0020] FIG. 8 depicts an example computing system suitable for performing workflows according to one or more embodiments of the present disclosure.DETAILED DESCRIPTION OF EMBODIMENTS
[0021] The terminology used below may be interpreted in its broadest reasonable manner, even though it is being used in conjunction with a detailed description of certain specific examples of the present disclosure. Indeed, certainterms may even be emphasized below; however, any terminology intended to be interpreted in any restricted manner will be overtly and specifically defined as such in this Detailed Description section. Both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the features, as claimed.
[0022] Nucleic acid sequencing has become a fundamental technique in various fields, enabling the analysis of genetic information at more granular scales. As the field has progressed, the need for accurate and efficient analytical tools has grown, particularly when the sequenced data is utilized in downstream applications such as differential expression analysis and accurate gene expression quantification. In this regard, duplicate marking may be performed to identify and remove duplicate reads of nucleic acid fragments that arise during the sequencing process. More specifically, duplicate reads correspond to those identical copies of the original nucleic acid fragments that may be generated due to various factors in the sequencing workflow. The process of duplicate marking ensures the accuracy of downstream analyses and prevents biases from being introduced by these duplicated sequences.
[0023] Traditional duplicate marking tools have been widely used in DNA sequencing pipelines, relying on grouping reads by unclipped 5' positions, orientation, and references. These tools have proven effective for short reads of fixed maximum length, as commonly found, e.g., in circulating cell-free DNA (“cfDNA”) data. One commonly used duplicate marking tool, for instance, utilizes a sorting process to identify and mark duplicate reads efficiently. This process first involves sorting Binary Alignment Map (BAM) records (i.e. , raw data of genome sequencing) based on the read positions, ensuring that pairs with the sameendpoints are grouped together. After sorting, the duplicate marking tool inserts read signatures, such as endpoints and other relevant information, into a sorted collection structure. The sorted collection structure is then iterated through to identify duplicate reads, whereby reads with similar signatures are grouped together as potential duplicates. The identified duplicates are then copied over to a new BAM file, and the duplicate marks are applied.
[0024] Although the foregoing sorting process may be an effective way to organize the data, its time complexity may be a limiting factor, especially as the dataset size increases. In response to these efficiency shortcomings, high- performance duplicate sequencing marking tools were developed for marking PCR and optical duplicate reads. One such high-performance duplicate marking tool is configured to divide the BAM file into smaller subsets of the original data, known as “shards.” Within each shard, the tool utilizes a hashmap to identify pairs of reads with the same endpoints. Each shard processes its data independently, e.g., by scanning its BAM records and appending them to a hash map. After duplicate marking is completed within each shard, the results (marked and unmarked reads) from each shard are output to new BAM files. This independent processing of shards may allow for parallelism, in which each shard may be processed concurrently, which may correspondingly improve the overall efficiency of the marking process, especially in distributed computing environments.
[0025] In the sharding approach described above, duplicates are matched using their 5' positions, and each shard has “clip-padding” on each side to account for potential variations in read positions. For instance, referring now to FIG. 1 , a diagram 100 is provided showing the alignment positions of a pair of sequence reads. In this example, readl and read2 are duplicates according to their 5' position,but the two reads reside in different shards based on their alignment begin positions. To address these situations, the clip-padding of each shard may be leveraged to address this overlap. More particularly, when searching for duplicates in each shard, each read that is either in the clip-padding or in the shard is considered. So in this example, readl is in shard2, and read2 is in shard2's clip-padding, but the two reads are still compared with each other, it is decided whether they are duplicates, and one may be marked as a duplicate based on their scores. Note that the processing algorithm for shardl also compares readl and read2 and also decides they are duplicates, and marks one based on their scores. Since the scoring is deterministic, shardl and shard2 agree on which reads to mark as duplicates.
[0026] In essence, clip-padding may act as a buffer zone around each shard, allowing for the inclusion of reads that straddle the shard boundaries. By incorporating clip-padding, duplicate marking algorithms may more comprehensively assess read duplicates, thereby improving the accuracy of downstream analysis. As another illustration of clip-padding, a shard may be responsible for processing reads within a specific genomic region defined by start positions on a chromosome (e.g., as defined by ch1 : [s_1 , s_2}, where s_1 and s_2 represent the start positions of the first and last reads covered by the shard, respectively). Clip-padding, denoted as “x”, may extend the coverage of the shard beyond its immediate genomic region. Specifically, it encompasses additional genomic positions on both ends (upstream and downstream) of the shard’s range. In other words, clip-padding expands the effective boundary of the shard to include reads that may partially overlap with its defined region. Accordingly, all reads with start positions falling within the interval chr1 : [s_1 -x, s_2+x] may be considered when processing the shard. Here, the subtraction of x from s_1 and addition of x to s_2 effectively widen the genomicrange covered by the shard, ensuring that reads slightly outside the original boundary are still accounted for during the duplicate marking process.
[0027] Although efficient for examining shorter DNA fragments, e.g., in cfDNA, the foregoing sharding approach may encounter challenges when tasked with marking duplicates in comparatively longer reads, such as in RNA data. More particularly, RNA sequence data often contains introns, which are non-coding regions in genes. Reads spanning introns result in alignments with large genomic distances, e.g., resulting at least partially due to skips or large gaps, such as those occurring during splicing events. These inherent characteristics of RNA may result in shard conflicts resulting from the inefficient handling of distant mate pairs, which may impact the accuracy of the duplicate marking process. More particularly, the term “distant mate pairs” refers to read pairs for which one read is in one shard, and its mate (i.e. , “distant mate”) is in another shard due to the presence of the large, intronspanning genomic distances. To correctly mark duplicates, both shards need to be aware of all the reads in a pair to correctly mark the duplicates. If one shard is not aware of all the reads in a pair, then a conflict may arise, which may lead to inaccurate duplicate marking. Conventional sorting approaches, such as those duplicate marking tools described above, may also struggle to efficiently handle such large genomic distances, leading to increased computational complexity and potential memory issues.
[0028] FIG. 2 provides diagram 200 that illustrates a set of shards that have a clip-pad size conventionally utilized for processing cfDNA. In this example, using conventional marking tools, shardl can ‘see’ and is aware of P1 R1 and P1 R2 because they both start in shardl . Shardl can also ‘see’ P2R1 and P2R2 because P2R1 starts in shardl and P2R2 is in shardTs padding. Shard2 can ‘see’ P2R2because that read starts in shard2’s padding. However, shard2 cannot ‘see’ P1 R1 and P1 R2 because neither read is in shard2 or shard2’s padding. Shard2 also cannot ‘see’ P2R1 because it does not start in shard2 or its padding (but this is rescued by the fact that P2R2 is a distant mate of shard2). Because shardl and shard2 cannot ‘see’ the same reads, a conflict arises between the shards. Resolving these conflicts, as described below, requires consideration of which reads may be injected into a shard to facilitate comprehensive duplicate marking.
[0029] One option to address the foregoing issues for RNA data is to significantly increase the clip-padding. For instance, FIG. 3 provides diagram 300 that illustrates a set of shards that each have larger clip-pads. In this example, shardl can ‘see’ P1 R1 and P1 R2 because they both start in shardl . Shardl can also ‘see’ P2R1 and P2R2 because P2R1 starts in shardl and P2R2 is in shardl ’s padding. Shard2 can see P1 R1 and P1 R2 because they are in shard2’s padding. Shard2 can also see P2R1 because it is in shard2’s padding and can see P2R2 because it starts in shard2. In this example, there is no conflict because both shards can see all of the read pairs between them.
[0030] In the context of this application, the differentiation between what constitutes a “large" versus “small” clip-pad may be influenced by various factors, including the characteristics of the sequencing data and the specific requirements of the data-processing workflows. For instance, in workflows involving DNA and methylation data processing, the maximum read-length and the maximum net deletion length may be relatively modest, totaling less than 250 bases. This is because skips, or large gaps in alignment, are not possible in these types of data sets. Therefore, the clip-padding required for these workflows may be considered “small,” typically limited to a small multiple, e.g., under 10 times, of the maximumread-length. In contrast, RNA data processing workflows present unique challenges due to the presence of skips, which may span hundreds of thousands of bases. While the maximum read-length and net deletion length in RNA data may be similar to those in DNA / methylation data, the inclusion of skips may necessitate significantly larger clip-padding to accommodate the extended range of potential alignments. In this context, “long” clip-padding may refer to multiples of 100 times or more the maximum read-length, reflecting the substantial buffer required to address skips effectively. Overall, the clip-padding size may be selected so that the clip-padding is large enough to accommodate the fixed maximum read-length, plus the maximum net of deletion and skip length the aligner is able to report.
[0031] Increasing the size of the clip-padding in the shards, however, may lead to a significant increase in the amount of data that needs to be buffered and processed simultaneously. In effect, these requirements may contribute to the exhaustion of computational resources afforded to the comparison process including memory, storage, and compute power. More particularly, as the clip padding grows to accommodate the characteristics of long RNA reads, the memory requirements for duplicate computation between shards increases proportionally. This may become an issue, especially when working with large datasets and attempting to process extensive clip-padding concurrently. Naively increasing the clip-padding therefore can negate the advantages of parallel and / or concurrent processing. A solution is therefore needed that may leverage smaller clip-padding while concurrently addressing shard conflicts.
[0032] Accordingly, the present disclosure is directed to addressing the computational challenges associated with performing duplicate marking in long nucleic acid (e.g., RNA) reads by leveraging smaller clip-padding in the shards whilesimultaneously addressing shard conflicts via the utilization of specialized indices. More particularly, a two-pass approach on the BAM file may be utilized. During the initial pass, distant mates and their unclipped 5' positions may be identified and stored in a first data set. In the second pass, reads sharing the unclipped 5’ positions but that are not distant mates themselves, e.g., known as “distant mates friends,” are collected and stored in a second data set. This information is utilized during the sharding process to inject relevant reads into the appropriate shards, thereby resolving shard conflicts. Furthermore, because smaller clip-padding is utilized, the computational stress on the memory may not result in memory exhaustion.
[0033] The concepts described herein introduce a novel approach to duplicate marking in the context of long nucleic acid reads, e.g., long RNA reads, thereby addressing at least some of the practical challenges related to memory exhaustion in existing algorithms, especially when dealing with large datasets and long RNA reads. In an aspect, the injections of reads into shards, along with adjustments to clip-padding, is a specific and technical improvement in computer technology to optimize memory usage during the duplicate marking process. Correspondingly, as a result of the more efficient use of the memory due to the utilization of the small clip-padding in each shard, independent and parallel processing of the raw data may be practically realized. Furthermore, these sharding processes require the utilization of machine-based, computational algorithms to concurrently analyze large amounts of data in a time-efficient manner. Therefore, the processes described herein cannot be performed by a human user in their mind while still meeting timing expectations inherent in processing nucleic acid data with an actionable goal in mind (e.g., providing test results to a patient or clinical user).
[0034] The subject matter of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific exemplary embodiments. An embodiment or implementation described herein as “exemplary” is not to be construed as preferred or advantageous, for example, over other embodiments or implementations; rather, it is intended to reflect or indicate that the embodiment(s) is / are “example” embodiment(s). Subject matter may be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any exemplary embodiments set forth herein; exemplary embodiments are provided merely to be illustrative. Likewise, a reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, or systems. Accordingly, embodiments may, for example, take the form of hardware, software, firmware, or any combination thereof. The following detailed description is, therefore, not intended to be taken in a limiting sense.
[0035] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment” or “in some embodiments,” or “in one aspect” or “in some aspects” as used herein does not necessarily refer to the same embodiment or aspect, and the phrase “in another embodiment” or “in another aspect” as used herein does not necessarily refer to a different embodiment or aspect. It is intended, for example, that claimed subject matter include combinations of exemplary embodiments in whole or in part.
[0036] Non-limiting cancer types that the concepts described herein may be applied to include, for example, breast cancer, lung cancer (e.g., non-small cell lungcancer (NSCLC)), prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, cancer of the esophagus, a lymphoma, head and neck cancer, ovarian cancer, a hepatobiliary cancer, a melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer. Additionally, it is also important to note that although the concepts described throughout this disclosure are made in reference to cancer, these designations are for exemplary purposes only and are not intended to be limiting. Specifically, duplicate marking may be performed on nucleic acids such as DNA and / or RNA in relation to other disease types or sequencing for any reason, whether or not related to a disease state.
[0037] FIG. 4A depicts an exemplary system for performing duplicate marking in long RNA sequence reads. Exemplary system 400 includes a data collection component 10, a database 20, and device data intelligence component 30, operably connected to each other via network 40. Alternatively, or additionally, one or more of the components may be connected with another component locally without reliance on network connection, e.g., through a wired connection. In many aspects described herein, sequencing data of nucleic acids are used to illustrate the concepts. However, one of skill in the art would understand that the current method may be applied to sequencing data of DNA or other materials, as well, from a variety of sample types, e.g., a blood sample (e.g., a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc.
[0038] As disclosed herein, data collection component 10 may include a device or machine with which sequencing data may be generated. In some embodiments, data collection component 10 may include one or more sequencingdevices or a facility that uses one or more sequencing devices to generate nucleic acid (e.g., DNA or RNA) sequence data of biological samples. In some aspects, data collection component 10 may be a database that receives sequencing information generated from one or more sequencing devices. Any suitable liquid or solid biological samples may be used for sequencing. In some embodiments, a biological sample may be cell-based, for example, one or more types of tissue. In some embodiments, a biological sample may be a sample that includes cell-free nucleic acid fragments. Examples of biological samples include, but are not limited to, a blood sample (e.g., a cell-free DNA (cfDNA) sample, a cell-free RNA (cfRNA) sample, a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc. Further, although sequencing of RNA from these samples is discussed herein, DNA from these samples may alternatively or additionally be sequenced.
[0039] Examples of sequencing data may include, but are not limited to, sequence read data of targeted genomic locations, partial or whole genome sequencing data of the genome represented by nucleic acid fragments in cell-free or cell-based samples, partial or whole genome sequencing data including one or more types of epigenetic modifications (e.g., methylation), or combinations thereof.
[0040] Data acquired by the data collection component 10 may be transferred to database 20 via network 40 or a local or network connection. In some embodiments, data collection component 10 may alternatively receive data from one or more sequencing devices. In some embodiments, the collected data may be analyzed by data intelligence component 30, via network 40 or a local or network connection. FIG. 4B depicts exemplary functional modules that may be implemented to perform tasks of data intelligence component 30.
[0041] FIG. 4B depicts an exemplary computer system 410 for performing marking of duplicate reads in long RNA sequence data. Exemplary system 410 achieves such functionalities by implementing, on one or more computer devices, user input and output (I / O) module 420, memory or database 430, data processing module 440, data analysis module 450, classification module 460, network communication module 470, and any other functional modules that may be needed for carrying out a particular task (e.g., an error correction or compensation module, a data compression module, etc.). As disclosed herein, user I / O module 420 may further include an input sub-module, such as a keyboard, and an output sub-module, such as a display (e.g., a printer, a monitor, or a touchpad). In some embodiments, all functionalities may be performed by one computer system. In some embodiments, the functionalities are performed by more than one computer system.
[0042] Also disclosed herein, a particular task may be performed by implementing one or more functional modules. In particular, each of the enumerated modules itself may, in turn, include multiple sub-modules. For example, data processing module 440 may include a sub-module for data quality evaluation (e.g., for discarding very short sequence reads or sequence reads including obvious errors), a sub-module for normalizing numbers of sequence reads that align to different regions of a reference genome, a sub-module to compensate / correct GO biases, a sub-module for marking duplicate reads in a raw dataset, etc.
[0043] In some embodiments, a user may use I / O module 420 to manipulate data that is available either on a local device or can be obtained via a network connection from a remote service device or another user device. For example, I / O module 420 may allow a user, e.g., via a keyboard, a mouse, or a touchpad, to initiate or perform data analysis via a graphical user interface (GUI). In someembodiments, a user may manipulate data via voice control. In some embodiments, user authentication may be required before a user is granted access to the data being requested. In some embodiments, user I / O module 420 may be used to manage various functional modules. For example, a user may request via user I / O module 420 input data while an existing data processing session is in process. A user may do so by selecting a menu option or type in a command discretely without interrupting the existing process. In another example, a user may utilize user I / O module 420 to set various thresholds, configure duplicate marking settings (e.g., adjust the size of the clip-pad, modify a function of one or more indices, etc.) and / or provide other instructions to computer system 410 that facilitate how sequence reads are identified, aligned, processed, and / or marked as duplicates, for example. As disclosed herein, a user may use any suitable type of input to direct and control data processing and analysis via I / O module 420.
[0044] In some embodiments, system 410 further comprises a memory or database 430. In some embodiments, database 430 comprises a local database that may be accessed via user I / O module 420. In some embodiments, database 430 comprises a remote database that may be accessed by user I / O module 420 via network connection. In some embodiments, database 430 is a local database that stores data retrieved from another device (e.g., a user device or a server). In some embodiments, memory or database 430 may store data retrieved in real-time from internet searches. In some embodiments, memory or database 430 may store data received from the duplicate marking tool (e.g., shard output data representing identified duplicates, etc.). In some embodiments, database 430 may send data to and receive data from one or more of the other functional modules, including, but not limited to, a data collection module (not shown), data processing module 440, dataanalysis module 450, classification module 460, network communication module470, and etc.
[0045] In some embodiments, database 430 may be a database local to the other functional modules. In some embodiments, database 430 may be a remote database that may be accessed by the other functional modules via wired or wireless network connection (e.g., via network communication module 470). In some embodiments, database 430 may include a local portion and a remote portion.
[0046] In some embodiments, system 410 comprises a data processing module 440. Data processing module 440 may receive data from I / O module 420 or database 430. In some embodiments, data processing module 440 may perform standard data processing algorithms, such as one or more of noise reduction, signal enhancement, normalization of counts of sequence reads, correction of GC bias, etc. In some embodiments, data processing module 440 may be configured to leverage one or more types of duplicate marking tools, e.g., those tools configured to perform duplicate marking by sorting, sharding, or some combination of the foregoing. Upon receiving indications of duplicate reads, data processing module 440 may apply marking tags designating the relevant read pairs as duplicate and / or may computationally remove those reads from a dataset and / or prevent the duplicate reads from being considered in one or more downstream processes (e.g., feature selection, classifier training, etc.).
[0047] In some embodiments, given sequencing data, data processing module 440 may be configured to stream the sequencing data in batches such that a limited number of records of the sequencing data are loaded in the memory or database 430 that can be processed in parallel. Records may include a sequenceread, an example of which is a pair of sequence reads, hereafter referred to as read pair (e g., n and r2).
[0048] In some embodiments, data processing module 440 may be configured to pre-filter sequence reads, e.g., to reduce the overall size of the data pool for downstream analysis (e.g., as further described with reference to FIG. 6). The pre-filtering process may ultimately reduce the consumption of computer resources and may improve the accuracy of a downstream analysis. More particularly, FIG. 5 depicts a pre-filtering process 500 for removing sequence reads that meet one or more criteria, in accordance with an embodiment. Examples of the criteria include, but are not limited to: identifying whether a sequence read is a singleton, identifying whether a sequence read is a hard clip, filtering based on a template length (TLEN) (e.g., a threshold TLEN), filtering based on an alignment score (e.g., a threshold alignment score), or filtering based on a base quality score (e.g., a threshold of a median or mean base quality score). Another criteria includes determining that if a sequence read pair meets the criterion that the reads of the read pair are from differing chromosomes, then the sequence read pair is maintained and not filtered out. Additional examples of criteria include filtering based on a bit flag, a cigar, an edit distance (e.g., a minimum or maximum edit distance), a suboptimal alignment score, or a supplementary alignment measure.
[0049] In some embodiments, the pre-filtering process 500 may involve removing sequence reads that meet one or more criteria to identify a subset of filtered sequence reads. In some embodiments, the pre-filtering process 500 additionally or alternatively involves maintaining sequence reads that meet one or more criteria. In various embodiments, additional or fewer steps than shown in FIG. 5 for filtering sequence reads that meet one or more criteria may beimplemented. Additionally, although FIG. 5 depicts a flow process where each of steps 505-525 are shown to be performed in series, in various embodiments, one or more of the different steps 505-525 may be performed in parallel. For example, steps 510, 515, and 520 may be performed in parallel and the resulting, non-filtered sequence reads may then be further analyzed at step 525.
[0050] As shown in FIG. 5, at step 505, singletons and hard clips are filtered out. A singleton refers to a single read of a read pair. Without the second read of a read pair, a singleton may not be informative for various downstream analytical processes and may be readily filtered out. A hard clip refers to a sequence read that has been truncated.
[0051] At step 510, sequence read pairs where the first read and the second read are from differing chromosomes are maintained. In various embodiments, the first read and the second read may be determined to be from differing chromosomes by aligning the first read and the second read to known sequences of chromosomes (e.g., a reference genome). In various embodiments, sequence read pairs that are identified and maintained here at step 510 are further maintained through steps 515, 520, and 525.
[0052] At step 515, sequence reads may be filtered based on the template length (TLEN) for each sequence read. Generally, sequence reads with TLENs greater than a threshold TLEN are maintained whereas sequence reads with TLENs less than the threshold TLEN are filtered and removed. The threshold TLEN can be one of 50, 100, 200, 300, 400, or 500.
[0053] At step 520, sequence reads may be filtered based on the alignment score for each sequence read. The alignment score for a sequence read captures information about how well a read aligns to a reference genome by capturinginformation such as matches, mismatches, open gaps (e.g., extends), clipping and improper pairs. As an example, the alignment score (AS) may be expressed as:AS=(1 *matches)-(4*mismatches)-(6*opens)-(1 *extends)-(5*clipped)-(17*im proper_pair)
[0054] Therefore, a high alignment score represents a well-aligned sequence read, whereas a low alignment score represents a poorly aligned sequence read.
[0055] A sequence read may be filtered by comparing the alignment score for the sequence read to a threshold alignment score. Generally, sequence reads with alignment scores less than a threshold alignment score are maintained whereas sequence reads with alignment scores greater than the threshold alignment score are filtered and removed. In various embodiments, the threshold alignment score is one of 50, 75, 100, 125, or 140. In particular embodiments, the threshold alignment score is 130.
[0056] At step 525, sequence reads may be filtered based on the base quality score for each sequence read. Here, the base quality for a sequence read refers to the Phred quality score of nucleotide bases within the sequence read. In one embodiment, the base quality score is represented by a median base quality score, which is calculated as the median of Phred quality scores of nucleotide bases of the sequence read. Generally, sequence reads that each have a median base quality score greater than a threshold median base quality score are maintained, whereas sequence reads that each have a median base quality score less than the threshold base quality score are filtered and removed. In various embodiments, the threshold alignment score is one of 10, 25, 35, or 40. In particular embodiments, the threshold alignment score is 30.
[0057] In some embodiments, system 410 comprises a data analysis module 450. In some embodiments, data analysis module 450 includes identifying and treating systematic errors in sequencing data, as described in connection with data processing module 440.
[0058] In some embodiments, system 410 comprises a classification module 460, which may embody a “machine-learning model” or “trained classifier.” As used herein, a “machine-learning model” or “trained classifier” generally encompasses instructions, data, and / or a model configured to receive input, and apply one or more of a weight, bias, classification, or analysis on the input to generate an output. The output may include, for example, a classification of the input, an analysis based on the input, a design, process, prediction, or recommendation associated with the input, or any other suitable type of output. A machine-learning model is generally trained using training data, e.g., experiential data and / or samples of input data, which are fed into the model in order to establish, tune, or modify one or more aspects of the model, e.g., the weights, biases, criteria for forming classifications or clusters, or the like. Aspects of a machine-learning model may operate on an input linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration.
[0059] The execution of the machine-learning model may include deployment of one or more machine-learning techniques, such as k-nearest neighbors, linear regression, logistic regression, random forest, gradient boosted machine (GBM), deep learning, a deep neural network, and / or any other suitable machine-learning technique that solves problems in the field of bioinformatics. Supervised, semi-supervised, and / or unsupervised training may be employed. For example, supervised learning may include providing training data and labelscorresponding to the training data, e.g., as ground truth. Unsupervised approaches may include clustering, classification or the like. K-means clustering or K-Nearest Neighbors may also be used, which may be supervised or unsupervised. Combinations of K-Nearest Neighbors and an unsupervised cluster technique may also be used. Any suitable type of training may be used, e.g., stochastic, gradient boosted, random seeded, recursive, epoch or batch-based, etc.
[0060] In an exemplary use case, a machine-learning model may be trained to analyze data from a test sample from a test subject whose status with respect to a medical condition is unknown and subsequently classifies the unknown test sample from the test subject based on the likelihood of the subject fitting into a particular category. In some embodiments, the one or more parameters may include a binomial probability score that is calculated based on logistic regression analysis. As disclosed herein, the binomial probability score may correspond to the likelihood of a subject having a certain medical condition, such as cancer. For example, a score of over a predefined threshold may indicate that the subject associated with a test sample is more likely to have cancer than not have cancer. In some embodiments, the one or more parameters may include a sequencing or methylation data distribution pattern correlating with the presence of cancer. A subject associated with a test sample having sequencing or methylation data with a pattern resembling the cancer pattern to a sufficient degree may be predicted as having cancer. In some embodiments, a sequencing or methylation data distribution pattern may be identified in connection with a specific type of cancer, determining a tissue of origin or cancer signal origin, thus allowing a test sample to be classified as indicative of a certain cancer type.
[0061] As disclosed herein, network communication module 470 may be used to facilitate communications between a user device, one or more databases, and any other suitable system or device through a wired or wireless network connection. Any communication protocol / device may be used, including, without limitation, a modem, an Ethernet connection, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication facilities, etc.), a near-field communication (NFC), a Zigbee communication, a radio frequency (RF) or radio-frequency identification (RFID) communication, a PLC protocol, a 3G / 4G / 5G / LTE based communication, and / or the like. For example, a user device having a user interface platform for processing / analyzing sequence reads may communicate with another user device with the same platform, a regular user device without the same platform (e.g., a regular smartphone), a remote server, a physical device of a remote loT local network, a wearable device, a user device communicably connected to a remote server, and etc.
[0062] The functional modules described herein are provided by way of example. It will be understood that different functional modules may be combined to create different utilities. It will also be understood that additional functional modules or sub-modules may be created to implement a certain utility.
[0063] Referring now to FIG. 6, an exemplary workflow 600 is provided for performing duplicate marking on long nucleotide reads, e.g., long RNA reads, utilizing a sharding approach with clip-padding. Aspects of the exemplary workflow 600 may be performed in accordance with some or all components described inFIGS. 4A and 4B.
[0064] At step 605, a BAM file that stores sequence alignment data may be accessed and scanned. The BAM file is a binary file that is a compressed and indexed representation of DNA or RNA sequence alignments against a reference genome. For the purposes of this discussion, RNA sequences present in the BAM file are the focus of the concepts described herein, but this designation is not limiting. In an aspect, the BAM file may be accompanied by an index file, e.g., a Binary Alignment Index (“BAI”), which allows for rapid random access to specific regions of the alignment. This is particularly important when working with large datasets, as it facilitates efficient retrieval of relevant portions of the alignment without having to read the entire file. In an aspect, each alignment in the BAM file may be represented by a Compact Idiosyncratic Gapped Alignment Report (“CIGAR”) string, which describes how the sequence aligns to the reference genome, including matches, mismatches, insertions, and / or deletions.
[0065] At step 610, computer system 410 may leverage a duplicate marking algorithm to initiate a sharding process on the RNA alignment data present in the BAM file. In an aspect, before sharding begins, the genomic data may be divided into segments known as shards. The definition of these shards is based on specific criteria, e.g., genomic coordinates, to ensure that each shard encompasses a distinct genomic region. In some aspects, the algorithm may determine the genomic boundaries (e.g., utilizing genomic coordinates) for each shard, ensuring that the partitioned data is logically separated based on the specific requirement of the duplicate marking process. The genomic data is then distributed among different shards based on the determined boundaries. Each shard is responsible for processing a specific portion of the overall genomic dataset.
[0066] In an aspect of the present disclosure, shards may be processed independently and in parallel. This parallel processing allows for efficient utilization of computational resources, enabling the analysis of large datasets in a more timeeffective manner. In some aspects, while shards operate independently, there may be a need for communication or data sharing between shards in specific steps of the algorithm. For example, in resolving conflicts arising from distant mates, information about mates and their friends may need to be shared between shards.
[0067] More particularly to the foregoing, the operation of shards may involve multiple threads within the same process. This means that the threads may share memory and execute concurrently on the same virtual machine. While they do not necessarily run on the same processor core, they may utilize shared memory for communication and coordination. With respect to information sharing among threads, the parallel execution may be designed to run on a single machine without the need for real-time or supervisory processes to ensure information sharing. Rather, sharing of inter-thread information may rely on standard thread synchronization mechanisms, such as locks or mutexes, to manage access to shared data structures. These mechanisms facilitate the maintenance of data integrity while allowing efficient parallel processing of the input data.
[0068] At step 615, computer system 410 may identify distant mate pairs in the RNA alignment data present in the BAM file. As defined above, distant mate pairs refer to pairs of reads in which one read starts in a different shard from its mate due, for instance, to the presence of the large, intron-spanning genomic distances. In the context of RNA data with long reads and potential genomic skips (e.g., resulting from introns), it’s common for read pairs to span significant genomic distances and start in separate shards from one another.
[0069] In an aspect, the identification of the distant mate pairs may begin during the initial pass through the BAM file. As the duplicate marking algorithm traverses through the file, it keeps track of distant mates and their corresponding distant mate pairs. In an aspect, the algorithm may scan the BAM file to identify cases in which a pair of reads has one read starting in one shard while its distant mate pair starts in a different shard (e.g., a different adjacent or non-adjacent shard). This scanning may involve inspecting the positions, orientations, and other relevant information of the reads in the pair. In an aspect, during the identification process, the unclipped 5' positions (or other relevant positions) of the distant mates in the distant mate pairs may be recorded. The unclipped 5' position refers to the genomic position before any clipping occurs (e.g., removal of bases due to alignment adjustments). These recorded positions may be later used for the injection of reads into the respective shards during the sharding process. Once identified, the distant mate pairs, along with their recorded positions, are collected into a first data set or data structure, e.g., known as a “distantMatesPositions” data set, for subsequent use.
[0070] An example data structure for storing the distant mate pairs may be one that offers efficient lookup and retrieval characteristics while avoiding memory overheard. For instance, in one aspect, the data structure may be a hash table or hashmap. Hash tables offer constant-time average-case lookup, insertion, and deletion operations, making them efficient for storing key-value pairs. In the instant case, the distant mate pairs’ positions, orientations, and / or associated information may be stored as keys in the hashmap, with the corresponding values representing any associated metadata or auxiliary information. In another aspect, the data structure may be a balanced tree structure, such as a balanced binary search tree.While insertion and lookup operations in balanced trees may have logarithmic time complexity, they provide deterministic performance guarantees and may be more suitable if memory usage needs to be controlled.
[0071] At step 620, computer system 410 may identify “friends” of the distant mates in the RNA alignment data present in the BAM file, as described below. In an aspect, the duplicate marking algorithm may perform a second pass through the BAM file to identify reads that share unclipped 5' positions with the identified distant mate pairs present in the distantMatesPositions set. More particularly, for each read in the BAM file, the algorithm may compare its unclipped 5' position and orientation with the unclipped 5' position of the reads in the distantMatesPositions set. Reads that share both unclipped 5' positions and orientations with the distant mate pairs are identified as “distant mate friends.” These reads are not distant mates themselves, but have similarities with the distant mate pairs. In an aspect, the duplicate marking algorithm records or stores these identified “distant mates’ friends” in a second and separate data structure, e.g., a “distantMatesFriends” data set. The second data structure may be of the same, or different type, as the first data structure. In some aspects, the algorithm may record additional information about the distant mates’ friends, such as their read IDs, genomic locations, or any other relevant details that may assist in subsequent steps of the duplicate marking process.
[0072] Both the “distantMatesPositions” and “distantMatesFriends” data sets are accessible to all shards. This accessibility ensures that information about distant mates and their distant mate friends is available across the parallel processing shards during duplicate marking.
[0073] At step 625, computer system 410 may inject certain reads into relevant shards to address conflicts that may arise when distant mate pairs areencountered. In an aspect, the injection of reads into a shard may be governed by specific criteria. Specifically, only reads that share endpoints (e.g., unclipped 5' left, unclipped 5' right) with the distant mate pair are injected into the shard. For example, referring to FIG. 3, if read P2R1 in shardl is a distant mate of read P2R2 in shard2, then all reads that share endpoints with P2R1 + P2R2 may be injected into shard2. The shards that require pulling information from the “distantMatesPositons” data set also need to pull information from the “distantMatesFriends” data set. In an aspect, the injection may be targeted and selective, so that only relevant reads are brought into the shard for conflict resolution. In an aspect, the injection process may be performed iteratively for each distant mate pair, resolving conflicts across the dataset and allowing for comprehensive duplicate marking.
[0074] In an aspect, the injection process in the context of the duplicate marking algorithm may involve incorporating additional position and orientation information from distant mates and distant mates’ friends into the shard’s data structure. During the processing of each shard, the full distant mates and distant mates’ friends data structures are utilized to construct a comprehensive table containing position and orientation information for all reads within the shard. This table serves as the primary input to the final duplicate-marking function. It is important to note that although the full distant mates and distant mates’ friends data structures are accessed during the construction of the table, they are not directly utilized by the final duplicate-marking function itself. Instead, the injection position and orientation information from these data structures may be integrated into the table, enhancing the accuracy and completeness of the duplicate marking process without directly involving the distant mates and distant mates’ friends data structures in the final function’s operations.
[0075] After the injection of reads into the relevant shard(s), computer system 410 may continue with the normal duplicate marking process within the shard. The injected reads, marked as duplicates, may be handled independently by their own shard during the subsequent steps. In an aspect, if a read is identified as a duplicate, the algorithm may flag or mark it as such. This may involve adding a specific tag or flag to the read’s alignment record to indicate its duplicate status. Ultimately, the output of the duplicate marking process within a shard may include information about which reads are marked as duplicates. This information may be used for downstream analysis, quality control, and interpretation of the sequencing data.
[0076] Referring now to FIG. 7, a table 700 is provided that illustrates the performance improvements resulting from the duplicates marking process described herein. More particularly, an objective of the underlying disclosure is to reduce the collective size (e.g., shard + clip-padding) of the shards undergoing processing. Accomplishing that objective with known techniques was not possible. For instance, table 700 illustrates column 705, which presents data for processing RNA data using conventionally long clip-pads, and column 710, which presents data for processing RNA data using much smaller default parameters. Upon examination of the table, significant reductions in both processing time and computational resource utilization were witnessed across the board. For instance, a shard size of 5 million bases that was processed utilizing conventional means contained a clip-padding of approximately 1 .5 million bases, took approximately 28 minutes to process, required 23.1 cpu to facilitate processing, and required a storage capacity in memory of 244.2 gibibytes (GiB). Through the implementation of the novel duplicates marking process described herein, an RNA data set is able to be processed more efficiently. Forinstance, column 710 indicates that a shard size of 5 million bases that was processed utilizing the systems and methods described herein was able to be processed in approximately 14 minutes, required 16.3 cpu to facilitate processing, and required a storage capacity in memory of just 37.3 GiB, i.e., a nearly 6-fold memory requirement reduction at a faster speed.
[0077] In general, any process discussed in this disclosure that is understood to be computer-implementable may be performed by one or more processors of a computer system, such as system environment 110, as described above. A process or process step performed by one or more processors may also be referred to as an operation. The one or more processors may be configured to perform such processes by having access to instructions (e.g., software or computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The instructions may be stored in a memory of the computer server. A processor may be a central processing unit (CPU), a graphics processing unit (GPU), or any suitable types of processing unit.
[0078] A computer system, such as system environment 110, may include one or more computing devices. If the one or more processors of the computer system are implemented as a plurality of processors, the plurality of processors may be included in a single computing device or distributed among a plurality of computing devices. If a system environment comprises a plurality of computing devices, the memory of the computer system may include the respective memory of each computing device of the plurality of computing devices.
[0079] FIG. 8 is a simplified functional block diagram of a computer system800 that may be configured as a computing device for executing the processesdescribed herein, according to exemplary embodiments of the present disclosure. FIG. 8 is a simplified functional block diagram of a computer that may be configured according to exemplary embodiments of the present disclosure. In various embodiments, any of the systems herein may be an assembly of hardware including, for example, a data communication interface 820 for packet data communication. The platform also may include a central processing unit (“CPU”) 802, in the form of one or more processors, for executing program instructions. The platform may include an internal communication bus 608, and a storage unit 806 (such as ROM, HDD, SDD, etc.) that may store data on a computer readable medium 822, although the system 800 may receive programming and data via network communications via a network 825 (e.g., voice, video, audio, images, or any other data over the network 825). The system 800 may also have a memory 804 (such as RAM) storing instructions 824 for executing techniques presented herein, although the instructions 824 may be stored temporarily or permanently within other modules of system 800 (e.g., processor 802 and / or computer readable medium 822). The system 800 also may include input and output ports 812 and / or a display 810 to connect with input and output devices such as keyboards, mice, touchscreens, monitors, displays, etc. The various system functions may be implemented in a distributed fashion on a number of similar platforms, to distribute the processing load. Alternatively, the systems may be implemented by appropriate programming of one computer hardware platform.
[0080] In this disclosure, the term “based on” means “based at least in part on.” The singular forms “a,” “an,” and “the” include plural referents unless the context dictates otherwise. The term “exemplary” is used in the sense of “example” rather than “ideal.” The terms “comprises,” “comprising,” “includes,” “including,” orother variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, or product that comprises a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent to such a process, method, article, or apparatus. Relative terms, such as “about,” “approximately,” “substantially,” and “generally,” are used to indicate a possible variation of ±10% of a stated or understood value. In addition, the term “between” used in describing ranges of values is intended to include the minimum and maximum values described herein. The use of the term “or” in the claims and specification is used to mean “and / or” unless explicitly indicated to refer to alternatives only if the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” As used herein “another” may mean at least a second or more.
[0081] As used herein, the term “user” generally encompasses any person or entity, such as a researcher and / or a care provider (e.g., a doctor, etc.), that may desire information, resolution of an issue, or engage in any other type of interaction with a provider of the systems and methods described herein (e.g., via an application interface resident on their electronic device, etc.). The term “electronic application” or “application” may be used interchangeably with other terms like “program,” or the like, and generally encompasses software that is configured to interact with, modify, override, supplement, or operate in conjunction with other software.
[0082] Program aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of executable code and / or associated data that is carried on or embodied in a type of machine-readable medium. “Storage” type media include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as varioussemiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer of the mobile communication network into the computer platform of a server and / or from a server to the mobile device. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0083] Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0084] Thus, while certain embodiments have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as falling within the scope of the invention. For example,functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present invention.
[0085] The above disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other implementations, which fall within the true spirit and scope of the present disclosure. Thus, to the maximum extent allowed by law, the scope of the present disclosure is to be determined by the broadest permissible interpretation of the following claims and their equivalents, and shall not be restricted or limited by the foregoing detailed description. While various implementations of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more implementations are possible within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented method, the computer-implemented method comprising: accessing, using a processor of a computing device, a data file containing aligned sequencing data; initiating, using the processor, a duplicates marking process on the aligned sequencing data, wherein the initiating comprises segmenting the aligned sequencing data into a plurality of shards; identifying, using the processor and in a first scan of the data file, a first set of data associated with a distant mate pair in the aligned sequencing data; storing, using the processor, the first set of data into a first data structure; identifying, using the processor and in a second scan of the data file, a second set of data corresponding to a distant mate friend associated with the distant mate pair; storing, using the processor, the second set of data into a second data structure; identifying, using the processor and by accessing the first set of data in the first data structure, a conflict between two shards of the plurality of shards; and resolving the conflict by accessing the second data structure and introducing sequence information contained in the second set of data into one of the two shards needing the sequence information.
2. The computer-implemented method of claim 1 , wherein the data file is a Binary Alignment Map (BAM) file.
3. The computer-implemented method of claim 1 , wherein the distant mate pair corresponds to a pair of sequence reads where one read in the pair is aligned to a first shard in the plurality of shards and another read is aligned to a second shard in the plurality of shards.
4. The computer-implemented method of claim 1 , wherein the first set of data includes unclipped 5' positions of each pair in the distant mate pair.
5. The computer-implemented method of claim 4, wherein the distant mate friend shares the unclipped 5' position for at least one pair in the distant mate pair.
6. The computer-implemented method of claim 1 , wherein each of the plurality of shards comprises a clip-padding on each end, wherein the clip-padding extends coverage of each of the plurality of shards to encompass additional upstream and downstream genomic positions.
7. The computer-implemented method of claim 1 , wherein the first data structure and the second data structure are accessible to each of the plurality of shards.
8. The computer-implemented method of claim 1 , wherein the introducing comprises introducing the sequence information iteratively for each of the distant mate pairs.
9. The computer-implemented method of claim 1 , wherein the introducing comprises introducing according to predetermined criteria.
10. The computer-implemented method of claim 9, wherein the predetermined criteria include a shared 5’ endpoint and a matching orientation between sequence reads in the first data structure and the second data structure.11 . The computer-implemented method of claim 1 , further comprising providing, at a conclusion of the duplicates marking process, an output comprising duplicate information contained in the aligned sequencing data.
12. A system, comprising: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations comprising: accessing, using the one or more processors of a computing device associated with the system, a data file containing aligned sequencing data; initiating, using the one or more processors, a duplicates marking process on the aligned sequencing data, wherein the initiating comprises segmenting the aligned sequencing data into a plurality of shards; identifying, using the one or more processors and in a first scan of the data file, a first set of data associated with a distant mate pair in the aligned sequencing data; storing, using the one or more processors, the first set of data into a first data structure;identify ing , using the one or more processors and in a second scan of the data file, a second set of data corresponding to a distant mate friend associated with the distant mate pair; storing, using the one or more processors, the second set of data into a second data structure; identifying, using the one or more processors and by accessing the first set of data in the first data structure, a conflict between two shards of the plurality of shards; and resolving the conflict by accessing the second data structure and introducing sequence information contained in the second set of data into one of the two shards needing the sequence information.
13. The system of claim 12, wherein the data file is a Binary Alignment Map (BAM) file.
14. The system of claim 12, wherein the distant mate pair corresponds to a pair of sequence reads where one read in the pair is aligned to a first shard in the plurality of shards and another read is aligned to a second shard in the plurality of shards.
15. The system of claim 12, wherein the first set of data includes unclipped 5' positions of each pair in the distant mate pair, wherein the distant mate friend shares the unclipped 5' position for at least one pair in the distant mate pair.
16. The system of claim 12, wherein each of the plurality of shards comprises a clip-padding on each end, wherein the clip-padding extends coverage of each of the plurality of shards to encompass additional upstream and downstream genomic positions.
17. The system of claim 12, wherein the first data structure and the second data structure are accessible to each of the plurality of shards.
18. The system of claim 12, wherein the introducing comprises introducing the sequence information iteratively for each of the distant mate pairs.
19. The system of claim 12, wherein the introducing comprises introducing according to predetermined criteria.
20. The system of claim 19, wherein the predetermined criteria include a shared 5’ endpoint and a matching orientation between sequence reads in the first data structure and the second data structure.
21. The system of claim 12, further comprising providing, at a conclusion of the duplicates marking process, an output comprising duplicate information contained in the aligned sequencing data.
22. A non-transitory computer-readable medium storing computer-executable instructions which, when executed by a system, cause the system to perform operations comprising:accessing, using a processor of a computing device, a data file containing aligned sequencing data; initiating, using the processor, a duplicates marking process on the aligned sequencing data, wherein the initiating comprises segmenting the aligned sequencing data into a plurality of shards; identifying, using the processor and in a first scan of the data file, a first set of data associated with a distant mate pair in the aligned sequencing data; storing, using the processor, the first set of data into a first data structure; identifying, using the processor and in a second scan of the data file, a second set of data corresponding to a distant mate friend associated with the distant mate pair; storing, using the processor, the second set of data into a second data structure; identifying, using the processor and by accessing the first set of data in the first data structure, a conflict between two shards of the plurality of shards; and resolving the conflict by accessing the second data structure and introducing sequence information contained in the second set of data into one of the two shards needing the sequence information.
Citation Information
Patent Citations
Methods and apparatus for processing circulating tumor DNA repetitive sequences
CN108229103B
Compositions and methods for identification of a duplicate sequencing read
US20150132763A1