Variant Calling for High-Coverage Samples with Limited Memory
The variant caller subsystem addresses memory allocation issues in sequencing determination platforms by dividing or spilling callable regions, ensuring efficient and accurate variant calling operations without crashes.
Patent Information
- Application Number
- JP2024575576
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-24
- Filing Date
- 2023-06-23
- Publication Date
- 2025-07-23
AI Technical Summary
Sequencing determination platforms face memory allocation issues when processing large sequencing determination data, leading to delays or crashes due to exceeding memory limits, which affects the processing of sequencing data and other applications on computing systems.
Implementing a variant caller subsystem that identifies callable regions and performs variant calling within allocated memory by dividing or spilling these regions to manage memory usage effectively, utilizing a calling subsystem to monitor memory thresholds and split or spill data to disk when necessary, ensuring efficient processing without exceeding memory limits.
This approach allows for effective management of memory resources, preventing crashes and maintaining processing efficiency by dividing or spilling callable regions, thereby ensuring accurate and timely variant calling and base calling operations.
Smart Images

Figure 2025523521000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 355,541, filed on June 24, 2022, which is hereby incorporated by reference in its entirety.
Background Art
[0002] In recent years, biotechnology companies and research institutions have been improving hardware and software platforms for determining the sequence of nucleotide bases (or entire genomes) and identifying variant calls for nucleotide bases that differ from the reference bases of a reference genome. Sequencing platforms can monitor tens of thousands or more oligonucleotides to detect more accurate nucleotide base calls from larger base call data sets. For example, the cameras of such sequencing platforms can capture images of fluorescent tags irradiated from nucleotide bases incorporated into such oligonucleotides. After capturing such images, the sequencing platform transmits the data to a computing device having sequencing data analysis software that aligns nucleotide reads to the reference genome. Based on the aligned nucleotide fragment reads, the sequencing platform can determine nucleotide base calls for genomic regions and identify variants within the nucleic acid sequence of a sample.
[0003] As the amount of sequencing determination data that can be analyzed continues to increase, there are issues when the sequencing determination platform operates within a specific hardware allocation of the computing system on which it runs. For example, when the sequencing determination data is loaded into memory, the size of the sequencing determination data may exceed a specific memory allocation, potentially causing delays or crashes in the processing of the sequencing determination platform and / or other applications running on the computing system. The sequencing determination platform will need to continue to expand to operate within such an allocation and handle the increasing sequencing determination data being analyzed.
Summary of the Invention
Means for Solving the Problem
[0004] Systems, methods, and apparatuses are described herein for identifying callable regions and performing variant calling while operating within an allocated memory. The sequencing determination subsystem may comprise a secondary analysis subsystem implemented on one or more devices for performing secondary analysis of the sequencing determination data. For example, the sequencing determination subsystem may comprise a variant caller or a variant caller subsystem. The variant caller may comprise a calling subsystem configured to identify callable regions to be processed for performing variant calling and / or base calling. The calling subsystem can transmit the callable regions to a genotyping subsystem downstream of the variant caller. The genotyping subsystem can perform variant calling and / or base calling within the callable regions.
[0005] The calling subsystem of the variant caller can receive sequencing data that includes multiple reads of a genomic sequence. The calling subsystem of the variant caller may be configured to detect callable regions of the sequencing data when the depth of the multiple reads exceeds a callable region depth threshold. The calling subsystem of the variant caller can monitor its memory usage. The memory threshold may be a fixed threshold or a dynamic threshold. When the memory used by the calling subsystem of the variant caller exceeds the memory threshold of the total amount of memory allocated to the calling subsystem, the calling subsystem can divide the callable region and send the divided portions of the callable region to the genotyping subsystem of the variant caller for variant calling based on the divided portions. Sending the divided portions of the callable region to the genotyping subsystem increases the availability of the memory used by the calling subsystem.
[0006] When the memory used by the calling subsystem of the variant caller is within the memory threshold of the total amount of memory allocated to the calling subsystem of the variant caller, the calling subsystem can analyze the sequencing data to identify insertions, deletions, or other variants or mutations in the sequencing data within a pre-defined proximity of the identified splits. After identifying a variant or mutation in the sequencing data, the calling subsystem can divide callable regions outside the pre-defined proximity of the variant or mutation. The variant or mutation and / or the pre-defined proximity of the identified splits can be determined based on population data accessed by the calling subsystem.
[0007] The calling subsystem can analyze the buffered sequencing data to identify positions for the division of callable regions within the buffered sequencing data. The calling subsystem can identify a portion of the buffered sequencing data having a read depth below a division threshold and perform a division of the buffered sequencing data at the identified portion having a read depth below the division threshold. The divided portion of the callable region may be the entire callable region currently present in the memory used by the calling subsystem. In another example, the calling subsystem can maintain a predefined amount of the sequencing data within the first divided portion of the callable region in the memory used by the variant caller's calling subsystem, such that the first and second divided portions of the callable region have an overlap in the sequencing data. The overlap may include a predefined number of bases. The overlap can be determined based on population data or user input accessed by the calling subsystem. The genotyping subsystem can remove the overlap in the sequencing data between the divided portions of the callable region before performing variant calling on the callable region.
[0008] In another exemplary embodiment, if the memory used by the variant caller's calling subsystem is within a memory threshold of the total amount of memory allocated to the variant caller's calling subsystem, the calling subsystem can spill the callable region to disk storage. The spilled callable region can then be streamed back from disk into the memory used by the genotyping subsystem for processing. The entire callable region may be spilled to disk, or a portion of the callable region may be spilled to disk while a second portion of the callable region remains in memory. The genotyping subsystem can analyze the spilled callable region streamed back from disk and discard one or more portions of the spilled callable region from memory not used for variant calling before streaming additional portions of the spilled callable region into memory. The genotyping subsystem may have a separate allocation of memory and can monitor a second memory threshold associated with the genotyping subsystem while streaming the callable region spilled from disk to prevent exceeding the second memory threshold.
Brief Description of the Drawings
[0009]
Figure 1A
Figure 1B
Figure 2A
Figure 2B
Figure 2C
Figure 2D
Figure 2E
Figure 3
Figure 4
Figure 5
DETAILED DESCRIPTION OF THE INVENTION
[0010] FIG. 1A shows a schematic diagram of a system environment (or “environment”) 100 described herein. As shown, environment 100 includes one or more server devices 102 connected to client device 108 and sequencing device 114 via network 112.
[0011] As shown in FIG. 1A, server device 102, client device 108, and sequencing device 114 can communicate with each other via network 112. Network 112 can include any suitable network over which computing devices can communicate. Network 112 can include wired and / or wireless communication networks. Exemplary wireless communication networks can be composed of one or more types of radio frequency (RF) communication signals using one or more wireless communication protocols such as cellular communication protocols, wireless local area network (WLAN) or WIFI communication protocols, and / or other wireless communication protocols. In addition to, or instead of, communicating through network 112, server device 102, client device 108, and / or sequencing device 114 may bypass network 112 and communicate directly with each other.
[0012] As shown by FIG. 1A, the sequencing device 114 may comprise a device for sequencing a biological sample. The biological sample may include human and non-human deoxyribonucleic acid (DNA) for determining the individual nucleotide bases of a nucleic acid sequence (e.g., sequencing by synthesis). The biological sample may include human and non-human ribonucleic acid (RNA). The sequencing device 114 may utilize the computer-implemented methods and systems described herein to analyze nucleic acid segments and / or oligonucleotides extracted from a sample, either directly or indirectly on the sequencing device 114, to generate nucleotide reads and / or other data. More specifically, the sequencing device 114 may receive and analyze nucleic acid sequences extracted from a sample within a nucleotide sample slide (e.g., a flow cell). The sequencing device 114 may utilize sequencing by synthesis (SBS) to sequence nucleic acid segments into nucleotide reads.
[0013] As further shown by FIG. 1A, the server device 102 may generate, receive, analyze, store, and / or transmit digital data such as data for determining nucleotide base calls or for sequencing nucleic acid polymers. As shown in FIG. 1A, the sequencing device 114 may generate and transmit nucleotide reads and / or other data for analysis by the server device 102 for base calling and / or variant calling (which the server device 102 may receive). The server device 102 may also communicate with the client device 108. Specifically, the server device 102 may transmit data including sequencing data or other information to the client device 108, and the server device 102 may receive input from a user via the client device 108.
[0014] Server device 102 may comprise a distributed aggregate of servers, in which case server device 102 comprises several server devices that are distributed through network 112 and are located in the same or different physical locations. Further, server device 102 may include a content server, an application server, a communication server, a web hosting server, or another type of server.
[0015] As further shown in FIG. 1A, server device 102 and / or sequencing device 114 may comprise a sequencing subsystem 104. Sequencing subsystem 104 may be implemented as hardware and / or software on one or more devices for performing secondary analysis of sequencing data. The sequencing subsystem may be implemented as a secondary analysis subsystem on one or more devices for performing secondary analysis. For example, sequencing subsystem 104 can analyze nucleotide reads such as sequencing metrics received from sequencing device 114 and / or other data to determine the nucleotide base sequence for a nucleic acid polymer. For example, sequencing subsystem 104 can receive raw data from sequencing device 114 and can determine the nucleotide base sequence for a nucleic acid segment. The raw data can be received from sequencing device 114 in a file format recognizable for processing (e.g., a FASTQ file). A FASTQ file may include a text file containing sequence data from clusters passing through a filter on a flow cell. The FASTQ format is a text-based format for storing both biological sequences (e.g., nucleotide sequences, etc.) and their corresponding quality scores. Sequencing subsystem 104 can process the sequencing data to determine the sequence of nucleotide bases in DNA and / or RNA segments or oligonucleotides.
[0016] In addition to processing and determining sequences for biological samples, the sequencing subsystem 104 can generate files for processing and / or transmission to other devices. The files generated can be in sequence alignment / map (SAM) format, binary alignment / map (BAM) format, compressed reference-oriented alignment map (CRAM) format, and / or another file format for processing and / or transmission to other devices. The SAM format may be an alignment format for storing reads aligned to a reference genome. The SAM format can support short and long reads (e.g., up to 128 Mb) generated by different sequencing devices 114. The SAM format may be a text format file that is human-readable. The BAM format can maintain the same information within a SAM file but is a compressed binary format that is machine-readable. A BAM file can indicate the alignment of reads received with data received from a sequencing device 114. A CRAM file can be stored in a compressed column file format for storing biological sequences.
[0017] The client device 108 can generate, store, receive, and / or transmit digital data. Specifically, the client device 108 can receive sequencing metrics from the sequencing device 114. Further, the client device 108 can communicate with the server device 102 to receive one or more files including nucleotide base calls and / or other metrics. The client device 108 can present or display information regarding nucleotide base calls within a graphical user interface to a user associated with the client device 108.
[0018] The client device 108 shown in FIG. 1A can include various types of client devices. In the example, the client device 108 can include a non-mobile device such as a desktop computer or server, or other types of client devices. In other examples, the client device 108 can include a mobile device such as a laptop, tablet, mobile phone, or smartphone.
[0019] As further shown in FIG. 1A, the client device 108 can include a sequencing application 110. The sequencing application 110 can be a web application or a native application (e.g., a mobile application, a desktop application) stored and executed on the client device 108. The sequencing application 110 can include instructions that, when executed, cause the client device 108 to receive data from the sequencing device 114 and present data such as data from a variant call file to the user of the client device 108 for display on the client device 108.
[0020] As further shown in FIG. 1A, the environment 100 can include a database 116. The database 116 can store information such as, for example, a variant call file, a sample nucleotide sequence, a nucleotide read, a nucleotide base call, a sequencing metric, population data, and / or other data described herein. The server device 102, the client device 108, and / or the sequencing device 114 can communicate with the database 116 (e.g., via the network 112) to store and / or access information such as, for example, a variant call file, a sample nucleotide sequence, a nucleotide read, a nucleotide base call, a sequencing metric, population data, and / or other data described herein.
[0021] The environment 100 may be included in a local network or a local high-performance computing (HPC) system. The environment 100 may be included in a cloud computing environment that includes a plurality of server devices, such as server device 102 having distributed software and / or data. The sequencing subsystem 104 may be implemented to operate one or more subsystems as described herein, may be implemented on a single device such as server device 102 or sequencing device 114, or may be distributed across multiple devices such as server device 102 and / or sequencing device 114. The server device 102 and / or the sequencing device 114 can access the database 116 via the network 112, for example, in a cloud-based computing system.
[0022] FIG. 1A illustrates the components of the environment 100 that communicate via the network 112, but it will be understood that the components of the environment 100 may communicate directly with each other, for example, by bypassing the network 112. For example, the client device 108 may communicate directly with the sequencing device 114.
[0023] The array determination subsystem 104 may comprise one or more array determination subsystems used to analyze the array determination data received from the array determination device 114 and / or to perform secondary analysis to identify variants in the array determination data. Nucleotide base calls can indicate the determination or prediction of the type of nucleotide base incorporated within an oligonucleotide on a nucleotide sample slide (e.g., read-based nucleotide base calls), or the determination or prediction of the type of nucleotide base present at a genomic coordinate or genomic region within a sample genome. For example, a nucleotide base call can include a base call corresponding to a genomic coordinate and a reference genome, e.g., an indication of a variant or non-variant at a specific position corresponding to the reference genome. A nucleotide base call can refer to the base detected at a position within a read, along with a quality score indicating the confidence in that call. Base calls can enable the detection of mutations or variants based on a comparison between the base calls in each read spanning a position and the base presented in the reference genome at the same position. Variants can include, but are not limited to, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), or base calls that are part of a structural variant. An insertion changes the DNA sequence by adding one or more nucleotides to the sequence compared to the reference genome. A deletion changes the DNA sequence by removing at least one nucleotide from the sequence compared to the reference genome. The deleted DNA can change the function of the affected protein or proteins. A single nucleotide base call can include an adenine call, a cytosine call, a guanine call, or a thymine call (abbreviated as A, C, G, T) for DNA, or a uracil call (abbreviated as U, instead of a thymine call) for RNA. A mutation can include a single change or difference in a gene sequence. A variant can include a sequence containing one or more mutations.
[0024] FIG. 1B shows an example of one or more subsystems that may be implemented by the sequencing subsystem 104 to identify variants. As shown in FIG. 1B, the sequencing subsystem 104 may implement a mapper subsystem 122, a sorter subsystem 124, and / or a variant caller subsystem 126. The one or more subsystems may be implemented to perform secondary analysis. The mapper subsystem 122 may be implemented to align reads in the sequencing data received from the sequencing device 114 and / or stored in the server device 102. The reads in the sequencing data generated by the sequencing device 114 and / or generated and stored in a file by the server device 102 may not be included in a single sequence having all the DNA information. Instead, the sequencing data generated by the sequencing device 114 and / or generated and stored in a file by the server device 102 may include several short subsequences or reads having partial DNA information. Read alignment may be performed by the mapper subsystem 122 to map the reads to a reference genome and identify the position of each individual read on the reference genome. The mapper subsystem 122 may stream the unaligned reads from the sequencing data as individual base call (BCL) files of FASTQ or ILLUMINA and perform read alignment on the sequencing data therein. The FASTQ file can contain up to millions of entries and can be several megabytes (Mb) or gigabytes (GB) in size. The mapper subsystem 122 can output the aligned reads as an aligned BAM file, as described herein.
[0025] A BAM file can include a header section and an alignment section. The header section can contain information about the file, such as the sample name, sample length, and alignment method. The alignment section can contain read names, read sequences, read qualities, alignment information, and other custom tags for the reads. For each read or read pair, the alignment section can include a read group. A read group can include a subset of reads on a flow cell from the same lane, sample, and / or library preparation. Different read groups can have different coverage or different depths. Depth can be determined by the number of reads aligned to a position in the sequence with a particular quality. Depth can be determined by the number of reads aligned to a position in the sequence with a particular quality. The number of reads can be determined for one or more read groups. The alignment section can include a barcode tag indicating a multiplexed sample identifier associated with the read. The alignment section can include single-end alignment quality. The alignment section can include an edit distance tag recording the Levenshtein distance between the read and the reference.
[0026] Read alignment can be performed using a hash table. The hash table can be constructed for genomic reference, which may enable the lower part of a read or a seed to be mapped to the genome. The position of a read can be determined from the result of seed extension at each of its mapping positions. The mapper subsystem 122 can map many overlapping seeds from each read to perfect matches in the reference using the hash table index of the reference genome. The hash table can be constructed from any selected reference using a multi-threaded tool and loaded into the random access memory (RAM) 125. For example, the RAM 125 may comprise board dynamic RAM (DRAM) of a field programmable gate array (FPGA) on the server device 102. The hash table can be stored in the RAM 125 before the mapping operation performed by the mapper subsystem 122. The read mapping process can be performed by the FPGA logic on the RAM 125.
[0027] After the read alignment is performed by the mapper subsystem 122, the aligned sequencing data can be passed downstream of the sorter subsystem 124 to sort the reads for each reference position, and polymerase chain reaction (PCR) or optical duplication can be optionally flagged. The initial sorting stage can be performed by the sorter subsystem 124 on the aligned reads coming back from the RAM 125. When the mapping is complete, the final sorting and duplication marking can be started. The sorter subsystem 124 can write another BAM file containing the sorted sequencing data to the RAM 125 for downstream access by the variant caller subsystem 126.
[0028] Using the variant caller subsystem 126, variants can be called from the alignment and sorted reads in the sequencing data. For example, the variant caller subsystem can receive a sorted BAM file as input, process the reads, and generate variant data included in a variant call file (VCF) or a genomic variant call format (gVCF) file as the output from the variant caller subsystem 126.
[0029] The variant caller subsystem 126 may comprise a calling subsystem 128 and / or a genotyping subsystem 130. When the variant caller subsystem 126 receives sequencing data, the calling subsystem 128 can identify callable regions having sufficiently aligned coverage. The callable regions can be identified based on the read depth. The read depth can represent the number of times a particular base appears within each read of the sequencing data. Sometimes, incorrect bases may be incorporated into the DNA fragments identified in the sequencing data. For example, the camera within the sequencing device 114 may pick up incorrect signals, the mapper subsystem 122 may misplace the reads, or the sample may be contaminated, resulting in incorrect bases being called in the sequencing data. By sequencing each fragment multiple times to generate multiple reads, there is a confidence or likelihood that the identified variant is a true variant and not an artifact from the sequencing process. The read depth represents the number of times each individual base has been sequenced, or the number of reads in which an individual base appears in the sequencing data. The higher the read depth, the higher the level of reliability in variant calling.
[0030] The callable region may be a region passed downstream to the genotyping subsystem 130 to call variants from the callable region. For example, the genotyping subsystem 130 can compare the callable region with the reference genome for variant calling. The calling subsystem 128 can identify the callable region when the read depth of the sequencing data exceeds the callable region depth threshold. For example, the calling subsystem 128 can identify the callable region in the sequencing data when the read depth of one or more sequence fragments exceeds one depth threshold. After the callable region is identified, the calling subsystem 128 can pass the callable region to the genotyping subsystem 130, and then the callable region can be changed to the active region to generate potential positions where variants may exist within the active region. The genotyping subsystem 130 can identify the probability or call score of whether the potential position contains a variant.
[0031] FIG. 2A includes a graph 200 showing an example of a callable region based on the read depth of a genomic region within sequencing data. As shown in FIG. 2A, the sequencing data may include several short subsequences or reads 202 having partial DNA information. The reads 202 may overlap at a given reference base position, or nucleotide. For example, in a given genomic region 204 within the sequencing data, 8 overlapping reads 202 may exist. Thus, the genomic region 204 within the sequencing data may have a read depth of 8. In another genomic region 206 within the sequencing data, 4 overlapping reads 202 may exist. Thus, the genomic region 206 within the sequencing data may have a read depth of 4. The read depth may be the number of mapped reads at each base position within the sequencing data. The genomic regions 204, 206 may each include, for example, one or more reference base positions.
[0032] As described herein, callable region 212 can be identified by the calling subsystem of the variant caller based on the read depth of the sequencing data. The calling subsystem of the variant caller can identify that the read depth of callable region 212 reaches or exceeds the callable region depth threshold at position 208 in the sequencing data. The callable region depth threshold may be, for example, a read depth of 0 or 1 so that the callable region 212 can be started when the read depth is detected. However, other depth thresholds may be implemented. Callable region 212 may continue to be buffered in memory (e.g., RAM 125 shown in FIG. 1B) by the calling subsystem so as to be sent to the genotyping subsystem of the variant caller to detect variants within callable region 212. The calling subsystem can identify that the read depth reaches or falls below the callable region depth threshold at position 210 in the sequencing data and identify the end of callable region 212. Then, callable region 212 can be sent to or accessed by the genotyping subsystem of the variant caller to detect variants in callable region 212.
[0033] Referring again to FIG. 1B, the calling subsystem 128 of the variant caller subsystem 126 can continue to buffer callable regions in the RAM 125 as long as the read depth of the callable regions exceeds the callable region depth threshold in the alignment data. The calling subsystem 128 can attempt to buffer the entire callable region in the RAM 125 so that the alignment context within the alignment data that can be used to make variant calls in any of the genomic regions of the alignment data in the genotyping subsystem 130 is not lost. The alignment context can include one or more reads and / or bases upstream and / or downstream within the sequence from a certain position on the sequence. The callable region may include a set of sequences that align to the same region of the reference genome or sequence, which may be referred to as a pileup. The pileup may include a group of overlapping reads at the same position with a read depth. The callable region may include a range of reference positions where the sample sequence is piled up to a specific depth that aligns with the reference position. The depth may be indicated by a threshold or value indicating the pileup. For example, the callable region may include a range of reference positions where the sample sequence is piled up to a depth greater than the threshold or value. As the sample sequence within the callable region is piled up to a greater depth, there may be a greater confidence in the variant calling performed.
[0034] Since these callable regions continue to be buffered and / or piled up, they can become larger than the space allocated to the calling subsystem 128 in the RAM 125. In one example, the space allocated in the RAM 125 for the callable region may be 14GB, 20GB, or another level of allocated memory in the RAM 125. In one example, the entire memory available for allocation in the RAM 125 may be 40GB of RAM. The callable region may exceed this allocation and / or other allocations.
[0035] When processing the array determination data, the array determination subsystem 104 may be required to be within the limits of the global memory in the RAM 125 and / or placed on a hard disk drive (HDD) or disk 123. The RAM 125 and the disk 123 may be different types of memory or storage. The RAM 125 can be used to store programs and data that can be used in real time by a processor on the server device 102 that operates the array determination subsystem 104. The RAM 125 may be volatile and may be erased when the switch of the computing device is turned off. The disk 123 may be a persistent storage or non-volatile memory used to store user-specific data, programs, and files that can be accessed when the switch of the computing device is turned on or after the switch is turned off. For example, the array determination subsystem 104 and / or its subsystems may be stored on the disk 123 as computer-executable instructions that can be loaded into the RAM 125 to operate as described herein. The disk 123 may be a network persistent storage including persistent storage shared on one or more computing devices on a network (e.g., a cloud storage system).
[0036] To keep the array determination subsystem 104 within the limits of global memory during operation, each subsystem can have its memory limits allocated or imposed within the RAM 125 and / or on the disk 123. Exceeding these memory limits can cause the operation of individual subsystems to run slower or crash, and the array determination subsystem 104 as a whole to run slower or crash, and / or other applications running on the server device 102 on which the array determination subsystem 104 is operating to run slower or crash. To prevent exceeding these memory limits for a given subsystem, the array determination subsystem 104 and / or the operating system running on one or more server devices 102 monitors the amount of memory utilized by a given subsystem and can cancel or pause the operation of the subsystem to prevent exceeding the memory limits.
[0037] These memory limitations can be particularly challenging for the calling subsystem 128 of the variant caller subsystem 126. As described herein, as the size of the callable regions identified by the calling subsystem 128 continues to increase, the memory limits allocated and / or imposed by the calling subsystem 128 within the RAM 125 may be exceeded by the size of the callable regions before being passed to the genotyping subsystem 130. The size of the callable regions may increase due to the number of reads in various portions of the callable regions. For example, if the genome has been sequenced to an average depth of 300, a portion of the callable region may occupy a large amount of memory allocated to the calling subsystem 128 within the RAM 125. Similarly, as the genomic length of the callable regions increases, the memory allocated to the calling subsystem 128 within the RAM 125 may be similarly occupied and / or exceeded. Thus, the calling subsystem 128 may not be able to fit the entirety of each callable region in the memory buffer allocated within the RAM 125, especially as the depth of these callable regions continues to increase for genome sequencing analysis. When the callable regions reach the allocated memory within the RAM 125 (e.g., 20GB), the calling subsystem 128 may continue to search for available memory, and due to the limited memory allocated for the subsystem, it may stall the calling subsystem 128 and / or the variant caller subsystem 126. After a certain period, the operation of the variant caller subsystem 126 can be cancelled or paused to free up the memory resources within the RAM 125.
[0038] In addition to the RAM 125, each of the subsystems can access storage in the disk 123. The disk 123 may comprise non-volatile memory or persistent storage. Thus, the variant collator subsystem 126 and / or the collating subsystem 128 can spill the callable regions to the disk 123. However, each time data enters and leaves the disk, the cost of processing resources for writing data to the disk 123 and reading data from the disk 123 is greater than when accessing the same data from the RAM 125. Further, the processing cost may be even greater when data compression / decompression is performed for storage on the disk 123. These processing costs may degrade the performance of other subsystems within the array determination subsystem 104 and / or other applications operating on the server device 102.
[0039] In an attempt to reduce the processing costs associated with reading from and writing to disk 123, each subsystem within the array determination subsystem 104 can attempt to operate within the memory allocated to the subsystem. In order to operate within the memory allocated to the calling subsystem 128, the calling subsystem 128 can monitor the memory used within the RAM 125 to buffer the callable region. When the buffered memory reaches the memory threshold of the total amount of RAM 125 allocated to the calling subsystem 128, the calling subsystem 128 can divide the callable region into two or more genomic regions. For example, the memory threshold can be set to a certain percentage (e.g., 70% or 75%) of the total memory allocated within the RAM 125 for the calling subsystem 128. In another example, the memory threshold can be set to a predefined amount of memory (e.g., 1GB), and the calling subsystem 128 can be made to perform the division at or near each threshold amount of memory. The memory threshold can be fixed or dynamic in order to cause the division of the callable region at different positions, as further described herein. The division can force the end of the callable region within the RAM 125 even if the callable region does not drop to the callable region depth threshold. The calling subsystem 128 can scan the reads before and / or after the position within the array data at which the memory threshold was reached and can move the division earlier or later in an attempt to prevent the disappearance of variants (e.g., due to the pre- and post-alignment context). After dividing a region, the calling subsystem 128 can pass the divided region downstream to the genotyping subsystem 130. In one example, the divided region can be passed by sending pointers to the reads of the divided portions to the genotyping subsystem 130. The memory (e.g., RAM) can remain occupied with the reads of the divided portions for processing by the genotyping subsystem 130.The genotyping subsystem 130 can change a portion of the callable region to an active region to generate potential positions where variants may exist within the active region. The genotyping subsystem 130 can identify the probability or call score as to whether a potential position contains a variant in the active region. After the genotyping subsystem 130 processes the reads in the sequencing data, the memory (e.g., RAM) may be freed to receive additional sequencing data.
[0040] When the buffered memory reaches the memory threshold of the total amount of RAM 125 allocated to the calling subsystem 128, the calling subsystem 128 can send the entire buffered portion of the callable region to the genotyping subsystem 130 to free the entire buffer. In another example, the calling subsystem 128 can identify another position within the buffered sequencing data to create a split. For example, the calling subsystem 128 can analyze the read depth of the buffered sequencing data to identify the position to create a split. The calling subsystem 128 can identify a portion of the sequencing data having a read depth less than a split threshold for creating a split. The split threshold may be set to a read depth greater than the callable region depth threshold. However, the split threshold may be set to a read depth that prevents the loss of additional sequencing context around the position of the split that can be used by the genotyping subsystem 130 to more accurately predict variants during variant calling.
[0041] FIG. 2B includes a graph 200a showing an example of a position where callable regions can be divided based on the read depth of genomic regions within the sequencing data. The sequencing data shown in graph 200a of FIG. 2B may be the same as the sequencing data in graph 200 shown in FIG. 2A. However, the calling subsystem can monitor the buffered memory in the RAM and identify that the memory threshold has been reached at position 216 within the sequencing data. The calling subsystem can then analyze the buffered sequencing data to identify position 214 where the read depth is below the division threshold. Callable region 212 may be divided at position 214 (e.g., when the division threshold is set to a read depth of 3), and the divided portion 218 may be transmitted downstream to the genotyping subsystem for variant calling. Since the division threshold can be met at various positions within the buffered sequencing data, the calling subsystem can identify position 214 that clears the maximum amount of data from the buffer. The remaining portion 220 of callable region 212 may be left in the buffer, and the calling subsystem may continue to buffer the sequencing data until the callable region depth threshold or another memory threshold is reached.
[0042] The calling subsystem may use additional logic that allows the calling subsystem to make smarter divisions in callable regions within the sequencing data. For example, the calling subsystem can identify that the memory threshold has been reached or that the sequencing data is within a predefined amount of occupied memory since reaching the memory threshold, and scan the buffered sequencing data to identify positions where divisions can be made to prevent the disappearance of variants in the genotyping subsystem. The calling subsystem may implement dynamic memory thresholds and / or dynamic division thresholds. The dynamic thresholds may enable smarter divisions based on the amount of memory occupied in the buffer and the read depth of the sequencing data.
[0043] FIG. 2C includes a graph 200b showing another example of positions for dividing callable regions based on the read depth of genomic regions within the array determination data. The array determination data shown in graph 200b of FIG. 2C may be the same as the array determination data in graphs 200 and 200a shown in FIGS. 2A and 2B, respectively. As shown in FIG. 2C, the calling subsystem can analyze the buffered array determination data to identify one or more positions 214a, 214b for dividing the callable region 212. For example, the calling subsystem can divide the callable region 212 into one or more genomic regions based on a predefined lower division threshold. The calling subsystem can divide the callable region 212 each time the lower division threshold is met or exceeded. The lower division threshold may be greater than the callable region depth threshold. As shown in FIG. 2C, the lower division threshold may be set to depth threshold 3, and the calling subsystem can be made to divide the callable region 212 each time the division threshold is met or exceeded in the array determination data. The predefined lower division threshold can prevent the loss of the array determination context by dividing the callable region 212 at positions in the array determination data where the read depth is relatively lower than other regions in the array determination data. Therefore, there is less loss of data when performing variant calling in the genotype determination subsystem due to a smaller depth compared to when performing the division at positions in the array determination data having a larger depth, and the confidence level of the variants in the division is smaller. To avoid the loss of the array determination context and / or to perform multiple divisions within genomic regions that do not occupy at least a minimum level of buffer available for storage, the calling subsystem can set a lower memory threshold for creating the divisions. For example, the division at position 214a can be performed after the lower memory threshold is met or exceeded in the buffer. After the division at 214a, the divided portion 218a can be sent to the genotype determination subsystem for identifying variants.
[0044] The lower split threshold and / or the lower memory threshold may be configured by the user and / or may be updated dynamically during a particular callable region 212. For example, the lower split threshold and / or the lower memory threshold can be adjusted based on user input (e.g., received from an array determination application 110 running on client device 108 as shown in FIG. 1A). The lower split threshold can also, or alternatively, be updated dynamically based on the amount of available memory remaining in the buffer. For example, the lower split threshold may increase up to a predefined read depth as the amount of space occupied in the buffer by callable region 212 increases. As shown in FIG. 2C, the calling subsystem can determine to split callable region 212 at location 214b of the array determination data having a read depth greater than location 214a. After the split at 214b, split portion 218b can be sent to the genotyping subsystem to identify variants.
[0045] The calling subsystem may include at least a portion of the logic within the genotyping subsystem that is used to identify variants in the genotyping subsystem to prevent the loss of variants or loss of context and improve the ability to predict variants in the genotyping subsystem. For example, the calling subsystem can identify insertions, deletions, or other variants or mutations in a portion of the array determination data and prevent the splitting of the callable region within a predefined region of indels.
[0046] FIG. 2D includes a graph 200c showing another example of positions for dividing callable regions into genomic regions in the array determination data. The array determination data shown in graph 200c of FIG. 2D may be the same as the array determination data in graphs 200, 200a, and / or 200b shown in FIGS. 2A, 2B, and 2C, respectively. As shown in FIG. 2D, the calling subsystem can detect that a memory threshold and / or a division threshold has been reached for callable region 212 and identify a position 214d for performing a division in callable region 212. For example, the calling subsystem can identify position 214d based on the read depth of the array determination data at position 214d. The calling subsystem can then analyze the buffered array determination data to determine whether there are insertions, deletions, or other variants or mutations within a pre-defined proximity (e.g., genomic region 226) of the identified division position 214d. Genomic region 226 may be a pre-defined number of bases (e.g., 100 bases, 1,000 bases, etc.), a pre-defined number of buffer positions, and / or a pre-defined amount of memory. For example, the calling subsystem can analyze each read 202 and compare read 202 to a reference sequence to detect variants or mutations in one or more genomic portions 222 of read 202. Genomic portion 222 may include one or more bases themselves. The bases within genomic portion 222 can be specified by a start base or start index and / or an end base or end index.
[0047] The calling subsystem can utilize population data to identify variants or mutations in genomic portions 222 that are more likely (or more common) for a population corresponding to a sample genome. Thus, the calling subsystem can utilize various frequencies and / or population data indicating the likelihood of insertions, deletions, structural variants, copy number polymorphisms, or other variants or mutations occurring within genomic region 222, and / or the size of genomic region 222. The calling subsystem can access a database containing population data. The database can indicate genomic sequences and corresponding genomic coordinates for a given population or ethnic group. The database can also include metadata indicating surrounding nucleotide bases common to the population or ethnic group. The calling subsystem can determine and utilize population data corresponding to a sample genome. By way of example, in some embodiments, the calling subsystem can identify or receive data regarding a population and / or ethnic group corresponding to a particular sample genome. Thus, the calling subsystem can identify variants or mutations from bases or genomic portions common to the population or ethnic group. By way of example, in one or more embodiments, the calling subsystem can utilize a reference genome corresponding to a specified population or ethnic group corresponding to a sample genome. The calling subsystem can identify insertions, deletions, structural variants, copy number polymorphisms, or other variants or mutations at genomic coordinates within a genomic region (e.g., genomic region 222), or genomic regions containing other variants or mutations. The calling subsystem can also identify the size of genomic region 222 based on a reference genome corresponding to the specified population or ethnic group.
[0048] The calling subsystem can utilize data regarding populations and / or ethnic groups to identify insertions, deletions, structural variants, copy number polymorphisms, or other variants or mutations in the genomic coordinates of a sample genome or a target genomic region. By way of example, the calling subsystem can identify variants or mutations within genomic region 222 based on the surrounding nucleotide bases within the genomic region. Based on the nucleotide bases within the genomic region, the calling subsystem can identify the likelihood of an insertion, deletion, structural variant, copy number polymorphism, or other variant or mutation for a genomic coordinate or region.
[0049] If the calling subsystem detects that a variant or mutation in genomic portion 222 is within a pre-defined genomic region 226 of the identified split position 214d, the calling subsystem can prevent a split from occurring at position 214d. The calling subsystem can identify another position 214c for performing a split within the buffered portion of callable region 212. Position 214c can be at least a pre-defined genomic region 224 from genomic portion 222 (e.g., from the start index or start base of genomic portion 222). Genomic region 224 can be a pre-defined number of bases (e.g., 100 bases, 1,000 bases, etc.), a pre-defined number of buffer positions, and / or a pre-defined amount of memory. Genomic region 224 can be the same as or different from genomic region 226. A split of callable region 212 at position 214c, which is at least a pre-defined genomic region 224, can prevent the loss of the sequencing context before and after a variant or mutation for variant calling in the genotyping subsystem.
[0050] Genomic regions 224, 226 may vary dynamically or in other ways based on the type of sequencing data being analyzed. For example, genomic regions 224, 226 can be determined based on population data. The calling subsystem can identify some common bases used as a context before and after sequencing for insertions, deletions, structural variants, copy number polymorphisms, or other variants or mutations based on population data. The calling subsystem can update genomic regions 224, 226 based on the number of bases shown in the population data. Variants or mutations may not spread equally across the genome. Population data can be used to save positions (e.g., positions with few mutations), and in an attempt to avoid splitting within genomic regions 224, 226 of variants or mutations, splitting can be done at these positions.
[0051] The calling subsystem may attempt to identify position 214c with a predefined read depth of the splitting threshold from genomic portion 222 to identify another split. If the calling subsystem cannot find a position in the sequencing data with a predefined read depth of the splitting threshold, the splitting threshold may be dynamically updated to allow a greater or smaller read depth, or the calling subsystem may not be able to perform another split between position 208 at the start of the callable region 212 and genomic portion 222 containing the variant or mutation. Then, the calling subsystem may attempt to split callable region 212 after genomic portion 222 containing the variant or mutation. The calling subsystem may attempt to split callable region 212 from genomic portion 222 at a position outside genomic region 224. The memory threshold may have to be increased to accommodate splitting while maintaining the context around the variant or mutation.
[0052] As described herein, the partitioning of callable region 212 can cause the loss of sequencing context that can be utilized by the genotyping subsystem to create accurate variant calls. When partitioning callable region 212, in order to preserve the sequencing context and prevent other loss of sequencing data, the calling subsystem can maintain duplication of data from callable region 212 in a buffer and send multiple partitioned portions of the callable region to the genotyping subsystem that includes the duplicate data. The genotyping subsystem can identify the duplicate data and discard any multiple data. However, the duplicate data can help ensure that the sequencing context and / or other data around the partition are preserved for variant calling.
[0053] FIG. 2E includes a graph 200d showing another example of a position for partitioning a callable region into genomic regions in sequencing data. The sequencing data shown in graph 200d of FIG. 2E may be similar to the sequencing data in graphs 200, 200a, 200b, and / or 200c shown in FIGS. 2A, 2B, 2C, and 2D, respectively. As shown in FIG. 2E, the calling subsystem can detect that a memory threshold and / or a partitioning threshold has been reached for callable region 212 and identify a position 214f for performing a partition in callable region 212. For example, the calling subsystem can identify position 214f based on the read depth of the sequencing data at position 214f. The calling subsystem can send partitioned portion 218c to the genotyping subsystem to identify variants. Partitioned portion 218c may include the sequencing data from the position 214f of the partition and the start of callable region 212 or the position of a previous partition.
[0054] After transmitting the split portion 218c to the genotyping subsystem, the calling subsystem can maintain the overlapping portion 230 of the sequencing data in memory. The overlapping portion 230 may be a predefined portion of the sequencing data from position 214f to position 214e. The predefined portion may be a predefined number of bases (e.g., 100 bases, 1,000 bases, etc.), a predefined number of buffer positions, and / or a predefined amount of memory. The predefined number of bases, the number of buffer positions, and / or the amount of memory may be predefined and / or defined by the user based on user input. The overlapping portion 230 may change dynamically. For example, the overlapping portion 230 may change based on population data. The overlapping portion 230 can be defined by read depth. For example, the calling subsystem can identify position 214e to define the overlapping portion 230 of the sequencing data from split position 214f based on a predefined read depth at position 214e. The genotyping subsystem can maintain or update a pointer to read 202 within the overlapping portion 230 so that the memory of the calling subsystem 128 continues to be occupied by read 202 within the overlapping portion 230. When the calling subsystem identifies a subsequent position for splitting the callable region 212 or reaches position 210 at the end of the callable region 212, the calling subsystem can transmit the remaining portion 220 of the callable region 212 to the genotyping subsystem for variant calling. The calling subsystem can analyze the split portion 218c and determine not to maintain the overlap based on the data in the split portion 218c. The genotyping subsystem can identify the overlapping portion and remove duplicate bases before performing variant calling. The overlapping data maintained near the split can prevent the loss of data at or near the split.
[0055] FIG. 3 is a flowchart of a procedure 300 for dividing callable regions in sequencing data into genomic regions based on a memory threshold. Procedure 300 can be performed by one or more computing devices. For example, procedure 300 may be performed by one or more server devices (e.g., server device 102 shown in FIG. 1A) and / or a sequencing device (e.g., sequencing device 114) that execute one or more parts of a sequencing subsystem (e.g., sequencing subsystem 104 shown in FIGS. 1A and 1B). Procedure 300 may be performed by one or more subsystems of the sequencing subsystem. For example, procedure 300 may be performed by a calling subsystem (e.g., calling subsystem 128 shown in FIG. 1B). The calling subsystem may be provided as an exemplary subsystem for performing procedure 300, but one or more parts of procedure 300 may be performed by one or more systems and / or subsystems, such as the sequencing subsystem and / or one or more subsystems therein. Procedure 300 may be performed on a single computing device or may be distributed across multiple computing devices. For example, procedure 300 may be executed as computer-executable instructions retrieved from memory by one or more processors on one or more computing devices.
[0056] As shown at 302 in FIG. 3, the calling subsystem can receive sequencing data that includes a plurality of reads of a genomic sequence. At 304, the calling subsystem can identify the start of a callable region within the sequencing data. For example, the callable region can be identified at the start of the sequence data and / or when the sequencing data reaches a callable region depth threshold. The callable region depth threshold may be a value of 0 or 1. The callable region depth threshold may be predetermined or dynamic as described herein. The callable region depth threshold for the start of the callable region may be the same as or different from the callable region depth threshold for the end of the callable region.
[0057] In 306, the colling subsystem can monitor the memory used to store the alignment determination data in the callable area. The callable area can be buffered in the RAM by the colling subsystem in an attempt to identify the callable area within the RAM and send the callable area to the genotyping subsystem for variant calling within the callable area. The memory threshold may be predetermined or dynamic as described herein. For example, the memory threshold may be set to a relatively large portion (e.g., bytes, percentage, etc.) (e.g., 70%, 75%, etc.) of the total memory allocated to store the callable area of the alignment determination data, and send a relatively large portion of that callable area to the genotyping subsystem, or a relatively small portion (e.g., bytes, percentage, etc.) (e.g., 25%, 30%, etc.) of the total memory allocated to store the callable area of the alignment determination data may be sent to send a relatively small portion of the callable area to the genotyping subsystem.
[0058] In 308, the calling subsystem can determine whether the memory used to store the callable region has reached the memory threshold of the memory allocated to store the callable region. If the memory threshold has not been reached, the calling subsystem can, in 314, monitor the alignment determination data to determine whether the end of the callable region has been reached. For example, the calling subsystem can identify the end of the callable region in the alignment determination data when the alignment determination data reaches the callable region depth threshold. For example, the callable region depth threshold may be a value of 0 or 1. Once the end of the callable region is identified, the calling subsystem can, in 316, transmit the callable region downstream to the genotyping subsystem for variant calling. If the end of the callable region has not been reached, the calling subsystem can, in 315, continuously receive the alignment determination data, continuously monitor the memory used to store the callable region 306, and / or monitor the end of the callable region.
[0059] When the collating subsystem determines at 308 that a memory threshold has been reached, the collating subsystem can determine at 310 the position within the array determination data for dividing the callable region. The callable region can be divided to store the memory resources allocated to the collating subsystem, the array determination subsystem, and / or the computing device on which the collating subsystem can operate. The collating subsystem can determine at 310 to divide the callable region at the position within the array determination data where the memory threshold has been reached. As a result, the collating subsystem transmits the entire buffered array determination data including a part of the callable region to the genotyping subsystem. As described herein, the callable region can be divided by analyzing the buffered portion of the callable region to identify another position for division. For example, the division can reduce the amount of array determination context that can be lost, avoid loss of variants, avoid division within a pre-defined region of a variant such as an insertion, deletion, structural variant, copy number polymorphism, or another variant or mutation, and / or be performed at a position that stores the data in another manner when the division is made.
[0060] The position of the division can be determined at 310 based on a division threshold. The division threshold may be set to a read depth greater than the callable region depth threshold. As described herein, the collating subsystem can analyze the region where the division can be made to determine whether to divide the callable region at another position. For example, the collating subsystem can identify an insertion, deletion, or other variant or mutation within a pre-defined proximity of the position of the division and identify another position for division. The position of the division can be made to store the array determination context data that can be used by the genotyping subsystem to identify variants. The position of the division may be made based on population data as described herein.
[0061] The divided portion of the callable region can be sent downstream to the genotyping subsystem for variant calling at 312. The genotyping subsystem can start analyzing the divided portion of the callable region to identify variants. The calling subsystem can maintain a portion of the divided portion of the callable region in a buffer, such that the overlapping portion of two consecutive portions of the callable region has an overlap when received by the genotyping subsystem. The genotyping subsystem can identify the overlapping portion and remove duplicate bases before performing variant calling. The overlapping data maintained near the division can prevent loss of data at or near the division. If the position where the division is made is not at the end of the callable region, the calling subsystem can continuously receive sequencing data at 315, monitor the memory used to store the callable region 306, and / or monitor the end of the callable region.
[0062] Embodiments are described herein for dividing a callable region into one or more parts and operating within a memory allocation in a RAM for storing the callable region. Alternatively, or in addition thereto, the callable region, or a portion thereof, can be spilled to disk and operated within a memory allocation in the RAM. The disk may comprise non-volatile memory or persistent storage. To avoid incorrect variant calls and preserve the sequencing context for making variant calls in a genotyping subsystem, a portion of the callable region can be spilled to disk by a calling subsystem and streamed back from the disk by the genotyping subsystem of the variant caller for further analysis. For example, if the callable region reaches a spill threshold (which may be a memory threshold or a threshold number of bases), the callable region can be spilled to disk. To avoid the risk of losing variant calls at the split positions, the callable region can be spilled to disk. Reads within the callable region can be written to disk to save the callable region.
[0063] Referring again to FIG. 1B, the calling subsystem 128 of the variant caller subsystem 126 can receive the aligned and sorted sequencing data and identify a callable region having sufficient aligned coverage. The callable region can be identified based on read depth as described herein. The callable region may be a region passed downstream to the genotyping subsystem 130 for calling variants from the callable region. For example, the genotyping subsystem 130 can compare the callable region to a reference genome to call variants.
[0064] The collating subsystem 128 can continue to buffer callable regions in the RAM 125 as long as the read depth of the callable regions exceeds the callable region depth threshold in the alignment determination data. To operate within the memory allocated to the collating subsystem 128, the collating subsystem 128 can monitor the memory used in the RAM 125 for buffering callable regions. When the buffered memory reaches the memory threshold of the total amount of the RAM 125 allocated to the collating subsystem 128, the collating subsystem 128 can spill the callable regions to the disk 123 (e.g., HDD). The memory threshold may be the same as or different from the memory threshold described herein for dividing callable regions.
[0065] The calling subsystem 128 can spill the entire callable region to disk 123 and store it as a file that can be accessed by the downstream genotyping subsystem 130 when performing variant calling. The calling subsystem 128 can call functions within the genotyping region, provide the calling subsystem 128 with the start pointer and end pointer of the callable region, and instruct the calling subsystem 128 to analyze the reads within the callable region. When a portion of the callable region is spilled to disk, the calling subsystem 128 can pass information on how to access the reads stored on disk (e.g., using a file name or other identifier). The genotyping subsystem 130 can load the entire callable region into RAM 125, or stream back a portion of the spilled callable region from disk 123 to RAM 125. In another example, the calling subsystem 128 can maintain a portion of the callable region within buffered RAM 125 before reaching a memory threshold and spill the remainder of the callable region to disk. The calling subsystem 128 can update the pointer within RAM 125 and send a portion of the callable region within RAM 125 to the genotyping subsystem 130. The genotyping subsystem 130 can start processing the portion of the callable region first received in RAM 125, and then stream a portion of the callable region from disk 123 to RAM 125 when the RAM 125 allocated to the genotyping subsystem is freed. Since the genotyping subsystem 130 can allocate a larger portion of RAM 125 than the calling subsystem, the genotyping subsystem 130 may be able to load a larger portion of the callable region into RAM 125 when processing the callable region.
[0066] The genotyping subsystem 130 can perform an initial analysis to determine whether to maintain sequencing data within the streamed portion in RAM 125 before streaming a portion of the spillable callable region from disk 123 to RAM 125 and accessing an additional portion of the spillable callable region from disk 123. For example, the genotyping subsystem 130 can stream a first portion (e.g., 1GB, 2GB, etc.) of a read within the callable region from disk 123 and perform an initial analysis to determine whether this first portion of the read contains an insertion, deletion, or another variant or mutation. The genotyping subsystem 130 can discard a portion of the read within the callable region in RAM 125 that is not used for variant calling (e.g., the context before and after the variant or the sequencing of the variant calling) and stream a second portion of the callable region from disk 123. The genotyping subsystem 130 can continuously stream a portion of the callable region from disk 123 to load a portion of the callable region in RAM 125 that is used for variant calling. The genotyping subsystem 130 can continuously monitor the amount of the callable region loaded into RAM 125 and use a memory threshold of the amount of memory allocated to the genotyping subsystem 130 to limit the amount of the callable region streamed from disk 123 before processing the callable region and free up a portion of RAM 125 before uploading an additional portion of the callable region. The memory threshold can be a separate memory threshold that is predefined data less than the memory allocated for the genotyping subsystem 130.
[0067] Spilling the callable region to disk 123 can enable the colling subsystem 128 to operate within the memory allocation for the callable region without losing the sequencing context that can be used to perform variant calling. Maintaining a portion of the callable region in RAM 125 reduces the data to be sent to disk 123 and can also save processing / compression that can be used for writing to / reading from disk 123.
[0068] Figure 4 is a flowchart of a procedure 400 for spilling a callable region to disk based on a memory threshold. Procedure 400 can be performed by one or more computing devices. For example, procedure 400 can be performed by one or more server devices (e.g., server device 102 shown in FIG. 1A) and / or a sequencing device (e.g., sequencing device 114) that execute one or more portions of a sequencing subsystem (e.g., the sequencing subsystem 104 shown in FIGS. 1A and 1B). Procedure 400 can be performed by one or more subsystems of the sequencing subsystem. For example, procedure 400 can be performed by a colling subsystem and / or a genotyping subsystem (e.g., colling subsystem 128 and / or genotyping subsystem 130 shown in FIG. 1B). The colling subsystem and the genotyping subsystem can be provided as exemplary subsystems for performing a portion of procedure 400, but one or more portions of procedure 400 can be performed by one or more systems and / or subsystems, such as the sequencing subsystem and / or one or more subsystems thereof. Additionally, one of the colling subsystem or the genotyping subsystem can be described herein as performing one or more features, but these features can be performed by other subsystems. Procedure 400 can be performed on a single computing device or can be distributed across multiple computing devices. For example, procedure 400 can be executed as computer-executable instructions retrieved from memory by one or more processors on one or more computing devices.
[0069] As shown at 402 in FIG. 4, the calling subsystem can receive sequencing data including a plurality of reads of a genomic sequence. At 404, the calling subsystem can identify the start of a callable region within the sequencing data. For example, the callable region can be identified at the start of the sequence data and / or when the sequencing data reaches a callable region depth threshold.
[0070] At 406, the calling subsystem can monitor the memory used to store the sequencing data in the callable region. The callable region can identify the callable region in the RAM and can be buffered in the RAM by the calling subsystem in an attempt to send the callable region to a genotyping subsystem for variant calling within the callable region. The memory threshold may be set to a portion (e.g., bytes, percentage, etc.) (e.g., 70%, 75%, etc.) of the total memory allocated to store the callable region of the sequencing data.
[0071] At 408, the calling subsystem can determine whether the memory used to store the callable region has reached the memory threshold of the memory allocated to store the callable region. If the memory threshold has not been reached, the calling subsystem can, at 414, monitor the alignment determination data to determine whether the end of the callable region has been reached. The calling subsystem can identify the end of the callable region in the alignment determination data when the alignment determination data reaches the callable region depth threshold. For example, the callable region depth threshold may be a value of 0 or 1. Once the end of the callable region is identified, the calling subsystem can, at 416, transmit the callable region downstream to the genotyping subsystem for performing variant calling at 418. If the end of the callable region has not been reached, the calling subsystem can, at 415, continuously receive the alignment determination data, continuously monitor the memory used to store the callable region at 406, and / or monitor the end of the callable region.
[0072] If the calling subsystem determines at 408 that the memory threshold has been reached, the calling subsystem can, at 410, determine to spill the callable region to disk. The entire callable region can be spilled to disk, or a portion of the callable region may be maintained in RAM. For example, a portion of the callable region that was buffered before reaching the memory threshold at 408 can be maintained in RAM and transmitted downstream to the genotyping subsystem for processing. The calling subsystem can transmit a portion of the callable region maintained in RAM downstream to the genotyping subsystem by updating a pointer in the buffer to maintain a portion of the callable region in RAM. The spilled portion of the callable region may be compressed and / or stored on disk.
[0073] At 412, the genotyping subsystem can stream callable regions from disk. For example, at 412, the genotyping subsystem can stream the entire callable region or a portion thereof from disk to RAM. Streaming the callable region may include decompressing the callable region. The genotyping subsystem can stream a portion of the callable region from disk to RAM, analyze a portion of the callable region to determine whether to discard one or more portions of the callable region, and free up additional RAM that cannot be used for variant calling. For example, the genotyping subsystem can analyze each read within the region. If a read is equal to the reference, the read can be determined to contain no variants and can be discarded or ignored. The genotyping subsystem can also identify and discard or ignore sequencing errors. The genotyping subsystem can monitor another memory threshold while streaming the callable region from disk to ensure that the memory allocation within RAM for the genotyping subsystem does not exceed. The genotyping subsystem may perform variant calling and / or genotyping using the callable region or a portion thereof uploaded to RAM. One or more portions of the callable region may be used for performing variant calling at 418.
[0074] FIG. 5 is a block diagram illustrating an exemplary computing device 500. One or more computing devices, such as computing device 500, may implement one or more features for generating and / or processing callable regions, as described herein. For example, computing device 500 may include one or more of the sequencing device 114, client device 108, and / or server device 102 shown in FIG. 1A. As shown by FIG. 5, computing device 500 may comprise a processor 502, a memory 504, a storage device 506, an I / O interface 508, and / or a communication interface 510, which may be communicatively coupled by a communication infrastructure 512. It should be understood that computing device 500 may include fewer or more components than those shown in FIG. 5.
[0075] Processor 502 may include hardware for executing instructions, such as instructions that make up a computer program. In an embodiment, to execute instructions for dynamically changing a workflow, processor 502 may retrieve (or fetch) instructions from an internal register, internal cache, memory 504, or storage device 506 and decode and execute these instructions. Memory 504 may be volatile or non-volatile memory used to store data, metadata, computer-readable or machine-readable instructions, and / or programs for execution by a processor operating as described herein. Memory 504 may include RAM. Storage device 506 may include storage such as a hard disk, flash disk drive, or other digital storage device for storing data or instructions for performing the methods described herein.
[0076] The I / O interface 508 can enable a user to provide input to, receive output from, and / or otherwise transfer data to and receive data from the computing device 500. The I / O interface 508 can include a mouse, keypad or keyboard, touch screen, camera, optical scanner, network interface, modem, other known I / O devices, or combinations of such I / O interfaces. The I / O interface 508 can include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. The I / O interface 508 can be configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content.
[0077] The communication interface 510 can include hardware, software, or both. In any case, the communication interface 510 can provide one or more interfaces for communication (e.g., packet-based communication, etc.) between the computing device 500 and one or more other computing devices or networks. The communication can be wired communication or wireless communication. By way of example and not limitation, the communication interface 510 can include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as WI-FI.
[0078] Additionally, communication interface 510 can facilitate communication with various types of wired or wireless networks. Communication interface 510 can also facilitate communication using various communication protocols. Communication infrastructure 512 can also include hardware, software, or both that couple the components of computing device 500 to each other. For example, communication interface 510 can enable multiple computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein using one or more networks and / or protocols. By way of example, a sequencing process can enable multiple devices (e.g., client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.
[0079] In addition to what has been described herein, the methods and systems may also be implemented, for example, by a computer program, software, or firmware incorporated into one or more computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted by a wired or wireless connection) and tangible / non-transitory computer-readable storage media. Examples of tangible / non-transitory computer-readable storage media include, but are not limited to, read only memory (ROM), random-access memory (RAM), removable disks, and optical media such as CD-ROM disks and digital versatile disks (DVDs).
[0080] The present disclosure has been described with respect to specific embodiments and generally related methods, but modifications and substitutions of the embodiments and methods will be apparent to those skilled in the art. Accordingly, the above description of exemplary embodiments does not limit the present disclosure. Other changes, substitutions, and modifications are also possible without departing from the spirit and scope of the present disclosure.
[0081] Clause 1. A computer-implemented method, comprising: receiving sequencing data including a plurality of reads of a genomic sequence in a calling subsystem of a variant caller, wherein the calling subsystem of the variant caller is configured to detect a callable region of the sequencing data when the depth of the plurality of reads exceeds a callable region depth threshold, and the calling subsystem of the variant caller is configured to transmit at least a portion of the callable region to a genotyping subsystem of the variant caller for variant calling of the callable region; monitoring memory used by the calling subsystem of the variant caller; when the memory used by the calling subsystem of the variant caller exceeds a memory threshold of a total amount of memory allocated to the calling subsystem of the variant caller, dividing the callable region; and transmitting the divided portions of the callable region to the genotyping subsystem of the variant caller for variant calling based on the divided portions. 2. The computer-implemented method according to clause 1, wherein transmitting the divided portions of the callable region to the genotyping subsystem increases the availability of memory used by the calling subsystem. 3. The computer-implemented method according to clause 1, wherein the memory threshold is a fixed threshold. 4. The computer-implemented method according to clause 1, wherein the memory threshold is a dynamic threshold. 5. The method further comprises: In the variant caller's calling subsystem, when the memory used by the variant caller's calling subsystem is within the memory threshold of the total amount of memory allocated to the variant caller's calling subsystem, analyzing the alignment determination data and identifying a variant or mutation in the alignment determination data within a pre-defined proximity of a specified partition, after identifying a variant or mutation in the alignment determination data, partitioning callable regions outside the pre-defined proximity of the variant or mutation, the computer-implemented method according to claim 1. 6. The computer-implemented method according to claim 5, wherein at least one of the variant or mutation, or the pre-defined proximity of the specified partition, is determined based on population data accessed by the calling subsystem. 7. The computer-implemented method according to claim 5, wherein the variant or mutation includes an insertion or deletion. 8. The computer-implemented method according to claim 1, further comprising analyzing buffered alignment determination data to identify positions for partitioning callable regions within the buffered alignment determination data. 9. Identifying a portion of the buffered alignment determination data having a read depth less than a partitioning threshold, and partitioning the buffered alignment determination data in the identified portion having a read depth less than the partitioning threshold, the computer-implemented method according to claim 8. 10. The computer-implemented method according to claim 1, wherein when the depth of a plurality of reads is less than a callable region depth threshold used to detect a callable region, the callable region is transmitted to a genotyping subsystem. 11. The computer-implemented method according to claim 1, wherein the partitioned portion of the callable region is the entire callable region currently present in the memory used by the calling subsystem. 12. The partitioned portion is the first partitioned portion, and the method further Maintaining a predefined amount of sequencing data within the memory used by the variant caller's calling subsystem, within a first partition of the callable region, and Detecting a second partition of the callable region based on a callable region depth threshold or a memory threshold, and Transmitting the second partition of the callable region to the genotyping subsystem of the variant caller for variant calling based on the second partition, where the second partition includes an overlap with the first partition of the callable region that contains a predefined amount of sequencing data within the sequencing data. The computer-implemented method according to claim 1 includes 13. The computer-implemented method according to claim 12, wherein the predefined amount of overlapping sequencing data is a predefined number of bases. 14. The computer-implemented method according to claim 12, wherein the predefined amount of overlapping sequencing data is determined based on population data accessed by the calling subsystem. 15. The computer-implemented method according to claim 12, wherein the predefined amount of overlapping sequencing data is determined based on user input. 16. In the genotyping subsystem of the variant caller, identifying the overlap in the sequencing data between the first and second partitions of the callable region, and Removing the overlap in the sequencing data between the first and second partitions of the callable region before performing variant calling on the callable region. The computer-implemented method according to claim 12 further includes 17. The callable region depth threshold is a first callable region depth threshold, the partition is a first partition, and the method further includes Monitoring the depth of a plurality of reads in the calling subsystem of the variant caller, and Dividing the callable region if the depth of the plurality of reads is less than a second callable region depth threshold. Sending a second divided portion of the callable region to a genotyping subsystem of a variant caller for variant calling based on the second divided portion, the computer-implemented method according to claim 1. 18. A sequencing system, a memory, at least one processor, in a calling subsystem of a variant caller, receiving sequencing data including a plurality of reads of a genomic sequence, the calling subsystem of the variant caller being configured to detect a callable region of the sequencing data when the depth of the plurality of reads exceeds a callable region depth threshold, the calling subsystem of the variant caller being configured to transmit at least a portion of the callable region to a genotyping subsystem of the variant caller for variant calling of the callable region, monitoring a memory used by the calling subsystem of the variant caller, when the memory used by the calling subsystem of the variant caller exceeds a memory threshold of a total amount of memory allocated to the calling subsystem of the variant caller, dividing the callable region, a sequencing system comprising at least one processor configured to transmit a divided portion of the callable region to a genotyping subsystem of the variant caller for variant calling based on the divided portion. 19. The sequencing system according to claim 18, wherein the at least one processor is configured to transmit a divided portion of the callable region to the genotyping subsystem to increase the availability of the memory used by the calling subsystem. 20. The sequencing system according to claim 18, wherein the memory threshold is a fixed threshold. 21. The sequencing system according to claim 18, wherein the memory threshold is a dynamic threshold. 22. The at least one processor is In the colling subsystem of the variant caller, when the memory used by the colling subsystem of the variant caller is within the memory threshold of the total amount of memory allocated to the colling subsystem of the variant caller, analyze the alignment determination data, Identify a variant or mutation in the alignment determination data within a pre-defined proximity of a specified partition, The alignment determination system according to clause 18, configured to perform a partition of a callable region outside the pre-defined proximity of the variant or mutation after identifying the variant or mutation in the alignment determination data. 23. The alignment determination system according to clause 22, wherein at least one of the variant or mutation, or the pre-defined proximity of the specified partition, is determined based on population data accessed by the colling subsystem. 24. The alignment determination system according to clause 22, wherein the variant or mutation includes an insertion or deletion. 25. At least one processor is Further configured to analyze the buffered alignment determination data to identify positions for partitioning a callable region within the buffered alignment determination data, the alignment determination system according to clause 18. 26. At least one processor is Identify a portion of the buffered alignment determination data having a read depth less than a partition threshold, The alignment determination system according to clause 25, further configured to perform a partition of the buffered alignment determination data in the identified portion having a read depth less than the partition threshold. 27. At least one processor is configured to transmit a callable region to a genotyping subsystem when the depth of a plurality of reads is less than a callable region depth threshold used to detect the callable region, the alignment determination system according to clause 18. 28. The partitioned portion of the callable region is the entire callable region currently present in the memory used by the colling subsystem, the alignment determination system according to clause 18. The splitting part is the first splitting part, and at least one processor maintains a predefined amount of sequencing data within the first splitting part of the callable region in the memory used by the colling subsystem of the variant caller, detects a second splitting part of the callable region based on a callable region depth threshold or a memory threshold, sends the second splitting part of the callable region to the genotyping subsystem of the variant caller for variant calling based on the second splitting part, and the second splitting part is further configured to include an overlap with the first splitting part of the callable region that contains a predefined amount of sequencing data within the sequencing data, the sequencing system according to clause 18. 30. The predefined amount of overlapping sequencing data is a predefined number of bases, the sequencing system according to clause 29. 31. The predefined amount of overlapping sequencing data is determined based on the population data accessed by the colling subsystem, the sequencing system according to clause 29. 32. The predefined amount of overlapping sequencing data is determined based on user input, the sequencing system according to clause 29. 33. At least one processor identifies the overlap within the sequencing data between the first splitting part and the second splitting part of the callable region in the genotyping subsystem of the variant caller, is further configured to remove the overlap within the sequencing data between the first splitting part and the second splitting part of the callable region before performing variant calling on the callable region, the sequencing system according to clause 29. 34. The callable region depth threshold is the first callable region depth threshold, the splitting part is the first splitting part, and at least one processor monitors the depth of multiple reads in the colling subsystem of the variant caller, If the depths of the plurality of reads are less than the second callable region depth threshold, divide the callable region, The sequencing system according to clause 18, further configured to transmit a second divided portion of the callable region to a genotyping subsystem of the variant caller for variant calling based on the second divided portion. 35. At least one computer-readable recording medium having computer-executable instructions recorded thereon, which, when executed by at least one processor, cause the at least one processor to, In a calling subsystem of a variant caller, receive sequencing data including a plurality of reads of a genomic sequence, the calling subsystem of the variant caller being configured to detect a callable region of the sequencing data when the depths of the plurality of reads exceed a callable region depth threshold, the calling subsystem of the variant caller being configured to transmit at least a portion of the callable region to a genotyping subsystem of the variant caller for variant calling of the callable region, Monitor the memory used by the calling subsystem of the variant caller, If the memory used by the calling subsystem of the variant caller exceeds a memory threshold of the total amount of memory allocated to the calling subsystem of the variant caller, divide the callable region, At least one computer-readable recording medium that causes a divided portion of the callable region to be transmitted to a genotyping subsystem of the variant caller for variant calling based on the divided portion. 36. The computer-executable instructions are configured to cause at least one processor to transmit a divided portion of a callable region to a genotyping subsystem to increase the availability of memory used by a calling subsystem, the at least one computer-readable recording medium according to clause 35. 37. The memory threshold is a fixed threshold, the at least one computer-readable recording medium according to clause 35. 38. The memory threshold is a dynamic threshold, at least one computer-readable recording medium according to clause 35. 39. The computer-executable instructions cause at least one processor to In the calling subsystem of the variant caller, when the memory used by the calling subsystem of the variant caller is within the memory threshold of the total amount of memory allocated to the calling subsystem of the variant caller, analyze the alignment determination data, Identify a variant or mutation in the alignment determination data within a pre-defined proximity of the identified partition, After identifying a variant or mutation in the alignment determination data, the alignment determination system according to clause 35, which is configured to perform a partition of the callable region outside the pre-defined proximity of the variant or mutation. 40. At least one computer-readable recording medium according to clause 39, wherein at least one of the variant or mutation, or the pre-defined proximity of the identified partition, is determined based on population data accessed by the calling subsystem. 41. The variant or mutation includes an insertion or deletion, at least one computer-readable recording medium according to clause 39. 42. The computer-executable instructions cause at least one processor to Analyze the buffered alignment determination data to identify positions for partitioning the callable regions within the buffered alignment determination data, at least one computer-readable recording medium according to clause 35. 43. The computer-executable instructions cause at least one processor to Identify a portion of the buffered alignment determination data having a read depth less than a partitioning threshold, At least one computer-readable recording medium according to clause 42, which is configured to cause a partition of the buffered alignment determination data in the identified portion having a read depth less than the partitioning threshold. 44. The computer-executable instructions are configured to cause at least one processor to transmit a callable region to a genotyping subsystem if the depth of a plurality of reads is less than a callable region depth threshold used to detect the callable region, for at least one computer-readable recording medium according to clause 35. 45. A divided portion of the callable region is the entire callable region currently present in the memory used by the calling subsystem, for at least one computer-readable recording medium according to clause 35. 46. The divided portion is a first divided portion, and the computer-executable instructions cause at least one processor to maintain a predefined amount of sequencing data within a first divided portion of the callable region in the memory used by the calling subsystem of the variant caller, detect a second divided portion of the callable region based on a callable region depth threshold or a memory threshold, transmit the second divided portion of the callable region to the genotyping subsystem of the variant caller for variant calling based on the second divided portion, and the second divided portion is further configured to include an overlap with the first divided portion of the callable region that contains a predefined amount of sequencing data within the sequencing data, for a sequencing system according to clause 35. 47. The predefined amount of overlapping sequencing data is a predefined number of bases, for at least one computer-readable recording medium according to clause 46. 48. The predefined amount of overlapping sequencing data is determined based on population data accessed by the calling subsystem, for at least one computer-readable recording medium according to clause 46. 49. The predefined amount of overlapping sequencing data is determined based on user input, for at least one computer-readable recording medium according to clause 46. 50. The computer-executable instructions cause at least one processor to In the genotyping subsystem of the variant caller, identify duplicates in the sequencing data between the first and second split portions of the callable region. The at least one computer-readable recording medium according to clause 46, configured to remove duplicates in the sequencing data between the first and second split portions of the callable region before performing variant calling on the callable region. 51. The callable region depth threshold is the first callable region depth threshold, the split portion is the first split portion, and the computer-executable instructions are for at least one processor In the calling subsystem of the variant caller, monitor the depths of a plurality of reads. If the depths of the plurality of reads are less than a second callable region depth threshold, split the callable region. The at least one computer-readable recording medium according to clause 35, configured to send the second split portion of the callable region to the genotyping subsystem of the variant caller for variant calling based on the second split portion. 52. A computer-implemented method, Receiving, in the calling subsystem of the variant caller, sequencing data including a plurality of reads of a genomic sequence, wherein the calling subsystem of the variant caller is configured to detect a callable region of the sequencing data when the depths of the plurality of reads are less than a depth threshold, and the calling subsystem of the variant caller is configured to send at least a portion of the callable region to the genotyping subsystem of the variant caller for variant calling of the callable region; Monitoring the memory used by the calling subsystem of the variant caller; Spilling the callable region to disk if the memory used by the calling subsystem of the variant caller is within a memory threshold of the total amount of memory allocated to the calling subsystem of the variant caller; A computer-implemented method including streaming back a spilled callable region from a disk to memory for use by a genotyping subsystem for processing. 53. Spilling the callable region to disk includes spilling a first portion of the callable region to disk and maintaining a second portion of the callable region in memory, the computer-implemented method of clause 52. 54. The second portion of the callable region is sent to the genotyping subsystem for processing by a pointer in memory, maintaining the second portion of the callable region in memory, and the first portion of the callable region is streamed from disk, the computer-implemented method of clause 53. 55. Analyzing the spilled callable region streamed back from disk; Before streaming an additional portion of the spilled callable region to memory, discarding one or more portions of the spilled callable region from memory not used for variant calling, further including the computer-implemented method of clause 53. 56. The memory threshold is a first memory threshold associated with the calling subsystem, and the method Monitors a second memory threshold associated with the genotyping subsystem; Prevents exceeding the second memory threshold while streaming the callable region spilled from disk, further including the computer-implemented method of clause 53. 57. An array determination system comprising Memory; At least one processor, In a variant caller's calling subsystem, cause sequencing data including a plurality of reads of a genomic sequence to be received, the variant caller's calling subsystem being configured to detect a callable region of the sequencing data when the depth of the plurality of reads is less than a depth threshold, the variant caller's calling subsystem being configured to transmit at least a portion of the callable region to a genotyping subsystem of the variant caller for variant calling of the callable region, Cause the memory used by the variant caller's calling subsystem to be monitored, When the memory used by the variant caller's calling subsystem is within a memory threshold of the total amount of memory allocated to the variant caller's calling subsystem, spill the callable region to disk, At least one processor configured to stream the spilled callable region back from disk to the memory used by the genotyping subsystem for processing, a sequencing system comprising the at least one processor. 58. The sequencing system according to clause 57, wherein the at least one processor is configured to spill the callable region to disk by spilling a first portion of the callable region to disk and further configured to maintain a second portion of the callable region in memory. 59. The sequencing system according to clause 58, wherein the at least one processor is further configured to transmit a second portion of the callable region to the genotyping subsystem for processing by a pointer in memory and to maintain the second portion of the callable region in memory, and the at least one processor is further configured to stream a first portion of the callable region from disk. 60. The at least one processor is Analyze the spilled callable region streamed back from disk, The alignment determination system according to clause 58, further configured to discard one or more portions of the spilled callable region from memory that is not used for variant calling before streaming the additional portion of the spilled callable region to memory. 61. The memory threshold is a first memory threshold associated with the calling subsystem, and at least one processor monitors a second memory threshold associated with the genotyping subsystem, and is further configured to prevent exceeding the second memory threshold while streaming the callable region spilled from disk. The alignment determination system according to clause 58. 62. At least one computer-readable recording medium having computer-executable instructions recorded thereon, which, when executed by at least one processor, cause the at least one processor to receive alignment data including a plurality of reads of a genomic sequence in the calling subsystem of the variant caller, the calling subsystem of the variant caller being configured to detect a callable region of the alignment data when the depth of the plurality of reads is less than a depth threshold, the calling subsystem of the variant caller being configured to transmit at least a portion of the callable region to the genotyping subsystem of the variant caller for variant calling of the callable region, monitor the memory used by the calling subsystem of the variant caller, spill the callable region to disk when the memory used by the calling subsystem of the variant caller is within a memory threshold of the total amount of memory allocated to the calling subsystem of the variant caller, and stream the spilled callable region back from disk to the memory used by the genotyping subsystem for processing. At least one computer-readable recording medium. 63. The computer-executable instructions are configured to spill the callable region to disk by spilling at least a first portion of the callable region to disk and further configured to maintain a second portion of the callable region in memory, for at least one of the computer-readable recording media according to clause 62. 64. The computer-executable instructions are configured to cause at least one processor to send a second portion of the callable region to a genotyping subsystem for processing by a pointer in memory, to maintain the second portion of the callable region in memory, and to further configure at least one processor to stream a first portion of the callable region from disk, for at least one of the computer-readable recording media according to clause 63. 65. The computer-executable instructions are for at least one processor to analyze the spilled callable region streamed back from disk, and to discard one or more portions of the spilled callable region from memory not used for variant calling before streaming an additional portion of the spilled callable region to memory, for at least one of the computer-readable recording media according to clause 63. 66. The memory threshold is a first memory threshold associated with a calling subsystem, and the computer-executable instructions are configured to cause at least one processor to monitor a second memory threshold associated with a genotyping subsystem, and to prevent exceeding the second memory threshold while streaming the callable region spilled from disk, for at least one of the computer-readable recording media according to clause 63.
Explanation of Signs
[0082] 100 Environment 102 Server Device 104 Array Determination Subsystem 108 Client Device 110 Array Determination Application 112 Network 114 Array Determination Device 116 Database 122 Mapper Subsystem 123 Disk 124 Sorter Subsystem 126 Variant Caller Subsystem 128 Calling Subsystem 130 Genotype Determination Subsystem 200, 200a, 200b, 200c, 200d Graphs 202 Read 204, 206 Genome Regions 208 Location 210 Location 212 Callable Region 214 Location 216 Location 218 Split Portion 220 Portion 222 Genome Region 224, 226 Genome Regions 230 Duplicate Portion 300 Procedure 306 Callable Region 400 Procedure 500 Computing Device 502 Processor 504 Memory 506 Storage Device 508 I / O Interface 510 Communication Interface 512 Communication Infrastructure
Claims
1. A method implemented by a computer, comprising: receiving sequencing data including a plurality of reads of a genomic sequence in a calling subsystem of a variant caller, wherein the calling subsystem of the variant caller is configured to detect a callable region of the sequencing data when the depth of the plurality of reads exceeds a callable region depth threshold, and the calling subsystem of the variant caller is configured to transmit at least a portion of the callable region to a genotyping subsystem of the variant caller for variant calling of the callable region; monitoring a memory used by the calling subsystem of the variant caller; when the memory used by the calling subsystem of the variant caller exceeds a memory threshold of a total amount of memory allocated to the calling subsystem of the variant caller, dividing the callable region; transmitting the divided portions of the callable region to the genotyping subsystem of the variant caller for variant calling based on the divided portions.
2. The method implemented by a computer according to claim 1, wherein transmitting the divided portions of the callable region to the genotyping subsystem increases the availability of the memory used by the calling subsystem.
3. The method implemented by a computer according to claim 1, wherein the memory threshold is a fixed threshold.
4. The method implemented by a computer according to claim 1, wherein the memory threshold is a dynamic threshold.
5. The method further comprises: in the calling subsystem of the variant caller, when the memory used by the calling subsystem of the variant caller is within the memory threshold of the total amount of memory allocated to the calling subsystem of the variant caller, analyzing the sequencing data; identifying variants or mutations in the sequencing data within a pre-defined proximity of a specified division. After identifying the variant or mutation in the array determination data, dividing the callable region outside the pre-defined proximity of the variant or mutation. The method implemented by a computer according to claim 1 includes this step.
6. The method implemented by a computer according to claim 5, wherein at least one of the variant or mutation, or the pre-defined proximity of the identified division, is determined based on population data accessed by the calling subsystem.
7. The method implemented by a computer according to claim 5, wherein the variant or mutation includes an insertion or deletion.
8. The method implemented by a computer according to claim 1 further includes analyzing the buffered array determination data to identify positions for dividing the callable region within the buffered array determination data.
9. Identifying a portion of the buffered array determination data having a read depth less than a division threshold; Dividing the buffered array determination data in the identified portion having the read depth less than the division threshold. The method implemented by a computer according to claim 8 includes these additional steps.
10. The method implemented by a computer according to claim 1, wherein if the depth of the plurality of reads is less than the callable region depth threshold used to detect the callable region, the callable region is transmitted to the genotyping subsystem.
11. The method implemented by a computer according to claim 1, wherein the divided portion of the callable region is the entire callable region currently present in the memory used by the calling subsystem.
12. The divided portion is the first divided portion, and the method further includes: Maintaining a pre-defined amount of the array determination data within the first divided portion of the callable region in the memory used by the calling subsystem of the variant caller; Detecting a second divided portion of the callable region based on the callable region depth threshold or the memory threshold. Transmitting the second divided portion of the callable region to the genotyping subsystem of the variant caller for the variant calling based on the second divided portion, wherein the second divided portion includes an overlap with the first divided portion of the callable region that includes the predefined amount of sequencing data within the sequencing data, and the method implemented by a computer according to claim 1 includes transmitting.
13. The method implemented by a computer according to claim 12, wherein the predefined amount of sequencing data of the overlap is a predefined number of bases.
14. The method implemented by a computer according to claim 12, wherein the predefined amount of sequencing data of the overlap is determined based on population data accessed by the calling subsystem.
15. The method implemented by a computer according to claim 12, wherein the predefined amount of sequencing data of the overlap is determined based on user input.
16. In the genotyping subsystem of the variant caller, identifying the overlap in the sequencing data between the first divided portion and the second divided portion of the callable region; Before performing the variant calling on the callable region, further removing the overlap in the sequencing data between the first divided portion and the second divided portion of the callable region, and the method implemented by a computer according to claim 12 includes this.
17. The callable region depth threshold is a first callable region depth threshold, the divided portion is a first divided portion, and the method implemented by a computer further includes monitoring the depth of the plurality of reads in the calling subsystem of the variant caller; when the depth of the plurality of reads is less than a second callable region depth threshold, dividing the callable region; transmitting the second divided portion of the callable region to the genotyping subsystem of the variant caller for the variant calling based on the second divided portion, and the method implemented by a computer according to claim 1 includes this.
18. A method implemented by a computer, comprising In a variant caller's calling subsystem, receiving sequencing data including a plurality of reads of a genomic sequence, wherein the calling subsystem of the variant caller is configured to detect a callable region of the sequencing data when the depth of the plurality of reads is less than a depth threshold, and the calling subsystem of the variant caller is configured to transmit at least a part of the callable region to a genotyping subsystem of the variant caller for variant calling of the callable region, the receiving; monitoring memory used by the calling subsystem of the variant caller; spilling the callable region to disk when the memory used by the calling subsystem of the variant caller is within a memory threshold of the total amount of memory allocated to the calling subsystem of the variant caller; streaming back the spilled callable region from the disk to the memory used by the genotyping subsystem for processing, a method implemented by a computer.
19. Spilling the callable region to disk includes spilling a first portion of the callable region to disk and maintaining a second portion of the callable region in the memory, the method implemented by a computer according to claim 18.
20. The second portion of the callable region is transmitted to the genotyping subsystem for processing by a pointer in the memory, maintaining the second portion of the callable region in the memory, and the first portion of the callable region is streamed from the disk, the method implemented by a computer according to claim 19.
21. analyzing the spilled callable region streamed back from the disk; discarding one or more portions of the spilled callable region from memory not used for the variant calling before streaming an additional portion of the spilled callable region to the memory, the method implemented by a computer according to claim 19.
22. The memory threshold is a first memory threshold associated with the colling subsystem, and the method comprises: monitoring a second memory threshold associated with the genotyping subsystem; and preventing the second memory threshold from being exceeded while streaming the spillable callable region from the disk. The computer-implemented method according to claim 19 further comprises.