Device and method for locating sample reads in a reference genome

By using reference guide devices in genome sequencing, identifying the matching locations of sample substring sequences and reference sequences, the scalability limitations of genome sequencing memory resources and computational costs in the prior art are solved, and more efficient genome sequencing is achieved.

CN114730615BActive Publication Date: 2025-05-27WESTERN DIGITAL TECHNOLOGIES INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080080230.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-17
Filing Date
2020-07-01
Publication Date
2025-05-27
Estimated Expiration
2040-07-01

AI Technical Summary

Technical Problem

Existing genome sequencing technologies have scalability limitations in memory resources and computational costs, especially in de novo sequencing and reference alignment sequencing, requiring a large amount of memory resources and high computational costs to process and compare large amounts of DNA sample reads.

Method used

Using a reference guide device, by storing the overlapping reference sequence of the reference genome in the reference guide device and using a circuit to identify the matching positions of the sample substring sequence and the reference sequence, a probability position index of the sample read segments within the reference genome is provided to replace the algorithm executed by the host and improve sequencing efficiency.

Benefits of technology

Reduces memory resources and computational costs for genome sequencing, improves sequencing scalability, and reduces the time and cost of performing de novo or reference alignment genome sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114730615B_ABST
    Figure CN114730615B_ABST
Patent Text Reader

Abstract

The present invention provides an apparatus for localizing sample reads relative to a reference genome, the apparatus including a plurality of unit groups. Each unit group stores a reference sequence representing reference bases from the reference genome, the reference sequence corresponding to the order of units in the corresponding unit group. Each unit group also stores a current substring sequence representing sample bases from the sample reads, the current substring sequence corresponding to the order of units in the corresponding unit group. Each unit group stores the same current substring sequence and a reference sequence representing a portion of the reference genome, the portion of the reference genome partially overlapping at least one other portion of the reference genome represented by one or more other reference sequences stored in one or more other unit groups. Identify, among the plurality of unit groups, the unit groups in which the stored reference sequences match the current substring sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of co - pending U.S. application Ser. No. 16 / 821,849, filed Mar. 17, 2020, and entitled "REFERENCE - GUIDED GENOME SEQUENCING" (Attorney Docket No. WDA - 4724 - US), the entire content of which is incorporated herein by reference. This application also claims the benefit of co - pending U.S. application Ser. No. 16 / 822,010, filed Mar. 18, 2020, and entitled "REFERENCE - GUIDED GENOME SEQUENCING" (Attorney Docket No. WDA - 4725 - US), the entire content of which is incorporated herein by reference. Background of the Invention

[0003] Current limitations in DNA (deoxyribonucleic acid) sample processing result in sample reads or portions of the sample genome having generally unknown positions within the sample genome. For de novo sequencing that does not use a reference genome when comparing sample reads to each other to localize the sample reads within the sample genome, the sample reads are typically analyzed as a single large group, which requires substantial memory resources and high computational costs to compare the sample reads within the large group to determine the positions of the sample reads within the sample genome. Such conventional methods of de novo sequencing are not scalable with respect to the large amounts of data that need to be processed for genome sequencing. More specifically, conventional de novo sequencing methods typically store a large group of sample reads in shared memory such as expensive 2TB DRAM. Since the number of compute cores that can be connected to the shared DRAM via independent high - bandwidth channels is limited (e.g., at most 24 cores), this arrangement limits the number of independent computational threads available for de novo sequencing (e.g., at most 128 computational threads).

[0004] For reference alignment sequencing that locates sample reads within a sample genome using a reference genome, typically each sample read is searched against the entire reference genome to locate the sample read within the reference genome. Such reference alignment sequencing also requires a large amount of memory resources to store the entire reference genome and high computational costs to compare each sample read with the entire reference genome. Conventional methods of reference alignment sequencing also have limited scalability. More specifically, conventional methods of reference alignment sequencing can randomly divide sample reads into groups that are processed by corresponding computational threads. However, each computational thread typically requires a large dedicated memory such as 16GB DRAM to store the entire reference genome. In other techniques, the reference genome can be stored in a single shared 16GB DRAM, but as noted above for conventional de novo sequencing, this shared memory arrangement limits the number of cores and computational threads that can access the shared memory. Accordingly, there is a need for improvement in genomic sequencing in terms of computational cost, memory resources, and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The features and advantages of the embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. The accompanying drawings and the associated description are provided to illustrate the embodiments of the present disclosure and are not intended to limit the scope of the claimed subject matter.

[0006] Figure 1 is a block diagram of a system for genomic sequencing including a reference-guided device according to one or more embodiments.

[0007] Figure 2 illustrates an example of multiple unit groups in a reference-guided device according to one or more embodiments.

[0008] Figure 3 is a graph depicting the uniqueness of substrings of different lengths in the human reference genome H38.

[0009] Figure 4A illustrates an example of identifying a unit group in a reference-guided device in which a currently stored substring sequence matches a reference sequence according to one or more embodiments.

[0010] Figure 4B is an example of a circuit for comparing substring base values with reference base values stored in a unit according to one or more embodiments.

[0011] Figure 4C is an example of a circuit for comparing unit output values in a unit group according to one or more embodiments.

[0012] Figure 5 is a flowchart of a sample read location process according to one or more embodiments.

[0013] Figure 6 is a flowchart of a matching recognition sub - process using logical operations according to one or more embodiments.

[0014] Figure 7 is a flowchart of a matching recognition sub - process using the inner product of a reference vector and a substring vector according to one or more embodiments. Detailed Description of the Invention

[0015] Numerous specific details are set forth in the following detailed description in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that various embodiments may be practiced without some of these specific details. In other instances, well - known structures and techniques have not been shown in detail to avoid unnecessarily obscuring the various embodiments.

[0016] System Example

[0017] Figure 1 is a block diagram of a system 100 for genome sequencing according to one or more embodiments. The system includes a host 101 and a reference - guiding device 102. The host 101 communicates with the reference - guiding device 102 to determine the probabilistic location of a sample read within a reference genome. In some specific implementations, the device 102 may provide an index 10 indicating the probabilistic location of the sample read stored in the memory 108 of the device 102 to the host 101. In other specific implementations, the device 102 may provide another data structure or indication of the probabilistic location of the sample read to the host 101.

[0018] The sample read or a sample substring sequence obtained from the sample read may initially be provided to the reference - guiding device 102 by the host 101 and / or by Figure 1 another device not shown (such as an additional host) to determine the probabilistic location of the sample read within a reference genome stored in one or more arrays 104 of the device 102. In some specific implementations, a read device that generates the sample read, such as an Illumina device (obtained from Illumina, Inc., San Diego, California) or a nanopore device, may provide the sample read to the reference - guiding device 102.

[0019] For ease of description, exemplary embodiments in the present disclosure will be described in the context of DNA sequencing. However, the embodiments of the present disclosure are not limited to DNA sequencing and can generally be applied to any nucleic - acid - based sequencing, including RNA (ribonucleic acid) sequencing.

[0020] The host 101 may include, for example, a computer, such as a desktop or a server, which may implement genomic sequencing algorithms, such as a seed and extend algorithm for exact matching and / or a computationally more complex algorithm for approximate matching of sample reads in a genome, such as the Burrows-Wheeler algorithm or the Smith-Waterman algorithm. As discussed in more detail below, the device 102 may be used to preprocess sample reads prior to de novo or reference alignment sequencing. In this regard, the probabilistic positions provided by the reference-guided device 102 may replace or improve the efficiency of the algorithms executed by the host 101 in terms of memory resources and computational costs. Additionally, and as described in related co-pending applications 16 / 821,849 and 16 / 822,010, both of which are incorporated by reference above, the probabilistic positions of the sample reads provided by the device 102 may allow for an improvement in the scalability of genomic sequencing, thereby reducing the cost and time of performing de novo or reference alignment genomic sequencing.

[0021] In some specific embodiments, the reference-guided device 102 may include, for example, one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) for generating an index 10 indicating the probabilistic position of a sample substring sequence from a sample read relative to a reference genome. The probabilistic position of the sample substring sequence may provide the probabilistic position of the sample read from which the sample substring sequence was obtained to the host 101. In some specific embodiments, the host 101 or another device may provide the current sample substring sequence to the reference-guided device 102 to load into one or more arrays 104 of the device 102. In other specific embodiments, the host 101 or another device may provide the sample read to the reference-guided device 102, and the reference-guided device 102 may determine the sample substring sequence to load into one or more arrays 104 from the sample read.

[0022] The host 101 and the device 102 may be physically co-located or may not be physically co-located. For example, in some specific embodiments, the host 101 and the device 102 may communicate via a network, such as by using a local area network (LAN) or a wide area network (WAN), such as the Internet, or a data bus or a data network structure. Additionally, those of ordinary skill in the art will understand that other specific embodiments may include multiple hosts 101 and / or multiple devices 102 for providing the probabilistic positions of sample reads. In certain embodiments, the host 101 and the device 102 (or multiple hosts and devices) are integrated into a single device or system.

[0023] As Figure 1As shown in the example of, device 102 includes one or more arrays 104. As used herein, a cell generally refers to a memory location for storing one or more values representative of one or more nucleotides (referred to as bases in the present disclosure). In some specific embodiments, one or more arrays 104 may include such cells that also include logic for performing one or more operations on the values stored in the cells. In such examples, each cell in one or more arrays may store a reference value representative of a reference base from a reference genome and a sample value representative of a base from a sample substring sequence. The cell may perform one or more operations to output a value that can be used by circuit 106 or the circuits of one or more arrays 104 to determine whether a group of cells in the one or more arrays 104 stores a reference sequence that matches the substring sequence stored in the group of cells. In some specific embodiments, one or more arrays 104 may include one or more systolic arrays in which reference values representative of reference bases from a reference genome are loaded, and sample values representative of bases from a sample substring sequence may be loaded into the cells for comparison with the reference values before passing the sample values to the next cell in another group of cells of the one or more arrays 104.

[0024] In other specific embodiments, one or more arrays 104 may include solid-state memory cells that may not perform operations to determine whether the values stored in the cells match. For example, in some specific embodiments, circuit 106 may determine whether the values stored in each cell match. As another variation, one or more arrays 104 may each store a reference value representative of a reference base or a sample value representative of a sample base. In such specific embodiments, cells storing reference values may be paired with cells storing sample values for comparison of the reference bases with the sample bases. In other specific embodiments, the cells in one or more arrays 104 may include circuit elements such as registers, latches, or flip-flops.

[0025] Although the description herein generally refers to solid-state memory, it should be understood that solid-state memory may include one or more of various types of memory devices, such as flash integrated circuits, chalcogenide RAM (C-RAM), phase change memory (PC-RAM or PRAM), programmable metallization cell RAM (PMC-RAM or PMCm), Ovonic unified memory (OUM), resistive RAM (RRAM), NAND memory (e.g., single-level cell (SLC) memory, multi-level cell (MLC) memory (i.e., two or more levels), or any combination thereof), NOR memory, EEPROM, ferroelectric memory (FeRAM), magnetoresistive RAM (MRAM), other discrete non-volatile memory (NVM) chips, or any combination thereof.

[0026] The circuit 106 may include, for example, hardwired logic, analog circuits, and / or combinations thereof. In other specific embodiments, the circuit 106 may include one or more ASICs, microcontrollers, digital signal processors (DSPs), FPGAs, and / or combinations thereof. In some specific embodiments, the circuit 106 may include one or more system-on-chips (SoCs) that may be combined with the memory 108. As discussed in more detail below, the circuit 106 is configured to identify a cell group in one or more of the arrays 104 in which the stored reference sequence matches the current substring sequence stored in the cell group.

[0027] More specifically, for each cell group in one or more of the arrays 104, a reference sequence of reference bases from the reference genome may be stored in the cell group. The reference sequence corresponds to the order of the cells in the corresponding cell group. Each cell group is configured to store a reference sequence representing a portion of the reference genome that partially overlaps with at least one other portion of the reference genome represented by one or more other reference sequences stored in one or more other cell groups. Examples of the storage of such overlapping reference sequences in the array are discussed in more detail below with reference to Figure 2 more detail.

[0028] In addition, each cell group in one or more of the arrays 104 is configured to store the same current substring sequence corresponding to the order of the corresponding cell group. As described above, the circuit 106 is configured to identify a cell group in one or more of the arrays 104 in which the stored current substring sequence matches the reference sequence stored in the cell group. In some specific embodiments, the circuit 106 may identify the cell group with the matching sequence based on the value output from the cell after performing at least one logical operation (such as an XNOR operation). In other specific embodiments, the circuit 106 may identify the cell group with the matching sequence based on the value output from the cell after multiplying a reference value representing a reference base and a sample value representing a sample base. In other specific embodiments, the circuit 106 may perform all operations on the values stored in the cells, rather than having some operations performed by the cells themselves.

[0029] The memory 108 of the device 102 may include, for example, volatile memory such as DRAM for storing the index 10. In other specific embodiments, the memory 108 may include non-volatile memory such as MRAM. As Figure 1As shown, the memory 108 stores an index 10 that can be used by the host 101 to determine the probabilistic location of a sample read segment within a reference genome represented by overlapping reference sequences loaded into or stored in one or more arrays 104. In some embodiments, the index 10 can include a data structure, such as a bitmap or other data structure, that indicates the indices or locations in the reference genome corresponding to groups of units identified as storing matching sequences. The circuit 106 can update the index 10 for different sample substring sequences loaded into each group of units of one or more arrays 104. In some embodiments, the circuit 106 can indicate the average location in the index 10 of a substring sequence having multiple matching groups of units. In other embodiments, for a particular substring sequence, only the first matching group of units can be used, or for a substring sequence having more than a single group of units storing matching sequences, the circuit 106 can not update the index 10 at all.

[0030] In addition, some embodiments may not use an index or other data structure to indicate the location of groups of units having matching sequences. For example, in some embodiments, the circuit 106 can directly output to the host 101 data indicating groups of units having matching sequences.

[0031] As those of ordinary skill in the art will understand with reference to this disclosure, other embodiments can include a different number or arrangement of components than those shown in the example of Figure 1 system 100. For example, other embodiments can combine the host 101 and the device 102, or can include a different number of devices 102 and / or hosts 101.

[0032] Figure 2 An example of multiple groups of units in a reference-guided device 102 according to one or more embodiments is shown. As Figure 2 shown in the example, the array 104 includes groups of units 110 1 through 110 L-19 . Although Figure 2 the groups of units 110 in 1 through 110 L-19are shown as columns, but other embodiments may include groups of units that are not physically arranged as columns. In some embodiments, the array 104 may replace a defective unit from one group of units with another unit located in a different part of the same array or in a pool of spare units in a different array. In embodiments where each group of units stores an overlapping reference sequence that has been shifted one reference base relative to the previous group of units, L may be equal to the full length of the reference genome, such as 3.2 billion groups of units or unit columns, as in the case of the complete reference human genome H38. In contrast, other embodiments may store overlapping reference sequences that have been shifted a different number of reference bases (e.g., shifted two reference bases), such that fewer groups of units or columns are needed, which allows for a smaller size of the array 104. However, shifting the overlap by more than one reference base may come at the cost of reducing the likelihood of finding a substring sequence match.

[0033] As Figure 2 shown in the example of, each group of units 110 stores a reference value (e.g., R1, R2, R3, etc.) representing a reference base and a sample value (S1, S2, S3, etc.) representing a sample base. For example, in the case of DNA sequencing, since there are four possible bases: adenine (A), guanine (G), cytosine (C), and thymine (T), each reference value and each sample value may be represented by two bits. When each group of units 110 stores the same sample sequence of sample values S1 to S20, each group of units 110 stores a different partially overlapping reference sequence that is shifted one reference base relative to the reference sequence stored in an adjacent group of units. For example, the group of units 110 1 stores a first reference sequence having reference values R1 to R20, while the group of units 110 2 stores a second reference sequence having reference values R2 to R21. In other embodiments, the shift offset and the resulting overlap may be different between groups of units compared to what is shown in the example of Figure 2 .

[0034] The arrangement of storing partially overlapping reference sequences and substring sequences in the array 104 generally allows for the efficient localization of the probabilistic positions of sample reads within the reference genome. Additionally, the reference sequences only need to be loaded into or stored in the array 104 once. Then, the iteration of loading or storing different substring sequences from the sample reads can provide the probabilistic positions of the sample reads within the reference genome, which can be used by the host 101 to intelligently classify the sample reads into read groups for more efficient de novo or reference alignment sequencing, as discussed in co-pending related applications 16 / 821,849 and 16 / 822,010, which are incorporated herein by reference. In this regard, different embodiments may use a first type of cell (such as ROM or NAND flash cells) to store the reference sequences, and a second type of cell (such as MRAM cells) to store the substring sequences, which second type of cell is more suitable for repeated rewriting with better write durability.

[0035] In Figure 2 the example used, the substring sequence length is 20 and includes sample values S1 to S20. As discussed in more detail below with reference to Figure 3 the length of the substring sequence corresponding to the number of cells in the cell group or column can be selected based on the desired uniqueness of the substring sequence within the reference genome relative to the number of cells and the operations required to identify the cell group or column storing the matching sequence.

[0036] Figure 3 is a graph depicting the uniqueness of substring sequences of different lengths in the human reference genome H38. Figure 3 The dashed line in Figure 3 represents the expected distribution if each base in the reference genome H38 is randomly selected uniformly for different substring lengths indicated along the x-axis. Figure 3 The solid line in

[0037] As Figure 3 shown by the solid line in Figure 3As shown, when the substring length is shorter than 15 bases, it may not be possible to identify any unique match within H38 for almost all of the attempted substring sequences.

[0038] On the other hand, when the substring length is greater than 25 bases, it will result in additional storage costs for the cells in one or more of the arrays 104, and greater computational costs due to the increased operations, while there is little improvement in the number of unique matches. As a result, the example discussed above Figure 2 uses a substring length of 20 bases, which means Figure 2 each cell group 110 in

[0039] Figure 4A includes a predetermined number of 20 cells. Those of ordinary skill in the art, referring to the present disclosure, will understand that for other examples, different substring lengths or different predetermined numbers of cells in each cell group may be preferred, such as when using a different reference genome or a portion of a reference genome, as may be the case for medical diagnostics of genetic conditions related to a specific portion of the reference genome. Additionally, different trade - offs between computational cost, the number of cells, and accuracy in terms of a greater number of unique matches may also affect the number of cells used for each cell group in one or more of the arrays 104. Figure 4A As shown, the array 104 includes a plurality of cell groups, as in the example of Figure 2 discussed above. In the example of Figure 4A , each cell group is represented by a column number i ranging from 1 to L-(M - 1). Each cell in each cell group or column is also represented by a row number j ranging from 1 to M. As discussed above, L-(M - 1) may correspond to the number of overlapping reference sequences from the reference genome, and M may correspond to the number of bases in the substring sequence, such as 20 bases, as in the exemplary array 104 of Figure 2 .

[0040] The reference sequences of the reference genome may be loaded or stored in the cell groups, where each cell stores a reference value representing a reference base from the reference sequence. As described above, the reference sequences from one cell column or group to the next group or column may overlap a predetermined number of reference values or reference bases, such as overlapping one, two, or three reference values or bases. The order of the cells in the group or column corresponds to the order of the reference bases in the reference sequence. In some specific embodiments, before transporting the reference - guiding device to the customer, the reference sequences may be initially loaded or stored by the manufacturer of the reference - guiding device for a specific reference genome. In other specific embodiments, the reference sequences may be loaded or stored by the customer on - site.

[0041] The current substring sequence is loaded or stored in a group of cells, where each cell stores a sample value representing a sample base from the current substring sequence. Each group of cells or column may store the same current substring sequence. Additionally, the order of the cells in the group or column corresponds to the order of the sample bases in the current substring sequence. In some embodiments, the array 104 may include a systolic array, where the current substring sequence is passed from one group of cells or column to the next.

[0042] As discussed below with reference Figure 4B and Figure 4C more particularly, a comparison is made between the reference value and the sample value in each cell such as cell i,j, and each cell provides a cell output value to the circuit 106 to identify the column or group of cells in which all reference values match all substring values. Then, the position of the matching column or group of cells can be used to update a data structure such as Figure 1 the index 10 in. In other embodiments, the position of the matching column or group of cells may alternatively be provided to another device such as Figure 1 the host 101 in, without updating the data structure.

[0043] Figure 4B is an example of a circuit for comparing substring base values with reference base values stored in cells according to one or more embodiments. As described above, each substring base and reference base may be represented by two bits. For example, an A base may be represented by the binary value 00, a C base may be represented by the binary value 01, a G base may be represented by the binary value 10, and a T base may be represented by the binary value 11. In other embodiments, these bases may be represented by other values, as in the example of using the inner product discussed below with reference Figure 7 where the bases may have values including 1 or -1.

[0044] As Figure 4B the example of shows, the circuit within cell i,j includes two XNOR gates that output to an AND gate. More particularly, the first bit of the substring base value i,j stored in cell i,j is input to the first XNOR gate together with the first bit of the reference base value i,j stored in cell i,j. The second bit of the substring base value i,j is input to the second XNOR gate together with the second bit of the reference base value i,j. If the two inputs to the XNOR gate match, the output has the high binary value 1. On the other hand, if the two inputs to the XNOR gate do not match, the output has the low binary value 0.

[0045] The output values from each XNOR gate are input to an AND gate. If both inputs are 1, indicating that each of the first and second bits of the reference base value and the substring base value match, the cell comparison output value from the AND gate is the high binary value 1. Otherwise, the cell comparison output value from the AND gate is the low binary value 0. This high or low binary value is output from the cell to a circuit, such as to Figure 1 circuit 106 in, to identify a column or group of cells in which all reference base values match all substring base values stored in the group of cells.

[0046] Figure 4C is an example of a circuit for comparing cell output values in a group of cells according to one or more embodiments. As Figure 4C shown, the cell comparison output values from each cell in the group of cells are input to an AND gate to produce a column output value for column i. If the cell comparison output values of all cells 1 to M in column i all have the high binary value 1 indicating a match, the column output value from the AND gate for that column is the high binary value 1. This column output value can be used to identify the column or group of cells as having a matching substring sequence and reference sequence. Figure 4C The circuit shown can be part of a circuit external to the array 104 or can be part of the array 104.

[0047] In some cases, there may be multiple groups of cells identified as storing reference sequences that match the current substring sequence. In such cases, circuit 106 can use only the first matching position, the first matching position and other matching positions, or can use all matching positions to locate the current substring sequence within the reference genome. In other cases, the current substring sequence may result in a mismatch. For example, a mutation or a read error in the sample read from which the substring sequence is taken may prevent a match or may result in an error in the match.

[0048] Other embodiments may use different circuits or different processes for identifying groups of cells in which the stored reference sequences match the substring sequences stored in the group of cells. For example, an inner product or dot product operation may alternatively be used to identify the group of cells storing the matching sequences, rather than logic gates, as discussed in more detail below with reference to Figure 7 the match identification sub-process. Another example is that Figure 4C the NAND gates in can be replaced by a circuit for summing the cell comparison output values of the group of cells and then comparing the sum with the number of cells in the group of cells. In such examples, if the sum of the cells from the group equals the number of cells in the group, the reference sequence of the group of cells matches the substring sequence.

[0049] Exemplary Recognition Process

[0050] Figure 5It is a flowchart of a sample read segment positioning process according to one or more embodiments. Figure 5 The process can be performed, for example, by Figure 1 device 102 and / or host 101 in

[0051] In block 502, the reference sequence is stored in a corresponding unit group among multiple unit groups for reference bases from the reference genome. As referred to above with reference to Figure 2 described, the storage location of the reference sequence corresponds to the order of the units in the unit group. In addition, each reference sequence represents a part of the reference genome, and this part of the reference genome partially overlaps or is shifted with respect to at least one other part of the reference genome represented by one or more other reference sequences stored in one or more other unit groups.

[0052] In some specific embodiments, the reference guiding device 102 can receive the reference sequence or the reference genome from the host 101. In other specific embodiments, the reference guiding device 102 can be preconfigured by the manufacturer, where the reference sequence is programmed or stored in a unit group of a specific genome (such as the human genome H38).

[0053] In block 504, the current substring sequence is stored in each unit group among multiple unit groups for sample bases from the sample read segment. The storage location of the current substring sequence within each unit group corresponds to the order of the unit group. The current substring sequence can be received from the host 101, or can be selected by the device 102 from the sample read segment provided by the host 101. In some specific embodiments, the circuit 106 of the device 102 or the host 101 can randomly select a substring sequence from the sample read segment. In other specific embodiments, the circuit 106 or the host 101 can select substring sequences spaced apart in the middle of the entire sample read segment.

[0054] In block 506, the circuit 106 identifies the unit group among multiple unit groups in which the stored reference sequence matches the current substring sequence stored in the unit group. In some specific embodiments, logic gates can be used to perform the identification of the unit group, as in the example discussed above for Figures 4A to 4C discussed. In other specific embodiments, the identification of the unit group can be performed by performing calculations using the stored reference values and sample values, as in the Figure 7 exemplary matching identification sub - process discussed below.

[0055] In block 508, circuit 106 or host 101 determines whether the subsequence sequence stored in block 504 is the last subsequence sequence from the sample read segments to be stored in the cell group. In some embodiments, a predetermined number of subsequence sequences may be iteratively stored in the cells of device 102 for comparison with a reference sequence from a reference genome. The number of different subsequence sequences obtained from the sample read segments may depend on, for example, the length of the subsequence sequence (e.g., 20 bases in Figure 2 ), the length of the reference genome, the length of the sample read segments (e.g., short read segments of 250 or 300 bases from an Illumina device versus long read segments of 5,000 bases from a nanopore device), the accuracy of the method used to create the sample read segments, and the desired accuracy for locating the sample read segments within the reference genome. In one example, short read segments of 250 or 300 bases may be located in a reference genome with only a few matching subsequence sequences. Such examples may use only ten subsequence sequences from the sample read segments to generate sufficient matches to locate the sample read segments in the reference genome. Figure 2 In the case of 20 bases in Figure 2

[0056] If it is determined in block 508 that the current subsequence sequence is not the last subsequence sequence from the sample read segments, the process proceeds to block 510 to rewrite the current subsequence sequence with the next subsequence sequence from the sample read segments for storage in the plurality of cell groups. Then, Figure 5 the process returns to block 506 to identify the cell groups in which the reference sequence matches the next subsequence sequence. It is noted that since the same reference sequence may be reused for the next subsequence sequence, block 502 is not repeated. For multiple iterations of the subsequence sequences from the sample read segments, the reference sequence or reference genome only needs to be loaded or stored once to improve the efficiency of the sample read segment position identification process.

[0057] In some embodiments, circuit 106 or host 101 may determine in block 508 whether another subsequence sequence is needed to locate the sample read segments based on the number of previously tested subsequence sequences. For example, if the first four subsequence sequences have resulted in a match, then the sixth subsequence sequence does not need to be tested. On the other hand, if the first four subsequence sequences have not resulted in any matches, the fifth subsequence sequence may be loaded.

[0058] If it is determined in block 508 that the current substring sequence is the last substring sequence from the sample read, the process proceeds to block 512 to determine the probabilistic location of the sample read within the reference genome based on the identified groups of units from the different substring sequences of the sample read. As described above for block 506, the first matching group of units can be used as the location for each substring sequence, or alternatively, if some substring sequences result in multiple matching groups of units, the multiple matching groups of units can be used as the possible locations for the substring sequences. In other cases, due to errors in the read sample or mutations in the sample, the substring sequence may not have a matching location. The location of the sample read determined by circuit 106 or host 101 in block 512 can be probabilistic in the sense that multiple possible locations can be identified for different substring sequences from the sample read, and the consensus or statistics derived from the matching locations can be used to probabilistically position the sample read within the reference genome.

[0059] In one example, the average of all the locations of all the matching groups of units of all the substring sequences is used to identify the most likely location of the sample read within the reference genome. In another example, only one location of each substring sequence having a matching group of units is used in the average. In yet another example, the probabilistic location of the sample read can be determined by identifying the most distant interval location within the reference genome corresponding to the matching group of units of the substring sequence. In other examples, one or more outlier locations relative to a set of matching locations can be discarded when determining the probabilistic location of the sample read within the reference genome.

[0060] Figure 6 is a flowchart of a matching recognition sub - process using logical operations according to one or more embodiments. Figure 6 The sub - process can be performed by the units in array 104 and / or circuit 106 as part of block 506 in the sample read location process discussed above to identify the groups of units in which the stored reference sequence matches the current substring sequence stored in the group of units. Figure 5 In block 602, at least one XNOR operation is performed in each unit of the multiple groups of units to compare the sample base from the current substring sequence with the reference base from the reference sequence. As discussed above with reference to

[0061] Two XNOR gates and one AND gate can be used in the unit to compare the values of the reference base and the sample base stored in the unit. Figure 4A As discussed above, two XNOR gates and one AND gate can be used in the unit to compare the values of the reference base and the sample base stored in the unit.

[0062] In block 604, a comparison value is output from each unit of the multiple groups of units, which indicates whether the sample base of the unit matches the reference base of the unit. The comparison value can be the high binary value 1 or the low binary value 0 indicating whether the reference value and the sample value stored in the unit match.

[0063] In block 608, circuit 106 identifies the cell group in which the reference sequence stored therein matches the current substring sequence by performing AND operation to the comparison value output from the cell in the corresponding cell group. If any one of the comparison values ​​has a low binary value of 0, the result of the AND operation will have a low binary value of 0, thereby indicating that the cell group does not store a matching sequence. On the other hand, if all comparison values ​​have a high binary value of 1, the result of the AND operation will have a high binary value of 1, thereby indicating that the cell group stores a matching sequence. In other specific implementations, circuit 106 can identify the cell group in which the reference sequence stored therein matches the current substring sequence by summing the comparison value and comparing the sum with the predetermined number of cells in the cell group. In such specific implementations, if all comparison values ​​from the cell have a value of 1, then when all cells have a matching value, the comparison value sum of the cell group will equal the total number of cells in the cell group. Although XNOR and AND are mentioned as examples, those of ordinary skill in the art will recognize that in other embodiments, the same result can be achieved by other logical combinations.

[0064] As described above, other processes may be used to identify groups of cells in which a stored reference sequence matches a substring sequence stored in the group of cells. In this regard, Figure 7 is a flow chart of a match identification subprocess using an inner product or a dot product of a reference vector and a substring vector according to one or more embodiments. Figure 7 The sub-process of may be performed by units in array 104 and / or circuit 106, as discussed above. Figure 5 The sample read segment positioning process is performed as part of block 506 to identify a unit group in which the stored reference sequence matches the current substring sequence stored in the unit group.

[0065] In box 702, the product of a first stored value representing a substring base and a second stored value representing a reference base is calculated for each cell. The substring values ​​stored in the cell group may represent a substring vector, and the reference values ​​stored in the cell group may represent a reference vector for the cell group. For example, each reference value and each sample value may be represented by two numbers including 1 and / or -1. In such examples, base C may have a value of 1,1, base G may have a value of -1,-1, base T may have a value of 1,-1, and base A may have a value of -1,1. As will be understood by those of ordinary skill in the art with reference to this disclosure, different combinations of 1 and -1 may be used to represent bases.

[0066] In block 704, the product calculated for each cell in the group of cells is output from each cell to circuit 106. In other implementations, circuit 106 may calculate the product of the values ​​stored in the cells.

[0067] For each group of cells, the products output from the cells are summed in block 706. The sum of the products for each group of cells is then compared in block 708 to twice the number of cells in the group of cells or twice the substring sequence length. In other specific implementations, the sum of the products for each group of cells may be compared to a different predetermined multiple of the number of cells in the group. For example, in a specific implementation where the cell outputs indicate a matching value of 1 and a non - matching value of 0, the sum is compared to 1 times the total number of cells, rather than twice the number of cells in the group. Similarly, in a specific implementation where the cell outputs indicate a matching value of 0, the sum is compared to 0 times the number of cells.

[0068] In block 710, circuit 106 or host 101 identifies the groups of cells in which the sum of the products is equal to twice the number of cells in the group of cells or twice the substring sequence length. Such groups of cells have matching sequences because each product of the cells from such groups equals 1 and thus sums to twice the number of cells (or twice the substring sequence length).

[0069] For example, using only four bases as the substring sequence length, for illustrative purposes where the substring sequence length is shorter than the range of 17 to 25 bases discussed above, the reference sequence for a group of cells can be represented as R = CCAG, the matching substring sequence can be represented as S1 = CCAG, and the non - matching substring sequence can be represented as S2 = GGAG. Then, using the values assigned to the bases discussed above for block 702, the encoded reference sequence or reference vector is [1,1,1,1, - 1,1, - 1, - 1]. The encoded matching substring sequence or matching substring sequence vector will also be [1,1,1,1, - 1,1, - 1, - 1]. The encoded non - matching substring sequence or non - matching substring sequence vector will be [-1, - 1, - 1, - 1, - 1,1, - 1, - 1].

[0070] Taking the dot product or inner product of the reference vector and the matching substring sequence vector gives 8, which is twice the number of cells in the group of cells or twice the substring sequence length of 4 bases. On the other hand, taking the dot product or inner product of the reference vector and the non - matching substring sequence vector gives 0, which is less than twice the number of cells in the group or the substring sequence length. Thus, an inner product or dot product that produces a value less than twice the number of cells in the group or twice the substring sequence length does not correspond to a group of cells storing a matching sequence.

[0071] As described above, the foregoing reference guiding device and method generally allow sample reads to be probabilistically positioned within a reference genome. This can improve the efficiency of de novo and reference alignment sequencing by preprocessing sample reads into groups based on their positions in the reference genome for further sequencing. In the case of de novo sequencing, this can improve the scalability and efficiency of de novo sequencing by allowing more computational threads to access each positioned group of sample reads in a smaller shared memory compared to conventional methods where a larger and more expensive memory is used to access all sample reads by a limited number of computational threads. In the case of reference alignment sequencing, compared to conventional reference alignment sequencing that may use a single shared memory to store the entire reference genome, the positioned groups of sample reads can allow only a smaller and more relevant portion of the reference genome to be stored in the smaller and less expensive memory of each positioned group, while allowing more computational threads to access the multiple smaller memories to improve scalability.

[0072] Other Embodiments

[0073] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks, modules, and processes described in connection with the examples disclosed herein can be implemented as electronic hardware, software, or combinations of both. Additionally, the foregoing processes can be embodied on a computer-readable medium that causes a processor, controller, or other circuitry to perform or implement certain functions.

[0074] To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, and modules have been described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those of ordinary skill in the art can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0075] The various illustrative logical blocks, units, modules, and circuits described in connection with the examples disclosed herein can be implemented or performed with a general-purpose processor, GPU, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor or controller circuitry can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, an SoC, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0076] The activities of the methods or processes described in connection with the examples disclosed herein can be embodied directly in hardware, in software modules executed by a processor or controller circuit, or in a combination of both. The steps of the method or algorithm can also be executed in an order alternative to the order provided in the examples. The software modules can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable medium, an optical medium, or any other form of storage medium known in the art. The exemplary storage medium is coupled to the processor or controller circuit such that the processor or controller circuit can read information from, and write information to, the storage medium. In an alternative, the storage medium can be part of the processor or controller circuit. The processor or controller circuit and the storage medium can reside in an ASIC or an SoC.

[0077] The foregoing description of the exemplary embodiments of the present disclosure has been provided to enable any person of ordinary skill in the art to make or use the embodiments of the present disclosure. Various modifications to these examples will be readily apparent to those of ordinary skill in the art, and the principles disclosed herein can be applied to other examples without departing from the scope of the present disclosure. The embodiments are to be considered in all respects only as exemplary and not restrictive. Additionally, the language used in the following claims in the form of "at least one of A and B" should be understood to mean "only A, only B, or both A and B".

Claims

1. A device, the device comprises: at least one two-dimensional systolic array including a plurality of cell groups, wherein each cell group of the plurality of cell groups is configured to: store a reference sequence representing a reference base from a reference genome, the reference sequence corresponding to the order of cells in the corresponding cell group; and store the current substring sequence representing sample bases from a sample read by passing the current substring sequence from one cell group in the at least one two-dimensional systolic array to the next cell group, wherein the current substring sequence corresponds to the order of the cells in the corresponding cell group; wherein each cell group of the plurality of cell groups is further configured to store the same current substring sequence and a reference sequence representing a part of the reference genome, the part of the reference genome partially overlapping with at least one other part of the reference genome represented by one or more other reference sequences stored in one or more other cell groups in the at least one two-dimensional systolic array; and a circuit configured to identify a cell group in the at least one two-dimensional systolic array in which the stored reference sequence matches the current substring sequence stored in the cell group.

2. The device according to claim 1, wherein at least one of the circuit and each cell group is further configured to perform one or more logical operations to determine whether the stored reference sequence matches the current substring sequence stored in the cell group.

3. The device according to claim 1, wherein each cell of the plurality of cell groups is further configured to: perform at least one XNOR operation to compare a first value stored in the cell representing a sample base from the current substring sequence with a second value stored in the cell representing a reference base from the reference sequence stored in the corresponding cell group; and output the comparison value of the at least one XNOR operation to the circuit, the comparison value indicating whether the sample base of the cell matches the reference base of the cell.

4. The device according to claim 3, wherein the circuit is further configured to identify a cell group in which the stored reference sequence matches the current substring sequence stored in the cell group by performing an AND operation on the comparison values output from the cells of the corresponding cell group.

5. The device according to claim 1, wherein each cell of the plurality of cell groups is further configured to: calculate the product of the values stored in the cell representing the sample base and the reference base; and output the product to the circuit; and wherein the circuit is further configured to identify a cell group in which the stored reference sequence matches the current substring sequence stored in the cell group at least partially based on the product output by the cell.

6. The device according to claim 5, wherein the circuit is further configured to, for each cell group of the plurality of cell groups: sum the products output by the cells in the cell group; Compare the sum to a predetermined multiple of the number of units in the unit group; and In response to the sum being equal to the predetermined multiple of the number of units in the unit group, identify the unit group as one in which the reference sequence stored therein matches the current substring sequence stored in the unit group.

7. The apparatus according to claim 1, wherein each of the plurality of unit groups consists of a predetermined number of units, the predetermined number of units being in the range of 17 to 25 units.

8. The apparatus according to claim 1, wherein each of the plurality of unit groups is further configured to: Rewrite the current substring sequence with a subsequent substring sequence representing a sample base from the sample read to store the subsequent substring sequence in the unit group; and Retain the corresponding reference sequence stored in the unit group; and wherein the circuit is further configured to identify a unit group among the plurality of unit groups in which the retained reference sequence stored in the unit group matches the subsequent substring sequence stored in the unit group.

9. The apparatus according to claim 1, wherein the circuit is further configured to determine a probabilistic position of the sample read within the reference genome based on an iteration of the following steps: Store different substring sequences of the sample read in the plurality of unit groups; and Identify a unit group among the plurality of unit groups in which the reference sequence stored therein matches the substring sequence stored in the unit group.

10. The apparatus according to claim 1, wherein the circuit includes at least one of a field programmable gate array (FPGA) and an application specific integrated circuit (ASIC).

11. The apparatus according to claim 1, wherein the units in the plurality of unit groups include at least one of registers, latches, and flip-flops.

12. A method of positioning a sample read relative to a reference genome, the method comprising: Store in a plurality of unit groups of at least one two-dimensional systolic array a reference sequence representing a reference base from the reference genome, the reference sequence corresponding to the order of units in the corresponding unit group of the plurality of unit groups, wherein each of the plurality of unit groups stores a reference sequence representing a portion of the reference genome, the portion of the reference genome overlapping at least in part with at least one other portion of the reference genome represented by one or more other reference sequences stored in one or more other unit groups; Store in each of the plurality of unit groups a current substring sequence of sample bases from the sample read by passing a first substring sequence from one unit group in the at least one two-dimensional systolic array to the next unit group, wherein the current substring sequence corresponds to the order of the units in the corresponding unit group; and Identify a unit group in the at least one two-dimensional systolic array in which the reference sequence stored therein matches the current substring sequence stored in the unit group.

13. The method according to claim 12, further comprising performing one or more logical operations on each cell group to determine whether the stored reference sequence matches the current substring sequence stored in the cell group.

14. The method according to claim 12, further comprising, for each cell of the plurality of cell groups: performing at least one XNOR operation on a first value stored in the cell for a sample base from the current substring sequence and a second value stored in the cell for a reference base from the reference sequence stored in the corresponding cell group; and outputting a comparison value of the at least one XNOR operation, the comparison value indicating whether the sample base of the cell matches the reference base of the cell.

15. The method according to claim 14, further comprising identifying a cell group in which the stored reference sequence matches the current substring sequence stored in the cell group by performing an AND operation on the comparison values output from the cells of the corresponding cell group.

16. The method according to claim 12, further comprising identifying a cell group in which the stored reference sequence matches the current substring sequence stored in the cell group by calculating, for at least each cell group of the plurality of cell groups, an inner product of a reference vector representing the reference bases stored in the cell group and a substring vector representing the sample bases stored in the cell group.

17. The method according to claim 12, wherein the current substring sequence is in the range of 17 to 25 bases.

18. The method according to claim 12, further comprising, for each cell group of the plurality of cells: rewriting the current substring sequence with a subsequent substring sequence of the sample bases from the sample read to store the subsequent substring sequence; retaining a corresponding portion of the reference genome as the reference sequence stored in the cell group; and and identifying a cell group of the plurality of cell groups in which the retained reference sequence stored in the cell group matches the subsequent substring sequence stored in the cell group.

19. The method according to claim 12, further comprising determining a probabilistic position of the sample read within the reference genome based on an iteration of the following steps: storing different substring sequences from the sample read in the cell groups of the at least one two-dimensional systolic array; and identifying a cell group of the at least one two-dimensional systolic array in which the stored reference sequence matches the substring sequence stored in the cell group.

20. A method of operating a device including at least one two-dimensional systolic array, the at least one two-dimensional systolic array including a plurality of cell groups, the method comprising: storing the first substring sequence of the sample bases from the sample read in each cell group of the plurality of cell groups by passing the first substring sequence from one cell group in the at least one two-dimensional systolic array to the next cell group such that the sample bases in the first substring sequence correspond to the order of the cells in the corresponding cell group of the plurality of cell groups; Each of the plurality of cell groups is configured to store a reference sequence representing a different part of a reference genome; Identify a cell group in the at least one two-dimensional systolic array in which the stored reference sequence matches the first substring sequence stored in the cell group; Store the second substring sequence of sample bases from another part of the sample read in each of the plurality of cell groups by passing the second substring sequence from one cell group in the at least one two-dimensional systolic array to the next cell group, such that the sample bases of the second substring sequence correspond to the order of the cells in the cell group, and rewrite the first substring sequence; And Identify a cell group in the at least one two-dimensional systolic array in which the stored reference sequence matches the second substring sequence stored in the cell group.

21. The method according to claim 20, further comprising, for each of the first substring sequence and the second substring sequence, performing one or more logical operations for each cell group to determine whether the stored reference sequence matches the substring sequence stored in the cell group.

22. The method according to claim 20, further comprising performing at least one logical XNOR operation for each cell of the plurality of cell groups to compare the sample bases from the substring sequence stored in the cell with the reference bases from the reference sequence stored in the cell; and Output a value indicating whether the sample bases stored in the cell match the reference bases stored in the cell from each cell of the plurality of cell groups.

23. The method according to claim 22, further comprising identifying a cell group in which the stored reference sequence matches the substring sequence stored in the cell group by performing a logical AND operation on the values output from the cells of the corresponding cell group.

24. The method according to claim 20, further comprising identifying a cell group in which the stored reference sequence matches the substring sequence stored in the cell group by calculating at least for each cell group of the plurality of cell groups the inner product of a reference vector representing the reference bases stored in the cell group and a substring vector representing the sample bases stored in the cell group.

25. The method according to claim 20, further comprising determining a probabilistic position of the sample read within the reference genome based on the identification of at least one of: One or more cell groups in which the stored reference sequence matches the first substring sequence stored in the one or more cell groups; and One or more cell groups in which the stored reference sequence matches the second substring sequence stored in the one or more cell groups.

Citation Information

Patent Citations

  • Reference-guided genome sequencing

    US20210292830A1

  • Reference-guided genome sequencing

    US20210295946A1

  • Biological sequence information processing method and device

    JP2005251192A