Devices and Methods for Genome Sequencing

By using nonvolatile memory cell arrays and programming solutions, the shortcomings in genome sequencing equipment in terms of processing efficiency, physical size and energy consumption are solved, and efficient precise matching and close matching stage switching is achieved, which improves the overall performance of genome sequencing.

CN114787930BActive Publication Date: 2025-07-08WESTERN DIGITAL TECHNOLOGIES INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180006731.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-22
Filing Date
2021-01-25
Publication Date
2025-07-08
Estimated Expiration
2041-01-25

AI Technical Summary

Technical Problem

Existing genome sequencing devices have shortcomings in processing efficiency, physical size and energy consumption, especially when performing the exact and approximate matching phases, the volatile nature of static random access memory leads to large power consumption and large area occupancy.

Method used

Using a non-volatile memory (NVM) cell array, combined with the arrangement of content addressable memory (CAM) and ternary CAM (TCAM), efficient switching of the precise matching and approximate matching stages is achieved through programming and coding schemes, reducing energy consumption and improving processing efficiency.

Benefits of technology

Improves the processing efficiency of genome sequencing, reduces the physical size of the device, and reduces energy consumption, while supporting fast switching of precise matching and approximate matching stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787930B_ABST
    Figure CN114787930B_ABST
Patent Text Reader

Abstract

The present invention discloses a device that includes an array of non-volatile memory (NVM) cells. A reference sequence representing a portion of a genome is stored in corresponding groups of NVM cells. An exact-match phase substring sequence representing a portion of at least one sample read is loaded into a group of NVM cells. The array is used as a content-addressable memory (CAM) to identify one or more groups of NVM cells in which the stored reference sequence matches the loaded exact-match phase substring sequence. An approximate-match phase substring sequence is loaded into a group of NVM cells. The array is used as a ternary CAM (TCAM) to identify one or more groups of NVM cells in which the stored reference sequence approximately matches the loaded approximate-match phase substring sequence. When the array is used as a TCAM, at least one of the reference sequence and the approximate-match phase substring sequence for each group of NVM cells includes at least one wildcard value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This patent application claims priority to U.S. Patent Application No. 16 / 908,581, filed on Jun. 22, 2020, and entitled "Devices and Methods for Genome Sequencing" (Attorney Docket No. WDA - 4949 - US), the entire content of which is hereby incorporated by reference. This patent application is related to co - pending U.S. Patent Application No. 16 / 821,849, filed on Mar. 17, 2020, and entitled "Reference - Guided Genome Sequencing" (Attorney Docket No. WDA - 4724 - US), the entire content of which is hereby incorporated by reference. This patent application is also related to co - pending U.S. Patent Application No. 16 / 822,010, filed on Mar. 18, 2020, and entitled "Reference - Guided Genome Sequencing" (Attorney Docket No. WDA - 4725 - US), the entire content of which is hereby incorporated by reference. This patent application is also related to co - pending U.S. Patent Application No. 16 / 820,711, filed on Mar. 17, 2020, and entitled "Devices and Methods for Locating a Sample Read in a Reference Genome" (Attorney Docket No. WDA - 4726 - US), the entire content of which is hereby incorporated by reference. Background of the Invention

[0003] Current limitations in DNA (deoxyribonucleic acid) and RNA (ribonucleic acid) sample processing result in sample reads or portions of the sample genome having generally unknown positions within the sample genome. The sample reads must then be sequenced or placed back into their positions within the sample genome. The two main types of genome sequencing are de novo sequencing and reference - alignment sequencing. Reference - alignment sequencing uses a reference genome to locate sample reads within the sample genome. On the other hand, de novo sequencing generally does not use a reference genome, but rather compares sample reads to each other to locate the sample reads within the sample genome.

[0004] Both de novo sequencing and reference alignment sequencing typically include an exact matching phase and an approximate matching phase. Exact matching can be performed by, for example, a seed and expand algorithm to find sample reads that exactly match within a reference genome to localize the sample reads within the reference genome. However, due to sample read errors and mutations, it is often not possible to exactly match or sequence all sample reads at their positions in the sample genome. Approximate matching can use algorithms such as the Smith-Waterman algorithm or an automaton-based algorithm to find the closest match or the best fit alignment in the sample genome.

[0005] Sequencing a sample genome typically requires a long processing time, potentially taking hundreds to thousands of processor hours. Current devices for genome sequencing can include different arrangements of static random access memory (SRAM) content addressable memory (CAM) or ternary CAM (TCAM) for performing different phases of exact and approximate matching. However, such devices consume a large amount of power due to their volatile nature and also consume a large area, such as by including six transistors per SRAM cell. The computational efficiency of genome sequencing systems still needs to be improved by orders of magnitude. Therefore, there is a need to improve devices for genome sequencing in terms of processing efficiency, physical size, and energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The features and advantages of embodiments of the present disclosure will become more apparent from the following detailed description and in conjunction with the accompanying drawings. The accompanying drawings and the associated description are provided to illustrate embodiments of the present disclosure and not to limit the scope of the claimed subject matter.

[0007] Figure 1 is a block diagram of a device for genome sequencing including a plurality of arrays of non-volatile memory (NVM) cells according to one or more embodiments.

[0008] Figure 2 is according to one or more embodiments Figure 1 of an array of NVM cells in a device.

[0009] Figure 3 is according to one or more embodiments Figure 1 of an exemplary circuit diagram of an NVM cell in a device.

[0010] Figure 4A is an exemplary truth table of an NVM cell for use when operating as a ternary content addressable memory (TCAM) according to one or more embodiments Figure 3 for.

[0011] Figure 4Bis for when operating as a content addressable memory (CAM) according to one or more embodiments Figure 3 exemplary truth table for NVM cells

[0012] Figure 5 is an example of a device for simultaneously processing substring sequences according to one or more embodiments Figure 1 of a device

[0013] Figure 6 is an example of a device for processing an exact match phase substring sequence and an approximate match phase substring sequence according to one or more embodiments Figure 1 of a device

[0014] Figure 7 illustrates an example of a pipelined substring sequence according to one or more embodiments

[0015] Figure 8 is a flowchart of a genomic sequencing process according to one or more embodiments

[0016] Figure 9 is a flowchart of an operation mode switching process for at least one array of NVM cells according to one or more embodiments

[0017] Figure 10 is a flowchart of a concurrent matching process according to one or more embodiments

[0018] Figure 11 is a flowchart of a sample read classification process according to one or more embodiments DETAILED DESCRIPTION

[0019] Numerous specific details are set forth in the following detailed description in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that various embodiments may be practiced without some of these specific details. In other instances, well-known structures and techniques have not been shown in detail to avoid unnecessarily obscuring the various embodiments.

[0020] System Example

[0021] Figure 1FIG. 0 is a block diagram of a device 100 for genomic sequencing including a plurality of arrays of non-volatile memory (NVM) cells according to one or more embodiments. A host 103 communicates with the device 100 via a control circuit 104 to sequence sample reads of a sample genome. In some embodiments, the device 100 may provide an index 10 to the host 103, which is stored in the memory 106 of the device 100 and indicates the location of the sample reads. In other embodiments, the device 100 may provide an indication of the location of another data structure or sample reads input to the device 100 to the host 103.

[0022] A sample read or a sample substring sequence taken from a sample read may initially be provided to the device 100 by the host 103 and / or Figure 1 another device not shown (such as an additional host) to determine the location of the sample read for sequencing or aligning the sample read in the sample genome from which the sample read is taken. In some embodiments, a read device such as an Illumina device (Illumina, Inc. of San Diego, California) or a nanopore device that generates the sample read may provide the sample read or the sample substring sequence to the device 100.

[0023] The substring sequence to be loaded into the array 101 is compared with a reference sequence stored in the array 101. The stored reference sequence represents a portion of the genome. In the case of sequencing for reference alignment, a reference sequence representing bases from a reference genome (e.g., the human reference genome H38) or a portion thereof may be stored in the NVM cells of one or more arrays 101 for comparison with the loaded substring sequence. In the case of de novo sequencing, the sample reads may be compared with each other to identify overlapping regions for sequencing or aligning the sample reads in the sample genome. In such cases, the NVM cells of one or more arrays 101 may store a reference sequence representing a portion of the sample genome for comparison with the loaded substring sequence.

[0024] For ease of description, the exemplary embodiments in the present disclosure will be described in the context of DNA sequencing. However, the embodiments of the present disclosure are not limited to DNA sequencing and may generally be applied to any nucleic acid-based sequencing including RNA (ribonucleic acid) sequencing.

[0025] The host 103 may include a computer, such as a desktop computer or a server, that can implement genomic sequencing algorithms using the device 100, such as seed and extend algorithms for exact matching of sample reads in a genome and / or computationally more complex algorithms such as the Burrow-Wheeler algorithm, automaton-based algorithms, or Smith-Waterman algorithms for approximate matching. Examples of automaton-based approximate matching algorithms may include, for example, the Levenshtein automaton or the string-independent local Levenshtein automaton algorithm.

[0026] The host 103 and the device 100 may or may not be physically co-located. For example, in some embodiments, the host 103 and the device 100 may communicate via a network, such as a local area network (LAN) or a wide area network (WAN) using the Internet, or a data bus or fabric. Additionally, one of ordinary skill in the art will understand that other embodiments may include multiple hosts 103 and / or multiple devices 100 for sequencing sample reads. In certain embodiments, the host 103 and the device 100 (or multiple hosts and devices) are integrated into a single device or system.

[0027] In Figure 1 example, each of the arrays 101 A 、101 B 、101 C and 101 D in the device 100 is configured to operate as a content-addressable memory (CAM) for exact matching and as a ternary CAM (TCAM) for approximate matching, depending on the encoding of the reference sequence stored in the array and the substring sequence loaded into the array. As discussed in more detail below with reference to Figure 3 、 Figure 4A and Figure 4B , the array 101 includes non-volatile memory (NVM) cells configured to allow the use of wildcard values (which may also be referred to as "don't care values" or "X values") when the array is used as a TCAM.

[0028] Compared with conventional static RAM (SRAM) TCAM and SRAM CAM, the arrangements of NVM TCAM and CAM disclosed herein include non-volatile solid-state memories that use less energy and consume less physical space. Although the description herein generally refers to solid-state memories, it should be understood that solid-state memories can include one or more memory devices of various types of memory devices, such as flash integrated circuits, chalcogenide RAM (C-RAM), phase change memory (PC-RAM or PRAM), programmable metallization cell RAM (PMC-RAM or PMCm), ovonic unified memory (OUM), resistive RAM (RRAM), NAND memory (e.g., single-level cell (SLC) memory, multi-level cell (MLC) memory (i.e., two or more layers), or any combination thereof), NOR memory, EEPROM, ferroelectric memory (FeRAM), magnetoresistive RAM (MRAM), other discrete non-volatile memory (NVM) chips, or any combination thereof. In some specific embodiments, the cells in the array 101 can include circuit elements, such as non-volatile registers, latches, or flip-flops.

[0029] As used herein, a cell generally refers to a memory location for storing a reference value or a portion of a reference value for representing a nucleotide (referred to as a base in the present disclosure). In Figure 1 the example of, the array 101 includes cells that also include logic circuitry for comparing a loaded sample value with a stored reference value. The cells are arranged in groups into the array 101 to store a reference sequence representing a portion of a genome. In this regard, the reference values of the reference sequence are stored corresponding to the order of the cells in the group. Similarly, the sample values of the substring sequences are loaded into the cells of the group corresponding to the order of the cells of the group. Comparing the stored reference values of each group with the loaded sample values provides an indication of whether the stored reference values of the group match (in the case of CAM operation) or approximately match (in the case of TCAM operation) the loaded substring sequence. In some specific embodiments, the array 101 can include one or more systolic arrays, where the reference values and / or sample values can be passed to the next cell of another group of cells of the array 101.

[0030] The memory 106 of the device 100 can include, for example, volatile memory, such as DRAM or SRAM, for storing the index 10 and the scoring matrix 20. In other specific embodiments, the memory 106 can include NVM, such as MRAM. As Figure 1As shown, the memory 106 stores an index 10 that can be used by the host 103 to indicate the positions of the sample read segments within a genome such as a reference genome or a sample genome where there are matches and / or approximate matches. In some embodiments, the index 10 can include a data structure, such as a bitmap or other data structure, indicating positions in the index or the genome that correspond to groups of units identified as storing sequences with matches and / or approximate matches.

[0031] The control circuit 104 can update the index 10 for different sample substring sequences loaded into groups of units of the array 101. In some embodiments, the circuit 104 can indicate the average position in the index 10 for a group of units having multiple matches or approximate matches for a substring sequence. In other embodiments, only the group of units corresponding to the first match or approximate match for a particular substring sequence can be used, or the control circuit 104 may not update the index 10 at all for a substring sequence of a group of units having more than one stored matching sequence.

[0032] In addition, some embodiments may not use an index or other data structure to indicate the positions of groups of units having sequences with matches or approximate matches. For example, in some embodiments, the control circuit 104 can directly output an indication of the positions of the sequences with matches or approximate matches to the host 103.

[0033] The scoring matrix 20 can be used as part of an approximate matching algorithm such as the Smith-Waterman algorithm. For example, the scoring matrix 20 can be updated by the control circuit 104 to score the matches and mismatches between a reference sequence stored in the array 101 and a substring sequence loaded into the array 101 to determine the best alignment for the substring sequence. As will be understood by those of ordinary skill in the art with reference to this disclosure, the scoring matrix 20 can include various forms of weighted data structures indicating, for example, the relative desirability of insertions, deletions, and mutations.

[0034] The control circuit 104 can include, for example, hardwired logic circuitry, analog circuitry, and / or combinations thereof. In other embodiments, the control circuit 104 can include one or more ASICs, microcontrollers, digital signal processors (DSPs), FPGAs, and / or combinations thereof. In some embodiments, the control circuit 104 can include one or more system-on-chips (SoCs) that can be combined with the memory 106. In this regard, in some embodiments, one or more arrays 101 and interconnects 107 can be integrated with the control circuit 104 and the memory 106 as a single component or chip.

[0035] In Figure 1In the example of FIG. 1 , control circuit 104 communicates with array 101 and memory 106 using interconnect 107. In some implementations, array 101 and memory 106 can communicate directly via interconnect 107 without involving control circuit 104. Interconnect 107 can include, for example, a bus for communicating between array 101, control circuit 104, and memory 106.

[0036] As described above, the device 100 may include control circuitry 104 and / or other circuitry (e.g., Figure 2 The sensing and latching circuit 114 A and encoder 116 A ) is configured to identify, among a plurality of groups of cells in array 101, a group of cells in which a stored reference sequence matches or approximately matches a loaded substring sequence. Figure 2 and Figure 3 As discussed in more detail, in some implementations, identification of a group of cells having a matching or nearly matching sequence can be made by the state or charge of a match line shared by the cells in the group of cells.

[0037] As will be understood by those skilled in the art with reference to this disclosure, other specific implementations may include Figure 1 For example, other implementations may combine host 103 and device 100, or may include different numbers of arrays 101, memories 106, devices 100, and / or hosts 103. In this regard, device 100 may include more than one device 100. Figure 1 More arrays 101 are shown to allow for more concurrent or parallel comparisons to improve processing efficiency.

[0038] Figure 2 is an array 101 depicting a device 100 according to one or more embodiments. A An exemplary block diagram of the NVM unit 102 in FIG. Figure 2 As shown, array 101 A The NVM cells 102 are arranged in rows and columns in a cross-point array structure. Figures 3 to 4B As discussed in more detail, array 101 AIt can be programmed to operate as a TCAM or a CAM. Each NVM cell 102 or pair of NVM cells is configured to store a reference value for comparison with a sample value loaded into the cell or cell pair. Each base of the reference sequence and each base of the substring sequence can be represented by two bits because, for example, in the case of DNA sequencing, there are four possible bases - adenine (A), guanine (G), cytosine (C), and thymine (T). In some specific embodiments, two adjacent cells 102 in the same row can store a reference value representing a reference base, which is compared with a sample value representing a sample base loaded into the two adjacent cells. In other specific embodiments, each cell can store a reference value representing a reference base, which is compared with a sample value representing a sample base loaded into the cell. As discussed in more detail below with reference to Figures 3 to 4B when operating as a TCAM, the array 101 A can use two cells to represent a reference value, but when operating as a CAM, one or two cells can be used to represent a sample value, depending on the use of wildcard values for the reference value and / or the sample value.

[0039] The control circuit 104 controls the word line (WL) driver 108 A and also controls the programming line and search line driver 110 A to program the cells 102 in a specific row to store reference values for the reference sequence. In such specific embodiments, the cells 102 of each row (e.g., cells 102 11 to 102 1n ) can form a group of cells storing the reference sequence, and the reference sequence is stored corresponding to the order of the cells 102 in the group of cells. In the example of Figure 2 the array 101 A includes m groups of cells 102 that can store different reference sequences.

[0040] The control circuit 104 also controls the programming line and search line driver 110 to load the substring sequence into each group of cells 102 for comparison with the stored reference sequence, where the substring sequence corresponds to the order of the group of cells 102. In the example of Figure 2 the array 101 A includes n columns of cells, which allows loading a substring sequence of length n when used as a CAM, or a substring sequence of length n / 2 when used as a TCAM or a CAM in some specific embodiments. The number of cells 102 in each group or row of cells and thus the length of the stored reference sequence and the loaded substring sequence can depend on the desired accuracy for positioning the substring sequence in the reference genome.

[0041] In some specific implementations, the number of units 102 in each group can be 40 units, allowing a substring sequence representing 20 bases to be compared with a reference sequence representing 20 bases in TCAM mode. As discussed in more detail in co-pending patent application Ser. No. 16 / 820,711, filed Mar. 20, 2020, and incorporated herein by reference, substring sequences between 17 and 25 bases in length can provide a sufficient number of unique matches for most substring sequences relative to the reference genome H38. Substring lengths shorter than 17 bases will require a greater number of substring sequences from the sample reads to determine the position of the sample reads within the reference genome H38, and substring lengths shorter than 15 bases may not identify any unique matches for almost all of the attempted substring sequences within the reference genome H38.

[0042] On the other hand, substring lengths greater than 25 bases will incur additional storage costs in terms of the units required in the array 101 and greater computational costs due to the associated increase in the operations required, while the number of unique matches is improved little. However, one of ordinary skill in the art, referring to the present disclosure, will understand that for other examples, such as when using a different reference genome or a portion of a reference genome, different substring lengths or different numbers of units in each group of units may be preferred, such as in the case of a medical diagnosis of a genetic condition related to a specific portion of a reference genome. Additionally, different trade-offs between computational cost, the number of units, and accuracy in terms of a greater number of unique matches can also affect the number of units in each group of units for the array 101. For example, some specific implementations may provide groups of units 102 with 200 units to correspond to a short sample read length of 200 bases. However, as discussed in more detail below with reference to Figure 5 the sample reads can be logically divided into multiple adjacent substring sequences for parallel or concurrent processing by multiple arrays 101.

[0043] Advantageously, the units in each row can be loaded simultaneously or almost simultaneously with the same substring sequence in a single load operation or clock cycle to quickly identify any group of units 102 in which the loaded substring sequence matches or approximately matches a reference sequence stored in a different group of units 102. In some specific implementations, if the array 101 A is used as a TCAM, each unit 102 can store two bits representing a reference value or half of a reference value, and for each column of units 102 (e.g., units 102 11 to 102 m1 ), each column can be loaded with two bits having a high value or a low value (i.e., "1" or "0" value) driven for the search lines SL1 and SL2.

[0044] In other specific implementations, each unit 102 may store 3 bits instead of 2 bits, where the third bit is used as a masking bit or "don't care" bit for a wildcard value for approximate matching. In such specific implementations, each unit 102 may store a reference value representing a reference base. When the masking bit is set, the match line will indicate a match regardless of the sample value loaded. In such specific implementations, a third search line (e.g., SL3) may be used to load a wildcard value as part of a substring sequence, which may result in a match regardless of the stored reference value.

[0045] Figure 2 The use of the groups or rows of units 102 shown in allows for a quick comparison of a substring sequence with a large number of reference sequences in a single operation or clock cycle to improve processing efficiency. Array 101 A may include rows or groups of, for example, hundreds or thousands of units to allow for a comparison of a substring sequence with hundreds or thousands of reference sequences in a single operation or clock cycle. In some specific implementations, each group of units 102 stores an overlapping reference sequence that has been shifted by one or another predetermined number of reference bases from the reference sequences stored in the adjacent groups of units.

[0046] Before loading the substring sequence, the control circuit 104 controls the precharge circuit 112 to charge the match lines (ML) for each group or row of units 102. In some specific implementations, the precharge circuit 112 may charge each match line to a high value (i.e., a "1" value). The precharge circuit 112 may include, for example, PMOS transistors. When the substring sequence is loaded into the groups of units by programming the search lines for each column, a single mismatch between the value stored in the unit and the value in the unit into which the substring sequence is loaded will drive the match line for that group of units to a low value (i.e., a "0" value). As discussed in more detail below, using a wildcard value will provide a match for the unit regardless of the value compared to the wildcard value.

[0047] Sense and latch circuit 114 A After the substring sequence has been loaded, the values of each match line are read, and the sensed or read match line values are provided to the encoder 116 A to identify the groups of units for which the match lines or the substring sequences loaded therein match the stored reference sequences. The encoder 116 A may provide an address or other indication of the location of the groups of units for which the stored reference sequences match the loaded substring sequence. In Figure 2In an example, if any, an indication of a group of the identified cells may be added to index 10, which may include, for example, a bitmap or other data structure for indicating the location of the group of the matching cells, which in turn may indicate an exact match or an approximate match between the stored reference sequence and the loaded substring sequence.

[0048] Those of ordinary skill in the art will understand that other specific implementations may include circuit arrangements or operations different from those shown and described for the Figure 2 example. For example, other specific implementations may alternatively set the match line to a low value and indicate a low value for a group of cells storing a reference sequence that matches or approximately matches the loaded substring sequence. As another exemplary variation, other specific implementations may not include the memory 106 but may alternatively provide an indication of a group of cells that match or do not match directly to the control circuit 104.

[0049] Figure 3 is for an array 101 from Figure 2 in accordance with one or more embodiments A of the NVM cells 102 12 exemplary circuit diagram. Figure 3 The arrangement of the 12 includes four transistors and two resistors. In some specific implementations, the NVM cells 102 12 may be four-transistor, two-resistor RRAM, which may be referred to as a 4T2R-RRAM structure. Those of ordinary skill in the art, referring to the present disclosure, will understand that other cell structures are possible, such as two-transistor, two-resistor PCM, referred to as a 2T2-PCM structure, or six-transistor, two-resistor MRAM structure, referred to as a 6T2-MTJ structure. As described above, in contrast to conventional volatile cell structures such as SRAM, using a non-volatile cell structure allows the array to consume less power since the stored reference values do not have to be refreshed. Additionally, using fewer transistors than a conventional SRAM TCAM cell that may use sixteen transistors allows for a higher storage density, such as more megabytes per square millimeter of the array. This may reduce the physical size of the device 100 or allow for more arrays 101 in the device 100 to reduce processing time.

[0050] In Figure 3In the example, resistors M1 and M2 can be programmed to a high resistance value or a low resistance value by driving word line WL1 and programming line PL to store a reference value. Search lines SL1 and SL2 can be driven to a high value or a low value corresponding to an input sample value, which can change the charge of match line ML1 according to whether the charge of SL1 matches that of M1 and whether the charge of SL2 matches that of M2. If the corresponding charge pairs match, match line ML1 retains its precharged state to indicate that the stored value matches the loaded value. If any of the corresponding charge pairs do not match, match line ML1 is pulled down and loses its precharged state to indicate that the stored value does not match the loaded value.

[0051] Figure 3 in cell 102 12 The example in stores two bits for comparison with two bits loaded into the cell via SL1 and SL2. Other specific implementations may use different cell structures. For example, other cell structures may allow a masked bit or a third bit (e.g., "ternary digit") to be stored in the cell and may also allow a masked bit to be loaded into the cell. In such examples, when array 101 is used as a TCAM, the masked bit can be used as a wildcard value in the stored reference sequence or the loaded substring sequence. Then the state of match line ML1 can be maintained when setting the masked bit, regardless of other bits stored or loaded into the cell. In Figure 3 In the exemplary structure of, the value stored in cell 101 12 or loaded into cell 101 12 The wildcard value in can be obtained by encoding the reference value and the sample value, which is discussed in more detail below with reference to Figure 4A and Figure 4B more detailed discussion.

[0052] As Figure 4A shown, the input of a sample value from a substring sequence can have three different values of 1, 0, or wildcard value X, depending on the programming of search lines SL1 and SL2. For example, when both search lines SL1 and SL2 have a value of 1, the loaded or input value is wildcard value X. Similarly, when both resistors M1 and M2 have a value of 1, the stored value is wildcard value X. A wildcard value for either the loaded value or the stored value will result in a match, which will maintain Figure 3 the precharged state of the match line in the exemplary cell structure.

[0053] As described above, in the case of DNA sequencing, each of the four bases - adenine (A), guanine (G), cytosine (C), and thymine (T) - can be represented by two bits, for example. For example, A can be represented by 00, G can be represented by 01, C can be represented by 10, and T can be represented by 11. When operating as a TCAM and using Figure 4AWhen encoding, each unit 102 can store a value of 1, 0, or X, such that two adjacent units 102 in a group or row of units represent a stored reference base, such as A, G, C, or T.

[0054] In Figure 4A the exemplary truth table, when M1 is programmed to a high value (i.e., value 1 for M1) and M2 is programmed to a low value (i.e., value 0 for M2), the programming of resistors M1 and M2 can provide the stored value 1. On the other hand, when M1 is programmed to a low value and M2 is programmed to a high value, the stored reference value 0 is provided. When both M1 and M2 are programmed to high values (i.e., 11), the stored reference value is the wildcard value X, which means that the loaded or input value will match the stored value, regardless of the loaded value (i.e., the values programmed for S1 and S2).

[0055] The loaded values controlled by S1 and S2 follow a similar pattern, where a high value of S1 and a low value of S2 result in an input value 1, while a low value of S1 and a high value of S2 result in an input value 0. Similar to the stored values controlled by M1 and M2, high values of both S1 and S2 result in a loaded wildcard value X, which means that the loaded value will match the stored value, regardless of the stored value (i.e., the values programmed for M1 and M2).

[0056] Figure 4B is an exemplary truth table of NVM cells 102 when operating as a CAM according to one or more embodiments. As Figure 4B shown, the programming of the cells is the same as the operation as a TCAM in Figure 4A except that the encoding scheme has been changed to remove the wildcard value X from the programming, since they are not used for exact matching. In other words, S1 and S2 are no longer programmed to high values (i.e., 11), and M1 and M2 are no longer programmed to high values, such that the wildcard value is no longer used. In such embodiments, a pair of cells 102 can still be used to represent a base, as in Figure 4A the example of the operation as a TCAM. This can facilitate the reuse of the stored values in the cells for the reference sequence, such that it is not necessary to store the reference sequence again when switching between the exact matching phase and the approximate matching phase for sequencing of the same genome or a portion thereof.

[0057] However, other embodiments may allow for reference sequences and substring sequences that are approximately twice as long when the array 101 operates as a CAM as when the array operates as a TCAM. In such embodiments, each cell 102 may represent a single base, where the loaded and stored values for each cell are two-bit values to represent four different bases. For example, instead of the stored values of two adjacent cells 102 representing a base (such as A with an input value of 00), the search lines S1 and S2 may each be programmed to 0 to represent loading the base value of 00 for A into a single cell 102. Similarly, M1 and M2 may each be programmed to represent the stored base for the cell 102.

[0058] The control circuit 104 may be configured to change the encoding scheme for the reference sequences and substring sequences to switch between operation of the array as a CAM and as a TCAM. Such a change in the encoding scheme may include, for example, designating a value (e.g., value 11) as a wildcard value and using that value to operate the array in TCAM mode while avoiding using that value to operate the array in CAM mode. In other embodiments, the control circuit 104 may change the encoding scheme by using a third bit or state for the wildcard value. By changing the encoding scheme, the array 101 can generally be selectively used as both a CAM and a TCAM, which can be used for the exact match phase and the approximate match phase of genome sequencing. Compared to conventional sequencing devices that use dedicated hardware for the exact match phase and the approximate match phase, this versatility of the device 100 can improve processing efficiency and reduce the size of the device 100 for a given performance latency.

[0059] One of ordinary skill in the art, referring to the present disclosure, will understand that other encoding schemes or programming are possible in addition to Figure 4A and Figure 4B the ones shown. For example, input 1 may alternatively be generated by a low value of SL1 and a high value of SL2, or the wildcard value X may be generated by two low values instead of being generated by two high values as in the exemplary truth tables of Figure 4A and Figure 4B .

[0060] Figure 5 is an example of a device 100 for simultaneously processing substring sequences according to one or more embodiments. As Figure 5 shown, the arrays 101 A 、101 B 、101 C and 101 D are simultaneously loaded with different substring sequences. In some embodiments, this may be performed as a single load operation during a single clock cycle of the control circuit 104. Then, sensing and latching circuits 114 such as Figure 2 inA and the encoder 116 A The circuitry can simultaneously identify groups of cells in which the stored reference sequences match the substring sequences loaded into the array. Such circuitry can then provide an indication of the location of such matching sequences to update an index stored in the memory 106, such as index 10.

[0061] In Figure 5 the example of, substring sequences A and B are two parts of the same sample read, and the two parts are processed in parallel or simultaneously to identify groups of cells in which the loaded substring sequences match or approximately match the stored reference sequences. The number of different substring sequences taken from the sample read for comparison with the reference sequence can depend on, for example, the length of the substring sequence, the length of the reference genome, the length of the sample read (e.g., a 250- or 300-base short read from an Illumina device versus a 5000-base long read from a nanopore device), the accuracy of the process used to generate the sample read, the reference genome for the stored reference sequence (e.g., reference genome H38), and the desired accuracy for positioning the sample read within the reference genome or the sample genome. In one example, a short read of 250 or 300 bases can be located in a reference genome with only a few matching substring sequences. Such an example can use only ten substring sequences from the sample read to generate sufficient matches to position the sample read within the reference genome.

[0062] Similarly, substring sequences C and D are two parts of different sample reads that are loaded simultaneously relative to each other and relative to substring sequences A and B. The arrays 101 A and 101 B can be combined into a first array group 1051 for finding matches for the first sample read, and the arrays 101 C and 101 D can be combined into a second array group 1052 for finding matches for the second sample read. As will be understood by those of ordinary skill in the art with reference to this disclosure, the number of arrays 101 and / or groups of arrays can be much greater than those provided for illustrative purposes as shown in Figure 5 . For example, the devices of this disclosure can include hundreds of arrays 101 to allow parallel or simultaneous comparison of many parts of a sample read with a reference sequence, parallel or simultaneous comparison of parts of many different sample reads with a reference sequence, or parallel or simultaneous comparison of a single substring sequence with a large number of reference sequences. Such parallel or concurrent operation of the arrays 101 can improve the processing efficiency of the device 100.

[0063] Figure 6 is an example of a device 100 for simultaneously processing substring sequences in an exact match phase and an approximate match phase according to one or more embodiments. The array 101A 、 101 B 、 101 C and 101 D Each of the arrays in A , B , C , and D can operate as a CAM and a TCAM, depending on the encoding scheme used to store and / or load the sequences in the array. As discussed above, programming of wildcard values or X values allows the array 101 to operate as a TCAM. These wildcard values can be left unused to allow the array 101 to operate as a CAM. In some embodiments, due to the ability of each cell in CAM mode to store two bits representing any one of four bases, the array 101 can switch from storing a reference value in two cells in TCAM mode to storing a reference value in one cell.

[0064] In Figure 6 the example of Figure 6 , the first group 1051 of arrays 101 and 101 simultaneously uses the array 101 A and 101 B to perform an exact match operation on the substring sequence A and uses the array 101 A to perform an approximate match operation on the substring sequence A. Some of the input or loaded values in the substring sequence A can be replaced with wildcard values, and / or some of the stored reference sequence values can be replaced with wildcard values to operate the array 101 B as a TCAM. B

[0065] Similarly, the second group 1052 of arrays 101 and 101 simultaneously uses the array 101 C and 101 D to perform an exact match operation on the substring sequence B and uses the array 101 C to perform an approximate match operation on the substring sequence B. Some of the input values in the substring sequence B can be replaced with wildcard values, and / or some of the stored reference sequence values can be replaced with wildcard values to operate the array 101 D as a TCAM. The operations of the first group 1051 and the operations of the second group 1052 can also be performed in parallel or simultaneously relative to each other. In other embodiments, the use of the arrays 101 within the array group 105 can occur in stages rather than simultaneously. In such embodiments, the performance of the approximate match stage of the sequencing can depend on the completion of the exact match stage or vice versa. D

[0066] The ability of the device 100 to use the array 101 in TCAM mode or CAM mode can allow different search granularities and allow the same hardware to be used for the exact match stage and the approximate stage of genome sequencing. This can generally reduce the footprint of the hardware required for genome sequencing.

[0067] Figure 7 shows an example of a pipelined substring sequence according to one or more embodiments. As Figure 7 shown, the substring sequence can be passed from one array to another to compare the substring sequence with a large number of reference sequences stored in the array. In addition, different substring sequences can cycle through the array to increase or maximize the number of arrays 101 used during a given clock cycle to further improve processing efficiency. In some specific implementations, the substring sequence can be passed via the control circuit 104, or can be passed more directly via the interconnect 107 from other circuits such as programming lines and search line driver circuits.

[0068] In Figure 7 the example of A the substring sequence E is supplied from the array 101 to the array 101 B in consecutive clock cycles of the device 100, supplied to the array 101 C and supplied to the array 101 D . After clock cycle 0, the array 101 A can be used to compare the next substring sequence with the reference sequence stored in the array 101 A . Then during clock cycle 1 the substring sequence F is loaded into the array 101 A , while the substring sequence E is loaded into the array 101 B .

[0069] In clock cycle 2, the substring sequence G is loaded into the array 101 A , while the substring sequence E is loaded into the array 101 C and the substring sequence F is loaded into the array 101 B . At clock cycle 3, all four arrays 101 are used concurrently, thus comparing the different substring sequences E, F, G, and H with the reference sequences stored in the arrays 101 D , 101 C , 101 B , and 101 A respectively. In seven clock cycles, four different substring sequences are compared with the reference sequences stored in all four arrays 101. As Figure 7 the example of

[0070] Exemplary Process

[0071] Figure 8A flowchart of a genomic sequencing process including an exact match phase and an approximate match phase according to one or more embodiments. Figure 8 The process may be performed by a control circuit 104 of, for example, device 100.

[0072] In block 802, the circuit stores a reference sequence in corresponding groups of NVM cells of one or more arrays. The groups of NVM cells may include cells along a match line in the array. The reference sequence may be stored as a series of reference values that can be represented by resistance values set by resistors in the cells, such as Figure 3 resistors M1 and M2 in.

[0073] In block 804, the circuit loads an exact match phase substring sequence into groups of NVM cells of one or more arrays. The loading of the exact match phase substring sequence may represent an exact match phase in which the loaded substring sequence and the stored reference sequence do not include any wildcard values. In some specific implementations, the exact match phase may be relatively faster than the approximate match phase, and the approximate match phase may require more comparisons and may require additional processing, such as filling and evaluating a scoring matrix such as scoring matrix 20.

[0074] In block 806, one or more groups of NVM cells are identified where the stored reference sequence matches the loaded exact match phase substring sequence. In some specific implementations, the exact match phase substring sequence may be loaded into multiple arrays for concurrent comparison with an entire reference genome or a relatively large portion of the reference genome represented by reference sequences stored in each group of cells in the multiple arrays. In such specific implementations, one or more positions in the reference genome that match the loaded exact match phase substring sequence may be identified. An indication of the identified groups of cells may be stored in the memory of the device, such as stored in Figure 1 memory 106 in as part of index 10 for the circuit to identify one or more groups of NVM cells. In other specific implementations, an indication of the identified groups of cells may be provided directly to the circuit. In the case of multiple matches, the first match position or the first match group of cells in the reference genome may be recorded, or alternatively, all match positions or match groups may be recorded, or since the match is not unique, no match positions or match groups may be recorded.

[0075] As discussed above, multiple substring sequences may be taken from a sample read to provide a probabilistic position within the reference genome based on the match positions or groups of matching cells identified for the substring sequences of the sample read. For example, in some specific implementations, the average of the match positions for a certain number of substring sequences (such as ten substring sequences) taken from a sample read may provide a probabilistic position for the sample read within the reference genome.

[0076] Based on the probabilistic location, an approximate matching phase can be performed on a smaller partition of the reference genome to determine the best alignment of the sample reads within the reference genome partition. In block 808, the approximate matching phase substring sequences are loaded into a group of NVM cells of one or more arrays for approximate matching using TCAM operations. At least one of the loaded approximate matching phase sequences and the stored reference sequences for each group of cells in the approximate matching phase includes at least one wildcard value. Such approximate matching can accommodate read errors or mutations in the sample reads. Additionally, the use of approximate matching algorithms such as the Smith-Waterman algorithm or automaton-based algorithms can be used by a host (such as Figure 1 host 103 in [reference] or circuit 104) to find the best or most suitable alignment of the sample reads within the sample genome by considering insertions, deletions, and mutations in the sample reads.

[0077] In some embodiments, the group of identified cells from the exact matching phase can be used as a series of groups of cells, which can be used as a subset of the groups of all the cells in the array, and which may include only, for example, one or two arrays, rather than the larger set of arrays used in the exact matching phase. In other embodiments, after an exact match is found for a substring sequence or a portion of the sample reads in the initial set, the large set of substring sequences or sample reads can be reduced to a smaller set of non-matching substring sequences or sample reads. The remaining substring sequences or sample reads can then be located within the genome by using approximate matching during the approximate matching phase.

[0078] In block 810, one or more groups of NVM cells are identified in which the stored reference sequences approximately match the approximate matching phase substring sequences loaded in block 808. In some embodiments, such as in the case of performing the Smith-Waterman approximate matching algorithm, the scoring matrix can be analyzed, such as by performing backtracking of the scores stored in the scoring matrix, to identify the best or most suitable alignment of the approximate matching phase substring sequences in the reference genome or the sample genome.

[0079] One of ordinary skill in the art, referring to this disclosure, will understand that other embodiments may include exact matching phases and approximate matching phases in a different order than those discussed above for Figure 8 example, other embodiments may perform the approximate matching phase before the exact matching phase, or may mix the exact matching phase and the approximate matching phase. In one such example, a particular substring sequence may undergo exact matching and then approximate matching, followed by an exact matching phase and an approximate matching phase for a different substring sequence. Additionally, and as discussed above with reference to Figures 5 to 7As discussed in the concurrent processing example, different arrays 101 in device 100 can perform exact matching as a CAM, while other arrays 101 in device 100 perform approximate matching as a TCAM.

[0080] Figure 9 is a flowchart of an operation mode switching process for at least one array of NVM cells according to one or more embodiments. Figure 9 The process can be performed by, for example, control circuit 104 of device 100. In some specific implementations, for the exact matching phase, Figure 9 the process can follow at least one array operating in CAM mode. In other specific implementations, for the approximate matching phase, Figure 9 the process can follow at least one array operating in TCAM mode.

[0081] In block 902, after identifying one or more groups of NVM cells in an earlier mode (e.g., CAM or TCAM mode) or operation phase (e.g., exact matching phase or approximate matching phase), the reference sequences stored in the groups of NVM cells in one or more arrays are retained. In this regard, the circuit does not reset or clear the reference values stored in the groups of NVM cells, and thus retains the reference sequences for use in another mode or operation phase. Compared to conventional genomic sequencing that may require separate hardware and / or separate storage or reprogramming of the reference sequences for each matching phase, this can improve the operational efficiency in some specific implementations when switching between the exact matching phase and the approximate matching phase.

[0082] In block 904, the circuit resets the match lines of one or more arrays. In some specific implementations, this can include recharging the match lines such that all match lines or previously unmatched match lines are raised to a charged state after a previous operation in another mode or different phase. For example, the match line corresponding to the group of cells where there was a mismatch between the reference sequence and the loaded substring sequence in a previous operation can be recharged. In other specific implementations, resetting the match lines can include discharging the match lines such that all match lines or previously matched match lines are lowered to ground.

[0083] In block 906, the circuitry loads a substring sequence of another matching stage that is different from the previously loaded substring sequence in a previous operation or stage. For example, if the previous operation was in the CAM mode for an exact matching stage, the previously loaded substring sequence may have included one or more wildcard values. The new substring sequence loaded in block 906 may then include wildcard values for an approximate matching stage for operations in the TCAM mode. On the other hand, if the previous operation was in the TCAM mode for an approximate matching stage, the previously loaded substring sequence may have included one or more wildcard values. The new substring sequence loaded in block 906 is then encoded to not include wildcard values for an exact matching stage for operations in the CAM mode.

[0084] In block 908, the circuitry identifies one or more groups of NVM cells in at least one array in which the stored reference sequence matches or approximately matches the newly loaded substring sequence. The identification of the one or more groups of NVM cells may be generated by receiving one or more indications from an encoding circuit (e.g., Figure 2 the encoder 116 in A ) regarding which match lines have a changed or maintained charge state. For example, the control circuit may directly receive an indication for identifying a matching group of NVM cells from the encoding circuit or an intermediate memory such as Figure 1 the memory 106 in Figure 1 In some embodiments, the indication of the identified group of cells may be stored in an index such as index 10 in

[0085] As described above, preserving the stored reference sequence from one operation mode to another (e.g., from CAM operation to TCAM operation or vice versa) may allow for relatively fast transition from a first type of matching stage to a second type of matching stage using the same array. This generally improves the operation or processing efficiency of a genomic sequencing device and may facilitate reusing the same array for both an exact matching stage and an approximate matching stage of sequencing.

[0086] Figure 10 is a flowchart of a concurrent matching process according to one or more embodiments. Figure 10 The process of Figures 5 to 7 may be performed by, for example, the control circuit 104 of device 100. As discussed above with reference to

[0087] In block 1002, the circuit simultaneously loads one or more subsequence sequences of one or more sample reads into an array including a group of NVM cells. In some embodiments, the circuit may load subsequence sequences representing different parts of the same sample read. In such examples, the subsequence sequences may provide a probabilistic location for the sample reads within a reference genome, which is represented in whole or in part by the reference sequences stored in the array. In other examples, the loaded subsequence sequences may represent parts from different sample reads. In such examples, one or more arrays may be loaded with subsequence sequences from a first sample read, while one or more other arrays may be loaded with subsequence sequences from a second sample read. In still other examples, a single subsequence sequence may be loaded into multiple arrays to perform a search of the loaded subsequence sequence across most or all of the genome represented by the reference sequences stored in the arrays.

[0088] In block 1004, the circuit simultaneously identifies groups of NVM cells in the array where the stored reference sequences match or approximately match the loaded subsequence sequences. In some embodiments, the array and / or one or more of the subsequence sequences may be programmed or encoded for exact matching. In other embodiments, the array and / or one or more of the subsequence sequences may be programmed or encoded for approximate matching to include wildcard values. The circuit may identify groups of NVM cells where the stored reference sequences match or approximately match the loaded subsequence sequences based on an indication received from an encoder of the array (e.g., Figure 2 the encoder 116 in A ) or based on an indication received from an intermediate memory (e.g., Figure 2 the memory 106 in

[0089] Figure 11 is a flowchart of a sample read classification process according to one or more embodiments. Figure 11 The process of

[0090] In block 1102, the circuit loads into the array the exact-match-phase substring sequences representing portions of multiple sample reads. The particular substring sequences can be loaded into multiple arrays in one time or clock cycle to identify potential match positions within the reference genome represented by the reference sequences stored across the multiple arrays. Then, different exact-match-phase substring sequences from the same or different sample reads can be loaded in subsequent times or clock cycles. In other specific embodiments, different exact-match-phase substring sequences from the same sample read or from different sample reads can be loaded into different arrays in one time or clock cycle, and the next set of exact-match-phase substring sequences from the same sample read or from different sample reads can then be loaded in subsequent times or clock cycles.

[0091] In block 1104, one or more groups of NVM cells in the array where the stored reference sequences match the loaded exact-match-phase substring sequences are identified. The identification of one or more groups of NVM cells can be performed by a circuit such as a sense and latch circuit for the array (e.g., Figure 2 the sense and latch circuit 114A in

[0092] In block 1106, the control circuit of the device (e.g., Figure 1 the control circuit 104 in Figure 1 ), or a host communicating with the device (e.g., Figure 1 the host 103 in Figure 1 ), determines the probable position of the sample read within the reference genome based on the one or more groups of the identified NVM cells. The control circuit or the host can use the indication of the groups of the identified cells stored in an index or other data structure to determine the probable or approximate position of the sample read. For example, referring to Figure 1 , the memory 106 can store an index 10 that can indicate the average of multiple positions identified for different substring sequences of the same sample read. The position of the sample read determined in block 1106 can be probabilistic in the sense that multiple possible positions can be identified for different substring sequences from the same read, and a consensus or statistic from the match positions can be used to probabilistically position the sample read within the reference genome.

[0093] In one example, the average of all identified positions of a group of NVM cells that match a substring sequence of a sample read is used to identify the most likely position of the sample read within the reference genome. In another example, only one position of a group of cells having one match for each substring sequence is used in the average. In yet another example, the probable position of the sample read can be determined by identifying the most distant spaced positions within the reference genome that correspond to a group of NVM cells that match a substring sequence from the sample read. In other examples, when determining the probable position of a sample read within the reference genome, one or more outlier positions regarding a set of matching positions can be discarded.

[0094] In block 1108, the circuit or host classifies the plurality of sample reads into sample groups to align the sample reads using approximate matching, at least in part based on the determined probable positions of the sample reads. In this regard, the approximate matching phase that can be performed by device 100 after the exact matching performed by device 100 in the Figure 11 course of the process finds the best or optimal alignment of the sample reads to a smaller partition of the reference genome that is associated with the sample group. In other embodiments, the following approximate matching phase performed by device 100 can use approximate matching to find the best or optimal alignment of the sample reads relative to other sample reads in the sample group. In both cases, since a smaller partition of the reference genome or a smaller number of sample reads are compared to each other within each sample group, the computational complexity and processor hours used to perform the approximate matching phase are significantly reduced by classifying the sample reads into sample groups.

[0095] Those of ordinary skill in the art will understand, upon reference to this disclosure, Figure 11 that the order of the blocks can be different in other embodiments. For example, blocks 1102 to 1106 can be performed as an iterative set for each sample read before the classification of the sample reads is performed in block 1108. In other examples, for a predefined or dynamically changing partition of the reference genome, the probable position of each sample read can be classified on the fly for successive iterations of a process that performs one sample read at a time. In still other embodiments, after determining the probable positions of all sample reads, the circuit can determine the groups of sample reads for classification such that the groups of sample reads include approximately the same number of sample reads or cover approximately the same range of bases in the reference genome.

[0096] As discussed above, the foregoing examples of devices including arrays of NVM cells that can operate as CAMs or TCAMs according to programming or encoding can generally improve processing efficiency and reduce the physical size of genomic sequencing devices. In this regard, the reuse of the same array for exact matching when operating as a CAM and for approximate matching when operating as a TCAM can reduce the storage amount required for different types of matching and can allow the stored reference sequences to be reused for different stages or for different subsequence sequences representing one or more sample reads. Additionally, compared to SRAM arrays, the non-volatile nature of the arrays disclosed herein can reduce the energy requirements for performing sequencing while still obtaining the benefits of the fast matching provided by a CAM.

[0097] Other Embodiments

[0098] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks, modules, and processes described in connection with the examples disclosed herein can be implemented as electronic hardware, software, or combinations of both. Additionally, the foregoing processes can be embodied on a computer-readable medium that causes a processor, controller, or other circuitry to perform or implement certain functions.

[0099] To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, and modules have been described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those of ordinary skill in the art can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0100] The various illustrative logical blocks, units, modules, and circuits described in connection with the examples disclosed herein can be implemented or performed with a general-purpose processor, GPU, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor or controller circuitry can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, an SoC, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0101] The acts of a method or process described in connection with the examples disclosed herein may be embodied directly in hardware, in a software module executed by a processor or other circuitry, or in a combination of both. The steps of the method or algorithm may also be performed in an order alternative to that provided in the examples. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable medium, an optical medium, or any other form of storage medium known in the art. The exemplary storage medium is coupled to the processor or controller circuitry such that the processor or controller circuitry can read information from, and write information to, the storage medium. In an alternative, the storage medium may be integral to the processor or controller circuitry. The processor or controller circuitry and the storage medium may reside in an ASIC or an SoC.

[0102] The foregoing description of the exemplary embodiments of the present disclosure has been provided to enable any person of ordinary skill in the art to make or use the embodiments of the present disclosure. Various modifications to these examples will be readily apparent to those of ordinary skill in the art, and the principles disclosed herein may be applied to other examples without departing from the spirit or scope of the present disclosure. The embodiments are to be considered in all respects only as illustrative and not restrictive. Additionally, the language used in the following claims in the form of "at least one of A and B" should be construed to mean "only A, only B, or both A and B".

Claims

1. An apparatus, the apparatus comprising: a plurality of arrays of non-volatile memory (NVM) cells; and circuitry configured to: store reference sequences in respective groups of NVM cells of the plurality of arrays, each reference sequence representing a part of a genome; load one or more exact match stage substring sequences into groups of NVM cells of the plurality of arrays for comparison with the reference sequences stored in the groups of NVM cells, the one or more exact match stage substring sequences representing parts of at least one sample read; use the plurality of arrays as a content addressable memory (CAM) to identify one or more groups of NVM cells in the plurality of arrays in which the stored reference sequences match the loaded exact match stage substring sequences, wherein when using the plurality of arrays as a CAM, neither the reference sequences being compared nor the one or more exact match stage substring sequences include wildcard values; load one or more approximate match stage substring sequences into groups of NVM cells of the plurality of arrays for comparison with the reference sequences stored in the groups of NVM cells, the one or more approximate match stage substring sequences representing parts of the at least one sample read; and use the plurality of arrays as a ternary CAM (TCAM) to identify one or more groups of NVM cells in the plurality of arrays in which the stored reference sequences approximately match the loaded approximate match stage substring sequences, wherein when using the plurality of arrays as a TCAM, at least one of the reference sequences being compared and the approximate match stage substring sequences for each group of NVM cells includes at least one wildcard value; and wherein the circuitry is further configured to: load one or more substring sequences into the arrays of the plurality of arrays simultaneously; and simultaneously identify groups of NVM cells in the arrays in which the stored reference sequences match or approximately match the loaded one or more substring sequences.

2. The apparatus according to claim 1, wherein the circuitry is further configured to change an encoding scheme for the reference sequences and the substring sequences to switch between operation of the plurality of arrays as a CAM and as a TCAM.

3. The apparatus according to claim 1, wherein the circuitry is further configured to: retain the reference sequences stored in a group of NVM cells of one of the plurality of arrays after operation of the array as a CAM; reset the match lines of the array; load an approximate match stage substring sequence into the group of NVM cells of the array; and identify one or more groups of NVM cells in the array in which the stored reference sequences approximately match the loaded approximate match stage substring sequence for operation of the array as a TCAM.

4. The apparatus according to claim 1, wherein the circuitry is further configured to: retain the reference sequences stored in a group of NVM cells of one of the plurality of arrays after operation of the array as a TCAM; Reset the match lines of the array; Load a sequence of exact match phase substrings into the groups of the NVM cells of the array; And Identify one or more groups of NVM cells in the array in which the stored reference sequences match the loaded sequence of exact match phase substrings for operation of the array as a CAM.

5. The apparatus of claim 1, wherein the one or more concurrently loaded substring sequences include substring sequences representing different portions of the same sample read.

6. The apparatus of claim 1, wherein the circuit is further configured to use an indication of one or more identified groups of NVM cells from a first array of the plurality of arrays for a first substring sequence for at least one of the following operations: Determine one or more reference sequences to be stored in a second array of the plurality of arrays; and Determine at least one additional substring sequence to be loaded into the second array.

7. The apparatus of claim 1, the apparatus further comprising at least one memory configured to store an index indicating the identified groups of NVM cells in which the stored reference sequences approximately match or exactly match the loaded substring sequences.

8. The apparatus of claim 1, the apparatus further comprising at least one memory configured to store a Smith-Waterman scoring matrix for use during an approximate match phase.

9. The apparatus of claim 1, wherein the circuit is further configured to: Load a sequence of exact match phase substrings into an array of the plurality of arrays, wherein the sequence of exact match phase substrings represents portions from a plurality of sample reads; Identify one or more groups of NVM cells in the array in which the stored reference sequences match the loaded sequence of exact match phase substrings, wherein each stored reference sequence represents a portion of a reference genome; Determine the probable positions of the plurality of sample reads within the reference genome based on the one or more identified groups of NVM cells; And Classify the plurality of sample reads into a plurality of sample groups at least in part based on the determined probable positions of the respective sample reads for aligning the sample reads using approximate matching.

10. A method of genome sequencing, the method comprising: Storing reference sequences in respective groups of non-volatile memory (NVM) cells of at least one array of NVM cells, each reference sequence representing a portion of a genome; Loading one or more sequences of exact match phase substrings into groups of NVM cells of the at least one array for operation of the at least one array as a content addressable memory (CAM), the one or more sequences of exact match phase substrings representing one or more portions of at least one sample read; Identify one or more groups of NVM cells in the at least one array where the stored reference sequences match the loaded exact match phase substring sequences, where neither the stored reference sequences nor the loaded one or more exact match phase substring sequences include wildcard values when the at least one array operates as a CAM; Load one or more approximate match phase substring sequences into a group of NVM cells in the at least one array for operation of the at least one array as a ternary CAM (TCAM), where the one or more approximate match phase substring sequences represent one or more portions of the at least one sample read segment; And Identify one or more groups of NVM cells in the at least one array where the stored reference sequences approximately match the loaded approximate match phase sequences, where when the at least one array operates as a TCAM, at least one of the stored reference sequences and the loaded approximate match phase substring sequences for each group of NVM cells includes at least one wildcard value; And Where the method further includes: Simultaneously load one or more substring sequences of one or more sample read segments into the array in the at least one array; And Simultaneously identify groups of NVM cells in the array where the stored reference sequences match or approximately match the loaded one or more substring sequences.

11. The method according to claim 10, the method further includes changing the encoding scheme for the reference sequences and the substring sequences to switch between operation of the at least one array as a CAM and as a TCAM.

12. The method according to claim 10, the method further includes: Retain the reference sequences in the group of NVM cells of one of the at least one array after operation of the array as a CAM; Reset the match lines of the array; Load the approximate match phase sequences into the group of NVM cells of the array; And Identify one or more groups of NVM cells in the array where the stored reference sequences approximately match the loaded approximate match phase sequences for operation of the array as a TCAM.

13. The method according to claim 10, the method further includes: Retain the reference sequences in the group of NVM cells of one of the at least one array after operation of the array as a TCAM; Reset the match lines of the array; Load the exact match phase substring sequences into the group of NVM cells of the array; And Identify one or more groups of NVM cells in the array where the stored reference sequences match the loaded exact match phase substring sequences for operation of the array as a CAM.

14. The method according to claim 10, where the one or more simultaneously loaded substring sequences include substring sequences representing different portions of the same sample read segment.

15. The method according to claim 10, the method further comprising using an indication of one or more identified groups of NVM cells of a first array of the at least one array for a first substring sequence for at least one of the following operations: determining one or more reference sequences to be stored in a second array of the at least one array; and determining at least one additional substring sequence to be loaded into the second array.

16. The method according to claim 10, the method further comprising storing in the at least one memory an indication of an identified group of NVM cells in which a reference sequence stored therein approximately matches or exactly matches a loaded substring sequence.

17. The method according to claim 10, the method further comprising: loading an exact match phase substring sequence into the at least one array, wherein the exact match phase substring sequence represents a portion from a plurality of sample reads; identifying one or more groups of NVM cells in the at least one array in which a stored reference sequence matches the loaded exact match phase substring sequence, wherein each stored reference sequence represents a portion of a reference genome; determining a probabilistic position of the plurality of sample reads within the reference genome based on the one or more identified groups of NVM cells; and classifying the plurality of sample reads into a plurality of sample groups at least in part based on the determined probabilistic positions of the respective sample reads.

18. An apparatus for genome sequencing, the apparatus comprising: at least one array of non-volatile memory (NVM) cells, the at least one array of non-volatile memory (NVM) cells being configured to store reference sequences in respective groups of the NVM cells, each reference sequence representing a portion of a genome; and means for: storing the reference sequences in the respective groups of NVM cells of the at least one array; loading one or more exact match phase substring sequences into a group of NVM cells of the at least one array for operation of the at least one array as a content addressable memory (CAM), the one or more exact match phase substring sequences representing one or more portions of at least one sample read; identifying one or more groups of NVM cells in the at least one array in which a stored reference sequence matches the loaded exact match phase substring sequence, wherein when the at least one array operates as a CAM, neither the stored reference sequence nor the loaded one or more exact match phase substring sequences include wildcard values; loading one or more approximate match phase substring sequences into a group of NVM cells of the at least one array for operation of the at least one array as a ternary CAM (TCAM), the one or more approximate match phase substring sequences representing one or more portions of the at least one sample read; Identify one or more groups of NVM cells in the at least one array where the stored reference sequences approximately match the loaded approximate match phase sequences, where when the at least one array operates as a TCAM, at least one of the stored reference sequences and the loaded approximate match phase substring sequences for each group of NVM cells includes at least one wildcard value; Simultaneously load one or more substring sequences of one or more sample reads into the arrays in the at least one array; And Simultaneously identify groups of NVM cells in the array where the stored reference sequences match or approximately match the loaded one or more substring sequences.

19. An apparatus, the apparatus comprising: Multiple arrays of non-volatile memory (NVM) cells, the multiple arrays of non-volatile memory (NVM) cells being capable of operating in a content-addressable memory (CAM) mode and a ternary CAM (TCAM) mode; And A circuit configured to: Store reference sequences in corresponding groups of NVM cells of the multiple arrays; Load one or more exact match phase substring sequences into groups of NVM cells of the multiple arrays; Set the multiple arrays in the CAM mode and identify one or more groups of NVM cells where the stored reference sequences match the loaded exact match phase substring sequences; Load one or more approximate match phase substring sequences into groups of NVM cells of the multiple arrays; and Set the multiple arrays in the TCAM mode and identify one or more groups of NVM cells where the stored reference sequences approximately match the loaded approximate match phase substring sequences, where at least one of the compared reference sequences and the approximate match phase substring sequences includes at least one wildcard value; and Wherein the circuit is further configured to: Simultaneously load one or more substring sequences into the arrays in the multiple arrays; and Simultaneously identify groups of NVM cells in the array where the stored reference sequences match or approximately match the loaded one or more substring sequences.

20. The apparatus according to claim 19, wherein the circuit is further configured to change the encoding scheme for the reference sequences and the substring sequences to switch between the operation of the multiple arrays in the CAM mode and the TCAM mode.

21. The apparatus according to claim 19, wherein the circuit is further configured to: Retain the reference sequences stored in the groups of NVM cells of one of the multiple arrays after the operation of the array in the CAM mode; Reset the match lines of the array; Load the approximate match phase substring sequences into the groups of NVM cells of the array; And Identify one or more groups of NVM cells in the array where the stored reference sequences approximately match the loaded approximate match phase substring sequences for the operation of the array in the TCAM mode.

22. The apparatus according to claim 19, wherein the circuit is further configured to: After operation of the array in the TCAM mode, a reference sequence is retained in a group of NVM cells of one of the plurality of arrays; Reset the match lines of the array; Load an exact match phase substring sequence into the group of NVM cells of the array; And Identify one or more groups of NVM cells of the array in which the stored reference sequence matches the loaded exact match phase substring sequence for operation of the array in the CAM mode.

23. The apparatus of claim 19, wherein the one or more substring sequences loaded simultaneously include substring sequences representing different portions of the same sample read segment.

Citation Information

Patent Citations

  • Reference-guided genome sequencing

    US20210292830A1

  • Reference-guided genome sequencing

    US20210295946A1

  • Devices and methods for locating a sample read in a reference genome

    US20210295949A1