Error correction method for a computational SSD that supports high-speed file semantic search

The computational SSD system addresses inefficiencies in conventional AI/ML semantic search by distributing calculations across NAND flash dies and using die-based computational logic for result aggregation, resulting in reduced bandwidth and energy waste and improved computational efficiency.

JP2025517740AActive Publication Date: 2025-06-10SANDISK TECHNOLOGIES LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024568300
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-14
Filing Date
2023-10-17
Publication Date
2025-06-10
Estimated Expiration
2043-10-17

AI Technical Summary

Technical Problem

Conventional AI/ML implementations of semantic search on storage devices face inefficiencies due to substantial bandwidth and energy waste from data movement between the host processor, host memory, and storage device, as well as bottlenecks in the storage device controller handling both I/O and computational operations.

Method used

A computational SSD system that distributes calculations across each NAND flash die of the SSD, using a new die-based computational logic circuit for result aggregation, allowing for file-semantic search by reading file feature vectors from multiple dies to the SSD controller and processing distance calculations in DRAM.

Benefits of technology

This approach reduces bandwidth and energy waste, alleviates bottlenecks in the storage device controller, and enhances computational speed by optimizing data processing within the SSD system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517740000001_ABST
    Figure 2025517740000001_ABST
Patent Text Reader

Abstract

In this specification, a device and method for performing semantic search on an SSD are disclosed via a computational SSD system that distributes calculations across each NAND flash die of the SSD and, at the same time, the SSD controller processes result aggregation using a new die-on computational logic circuit to provide file semantic search on the device. The computational SSD system can read file feature vectors from multiple dies to the SSD controller, and these feature vectors may be buffered in DRAM if necessary, and the controller processes distance calculations. A local die-on AI / ML processing unit can, for example, perform computational and comparison operations and pass the processing scores and results to the SSD controller. The SSD controller aggregates the results from all dies and returns the results to the host. The feature vector store size, circuitry, and the number of die-on AI / ML processing units can be configured as needed to adapt to different tasks, system constraints, and / or feature vector sizes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Non - Provisional Application No. 18 / 449,165, filed on August 14, 2023, entitled "Error Correction Methods for Computational SSD Supporting Rapid File Semantic Search", which claims the priority of U.S. Provisional Application No. 63 / 476,666, filed on December 22, 2022, and the entire contents thereof are incorporated herein by reference for all purposes.

[0002] This disclosure relates to storage systems. More particularly, this disclosure relates to utilizing a storage system to execute a machine learning process.

Background Art

[0003] Storage devices are ubiquitous within computing systems. In recent years, solid - state storage devices (SSDs) have become increasingly common. These non - volatile storage devices can communicate with and utilize various protocols, including non - volatile memory express (NVMe) and peripheral component interconnect express (PCIe), to reduce processing overhead and increase efficiency.

[0004] As SSDs have evolved, they have become more power efficient compared to traditional hard disk drives (HDDs), thus providing advantages to SSDs in the consumer and commercial markets. Over the years, many users have accumulated thousands of files documenting their lives, experiences, travels, and interests. To find specific media files based on that content, users often spend a lot of time navigating, organizing, and searching through albums and folders for specific photos, videos, and other media. Generally, user media is stored across various storage and computing devices, and users are limited in their methods for contextually searching for photos, videos, and other media to find the desired media file. In fact, users are limited to traditional search methods such as keyword searches to find media on storage systems that are limited to searching for the most recent files or files with specific file names or attributes. Furthermore, the vast amount of stored media files and the lack of contextual search methods make it difficult for users to find specific media on their storage and computing devices.

[0005] With the development of Artificial Intelligence (AI) and Machine Learning (ML), accurate semantic search has become possible, enabling users to find relevant content within their media files stored on their storage devices. However, conventional implementations of AI / ML semantic search on storage devices have several problems that make it inefficient and limited. One problem in conventional AI / ML implementations of semantic search is that data is repeatedly moved between the host processor, host memory, and the storage device, leading to substantial waste of bandwidth and energy from data movement. Another problem in conventional AI / ML implementations of semantic search is preparing data for AI / ML processing for on-host computing, i.e., reading data from multiple dies into the storage device controller, buffering the data in memory for the controller to process the AI / ML processing, and then executing the AI / ML processing. This causes a bottleneck for the storage device controller as the controller has to handle both I / O operations and computational operations, and such computations require additional memory buffers to store data for AI / ML processing.

Brief Description of the Drawings

[0006] The above and other aspects, features, and advantages of the various embodiments of the present disclosure will become more apparent from the following description presented in conjunction with the multiple figures of the following drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 11A

Figure 11B

Figure 11C

Figure 12

Figure 13

Figure 14

Figure 15

[0007] Corresponding reference characters indicate corresponding components throughout the several views of the drawings. The elements in the several views are shown for the sake of brevity and clarity and are not necessarily drawn to scale. For example, some dimensions of the elements in the figures may be emphasized relative to other elements to facilitate understanding of various embodiments of the present disclosure. Additionally, commonly understood elements that are useful or necessary in commercially realizable embodiments are often not drawn to make the more unobscured views of these various embodiments of the present disclosure easier.

DETAILED DESCRIPTION OF THE INVENTION

[0008] In response to the above problems, this specification describes a device and method that can perform semantic search on an SSD through a computational SSD system that distributes calculations across each NAND flash die of the SSD, and at the same time, the SSD controller processes result aggregation using a new die - based computational logic circuit to provide file - semantic search on the device. Specifically, many embodiments utilize a computational SSD system that reads file feature vectors from multiple dies to the SSD controller. When there are millions of feature vectors to be compared, these feature vectors can be buffered in DRAM, and the controller processes distance calculations. Local die - based AI / ML processing units can, for example, execute computational and comparison operations and pass the processing scores and results to the SSD controller. The SSD controller aggregates the results from all dies and returns the results to the host. In some embodiments, the feature vector store size of each die - based AI / ML processing unit can be configured as needed to adapt to different tasks and / or feature vector sizes. Based on this, the circuits of the AI / ML processing units and some die - based AI / ML processing units can be configured and distributed as needed to provide increased computational speed or to meet specific die - area constraints. Additional embodiments are discussed in more detail below.

[0009] These solutions can help reduce substantial waste of bandwidth and energy from data movement between the host processor, host memory, and storage device. Further, these solutions can help reduce the bottleneck of the storage device controller in AI / ML processing, where the storage device controller needs to read feature vectors from multiple dies to the controller and buffer the feature vectors in memory to perform AI / ML calculations and processing, and thus needs to handle both I / O and computational operations in AI / ML operations.

[0010] Aspects of the present disclosure may be embodied as an apparatus, system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software aspects and hardware aspects that are generally referred to herein as a “function,” “module,” “apparatus,” or “system.” Further, aspects of the present disclosure may take the form of a computer program product embodied in one or more non-transitory computer-readable storage media storing computer-readable and / or executable program code. Many of the functional units described herein are labeled as functions in order to more specifically emphasize their implementation independence. For example, a function may be implemented as a hardware circuit comprising off-the-shelf semiconductors such as a custom VLSI circuit or gate array, logic chip, transistor, or other discrete components. A function may also be implemented on a programmable hardware device via a field programmable gate array, programmable array logic, programmable logic device, or the like.

[0011] A function may also be implemented at least in part in software for execution by various types of processors. An identified function of executable code may comprise, for example, one or more physical or logical blocks of computer instructions that may be organized as an object, procedure, or function. Nevertheless, the executable files of an identified function need not be physically located together and may include disparate instructions stored at different locations, which, when logically bound together, constitute that function and achieve the specified purpose of that function.

[0012] In fact, the functionality of the executable code can include a single instruction or many instructions and can even be distributed over multiple different code segments, between different programs, over multiple storage devices, etc. When a functionality or a part of a functionality is implemented in software, the software part can be stored on one or more computer-readable and / or executable storage media. Any combination of one or more computer-readable storage media can be utilized. The computer-readable storage media can include, for example, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing, but does not include propagated signals. In the context of this document, a computer-readable and / or computer-executable storage medium may be a tangible and / or non-transitory medium that can include or store a program used by or in connection with an instruction execution system, apparatus, processor, or device.

[0013] The computer program code for performing the operations for aspects of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Python, Java, Smalltalk, C++, C#, Objective C, conventional procedural programming languages such as the "C" programming language, scripting programming languages, and / or other similar programming languages. The program code may be executed, partially or wholly, on one or more users' computers and / or via a data network, such as over a remote computer or server.

[0014] A component, as used herein, constitutes a tangible, physical, non-transitory device. For example, a component may be implemented as off-the-shelf semiconductors such as custom VLSI circuits, gate arrays, or other integrated circuit hardware logic circuits, logic chips, transistors, or other discrete devices, and / or other mechanical or electrical devices. A component may also be implemented in a programmable hardware device such as a field programmable gate array, programmable array logic, programmable logic device. A component may comprise one or more silicon integrated circuit devices (e.g., chips, dies, die planes, packages), or other discrete electrical elements, that communicate electrically with one or more other components via, for example, electrical lines of a printed circuit board (PCB). Each of the functions and / or modules described herein may alternatively be embodied by, or implemented as, a component in a particular embodiment.

[0015] A circuit, as used herein, includes a set of one or more electrical and / or electronic components that provide one or more paths for current. In certain embodiments, the circuit may include a return path for current such that the circuit is a closed loop. However, in other embodiments, a set of components that may not include a return path for current may be referred to as a circuit (e.g., an open loop). For example, an integrated circuit may be referred to as a circuit regardless of whether it is coupled to ground (as a return path for current). In various embodiments, a circuit may include a portion of an integrated circuit, an integrated circuit, a set of integrated circuits, or non-integrated electrical components and / or a set of electrical components with or without an integrated circuit device. In one embodiment, the circuit may include a custom VLSI circuit, a gate array, a logic circuit, or other integrated circuit. It may be implemented as an off-the-shelf semiconductor such as a logic chip, transistor, or other discrete device, and / or other mechanical or electrical device. The circuit may also be implemented as a synthesized circuit within a programmable hardware device such as a field programmable gate array, programmable array logic, programmable logic device (e.g., as firmware, netlist, etc.). The circuit may comprise one or more silicon integrated circuit devices (e.g., chips, dies, die planes, packages), or other discrete electrical elements that communicate electrically with one or more other components via, for example, electrical traces of a printed circuit board (PCB). Each of the functions and / or modules described herein may, in certain embodiments, be embodied by or implemented as a circuit.

[0016] Throughout this specification, references to "one embodiment", "an embodiment", or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrases "in one embodiment", "in an embodiment", and similar language throughout this specification are not necessarily referring to the same embodiment, and unless otherwise expressly specified, may mean "one or more but not all embodiments". The terms "including", "comprising", "having", and variations thereof mean "including but not limited to" unless expressly specified otherwise. An enumerated list of items does not necessarily mean that any or all of the items are mutually exclusive and / or mutually inclusive unless expressly specified otherwise. "A", "an", and "the" also represent "one or more" unless otherwise expressly specified.

[0017] Furthermore, as used herein, references to reading, writing, storing, buffering, and / or transferring data can include the entire data, a portion of the data, a set of data, and / or a subset of the data. Similarly, references to reading, writing, storing, buffering, and / or transferring non-host data can include the entire non-host data, a portion of the non-host data, a set of non-host data, and / or a subset of the non-host data.

[0018] Finally, as used herein, the terms "or" and "and / or" should be interpreted as inclusive or meaning any one or any combination. Thus, "A, B, or C" or "A, B, and / or C" means any of "A, B, C, A and B, A and C, B and C, A, B, and C". An exception to this definition occurs only when the combination of elements, functions, steps, or acts are mutually exclusive in some way.

[0019] Aspects of the present disclosure will be described below with reference to schematic flowcharts and / or schematic block diagrams of methods, apparatuses, systems, and computer program products according to embodiments of the present disclosure. Each block of the schematic flowchart and / or schematic block diagram, as well as combinations of blocks in the schematic flowchart and / or schematic block diagram, is understood to be implementable by computer program instructions. These computer program instructions may be provided to a processor of a computer or other programmable data processing apparatus to create a machine such that the instructions executed via the processor or other programmable data processing apparatus implement the functions and / or acts specified in the blocks of the schematic flowchart and / or schematic block diagram.

[0020] Also, note that in some alternative implementations, the functions shown within a block may occur in a different order than shown in the figures. For example, two blocks shown in succession may be executed substantially in parallel, or the blocks may be executed in the reverse order depending on the associated functions. As other processes and methods, equivalents of one or more blocks, or portions thereof, shown in the figures may be contemplated in terms of function, logic, or effect. Although various arrow types and line types may be employed in the flowchart and / or block diagram, they are understood not to limit the scope of the corresponding embodiments. For example, an arrow may indicate an unspecified waiting or monitoring period of duration between the recited steps of the illustrated embodiment.

[0021] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. The foregoing summary is merely exemplary and is not intended to be limiting in any way. In addition to the exemplary aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description. The description of the elements in each figure may refer to the elements of the previous figures. The same numbers may refer to the same elements in the figures, which may include alternative embodiments of the same element.

[0022] Referring to FIG. 1, a schematic block diagram of an exemplary host computing device 110 having a storage system 102 suitable for a computing SSD system that supports semantic search according to an embodiment of the present disclosure is shown. The computing SSD system 100 includes one or more storage devices 120 of the storage system 102 within the host computing device 110 that communicate via a controller 126. The host computing device 110 may include a processor 111, volatile memory 112, and a communication interface 113. The processor 111 may include one or more central processing units, one or more general-purpose processors, one or more application-specific processors, one or more virtual processors (e.g., the host computing device 110 may be a virtual machine operating within a host), one or more processor cores, and the like. The communication interface 113 may include one or more network interfaces configured to communicatively couple the controller 126 of the host computing device 110 and / or the storage device 120 to a communication network 115 such as an Internet Protocol (IP) network, a Storage Area Network (SAN), a wireless network, a wired network, or the like.

[0023] In various embodiments, the storage device 120 can be located at one or more different positions relative to the host computing device 110. In one embodiment, the storage device 120 includes one or more non-volatile memory devices 123, such as semiconductor chips or packages or other integrated circuit devices disposed on one or more printed circuit boards, storage housings, and / or other mechanical and / or electrical support structures. For example, the storage device 120 may comprise one or more direct inline memory module (DIMM) cards, one or more expansion cards and / or daughter cards, a solid state drive (SSD) or other hard drive device, and / or may have another memory and / or storage form factor. The storage device 120 may be integrated with and / or mounted on the motherboard of the host computing device 110, may be installed in a port and / or slot of the host computing device 110, may be installed on a dedicated storage appliance on a different host computing device 110 and / or network 115, may communicate with the host computing device 110 via an external bus (e.g., an external hard drive), or the like.

[0024] In one embodiment, the storage device 120 can be arranged on the memory bus of the processor 111 (e.g., on the same memory bus as the volatile memory 112, on a different memory bus from the volatile memory 112, instead of the volatile memory 112, etc.). In a further embodiment, the storage device 120 can be arranged on a peripheral bus of the host computing device 110, such as, but not limited to, a Non-Volatile Memory Express (NVMe) interface, a Serial Advanced Technology Attachment (SATA) bus, a Parallel Advanced Technology Attachment (PATA) bus, a Small Computer System Interface (SCSI) bus, a FireWire bus, a Fibre Channel connection, a Universal Serial Bus (USB), a Peripheral Component Interconnect Express (PCI Express or PCIe) bus, a PCIe Advanced Switching (PCIe-AS) bus, etc. In another embodiment, the storage device 120 can be arranged on a communication network 115, such as an Ethernet network, an InfiniBand network, SCSI RDMA over the network 115, a Storage Area Network (SAN), a Local Area Network (LAN), a Wide Area Network (WAN) such as the Internet, another wired and / or wireless network 115, etc.

[0025] The host computing device 110 may further include a computer-readable storage medium 114. The computer-readable storage medium 114 may include executable instructions configured to cause the host computing device 110 (e.g., the processor 111) to perform one or more of the methods disclosed herein. Additionally, or alternatively, the buffering component 150 may be embodied as one or more computer-readable instructions stored on the computer-readable storage medium 114.

[0026] The device driver and / or controller 126 may, in certain embodiments, present the logical address space 134 to the host client 116. As used herein, the logical address space 134 refers to a logical representation of memory resources. The logical address space 134 may include a plurality of logical addresses (e.g., ranges). As used herein, a logical address refers to any identifier for referencing a memory resource (e.g., data), and the identifier includes, but is not limited to, a logical block address (LBA), cylinder / head / sector (CHS), file name, object identifier, Universally Unique Identifier (UUID), Globally Unique Identifier (GUID), hash code, signature, index entry, range, extent, and the like.

[0027] The device driver of storage device 120 can maintain metadata 135 such as a logical-physical address mapping structure to map the logical addresses in the logical address space 134 to media storage locations on the storage device 120. The device driver may be configured to provide storage services to one or more host clients 116. The host clients 116 can include local clients operating on the host computing device 110 and / or remote clients 117 accessible via the network 115 and / or the communication interface 113. The host clients 116 can include, but are not limited to, an operating system, a file system, a database application, a server application, a kernel-level process, a user-level process, an application, etc.

[0028] In many embodiments, the host computing device 110 can include a plurality of virtual machines that can be instantiated or otherwise created based on user requests. As will be understood by those skilled in the art, the host computing device 110 can create a plurality of virtual machines configured as virtual hosts limited only by the available computing resources and / or demand. A hypervisor may be available to create, execute, and otherwise manage the plurality of virtual machines. Each virtual machine can include a plurality of virtual host clients similar to the host clients 116 that can utilize the storage system 102 to store and access data.

[0029] The device driver may be further communicatively coupled to one or more storage systems 102 that may include different types and configurations of storage devices 120, including but not limited to solid state storage devices, semiconductor storage devices, SAN storage resources, etc. One or more storage devices 120 may comprise one or more respective controllers 126 and non-volatile memory channels 122. The device driver can provide access to one or more storage devices 120 via any compatible protocol or interface 133, such as but not limited to SATA and PCIe. Metadata 135 may be used to manage and / or track data operations performed via protocol or interface 133. The logical address space 134 may include a plurality of logical addresses, each corresponding to a respective media location of one or more storage devices 120. The device driver can maintain metadata 135 that includes an any-to-any mapping between logical addresses and media locations.

[0030] The device driver may include, but is not limited to, a memory bus of the processor 111, a Peripheral Component Interconnect Express (PCI Express or PCIe) bus, a Serial Advanced Technology Attachment (ATA) bus, a Parallel ATA bus, a Small Computer System Interface (SCSI), FireWire, Fibre Channel, Universal Serial Bus (USB), a PCIe Advanced Switching (PCIe-AS) bus, a network 115, Infiniband, SCSI RDMA, etc., and further includes a storage device interface 139 configured to transfer data, commands, and / or queries to one or more storage devices 120 via a bus 125, and / or can communicate with it. The storage device interface 139 can communicate with one or more storage devices 120 using input-output control (IO-CTL) commands, IO-CTL command extensions, remote direct memory access, etc.

[0031] The communication interface 113 may include one or more network interfaces configured to communicatively couple the host computing device 110 and / or the controller 126 to the network 115 and / or one or more remote clients 117 (which can function as another host). The controller 126 may be part of and / or communicate with one or more storage devices 120. Although FIG. 1 shows a single storage device 120, the present disclosure is not limited in this regard and can be adapted to incorporate any number of storage devices 120.

[0032] The storage device 120 may include one or more non-volatile memory devices 123 of the non-volatile memory channel 122. The non-volatile memory devices 123 may include, but are not limited to, ReRAM, memristor memory, programmable metallization cell memory, phase change memory (PCM, PCME, PRAM, PCRAM, ovonic unified memory, chalcogenide RAM, or C-RAM), NAND flash memory (e.g., 2D NAND flash memory, 3D NAND flash memory), NOR flash memory, nano random access memory (nano RAM or NRAM), nanocrystal wire-based memory, silicon oxide-based sub-10 nanometer process memory, graphene memory, silicon-oxide-nitride-oxide-silicon (SONOS), programmable metallization cell (PMC), conductive-bridging RAM (CBRAM), magneto-resistive RAM (MRAM), magnetic storage media (e.g., hard disk, tape), optical storage media, etc. One or more non-volatile memory devices 123 of the non-volatile memory channel 122 may include, in certain embodiments, storage class memory (SCM) (e.g., write-in-place memory, etc.).

[0033] The non-volatile memory channel 122 may be referred to herein as a "memory medium," but in various embodiments, the non-volatile memory channel 122 may more generally be referred to as a non-volatile memory medium, non-volatile memory device, etc., and may include one or more non-volatile recording media capable of recording data. Further, the storage device 120 may include, in various embodiments, non-volatile recording devices, non-volatile memory arrays 129, a plurality of interconnected storage devices within an array, etc.

[0034] The non-volatile memory channel 122 may include one or more non-volatile memory devices 123, which may include, but are not limited to, chips, packages, planes, dies, etc. The controller 126 may be configured to manage data operations on the non-volatile memory channel 122 and may include one or more processors, programmable processors (e.g., FPGA), ASICs, microcontrollers, etc. In some embodiments, the controller 126 may be configured to store data in and / or read data from the non-volatile memory channel 122 and transfer data to / from the storage device 120, etc.

[0035] The controller 126 may be communicatively coupled to the non-volatile memory channel 122 via a bus 127. The bus 127 may include an I / O bus for communicating data with the non-volatile memory device 123. The bus 127 may further include a control bus for communicating addressing and other commands and control information to the non-volatile memory device 123. In some embodiments, the bus 127 may communicatively couple the non-volatile memory device 123 to the controller 126 in parallel. This parallel access may enable the non-volatile memory devices 123 to be managed as a group and form a non-volatile memory array 129. The non-volatile memory devices 123 may be divided into respective logical memory units (e.g., logical pages) and / or logical memory partitions (e.g., logical blocks). The logical memory units may be formed by logically combining each physical memory unit of the non-volatile memory devices 123.

[0036] In certain embodiments, the controller 126 may use the addresses of the word lines to organize blocks of word lines within the non-volatile memory device 123 such that the word lines can be logically organized into a monotonically increasing sequence (e.g., decoding and / or transforming the addresses of the word lines into a monotonically increasing sequence). In further embodiments, the word lines of a block within the non-volatile memory device 123 may be physically arranged in a monotonically increasing sequence of word line addresses, and consecutively addressed word lines may also be physically adjacent (e.g., WL0, WL1, WL2,...WLN).

[0037] The controller 126 may include and / or communicate with a device driver that runs on the host computing device 110. The device driver may provide storage services to the host client 116 via one or more interfaces 133. The device driver may further include a storage device interface 139 configured to transfer data, commands, and / or queries to the controller 126 via the bus 125 as described above.

[0038] Referring to FIG. 2, a schematic block diagram of an exemplary storage device 120 suitable for a computing SSD system that supports semantic search according to an embodiment of the present disclosure is shown. The controller 126 may include a front-end module 208 that interfaces with the host via a plurality of high-priority and low-priority communication channels, a back-end module 210 that interfaces with the non-volatile memory device 123, and various other modules that perform various functions of the storage device 120. In some examples, each module may be only a portion of a memory that includes instructions executable by a processor to implement the features of the corresponding module without the module including other hardware. Since each module includes at least some hardware even when the included hardware includes software, each module may be referred to interchangeably as a hardware module.

[0039] The controller 126 may include a buffer management / bus control module 214 that manages a buffer in a random access memory (RAM) 216 and controls internal bus arbitration for communication on an internal communication bus 217 of the controller 126. A read only memory (ROM) 218 can store and / or access system boot code. Although shown in FIG. 2 as being separately located from the controller 126, in other embodiments, one or both of the RAM 216 and the ROM 218 may be located within the controller 126. In still other embodiments, portions of the RAM 216 and the ROM 218 may be located both within and outside the controller 126. Further, in some implementations, the controller 126, the RAM 216, and the ROM 218 can be located on separate semiconductor dies. As described below, in one implementation, the submission queue and the completion queue can be stored in a controller memory buffer that can be housed in the RAM 216.

[0040] Additionally, the front end module 208 may include a host interface 220 and a physical layer interface 222 that provides an electrical interface between the host or the next level storage controller. The selection of the type of the host interface 220 may depend on the type of memory used. Exemplary types of the host interface 220 may include, but are not limited to, SATA, SATA Express, SAS, Fibre Channel, USB, PCIe, and NVMe. The host interface 220 can generally facilitate the transfer of data, control signals, and timing signals.

[0041] The backend module 210 may include an error correction engine 224 that encodes data bytes received from the host and decodes and corrects error for data bytes read from the non-volatile memory device 123 using an error-correction code (ECC). The backend module 210 may also include a command sequencer 226 that generates command sequences such as program, read, and erase command sequences to be transmitted to the non-volatile memory device 123. Additionally, the backend module 210 may include a redundant array of independent drive (RAID) module 228 that manages the generation of RAID parity and the recovery of failed data. The RAID parity may be used as an additional level of integrity protection for the data written to the storage device 120. In some cases, the RAID module 228 may be part of the error correction engine 224. The memory interface 230 provides command sequences to the non-volatile memory device 123 and receives status information from the non-volatile memory device 123. Along with the command sequences and status information, the data programmed to and read from the non-volatile memory device 123 may be communicated via the memory interface 230. The flash control layer 232 can control the overall operation of the backend module 210.

[0042] An additional module of the storage device 120 shown in FIG. 2 may include a media management layer 238 that performs wear leveling of the memory cells of the non-volatile memory device 123. The storage device 120 may also include other discrete components 240 such as an external electrical interface, external RAM, registers, capacitors, or other components that can interface with the controller 126. In an alternative embodiment, one or more of the RAID module 228, the media management layer 238, and the buffer management / bus control module 214 may be optional components that may not be necessary in the controller 126.

[0043] Finally, the controller 126 may also include intelligent memory array logic 234. In many embodiments, the intelligent memory array logic 234 may be configured to receive a query, extract context data from the received query, and determine a machine learning model for processing the query based on the extracted context data. The determination of the machine learning model for further processing the query can be performed by the controller 126. In some embodiments, the controller 126 can determine one or more appropriate machine learning models for further processing the query. The intelligent memory array logic 234 can receive data (e.g., a query), perform machine learning processes, calculations, comparisons, operations, or other processing to obtain and / or extract context data from the received query, generate one or more query vectors, and then determine one or more appropriate machine learning models for further processing the one or more query vectors, provided by hardware, firmware, software, electrical circuits and / or components, or any combination thereof. This can include accessing control data stored within the storage device 120, scanning the non-volatile memory device 123 for relevant pages, determining the relevant non-volatile memory device 123 associated with the generated query vector, and passing the query vector to the one or more determined relevant non-volatile memory devices 123.

[0044] Each of the non-volatile memory devices 123 may include feature data and feature vector stores related to one or more machine learning models and processes. The feature data may be evenly distributed within each of the one or more non-volatile memory devices 123. In some embodiments, the non-volatile memory devices 123 may be grouped into multiple sets, and each set of non-volatile memory devices 123 provides a specific machine learning model and executes a machine learning process based on the specific machine learning model. Further, each set of non-volatile memory devices 123 may comprise the same or different feature data and feature vector stores. Then, the controller 126 may pass the query vector to one or more related non-volatile memory devices 123.

[0045] Each of the non-volatile memory devices 123 of the memory array 129 may further include one or more machine learning processing units configured to process a query vector obtained from the controller 126 as a first input. The feature data within the one or more non-volatile memory devices 123 may be utilized as a second input, and the query vector and the feature data may be buffered in the memory device prior to processing by the one or more machine learning processing units. For example, the first and second inputs are processed to generate a comparison value. The one or more non-volatile memory devices 123 may be grouped into one or more memory sets based on similarity in the machine learning processing units or similarity in the feature data. The controller 126 may obtain either a comparison value or output data as a result of processing the first and second inputs from the one or more non-volatile memory devices 123. In some embodiments, one or more sets of non-volatile memory devices 123 may be grouped together based on a specific machine learning model or process, similarity in the machine learning processing units, or similarity in the feature data.

[0046] Referring to FIG. 3, a conceptual diagram of a page of memory cells, for example, configured in a NAND configuration, sensed or programmed in parallel according to an embodiment of the present disclosure is shown. FIG. 3 conceptually shows a bank of NAND strings 350 within the non-volatile memory device 123 of FIG. 1. A "page", such as page 360, is a group of memory cells that can be sensed or programmed in parallel. This is achieved in the peripheral circuitry by corresponding pages of sense amplifiers 310. The sensed results can be utilized in latches within corresponding sets of data latches 320. Each sense amplifier can be coupled to a NAND string, such as NAND string 350, via a bit line 336. For example, page 360 may be along a row and is sensed by a sense voltage applied to the control gates of the cells of the page commonly connected to word line WL3. Along each column, each memory cell, such as memory cell 311, may be accessible by a sense amplifier via a bit line 336. The data within data latch 320 can be toggled in or out between the memory controller 126 via data I / O bus 331.

[0047] The NAND string 350 can be a series of memory cells such as memory cell 311, which are daisy-chain connected by their sources and drains so as to form a source terminal and a drain terminal at their two ends respectively. A pair of select transistors S1, S2 can control the connection of the memory cell chain to an external source via the source terminal and the drain terminal of the NAND string respectively. In the memory array, when the source select transistor S1 is turned on, the source terminal is coupled to the source line 334. Similarly, when the drain select transistor S2 is turned on, the drain terminal of the NAND string is coupled to the bit line 336 of the memory array. Each memory cell 311 in the chain acts to accumulate charge. This has a charge storage element that accumulates a given amount of charge to represent the intended memory state. In many embodiments, the control gate in each memory cell can enable control for read and write operations. In many cases, the control gates of the corresponding memory cells in each row within a plurality of NAND strings can all be connected to the same word lines (such as WL0, WL1,...WLn342). Similarly, the control gates of each of the select transistors S1, S2 (accessed via the select lines 344SGS and SGD respectively) provide control access to the NAND string via their source terminals and drain terminals respectively.

[0048] The exemplary memory device described above includes physical page memory cells that store single-bit data. However, in most embodiments, each cell stores multi-bit data, and each physical page can have multiple data pages. Additionally, in further embodiments, a physical page can store one or more logical sectors of data. Typically, a host computing device 110 (see FIG. 1) operating with a disk operating system manages the storage of files by organizing the contents of the files in units of logical sectors, which are usually units of one or more 512-byte units. In some embodiments, a physical page can have 16 kB of memory cells that are sensed in parallel by corresponding 16 kB sense amplifiers via 16 kB of bit lines. An exemplary logical sector assigned by the host has a size of 2 kB of data. Thus, if the cells are each configured to store 1-bit of data (SLC), a physical page can store eight sectors. In the case of MLC, TLC, and QLC, and other dense structures, each cell can store 2, 3, 4, or more bits of data, and each physical page can store 16, 32, 64, or more logical sectors, depending on the structure being utilized. For the feature vector size and allocation in a computational SSD system that supports semantic search, the machine learning processing unit of the non-volatile memory device 123 can be configured to be adaptable to different feature vector sizes. As an example, a media file (e.g., an image, document, video, etc.) can be represented by any number of feature vectors. In some embodiments, the feature vectors of a media file can be 128 to 1024 features, and the size of the feature vector can be 256B to 2KB. Next, in the case of a 16 KB NAND flash page, when the page is read, 8 to 64 feature vectors are obtained.

[0049] One of the unique differences between flash memory and other types of memory is that the memory cells must be programmed from an erased state associated with no charge in the memory cell. This requires that the floating gate must first be emptied of charge before programming. Through programming, a desired amount of charge is returned to the floating gate. It may not support removing some of the charge from the floating gate to transition from a more programmed state to a less programmed state. Therefore, new data cannot overwrite existing data and must be written to a location that has not been previously written or erased. Additionally, it often takes a significant amount of time to erase all the charge from the floating gate. Thus, erasing on a per-cell, and even per-page, basis is cumbersome and inefficient. Therefore, in most embodiments, an array of memory cells is often divided into multiple blocks. As is common in many flash-based memory systems, a block is the unit of erasure. That is, each block contains the minimum number of memory cells that can be erased in a single operation. This, combined with the limited lifespan of the memory cells within the flash memory, increases the desire to limit the amount of erasing and programming that occurs within the storage device.

[0050] Referring to FIG. 4, a schematic block diagram of an exemplary computing SSD system that supports semantic search according to one embodiment of the present disclosure is shown. As described above, the controller 426 includes intelligent memory array logic 434 and a result aggregation unit 412. The controller 426 receives a semantic search query 401 and passes the query 401 to the intelligent memory array logic 434 for processing. In some embodiments, the controller 426 can receive data (e.g., a query), perform machine learning processes, calculations, comparisons, operations, or other processing to obtain and / or extract context data from the received query 401, and prepare the query 401 for the intelligent memory array logic 434. The content of the query 401 can include one or more texts, documents, images, audio, or other media. When receiving the query 401, the controller 426 processes the query 401 and extracts context data from the query 401. In addition to extracting context data from the query 401, other relevant context information for constructing the context data can be obtained by comparing the query 401 and / or the context data with one or more dictionaries, libraries, or databases, or combinations thereof. The relevant context information can be added to the query 401 or the context data extracted from the query 401 for further processing. Further, the context information may be stored locally in the controller 426, or may be remotely accessible to the controller 426 and the intelligent memory array logic 434. The controller 426 passes the context data to the intelligent memory array logic 434, and the intelligent memory array logic 434 determines a machine learning model for processing the query 401 based on the context data provided by the controller 426.

[0051] In some embodiments, the controller 426 may determine one or more appropriate machine learning models or machine learning processes within the intelligent memory array logic 434 to further process the query 401. The intelligent memory array logic 434 can receive data (e.g., a query), and perform machine learning processes, calculations, comparisons, operations, or other processing to obtain and / or extract context data from the received query 401.

[0052] The intelligent memory array logic 434 processes the query 401 and / or the context data from the query 401, and generates a query vector 403 based on the received query 401, the context data from the query 401, or a combination thereof. Further, the query vector 403 may be generated partially or wholly based on the determined machine learning model. The intelligent memory array logic 434 can determine one or more associated non-volatile memory devices 423 or a set of non-volatile memory devices 423 for processing the generated query vector 403. The determination may be based on the query 401, the context data from the query 401, the relevant context information from the query 401, the determined machine learning model, or any combination thereof. The intelligent memory array logic 434 then passes the generated query vector 403 to one or more associated non-volatile memory devices 423.

[0053] As described above, each of the one or more non-volatile memory devices 423 includes characteristic data and one or more machine learning processing units 421aa, 421ab,... 421an, 421ba, 421bb,... 421bn, 421na, 421nb,... 421nn, etc. (hereinafter, "machine learning processing unit 421"). Each machine learning processing unit 421 may include hardware, firmware, software, electrical circuits and / or components, or any combination thereof, for executing machine learning processes, calculations, comparisons, operations, or other processes to obtain results and / or values based on inputs from the characteristic data stored in the non-volatile memory device 423 and inputs provided by the intelligent memory array logic 434. Each of the one or more machine learning processing units 421 may be configured to process the passed query vector 403 as a first input and utilize the stored characteristic data as a second input. The one or more machine learning processing units 421 may then process the first and second inputs to generate results and / or values (e.g., comparison values) as output 405 to the controller 426. The intelligent memory array logic 434 may add characteristic metadata to the output 405 for further processing by the controller 426.

[0054] In some embodiments, the query vector 403 and the characteristic data may be buffered in the memory device before being processed by the one or more machine learning processing units 421. Further, the characteristic data may be evenly distributed within the one or more non-volatile memory devices 423. In some embodiments, each of the one or more non-volatile memory devices 423 may be grouped into one or more memory sets based on similarities in the machine learning processing units or in the characteristic data.

[0055] The result aggregation unit 412 obtains results and / or values as output 405 from all the machine learning processing units 421, executes machine learning processes, calculations, comparisons, operations, or other processing on the output 405, and for example, searches, collects, and presents the results from the output 405 in a report-based, or tabular format summarized as result 470. The result 470 can be passed to the host 110 for further processing.

[0056] Referring to FIG. 5, a schematic block diagram of an exemplary machine learning processing unit of a computational SSD system that supports semantic search according to an embodiment of the present disclosure is shown. In some embodiments, the machine learning processing unit 421 may be configured to trade off area for speed, and when the speed is limited, it may result in parallel processing with less hardware reuse or lower area efficiency. As an example, the machine learning processing unit 500 may be configured to be faster. The machine learning processing unit 500 may include hardware, firmware, software, electrical circuits and / or components, or any combination thereof for executing machine learning processes, calculations, comparisons, operations, or other processing. The machine learning processing unit 500 includes a module 571 that determines the feature vector size of the query vector V q and one or more distance calculation units 573a... 573n (hereinafter, "distance calculation unit 573") that calculate the distance (e.g., Euclidean / Hamming distance) between two input vectors (e.g., V q , V f(i) ), and receives the feature size extracted from the query vector V q by the module 571, and for example, compares the number of distance comparisons with the feature vector size and all the input feature vectors V f(i)... V f(i+n)By selecting the correct one as the final similarity score, it includes a control unit 574 that adapts to the variable feature vector size. Further, the machine learning processing unit 500 includes, for example, a module 575 having one or more machine learning operators, functions, or registers for performing machine learning processing that compares two distance inputs from a distance calculation unit 573 and outputs the smaller distance to a register. Then, the register of module 575 can provide the result, score, or value to the non-volatile memory device 423 and / or the controller 426.

[0057] Referring to FIG. 6, a schematic block diagram of an exemplary machine learning processing unit of a computing SSD system that supports semantic search according to an embodiment of the present disclosure is shown. In some embodiments, the machine learning processing unit 421 may be configured to trade off speed against area, where the area is limited, resulting in hardware reuse or sequential processing at a lower speed. As an example, the machine learning processing unit 600 may be configured to be highly area-efficient. The machine learning processing unit 600 may include hardware, firmware, software, electrical circuits and / or components, or any combination thereof for performing machine learning processes, calculations, comparisons, operations, or other processing. The machine learning processing unit 600 includes a module 671 that determines the feature vector size of the query vector V q and a distance calculation unit 673 for calculating the distance (e.g., Euclidean / Hamming distance) between two input vectors (e.g., V q , V f ). The module 671 receives the feature size extracted from the query vector V q and, for example, compares the number of distance comparisons with the feature vector size and all input feature vectors V fBy selecting the correct one as the final similarity score, it includes a control unit 674 that adapts to the variable feature vector size. Further, the machine learning processing unit 600 includes, for example, a module 675 having one or more machine learning operators, functions, or registers for performing machine learning processing that compares two distance inputs from the distance calculation unit 673 and outputs the smaller distance to a register. Then, the register of module 675 can provide the result, score, or value to the non-volatile memory device 423 and / or the controller 426.

[0058] Referring to FIG. 7, a flowchart showing a process 700 for utilizing an exemplary computing SSD system that supports semantic search according to an embodiment of the present disclosure is shown. Process 700 can start by receiving a query (block 710). In many embodiments, this occurs when the host or SSD controller receives a query for semantic search. In some embodiments, the user provides a semantic search request that is passed to the host or SSD controller. Accordingly, various embodiments of the controller can instruct a plurality of processes, such as a process within the controller 426 (see FIG. 4) for a computing SSD system that supports semantic search, to execute the semantic search process 700.

[0059] Process 700 can extract context data from the received query (block 715). In some embodiments, the context data may be obtained directly from the query. As described above, in some embodiments, context information related to the query can be obtained via other resources to help construct the context data of the query. The relevant context information for constructing the context data can be obtained by comparing the query and / or the context data with one or more machine learning processes having access to a dictionary, library, or database, or a combination thereof.

[0060] Process 700 can determine a machine learning model based on the extracted context data (block 720). As is known in the art, various machine learning models, for example, convolutional neural network (CNN), computer vision, supervised learning, semi-supervised learning, unsupervised learning, reinforcement learning (RL), deep learning (DL), artificial neural network (ANN), computer vision, speech recognition, natural language processing, machine translation, etc. can be used to extract and process features from a query as desired.

[0061] Once a machine learning model is determined based on the extracted context data, process 700 can use the machine learning model to process a query (block 725). By way of example and not limitation, when processing a query, various features can be obtained from the machine learning model, the features can be processed to create a vector, and the vector can be prepared for machine learning processes such as similarity search.

[0062] Based on the received query and the features obtained from the determined machine learning model, process 700 can generate a query vector (block 730) and can determine one or more associated non-volatile memory devices associated with the generated query vector (block 735). As described above, non-volatile memory device 123 can be part of any suitable storage device 120 that can have various different types of memory devices. By way of example and not limitation, the storage device can comprise a plurality of non-volatile NAND memory devices including single-level cell NAND memory and quad-level cell NAND memory.

[0063] Process 700 can pass the query vector to one or more determined associated non-volatile memory devices (block 740). Process 700 can process the passed query vector as a first input (block 745) and utilize the feature data on one or more associated non-volatile memory devices as a second input (block 750). As described above, one or more determined associated non-volatile memory devices have a machine learning processing unit for processing the query vector as a first input and the feature data stored on one or more non-volatile memory devices as a second input.

[0064] Process 700 can process the first and second inputs to generate a comparison value (block 755). In certain embodiments, process 700 may utilize a super-dimensional vector / vector symbol architecture to enable on-die calculations without error correction code (ECC). In this way, features can be stored as super-dimensional vectors rather than dense feature vectors that may be noise-resistant.

[0065] Referring to FIG. 8, a flowchart showing a process 800 for utilizing an exemplary computational SSD system that supports semantic search according to an embodiment of the present disclosure is shown. Process 800 can begin by processing a query using a machine learning model (block 810). In many embodiments, this occurs when a host or SSD controller receives a query for semantic search. In some embodiments, a user may provide a semantic search request that can be passed to the host or SSD controller. Accordingly, various embodiments of the controller can instruct a plurality of processes, such as a process within controller 426 (see FIG. 4) of a computational SSD system that supports semantic search, to execute semantic search process 800.

[0066] Process 800 can generate a query vector (block 815). When generating the query vector, the context data can be directly obtained from the query and the processing of the query by the machine learning model. By way of non-limiting example, when processing the query, various features can be obtained from the machine learning model, the features can be processed to create a vector, and prepared for machine learning processes such as similarity search. As is known in the art, various machine learning models, for example, convolutional neural network (CNN), computer vision, supervised learning, semi-supervised learning, unsupervised learning, reinforcement learning (RL), deep learning (DL), artificial neural network (ANN), computer vision, speech recognition, natural language processing, machine translation, etc. can be used to extract and process features from the query as desired.

[0067] Process 800 can pass the query vector to one or more machine learning processing units as a first input (block 820), and the machine learning processing units further utilize the feature data stored on one or more associated non-volatile memory devices as a second input (block 825). Process 800 can utilize one or more machine learning processing units to process the first and second inputs to determine output data retrieved from one or more non-volatile memory devices (block 830). Process 800 can assign feature metadata to the output data using intelligent memory array logic (block 835). In certain embodiments, process 800 can utilize a superdimensional vector / vector symbol architecture to enable on-die calculations without an error correction code (ECC). In this way, the features can be stored as superdimensional vectors rather than dense feature vectors that may be noise-resistant.

[0068] Referring to FIGS. 7-8, the hyperdimensional calculations for computing the similarity between data can be realized by, for example, three operations including addition, multiplication, and substitution. Hyperdimensional computing can be inherently robust because information is evenly distributed across all bits of the hypervector and provides high-speed learning ability, high energy efficiency, and acceptable accuracy in learning and classification tasks.

[0069] Referring to FIG. 9, a schematic block diagram of an exemplary machine learning process according to an embodiment of the present disclosure is shown. As shown, the input for the machine learning system 900 can include a semantic search request by one or more queries 901. The query 901 can include text, images, documents, audio, or other media. The query 901 can be passed to an intelligent memory array logic 934 that performs machine learning processes, calculations, comparisons, operations, or other processing to obtain and / or extract context data from the received query 901. The intelligent memory array logic 934 determines context data from the query 901, assigns features to the query 901, and forms a query vector 903 as a first input for the machine learning processing unit (MPLU) 921 of the machine learning system 900.

[0070] The second input of the machine learning system 900 includes data stored on an SSD device such as one or more media files 902. The media file 902 can include text, images, documents, audio, or other media. The stored media file 902 can be assigned feature data 904 by the intelligent memory array logic 934, and each media file 902 can be stored on one or more non-volatile memory devices 123. The feature data 904 can be scanned by the MPLU 921 from one or more non-volatile memory devices 123 as a second input for comparison with the query vector 903.

[0071] MPLU921 obtains output data 905 according to an iterative process of retrieving, extracting, calculating, and comparing the feature data 904 of the stored media file 902 with the query vector 903. The output data 905 can be provided as a result of one or more operations by MPLU921. In some embodiments, the output data 905 can be a similarity score of the comparison between the query vector 903 and one or more sets of the feature data 904. Then, the output data 905 can be passed to the controller 126 for further processing.

[0072] In some embodiments, the controller 426 may determine one or more appropriate machine learning models or machine learning processes within the intelligent memory array logic 934 to further process the query 401. The intelligent memory array logic 934 can receive data (e.g., a query), and execute machine learning processes, calculations, comparisons, operations, or other processing to obtain and / or extract context data from the received query 901.

[0073] As described above, each machine learning processing unit 921 may include hardware, firmware, software, electrical circuits and / or components, or any combination thereof, to execute machine learning processes, calculations, comparisons, operations, or other processing to obtain results and / or values based on inputs from the feature data stored on the non-volatile memory device 123 and inputs provided by the intelligent memory array logic 934.

[0074] In some embodiments, the query vector 903 and the feature data 904 can be buffered in a memory device before being processed by one or more machine learning processing units 921. Further, the feature data 904 can be evenly distributed within one or more non-volatile memory devices 123. In some embodiments, each of the one or more non-volatile memory devices 123 can be grouped into one or more memory sets based on similarity in the machine learning processing unit or similarity in the feature data.

[0075] Referring to FIG. 10A, a schematic block diagram of an exemplary error correction system according to an embodiment of the present disclosure is shown. An error correction system 1000 is shown for integrating an ECC function for encoding and decoding dense feature vectors.

[0076] The exemplary error correction system 1000 may include a host 1010, an SSD controller 1026, a NAND flash memory array 1029, and a volatile memory 1016 (e.g., DRAM). The error correction system 1000 may further include a host interface 1020, an embedded processor 1011, an ECC memory module 1024 having at least one encoding circuit 1024a and a decoding circuit 1024b and processing, an SSD controller 1026, and a flash interface 1032 that is partially or wholly included in, or communicates separately with, or is integrated therein. The embedded processor 1011 may communicate with the volatile memory 1016. The volatile memory 1016 may be integrated within the SSD controller 1026 and / or communicate with the embedded processor 1011.

[0077] The NAND flash memory array 1029 may include a plurality of NAND dies 1023, and each NAND die has a memory circuit and logic 1027 including an ECC memory module 1028 and a machine learning processing unit 1021. In some embodiments, an exemplary error correction system 1000 can operate in a storage mode where the SSD controller 1026 receives file data and one or more instructions from the host 1010 and stores the data on the NAND die 1023. Further, in the storage mode, the SSD controller 1026 can receive one or more instructions from the host 1010 to access the same or different file data stored on the NAND die 1023. As shown, in a write operation, the ECC memory module 1024 can encode an error correction code in the file data to facilitate storage, retrieval, and processing of error-free file data on the NAND die 1023. Further, in a read operation, to retrieve error-free file data for transmission to the host 1010, the ECC memory module 1024 can use the encoded error correction code for the file data to process and / or remove file data errors received from the file data stored on the NAND die 1023. In some embodiments, file data errors in the storage mode can be introduced, for example, but not limited to, at the flash interface 1032 or at any file data processing stage within the SSD controller 1026 from the receipt of file data and / or instructions by the host 1010 to the storage of the file data and / or instructions on the NAND flash memory array 1029.

[0078] In some embodiments, the exemplary error correction system 1000 can operate in a computing mode in which the NAND flash memory array 1029 utilizes the memory circuit and logic 1027 with the ECC memory module 1028 to encode, decode, verify error-free feature data (e.g., dense feature vectors) for comparison with query vectors to obtain scores and provide them to the machine learning processing unit 1021. Thus, each NAND die 1023 includes an ECC function for encoding and decoding feature data error correction codes to provide error-free feature data or error-free dense feature vectors. Thus, the NAND die 1023 stores at least two types of data, namely raw file data and feature data corresponding to the file (e.g., feature vectors). The memory circuit and logic 1027 can perform feature data management to maximize feature data processing efficiency, e.g., parallel processing of feature data and on the die. In some embodiments, similar to the machine learning processing unit 1021 disposed in each NAND die 1023, the memory circuit and logic 1027 may also be distributed across the NAND flash memory array 1029 as a single unit for each NAND die 1023 to process the file data and feature data stored in the corresponding NAND die 1023. Further, in some embodiments, the feature data may be evenly divided and distributed across each NAND die 1023 of the NAND flash memory array 1029. Further, the feature vectors can be repositioned within the same NAND die 1023 during storage management operations such as garbage collection, for example.

[0079] FIG. 10B is a chart of the exemplary error correction system of FIG. 10A, according to one embodiment of the present disclosure. To ensure data integrity, data scrubbing can be performed on the feature vectors (e.g., using FV scrubbing) and file data (e.g., using conventional data scrubbing). As shown, the number of bit flips per unit time for FV data scrubbing on the feature data can be performed at a higher frequency than the conventional data scrubbing on the file data, or a fraction of the frequency of the conventional data scrubbing, e.g., 1 / 6, 1 / 4, 1 / 3, 1 / 2, 2 / 3 of the conventional scrubbing time. Further, the FV scrubbing and the conventional scrubbing can be triggered when a predetermined threshold of bit flips is reached. For example, when the FV threshold is reached, FV scrubbing can be performed, and when the conventional threshold is reached, conventional scrubbing can be performed. In some embodiments, a combination of time / frequency and threshold can be used to perform at least one of the FV scrubbing or the conventional scrubbing. As shown, the FV data scrubbing can be performed in 2 / 3 of the time of the conventional scrubbing of the file data.

[0080] FIG. 10C is a table of the exemplary error correction system of FIG. 10A according to one embodiment of the present disclosure. Error correction of file data and feature data is shown as above. The file data and feature data may also undergo data scrubbing to reduce the possibility that a single correctable error accumulates into a large number of uncorrectable errors. As shown, the error correction engine may be used to verify file data, and conventional data scrubbing may be used for wear leveling to maintain data integrity of file data in the memory mode. Conventional data scrubbing may include error correction techniques that use one or more background tasks to periodically inspect main memory or storage for errors and then correct the detected errors using redundant data in the form of different checksums or copies of the data. The file data is purified through a process of detecting and correcting corrupted or inaccurate records from a record set, table, or database, identifying incomplete, incorrect, inaccurate, or irrelevant portions of the data, and then replacing, modifying, or deleting the wrong, corrupted, or rough data.

[0081] As an example, an exemplary error correction system 1000 may require on-die ECC when in a computational mode for feature vector data (e.g., dense feature vectors). In some embodiments, a weaker ECC such as Bose Chaudhuri Hocquenghem (BCH) may be used to reduce the area and latency overhead of on-die ECC. In some embodiments, other types of on-die ECC, such as Low Density Parity Check (LDPC), may be used. Additionally, on-die ECC may be further used for file data error correction. In some embodiments, feature vector data (e.g., dense feature vectors) may be stored in their own memory blocks separate from file data to facilitate easier data management. In some embodiments, feature vector data (e.g., dense feature vectors) may be stored together with related file data in the same memory block to facilitate faster computation time and data access. Data scrubbing of file data can use the conventional data scrubbing described above, while feature data can use a separate data scrubbing algorithm to clean up, for example, incorrect, inaccurate, incomplete, empty, or duplicate feature vector data.

[0082] In some embodiments, when the SSD storage device 120 is operating in semantic search mode, requests from the host 1010 to service read / write requests may be ignored, delayed, executed after some semantic search operations are completed, or queued to be executed after all semantic search operations are completed, or any combination thereof. Thus, the volatile memory 1016 may be used to load feature data related to the data management data structure instead of the logical-to-physical (L2P) table for file data access.

[0083] Regarding feature vector data management, the application needs to maintain at least one database within the SSD storage device 120 to perform fast semantic search. For example, the database may include file names, hash values, and attributes of the feature vector data. Further, at least one file can include the attribute file name and hash value that are loaded into the application when the application is launched. In some embodiments, the information of the hash value and the feature vector data can be included in one or more data structures (e.g., Table X) to index one or more feature vectors. In some embodiments, Table X may be stored in the SSD storage device 120 firmware, and Table X may include Addr, hash value, and attributes of the feature vector data. Further, according to the hash value returned from the semantic search, the application can retrieve the file name from the search results. In some embodiments, the size of Table X may be proportional to the number of files stored in the SSD storage device 120. Further, in some embodiments, Addr may be in the form of [Die, Plane, Block, Page, Segment], where the number of bits represents Segment.

[0084]

Number

[0085] In some embodiments, the hash value feature vector may be a concatenation of the hash value and its feature vector to simplify data management. Further, when the SSD storage device 120 receives a search request (e.g., a query), the SSD controller 1026 needs to generate special search commands and send them to each NAND flash die 1023. The SSD controller 1026 can notify the NAND flash die 1023 which block contains the feature vector data, and the NAND flash die 1023 should perform calculations accordingly. Further, each NAND flash die 1023 can execute a function of top_k hashvalue=Loop(search(query_vec,plane#,block#,last_page#,last segment#)).

[0086] FIG. 11A is a schematic block diagram of an exemplary error correction system according to an embodiment of the present disclosure. As shown, an error correction system 1100 for integrating an ECC function for encoding and decoding a Hyper Dimensional Vector (HDV) or a Vector Symbolic Architecture (VSA), where on-die ECC can be optional. In the present disclosure, HDV and VSA can be used interchangeably, but hereinafter HDV is used. In some embodiments, integrating ECC logic, circuitry, or processing within each NAND die 1123 can be prohibitively costly. For example, incorporating an ECC memory module 1128 into each NAND die 1123 can be prohibitively expensive or occupy too much area on the NAND flash memory array 1129. Instead, HDV / VSA can be stored for each file, and such context / semantic representation is robust to noise, can maintain similarity, can be read from the NAND flash die 1123, and can be calculated without error correction.

[0087] An exemplary error correction system 1100 may include a host 1110, an SSD controller 1126, a NAND flash memory array 1129, and volatile memory 1116 (e.g., DRAM). The error correction system 1100 may further include a host interface 1120, an embedded processor 1111, at least one encoding circuit 1124a and decoding circuit 1124b and an ECC memory module 1124 having processing, and a flash interface 1132 that is partially or wholly included in, or otherwise communicates with, or is integrated therein, with the SSD controller 1126. The embedded processor 1111 can communicate with the volatile memory 1116. The volatile memory 1116 may be integrated within the SSD controller 1126 and / or communicate with the embedded processor 1111.

[0088] The NAND flash memory array 1129 may include a plurality of NAND dies 1123, each NAND die having a memory circuit and logic 1127 that includes a machine learning processing unit 1121. In some embodiments, the exemplary error correction system 1100 can operate in a storage mode where the SSD controller 1126 receives file data and one or more instructions from the host 1110 and stores the data on the NAND die 1123. Further, in the storage mode, the SSD controller 1126 can receive one or more instructions from the host 1110 to access the same or different file data stored on the NAND die 1123. As shown, in a write operation, the ECC memory module 1124 can encode an error correction code in the file data to facilitate storage, retrieval, and processing of error-free file data on the NAND die 1123. Further, in a read operation, to retrieve error-free file data for transmission to the host 1110, the ECC memory module 1124 can use the encoded error correction code for the file data to process and / or remove file data errors received from the file data stored on the NAND die 1123. In some embodiments, file data errors in the storage mode can be introduced, for example, but not limited to, at the flash interface 1132 or at any file data processing stage within the SSD controller 1126 from the receipt of file data and / or instructions by the host 1110 to the storage of file data and / or instructions on the NAND flash memory array 1129.

[0089] In some embodiments, an exemplary error correction system 1100 can operate in a computing mode where a NAND flash memory array 1129 utilizes the memory circuitry and logic 1127 of an SSD controller 1126 and an ECC memory module 1128 to encode, decode, verify error-free feature data (e.g., HDV / VSA vectors), and provide it to a machine learning processing unit 1121 to obtain a score by comparing with a query vector. Each NAND die 1123 utilizes two copies of the HDV vector, the ECC function from the ECC memory module 1128, and an exemplary error correction method described below to encode and decode a feature data error correction code to provide error-free feature data, e.g., an error-free HDV vector.

[0090] In the exemplary error correction system 1100, the NAND die 1123 stores at least two types of data, namely raw file data and HDV / VSA vectors. Feature data corresponding to the file data (e.g., HDV vectors) can be very robust to noise, maintain similarity, be read from the NAND die 1123, and be computed without error correction. In some embodiments, two copies of the HDV vector can be stored in each NAND die 1123 corresponding to the file data. The first copy is the HDV without ECC coding, and the second copy is the HDV with SSD controller 1026 ECC coding. As shown, the first copy of the HDV vector without ECC coding can be used by the machine learning processing unit 1121. However, since there is an upper limit to its resilience, the blocks storing the HDV vectors also need to be wear-leveled.

[0091] The second copy of the HDV vector with ECC coding can be used to correct errors within the first copy of the HDV vector. If a scrubbing algorithm discovers that the bit error rate of the HDV copy with ECC coding exceeds a predetermined threshold, the corrected HDV without ECC coding can be stored on the memory circuit and logic 1127, and the HDV copy with ECC coding can be relocated on the memory circuit and logic 1127. In some embodiments, depending on the application, if the VSA vector is very long and ultra-robust, a separate scrubbing algorithm for the second copy of the HDV vector with ECC coding may not be required; otherwise, a separate scrubbing algorithm can be implemented for the HDV vector with ECC coding.

[0092] The memory circuit and logic 1127 can perform feature data management to maximize feature data processing efficiency, e.g., parallel processing of feature data and dies. In some embodiments, the memory circuit and logic 1127 can also be distributed within the NAND flash memory array 1129 as a single unit for each NAND die 1123 to process file data and HDV vector data stored on the corresponding NAND die 1123. Further, in some embodiments, the feature data may be evenly divided and distributed across each NAND die 1123 of the NAND flash memory array 1129. Additionally, the feature vector data can be relocated within the same NAND die 1123, e.g., during storage management operations such as garbage collection.

[0093] FIG. 11B is a chart of the exemplary error correction system of FIG. 11A according to one embodiment of the present disclosure. To ensure data integrity, data scrubbing can be performed on the HDV vector (e.g., using HDV scrubbing) and file data (e.g., using conventional data scrubbing). As shown, the number of bit flips per unit time for HDV data scrubbing on the feature data can be performed at a higher frequency than conventional data scrubbing on the file data, or a fraction of the frequency of conventional data scrubbing, e.g., 1 / 6, 1 / 4, 1 / 3, 1 / 2, 2 / 3 of the conventional scrubbing time. Further, HDV scrubbing and conventional scrubbing can be triggered when a predetermined threshold of bit flips is reached. For example, when the HDV threshold is reached, HDV scrubbing can be performed, and when the conventional threshold is reached, conventional scrubbing can be performed. In some embodiments, a combination of time / frequency and threshold can be used to perform at least one of HDV scrubbing or conventional scrubbing. As shown, HDV data scrubbing can be performed at 2 / 3 of the time of conventional scrubbing of the file data.

[0094] FIG. 11C is a table of the exemplary error correction system of FIG. 11A according to one embodiment of the present disclosure. Error correction of file data and feature data is shown as above. The file data and the feature data may also undergo data scrubbing to reduce the possibility that a single correctable error accumulates into a number of uncorrectable errors. As shown, the ECC of the SSD controller 1126 can be used to verify the file data, and conventional data scrubbing, for example, low-density parity check (LDPC) for wear leveling to maintain data integrity of the file data in the memory mode, can be used. Conventional data scrubbing may include error correction techniques that use one or more background tasks to periodically check the main memory or storage for errors and then correct the detected errors using redundant data in the form of different checksums or copies of the data. The file data can be purified through a process of detecting and correcting corrupted or inaccurate records from a record set, table, or database, identifying incomplete, incorrect, inaccurate, or irrelevant portions of the data, and then replacing, modifying, or deleting the wrong, corrupted, or rough data.

[0095] As an example, when the exemplary error correction system 1100 is in the compute mode, conventional data scrubbing may not be required for the first copy of the HDV vector (i.e., HDV without ECC coding, the HDV vector). In some embodiments, conventional data scrubbing may be used as needed for the first copy of the HDV vector to maintain data integrity of the feature data of the first copy of the HDV vector without ECC coding. Further, in the compute mode, the ECC of the SSD controller 1126 may be used to verify the feature data of the second copy of the HDV vector (i.e., HDV-COPY, the HDV vector with ECC coding), for example, using low density parity check (LDPC). Further, at least one of conventional data scrubbing or HDV scrubbing using a feature data scrubbing algorithm may be used to clean up, for example, incorrect, wrong, incomplete, empty, or duplicate feature vector data.

[0096] In some embodiments, when the SSD storage device 120 is operating in the semantic search mode, requests from the host 1110 to service read / write requests may be ignored, delayed, executed after some semantic search operations are completed, or queued to be executed after all semantic search operations are completed, or any combination thereof. Thus, the volatile memory 1116 may be used to load feature data related to the data management data structure rather than the logical to physical (L2P) table for file data access. Further, HDV vector data management may be performed as described above.

[0097] In some embodiments, the hash value feature vector may be a concatenation of the hash value and its feature vector in order to simplify data management. Further, when the SSD storage device 120 receives a search request (e.g., a query), the SSD controller 1126 needs to generate special search commands and send them to each NAND flash die 1123. The SSD controller 1126 can notify the NAND flash die 1123 which block contains the feature vector data, and the NAND flash die 1123 should perform calculations accordingly. Further, each NAND flash die 1123 can execute a function of top_k hashvalue=Loop(search(query_vec,plane#,block#,last_page#,last segment#)).

[0098] FIG. 12 is a schematic block diagram of an exemplary error correction mode and storage mode of a computing SSD system 1200 that supports semantic search using an error correction system according to an embodiment of the present disclosure. As shown, the exemplary computing SSD system 1200 supports semantic search in the computing mode and raw file data processing in the storage mode with the host 1210 providing respective instructions and file data or feature data to the SSD storage device 1220. In some embodiments, when the non-volatile memory 1316 is limited, the SSD storage device 1320 can switch between the computing mode and the storage mode to achieve retrieval of the result file. Further, requests from the host 1310 to service read / write requests may be ignored, delayed, queued to be executed after some semantic search operations are completed, or after all semantic search operations are completed, or any combination thereof. Thus, the volatile memory 1116 can be used to load feature data related to the data management data structure rather than the logical-to-physical (L2P) table for file data access.

[0099] Process 1: In the computing mode, the non-volatile memory 1216 (e.g., DRAM) can be used by the table X, and the query 1201 is sent to the SSD storage device 1220, processed (e.g., by the machine learning processing unit 1021 / 1121), and then the result or hash 1203 can be returned to the host 1210 for the semantic search request.

[0100] Process 2: In the storage mode, the non-volatile memory 1216 (e.g., DRAM) can be used for L2P table mapping, and the data request 1205 is sent to the SSD storage device 1220, processed (e.g., by the SSD controller 1026 / 1126), and then the file data 1207 can be returned to the host 1210 for the file search / read / write / access request.

[0101] FIG. 13 is a schematic block diagram of another exemplary error correction mode and storage mode of a computing SSD system 1300 that supports semantic search using an error correction system according to an embodiment of the present disclosure. In some embodiments, if the non-volatile memory 1316 can be sufficient for both the storage mode and the computing mode, the storage operation and the computing operation can be performed simultaneously or in parallel without requiring separate modes.

[0102] For example, non-volatile memory 1316 (e.g., DRAM) can be used by table X, query 1301 is sent to SSD storage device 1320, processed (e.g., by machine learning processing unit 1021 / 1121), and then the result or hash 1303 can be returned to host 1310 for semantic search requests. Further, in parallel, non-volatile memory 1316 (e.g., DRAM) can be used for L2P table mapping, data request 1305 is sent to SSD storage device 1320, processed (e.g., by SSD controller 1026 / 1126), and then file data 1307 can be returned to host 1310 for file search / read / write / access requests. Further, in some embodiments, when a hyperscale or storage array is used, semantic search and file data can be executed in parallel, and the file data can be retrieved from a conventional SSD using a file replica from SSD storage device 1320.

[0103] Referring to FIG. 14, a flowchart showing a process 1400 for utilizing an exemplary error correction method in a computational SSD system that supports semantic search according to one embodiment of the present disclosure is shown. Process 1400 can begin by storing raw file data in one or more memory structures (block 1410). Process 1400 can store feature data corresponding to the raw file data in the corresponding one or more memory structures (block 1415). Process 1400 can load feature data from at least one of the one or more memory structures, where the feature data includes at least one of a dense feature vector or a hyperdimensional vector (block 1420). Process 1400 can generate one or more databases from the feature data (block 1425). Process 1400 can load raw file data from at least one of the one or more memory structures (block 1430). Process 1400 can utilize an error correction algorithm for storing and accessing the feature data (block 1435).

[0104] Referring to FIG. 15, a flowchart is shown illustrating a process 1500 for utilizing an exemplary error correction method in a computing SSD system that supports semantic search according to one embodiment of the present disclosure. Process 1500 can begin by processing a query to retrieve raw file data stored on one or more memory structures using a machine learning model (block 1510). Process 1500 can generate one or more databases from feature data stored in one or more memory structures, where the feature data includes at least one of a dense feature vector or a hyperdimensional vector (block 1515). Process 1500 can load one or more databases and raw file data from one or more memory structures (block 1525). Process 1500 can process the raw file data and one or more databases to obtain a response to the query (block 1525).

[0105] The information shown and described in detail herein can fully achieve the above objects of the present disclosure and the currently preferred embodiments of the present disclosure, and thus represents the subject matter broadly contemplated by the present disclosure. The scope of the present disclosure should be construed to fully encompass other embodiments that may become apparent to those skilled in the art, and thus should not be limited by anything other than the appended claims. Any reference to an element in the singular is not intended to mean "one and only one" unless explicitly stated as such, but rather is intended to mean "one or more." All structural and functional equivalents to the elements of the above-described preferred embodiments and additional embodiments contemplated by those skilled in the art are expressly incorporated herein by reference and are intended to be encompassed by the present claims.

[0106] Furthermore, there is no requirement for the system or method to address every problem that is required to be solved by the present disclosure, and the solutions to such problems are encompassed by the claims of this patent. Additionally, elements, components, or method steps in the present disclosure are not intended to be dedicated to the public, whether or not such elements, components, or method steps are explicitly recited in the claims. As will be apparent to those skilled in the art, various changes and modifications can be made in the form, materials, workpieces, and details of the manufacturing materials without departing from the spirit and scope of the present disclosure as set forth in the appended claims, and these are also encompassed by the present disclosure.

Claims

1. A device comprising: a processor; a memory array comprising a plurality of memory devices; a controller communicatively coupled to the memory array; one or more memory structures configured to store raw file data and feature data corresponding to the raw file data; the feature data includes at least one of a dense feature vector or a hyperdimensional vector; each of the one or more memory structures stores the feature data and utilizes an error correction algorithm for accessing the feature data; each of the one or more memory structures comprises one or more machine learning processing units configured to process the feature data.

2. The device of claim 1, wherein the feature data includes a dense feature vector based on error-free data.

3. The device of claim 2, wherein the dense feature vector is stored in a block separate from the raw file data.

4. The device of claim 1, wherein the feature data is evenly distributed within the one or more memory structures.

5. The device of claim 1, wherein the error correction algorithm is implemented as an on-die error correction code (ECC), and each of the one or more memory structures has ECC.

6. The device of claim 1, wherein the feature data includes a hyperdimensional vector, at least two copies of the hyperdimensional vector are stored in each of the one or more memory structures, a first copy of the hyperdimensional vector is stored without using an error correction algorithm, and a second copy of the hyperdimensional vector is stored using an error correction algorithm.

7. The device of claim 6, wherein the first copy of the hyperdimensional vector is provided to the machine learning processing unit, and the second copy of the hyperdimensional vector is provided to correct errors in the first copy.

8. A method comprising: storing raw file data in one or more memory structures; storing feature data corresponding to the raw file data in the corresponding one or more memory structures; loading, from at least one of the one or more memory structures, feature data that includes at least one of a dense feature vector or a hyperdimensional vector; generating one or more databases from the feature data. ​ loading the raw file data from at least one of the one or more memory structures; storing the feature data and utilizing an error correction algorithm for accessing the feature data, a method comprising the steps of: **Claim 9** The method according to claim 8, wherein the feature data includes a dense feature vector based on error-free data. **Claim 10** The method according to claim 9, wherein storing the feature data includes storing the dense feature vector in a block separate from the raw file data. **Claim 11** The method according to claim 8, further comprising evenly dispersing the feature data within the one or more memory structures. **Claim 12** The method according to claim 8, wherein the error correction algorithm is implemented as an on-die error correction code (ECC), and each of the one or more memory structures has an ECC. **Claim 13** The method according to claim 8, wherein the feature data includes a super-dimensional vector, and storing the feature data includes storing at least two copies of the super-dimensional vector in each of the one or more memory structures, wherein the first copy of the super-dimensional vector is stored without using an error correction algorithm, and the second copy is stored using an error correction algorithm. **Claim 14** The method according to claim 8, wherein each of the one or more memory structures includes a machine learning processing unit, the first copy of the super-dimensional vector is provided to the machine learning processing unit, and the second copy of the super-dimensional vector is provided to correct errors in the first copy. **Claim 15** A method comprising: processing a query for retrieving raw file data stored on one or more memory structures using a machine learning model; generating one or more databases from feature data stored on the one or more memory structures, wherein the feature data includes at least one of a dense feature vector or a super-dimensional vector; loading the one or more databases and the raw file data from the one or more memory structures; and processing the raw file data and the one or more databases to obtain a response to the query. **Claim 16** ​ The method according to claim 15, wherein the controller switches between a computing mode for processing the query to obtain the response and a storage mode for retrieving one or more files from the one or more memory structures. **Claim 17** The method according to claim 16, wherein the controller utilizes a computing mapping table having volatile memory configured for table X in the computing mode and a storage mapping table having the volatile memory configured for a logical-to-physical (L2P) table in the storage mode. **Claim 18** The method according to claim 15, wherein an error correction algorithm is utilized to process the raw file data and the one or more databases to obtain the response to the query. **Claim 19** The method according to claim 15, wherein the controller determines which blocks on the one or more memory structures contain the feature data and whether they should be processed to obtain the response to the query. **Claim 20** The method according to claim 15, wherein each of the one or more memory structures includes a machine learning processing unit for processing the raw file data and one or more databases to obtain the response to the query.

Citation Information

Patent Citations

  • Method and apparatus for supporting machine learning algorithms and data pattern matching in ethernet SSD

    US20180189635A1

  • Object recognition devices, electronic devices and methods of recognizing objects

    US20200034666A1