Flash translation layer with hierarchical security

The Flash Translation Layer in computer systems uses a hierarchical authentication scheme to address flash memory limitations, ensuring secure and reliable access by distributing wear evenly and protecting against power attacks, maintaining data integrity and security.

JP7716526B2Active Publication Date: 2025-07-31SONY SEMICON SOLUTIONS CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024042937
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-13
Filing Date
2024-03-18
Publication Date
2025-07-31
Estimated Expiration
2039-12-10

AI Technical Summary

Technical Problem

Flash memory in computer systems faces limitations such as inherent weaknesses in programming and erasing cycles, vulnerability to security attacks, and uneven wear distribution leading to premature device failure.

Method used

A hierarchical authentication scheme is implemented through a Flash Translation Layer (FTL) that translates logical addresses into physical addresses, performs wear leveling, and includes a security structure resistant to power interruptions, using Message Authentication Codes (MAC) and Initialization Vectors (IV) to authenticate data and mapping entries.

Benefits of technology

The solution ensures secure and reliable flash access by evenly distributing wear, protecting against power attacks, and maintaining data integrity and security, even in the event of power failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007716526000001
    Figure 0007716526000001
  • Figure 0007716526000002
    Figure 0007716526000002
  • Figure 0007716526000003
    Figure 0007716526000003
Patent Text Reader

Abstract

To provide a secure flash-based computer system.SOLUTION: In a computer system 100, a CPU 102 includes a non-volatile memory (NVM) interface and a processor. The NVM interface is configured to communicate with the NVM. The processor is configured to store in the NVM at least (i) data entries including data and (ii) mapping entries including mapping information that indicate physical addresses in which the data entries are stored in the NVM, and to verify authenticity of the data entries and of the mapping entries using a hierarchical authentication scheme. In the hierarchical authentication scheme, data entries include first authentication information that authenticates the data, and the mapping entries include second authentication information that authenticates both mapping information and data entries.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to flash-based computer systems, and more particularly to secure flash-based computer systems. [Background technology]

[0002] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 778,918, filed December 13, 2018, the disclosure of which is incorporated herein by reference.

[0003] In computer systems, the Flash Translation Layer (FTL) is an intermediate system, consisting of software and hardware, that manages flash memory operations. The FTL performs tasks such as logical to physical address translation, garbage collection, and wear leveling. Some FTLs also perform error correction coding (ECC), bad block management, encryption / decryption, and authentication.

[0004] PCT International Patent Application Publication WO2014 / 123372 (Patent Document 2) describes an FTL design framework that includes a log of data, mapping, and checkpoints to support error recovery.

[0005] US Patent No. 8,589,700 (Patent Document 3) describes a system and method for whitening, encrypting, and managing data for storage in non-volatile memory. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] U.S. Provisional Patent Application No. 62 / 778,918 [Patent Document 2] PCT International Application Publication WO2014 / 123372 [Patent Document 3] U.S. Patent 8,589,700 SUMMARY OF THE INVENTION

[0007] Embodiments of the invention described herein provide a computing device having a non-volatile memory (NVM) interface and a processor. The NVM interface is configured to communicate with the NVM. The processor is configured to store in the NVM at least (i) a data entry including data, and (ii) a mapping entry including mapping information indicating a physical address at which the data entry is stored in the NVM, and to verify the reliability of the data entry and the mapping entry using a hierarchical authentication scheme, wherein in the hierarchical authentication scheme, (i) the data entry includes first authentication information for authenticating the data, and (ii) the mapping entry includes second authentication information for authenticating both the mapping information and the data entry.

[0008] In some embodiments, the processor is configured to verify the reliability of the data entry and the mapping entry in response to an initialization instruction. In one embodiment, in response to writing data to the NVM, the processor is configured to update the hierarchical authentication scheme with (i) updated first authentication information reflecting the written data, and (ii) updated second authentication information reflecting the written data and the mapping information of the written data.

[0009] In the disclosed embodiments, in response to reading data from the NVM, the processor is configured to verify the reliability of the read data using at least the first authentication information and the second authentication information. In an exemplary embodiment, the processor is configured to update the hierarchical authentication scheme in an order that guarantees the consistency of the data and the hierarchical authentication scheme upon power-off. In some embodiments, as part of the hierarchical authentication scheme, the processor is further configured to store third authentication information for authenticating the mapping entry in the NVM.

[0010] According to an embodiment of the present invention, there is provided a method of computing, comprising: storing in a non-volatile memory (NVM) a mapping entry including at least (i) a data entry containing data, and (ii) mapping information indicating a physical address where the data entry is stored in the NVM. The reliability of the data entry and the mapping entry is verified using a hierarchical authentication method, in which (i) the data entry includes first authentication information for authenticating the data, and (ii) the mapping entry includes second authentication information for authenticating both the mapping information and the data entry.

Brief Description of the Drawings

[0011] The present invention will be more fully understood from the following detailed description of its embodiments in conjunction with the drawings: [Figure 1] A block diagram schematically showing a computer system according to an embodiment of the present invention. [Figure 2] A block diagram schematically showing an interface of a flash translation layer (FTL) according to an embodiment of the present invention. [Figure 3] A block diagram schematically showing a page in a flash device according to an embodiment of the present invention. [Figure 4] A block diagram schematically showing the structure of a data page according to an embodiment of the present invention. [Figure 5] A block diagram schematically showing the structure of a PT page according to an embodiment of the present invention. [Figure 6] A block diagram schematically showing the structure of a hash message authentication code (HMAC) page according to an embodiment of the present invention. [Figure 7] A block diagram schematically showing the structure of a free page according to an embodiment of the present invention. [Figure 8] A flowchart schematically showing a method for formatting a flash device according to an embodiment of the present invention. [Figure 9]3 is a flow chart that schematically illustrates a method for initializing a flash device, according to an embodiment of the present invention. [Figure 10] FIG. 1 is a block diagram that schematically illustrates RAM data used by the FTL, in accordance with an embodiment of the present invention. [Figure 11] 3 is a flow chart that schematically illustrates a method for reading data from flash, according to an embodiment of the present invention. [Figure 12] 1 is a flow chart that schematically illustrates a method for writing data to flash, according to an embodiment of the present invention; [Figure 13] 1 is a flowchart that schematically illustrates a method for wear leveling flash, in accordance with an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] (overview) A computer system's secondary storage differs from its primary storage (such as random access memory - "RAM") in that it is not directly accessible by the central processing unit ("CPU"). Typically, a computer uses input / output channels to access secondary storage and transfer the desired data to primary storage. Secondary storage is typically non-volatile, and in some computer systems, it may contain a much larger storage capacity than primary storage.

[0013] Traditionally, secondary storage has been based on magnetic media (e.g., hard disk drives - HDDs). However, more recently, computer systems have relied on semiconductor non-volatile memory (such as flash) as secondary storage in addition to, or instead of, traditional magnetic storage.

[0014] While offering large storage capacity at a relatively low cost, flash memory has some inherent limitations. In a typical flash device, any bit can be individually programmed to a first binary value (such as logic 1), but programming to a second binary value (such as logic 0) must be done in larger blocks of memory called pages. (Programming a page with a second binary value is called erasing the page.) Thus, when a page in flash is to be programmed with data, the CPU typically erases the flash and then programs the desired data. When only part of a page is programmed with data, the CPU typically copies the page to random access memory (RAM), changes the portion to be programmed, erases the page, and then reprograms the erased page from RAM.

[0015] A second inherent weakness of flash memory is aging (also known as "wear"). The number of times a cell can be reliably programmed / erased ("P / E cycles") is typically limited to 100,000. If the same page is repeatedly programmed and erased, it may reach the end of its life cycle (and therefore the flash device may be considered non-functional), even though other pages may have lower P / E cycle counts.

[0016] A third weakness of flash memory, especially flash memory that is external to a computer system (e.g., flash devices that connect to a computer system via a Universal Serial Bus (USB) connector), is its vulnerability to security attacks. Flash file systems may contain cryptographic keys and signatures, and flash drivers may apply trusted authentication techniques. However, because programming and erasing can take a relatively long time, an attacker could remove power midway through an erase or program cycle, thereby setting the storage device in an insecure state.

[0017] Embodiments of the invention disclosed herein provide apparatus and methods for secure and reliable flash access of software programs executed by a CPU. According to embodiments, when a CPU executes a software program that accesses flash, the CPU can invoke flash interface software called a Flash Translation Layer ("FTL"), which translates the software program's access to flash into a series of flash access (and possibly maintenance) operations, transparent to the software program, that are designed to mitigate the weaknesses and bypass some of the limitations of flash memory.

[0018] An FTL according to embodiments of the present invention implements a hierarchical security structure, providing high levels of security, including resilience against power interruption attacks. According to some embodiments, the FTL distributes flash wear evenly among flash pages (referred to as "wear leveling").

[0019] In one embodiment, the flash interface translates addresses sent by the CPU when accessing flash (hereinafter referred to as "logical addresses") into physical addresses within the flash device. The physical address space is larger than the logical address space, and additional storage space is used to store, among other things, tables, authentication codes (sometimes referred to as metadata), and other data (described below).

[0020] In an exemplary embodiment, a 16-byte Message Authentication Code (MAC) is appended to each group of 64 bytes corresponding to a 64-byte segment in the logical address space to form an 80-byte Security Interface Code (SIC), where the MAC is used for data authentication. The MAC may be generated when the CPU encrypts the data segment (e.g., using Advanced Encryption Standard Galois / Counter Mode AES128-GCM).

[0021] In some embodiments according to the present invention, the flash device includes four types of pages: data pages, which are used to store CPU data; pointer table (PT) pages, which are used to store logical to physical address translation pointers (and some additional data described below); hash message authentication code (HMAC) pages, which are used to authenticate PT pages; and free pages, which are used in write operations, e.g., when a data page is full. According to embodiments of the present invention, the authentication structure is hierarchical, and any changes to any bits in data pages and / or other page types are easily identified as the CPU hierarchically authenticates the flash device.

[0022] The embodiments described herein relate primarily to exemplary implementations of the hierarchical authentication scheme described above. However, the disclosed techniques can be used to verify the authenticity of data entries and mapping entries using a variety of other hierarchical authentication schemes in which (i) a data entry includes a first authentication information that authenticates the data, and (ii) a mapping entry includes a second authentication information that authenticates both the mapping information and the data entry. Such hierarchical authentication schemes may also include a third authentication information that authenticates the mapping entry. In this example, purely by way of example, the first authentication information includes a MAC, the second authentication information includes an Initialization Vector (IV), and the third authentication information includes a Hashed Message Authentication Code (HMAC).

[0023] Thus, according to an embodiment of the present invention, a flash translation layer (FTL) translates CPU accesses into a sequence of flash access instructions, implementing physical to logical address translation, efficient wear leveling, and a hierarchical security structure that is protected against power interruptions.

[0024] (System Description) A computer system typically consists of a fast primary storage device and a slower secondary storage device. The primary storage device is often random access memory (RAM), which is volatile (i.e., it loses data when power is turned off), while the secondary storage device is non-volatile memory (NVM), such as flash memory or a hard disk drive (HDD). The following description generally refers to flash memory. However, embodiments of the present invention are not limited to flash memory. In alternative embodiments, any other suitable type of NVM may be used (e.g., electrically erasable programmable read-only memory (EEPROM)).

[0025] 1 is a block diagram that schematically illustrates a computer system 100, according to one embodiment of the present invention. The computer system 100 includes a central processing unit (CPU) 102, also referred to as a computing device. While the embodiments described herein primarily refer to a CPU, the disclosed techniques may also be implemented in a variety of other computing devices, such as memory controllers. The system 100 further includes a random access memory (RAM) 104 and a flash memory 106. The computer system 100 may further include other units 108, such as input / output, interfaces, and the like, that are not described or required in the following disclosure.

[0026] The CPU 102 typically directly accesses data stored in RAM 104 by issuing memory access instructions. In contrast, accesses to data stored in Flash 106 are not handled directly; memory access instructions targeting data stored in Flash must be translated into a series of Flash access and management instructions.

[0027] In this example, CPU 102 includes an NVM interface 105 for communicating with flash 106 and a processor 103 configured to perform the disclosed techniques.

[0028] According to embodiments of the present invention, flash memory can be erased or programmed. Erasing is done at the granularity of a page (typically 4096 bytes), where all bits in the page are set to a first binary value (e.g., logic 1). Programming is done in groups of up to blocks (e.g., 80 bytes), at the granularity of a single bit, where all specified bits are set to a second binary value (e.g., logic 0).

[0029] In an embodiment according to the present invention, each page of Flash 106 can be erased a limited number of times (e.g., up to 100,000 times). When any page reaches its erase count limit, Flash 106 is considered bad, even if some (or most or all) of the other pages have been erased very little (or not at all).

[0030] In some embodiments in accordance with the present invention, flash 106 is less vulnerable to security attacks, such as data theft or data modification, than RAM 104. Flash 106 retains data even after power is removed, which provides hackers with ample opportunity to manipulate the flash, for example, by copying it and then analyzing its contents with a powerful computer. Furthermore, in some embodiments, flash 106 can be manually removed and inserted into computer system 100, for example, using a universal serial bus (USB) connector. In contrast, RAM 104 loses data when power is removed and is therefore much more protected (in some embodiments, RAM 104 and CPU 102 are integrated into the same integrated circuit, further improving protection of RAM 104 data).

[0031] In an embodiment according to the present invention, the software program executed by the CPU 102 issues random access flash memory read and write commands without being aware of the limitations of the flash memory, and the flash translation layer (FTL) interface (actually, a software driver) responds to the read and write commands issued by the software program and sends commands to the flash, whereby the access commands issued by the software program are executed. The FTL executed by the CPU 102 maintains a hierarchically secure structure of the data stored in the flash, while evenly distributing the erase counts of the flash pages to all pages.

[0032] As will be appreciated, the structure of the computer system 100 described above is cited as an example. The computer system according to the disclosed technology is not limited to the above description. For example, in an alternative embodiment, the CPU 102 can be an aggregate of multiple CPUs. The RAM 104 may be an aggregate of memories, some of which are directly connected and others are connected via a bus. The flash 106 may be any other type of non-volatile memory (e.g., EEPROM) and NVM Express over Fabrics (NVMF). In some embodiments, the software is loaded from the flash 106 to the RAM 104. In other embodiments, the software may be loaded via a network (not shown), a serial link (not shown), or a combination of channels.

[0033] In some embodiments, the computer system 100 includes a hardware security accelerator configured to compute various security functions such as AES-128. In one embodiment, the hardware security accelerator is embedded within the CPU 102.

[0034] 2 is a block diagram 200 that schematically illustrates the interface of a flash translation layer (FTL), in accordance with an embodiment of the present invention. The FTL 202 is called by high-level software 204, which in turn calls a low-level flash driver (LLFD) 206, which communicates directly with a flash device 208. The FTL 202, high-level software 204, and LLFD 206 all execute on the processor 103 of the CPU 102.

[0035] According to embodiments of the present invention, high-level software 204 accesses flash 208 by calling FTL 202 with the following parameters: a logical address in flash, a read or write instruction for a read or write cycle, and, in the case of a write cycle, the data to be written (in some embodiments, high-level software 204 may indicate the width of the data to be read or written, e.g., 1 byte, 2 bytes, 4 bytes, etc.). Upon completion, the FTL returns the read data (in the case of a read operation) and a completion indication to the calling program.

[0036] FTL 202 translates commands from high-level software 204 into low-level flash operations. In the exemplary embodiment of Figure 2, the low-level operations are not atomic flash commands and therefore are not input directly to flash device 206. Rather, the FTL invokes LLFD 206, which further decomposes the flash operations issued by the FTL into atomic flash commands. For example, in some embodiments, the LLFD may decompose the program instructions issued by the FTL into a series of program / verify instructions that are repeated until verification passes.

[0037] When the FTL calls the LLFD, the FTL passes parameters to the LLFD. According to the example embodiment of FIG. 2, the parameters include a physical address, a read instruction, a program instruction, an erase instruction, and, in the case of programming, the data to be programmed ("write data"). According to the example embodiment of FIG. 2, only one of programming, erasing, or reading can be instructed (if the write data includes a logic-1 bit and a logic-0 bit, only one level (e.g., logic 0) is programmed, and the other bits are not programmed and are considered to have been pre-erased).

[0038] The LLFD 206 returns a read complete indication, a program complete indication, an erase complete indication, and, in the case of a read, the read data to the FTL.

[0039] In alternative embodiments, the interface from the high-level software to the FTL and / or from the FTL to the LLFD may include, instead of explicit write and read data, a pointer to a memory buffer from which the write data should be read and / or a pointer to a memory buffer to which the read data should be written, and a length indicator.

[0040] 2, high-level software executed by the CPU writes and reads flash data by passing read and write instructions along with data and addresses directly to the FTL, and receiving read data from the FTL. The operation of the FTL, including address translation, encryption, decryption, hierarchical authentication, and wear leveling, is transparent to the calling software.

[0041] As will be appreciated, the interface of the FTL202 above is cited as an example. The FTL driver according to the disclosed technology is not limited to the above description. For example, in an alternative embodiment, the function of the LLFD206 can be integrated into the FTL202. In another embodiment, the LLFD206 can be integrated into the flash device 208. In some alternative embodiments, other suitable parameters can be passed between the high-level software 204 and the FTL202, and between the FTL202 and the LLFD206.

[0042] In some embodiments, the FTL202 is wrapped by a higher-level driver such as a file system executed by the CPU102, so that a software program can perform file operations (such as opening a file) on the data stored in the flash. In other embodiments, the interface from the high-level software 204 to the FTL202 may include diagnostic commands. For example, in one embodiment, the FTL may indicate that the wear count of a page has reached (or is close to) the specified maximum erase count of the flash device.

[0043] FIG. 3 is a block diagram 300 schematically showing pages within a flash device according to an embodiment of the present invention.

[0044] As will be appreciated, the flash device 208 (FIG. 2) can include additional pages not defined in the exemplary embodiment of FIG. 3. Such pages are not shown and are outside the scope of the present disclosure. For example, boot code and security signatures can be stored in the flash and accessed using an interface different from the FTL (or, in some embodiments, the FTL can be extended to access additional pages). In the following disclosure, it is assumed that if such other pages exist, they are stored in a different partition of the flash device.

[0045] In the exemplary embodiment shown in Figure 3, the flash device includes four types of pages: data pages 302, which store encrypted user data and metadata; pointer table (PT) pages 304, which store mapping information; hash message authentication code (HMAC) pages 306, which store security signatures of PT pages; and free (empty) pages 308. During device operation, pages dynamically change function. For example, data pages are emptied to become free (empty) pages, which become HMAC pages, etc. In the exemplary embodiment shown in Figure 3, there are Nd data pages (e.g., 98), Np PT pages (e.g., 30), Nf free pages (e.g., 4), and one HMAC page.

[0046] A hierarchical security structure is implemented in the flash pages, where data pages are authenticated using data (e.g., initialization vectors) from PT pages, and PT pages are authenticated using signatures stored in HMAC pages.

[0047] FIG. 4 is a block diagram that schematically illustrates the structure of a data page 302, according to one embodiment of the present invention.

[0048] 4, each data page includes fifty 80-byte data SIC (Security Interface Code) fields 402, a single 1-byte page type field 404, and a single 4-byte erase count field 406. (In the exemplary embodiment shown in FIG. 4, for each data page totaling 4096 bytes, the data page also includes 91 unused bytes.)

[0049] Each data SIC (Security Interface Code) field 402 includes a 64-byte data field 408 and a 16-byte MAC (Message Authentication Code) field 410 calculated over the respective 64-byte data field. The data and MAC fields are generated with an Initialization Vector (IV) stored in the PT page, for example using AES128-GCM.

[0050] FIG. 5 is a block diagram that schematically illustrates the structure of a PT (Pointer Table) page 304, according to one embodiment of the present invention.

[0051] In the exemplary embodiment shown in FIG. 5, each PT page includes a number of Pointer Table Security Information Code (PT-SIC) fields 502; a number of Page Table Header (PTH) fields 504; a type field 506 that indicates that the current page is a PT page; and an erase count field 508 that tracks the number of times the current page has been erased.

[0052] Each PT-SIC 502 consists of 80 bytes: a 64-byte PT-Data field 510 and a 16-byte MAC (Message Authentication Code) field 512 used to authenticate the PT-Data field. Each PT-Data field 502 contains four 14-byte PT entry fields 514 (plus some unused bits). Each PT entry 514 contains a 2-byte physical address field 516 that points to the address of the data SIC within the data page (64-byte resolution); and a 12-byte Initialization Vector (IV) field 518 used to authenticate and encrypt / decrypt the corresponding data SIC.

[0053] Each PTH field 504 is associated with one of the PT-SIC fields 502 and stores metadata for the corresponding PT-SIC. The PTH field 504 contains: one valid bit 520 indicating whether the current PTH and its corresponding PT-SIC contain valid data; a 15-bit PT-ID field 522 that stores the current PT-SIC and the logical address (64-byte resolution) corresponding to the PTH; and a 12-byte PT IV (Initialization Vector) used in authentication and encryption / decryption of the PT-SIC.

[0054] In an alternative (more efficient) embodiment, two consecutive PT-SICs 502 point to a group of nine PT entries 512 occupying 9x14=126 bytes. The reserved field consists of two bytes.

[0055] 6 is a block diagram that schematically illustrates the structure of an HMAC (Hash Message Authentication Code) page 306, according to an embodiment of the present invention. The HMAC page includes an HMAC signature 602 calculated over all PT headers, free space 604 for additional HMAC signatures, a type field 606 that indicates the current page is an HMAC, and an erasure count field 608 that tracks the number of times the current page has been erased. According to the example embodiment shown in FIG. 6, each HMAC signature includes 32 bytes, the page type field includes 1 byte, and the erasure count field includes 4 bytes.

[0056] The new HMAC (Hash Message Authentication Code) entry is written to the free space of the page (along with other entries in other flash pages), and the old entry is marked as invalid.

[0057] The HMAC entry stores the security signature of the PT table (in RAM), which in turn stores the initialization vector (IV) of the data table. Thus, the security structure is hierarchical, and any modification of any page of data, PT, or HMAC is detected by authenticating the flash device hierarchically.

[0058] FIG. 7 is a block diagram schematically showing the structure of free page 308 according to an embodiment of the present invention. The free page can be changed to a data page, a PT page, or an HMAC page in terms of its type. Therefore, the free page needs to maintain an erase count (in the embodiment, since all free page type codes are -1 (erased), when a page is erased, the page becomes free until its type field is changed). Therefore, the free page 308 is composed of an empty field (e.g., 4091 bytes out of 4096 bytes of the page); a 1-byte type field indicating that the current page is free; and a 4-byte erase count for tracking the number of times the page has been erased.

[0059] Therefore, according to the block diagrams shown in FIGS. 3-7, the flash device is divided into 4096-byte pages including data pages, PT tables, HMAC pages, and free pages. To relax the flash limitations, new entries are programmed into the empty fields of data, PT, and HMAC pages, thereby avoiding frequent erasures.

[0060] Since writing is done by concatenating information to the page, power failures during a write transaction are ignored. When the power is restored, the interrupted write fails either MAC authentication, PT-MAC authentication, or HMAC authentication. Therefore, the interrupted write is discarded and the flash translation layer (FTL) recovers by verifying the previous HMAC.

[0061] Each page tracks the number of erasures, so that the wear leveling algorithm (described later) can be used to evenly distribute erasures among the pages. Since the authentication structure is hierarchical, the CPU can authenticate the entire flash device hierarchically.

[0062] It should be understood that the page structure of flash device 106 described above with reference to Figures 3-7 is cited as an example. Flash pages according to the disclosed technology are not limited to the above description. For example, in an alternative embodiment, a data page may include 51 data SICs, a 3-byte erase count field, a 1-byte type field, and 12 unused bytes. In other embodiments, a flash page may have 8192 bytes (in which case the other figures above would change accordingly).

[0063] In some embodiments, error correction code may be added to some or all pages. Finally, in one embodiment, the CPU may mark some pages as defective (e.g., if the CPU gets an indication from the low-level flash driver (LLFD) 206 (FIG. 2) that a program or erase failed). The CPU then uses other pages from the pool of free pages.

[0064] It will be appreciated that in embodiments in accordance with the present invention, the flash device may include additional pages not listed within the same or other partitions of the flash, storing any form of encrypted or clear data.

[0065] We now proceed to describe the methods for formatting, initializing, reading, writing, and wear-leveling a flash device in an embodiment according to the present invention. For clarity, some non-essential steps are omitted from the following description.

[0066] (format) According to an embodiment of the present invention, when the flash device is formatted, the CPU first erases all flash pages. Next, for each page, the correct page type (data / PT / HMAC or free) is set and the erase count is initialized. Next, the CPU constructs a PT table (in RAM), a valid SIC entry table for each page, and a next SIC pointer table. Finally, the CPU calculates the HMAC signature (of the PT table in RAM) and programs it at the first HMAC entry.

[0067] The number of data pages, PT pages, and free pages are denoted as Nd, Np, and Nf, respectively (the number of HMAC pages is 1). In an exemplary embodiment, Nd = 96, Np = 30, and Nf = 4.

[0068] FIG. 8 is a flowchart 800 schematically showing a method of formatting a flash device according to an embodiment of the present invention. This flowchart is executed by the CPU 102 (FIG. 1).

[0069] The flowchart begins with an all-page erase step 802 where the CPU erases all pages (starting from the first page). In some embodiments, the CPU first checks whether the flash has already been erased in order to save time (and erasure) if the flash is new.

[0070] Next, the CPU enters a data page format step 804 where the CPU programs the type field of the current page to indicate that the page is a data page, and the erase count field of that page indicates that the page has been erased once (the erasure was performed in step 802). The CPU loops step 804 Nd times to format Nd data pages.

[0071] The formatting of data pages by looping through step 804 is repeated for PT pages and free pages. To format Np PT pages, the CPU loops through PT page format step 806 Np times (the type field is set to indicate a PT page, and the erase count is set to 1). To format Nf free pages, the CPU loops through free page format step 808 Nf times, with the type field set to indicate a free page (or unchanged if the free page's code is all 1s), and the erase count is set to 1.

[0072] The format of a single HMAC page is similar - the CPU enters HMAC page format step 810, setting the type field and erase count. Finally, the CPU enters HMAC calculation and programming step 812, where the CPU calculates the HMAC signature of the PT table (in RAM) and programs the result into the first HMAC entry of the HMAC page.

[0073] Thus, according to the example flowchart shown in Figure 8, when formatting flash, the CPU prepares Nd data pages, Np PT pages, a single HMAC page, and Nf free pages, where Nd + Np + Nf + 1 is equal to the number of flash pages used by the flash translation layer (FTL) (which may be less than or equal to the number of available flash pages). The first entry in the HMAC page is programmed with the HMAC signature of the PT table that the CPU reads from RAM.

[0074] It is understood that the above formatting flowchart is cited as an example. The formatting flow according to the disclosed technology is not limited to the above description. In another embodiment, for example, formatting begins by testing the flash device, which indicates a bad page and skips it. In other embodiments, flash testing may be performed after formatting. In still other embodiments, the order of formatting flow steps may be changed.

[0075] (initialization) According to an embodiment of the present invention, upon system initialization (e.g., power-on reset), the flash driver scans the flash device and validates the PT pages. Then, the flash driver prepares an initial PT table and a table of valid entries per page in RAM (both described below). Finally, the flash driver finds the next empty location where the next data, next PT, and next HMAC will be written and saves the corresponding pointers in RAM.

[0076] FIG. 9 is a flowchart 900 that outlines a method for initializing a flash device, according to one embodiment of the present invention. The flowchart is executed by the CPU 102, which typically enters the initialization flow upon reset, such as a power-on reset, upon a software command, or in response to other appropriate initialization instructions. The flow begins with a find PT page step 902, in which the CPU scans all flash pages, looking for a type field indicating a PT page. Next, in a find valid PT entry step 904, the CPU searches each PT page (valid field 520 of PTH 504 in FIG. 5) for a valid PT entry. Next, in a get logical address step 906, the CPU reads the PT-ID field 524, which indicates the logical address (64-byte resolution) of the PT entry. Next, in a write RAM-PT-table step 908, the CPU writes the address of the PT-SIC in RAM to an index derived from the logical address found in step 906.

[0077] After step 908, the CPU enters the per-page valid entry writing step 910, and checks the number of valid entries for each data page and each PT page. The CPU derives the number of valid entries for each PT page by setting the valid bit and counting the PTH. For a data page, the CPU derives the number of valid entries for the specified data page by counting the number of valid bit diagrams in the PTH fields corresponding to all PT entries pointing to the specified data page. The CPU writes the number of valid entries to the RAM.

[0078] After step 910, the CPU enters the PT page authentication step 912, where the CPU calculates a signature (e.g., 128 bits of SHA256 hash) and compares the result with the valid HMAC from the HMAC page. If the signatures do not match, the CPU initially assumes that the power was cut off the last time the flash was written, and takes measures to recover (described below). If the measures fail, the CPU either terminates forcibly or notifies the user that the security has been breached. If the signatures match, the CPU enters the next entry pointer creation step 914, where the CPU creates pointers to the next data page, PT page, and HMAC page entries, and saves the pointers to the RAM. The CPU can create the pointers using, for example, the following algorithm. i) Scan all data pages in reverse order from the type field to find the last entry that has not been erased (since a valid MAC field will never be all 1s, it can always be distinguished from an empty field). ii) The pointer to the next data page entry is the last non-empty entry of the first data page when an empty entry is found. iii) Repeat steps i) and ii) for the PT page. iv) Repeat steps i) and ii) for a single HMAC page. After step 914, the flow ends.

[0079] Note that if power is interrupted during a write operation, an anomaly may occur in the data structure. As explained below, FTL detects the anomaly during authentication and reverts to the data before the write failed, losing the last written data but maintaining integrity.

[0080] Thus, according to the exemplary embodiment shown in FIG. 9, at initialization time, the CPU hierarchically authenticates the flash and builds RAM tables and pointers to enable efficient and secure reads and writes from the flash device.

[0081] It should be understood that the above initialization flowchart is cited as an example. The initialization flow according to the disclosed technology is not limited to the above description. In alternative embodiments, for example, the order of steps may be different, and some steps may be combined, omitted, or replaced with alternative appropriate steps.

[0082] (RAM tables and pointers) According to an embodiment of the present invention, the flash driver prepares (during initialization) and maintains (during execution) pointers and tables in RAM (the size of the tables is much smaller than the size of the corresponding information stored in flash).

[0083] Figure 10 is a block diagram that schematically illustrates RAM data 1000 used by the FTL, in accordance with one embodiment of the present invention. The CPU fills the RAM data at initialization, as described above (with reference to Figure 9), and updates the RAM data following write and wear-leveling operations, as described below.

[0084] The RAM data 1000 includes a PT table 1002 , a valid entry table per page 1004 , and a next SIC pointer 1006 . The PT table 1002 contains multiple entries, one for each possible logical address (in 64-byte increments). Each PT table entry includes a 2-byte physical address field 1008 that stores a pointer to the corresponding PT entry address, and a 12-byte PT-IV field 1010 that specifies the initial vector of the PT message authentication code (MAC).

[0085] The valid SIC entry 1004 per page includes multiple valid SIC field entries 1014 that indicate how many valid SIC fields are stored in the corresponding data page. The valid SIC entry 1004 per page further includes multiple entries, one for each PT page. Each entry is composed of a 1-byte valid SIC field entry 1018 and indicates the number of valid SIC fields stored in the corresponding PT page.

[0086] Finally, the next SIC pointer 1006 includes the next data SIC pointer 1020 that points to the next free data SIC, the next PT SIC pointer 1022 that points to the next free PT SIC, and the next HMAC pointer 1024 that points to the next free HMAC field.

[0087] As will be appreciated, the above block diagram of the RAM data is cited as an example. The RAM data structure according to the disclosed technology is not limited to the above description. In alternative embodiments, other suitable structures of tables and pointers may be used.

[0088] (Read flow) According to an embodiment of the present invention, when high-level software requests a read operation from a logical address, the FTL reads the address and IV of the corresponding PT entry from the RAM, decrypts and authenticates the PT in the flash, obtains the address and IV read from the corresponding PT entry in the PT, reads the security interface code (SIC) from the data page, decrypts / authenticates (using the MAC and IV), and returns the decrypted data.

[0089] Figure 11 is a flowchart 1100 that schematically shows a method for reading data from a flash according to an embodiment of the present invention. This flowchart is executed by the CPU 102 (FIG. 1) in response to a read instruction received by the FTL 202 from the high-level software 204 (FIG. 2). The flow starts at the PT address and IV read step 1102, where the CPU reads the PT physical address 1008 (corresponding to the logical address received by the CPU in 64-byte units) and the PT IV 1010 from the RAM table 1002 (FIG. 10). Next, in the PT read step 1104, the CPU reads the PT entry from the PT page corresponding to the physical address read by the CPU in step 1102.

[0090] Next, in the PT decryption / authentication step 1106, the CPU decrypts the PT entry using the PT data, PT MAC, and IV read in step 1104. The decrypted data includes the physical address of the data SIC corresponding to the logical address and the IV.

[0091] If authentication fails in step 1106, the CPU either forcibly terminates or notifies the user that the flash data is damaged. If authentication is successful, the CPU proceeds to the data SIC acquisition step 1108. Here, the CPU reads the SIC entry from the data page using the address (from step 1106), and then proceeds to the data SIC decryption / authentication step 1110.

[0092] In step 1110, the CPU decrypts and authenticates the data SIC using the data SIC MAC and the corresponding IV (decrypted in step 1106).

[0093] If authentication fails in step 1110, the CPU either forcibly terminates or notifies the user that the flash data is damaged. If authentication is successful, the CPU proceeds to the data return step 1112, where the FTL returns the read data to the calling high-level software, and the flow ends.

[0094] It should be understood that the above read flow chart is cited as an example. The read flow according to the disclosed technology is not limited to the above description. In alternative embodiments, for example, authentication and / or decryption may be performed by a hardware accelerator rather than a CPU. In some embodiments, reading a page may include error detection and correction.

[0095] (Write flow) We now proceed to describe the write flow according to an embodiment of the present invention. Since the written data is typically shorter than the data SIC (typically, writes are at most 8 bytes wide, while the SIC is 64 bytes wide), writing to the data SIC is a read-modify-write sequence, where the complete SIC is read, the written portion is modified, and the new SIC-modified data is written back to the SIC.

[0096] 12 is a flowchart 1200 that generally illustrates a method for writing data to flash, according to an embodiment of the present invention, which is executed by CPU 102 in response to a write instruction received by FTL 202 from high-level software 204 (FIG. 2).

[0097] In the example embodiment shown in Figure 12, after the CPU writes the last entry of a page, the CPU enters a wear leveling sequence, in which the page with the lowest erase count is emptied, defragmented, saved to a free page, and the emptied page is turned into a free page. In the next step of flow 1200, the CPU updates the number of valid entries in the page so that the new number is equal to the full capacity of the page, and in the next step, when the CPU updates the pointer to the next entry (in RAM), the pointer points to the first entry of the emptied page (if not entering wear leveling, the pointer update consists of a pointer increment). For clarity, the wear leveling sequence is not shown in Figure 12 (described below with reference to Figure 13).

[0098] The flow of FIG. 12 starts at read step 1202 where the CPU executes a read flow including decoding and authentication. The read flow is the same as read flow chart 1100 (FIG. 11), except that data is not returned to the caller (step 1112), and instead the decoded SIC is stored in the RAM.

[0099] Next, at data change step 1204, the CPU replaces the corresponding byte of the decoded SIC with the write data. Next, the CPU enters SIC coding step 1206 where the CPU generates an IV (e.g., using a random number generator) and codes the modified SIC, e.g., using AES128 - GCM.

[0100] After step 1206, at SIC write step 1208, the CPU obtains the next SIC address from the RAM (pointer 1020, FIG. 10) and programs the corresponding SIC of the flash device with the coded data obtained at step 1208. If the written SIC is the last available entry of the page, the CPU executes a wear leveling sequence (after writing the SIC).

[0101] Next, at per - page valid entry update step 1210, the CPU updates the valid entry table 1014 (FIG. 10) for each data page, and then enters the next SIC pointer update step 1212 where the CPU updates pointer 1020 (FIG. 10) to the next SIC.

[0102] After step 1212, the CPU enters write new PT entry step 1214, where the CPU updates the physical address and IV fields of the corresponding PT entry and PTH, including encoding the PT entry using, for example, AES128-GCM. The PT entry reflects the physical address where the data SIC was written in step 1208, while the PTH is updated with the logical address and IV (generated in step 1206). The CPU obtains the address of the next PT entry pointer from RAM (Figure 10, 1022). If the PT entry the CPU is writing is the last available entry in the current PT page, the CPU performs a wear leveling sequence.

[0103] Next, the CPU enters an update valid entry per PT page step 1216 to update the valid entry per PT page table 1018 (FIG. 10) in RAM, and then enters an update next PT entry step 1218, where the CPU updates the pointer 1022 (FIG. 10).

[0104] The CPU now enters write new HMAC step 1220, where it calculates the HMAC for the new PT page (using, for example, AES128) and stores the result in the next HMAC pointer 1024 (Figure 10) in the HMAC page. If the HMAC entry that the CPU is writing is the last entry available in the current page, the CPU performs a wear leveling sequence.

[0105] Finally, the CPU enters next HMAC pointer update step 1222, updates the next HMAC pointer 1024 (FIG. 10), invalidates the old PT header (clears the valid bit in the PT header), and ends the flow.

[0106] As can be understood, the above write flowchart is cited as an example. The write flow according to the disclosed technology is not limited to the above description. In another embodiment, for example, the order of some steps may be changed. In other embodiments, the encoding of the data SIC (step 1206) and / or the PT entry (step 1212) includes generating an ECC code for the encoded data.

[0107] (wear leveling) In an embodiment according to the present invention, each flash page includes an erase count, and the FTL is configured to evenly distribute erasures among the pages in order to avoid early wear of the flash device. Such an even distribution is called wear leveling.

[0108] When a flash page (data, PT, or HMAC) is full, usually one or more invalid entries (for example, entries replaced by new entries) are stored in the page. Therefore, in most cases, the valid data can be compressed to create space for more entries. The process of compressing valid data is called defragmentation below, and the corresponding verb is called to defragment.

[0109] However, since a full page may have been erased more times than other pages, it is preferable to defragment another page with fewer erase counts.

[0110] In an embodiment according to the present invention, when a page becomes full, the FTL defragments the page with the minimum number of erasures. If there are no invalid entries in that page, the FTL defragments an additional page (the page with the fewest erase counts among all pages having at least one invalid entry).

[0111] FIG. 13 is a flowchart 1300 schematically showing a method for wear leveling a flash device according to an embodiment of the present invention. This flowchart is executed by the CPU 102 and starts after the CPU writes the last valid entry of a page (for example, steps 1208, 1214, and 1218 in FIG. 12).

[0112] The flow begins at page selection step 1302, where the CPU reads the erase counts of all pages and selects the page with the smallest erase count as the page to be emptied. (If there are multiple pages with the fewest number of erasures and at least one invalid entry, the page with the most invalid entries is emptied.)

[0113] Next, in copy step 1304 of the emptied page, the CPU reads the content of the selected page and stores the content in the RAM. Next, the CPU enters defragmentation step 1306, defragments (eliminates fragments) the copied data (for example, writes all valid entries to consecutive locations in the RAM), and proceeds to free page programming step 1308.

[0114] In step 1308, the CPU programs one of the free pages with the data in the RAM. The programming includes changing the type field from "free" to the type of the emptied page and retaining the erase count (therefore, the CPU first reads the erase count field from the free page).

[0115] In some alternative embodiments, to reduce RAM usage, steps 1304, 1306, and 1308 can be integrated by continuously reading entries from the emptied page and writing valid entries to the free page. Thus, instead of allocating the RAM space for the entire page, the allocated space is equal to only one entry.

[0116] Next, the CPU enters a page erase step 1310, where the CPU reads the erase count field of the page to be emptied, erases the page, and then programs the erase count field of the page with the incremented erase count value and the page type field with the free page indication.

[0117] The CPU then enters the PT and HMAC update step 1312, where the CPU updates the PT entry and then recalculates and programs the HMAC. If the page becomes full when updating the PT or HMAC, the wear leveling flow 1300 is recursively called.

[0118] Next, the CPU enters page full check step 1314 to check whether the emptied page is completely full (i.e., whether the page contained any invalid entries before defragmentation). If the new page is not full, the flow ends. If, in step 1314, the new page is full, the CPU enters non-full page selection step 1316, where the CPU selects the page with the smallest number of erases among all pages that have at least one invalid entry. The CPU then re-enters step 1304 and repeats steps 1304-1314.

[0119] It should be understood that the above wear leveling flowchart is cited as an example. The wear leveling flow according to the disclosed technology is not limited to the above description. In an alternative embodiment, for example, the page selected in step 1316 may be selected according to a weight function calculated according to the number of erases and the number of invalid entries. In some embodiments, the CPU copies only valid entries in step 1304, and step 1306 is skipped.

[0120] (Power interruption resilience) The structures and methods disclosed above ensure that the security of the system is not compromised if the power supply to the computer system 100 (FIG. 1) and / or the flash device is interrupted at any point. The power interruption is detectable during the next startup initialization (step 912, FIG. 9), and authentication fails.

[0121] In general, any power interruption occurring during any write transaction is harmless or the consistency of the flash data structure is lost, which is detected and corrected during initialization. Some examples are shown below: 1. If the power is interrupted while the CPU is writing new data to a data page, the PT continues to point to the previous data. 2. If the power is interrupted after the CPU has written data but before the CPU writes a new PT, the PT header points to the previous PT that points to the previous data. 3. If a power interruption occurs between the writing of a PT entry and the writing of the PT header, the previous PT header points to the previous PT that points to the previous data. 4. If the power is interrupted while the CPU is writing (or immediately before) the HMAC, the HMAC authentication fails when the power is next updated, and the CPU uses the previous PT header, PT entry (and HMAC). 5. If the power is cut off during defragmentation, after the page type has been written (step 1308, FIG. 13), the number of pages of various types is incorrect. In this case, the FTL defragments the spare pages. 6. If the power is interrupted before the page type is written, the pages marked as free do not become empty. To handle this contingency, the CPU can verify during initialization that all free pages are actually empty. 7. If the power is interrupted after the HMAC is updated but before the previous PT header becomes invalid, the HMAC authentication fails and the CPU invalidates the PT header.

[0122] (Use of a Replay Protection Monotonic Counter (RPMC)) In some embodiments according to the invention, a flash device may include one or more RPMCs operable to protect the flash memory against rollback (sometimes called "replay"). The count values generated by the RPMCs are guaranteed to be unique, allowing for the use of smaller IV fields. (RPMCs are described, for example, in U.S. Patent 9,405,707, which describes a system including a flash memory device that includes an RPMC and a host device.)

[0123] In some embodiments, when the flash device includes at least one RPMC, the CPU is configured as follows: 1. The flash format reads the value of the RPMC and stores it in a non-volatile register. (In some embodiments, this non-volatile register can be stored in a different flash partition signed with a different MAC. In other embodiments, the non-volatile register is a second RPMC, and the CPU increments the second RPMC until it matches the value of the first RPMC.) The contents of the non-volatile register are called the format version. 2. Before HMAC entry write step 1220 (FIG. 12), increment RPMC; use the new value for HMAC. Thus, all HMAC authentications are performed with the current RPMC value, and rollbacks are detected. 3. Compress PT pages using a smaller IV field - the random 96 bits above (see Figure 5) can be replaced with a 32-bit counter that increments with each write. When encoding / decoding, the CPU concatenates the format version value with the counter value - this ensures that the IV is unique.

[0124] It will be appreciated that embodiments of the present invention utilizing one or more RPMCs in a flash device are not limited to the above descriptions cited as examples, and other suitable techniques can be used to enhance security or reduce the size of flash tables.

[0125] The computer system 100, FTL 202, the configuration of pages 300 of flash (or partitions thereof), the configuration of the various page types (302, 304, 306), format flow 800, the initialization flow 900 described above, RAM table 1000, read flow 1100, write flow 1200, and wear leveling flow 1300 are exemplary configurations and flows shown purely for conceptual clarity. In alternative embodiments, any other suitable configurations and flows may be used. Elements not essential to an understanding of the disclosed technology have been omitted from the figures for clarity.

[0126] The different computer system elements described above may be implemented using suitable hardware, such as one or more application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs), using software, or using a combination of hardware and software elements.

[0127] In some embodiments, CPU 102 comprises one or more general-purpose processors, which are programmed with software to perform the functions described herein. The software may be downloaded to the processor in electronic form, for example, over a network or from a host, or alternatively or additionally, may be provided and / or stored on non-transitory tangible media such as magnetic, optical, or electronic memory, including a partition of flash memory 106.

[0128] According to some embodiments of the present invention, computer system 100 includes a hardware security accelerator that may be used for authentication and / or encryption / decryption. In one embodiment, the hardware security accelerator is integrated into CPU 102.

[0129] The embodiments described in this specification mainly address flash-based secure computer systems, but the methods and systems described herein can also be used in other suitable systems or applications.

[0130] Accordingly, it will be understood that the above embodiments are cited by way of example and that the present invention is not limited to those specifically shown and described above. Rather, the scope of the present invention includes both the various combinations and sub-combinations of the features described above, as well as those variations and modifications thereof that are not disclosed in the prior art and would be apparent to those of ordinary skill in the art upon reading the foregoing description. Documents incorporated by reference in this patent application should be considered an integrated part of this patent application. When definitions made explicitly or implicitly in this specification conflict with those defined in these incorporated documents, the definitions in this specification shall prevail.

Claims

1. A computing device comprising: an NVM interface configured to communicate with a non-volatile memory (NVM) having a plurality of pages; and a processor; wherein the processor is configured to: update data in the NVM by (i) writing data to an unused segment of a given page and (ii) marking as invalid a used segment that stores a previous version of the data; and select a page to defragment according to wear leveling selection criteria, copy valid data from the selected page to a free page, erase the selected page to be defragmented, and repeat the defragmentation operation until the free page filled with the valid data copied therefrom is full, for an additional page selected according to the wear leveling selection criteria; configured as such; a computing device characterized by the above.

2. The computing device according to claim 1, wherein the selection of the page for defragmentation, the copying of the valid data from the selected page, and the erasing of the selected page are initiated in response to filling all segments of the given page.

3. The computing device according to claim 1 or 2, wherein the wear leveling selection criteria aim to select the page with the smallest number of previous erasures.

4. The computing device according to claim 1 or 2, wherein the wear leveling selection criteria also depend on the count of invalid segments within the page.

5. The computing device according to claim 1 or 2, wherein the processor is configured to write the valid data to a plurality of consecutive segments within the free page.

6. The computing device according to claim 1 or 2, wherein when the processor discovers that the page to which the valid data has been copied is full after the copying of the valid data, the processor is configured to select an additional page for defragmentation and defragment and erase the additional page.

7. The processor is configured to: update mapping information indicating an updated position of the data in the NVM; and When it is discovered that the update of the mapping information has filled a specific page, start wear leveling and a defragmentation operation to defragment the specific page to another page; The computing device according to claim 1 or 2, characterized in that it is configured as described above.

8. The processor is: In the NVM, update the authentication information for authenticating the data; and When it is discovered that the update of the authentication information has filled a specific page, start wear leveling and a defragmentation operation to defragment the specific page to another page; The computing device according to claim 1 or 2, characterized in that it is configured as described above.

9. The computing device according to claim 1 or 2, characterized in that the processor is further configured to return the erased page to a free page.

10. A method of computing by an information processing device, comprising: i) writing data to an unused segment of a given page in a non-volatile memory (NVM) having a plurality of pages, and (ii) invalidating and marking a used segment storing a previous version of the data, thereby updating the data in the NVM; and selecting a page to defragment according to wear leveling selection criteria, copying valid data from the selected page to a free page, and erasing the selected page; A method of computing, characterized by comprising the above steps.

11. The method according to claim 10, characterized in that the selection of the page for defragmentation, the copying of the valid data from the selected page, and the erasing of the selected page are started in response to all segments of the given page being filled.

12. The method according to claim 10 or 11, characterized in that the wear leveling selection criteria aim to select the page with the smallest number of previous erasures.

13. The method according to claim 10 or 11, characterized in that the wear leveling selection criteria also depend on the count of invalid segments within the page.

Citation Information

Patent Citations

  • Method and device for storing and controlling data of external memory using plural flash memories

    JP1999126488A

  • terminal

    JP2002032256A

  • Memory system

    JP2012141944A

  • Memory system and control method

    JP2018142236A

  • Data storage device

    JP2018156452A