Controller, method of using controller, and system including controller
By using artificial neural networks in the computer memory subsystem to generate a collection of data prefetch parameters, the problem of inaccurate prediction of data access mode in cache prefetch is solved, and lower memory access latency and higher system performance are achieved.
Patent Information
- Application Number
- CN202510102579.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-07
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to accurately predict data access patterns in cache prefetching, resulting in memory access latency and poor system performance.
An artificial neural network is used to generate M sets of data prefetch parameters, and retrieve data from nonvolatile memory by selecting N sets, and prefetch the prefetched data into the cache. The input to the artificial neural network includes the logical data unit address segment identification or memory read delay for the current data request.
By accurately predicting data access patterns, it significantly reduces memory access latency, improves system performance, and enhances cache hit rate.
Smart Images

Figure CN120029943A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an artificial neural network for improving cache prefetch performance in a computer memory subsystem. Background Art
[0002] Cache prefetching helps reduce latency in accessing data by proactively fetching data from memory to cache before the host (e.g., processor) actually needs it. By accurately predicting data access patterns and prefetching required data into cache, the processor can significantly reduce the time spent waiting for data, thereby speeding up the execution of instructions. This can result in improved overall system performance, reduced memory access latency, and increased throughput. Summary of the invention
[0003] A controller is disclosed herein, which is configured to: generate M sets of data prefetch parameters based on a current data request for data from a non-volatile memory using an artificial neural network, where M is a positive integer; select N sets from the M sets, where N is a non-negative integer not greater than M; retrieve data from the non-volatile memory based on the N sets; and prefetch the prefetched data as part or all of the retrieved data to a cache. At least one of the inputs of the artificial neural network when generating the M sets is (A) a current logical data unit address (LDA) segment identifier (ID) of a logical data unit address (LDA) segment containing part or all of the data requested by the current data request or (B) a memory read latency of the non-volatile memory.
[0004] Each of the M sets of data prefetch parameters may include: a predicted starting logical block address (LBA); and a predicted input / output (I / O) size.
[0005] Both M and N may be equal to 1.
[0006] M can be greater than 1, and the controller can be configured to select N sets from the M sets by the following operations: causing the artificial neural network to generate, for each set in the M sets, a cache hit probability that data corresponding to the each set is requested by a future data request; and selecting a set in the M sets whose cache hit probability exceeds a pre-specified probability threshold, thereby achieving the selection of N sets from the M sets.
[0007] The controller may be configured to implement an artificial neural network in generating the M sets of data pre-fetch parameters.
[0008] The non-volatile memory may be a flash memory.
[0009] The artificial neural network may be a feedforward neural network, a reinforcement learning network, a long short-term memory network, a recurrent neural network, a transformer model, or any combination thereof.
[0010] The controller may be on a single semiconductor die.
[0011] The input to the artificial neural network may be selected from the group consisting of: a current application ID of an application making a current data request, a current LDA segment ID, a current starting LBA of data requested by a current data request, a current I / O size of data requested by a current data request, a memory read latency of a non-volatile memory, and any combination thereof.
[0012] The input to the artificial neural network may be selected from the group consisting of: a current application ID of an application making a current data request, a current LDA segment ID, a current starting LBA of data requested by the current data request, a current I / O size of data requested by the current data request, a memory read latency of the non-volatile memory, a partition ID associated with the current data request, a placement identifier associated with the current data request, a namespace ID associated with the current data request, and any combination thereof.
[0013] The controller may include: an LBA to LDA converter configured to convert the LBA to an LDA of the nonvolatile memory; and an LDA to physical data address (PDA) converter configured to convert the LDA to the PDA of the nonvolatile memory.
[0014] The controller may be configured to determine whether the cache contains the data requested by the current data request.
[0015] The LDA space of the non-volatile memory may include non-overlapping LDA segments of different sizes.
[0016] The cache may be part of a solid state drive (SSD) that includes a controller.
[0017] The cache may be part of the controller.
[0018] The controller may be part of a solid state drive (SSD), flash drive, motherboard, processor, computer, server, gaming device, or mobile device.
[0019] A method of using a controller may include: generating, with the controller, M sets of data prefetch parameters using an artificial neural network based on a current data request for data from a non-volatile memory; selecting N sets from the M sets; retrieving data from the non-volatile memory based on the N sets; and prefetching prefetched data that is part or all of the retrieved data to a cache. At least one of the inputs to the artificial neural network when generating the M sets is (A) a current LDA segment ID of an LDA segment containing part or all of the data requested by the current data request or (B) a memory read latency of the non-volatile memory. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A computer system according to an embodiment is schematically illustrated.
[0021] Figure 2 A diagram for memory access according to an embodiment is shown.
[0022] Figure 3 and Figure 4 Diagram illustrating cache operation according to various embodiments.
[0023] Figure 5 A flow chart outlining cache prefetch operations according to an embodiment is shown. DETAILED DESCRIPTION
[0024] Computer system 100
[0025] Figure 1 A computer system 100 according to an embodiment is schematically shown. The computer system 100 may include a host 110 (eg, a processor), a cache 120, a controller 130, and a non-volatile memory 140.
[0026] Cache 120
[0027] The cache 120 may operate as a high-speed access layer between the host 110 and a memory subsystem 130 + 140 including a controller 130 and a nonvolatile memory 140 .
[0028] Controller 130
[0029] The controller 130 can manage the data flow between the non-volatile memory 140 and the host 110, thereby ensuring efficient and timely access to the stored information. The controller 130 can coordinate the reading and writing of data relative to the non-volatile memory 140, thereby converting the data requests of the host 110 into memory operations. The controller 130 can implement caching, cache prefetching, and pipelining. The controller 130 can perform error detection and correction, power management, and overall system stability.
[0030] Note that in this patent application, “data” refers to information stored or to be stored in the non-volatile memory 140 and / or instructions to the controller 130 .
[0031] The controller 130 may be part of a solid state drive (SSD), a flash drive, a motherboard, a processor, a computer, a server, a gaming device, or a mobile device (not shown).
[0032] The cache 120 may be part of a solid state drive (SSD) including the controller 130. The cache 120 may be part of the controller 130.
[0033] The controller 130 may be formed on a single semiconductor die.
[0034] Non-volatile memory 140
[0035] The nonvolatile memory 140 may be a flash memory. Flash memory is a nonvolatile storage device that retains data even when power is removed. Flash memory technology utilizes floating gate transistors or charge trapping technology to store data.
[0036] LBA, LDA, and PDA in memory access
[0037] The host 110 may use a logical block address (LBA) to specify a logical address of a data block, which may be 64 bytes, 128 bytes, 256 bytes, or 512 bytes of data stored in the nonvolatile memory 140. However, in order to improve performance, the data unit in the nonvolatile memory 140 may be 256 bytes, 512 bytes, 1K bytes, 2K bytes, 4K bytes, or even larger. Therefore, a logical data unit address (LDA) may be used to specify the logical address of the data unit in the nonvolatile memory 140. Therefore, there is a size mismatch between the data block used by the host 110 and the data unit used by the nonvolatile memory 140.
[0038] In order to access a physical unit of the nonvolatile memory 140 to read / write a data unit, a physical data address (PDA) may be used to specify a physical address of a data unit in the nonvolatile memory 140 .
[0039] For a read command from the host 110 to read a data block, the controller 130 may convert the LBA of the data block into an LDA and then convert the LDA into a PDA by looking up a table (not shown) to retrieve the data unit from the nonvolatile memory 140 .
[0040] After retrieving the data unit associated with the LBA from the nonvolatile memory 140 , since the retrieved data unit consists of a plurality of data blocks, the controller 130 may extract a requested data block from the retrieved data unit using the LBA.
[0041] As an example, refer to Figure 1 and Figure 2 , host 110 requests a data block with logical address LBA3 (i.e., the input / output (I / O) size is 1 data block). Assume that the requested data is not in cache 120. In response, controller 130: (A) converts LBA3 to LDA0 using an LBA to LDA converter; and (B) converts LDA0 to physical address PDA9527 using an LDA to PDA converter. Controller 130 then retrieves the data unit at physical address PDA9527 from non-volatile memory 140. Controller 130 then extracts a data block from the retrieved data unit based on LBA3 and sends the extracted data block to host 110.
[0042] LDA section in memory access
[0043] The LDA space can be divided into non-overlapping LDA segments, and each LDA segment can be assigned a unique LDA segment identifier (ID). Figure 2 , the data unit (8 data blocks) at the logical address LDA0 may be an LDA segment, and the data unit (8 data blocks) at the logical address LDA1 may be another LDA segment. An LDA segment may contain at least one LDA address.
[0044] The controller 130 may convert the LBA from the host 110 into LDA to access storage space within a specific LDA segment.
[0045] The controller 130 can group data from applications with similar I / O behavior into the same LDA segment. This minimizes the write amplification factor and enhances system performance.
[0046] Non-overlapping LDA segments can have the same size or different sizes.
[0047] Cache Operations
[0048] refer to Figures 1 to 3 When the host 110 requests data in the non-volatile memory 140 (referred to as a current data request), the host 110 may specify to the controller 130: (A) the current starting LBA of the requested data (LBA3 in the above example); and (B) the current I / O size of the requested data (1 data block in the above example).
[0049] In response, controller 130 may check cache 120 using the current starting LBA and the current I / O size (see Figure 3 2 arrows going from southwest to box 350).
[0050] If the requested data is found in cache 120 (cache hit), Figure 3 ), the requested data may be sent from the cache 120 to the host 110.
[0051] If the requested data is not found in cache 120 (a cache miss), Figure 3 354 in FIG. 1 ), the controller 130 may: (A) retrieve (in the form of (one or more) data units) from the non-volatile memory 140, Figure 3 (B) extracting the requested data from the retrieved data unit(s); (C) sending the requested data to the host 110; and (D) storing a copy of the retrieved data unit(s) in the cache 120 for potential future use (see Figure 3 14. (the arrow in FIG. 14 from box 356 to cache 120).
[0052] Cache prefetch operations
[0053] The controller 130 may perform cache prefetching to improve cache hit rates by predicting data or instructions that may be needed in the near future and proactively bringing such data and instructions from the nonvolatile memory 140 into the cache 120 .
[0054] Specifically, refer to Figures 1 to 3 , when the host 110 requests data in the nonvolatile memory 140 (ie, a current data request), the host 110 may specify the following to the controller 130:
[0055] (i) the current starting LBA of the requested data (LBA3 in the above example),
[0056] (ii) the current I / O size of the requested data (1 block in the above example), and
[0057] (iii) Current application identification (ID), which is the ID of the application making the current data request.
[0058] In response, the controller 130 may send the following inputs to the prefetch prediction engine 310 (eg Figure 3 shown):
[0059] (A) Current application ID,
[0060] (B) the current LDA segment ID (i.e., the ID of the LDA segment that contains part or all of the data requested by the current data request),
[0061] (C) the current starting LBA (i.e., the LBA of the starting data block of the data requested by the current data request),
[0062] (D) the current I / O size (i.e., the size of the data requested by the current data request), and
[0063] (E) Memory read latency of the non-volatile memory 140 (ie, the amount of time it takes to retrieve data from the non-volatile memory 140 after a read request is initiated).
[0064] The prefetch prediction engine 310 may use an artificial neural network (not shown) to generate a predicted starting LBA and a predicted I / O size of data to be prefetched to the cache 120 based on the above-mentioned inputs (A), (B), (C), (D), and (E). Specifically, the artificial neural network may receive the above-mentioned inputs (A), (B), (C), (D), and (E) as inputs, and generate a predicted starting LBA and a predicted I / O size as outputs based on the inputs (A), (B), (C), (D), and (E).
[0065] Next, the controller 130 may use the LBA to LDA converter 314 and the LDA to PDA converter 316 to retrieve (ie, read) the data unit(s) in the non-volatile memory 140 based on the predicted starting LBA and the predicted I / O size (see Figure 3 318 in FIG.
[0066] Next, the controller 130 may prefetch (A) the retrieved data unit(s) or (B) a portion of the retrieved data unit(s) based on the predicted starting LBA and the predicted I / O size (which is the predicted request data) into the cache 120 (see Figure 3 320 in FIG.
[0067] Artificial Neural Networks
[0068] The artificial neural network used by the prefetch prediction engine 310 when generating a predicted starting LBA and a predicted I / O size may include an input layer, a hidden layer, and an output layer, wherein the input layer receives the above-mentioned inputs (A), (B), (C), (D), and (E) as inputs, and the output layer generates a predicted starting LBA and a predicted I / O size as outputs.
[0069] The artificial neural network may be a feedforward neural network, a reinforcement learning network, a long short-term memory network, a recurrent neural network, a transformer model, or any combination thereof.
[0070] refer to Figures 1 to 4 , the prefetch prediction engine 310 may be part of the controller 130. As a result, the controller 130 implements an artificial neural network.
[0071] Alternative embodiments for cache prefetching
[0072] In the embodiments described above, reference is made to Figures 1 to 3 , the prefetch prediction engine 310 uses an artificial neural network to generate a single set of two data prefetch parameters: predicted starting LBA and predicted I / O size.
[0073] In an alternative embodiment, reference Figure 4 , the prefetch prediction engine 310 may use an artificial neural network to generate M sets of data prefetch parameters based on inputs (A), (B), (C), (D), and (E), where M is an integer greater than 1 (see Figure 4 310 in FIG.
[0074] Next, the controller 130 can select N sets from the M sets, where N is a non-negative integer not greater than M (see Figure 4 312 in FIG.
[0075] Next, if N>0, the controller 130 may retrieve data from the non-volatile memory 140 based on the N sets (see Figure 4 ), and then pre-fetching a portion or all of the retrieved data into cache 120 (see Figure 4 320 in FIG.
[0076] Specifically, each of the M sets of data prefetch parameters may include: (A) a predicted starting LBA of data from the non-volatile memory 140 to be prefetched into the cache 120; (B) a predicted I / O size of the data from the non-volatile memory 140 to be prefetched into the cache 120; and (C) a cache hit probability, wherein the cache hit probability is a probability that data corresponding to the predicted starting LBA and predicted I / O size of the respective set is requested by a future data request made by the host 110 (see Figure 4 310 in FIG.
[0077] In order to select N sets from M sets, the controller 130 may select a set whose cache hit probability exceeds a pre-specified probability threshold among the M sets, thereby achieving the selection of N sets from the M sets (see Figure 4 312 in FIG.
[0078] To retrieve data from the nonvolatile memory 140 based on the N sets, for each of the N sets, the controller 130 may retrieve data from the nonvolatile memory 140 based on the predicted starting LBA and predicted I / O size of the respective set (see Figure 4 314, 316 and 318 in FIG.
[0079] Flowchart outlining the operation of controller 130
[0080] Figure 5 is a flow chart 500 outlining cache prefetch operations according to an embodiment.
[0081] In step S510, the operation may include: based on the current data request for data from the non-volatile memory, using the controller to generate M sets of data pre-fetch parameters using an artificial neural network, where M is a positive integer. For example, in the above embodiment, referring to Figures 1 to 4 , the controller 130 uses an artificial neural network to generate M sets of data prefetch parameters based on the current data request for data from the non-volatile memory 140, where M is a positive integer (M=1 corresponds to Figure 3 ; and M>1 corresponds to Figure 4 ).
[0082] In step S520, the operation may include: selecting N sets from the M sets. For example, in the above embodiment, referring to Figures 1 to 4 , the controller 130 selects N sets from the M sets.
[0083] In step S530, the operation may include: retrieving data from the non-volatile memory based on the N sets. For example, in the above embodiment, referring to Figures 1 to 4 , the controller 130 retrieves data from the non-volatile memory 140 based on the N sets.
[0084] In step S540, the operation may include: pre-fetching pre-fetched data as part or all of the retrieved data to a cache, wherein at least one of the inputs of the artificial neural network when generating the M sets is: (A) a current LDA segment ID of an LDA segment containing part or all of the data requested by the current data request; or (B) a memory read latency of the non-volatile memory. For example, in the above embodiment, referring to Figures 1 to 4, the controller 130 pre-fetches part or all of the retrieved data to the cache 120, wherein at least one of the inputs to the artificial neural network when generating the M sets is: (A) the current LDA segment ID of the LDA segment containing part or all of the data requested by the current data request; or (B) the memory read latency of the non-volatile memory 140.
[0085] Other embodiments
[0086] ZNS Support
[0087] refer to Figures 1 to 4 , the memory subsystem 130+140 may support a partitioned namespace (ZNS). As a result, when generating M sets of data prefetch parameters (M is a positive integer, where M=1 corresponds to Figure 3 , and M>1 corresponds to Figure 4 ), the input to the artificial neural network may include a partition ID associated with the current data request.
[0088] FDP Support
[0089] refer to Figures 1 to 4 , the memory subsystem 130+140 can support flexible data placement (FDP). As a result, when generating M sets of data prefetch parameters (M is a positive integer, where M=1 corresponds to Figure 3 , and M>1 corresponds to Figure 4 ), the input to the artificial neural network may include a placement identifier and a namespace ID associated with the current data request.
[0090] Controller Functions
[0091] refer to Figures 1 to 4 , LBA to LDA converter 314 ( Figure 3 and Figure 4 ) can be part of the controller 130.
[0092] refer to Figures 1 to 4 , LDA to PDA converter 316 ( Figure 3 and Figure 4 ) can be part of the controller 130.
[0093] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the appended claims.
[0094] CROSS-REFERENCE TO RELATED APPLICATIONS
[0095] This application claims priority to U.S. Provisional Application No. 63 / 602,411, filed on November 23, 2023, the entire disclosure of which is incorporated herein by reference.
Claims
1. A controller, configured to: Based on a current data request for data from a non-volatile memory, an artificial neural network is used to generate M sets of data pre-fetch parameters, where: M is a positive integer; Select N sets from the M sets, where N is a non-negative integer not greater than M; retrieving data from the non-volatile memory based on the N sets; and prefetching data as part or all of the retrieved data into a cache, Among them, when generating the M sets, at least one of the inputs of the artificial neural network is (A) the current LDA segment identifier, i.e., the current LDA segment ID, of the logical data unit address segment, i.e., the LDA segment, which contains part or all of the data requested by the current data request, or (B) the memory read delay of the non-volatile memory.
2. The controller according to claim 1, wherein: Each of the M sets of data pre-fetch parameters includes: Predicting the starting logical block address, i.e. predicting the starting LBA; and Predicting input / output size is predicting I / O size.
3. The controller according to claim 1, wherein: M=N=1.
4. The controller according to claim 1, in, M>1, and The controller is configured to select the N sets from the M sets by performing the following operations: causing the artificial neural network to generate, for each of the M sets, a cache hit probability that data corresponding to the each set is requested by a future data request; and A set among the M sets whose cache hit probability exceeds a pre-specified probability threshold is selected, thereby achieving selection of the N sets from the M sets.
5. The controller according to claim 1, wherein: The controller is configured to implement the artificial neural network when generating the M sets of data pre-fetch parameters.
6. The controller according to claim 1, wherein: The non-volatile memory is a flash memory.
7. The controller according to claim 1, wherein: The artificial neural network is a feedforward neural network, a reinforcement learning network, a long short-term memory network, a recurrent neural network, a transformer model or any combination thereof.
8. The controller according to claim 1, wherein: The controller is on a single semiconductor die.
9. The controller according to claim 1, wherein: The input of the artificial neural network is selected from the group consisting of: the current application ID of the application making the current data request, the current LDA segment ID, The current starting LBA of the data requested by the current data request, the current I / O size of the data requested by the current data request, a memory read latency of the non-volatile memory, and Any combination thereof.
10. The controller according to claim 1, wherein: The input of the artificial neural network is selected from the group consisting of: the current application ID of the application making the current data request, the current LDA segment ID, The current starting LBA of the data requested by the current data request, the current I / O size of the data requested by the current data request, a memory read latency of the non-volatile memory, the partition ID associated with the current data request, a placement identifier associated with said current data request, the namespace ID associated with the current data request, and Any combination thereof.
11. The controller according to claim 1, comprising: an LBA to LDA converter configured to convert the LBA to the LDA of the non-volatile memory; as well as The LDA to physical data address converter, ie, the LDA to PDA converter, is configured to convert the LDA to the PDA of the non-volatile memory.
12. The controller according to claim 1, wherein: The controller is configured to determine whether the cache contains data requested by the current data request.
13. The controller according to claim 1, wherein: The LDA space of the non-volatile memory includes non-overlapping LDA segments of different sizes.
14. The controller according to claim 1, wherein: The cache is part of a solid state drive (SSD) that includes the controller.
15. The controller according to claim 14, wherein: The cache is part of the controller.
16. A system comprising the controller according to claim 1, wherein: The system is a solid state drive, ie, SSD, a flash drive, a motherboard, a processor, a computer, a server, a gaming device, or a mobile device.
17. A method of using the controller according to claim 1, comprising: generating, with the controller, M sets of the data pre-fetch parameters using the artificial neural network based on a current data request for data from the non-volatile memory; Selecting the N sets from the M sets; retrieving data from the non-volatile memory based on the N sets; as well as prefetching prefetched data as part or all of the retrieved data into the cache, Wherein, when generating the M sets, at least one of the inputs of the artificial neural network is (A) the current LDA segment ID of the LDA segment containing part or all of the data requested by the current data request or (B) the memory read delay of the non-volatile memory.
18. The method according to claim 17, wherein: Each of the M sets of data pre-fetch parameters includes: Predicting the starting LBA; and Predict I / O size.
19. The method according to claim 17, wherein: M=N=1.
20. The method according to claim 17, in, M>1, and Wherein, selecting the N sets from the M sets includes: causing the artificial neural network to generate, for each of the M sets, a cache hit probability that data corresponding to the each set is requested by a future data request; and A set among the M sets whose cache hit probability exceeds a pre-specified probability threshold is selected, thereby achieving selection of the N sets from the M sets.