Metadata caching integrated circuit device

By introducing metadata cache in the memory control device, separating the storage and processing of user data and metadata, the problem of low utilization and bandwidth efficiency of external memory in the prior art is solved, and more efficient data writing and reading operations are achieved.

CN120035817AActive Publication Date: 2025-05-23ASTERA LABS INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202380069265.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-10
Filing Date
2023-07-27
Publication Date
2025-05-23
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

The prior art uses and bandwidth efficiency of external memory when processing data write and read operations, especially when there are separate write transactions and read-revise-write operations between metadata storage and user data storage.

Method used

By implementing metadata cache in the memory control device, separating the storage and processing of user data and metadata, the user data is stored in external memory, while the metadata is stored in the metadata cache, and the storage and reading operations are optimized through the composite write data words and the composite read data words.

Benefits of technology

It effectively reduces the number of write transactions in external memory, improves the utilization of memory bandwidth, reduces power consumption, and improves the overall performance of the data processing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005331164590000011
    Figure HDA0005331164590000011
  • Figure HDA0005331164590000012
    Figure HDA0005331164590000012
  • Figure HDA0005331164590000021
    Figure HDA0005331164590000021
Patent Text Reader

Abstract

A memory control device implements split storage of user data components and metadata components of a composite write data word that outputs the user data components via a memory control interface for storage within an external memory subsystem while separately storing the metadata components within a metadata cache implemented within the memory control device.
Need to check novelty before this filing date? Find Prior Art

Description

BRIEF DESCRIPTION OF THE DRAWINGS

[0001] Various embodiments disclosed herein are illustrated by way of example and not limitation in the accompanying drawings and in which like reference numerals refer to similar elements and in which:

[0002] Figure 1 An embodiment of a data processing system having a host device, a metadata cache memory control device, and an external memory subsystem is illustrated;

[0003] Figure 2 The diagram shows how Figure 1 A more detailed embodiment of a metadata manager that is deployed as the metadata manager shown;

[0004] Figure 3 The diagram shows Figure 1 and Figure 2 The exemplary explicit storage operation mode implemented within the metadata manager and metadata cache is shown;

[0005] Figure 4 Picture shows Figure 1 / Figure 2 operation of the metadata manager in an inference mode in which no metadata values ​​matching default metadata values ​​are stored, such that as long as the metadata cache maintains a unique metadata storage location, any metadata cache miss enables a deterministic inference that the metadata sought by the host is the default metadata value;

[0006] Figure 5 The diagram shows the reference Figure 3 and Figure 4 discussed and can be Figure 1 / Figure 2 Embodiments of a metadata cache capable of performing search, load, and entry invalidation operations deployed like the metadata cache of the embodiments;

[0007] Figure 6 Picture shows Figure 5 illustrative block registers and register input multiplexers of , showing examples of their interconnection with a metadata-cache block store and an output multiplexer; and

[0008] Figure 7 The diagram corresponds to Figure 5 An exemplary control signal diagram for the metadata cache operations listed in . DETAILED DESCRIPTION

[0009] In various embodiments herein, a memory control device implements ramified / split storage of user-data and metadata components of a composite write data word, the device outputting the user-data component via a memory control interface for storage in an external memory subsystem, while storing the metadata component separately in a metadata cache implemented in the memory control device. In multiple embodiments, the user-data component of the write data word matches the native read / write width of the memory control interface, so that in-situ metadata storage (i.e., in the metadata cache) avoids separate write transactions to the external memory system (which would otherwise be required to store the user-data component and the metadata component), thereby reducing external memory utilization per memory write transaction by 50% (and possibly by 66% in the case where storage of relatively small metadata components requires read-modify-write and thus requires two external memory accesses per metadata write), and correspondingly increasing memory bandwidth availability by 100% (or 200% in the case where read-modify-write transactions to external memory are avoided). Memory read transactions are similarly bifurcated, such that the metadata component of the composite read data word is fetched from the metadata cache and the user data component of the read data word is fetched from the external memory (reducing the external memory transaction count by 50% and correspondingly increasing the external memory bandwidth by 100%). In various embodiments, the metadata cache ("m-cache") is implemented with a block size (read / write width) that matches the user data component of the read / write data word, and is therefore capable of storing a number of metadata components corresponding to the ratio between the sizes of the user data component and the metadata component (i.e., in the primary case where the user data component is larger than the metadata component by a factor of N, each cache block may store N metadata values). Thus, when the user data component of a read / write data word is read from or written to the external memory, the corresponding metadata value is fetched (or written) in the operation of reading the entire cache block (and thus reading the multiple metadata values ​​corresponding to the corresponding data word and its user data component) from the metadata cache - during metadata writes, the incoming metadata is merged into the address-specified subfield of the cache block, and during metadata reads, the outbound metadata is extracted from the address-specified subfield of the cache block. In such an embodiment, each cache block may store multiple entry valid bits (i.e., one entry valid bit for each metadata subfield), where a cache hit is determined not only by matching the tag of the incoming physical address with the cache block index field, but also by selecting and evaluating the entry valid bit of the offset-specified subfield of the indexed cache block.In these and other embodiments, a primary metadata value (i.e., one that occurs more frequently than other metadata values) may constitute a "default" metadata value that is (i) not actually / explicitly stored within the metadata cache and (ii) returned as a metadata component of a read data word in response to a metadata-cache miss. Through this inferred / implicit metadata approach, the size of the metadata cache and search power consumption may be reduced in approximate proportion to the default metadata prevalence (e.g., a system in which 90% of metadata values ​​are default metadata values ​​would have 90% fewer explicitly stored metadata values, resulting in a 90% reduction in the capacity / size of the metadata cache to have the same performance as a 10x larger cache for a system in which all metadata values ​​are explicitly stored). These and other features and embodiments are discussed in more detail below.

[0010] Figure 1 A data processing system 100 is illustrated having a host device 101, a metadata cache memory control device 103, and a memory subsystem 105, which is referred to herein as "external" memory given that the memory subsystem is implemented separately from the control device 103. In the depicted example, the host device 101 includes core functional circuitry 111 (e.g., a central processing unit, a graphics processing unit, a digital signal processing unit, a neural network, etc.) and a communication interface 113 (COM), the core functional circuitry optionally including a data cache 115, which, as discussed below, can affect implementation and / or operational aspects of a metadata cache within the control device 103. The communication interface 113 enables physical layer interconnection with a corresponding communication interface 121 within the control device 103, over which standardized or proprietary communication protocols can be layered, including, in various embodiments, cache coherence protocols that support memory read and write semantics (e.g., Compute Express Link (CXL), Gen-Z, Open CAPI, etc.) and require a specified number of metadata bits to be transferred with each user data transfer. That is, the host device 101 may issue read and write instructions and corresponding physical addresses to the control component 103, and in response, the control component 103 processes these read and write requests via bifurcated access to the in-situ (e.g., on-chip or in-package) metadata cache 123 and external memory as discussed above. In the case of CXL and possibly other protocols, the host device 101 may issue "host" physical addresses (HPA) that are translated / converted into device physical addresses (DPA) within the communication interface 121 (or by other circuitry within the metadata cache memory control device). References herein to incoming DPAs within the metadata cache memory control device 103 may include such HPA to DPA translation / conversion, where necessary.

[0011] In addition to the above-mentioned host-side communication interface 121 and metadata cache 123 ("m-cache"), the metadata cache control component also includes a metadata management engine 125 ("metadata manager") and a memory control interface 127 (MC), which is coupled to the external memory 105. For the purpose of example herein, it is assumed that the external memory is a dynamic random access memory (DRAM) accessible via a command / address signaling path (CA) and a bidirectional data path (DQ). In general, the data path coupled between the memory control interface 127 and the external memory 105 has a bit width that matches the width of the user data component of the write data word / read data word received / returned to the host device 101 (or vice versa), and it itself can be coupled to multiple discrete memory components (for example, discrete DRAM packages arranged on a dual in-line memory module (DIMM) and coupled to corresponding slices of the DQ path), accompanied by selection and / or other timing signal lines, including data mask control, etc. More generally, the external memory 105 may be implemented in or utilizing a variety of different form-factors, interconnect topologies, core storage technologies (e.g., any combination of flash memory or various other non-volatile memory technologies, static random access memory (SRAM), etc. (including hybrid combinations including DRAM)), access protocols, etc., and in some instances may be integrated with the metadata cache memory control device 103 in a multi-chip package. Further, although reference is made to various examples / embodiments discussed below, Figure 1 105, but in all cases, the memory control interface 127 (or other memory interface circuitry) can provide control / access to internal memory as an alternative or in addition to the external memory 105 - for example, a relatively larger memory (e.g., an SRAM having a storage capacity that is 2 times, 4 times, 8 times, 16 times, 32 times, 100 times, 1000 times, 10,000 times, 100,000 times, or more as large as the m-cache 123) implemented on the same IC die as the control device 103 or within the same multi-IC package as the control device.

[0012] Still refer to Figure 1, the read / write request issued by the host device 101 is accompanied by a device physical address (DPA) that specifies the physical address within the external memory to be accessed (for data retrieval or storage), and in the case of a memory write request, the request is also accompanied by a composite or synthetic write data word that has the above-described user data component and metadata component—the inbound user data word and the corresponding inbound metadata. In the depicted example, the control component 103 responds to a memory read request by performing a fork, concurrent (at least partially overlapping in time) metadata and user data retrieval operations, issuing a memory read command via the memory control interface 127 to retrieve the user data component from the external memory 105, and issuing a cache search command (providing the DPA therewith) to the metadata cache 123 to determine whether the requested metadata value is stored in the metadata cache, and if so, retrieving the metadata component from the cache (avoiding resorting to the external memory and thus saving power / memory bandwidth, and in some embodiments, obviating the need to reserve a portion of the memory capacity for metadata storage). More specifically, the metadata cache 123 responds to a search instruction from the metadata manager 125 by returning a cache hit / miss signal ("hit") depending on whether the DPA indexes to a valid metadata entry (i.e., the hit signal is asserted if it indexes to a valid entry and de-asserted if it does not index to a valid entry), and if so, outputting the metadata value to the MD manager (i.e., the metadata input / output is shown as "MD" in Figure 1 . The metadata manager responds to an m-cache hit (the hit signal is asserted) by returning the metadata value provided by the cache, together with the user data retrieved from the external memory, to the host component—the metadata and user data form the respective components of a composite or synthetic read data word returned to the host (i.e., shown as "U-data + M-data output" in the COM path to the host)—buffering / queuing the metadata as needed to await the return of the corresponding read data from the external memory so that the two data components can be returned as a combined or at least temporally unified data output word. In an alternative embodiment, either data component (metadata or user data) can be returned to the host when available, while the other component can be returned at a later time (e.g., returning the metadata to the host while the user data is being retrieved from the external memory 105 or other storage).

[0013] Control component 103 responds to host data write requests with similar bifurcating actions - storing the metadata component ("M-Data In") of the incoming write data word in metadata cache 123 and storing the user data component ("U-Data In") in external memory. In various embodiments, metadata storage within m-cache 123 begins with the same search operation as metadata readout - the host-issued DPA is provided to m-cache 123 along with a search instruction to determine whether m-cache 123 already contains a valid entry corresponding to the DPA, overwrite the entry if so (m-cache hit), and take one of various optional actions if not (m-cache miss), including creating a new entry within the m-cache in which the inbound metadata value is directly stored, and in some cases fetching a cache block entry corresponding to the DPA from external memory and storing the inbound metadata value in the m-cache as part of the cache block load.

[0014] In various embodiments, the metadata cache is used as the sole / sole metadata storage—all metadata read / write requests issued to the controller device 103 are fulfilled only by the metadata cache, and there is no metadata storage in the external memory. In some cases, such m-cache-only (MCO) metadata storage is ensured by sizing the m-cache to provide a dedicated metadata storage location for each user data location (memory row address) in the external memory (i.e., a different m-cache metadata entry for each valid DPA). In other embodiments, particularly where metadata is used to ensure consistency between a host data cache (e.g., as shown at 115) and data stored in the external memory 105, MCO operation may be ensured by sizing the metadata cache to include a different m-cache entry for each user data entry in the host cache—in this arrangement, host cache evictions will occur before eviction-triggering conflicts in the m-cache 123, thereby ensuring that m-cache evictions do not occur, and deterministic metadata storage is achieved only in the metadata cache. In other embodiments, the m-cache capacity is insufficient to meet the worst-case metadata storage requirements (i.e., cannot be used as the sole metadata repository), so m-cache conflicts may occur (i.e., attempting to store metadata at a DPA-indexed cache location occupied by metadata corresponding to a different DPA) and metadata needs to be evicted from the m-cache to external memory. In such embodiments, as shown at 131, a portion of the external memory address space may be reserved for metadata storage (e.g., sufficient to accommodate worst-case / maximum metadata storage).

[0015] Figure 2 The metadata manager 150 (i.e., Figure 1 1) is a more detailed embodiment of a metadata manager 125 of a memory controller 150 having a finite state machine 151, and memory-facing command, address, and data multiplexers (153, 155, 157, respectively), address shift logic 159, and outbound metadata multiplexer 161. The MD manager 150 receives commands ("Cmd") and memory addresses (device physical addresses DPA) provided by the host via the host communication interface 121, receives write data words via the same interface, and returns read data words. More specifically, when the user component and metadata component (WrDat, WrMD) of a write data word arrive via the host interface in association with a write command issued by the host, the metadata component and DPA are forwarded to the metadata cache 123, while the user data component and DPA are forwarded to the external memory via the memory control interface 127 through the data multiplexer 157 and the address multiplexer 155 (and the address shift logic 159). The incoming write command is provided to both the finite state machine 151 and the memory control interface 127 (provided to the memory control interface via the command multiplexer 153), where the memory control interface responsively issues a corresponding write instruction (and DPA) to the external memory subsystem and utilizes the memory control interface to forward the user data component of the write data word (e.g., at a specified timing, issuing control and timing signals as needed to achieve user data transfer and storage).

[0016] The state machine 151, which may alternatively be implemented by a sequencer, a processor, or any other feasible control circuit, responds to the incoming write command by initiating a DPA search within the m-cache 123, and more specifically by issuing a search instruction via the instruction path "Instr". If a cache hit occurs (the hit signal is asserted by the m-cache 123), the state machine 151 instructs the m-cache (issues an instruction signal to it) to merge the inbound metadata value (WrMD) into the cache block indexed by the DPA, storing the inbound metadata value in a subfield of the cache block corresponding to the offset bits of the DPA. If a cache miss occurs, the state machine 151 either performs a cache block load operation (i.e., fetches a data block containing a metadata entry corresponding to the DPA from an external memory and loads the block of metadata entries into the m-cache 123), or instructs the m-cache to merge the inbound metadata value into the empty cache block (i.e., not load from the external memory) if the m-cache indicates that the cache block indexed by the DPA is empty (i.e., the "Empty" signal is asserted).

[0017] exist Figure 2In the embodiment of the present invention, each m-cache search operation, together with the above-mentioned hit and miss signals, generates an outbound cache black (oCB), which contains all metadata entries corresponding to the index and tag fields of the DPA (any or all of these metadata entries may or may not include valid metadata). Therefore, an m-cache miss with respect to a non-empty cache block (both the hit signal and the empty signal are deasserted after the m-cache search) indicates that the outbound cache block includes at least one valid metadata entry. In this case, in order to avoid losing the valid cache resident metadata when the cache block is loaded, the state machine 151 stores the outbound cache block in the external memory and implements an eviction operation (evicts the resident cache block to the external memory) to prepare for the subsequent loading of the cache block corresponding to the DPA provided by the host. More specifically, FSM 151 asserts an internal-commandenable signal (iCen) to route memory write commands provided by the FSM to memory controller interface 127 via command multiplexer 153, and also asserts an evict-write enable signal (eWen) to route (i) physical addresses provided by the m-cache (evict-DPA or “eDPA”) to the memory controller interface via address multiplexer 155 and (ii) outbound cache blocks (from m-cache 123) to the memory controller interface via data multiplexer 157. As discussed below, the internal-command enable signals are also used to enable address shift logic 159 to generate addresses indexed into metadata reserved address space within external memory (e.g., shifted into metadata address space relative to eDPA via a pointer) and to relatively resize user data components and metadata components of write data words issued by the host. The memory controller interface 127 responds to the write command originating from the FSM by performing a memory write operation to store the outbound cache block at a location in the external memory corresponding to the eDPA (as modified by the address shift logic 159) (e.g., at a fixed offset from the memory location indexed by the eDPA according to the metadata address pointer adMD). Thereafter, the FSM 151 deasserts the evict-write enable signal while continuing to assert iCen and outputs a read command to the memory controller interface (i.e., via the multiplexer) when the DPA provided in the host-provided write command is passed to the memory controller interface via the multiplexer 155 and the iCen enable address shift logic 159, thereby providing a command and address to fetch the cache block corresponding to the host-provided DPA from the external memory. As shown, the fetched cache block (arriving from the memory controller interface 127 via the read data path) is routed to the metadata cache as an inbound cache block (iCB).Thus, at a predetermined time (i.e., based on memory read latency) after initiating the fetch of an inbound cache block from external memory, the FSM 151 issues an instruction signal to load the inbound cache block into the metadata cache 123 at the DPA provided by the host, thereby storing any tag components of the DPA in association with the newly loaded cache block within the m-cache. Although not specifically shown, the metadata manager may include various storage elements to store address values, metadata values, etc. as needed to ensure the availability of data / addresses for block loading, metadata storage, and other operations without otherwise interfering with pipelined memory and m-cache access operations.

[0018] Continue to refer to Figure 2 , FSM 151 also initiates an m-cache search in response to a memory read command issued by the host. An m-cache hit is accompanied by an m-cache provision of a metadata value corresponding to the DPA provided by the host (the cache-provided metadata value is referred to herein as the outbound metadata value oMD), wherein the metadata value is returned to the host (RdMD) along with user data (RdDat) retrieved from external memory - the metadata and user data together forming a read data word, wherein the metadata is buffered as needed to accommodate external memory read latency and thereby enable simultaneous and / or packetized uniform transmission of the read data words to the host via the communications interface 121). If the DPA provided with the host read command misses the m-cache (hit signal deasserted) and indexes into an empty cache block (empty signal asserted), the FSM 151 performs a cache block load as in the case of a metadata write - either extracting the metadata of interest (i.e., requested by the host) from the inbound cache block (since it is returned from external memory), or re-searching the m-cache after the inbound cache block has been loaded into it, the latter of which deterministically produces a cache hit and valid outbound metadata (which constitutes the metadata of interest). If the m-cache miss is accompanied by an occupied cache block signal (i.e., m-cache deasserted hit signal, deasserted empty signal), the FSM may perform the eviction-write operation discussed above to store the resident m-cache block (i.e., oCB) in external memory before loading the DPA-specified cache block into the m-cache.

[0019] As described above, optionally, a high-prevalence default metadata value ("dMD") may not be stored, where an m-cache miss means that the metadata value being searched for is actually the default metadata value (i.e., assuming the host issues a write to a certain address before reading from the same address). More specifically, assuming an embodiment / configuration in which (i) metadata is stored only within the m-cache, and (ii) only metadata values ​​other than the high-prevalence default metadata value are explicitly stored (i.e., metadata values ​​that occur less frequently are stored, and default metadata values ​​that occur more frequently are not stored), a memory read issued by the host will result in an m-cache hit for an infrequently occurring metadata value and a cache miss for the default metadata value. That is, a metadata cache miss means that the metadata sought by the host is the default metadata value. Thus, in such an embodiment / configuration, the FSM 151 responds to an m-cache miss by asserting an enable default signal (“enDef”) to return a default metadata value (e.g., the dMD value programmed within configuration register 165) to the host requestor via oMD multiplexer 161 instead of the outbound metadata value from the cache. Conversely, the FSM 151 deasserts enDef in response to a cache hit to return the cached metadata value (oMD) to the host. In response to a write command issued by the host, the FSM searches the m-cache in a similar manner to determine whether an infrequently occurring metadata value has been explicitly stored (a cache hit), and if so, overwrites the infrequently occurring metadata value or invalidates the metadata entry depending on whether the inbound metadata value is an infrequently occurring metadata value or a default metadata value—operations discussed in more detail below. In the event of a cache block eviction (meaning that m-cache-only metadata storage can no longer be assumed), the FSM can revert to explicitly storing the default metadata value within the m-cache and / or external memory (i.e., external memory in the event of an evicted cache block). This operation will be discussed in further detail below.

[0020] exist Figure 2 In an embodiment of the present invention, configuration register 165 includes various programmable fields, such as, but not limited to: a mode field for enabling one of various operating modes, the aforementioned default metadata field for storing a host-specified default metadata value (dMD), a metadata address / pointer field for establishing a host-defined metadata address space within an external memory (as needed), an offset size field for specifying the number of metadata values ​​(count) per cache block and, in effect, the size of the metadata (the number of constituent bits), and a write policy field for enabling programmable cache write policies (e.g., write-back, write-through, etc.). In one embodiment, such as shown in detailed view 170, the mode field is a multi-bit field that enables the specification of a host-defined metadata address space within an external memory. Figure 1 The metadata cache memory controls at least any of the following operating modes within the device:

[0021] Explicit: All metadata values ​​are explicitly stored in the m-cache and in external memory when necessary due to m-cache conflicts;

[0022] Inference: metadata values ​​matching the programmed (or fixed) dMD are not stored and are inferred on cache misses when the metadata cache is the only metadata repository; all other “infrequently occurring” metadata values ​​are explicitly stored in the metadata cache and in external memory when necessary due to m-cache conflicts;

[0023] • Conditional inference: Same as inference mode (no dMD storage) until an m-cache conflict / overflow requires metadata storage in external memory, then transitions to explicit mode where all metadata values ​​(including dMD) are explicitly stored.

[0024] exist Figure 2 In the example of FIG. 1 , the programmable metadata address and programmable offset size fields (i.e., adMD and szOfst, respectively) within configuration register 165 are selectively applied within address shift logic 159 to generate an address shifted DPA that indexes a metadata address space within the external memory. In one embodiment, as shown in detailed view 180, the offset size field (szOfst) specifies the number of DPA offset bits (N) required to resolve each metadata entry within a cache block (thereby specifying the number of metadata entries per cache block as 2). N ), where the size of the cache block itself is established by a separately programmed value or fixed by system design. As shown, when FSM 151 asserts the internal command enable signal (iCen), multiplexer 181 selects the address shift instance of DPA (i.e., DPA right shifted N bits) added with the programmed metadata address (e.g., within summing circuit 183 (which may be a most significant bit concatenation rather than an explicit addition circuit)) to form the cache block storage address within the external memory - shown offset from DPA by adMD+DPA / 2 N(Integer division) shift DPA (sDPA). Taking a 1 megabyte (TB) external memory with a 512-bit (64-byte) data I / O width (i.e., user data size = memory row size = m-cache block size = 64B) and thus a 34-bit physical address (the least significant 6 bits of the 40-bit byte resolution address are not used in view of the 64-byte memory row size) and a 16-bit metadata size, each 64B m-cache block will store 32 metadata values ​​corresponding to DPAs with the same 29 most significant bits and corresponding offsets in 32 different 5-bit offsets (the 5 least significant bits of the DPA). At 190, an association between such a DPA and 32 corresponding user data values ​​(each of the 32 DPAs having a unique 5-bit offset and the same 29 MSBs) is shown, wherein each such DPA generates within the address shift logic 159, in response to an iCen assertion, a shifted physical address (sDPA) to a cache block storage location within an external memory for storing 32 metadata values ​​corresponding to the 32 user data values. At 191, an example of such a stored cache block (CB) is shown, the cache block containing 32 metadata entries (MD0, MD1, ..., MD31) of each exemplary 5-bit offset field size. Although the foregoing example of a 1TB external memory space with a 64B memory row size, a 64B cache block size, and a 16-bit metadata size is continued in the embodiments discussed below, in all cases, the external memory size, metadata storage address calculation (and implementation circuitry), memory row size, etc. may vary, with any or all of such parameters being present in the example embodiment. Figure 1 The metadata cache memory control device is shown to be programmed within one or more configuration registers (eg, as explicitly shown with respect to szOfst and adMD).

[0025] Figure 3 The diagram shows Figure 1 and Figure 2 1 and 2. An exemplary explicit store mode of operation implemented within the metadata manager and metadata cache is shown - i.e., a mode of operation in which all metadata values ​​are explicitly stored within the metadata cache, or, if desired, within external memory. Upon receiving a read command at 211, the metadata manager (shown at 125 with configuration registers 165) initiates an m-cache search at 213 to determine if the metadata value has already been cached for the incoming DPA. If so (cache hit signal asserted at 215), the metadata value outbound from the m-cache (i.e., oMD) is returned to the host requestor at 217 along with the user data value read back from external memory. As discussed, the metadata manager can buffer the oMD value as needed to enable both the metadata component and the user data component of the returned data word to be returned to the host simultaneously.

[0026] If the search at 213 results in an m-cache miss (negative determination at 215) and the m-cache indicates that the indexed cache block is empty (affirmative determination at 219), the metadata manager fetches the cache block from external memory (e.g., within the address space indexed via the shifted DPA as discussed above) and loads the cache block into the metadata cache, storing the new tag field components of the DPA into the cache (the overall "block load" operation shown at 223). In one embodiment, as shown at 225, the metadata manager extracts metadata from the fetched cache block before (or simultaneously with) loading the cache block into the m-cache, returning the extracted metadata value (along with the user data fetched from external memory) to the host without re-searching the metadata cache. In other embodiments, the metadata manager may instead re-search the m-cache after the cache block load (i.e., the process flow loops back to the cache search 213), thereby deterministically producing a cache hit at 215 and an oMD return at 217.

[0027] If the m-cache signals both a cache miss and that the DPA-indexed cache block is not empty (i.e., negative determinations at 215 and 219), the metadata manager evicts the DPA-indexed cache block (the resident cache block) at 227, thereby writing the resident cache block to the external memory at the eviction address (the eviction address is formed by concatenating the index field of the host-provided DPA with the tag field of the resident cache block stored by the m-cache) (and, at least in accordance with Figure 2 In the embodiment of the present invention, the eviction DPA is shifted within the address shift logic according to the offset field size (szOfst) and the metadata storage pointer (adMD). After the cache block eviction at 227, the cache block load at 223 is performed using either the MD fetch at 225 (without performing an m-cache re-search) or the m-cache re-search at 213 (and then a deterministic cache hit at 215 and an oMD return at 217).

[0028] continue Figure 3In the illustrated explicit store operation mode, the metadata manager responds to a write command issued by the host by initiating an m-cache search at the DPA (233), and in response to a cache hit (affirmation at 235), merges the inbound metadata value (iMD)—arriving with the user data component of the write data word—into the indexed m-cache block at 237, thereby storing the updated cache block in the m-cache. If a cache miss occurs (negative determination at 235) and the cache block indexed by the DPA is empty (affirmation at 239), then at 241, the metadata manager performs an iMD-merged cache block load—generally loading the cache block from external memory as discussed with reference to operation 223, but in the case of a metadata write, merging the iMD with the fetched cache block and then loading it into the m-cache (and also updating the tag field according to the incoming DPA as in the block load at 223). If a cache miss occurs with respect to an occupied / non-empty cache block (i.e., the DPA-indexed cache block contains at least one valid metadata entry, resulting in a negative determination at 239), the metadata manager evicts the resident cache block to external memory at 243 (as discussed with reference to 227) and then performs a cache block load merged with iMD at 241.

[0029] exist Figure 3 In the example of FIG. 1 , the iMD merge operations at 237 (after a cache hit) and 241 (after a cache miss) constitute m-cache write operations that follow a programmable write policy, such as in Figure 2 The wrP field of the configuration register is programmed (see Figure 2 Detailed view 170 of the external memory) and typically includes at least write-through and write-back options. Under the write-back policy setting, updates to the subordinate cache blocks are made within the metadata cache (i.e., to include the incoming metadata value) while deferring updates to the corresponding cache blocks in the external memory (e.g., until an eviction or other event occurs that requires restoration of consistency between the backing store in the external memory and the metadata cache), thereby making the external memory instance of the cache block obsolete and the m-cache instance of the cache block "dirty" (e.g., modified relative to the external memory instance). This loss of consistency between the instance of the cache block and the cache block instance stored in the external memory is eventually resolved, for example, when the m-cache instance of the cache block is evicted.

[0030] Under the write-through strategy, the metadata manager writes incoming metadata values ​​to both the metadata cache and the external memory in response to a write command issued by the same host. In the case of a single-channel memory system (i.e., a single command / address stream from the memory controller to the memory subsystem and a single corresponding data path), the metadata manager may perform a write-through of the cache block to the external memory after completing the bifurcated storage of the user data component and the metadata component of the host-provided write data word (i.e., storing the user data in the external memory and storing the metadata in the m-cache), or may even postpone such a write-through until an unused data access time slot is detected (i.e., so as not to interfere with ongoing host read / write requests). In a multi-channel memory subsystem, metadata write-through may be performed simultaneously or at least in parallel (at least partially overlapping in time) with a user data write via one memory channel via another memory channel.

[0031] Figure 4 Picture shows Figure 1 / Figure 2 The metadata manager of the embodiment of the present invention operates in an inferential mode - i.e., does not explicitly store metadata values ​​that match preprogrammed (or otherwise predetermined) default metadata values, such that as long as the m-cache maintains the only metadata storage location, any metadata cache miss means (i.e., can be deterministically inferred or deterministically indicated) that the metadata value sought by the host is the default metadata value. As discussed above, in the case where m-cache-only operation (i.e., the m-cache is the only metadata storage) is ensured by system design and / or the sheer size of the m-cache, evictions to external memory will not occur or need to be provided. In the more general case where m-cache-only operation cannot be ensured in all cases (i.e., an incoming DPA may resolve to m-cache location(s) already occupied by non-empty cache blocks corresponding to different DPA(s), the metadata manager may evict a cache block to external memory at some point, and thereby create two independent possibilities for any subsequent m-cache misses: (i) the metadata sought is the default metadata value and is therefore not stored in the m-cache, or (ii) the metadata sought (whether or not the default value) resides in the cache block evicted to external memory. Thus, after an eviction occurs and as long as an m-cache block containing at least one valid metadata entry remains in external memory, an m-cache miss means that the assumption of the default metadata value no longer holds, thereby requiring a cache block load on an m-cache miss in at least some embodiments, thereby substantially different than the operational flow in the pre-eviction, m-cache-only (m-cache as the only storage) state. Figure 4In a generalized embodiment of the present invention, the metadata manager maintains a state variable, referred to herein as an eviction-flag (evFlag), for indicating whether an eviction has occurred, and thus signals whether the metadata is stored only in the m-cache (implementing an m-cache-only (MCO) inference / assumption of default metadata on an m-cache miss) or stored in both the m-cache and the external memory (split storage). In various embodiments, one or more background processes may be executed by the metadata manager or a host processor to track cache blocks evicted to the external memory, thereby returning these cache blocks to the metadata cache (and / or restoring metadata values ​​within evicted cache blocks to default metadata values) when conditions permit, thereby ultimately cleaning all cache blocks from the external memory and resetting the eviction flag (restoring the MCO mode of operation). In a more specific example, a direct memory access engine (e.g., implemented by software execution) may execute a background restoration of metadata values ​​of evicted blocks to default values, thereby resetting the eviction flag when no non-default metadata remains in the external memory.

[0032] At system boot (and / or soft reset, etc.), the eviction flag is reset so that any incoming read or write command follows the m-cache only operation flow generally shown at 281. More specifically, the evFlag evaluation at 281 produces a negative determination so that the read or write command (branch at 285) triggers an m-cache search at 287 or 289, respectively. In the metadata read flow, a cache miss (negative determination at 291) means an unstored default metadata value (dMD), which is returned to the host requestor at 293. Figure 2 In an embodiment of the present invention, for example, metadata manager 150 asserts an enable default signal (enDef) to pass a default metadata value (e.g., provided by a pre-programmed field within configuration register 165) onto the RdMD path for return to the host via oMD multiplexer 161. Return to Figure 4 If an m-cache hit occurs after the search at 287 (affirmative determination at 291, which applies only to non-default "infrequently occurring" metadata values), outbound metadata (oMD) from the m-cache is returned to the host (i.e., the metadata manager deasserts Figure 2 enDef in the embodiment).

[0033] Still refer to Figure 4280 ), splits metadata writes following an m-cache search at 289 based on whether the inbound metadata value (iMD) matches a default metadata value. If iMD matches the default value (affirmative determination at 301 ) and no m-cache hit occurs (negative at 303 ), no further action is required because the default value iMD is not stored in the m-cache. If an m-cache hit occurs with respect to a DPA provided with the default value iMD (i.e., the affirmative branch at 303 ), then a non-default metadata value with respect to the DPA was previously stored in the m-cache. In this case, the metadata manager instructs the m-cache at 305 to invalidate the previously stored metadata entry, thereby reverting to an implicit instance of the default metadata with respect to the invalidated entry in response to future searches at the DPA by ensuring an m-cache miss. If the inbound metadata value does not match dMD, then the cache hit (affirmative at 307 ) triggers merging iMD into the m-cache block at 309 (as with respect to Figure 3 ), and a cache miss (negative branch at 307) with respect to an empty cache block (affirmative determination at 311) similarly triggers the merging of the iMD into the empty m-cache block (making the cache block non-empty) and the addition of the tag field storing the DPA (i.e., the iMD+tag load operation as shown at 313), thereby ensuring a cache hit for the subordinate DPA in any downstream m-cache search.

[0034] An MCO mode non-default metadata write operation that generates an m-cache miss on a non-empty block (negative determinations at 301, 307, and 311) constitutes an eviction trigger conflict. In this case (which may never occur in some m-cache implementations and applications), the metadata manager evicts the resident cache block to external memory (e.g., as described in relation to Figure 3 227 and 243 of the embodiment of the present invention), and sets the eviction flag to reflect the transition from m-cache-only (MCO) metadata storage to split metadata storage. After the eviction at 315, the metadata manager merges the inbound metadata value into a null (empty) cache block at 317 - i.e., an iMD-merged null-cache-block load, wherein the cache block containing the merged inbound metadata value as its only valid entry (i.e., the valid bit is cleared for all other entries) is stored in the m-cache along with the DPA tag field.

[0035] Referring now to the split store operation after the evFlag setting at 315 (i.e., the operation generally shown at 330), an incoming read command (Cmd=Read at 331) triggers an m-cache search (333), followed by a reference to Figure 3 The operations discussed are related to an m-cache miss in response to a read triggered search. These operations are, i.e., oMD returns (337) in response to an m-cache hit (affirmative at 335), or a cache miss triggered eviction 341 (if a non-empty block at 339), a block load 343, and either a metadata fetch 345 from the retrieved cache block at 345, or looping back to 333 for an m-cache re-search (i.e., a re-search at 333 after the cache block load at 343).

[0036] Still refer to the split storage operation process ( Figure 4 ), when the m-cache search triggered by the write command at 353 produces a cache hit (affirmative determination at 355), the m-cache controller takes one of two actions depending on whether the inbound metadata value matches the default value (determination at 357): either merge the non-dMD inbound metadata value into the cache block at 359; or, in the case of iMD=dMD, invalidate the metadata entry within the DPA-indexed cache block (i.e., the metadata entry specified by the DPA offset field) at 361. Keeping the storage data write path split, in the case of a cache miss on an empty cache block (negative at 355, affirmative at 363), the m-cache controller again takes alternative actions depending on whether the inbound metadata value matches the default value (determination at 367) - merge the cache block from the external memory with the non-default value iMD at 369; or, in the case of the default value iMD (affirmative at 367), load the cache block from the external memory at 371 (storing the new tag field in the m-cache), and then invalidate the metadata entry specified by the DPA offset field at 373. In the event that the split storage write operation indexes a non-empty block (resulting in negative determinations at 355 and 363), the indexed cache block is evicted to external memory at 375, followed by the operations discussed above for the empty block determination (i.e., the operations at 369 for the non-default value iMD, or the operations at 371 and 373 in the case where iMD=dMD).

[0037] Still refer to Figure 4 , wherein the configuration register 165 is programmed with a conditional inference operation mode (i.e., an alternative inference mode), after setting evFlag at 315 and completing the load of the null value block merged by iMD at 317, the m-cache manager switches from the inference operation flow shown at 280 to Figure 3, thereby explicitly storing all metadata values ​​and no longer inferring default value metadata in response to cache misses. In the event that the host component and / or the metadata cache memory control component clears evicted cache blocks from external memory (e.g., returns them to the metadata cache in a background operation, or converts / changes all metadata values ​​within these cache blocks to dMD so that they do not need to be returned to the m-cache), the metadata manager can automatically transition back to inferential operation after all metadata has been evicted from external memory, although some cache-stored default metadata values ​​will eventually be invalidated. Additionally, in relation to Figure 4 In all of the operational flows discussed, the eviction at 315 may trigger an error reporting operation in which the host (or other system management component) is notified of an impending / unexpected exit from the metadata cache-only operating mode. Following such an error reporting, the metadata cache control device may optionally Figure 4 The split storage process shown continues to operate, and / or if the metadata cache control device is programmed for conditional reasoning operations, it Figure 3 The explicit store flow shown (explicitly storing all metadata values) continues to operate. In addition, to avoid Figure 4 The indeterminate / previously unwritten cache block is loaded from external memory at operations 369 and 371 in the above description. The external memory (or any portion of the external memory) storing the metadata cache block may be initialized at system startup, or other response to other events (e.g., transitioning from MCO mode to split storage mode) with benign or host-defined initial metadata values.

[0038] Figure 5 The diagram shows the reference Figure 3 and Figure 4 discussed and can be Figure 1 / Figure 2 380 is an embodiment of a metadata cache 380 that is deployed like the metadata cache in the embodiment and is capable of performing search, load, and entry invalidation operations. In the depicted example, the m-cache 380 includes a cache controller 381 and a cache storage 383, the cache controller issues control signals to perform the operation ("Instr") instructed by the MD manager, and the cache storage includes an index field decoder 391, a tag storage 393, a tag comparator 395, a block storage 397, a block register 401, an input / output multiplexer circuit 403, and a hit signal generator 405 and an empty signal generator 407. In one embodiment, the incoming instruction signal is Figure 3 and Figure 4The superset of operations that can be indicated by the metadata manager within the operation flow of , which includes search, inbound metadata (iMD) load, cache block load, cache block load merged with iMD ("block + iMD load"), tag field / iMD load, null value block load merged with iMD, and entry invalidation. In applications or embodiments that do not require a superset of operations (i.e., such as Figure 3 A more limited set of operations in the operational flow of the m-cache 380 is sufficient—no entry invalidation or empty value block loads are performed after an m-cache miss, and no iMD is merged into an empty cache block), and hardware / circuit elements provided to perform unused operations can be omitted. In addition, to simplify the illustration, the m-cache storage is illustrated as a single-way set associative—i.e., direct mapping, where a given index into the incoming DPA resolves to only one storage tag field (single-way) rather than multiple tag fields (multi-way). In alternative embodiments, the m-cache 380 can be implemented using multi-way associativity (i.e., n-way associativity, where n>1), or can even be fully associative. More generally, the m-cache can be implemented using any feasible cache architecture, with the operational circuits / features discussed below changing accordingly.

[0039] The m-cache controller 381 responds to incoming cache search instructions (e.g., Figure 3 213 and 233 of Figure 4 287, 289, 333 and 353 in the operations shown): assert an enable index instruction (enIndx) to decode the index field of the incoming DPA in the index decoder 391 and apply the decoded index (the one-hot output of the decoder 391) to select the tag field value stored in the tag storage 393 and the cache block entry stored in the block storage 395 (referred to herein as the indexed tag and the indexed cache block, respectively). The tag comparator 397 compares the indexed tag with the tag field of the incoming DPA to determine whether the DPA indexes into the cache block containing the metadata entry of interest, thereby asserting or deasserting a tag match signal ("t-match") to indicate the comparison result. The indexed cache block - containing 'n' metadata entries, each metadata entry including storage of a corresponding metadata value, and in some embodiments including a valid bit for authenticating the metadata value as shown at 400 - is routed to the input of the block register 401 via the register input multiplexer component of the multiplexing circuit 403, and is also driven onto the cache block output path as an outbound cache block (oCB). As shown, the valid bit associated with the corresponding metadata entry within the output cache block is provided to the NOR gate implementation of the empty signal generator 407 to generate an empty signal ("empty") provided to the metadata manager (e.g., at Figure 3 Decisions 219, 239 and Figure 4The m-cache controller asserts a register-load signal (RegLd) to load the oCB into the block register 401, which in turn registers the constituent metadata entries within the cache block (i.e., 'n' metadata entries, where n=2 N and N is the szOfst value discussed above) is fed to an output multiplexer component of the multiplexer circuit 403. The output multiplexer outputs the metadata entry specified by the offset field of the incoming DPA as an offset-selected metadata entry, where the entry (similar to all other entries within the registered cache block) contains the metadata value and a valid bit, which indicates whether the metadata value is valid. The hit signal generator 405 (conceptually shown as a logic AND gate) is enabled by a compare-enable pulse (enCmp) from the m-cache controller 381 to assert a cache hit signal ("hit") in response to a valid metadata indication (valid bit asserted) and matching index tag and DPA tag fields (tag match signal asserted), and de-assert the hit signal to indicate a cache miss if either the valid bit or the tag match signal is de-asserted.

[0040] During an inbound metadata (iMD) load operation—for example, Figure 3 237 of them and Figure 4 As shown at 309 and 359 in the example, the metadata entry selected by the offset in the block register 401 is overwritten with the inbound metadata value, and the contents of the block register are then written back to the block storage location indexed by the DPA. In one embodiment, the cache controller 381 implements this operation by setting the register input multiplexer (i.e., the multiplexer that controls the source of the signal written into the block register) to a hold state (such as "iMxCtrl=01" shown in the exemplary instruction / operation table at 410), while also asserting the enable-merge signal (enMrg), which is used to override the hold-state multiplexer setting of the block register entry corresponding to the offset field of the DPA, and thereby enable iMD to overwrite the contents of this particular entry in the block register when the cache controller asserts the RegLd signal. Then, in response to the m-cache controller asserting the register-store signal (RegStr), the contents of the block register (now containing iMD in the entry specified by the DPA offset field) are written back to the block storage.

[0041] The m-cache controller 381 implements the cache block load operation (the third entry in table 410, for example, in Figure 3 223 of them and Figure 4 343 in ): setting the register input multiplexer to route the incoming cache block (iCB) to the per-entry input of the block register ("iMxCtrl=10"), loading the iCB into the block register by asserting the register load signal (RegLd), and then asserting the register store control signal and the tag store control signal (RegStr, TagStr) to store the contents of the block register (i.e., iCB) in the block storage 395 and the tag field of the host-provided DPA in the tag storage 393, respectively. The m-cache controller 381 implements the cache block load with the merged inbound metadata load (e.g., as in Figure 3 Operation 241 and Figure 4 410), which is implemented in much the same way as a cache block load, but in which a merge enable signal (enMrg) is additionally asserted to store the inbound metadata value in the block register entry specified by the offset field of the DPA provided by the host (i.e., iCB is stored in the block register, but the offset specified entry is overwritten with iMD, thereby effectively merging iMD into iCB as part of the block load). Thereafter, the contents of the block register are written back to the block storage (RegStr is asserted) and the tag field of the DPA is written to the tag storage (TagStr is asserted) to complete the block load operation.

[0042] As discussed above with respect to iMD loading, joint storage of the DPA tag field and inbound metadata is implemented (e.g., as in Figure 4 410), wherein the tag storage control signal (TagStr) is additionally asserted to write the DPA tag field into the tag storage. The m-cache controller 381 loads the empty / null value cache block merged with the inbound metadata via the same control signal sequence as in the block+iMD load operation (e.g., Figure 4 410 ), but additionally asserts the block nullification signal (Nul) before loading the block register, which clears the valid bits (i.e., indicates non-valid entries) of all metadata entries loaded into the block register, except for entries overwritten by the inbound metadata (i.e., by means of the merge enable signal). Finally, the m-cache controller 381 implements the iMD load in much the same manner as the iMD load. Figure 4305 and 369 of the metadata manager, wherein an invalidation signal (Inv) is additionally asserted prior to register loading—the invalidation signal is used to clear the valid bit within the block register entry specified by the DPA offset field and thereby invalidate the entry. In an alternative embodiment, the metadata cache may respond to a “flush” instruction from the metadata manager by invalidating the contents stored in the metadata cache at a variable granularity (e.g., invalidating a specific metadata entry according to the DPA offset field, invalidating the contents of the entire cache block indexed by the DPA, invalidating the entire m-cache contents), and evicting / writing the corresponding m-cache contents (including the entire m-cache contents) to / from an external memory. Such cache flushing operations may be performed to implement / support, for example, null value block loading (flush granularity = cache block), evicting the entire cache contents to an external memory, MD entry invalidation (flush granularity = MD entry), and the like. In such an embodiment, separate instructions and circuits for null value block loading and entry invalidation may be omitted, and circuits supporting variable granularity cache flushing operations may be employed. In addition, although reference is made to Figure 5 1 and other embodiments discussed herein illustrate and describe explicit valid bit storage (i.e., one valid bit per metadata entry stored in the m-cache), but alternatively, in all embodiments herein, the validity / invalidity of the metadata may be encoded within the metadata value itself (e.g., a host-specified / reserved 2 q In such an embodiment, an explicit valid bit (or set of bits) may be generated or applied within the m-cache as needed, for example, by circuitry that detects the validity / invalidity status of the encoded metadata within a given metadata entry and generates a valid bit accordingly for the purpose of evaluating cache hits / misses, and / or by circuitry that encodes the validity / invalidity status into one or more metadata entries within a cache block for the purpose of entry invalidation, empty value block loading, etc.

[0043] Figure 6 Picture shows Figure 5 4 and 5. An exemplary slice of the block register 401 and register input multiplexer 421 of FIG. 4 (which corresponds to the width of a single metadata entry within a multi-entry (n-entry) cache block), showing exemplary interconnections with the block storage 395 and output multiplexer 423 (register input multiplexer 421 and output multiplexer 423 are Figure 5403). In the depicted embodiment, a block register slice (slice 'i' is used to specify the storage of the i-th metadata entry of n metadata entries within a given m-cache block) is implemented by a flip-flop stage 425 having: (i) sufficient storage space and I / O width to store multi-bit metadata values ​​and corresponding valid bits; (ii) an output (Q) that drives the slice 'i' cache block contents to the input block storage 395; and (iii) an input (D) that is used to receive incoming metadata entries from the source selected by slice 'i' of data source multiplexer 427 (i.e., component of register input multiplexer 421).

[0044] In an m-cache search operation (i.e., triggered by the m-cache controller assertion of the index enable signal enIndx), each slice (Srch[0], Srch[1], ..., Srch[i], ..., Srch[n-1]) of the cache block indexed by the DPA is provided to the '00' port ( Figure 6 Only input multiplexer slice 'i' is shown). Thus, the m-cache controller stores the indexed cache block in the block register by driving the data source multiplexer control to '00' (iMxCtrl=00, selecting the '00' data source port), and then issuing a register load pulse (asserting and then deasserting RegLd), thereby selecting the indexed cache block into the block register 401. This control signal sequence is Figure 7 This is illustrated in the exemplary waveform diagram labeled “Search” in FIG. 1—as part of an m-cache search, assertion of a compare enable signal is additionally shown to enable conditional generation of a hit signal (based on the state of t-match and the entry valid bit).

[0045] Continue to refer to Figure 6 , the output of each block register slice after the m-cache search (i.e., constituting the corresponding metadata entry and thus storing the metadata value and the corresponding valid bit indicating whether the metadata value is valid) is provided to the corresponding input port of the output multiplexer 423, which in turn is based on the N-bit offset field (N=log 2 The metadata field within the offset-selected block register entry (i.e., output from multiplexer 423) constitutes the outbound metadata value (oMD) described above, and the valid bits within the entry correspond to Figure 5- which will be logically ANDed with the tag comparator output to generate a hit / miss signal (where such operation is alternatively performed within the metadata manager). In the depicted embodiment, the complete block register output (i.e., the collective output from all block register slices) constitutes the cache block (oCB) output from the m-cache for storage in external memory in an eviction operation.

[0046] A metadata load following a cache block store triggered by a search within block register 401 (e.g., as in Figure 3 Operation 237 and Figure 4 The operation 309, 359 of the embodiment of the present invention is implemented by overwriting the offset selected entry of the block register with the inbound metadata value (iMD). More specifically, the m-cache controller asserts the merge enable signal (enMrg) while setting the data source multiplexer control to the register retention state (iMxCtrl='01') so that when the register load signal is pulsed, all entries in the block register are maintained (reloaded) except for the entry corresponding to the DPA offset field. In the depicted example, the offset field of the DPA is provided to a decoder 429 which generates a one-hot output on n offset decode lines (odc) - one of the odc lines corresponding to the DPA offset field being activated / asserted while all other lines remain deasserted - such that assertion of the merge enable signal (via AND gate 431) activates the merge multiplexer 433 within the offset selected input multiplexer slice, thereby passing the iMD load control value of '11' to the control input of the data source multiplexer 427, rather than the '01' register hold value passed (from the m-cache controller) to the data source multiplexers within all other input multiplexer slices. Thus, when the m-cache controller pulses the register load signal, iMD is loaded into the offset selected entry within the block register 401 - and thereafter stored within the block register in response to the cache controller asserting the register store signal (RegStr). Figure 7 An example of this control signal sequence is illustrated in the "iMD Load" waveform diagram. Note that at least Figure 3 and Figure 4 In the operational flow of , an inbound metadata load follows an m-cache hit, which means that no tag storage is required (hence the TagStr control signal is deasserted) because it has been confirmed that the tag field stored in the m-cache tag storage matches the tag field of the DPA provided by the host.

[0047] In a cache block load operation (e.g., after a cache miss and possible eviction operation), the m-cache controller sets the data source multiplexer control value to pass the inbound cache block (iCB) to the block register (i.e., Figure 6 In the example, iMxCtrl='10'), the Register Load signal is then asserted to capture iCB in the block register, and the Register Store signal and the Tag Store signal are then pulsed to store the block register contents (iCB) in the block storage and the tag field of the incoming DPA in the tag storage. Figure 7 An exemplary block loading control signal sequence is shown under the heading "Block Loading" in FIG.

[0048] A block load operation merged with inbound metadata (block+iMD load) may be implemented generally as in the cache block load described above, with the m-cache controller additionally asserting a merge enable signal (e.g., Figure 7 ) to load the inbound metadata into the slice of the block register specified by the DPA offset field. A tag field load in conjunction with an inbound metadata store (e.g., as Figure 4 In operation 313) using Figure 7 The same signal sequence as shown in the iMD loading is implemented, except that, as Figure 7 As shown in the example entitled "Tag+iMD Load" in FIG. 1 , the m-cache controller additionally (eg, in parallel with RegStr) asserts the tag store signal.

[0049] Null value chunk load merged with inbound metadata (e.g. Figure 4 The operation 317) can be Figure 7 The same signal sequence as shown in the block +iMD load is implemented, where the m-cache controller additionally asserts the null block signal (Nul) during the register load operation. Figure 6 In the exemplary embodiment of , the block null signal is applied to the inverting input of AND gate 435, thereby clearing the valid bits (i.e., to indicate non-valid entries) of all entries loaded into the block register via the inbound cache block port (i.e., input multiplexer port '10'), while the merge enable signal enables the inbound metadata value (and therefore the valid bit) to be stored in the block register slice specified by the DPA offset field. Figure 7 An exemplary null block / iMD load control signal sequence is shown under the heading "N-Blk+iMD Load" in FIG.

[0050] exist Figure 6 and Figure 7 In an embodiment, the m-cache controller issues a Figure 7to implement the entry invalidation operation (e.g., as shown in Figure 4 305, 369 in the example), wherein the m-cache controller additionally asserts an entry invalidation signal (Inv). Figure 6 In the example of , the inverted version of the invalid signal (i.e., generated by inverter 437) constitutes a valid bit that is otherwise loaded with the inbound metadata value, such that assertion of the invalid signal effectively clears the valid bit associated with iMD, thereby invalidating the entry. Figure 7 The exemplary entry invalidation control signal sequence in is titled "Invalidate Entry".

[0051] Although the reference Figures 5 to 7 A block register based m-cache implementation and operation is described, but in alternative m-cache embodiments having block loads (iCB or nulls), iMD merges (including entry invalidations), and similar operations performed in-situ within the block storage 395 (i.e., no cache block needs to be read from the block storage for iMD merge, nullification, invalidation, and / or block load operations), the block register 401 may be omitted. More generally, as discussed above, the metadata cache (including the controller and its storage components) may be implemented in any feasible architecture, including architectures in which the cache block size matches the metadata entry size (rather than the memory line / user data size), architectures that maintain a per-entry and / or per-cache block dirty bit (such a bit is set to mark modified entry / cache block content), architectures that maintain additional status bits (e.g., access counters) that implement one of a variety of programmably selectable eviction / replacement policies (e.g., least recently used, first-in, first-out, last-in, last-out, most recently used, time-aware least recently used, least frequently used, etc.), and the like.

[0052] Overall reference Figures 1 to 7, the metadata cache memory control device described above can be implemented in a stand-alone integrated circuit package (e.g., having one or more IC dies) or in one or more discrete IC packages. Conversely, the metadata cache memory control device can be integrated with host components and / or external memory components in a multi-chip package. One or more programmed microcontrollers and / or dedicated hardware circuits (e.g., finite state machines, sequencers, registers or combinational circuits, etc.) can implement and / or control all or part of the various architectural and functional elements within the metadata cache memory control device presented herein (e.g., to implement any one or more of the metadata manager (e.g., the FSM therein), the m-cache controller, etc.). Additionally, any or all of these architectural / functional elements (including the entire metadata cache memory control device and / or a host device having core circuits for programming / configuring and operating the metadata cache memory control device) can be described using computer-aided design tools and expressed (represented) as data and / or instructions embodied in various computer-readable media in terms of their behavior, register transfers, logic components, transistors, layout geometry, and / or other characteristics. The formats of files and other objects that can implement such circuit expressions include, but are not limited to, formats that support behavioral languages ​​such as C, Verilog, and VHDL, formats that support register-level description languages ​​such as RTL, and formats that support geometric description languages ​​such as GDSII, GDSIII, GDSIV, CIF, MEBES, and any other suitable formats and languages. Computer-readable media that can embody such formatted data and / or instructions include, but are not limited to, various forms of computer storage media (e.g., optical, magnetic, or semiconductor storage media).

[0053] Such data and / or instruction-based representations of the circuits described above, when received within a computer system via one or more computer-readable media, may be processed by a processing entity (e.g., one or more processors) within the computer system in conjunction with the execution of one or more other computer programs to generate a representation or image of the physical manifestation of such circuits, the one or more other computer programs including, but not limited to, netlist generation programs, placement and routing programs, etc. Such representations or images may thereafter be used in device fabrication, for example, by causing the generation of one or more masks used to form various components of the circuits during device fabrication.

[0054] In the foregoing description and accompanying drawings, specific terms and reference numerals have been set forth to provide a comprehensive understanding of the disclosed embodiments. In some instances, the terms and reference numerals may imply specific details that are not required to practice these embodiments. For example, the expression "user data" is used herein in distinction from the metadata component of the composite / synthetic data word, and may include data from almost any source and / or data used by any execution entity (e.g., variable or static data (including program code itself) associated with an application, hardware driver, operating system code, dynamic link library, etc. executed by a processor). For exemplary purposes only, various widths of user data components, metadata sizes, external memory storage sizes, specific numbers of physical address bits and their subfields, cache architectures (including association levels ranging from direct mapping to full association), external memory architectures and / or storage technologies, signaling path widths, cache block sizes, command protocols, etc. are provided - in all cases, any feasible alternatives may be implemented. Similarly, signaling link parameters, protocols, configurations may be implemented according to any feasible open or proprietary standards and any versions of such standards. Although a memory subsystem coupled to a metadata cache memory control device has been described as an "external" memory, such a memory subsystem or any portion thereof may be implemented on the same die or within the same integrated circuit package (e.g., a system-in-package, a three-dimensional IC, etc.) as the metadata cache memory control component. Links or other interconnections between integrated circuit devices or internal circuit elements or blocks may be shown as buses or single signal lines. Alternatively, each bus may be a single signal line (e.g., on which digital or analog signals are time-division multiplexed), and each single signal line may alternatively be a bus. Regardless of how it is shown or described, signals and signaling links may be single-ended or differential. In alternative embodiments, a logic signal shown as having an effective high level assertion or "true" state may have an opposite assertion state. When a signal driver circuit asserts (or deasserts, if the context clearly states or indicates) a signal on a signal line coupled between the signal driver circuit and the signal receiving circuit, the signal driver circuit is considered to "output" a signal to the signal receiving circuit. The term "coupled" is used herein to express direct connections as well as connections made through one or more intermediate circuits or structures. Integrated circuit device or register "programming" may include, for example, but not limited to, loading control values ​​into configuration registers or other storage circuits within the integrated circuit device in response to host instructions (and thereby controlling operational aspects of the device and / or establishing a device configuration) or by a one-time programming operation (e.g., blowing fuses within the configuration circuit during device production), and / or connecting one or more selected pins or other contact structures of the device to a reference voltage line (also referred to as strapping) to establish a particular device configuration or operational aspect of the device. The terms "exemplary" and "embodiment" are used to convey examples, not preferences or requirements.Additionally, the terms "may" and "can" are used interchangeably to indicate optional (permissible) subject matter. The absence of either term should not be construed to mean that a given feature or technique is required.

[0055] Various modifications and changes may be made to the embodiments presented herein without departing from the broader spirit and scope of the present disclosure. For example, the features or aspects of any embodiment may be applied in combination with any other embodiment, or in place of its corresponding features or aspects. Therefore, the description and drawings should be regarded as illustrative rather than restrictive.

Claims

1. An integrated circuit component, include: a host interface, the host interface being used to receive a host command, a physical address, and write data, the write data comprising a first component data value and a second component data value; Memory control interface; Storage cache; A control circuit, the control circuit being configured to perform the following operations in response to the host command: outputting the first component data value and the physical address via the memory control interface to enable storage of the first component data value in a memory device external to the integrated circuit component; as well as The second component data value is stored within the storage cache at the location indicated by the physical address.

2. The integrated circuit component according to claim 1, in, The control circuitry for storing the second component data value at the location indicated by the physical address within the storage cache includes circuitry for: initiating a search within the storage cache to determine whether the storage cache contains a cache block corresponding to the physical address; And if so, storing the second component data value in the storage cache location occupied by the cache block.

3. The integrated circuit component according to claim 2, in, The cache block includes multiple entry storage locations, and the control circuit for storing the second component data value in the storage cache location occupied by the cache block includes a circuit for storing the second component data value in an entry storage location in the multiple entry storage locations indicated by the offset field of the physical address.

4. The integrated circuit component as claimed in claim 3, in, The storage cache includes a block storage memory and a block register, and wherein the circuitry for storing the second component data value in one of the plurality of entry storage locations indicated by the offset field includes circuitry for: outputting the cache block from the block storage memory to the block register; storing the second component of the user data in a portion of the block register corresponding to the one of the plurality of entry storage locations to generate an updated cache block in the block register; and writing the updated cache block back to the block storage memory.

5. The integrated circuit component according to claim 2, in, The first component data value and the second component data value are respectively composed of a first number of bits and a second number of bits, the first number of bits being larger than the second number of bits by an integer factor N, and wherein the cache block comprises N different subfields, each subfield having a storage capacity according to the second number of bits, such that the second component data value can be stored in one of the N different subfields, and up to Nl other component data values, each composed of the second number of bits, can be stored in other subfields of the N different subfields.

6. The integrated circuit component according to claim 1, in, The control circuitry for storing the second component data value at the location indicated by the physical address within the storage cache includes circuitry for searching the storage cache to determine whether the storage cache contains an entry corresponding to the physical address, including circuitry for: selecting a cache block from a plurality of cache blocks stored in the storage cache based on the first portion of the physical address; selecting a subfield of a plurality of subfields within the cache block based on a second portion of the physical address; as well as A cache hit signal is asserted in a first logic state or a second logic state based at least in part on contents within the one of the plurality of subfields to indicate whether the storage cache contains the entry corresponding to the physical address.

7. The integrated circuit component as claimed in claim 6, in, The circuitry for selecting the one of the multiple cache blocks based on the first portion of the physical address includes circuitry for selecting the one of the multiple cache blocks based on an index field within the first portion of the physical address, and wherein the circuitry for asserting the cache hit signal in the first state or the second state to indicate whether the storage cache contains the entry corresponding to the physical address based at least in part on the contents within the one of the multiple subfields includes circuitry for asserting the cache hit signal in the first state or the second state additionally based on whether a tag value stored in the storage cache matches a tag field within the first portion of the physical address.

8. The integrated circuit component according to claim 1, in, The host interface for receiving the host command, physical address, and write data comprises an interface compatible with a cache coherence communication standard, and wherein the second component data value comprises a metadata value required by the cache coherence communication standard.

9. The integrated circuit component according to claim 1, in, The host command and the physical address received via the host interface include a first host command and a first physical address, and wherein the host interface is further configured to receive a second host command and a second physical address, the second host command requesting a read data word consisting of a third component data value and a fourth component data value, and wherein the control circuitry includes circuitry for, in response to the second host command, performing the following operations: outputting the second physical address via the memory control interface to retrieve the third component data value from the memory device external to the integrated circuit component; conditionally retrieving the fourth component data value from the location indicated by the second physical address within the storage cache; and The third component data value and the fourth component data value are output via the host interface as the read data word requested by the second host command.

10. The integrated circuit component of claim 9, in, The circuitry for conditionally retrieving the fourth component data value from the location indicated by the second physical address within the storage cache includes circuitry for: searching the storage cache to determine whether the storage cache contains a cache block specified by a first portion of the second physical address; And if so, determining whether the cache block contains valid data within the entry specified by the second portion of the second physical address.

11. The integrated circuit component of claim 9, in, The circuitry for conditionally retrieving the fourth component data value from the location indicated by the second physical address within the storage cache includes circuitry for: searching the storage cache with respect to the second physical address; and obtaining, from the storage cache, a data value within an entry specified by an offset subfield of the second physical address as the fourth component data value when: (i) one or more other subfields of the second physical address indicate that a cache block containing the entry is stored within the storage cache, and (ii) the data value within the entry is indicated as valid by one or more associated bits.

12. The integrated circuit component of claim 11, further comprising a programmable register for storing a default value; and Circuitry for performing the following operations: outputting the default value from the programmable register as the fourth component data value so that the read data word output via the host interface includes the default value when: (i) the one or more other subfields of the second physical address indicate that no cache block containing the entry is stored in the storage cache, or (ii) the one or more associated bits indicate that no valid data value is within the entry.

13. A method of operating in an integrated circuit component, the integrated circuit component having a host interface, a memory control interface and a memory cache, the method include: receiving, via the host interface, a host command, a physical address, and write data, the write data comprising a first component data value and a second component data value; as well as In response to the host command: outputting the first component data value and the physical address via the memory control interface to store the first component data value in a memory device external to the integrated circuit component; as well as The second component data value is stored within the storage cache at a storage cache location indicated by the physical address.

14. The method according to claim 13, in, Storing the second component data value at the location indicated by the physical address within the storage cache includes: searching the storage cache to determine whether the storage cache contains a cache block corresponding to the physical address; and if so, storing the second component data value within the location occupied by the cache block.

15. The method of claim 14, in, The cache block includes multiple entry storage locations, and storing the second component data value in the storage cache location occupied by the cache block includes: storing the second component data value in an entry storage location in the multiple entry storage locations indicated by the offset field of the physical address.

16. The method of claim 15, in, The storage cache comprises a block storage memory and a block register, and wherein storing the second component data value in one of the plurality of entry storage locations indicated by the offset field comprises: outputting the cache block from the block storage memory to the block register; storing the second component of the user data in a portion of the block register corresponding to the one of the plurality of entry storage locations to generate an updated cache block in the block register; and writing the updated cache block back to the block storage memory.

17. The method of claim 14, in, The first component data value and the second component data value are respectively composed of a first number of bits and a second number of bits, the first number of bits being larger than the second number of bits by an integer factor N, and wherein the cache block comprises N different subfields, each subfield having a storage capacity according to the second number of bits, such that the second component data value can be stored in one of the N different subfields, and up to Nl other component data values, each composed of the second number of bits, can be stored in other subfields of the N different subfields.

18. The method of claim 13, in, Storing the second component data value at the location indicated by the physical address within the storage cache includes searching the storage cache to determine whether the storage cache contains an entry corresponding to the physical address, including: selecting a cache block from a plurality of cache blocks stored in the storage cache based on the first portion of the physical address; selecting a subfield of a plurality of subfields within the cache block based on a second portion of the physical address; and A cache hit signal is asserted in a first state or a second state based at least in part on content within the one of the plurality of subfields to indicate whether the storage cache contains the entry corresponding to the physical address.

19. The method of claim 18, in, Selecting the one of the multiple cache blocks based on the first portion of the physical address includes selecting the one of the multiple cache blocks based on an index field within the first portion of the physical address, and wherein asserting the cache hit signal to be in the first state or the second state to indicate whether the storage cache contains the entry corresponding to the physical address based at least in part on the content within the one of the multiple subfields includes asserting the cache hit signal to be in the first state or the second state based additionally on whether a tag value stored in the storage cache matches a tag field within the first portion of the physical address.

20. The method of claim 13, in, Receiving the host command, physical address, and write data via the host interface includes receiving the host command, physical address, and write data via an interface compatible with a cache coherence communication standard, and wherein the second component data value includes a metadata value required by the cache coherence communication standard.

21. The method of claim 13, in, Receiving the host command and the physical address includes receiving a first host command and a first physical address, the method further comprising: receiving, via the host interface, a second host command and a second physical address, the second host command requesting a read data word comprised of a third component data value and a fourth component data value; and In response to the second host command: outputting the second physical address via the memory control interface to retrieve the third component data value from the memory device external to the integrated circuit component; conditionally retrieving the fourth component data value from the location indicated by the second physical address within the storage cache; and The third component data value and the fourth component data value are output via the host interface as the read data word requested by the second host command.

22. The method of claim 21, in, Conditionally obtaining the fourth component data value from the location indicated by the second physical address within the storage cache includes: searching the storage cache to determine whether the storage cache contains a cache block specified by the first portion of the second physical address; and if so, determining whether the cache block contains valid data within an entry specified by the second portion of the second physical address.

23. The method of claim 21, in, Conditionally obtaining the fourth component data value from the location indicated by the second physical address within the storage cache includes: searching the storage cache with respect to the second physical address; and obtaining data within an entry specified by an offset subfield of the second physical address from the storage cache as the fourth component data value under the following circumstances: (i) one or more other subfields of the second physical address indicate that a cache block containing the entry is stored within the storage cache, and (ii) the data value within the entry is indicated as valid by one or more associated bits.

24. The method according to claim 23, further comprising: storing a default value in a programmable register of the integrated circuit component; and outputting the default value from the programmable register as the fourth component data value such that the read data word output via the host interface includes the default value when: (i) one or more other sub-fields of the second physical address indicate that no cache block containing the entry is stored in the storage cache, or (ii) one or more associated bits indicate that there is no valid data value in the entry.

25. An integrated circuit component, comprising: a host interface for receiving a host command, a physical address, and write data, the write data including a first component data value and a second component data value; a memory control interface; a storage cache; means for performing the following operations in response to the host command: outputting the first component data value and the physical address via the memory control interface so that the first component data value can be stored in a memory device external to the integrated circuit component; and storing the second component data value at a location indicated by the physical address in the storage cache.

Citation Information

Patent Citations

  • Storage device, storage system and computing device

    CN107957961A

  • Method for reading data and hybrid memory module

    CN108427647A

  • Storage devices, storage systems and methods of operating storage devices

    CN112988627A

  • Integrity protected access control mechanism

    CN114692177A

  • Techniques for directed data migration

    US10552085B1