Compress data for storage in a cache memory of a cache memory hierarchy

A hybrid cache compression method using lightweight and heavyweight techniques addresses inefficiencies in cache memory utilization by balancing data size reduction and response time across different cache levels, enhancing device performance.

CN113227987BActive Publication Date: 2025-07-15ADVANCED MICRO DEVICES INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201980086040.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-26
Filing Date
2019-06-19
Publication Date
2025-07-15
Estimated Expiration
2039-06-19

AI Technical Summary

Technical Problem

When compressing program code and data in cache memory, the prior art cannot reduce the data size while avoiding response delays to data access requests, resulting in inefficient storage capacity utilization.

Method used

Different levels of cache memory are used to use different types of compression methods. Lighter compression is used for cache memory with faster response times, and heavier compression is used for cache memory with slower response times. Combining lighter and heavier compression sequences to compress data to ensure fast response to data access requests.

Benefits of technology

It improves the storage capacity utilization of cache memory, reduces data access latency, and improves the overall performance and user satisfaction of electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113227987B_ABST
    Figure CN113227987B_ABST
Patent Text Reader

Abstract

An electronic device includes at least one compression - decompression functional block and a cache memory hierarchy having a first cache memory and a second cache memory. The at least one compression - decompression functional block receives data in an uncompressed state, compresses the data using one of a first compression or a second compression, and after compressing the data, provides the data to the first cache memory for storage therein. When retrieving data from the first cache memory for storage in the second cache memory, when the data was compressed using the first compression, the compression - decompression functional block decompresses the data to reverse the effect of the first compression on the data, thereby restoring the data to the uncompressed state, and provides the data compressed using the second compression or in the uncompressed state to the second cache memory for storage therein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Government Rights

[0002] This invention was made with government support under the Lawrence Livermore National Security PathForward program awarded by the Department of Energy (prime contract No. DE-AC52-07NA27344, subcontract No. B620717). The government has certain rights in this invention. Background Art

[0003] Related Art

[0004] Some electronic devices include a processor that executes program code for performing various operations. For example, an electronic device may include one or more central processing unit cores (CPU cores) or graphics processing unit cores (GPU cores) that execute program code for software applications, operating systems, etc. In addition to a memory (e.g., a "main" memory) and a mass storage device, many of these electronic devices also include one or more cache memories for storing program code and / or data. A cache memory is a fast-access memory that is used to locally store a copy of program code and / or data so that the processor can quickly retrieve it for use when executing the program code or performing other operations. Accessing a copy of program code and / or data in a cache memory is typically at least an order of magnitude faster than accessing program code and / or data in a memory and a mass storage device.

[0005] Cache memories generally have only a limited capacity to store copies of program code and / or data, and the capacity is typically much less than that of a memory or a mass storage device. For example, some electronic devices include a cache memory hierarchy, where the highest or first-level cache memory (i.e., the "L1" cache) has a capacity of 32 - 64 kilobytes (kB), the middle or second-level cache memory (i.e., the "L2" cache) has a capacity of 512 - 1024 kB, and the lowest or third-level cache memory (i.e., the "L3" cache) has a capacity of 2 - 4 megabytes (MB). In such electronic devices, the memory may have a capacity of 32 - 64 gigabytes (GB), and the mass storage device may have a capacity of 4 - 8 terabytes (TB). Because cache memories have a limited storage capacity, during operation, the processor periodically requests to store copies of sufficient program code and / or data that exceed the capacity of the cache memory available for simultaneously storing the copies. As a result, it may be necessary to evict or otherwise discard existing stored copies of program code and / or data in the cache memory to free up space for incoming copies of program code and / or data. For example, the cache memory can write a modified copy of the data back to the memory or a lower-level cache memory, and / or invalidate an unmodified data copy to make room for an incoming copy of the data.

[0006] To better utilize the available storage capacity of cache memories, designers have proposed many techniques to improve the arrangement of program code and / or data stored in cache memories. For example, some designers have proposed compressing program code and / or data so that the program code and / or data can be stored in a smaller space in the cache memory. For this technique, the program code and / or data are compressed to reduce the data size before being stored in the cache memory. When the compressed program code and / or data are subsequently provided from the cache memory to the requesting processor, the program code and / or data are decompressed to restore the program code and / or data to an uncompressed state.

[0007] Although compression results in a reduction in the size of program code and / or data stored in the cache memory, compression is not always an ideal solution. This is attributed to the characteristics of the different compression types that may be used to compress program code and / or data for storage in the cache memory. For example, using a heavier-weight compression (i.e., a compression with relatively more, slower, and / or more complex compression operations) may result in a larger proportion of the program code and / or data being compressed, but may take a longer time to compress and decompress. Thus, for higher-level cache memories where response latency is an important consideration, heavier-weight compression is less desirable. As another example, a lighter-weight compression can quickly compress program code and / or data using relatively simple compression operations, but is only able to compress program code and / or data that includes a limited set of patterns or values. While providing at least a small reduction in the size of the program code and / or data without incurring significant compression latency, the lighter-weight compression may not be able to compress enough data to free up significant space in a larger cache memory. Given the different performance requirements associated with the different cache memories in the cache memory hierarchy, a designer cannot find a single compression for compressing the program code and / or data of all cache memories. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A block diagram illustrating an electronic device in accordance with some embodiments is presented.

[0009] Figure 2 A block diagram illustrating performing compression and decompression on data when copying data up through the cache memories in a cache memory hierarchy in accordance with some embodiments is presented.

[0010] Figure 3 A block diagram illustrating performing compression and decompression on data when copying data down through the cache memories in a cache memory hierarchy in accordance with some embodiments is presented.

[0011] Figure 4 A flowchart illustrating a process for compressing data for storage in a cache memory in accordance with some embodiments is presented.

[0012] Figure 5 A flowchart illustrating a process for compressing data for storage in a cache memory in accordance with some embodiments is presented.

[0013] Throughout the drawings and the description, like reference numerals refer to like elements. DETAILED DESCRIPTION

[0014] The following description is presented to enable any person skilled in the art to make and use the described embodiments, and is provided in the context of a particular application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications. Thus, the described embodiments are not limited to the embodiments shown, but are to be accorded the widest scope consistent with the principles and features disclosed herein.

[0015] Term

[0016] In the following description, various terms are used to describe embodiments. The following is a simplified and general description of several of these terms. Note that these terms may have important additional aspects, which are not enumerated herein for the sake of clarity and brevity, and thus the description is not intended to limit these terms.

[0017] Functional block: A functional block refers to a group, collection, and / or set of one or more interrelated circuit elements (such as integrated circuit elements, discrete circuit elements, etc.). Circuit elements are "interrelated" because the circuit elements share at least one property. For example, interrelated circuit elements may be included in a particular integrated circuit chip or a portion thereof, fabricated on a particular integrated circuit chip or a portion thereof, or otherwise coupled to a particular integrated circuit chip or a portion thereof, may be involved in the execution of a given function (computing or processing function, memory function, etc.), may be controlled by a common control element, etc. A functional block can include any number of circuit elements, from a single circuit element (e.g., a single integrated circuit logic gate) to millions or billions of circuit elements (e.g., an integrated circuit memory).

[0018] Data: Data refers to information or values that can be stored in a cache memory or a memory (such as a main memory). For example, data can be or include information or values to be used in operations such as computing operations, control operations, sensing or monitoring operations, memory access operations, input / output device communications, or generated by such operations. As another example, data can be or include program code instructions or values from program code (e.g., variable values, constants, etc.) that are obtained from or go to a computer-readable storage medium, a mass storage device, a network interface, a memory, etc. As another example, data can be or include information obtained from or going to a functional block (such as an input / output device, a sensor, a human-machine interface device, etc.).

[0019] Overview

[0020] In the described embodiments, an electronic device includes a cache memory hierarchy having a first cache memory and a second cache memory. The first cache memory is lower in the hierarchy than the second cache memory and, thus, may be slower to access, have a higher capacity, and / or be located further from the execution circuitry accessing data in the cache memory hierarchy than the second cache memory. For example, in some embodiments, the first cache memory is a tertiary (“L3”) cache memory and the second cache memory is a secondary (“L2”) cache memory. In the described embodiments, data is compressed for storage in the cache memory in order to reduce the size of the data and, thus, more effectively utilize the available capacity of the cache memory (recall that “data” is a general term describing any value or information that can be stored in the cache memory, as described above). Because the cache memory and data access entities (e.g., CPU cores, GPU cores, etc.) have different tolerances for the latency associated with compressing and decompressing data, the described embodiments use a compression combination for the data to be stored in the cache memory. The specific compression selected for compressing the data of each cache memory is chosen to provide the benefit of reducing the size of the data for that cache memory while also avoiding decompression operations that would undesirably delay the response to data access requests.

[0021] It is desirable for a second cache memory that is higher in the hierarchy to respond to an access request in a smaller number of control clock cycles (e.g., 12 - 15 clock cycles). Thus, it is not desirable to add more than a few latency clock cycles to the response time of the second cache memory. For this reason, in some embodiments, a lighter-weight compression that involves fewer, simpler, and / or faster compression operations is used to compress data in the second cache memory. Since only lighter-weight compression is used for the data in these embodiments, compressed data can be retrieved from the second cache memory, decompressed, and returned to the access entity in a relatively short time. It is desirable for a first cache memory that is lower in the hierarchy to respond quickly to an access request, but accessing the first cache memory requires many control clock cycles (e.g., 30 - 40 clock cycles). Thus, adding several latency clock cycles to the response time of the first cache memory has a proportionally smaller impact – and the access entity may tolerate the additional latency associated with decompression. For this reason, in some embodiments, a heavier-weight compression that involves more, more complex, and / or slower compression operations compared to lighter-weight compression is used to compress data in the first cache memory. When heavier-weight compression is used for the data in the second cache memory, more patterns and / or values in the data can be compressed compared to when lighter-weight compression is used, but compressing and decompressing the data takes a relatively longer time.

[0022] In some embodiments, although heavier-weight compression is allowed to compress data in the first cache memory, either lighter-weight compression or heavier-weight compression is used to compress data in the first cache memory. In these embodiments, lighter-weight compression is preferentially used to compress data for storage in the first cache memory, and when lighter-weight compression cannot be used to compress the data, heavier-weight compression is used to compress the data. In other words, when an attempt to compress data using lighter-weight compression is unsuccessful (e.g., due to lost patterns or values in the data that might be compressed using lighter-weight compression), these embodiments attempt to use heavier-weight compression to compress the data. These embodiments preferably use lighter-weight compression because data that has been compressed using lighter-weight compression can be copied from the first cache memory to the second cache memory as is, without intermediate decompression, and thus can be done more quickly. Additionally, although using heavier-weight compression means that the data must be decompressed before being stored in the second cache memory, heavier-weight compression may result in a favorable size reduction to store data in the first cache memory that cannot be compressed using lighter-weight compression.

[0023] In some embodiments, a sequence using both a lighter-weight compression and a heavier-weight compression is used, and thus combined, to compress data in a first cache memory. In these embodiments, the lighter-weight compression is initially used to compress the data, and then the heavier-weight compression is used to compress the resulting compressed data again. A reverse sequence of heavier-weight decompression followed by lighter-weight decompression is used to perform a full decompression of the data output from the compression sequence. In these embodiments, before copying the data from the first cache memory to a second cache memory, it is decompressed to reverse the effect of the heavier-weight compression, leaving the data compressed with the lighter-weight compression to be stored in the second cache memory. Although this is at the cost of decompressing the data to prepare it for storage in the second cache memory, this is done to take advantage of the favorable size reduction of the heavier-weight compression of the data in the first cache memory. However, it should be noted that when copying the data from the first cache memory to the second cache memory, only the data needs to be decompressed to reverse the effect of the heavier-weight compression – the resulting decompressed data retains the effect of the lighter-weight compression. This makes the operation of copying the data from the first cache memory to the second cache memory faster than the case where the lighter-weight compression is also performed when storing the copy in the second cache memory.

[0024] By using lighter-weight compression and heavier-weight compression as described herein to compress the data in the first cache memory and the second cache memory, the described embodiments reduce the size of the data stored in the cache memory. Additionally, by using a specified compression for each cache memory, the latency of compressing and decompressing the data is arranged between the cache memories such that an access entity is less exposed to an unacceptable latency in response to an access request. Thus, the described embodiments are able to store more data in the cache memory, but do not introduce an unacceptable latency in response to a data access request. Storing more data in the cache memory means that an access entity can access more data in the cache memory, thus helping to avoid the need and latency of accessing data in a memory or a mass storage device. Therefore, accessing more data in the cache memory can improve the performance of the access entity, and thus improve the overall performance of the electronic device, thereby increasing user satisfaction.

[0025] Electronic device

[0026] Figure 1 A block diagram showing an electronic device 100 is presented in accordance with some embodiments. As Figure 1As can be seen, the electronic device 100 includes a processor 102 and a memory 104. Generally, the processor 102 and the memory 104 are implemented in hardware (i.e., using various circuit elements and devices). For example, the processor 102 and the memory 104 can all be fabricated on one or more semiconductor chips (including being fabricated on one or more separate semiconductor chips), can be made from a combination of semiconductor chips and discrete circuit elements, can be fabricated solely from discrete circuit elements, etc. As described herein, the processor 102 and the memory 104 perform operations associated with compressing data for storage in the cache memory.

[0027] The processor 102 is a functional block that performs computational operations and other operations (e.g., control operations, configuration operations, etc.) in the electronic device 100. For example, the processor 102 can be or include one or more microprocessors, a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), and / or other processing functional blocks. As Figure 1 can be seen, the processor 102 includes cores 106-108, each of which is a functional block that performs computational operations and other operations, such as a microprocessor core, a graphics processor core, an application-specific integrated circuit (ASIC), etc. Within the cores 106-108, the execution subsystems 110-112 include functional blocks and circuit elements, such as an instruction fetch / decoding unit, an instruction scheduling unit, an arithmetic logic unit (ALU), a floating-point operation unit, a computational unit, a programmable gate array, etc., for executing program code instructions and performing other operations.

[0028] Processor 102 includes a cache memory hierarchy that, proceeding from the lowest level to the highest level of the hierarchy, has first-level cache memories (L1 caches) 114-116 and second-level cache memories (L2 caches) 118-120 in respective cores 106-108, and a third-level cache memory (L3 cache) 122 shared by the cores. Each cache memory is a functional block that includes volatile memory circuitry, such as static random access memory (SRAM) circuitry for storing copies of data, and control circuitry for processing operations such as data access. At each level of the hierarchy, the respective one or more cache memories can be characterized at least in part by the capacity to store copies of data and the response time for responding to access requests, where smaller cache memories are higher in the hierarchy and the response time of cache memories higher in the hierarchy is shorter. For example, in some embodiments, the L1 caches 114-116 have a capacity of 32 kB and a response time of 4-8 cycles (of a control clock), the L2 caches 118-120 have a capacity of 512 kB and a response time of 12-15 cycles, and the L3 cache 122 has a capacity of 4 MB and a response time of 35-40 cycles. Thus, at the first level, assuming 64B cache lines, the L1 caches 114-116 can each simultaneously store up to 512 cache lines and are expected to respond to a data access request within 4-8 control clock cycles.

[0029] In the described embodiments, data is compressed before being stored in at least some of the cache memories to reduce the size of the data, thereby increasing the amount of data (e.g., the number of cache lines) that can be stored simultaneously in the cache memories. For example, depending on the compression used for a given cache memory, a 64B cache line can be compressed to 48B, 32B, or another value. In these embodiments, the cache memories include control mechanisms, lookup mechanisms, etc. that can operate with compressed and uncompressed data and combinations thereof. For example, these mechanisms can store data of different sizes (e.g., 64B uncompressed cache lines and 32B compressed cache lines, etc.) within the cache memory, perform lookups on compressed and uncompressed data, retrieve both compressed and uncompressed data to be provided to an access entity from within the cache memory, and so on. Generally, in the described embodiments, the cache memories can operate with both compressed and uncompressed data and combinations thereof.

[0030] In some embodiments, different compression is used for the cache memories in each level of the cache memory hierarchy (or for each individual cache memory). In these embodiments, the specific compression or the most heavyweight compression that is allowed for each cache memory is selected for the cache memory at each level of the hierarchy (or for each cache memory), at least in part based on the required response time of these cache memories to access requests. Generally, the response time of cache memories that are lower in the hierarchy and thus have longer response times is proportionally reduced by the addition of the latency for decompressing higher heavyweight compression. For example, adding 8 clock cycles for data decompression to a 40-clock-cycle response time of a given cache memory will increase the response time by 20%, but the impact on shorter response times is more pronounced. Thus, heavier compression is allowed for lower cache memories in the hierarchy. For example, in some embodiments, the data in the L1 caches 114-116, which are the highest in the hierarchy and have the lowest response time, is not compressed, i.e., the data is stored in these cache memories in uncompressed form. In these embodiments, the data in the L2 caches 118-120, which have an intermediate response time, is allowed to be compressed using only a lighter-weight compression that adds a smaller amount of decompression latency to the response time. In these embodiments, the data in the L3 cache 122, which has the highest response time, is allowed to be compressed using a heavier-weight compression (but may not be compressed as described herein), which adds a larger amount of decompression latency to the response time.

[0031] In some embodiments, the compression and / or decompression of data is performed locally for the cache memories in the cache memory hierarchy. In these embodiments, some or all of the cache memories are associated with a compression-decompression functional block and may be located near the compression-decompression functional block. For example, the compression-decompression functional block can be incorporated into or associated with an interface for data signal routing or a bus on which data is received by the cache memory and provided to other cache memories via the interface. In these embodiments, the compression-decompression functional block includes circuit elements for performing compression and decompression operations on the compression of data for the corresponding cache memory. In Figure 1In the illustrated embodiments, the compression - decompression function blocks include compression - decompression (COM - DECOM) function blocks 124 - 126 respectively associated with L2 caches 118 - 120, and a compression - decompression function block 128 associated with L3 cache 122. In these embodiments, when data is received by or provided by the corresponding cache memory, the compression - decompression function blocks perform operations for compressing the received data, if necessary, for storage in the cache memory or decompressing data from the cache memory to be provided to another cache memory (or other access entity). Continuing with the above example, when L2 cache 118 receives uncompressed data from L1 cache 114 for storage therein (such as when data from execution subsystem 110 is propagating through the cache memory hierarchy), compression - decompression function block 124 compresses the data using a lighter - weight compression and then provides the compressed data for storage in L2 cache 118. Additionally, and continuing with this example, when L1 cache 114 requests compressed data from L2 cache 118, compression - decompression function block 124 receives the compressed data from L2 cache 118, decompresses the data from the lighter - weight compression to the uncompressed state, and then provides the uncompressed data to L1 cache 114.

[0032] Memory 104 is a function block that performs operations on the memory of electronic device 100 (e.g., the "main" memory). Memory 104 includes: volatile memory circuits, such as fourth - generation double - data - rate synchronous DRAM (DDR4 SDRAM) and / or other types of memory circuits, which are used to store data and instructions used by function blocks in electronic device 100; and control circuitry that is used to handle access to the data and instructions stored in the memory circuits and perform other control or configuration operations.

[0033] In some embodiments, data is compressed not only for storage in the cache memories that are compressed for storage in the cache memory hierarchy, but also for storage in the memory 104. In some embodiments, data is transferred between one or more functional blocks in a compressed state and can be compressed for communication. In some embodiments, data is in a compressed state for communication and is decompressed at the destination cache memory or access entity. For example, the compressed data retrieved from the L2 cache 118 may not be immediately decompressed to undo the effects of the above lighter-weight compression (i.e., in the compression-decompression functional block 124), but may be transferred to the L1 cache 114 before decompression. In these embodiments, the arrangement of the compression-decompression functional blocks may be different, with local compression and / or decompression mechanisms at one or more other locations in the electronic device 100 (e.g., the L1 cache 114, etc.). For example, data going from the L1 cache 114 to the L2 cache 118 may be compressed at the L1 cache 114 before being transferred to the L2 cache 118 and thus may arrive at the L2 cache 118 in a compressed state (where the lighter-weight compression has already been applied).

[0034] Although a particular arrangement of cache memories is shown in Figure 1 , in some embodiments, there are different numbers and / or arrangements of cache memories. For example, in some embodiments, each of the L1 caches 114 - 116 is divided into two separate cache memories in each of the cores 106 - 108, one cache memory for storing program code instructions and one cache memory for storing data. Generally, the described embodiments can operate with any arrangement of cache memories for which the data compression operations described herein can be performed. Additionally, although a number of compression-decompression functional blocks are shown in Figure 1 , in some embodiments, different arrangements of functional blocks are used to perform the compression and decompression of data as described herein. For example, the cache controllers in some or all of the cache memories can be modified to incorporate circuit elements for performing compression and / or decompression.

[0035] For illustrative purposes, the electronic device 100 is simplified. However, in some embodiments, the electronic device 100 includes additional or different functional blocks and elements. For example, the electronic device 100 may include a display subsystem, a power subsystem, an input-output (I / O) subsystem, etc. The electronic device 100 generally includes sufficient functional blocks and elements to perform the operations described herein.

[0036] The electronic device 100 can be or can be included in any device that performs computing operations. For example, the electronic device 100 can be or can be included in the following: a desktop computer, a laptop computer, a wearable computing device, a tablet computer, a piece of virtual or augmented reality equipment, a smart phone, an artificial intelligence (AI) or machine learning device, a server, a network device, a toy, a piece of audio-visual equipment, a household appliance, a vehicle, etc. and / or a combination thereof.

[0037] Data compression

[0038] In the described embodiments, data is compressed for storage in a cache memory within a cache memory hierarchy and is decompressed as needed to store the data in the cache memory or to use the data (recall that "data" as used herein refers to any information or value that can be stored in a cache or memory). Generally, "compressing" data involves operations for losslessly reducing the size of the data (e.g., number of bits or bytes) for data storage, data communication via various communication mechanisms, and other operations. Since the data is losslessly compressed, the compressed data can be "decompressed" to recover the full uncompressed value of the data.

[0039] There are various types of data compression and various implementations of each compression type known in the art. For example, pattern matching compression, value matching compression, zero-content compression, Lempel-Ziv (LZ) (and variants) compression, Markov compression, and delta compression are several of the many compression types, and different implementations have been proposed for some or all of the compression types. The described embodiments are not limited to any particular compression type, but can operate with any combination of compression that can be performed on data to be stored in the cache memory.

[0040] As described herein, in some embodiments, a combination of "lighter-weight" compression and "heavier-weight" compression is used. Generally, lighter-weight compression is or includes one or more types of compression that compress and / or decompress faster (and possibly significantly faster) than one or more types of compression that are or are included in heavier-weight compression. For example, in some embodiments, lighter-weight compression can be or include a compression type that can compress and / or decompress in 1 - 3 cycles (of a control clock), while heavier-weight compression can be or include a compression type that can compress and / or decompress in 6 - 10 cycles. The difference between lighter-weight compression and heavier-weight compression may be attributed to differences in one or more of the number of compression operations performed, the complexity of the compression operations, and / or the speed at which the compression operations are performed.

[0041] In some embodiments, although an attempt is made to compress data, there may be no compressible patterns, values, etc. in the data. For example, in some embodiments, leading zeros or ones may be removed for compression, and the data may not have leading zeros or ones. In such cases, compression may not work, and thus the data may remain in its initial state as it existed prior to the attempt to compress it (i.e., may have the same number of bits or bytes and the same values). In these embodiments, the data may be stored in some or all of the cache memories in an uncompressed state. For example, and continuing with the above example, when lightweight compression does not result in a reduction / compression of a given data slice, that data slice may be stored in the L2 cache 118 in an uncompressed state.

[0042] In some embodiments, the compression-decompression functional block may attempt to perform two or more compressions on the data in sequence until the data is compressed or the last compression in the sequence has been attempted. Continuing with the above example, the compression-decompression functional block 128 may first attempt to compress a given data slice using lightweight compression for storage in the L3 cache 122. When an attempt to compress the given data slice using lightweight compression is successful, the compressed data with only lightweight compression applied may be stored in the L3 cache 122. In such a case, when and if the compressed given data slice is provided from the L3 cache 122 to the L2 cache (e.g., the L2 cache 118), no decompression may be provided for the given data slice because lightweight compression was used for the data in the L2 cache. However, when lightweight compression is not successful (and thus the data remains uncompressed), the compression-decompression functional block 128 may attempt to compress the given data slice using heavyweight compression. Then, the given data slice compressed with heavyweight compression or uncompressed (if heavyweight compression is not successful) is stored in the L3 cache 122. When using heavyweight compression to compress the given data slice, the compression-decompression functional block 128 will decompress the given data slice to reverse the effect of the heavyweight compression before providing the given data slice to the L2 cache.

[0043] In some embodiments, and different from the examples in the previous paragraph, the compression-decompression functional block can apply two or more compressions to data in sequence. Generally, in these embodiments, a lighter-weight compression is used to compress a given data slice, and then a heavier-weight compression is used to compress the given data slice again (which includes the effect of the lighter-weight compression). The combined compression is performed on the data to be stored in the cache memory, where both the heavier-weight compression and the lighter-weight compression are allowed. For example, in some of these embodiments, the combined compression is allowed in the L3 cache 122. In these embodiments, the compression-decompression functional block that performs the combined compression can be configured to process variable-length data for each compression. In other words, because the lighter-weight compression may or may not result in compression of a given data slice, the heavier-weight compression should be able to handle both uncompressed data and data that has been compressed using the lighter-weight compression.

[0044] Compress data for storage in a cache memory of a cache memory hierarchy

[0045] In the described embodiments, data is compressed for storage in cache memories (e.g., L1 caches 114-116, L2 caches 118-120, and L3 cache 122) in the cache memory hierarchy. When data is copied upward through the cache memories in the cache memory hierarchy from the memory (e.g., memory 104), the compression-decompression functional blocks (e.g., compression-decompression functional blocks 124-128) perform operations for compressing and decompressing the data such that the data is stored in the cache memory using the specified compression. Additionally, when data is provided downward through the cache memories in the cache memory hierarchy by a core (e.g., core 106), the compression-decompression functional blocks perform operations for compressing and decompressing the data such that the data is stored in the cache memory using the specified compression. Figure 2 A block diagram is presented that shows the compression and decompression performed on data when the data is copied upward through the cache memories in the cache memory hierarchy, according to some embodiments. Figure 3 A block diagram is presented that shows the compression and decompression performed on data when the data is copied downward through the cache memories in the cache memory hierarchy, according to some embodiments. Note that Figures 2 to 3 The operations shown are presented as general examples of operations performed by some embodiments. The operations performed by other embodiments include different operations, operations performed in a different order, and / or operations performed by different entities or functional blocks.

[0046] For Figures 2 to 3In the example of, it is assumed that only one compression is applied to the data at a time. As described elsewhere herein, in some embodiments, in a specified cache memory (e.g., L3 cache 122), a sequence of two compressions can be applied to the data such that the data is first compressed using a first compression (e.g., a lighter-weight compression) and then the already-compressed data is compressed again using a second compression (e.g., a heavier-weight compression). Although no examples of compression combinations are provided, the operation is similar, except for the fact that two compressions can be applied to the data in sequence and the data is decompressed using the corresponding decompression operations.

[0047] For Figures 2 to 3 In the example of, when data is propagated to or from memory, various compressions and decompressions are described as being applied to the data at a specified time / in the corresponding cache memory. In some embodiments, some or all of the compression and / or decompression is not applied to the data and / or the compression and / or decompression is applied to the data in a different manner. For example, in some embodiments, for a specified type of data access (e.g., a higher-priority data access, etc.), uncompressed data is provided directly to the requesting entity (e.g., a core, etc.) without being stored in the cache memory. As another example, in some embodiments, uncompressed data fetched from memory is routed around to the requesting entity (such that the data is provided to the requesting entity as soon as possible), but a copy is also provided to some or all of the cache memories (possibly in parallel) and compressed before being stored therein.

[0048] For Figure 3 In the example of, it is assumed that a copy of the data to be stored in memory provided by a data source such as core 106 is stored in the cache memories down through the cache memory hierarchy (i.e., like a write-through cache memory). Thus, in this example, as the data copy moves down through the hierarchy towards memory, the copy is stored in each cache memory. In some embodiments, the copy of the data does not propagate down through the hierarchy in this manner, but is first written to the highest cache in the hierarchy (e.g., L1 caches 114 - 116) and subsequently propagated down via eviction from the cache memories (i.e., like a write-back cache memory). However, the compression-decompression operations described for Figure 3 are similar in these embodiments and are performed when and if the data copy is stored in each cache memory.

[0049] When a copy of the data in an uncompressed state (e.g., a 64B cache line from a specified address in memory) is fetched from memory 104 for storage in L3 cache 122, Figure 2The operation in [description] begins. Before storing the data in the L3 cache 122, the compression-decompression functional block 128 first attempts to compress the data using the lighter-weight compression 200. For this operation, the compression-decompression functional block 128 applies a compression operation according to the specific lighter-weight compression in use, such as pattern matching, value replacement, etc. If the lighter-weight compression is successful, the compression-decompression functional block 128 updates the metadata associated with the data to mark the data as having been compressed using the lighter-weight compression, and provides the compressed data to the L3 cache 122 for storage therein (shown by the line bypassing the heavier-weight compression 202).

[0050] When the attempt to compress the data using the lighter-weight compression is unsuccessful, and thus the data remains uncompressed (i.e., the data lacks the patterns, values, etc. to which the lighter-weight compression is applied), the compression-decompression functional block 128 attempts to compress the data using the heavier-weight compression 202. For this operation, the compression-decompression functional block 128 applies a compression operation according to the specific heavier-weight compression in use, such as pattern matching, value replacement, etc. If the heavier-weight compression is successful, the compression-decompression functional block 128 updates the metadata associated with the data to mark the data as having been compressed using the heavier-weight compression, and provides the compressed data to the L3 cache 122 for storage therein. Otherwise, if the attempt to compress the data using the heavier-weight compression is unsuccessful, and thus the data remains uncompressed (i.e., the data lacks the patterns, values, etc. to which the heavier-weight compression is applied), the compression-decompression functional block 128 updates the metadata associated with the data to mark the data as uncompressed, and provides the uncompressed data to the L3 cache 122 for storage therein.

[0051] When receiving a request to provide a data copy to the L2 cache 118, when the data is marked as having been compressed using the heavier-weight compression, the L3 cache 122 provides the data to the compression-decompression functional block 128. The compression-decompression functional block 128 applies the heavier-weight decompression 204 to reverse the heavier-weight compression of the data, thereby restoring the data to an uncompressed state. The compression-decompression functional block 128 then updates the metadata associated with the data to mark the data as uncompressed, and provides the data in the uncompressed state to the L3 cache 122 for forwarding to the L2 cache 118. Otherwise, when the data is marked as compressed using the lighter-weight compression or uncompressed, the L3 cache 122 immediately provides the data to the L2 cache 118 without involving the compression-decompression functional block 128 (shown by the line bypassing the heavier-weight decompression 204). When receiving data in an uncompressed state or compressed using the lighter-weight compression, the L2 cache 118 stores the data.

[0052] Upon receiving a request to provide a data copy to the L1 cache 114, when the data is marked as having been compressed using a lightweight compression, the L2 cache 118 provides the data to the compression-decompression functional block 124. The compression-decompression functional block 124 applies a lightweight decompression 206 to reverse the lightweight compression of the data, thereby restoring the data to an uncompressed state. The compression-decompression functional block 124 then updates the metadata associated with the data to mark the data as uncompressed and provides the uncompressed-state data to the L2 cache 118 for forwarding to the L1 cache 114. Otherwise, when the data is marked as uncompressed, the L2 cache 118 immediately provides the data to the L1 cache 114 without involving the compression-decompression functional block 124 (which is shown by the line bypassing the lightweight decompression 206). Upon receiving the data in the uncompressed state, the L1 cache 114 stores the data. Then, the L1 cache 114 can provide the data to an access entity, such as the core 106, which may have generated a request to store a data copy in the L1 cache 114.

[0053] When a data copy in the uncompressed state destined to be stored in the memory 104 is provided to the L1 cache 114 from a data source such as the core 106, Figure 3 the operations in begin. The L1 cache 114 stores the data copy in the uncompressed state (e.g., in a 64B cache line including the data) in the L1 cache 114, where the metadata associated with the data set indicates that the data is in the uncompressed state. The L1 cache 114 also forwards the data copy in the uncompressed state to the L2 cache 118 for storage therein.

[0054] Upon receiving the data copy, the L2 cache 114 provides the data to the compression-decompression functional block 124. The compression-decompression functional block 124 attempts to compress the data using a lightweight compression 300. For this operation, the compression-decompression functional block 124 applies a compression operation, such as pattern matching, value replacement, etc., according to the specific lightweight compression in use. If the lightweight compression is successful, the compression-decompression functional block 124 updates the metadata associated with the data to mark the data as having been compressed using a lightweight compression and provides the compressed data to the L2 cache 118 for storage therein. When the attempt to compress the data using a lightweight compression is unsuccessful, and thus the data remains uncompressed (i.e., the data lacks the patterns, values, etc. to which the lightweight compression is applied), the compression-decompression functional block 124 provides the uncompressed data to the L2 cache 118 for storage therein. The L2 cache 118 also forwards the data copy in the uncompressed state or compressed using a lightweight compression to the L3 cache 122 for storage therein.

[0055] When a data copy is received, when the metadata associated with the data indicates that the data is compressed using a lighter-weight compression, the L3 cache 122 stores the compressed data (as shown by the line bypassing the heavier-weight compression 302). Otherwise, when the metadata associated with the data indicates that the data is not compressed, the L3 cache 122 provides the data to the compression-decompression functional block 128. The compression-decompression functional block 128 attempts to compress the data using the heavier-weight compression 302. For this operation, the compression-decompression functional block 128 applies a compression operation according to the specific lighter-weight compression in use, such as pattern matching, value replacement, etc. If the heavier-weight compression is successful, the compression-decompression functional block 128 updates the metadata associated with the data to mark the data as having been compressed using the heavier-weight compression and provides the compressed data to the L3 cache 122 for storage therein. When the attempt to compress the data using the heavier-weight compression is unsuccessful, and thus the data remains uncompressed (i.e., the data lacks the patterns, values, etc. to which the lighter-weight compression was applied), the compression-decompression functional block 128 provides the uncompressed data to the L3 cache 122 for storage therein.

[0056] The L3 cache 122 also provides a data copy to the memory 104. In some embodiments, the data is stored in the memory 104 in an uncompressed state. In these embodiments, when data compressed using a lighter-weight compression is received from the L2 cache 118, the L3 cache 122 provides the data to the compression-decompression functional block 128. The compression-decompression functional block 128 applies the lighter-weight decompression operation 304 to reverse the lighter-weight compression of the data, thereby restoring the data to an uncompressed state. The L3 cache 122 then provides the data in the uncompressed state to the memory. Otherwise, when the L3 cache 122 receives uncompressed data, the L3 cache 122 immediately provides the data to the memory without using the compression-decompression functional block 128 to decompress the data (as shown by the line bypassing the lighter-weight decompression 304). Upon receiving the data in the uncompressed state, the memory stores the data. In some embodiments, the data is stored in the memory in a compressed state. In these embodiments, and depending on the specific compression used in the memory, the L3 cache 122 may provide the data in the compressed state to the memory 104 for storage therein, or the L3 cache 122 may provide the data in the uncompressed state as described, and the compression-decompression mechanism in the memory may compress the data using the corresponding compression before storing the data in the memory 104.

[0057] Process for compressing data for storage in a cache memory

[0058] In the described embodiments, the electronic device performs an operation for compressing data using a specified compression for storage in a cache memory in a cache memory hierarchy of the electronic device. Figure 4 A flowchart is presented that shows a process for compressing data for storage in a cache memory according to some embodiments. Note that Figure 4 the operations shown are presented as general examples of operations performed by some embodiments. Operations performed by other embodiments include different operations, operations performed in a different order, and / or operations performed by different entities or functional blocks.

[0059] For Figure 4 the example in, it is assumed that a first cache memory (e.g., L3 cache 122) is allowed to store (along with uncompressed data) data compressed using either a first compression (e.g., a heavier-weight compression) or a second compression (e.g., a lighter-weight compression), and thus the corresponding compression-decompression functional block (e.g., compression-decompression functional block 128) is capable of compressing and decompressing data using the first compression and the second compression. It is also assumed that a second cache memory (e.g., L2 cache 118) is allowed to store (along with uncompressed data) data compressed using only the second compression, and thus the corresponding compression-decompression functional block (e.g., compression-decompression functional block 124) is capable of compressing and decompressing data using only the second compression. For this example, the second cache memory is not allowed to store data that has been compressed using the first compression. It is also assumed that a third cache memory (e.g., L1 cache 114) stores only uncompressed data.

[0060] Additionally, for Figure 4 the example in, only one of the first compression and the second compression is applied to the data for storage in the first cache memory. However, this is not required. Figure 5 A flowchart is presented that shows an embodiment in which a compression sequence is applied to the data for storage in the first cache memory.

[0061] When the compression-decompression functional block receives data in an uncompressed state, Figure 4 the operation in begins (step 400). For this operation, the compression-decompression functional block receives data in an uncompressed state, such as a 64B cache line, from a memory (e.g., memory 104) or another source (e.g., a network interface, a processor, etc.). In the "uncompressed" state, the full set of bits / bytes of the data according to the data type is presented and configured to represent the value of the data, and the data has not been compressed. For example, if the data includes one or more 32-bit floating-point values, all 32 bits representing each of the one or more values (e.g., sign, exponent, and fraction) are presented.

[0062] Then, the compression-decompression functional block compresses the data using one of a first compression and a second compression (step 402). Each of the first compression and the second compression includes one or more compression operations during which some or all of the bits / bytes of the uncompressed data are deleted, reduced, replaced, or otherwise altered. The specific one or more compression operations of each of the first compression and the second compression and thus the changes made to the uncompressed data during the respective compression depend on the type of compression. For example, the first compression and / or the second compression may be or include one or more of pattern matching compression, value matching compression, zero content compression, Lempel-Ziv (LZ) (and variants) compression, Markov compression, delta compression, etc.

[0063] For the operation in step 402, when compressing the data using one of the first compression and the second compression, the compression-decompression functional block first attempts to compress the data using the second compression (e.g., the lighter-weight compression). The second compression is attempted before the first compression (e.g., the heavier-weight compression) because both the first cache memory and the second cache memory are allowed to store and can decompress the data compressed using the second compression. Thus, when copying data from the first cache memory to the second cache memory, the data can be copied without decompressing the data, which avoids the need for the compression-decompression functional block to perform the corresponding decompression operation. When compressing the data using the second compression, the compression-decompression functional block sets metadata (e.g., one or more bits) associated with the data to indicate that the data is compressed using the second compression and provides the compressed data to the first cache memory for storage therein (step 404).

[0064] Despite the attempt of the second compression, the second compression may not result in compressed data because the data lacks the corresponding patterns, values, etc., and thus the data remains uncompressed. In this case, the compression-decompression functional block attempts to compress the data using the first compression. The compression-decompression functional block attempts the first compression because, although the data needs to be decompressed when copying the data compressed using the first compression from the first cache memory to the second cache memory, it is still desirable to compress the data for storage in the first cache memory to better utilize the available storage in the first cache memory. When compressing the data using the first compression, the compression-decompression functional block sets the metadata associated with the data to indicate that the data is compressed using the first compression and provides the compressed data to the first cache memory for storage therein (step 404).

[0065] For Figure 4In the example of [[ID=]], it is assumed that the data is successfully compressed using one of the first compression and the second compression. If the data cannot be compressed despite attempts to compress it using the first compression and the second compression, the data will be stored in the first cache memory in an uncompressed state. In this case, in step 408, only the uncompressed data is copied from the first cache memory to the second cache memory.

[0066] When retrieving data from the first cache memory for storage in the second cache memory, when the first compression is used to compress the data, the compression - decompression functional block decompresses the data to reverse the effect of the first compression on the data, thereby restoring the data to an uncompressed state (step 406). For this operation, the compression - decompression functional block performs one or more decompression operations during which some or all of the bits / bytes of the data that were previously deleted, reduced, replaced, or otherwise altered are restored to their original full values. As with compression, the specific one or more decompression operations performed depend on the type of the first compression. During this operation, the compression - decompression functional block places the data in a state (uncompressed) in which the data can be stored in the second cache memory. Then, the data that has been compressed using the second compression (compressed in step 402) or is in an uncompressed state is provided to the second cache memory for storage therein (step 408).

[0067] When retrieving data from the second cache memory for storage in the third cache memory, when the second compression is used to compress the data, the compression - decompression functional block decompresses the data to reverse the effect of the second compression on the data, thereby restoring the data to an uncompressed state (step 410). For this operation, the compression - decompression functional block performs one or more decompression operations during which some or all of the bits / bytes of the data that were previously deleted, reduced, replaced, or otherwise altered are restored to their original full values. As with compression, the specific one or more decompression operations performed depend on the type of the second compression. During this operation, the compression - decompression functional block places the data in a state (uncompressed) in which the data can be stored in the third cache memory. Then the data in an uncompressed state is provided to the second cache memory for storage therein (step 412).

[0068] Now turning to Figure 5 , Figure 5 FIG. [FIG. NUMBER] presents a flowchart showing a process for compressing data for storage in a cache memory according to some embodiments. Note that Figure 5 the operations shown are presented as general examples of operations performed by some embodiments. Operations performed by other embodiments include different operations, operations performed in a different order, and / or operations performed by different entities or functional blocks.

[0069] For Figure 5 the example in, assume that the first cache memory (e.g., L3 cache 122) is allowed to store data compressed using a sequence of a second compression (e.g., a lighter - weight compression) and a first compression (e.g., a heavier - weight compression) (along with the uncompressed data), such that the corresponding compression - decompression functional block (e.g., compression - decompression functional block 128) can compress and decompress data using the first compression and the second compression. Also assume that the second cache memory (e.g., L2 cache 118) is allowed to store data compressed using only the second compression (along with the uncompressed data), such that the corresponding compression - decompression functional block (e.g., compression - decompression functional block 124) can compress and decompress data using only the second compression. For this example, the second cache memory is not allowed to store data that has been compressed using the first compression. Also assume that the third cache memory (e.g., L1 cache 114) stores only uncompressed data.

[0070] Additionally, and different from the example in Figure 4 For the example in Figure 5 a sequence of both the second compression and the first compression is applied to the data for storage in the first cache memory. In other words, the data is compressed using the second compression, and the already - compressed data is compressed again using the first compression. This combination of compressions may result in a greater reduction in data size than either compression alone. In some embodiments, the latter compression (i.e., the first compression) includes one or more operations for handling data of different lengths, because the data may have been successfully compressed using the second compression or may not have been compressed (in cases where the data lacks patterns, values, etc. that can be compressed by the second compression).

[0071] Furthermore, for the example in Figure 5 assume that the data has been successfully compressed using a sequence of the second compression and the first compression. If the data cannot be compressed despite attempts to compress it using the first compression and / or the second compression, then the data will be stored in the first cache memory in an uncompressed state or some other single - compressed state. In such a case, Figure 5 the operation of Figure 4 is changed to operate on the compressed data using single - compression (similar to

[0072] When the compression - decompression functional block receives data in an uncompressed state, Figure 5The operation in [it] starts (step 500). For this operation, the compression-decompression functional block receives data in an uncompressed state, such as a 64B cache line, from a memory (e.g., memory 104) or another source (e.g., network interface, processor, etc.). In the "uncompressed" state, the complete set of bits / bytes of the data according to the data type is presented and configured to represent the value of the data, and the data has not been compressed. For example, if the data includes one or more 32-bit floating-point values, all 32 bits representing each of the values (e.g., sign, exponent, and fraction) are presented.

[0073] Then, the compression-decompression functional block compresses the data using a sequence of a second compression followed by a first compression (step 502). Each of the first compression and the second compression includes one or more compression operations during which some or all of the bits / bytes of the data are deleted, reduced, replaced, or otherwise altered. The specific one or more compression operations of each of the first compression and the second compression and thus the changes made to the data during the respective compressions depend on the type of compression. For example, the first compression and / or the second compression can be or include one or more of pattern matching compression, value matching compression, zero-content compression, Lempel-Ziv (LZ) (and variants) compression, Markov compression, delta compression, etc.

[0074] For the operation in step 502, both the second compression and the first compression are performed on the data. The second compression is performed first, thus initially compressing the data using the second compression because the second cache memory is allowed to store data compressed using the second compression and the data can be decompressed therefrom. Then the first compression is performed to attempt to further reduce the data size. This compression order enables the data to be decompressed to reverse the effect of the first compression and restore the data to have only the effect of the second compression, thus preparing to store a copy in the second cache memory. After compressing the data using the sequence of the second compression and the first compression, the compression-decompression functional block sets metadata (e.g., one or more bits) associated with the data to indicate that the data is compressed using the sequence of the second compression and the first compression, and provides the compressed data to the first cache memory for storage therein (step 504).

[0075] When retrieving data from the first cache memory for storage in the second cache memory, the compression-decompression functional block decompresses the data to reverse the effect of the first compression on the data, thereby restoring the data to the state in which the data was compressed using only the second compression (step 506). For this operation, the compression-decompression functional block performs one or more decompression operations during which some or all of the bits / bytes of the data that were previously deleted, reduced, replaced, or otherwise altered are restored to the values after the second compression. As with compression, the particular one or more decompression operations performed depend on the type of the first compression. During this operation, the compression-decompression functional block places the data in a state in which the data can be stored in the second cache memory (compressed using only the second compression), and updates the metadata associated with the data to indicate that the data is compressed using only the second compression. Then, the data compressed using the second compression (compressed in step 502) is provided to the second cache memory for storage therein (step 508).

[0076] When retrieving data from the second cache memory for storage in the third cache memory, the compression-decompression functional block decompresses the data to reverse the effect of the second compression on the data, thereby restoring the data to the uncompressed state (step 510). For this operation, the compression-decompression functional block performs one or more decompression operations during which some or all of the bits / bytes of the data that were previously deleted, reduced, replaced, or otherwise altered are restored to the original full values. As with compression, the particular one or more decompression operations performed depend on the type of the second compression. During this operation, the compression-decompression functional block places the data in a state in which the data can be stored in the third cache memory (uncompressed). Then the data in the uncompressed state is provided to the second cache memory for storage therein (step 512).

[0077] In some embodiments, an electronic device (e.g., electronic device 100 and / or a portion thereof) uses code and / or data stored on a non-transitory computer-readable storage medium to perform some or all of the operations described herein. More specifically, when performing the described operations, the electronic device reads the code and / or data from the computer-readable storage medium and executes the code and / or uses the data. The computer-readable storage medium can be any device, medium, or combination thereof that stores code and / or data for use by the electronic device. For example, the computer-readable storage medium can include, but is not limited to, volatile memory and / or non-volatile memory, including flash memory, random access memory (e.g., eDRAM, RAM, SRAM, DRAM, DDR4 SDRAM, etc.), read-only memory (ROM), and / or magnetic or optical storage media (e.g., disk drives, tapes, CDs, DVDs, etc.).

[0078] In some embodiments, one or more hardware modules perform the operations described herein. For example, the hardware modules can include, but are not limited to, one or more processors / cores / central processing units (CPUs), application specific integrated circuit (ASIC) chips, neural network processors or accelerators, field programmable gate arrays (FPGAs), computing units, embedded processors, graphics processing units (GPUs) / graphics cores, pipelines, accelerated processing units (APUs), functional blocks, controllers, compression-decompression functional blocks, and / or other programmable logic devices. When activated, the hardware modules perform some or all of the operations. In some embodiments, the hardware modules include one or more general-purpose circuits configured to perform the operations by executing instructions (program code, firmware, etc.).

[0079] In some embodiments, data structures representing some or all of the structures and mechanisms described herein (e.g., electronic device 100, compression-decompression functional blocks 124-128, cache memories in the cache memory hierarchy, and / or a portion thereof) are stored on a non-transitory computer-readable storage medium that includes a database or other data structure readable by the electronic device and directly or indirectly used to fabricate hardware including the structures and mechanisms. By way of example, the data structure can be a behavioral-level description or a register transfer level (RTL) description of a hardware function in a high-level design language (HDL) (e.g., Verilog or VHDL). The description can be read by a synthesis tool that can synthesize the description to produce a netlist that includes a listing of gates / circuit elements from a synthesis library that represents the functionality of the hardware including the above structures and mechanisms. The netlist can then be placed and routed to produce a data set that describes the geometry to be applied to a mask. The mask can then be used in various semiconductor manufacturing steps to produce one or more semiconductor circuits (e.g., integrated circuits) corresponding to the above structures and mechanisms. Alternatively, the database on the computer-accessible storage medium can be a netlist (with or without a synthesis library) or a data set (as needed), or a graphics data system (GDS) II data.

[0080] In this specification, variables or unspecified values (i.e., general descriptions of values in the absence of a particular instance of a value) are represented by letters such as N. As used herein, although similar letters may be used at different locations in this specification, the variables and unspecified values in each case are not necessarily the same, i.e., there may be different variable quantities and values intended for some or all of the general variables and unspecified values. In other words, in this specification, N and any other letter used to represent variables and unspecified values are not necessarily related to each other.

[0081] As used herein, the expressions “etc.” or “and so forth” are intended to present one and / or a situation, that is, an equivalent of “at least one” among the elements associated with etc. in the list. For example, in the statement “The electronic device performs a first operation, a second operation, etc.”, the electronic device performs at least one of the first operation, the second operation, and other operations. In addition, the elements in the list associated with and so forth are merely examples in a set of examples - and at least some of the examples may not appear in some embodiments.

[0082] The previous description of the embodiments has been given for purposes of illustration and description only. The previous description is not intended to be exhaustive or to limit the embodiments to the disclosed form. Accordingly, many modifications and variations will be obvious to practitioners in the art. Additionally, the above disclosure is not intended to limit the embodiments. The scope of the embodiments is defined by the appended claims.

Claims

1. An electronic device for compressing data for storage in a cache memory, the electronic device comprising: A cache memory hierarchy including a first cache memory and a second cache memory, the first cache memory being lower in the hierarchy than the second cache memory; And A processor configured to: Receive data in an uncompressed state; Compress the data by using a second compression and then, when using the second compression does not result in compression of the data, using a first compression; After compressing the data, provide the data to the first cache memory for storage therein; When retrieving the data from the first cache memory for storage in the second cache memory, when the data was compressed using the first compression, decompress the data to reverse the effect of the first compression on the data, thereby restoring the data to the uncompressed state; And Provide the data compressed using the second compression or in the uncompressed state to the second cache memory for storage therein.

2. The electronic device according to claim 1, further comprising: A third cache memory in the cache memory hierarchy, the second cache memory being lower in the hierarchy than the third cache memory, wherein: When retrieving the data from the second cache memory for storage in the third cache memory, when the data was compressed using the second compression, the processor: Decompresses the data to reverse the effect of the second compression on the data, thereby restoring the data to the uncompressed state; and Provides the data to the third cache memory for storage therein.

3. The electronic device according to claim 2, wherein: The first compression is a heavier-weight compression and the second compression is a lighter-weight compression; and Decompressing the data using the lighter-weight compression is faster than decompressing the data using the heavier-weight compression.

4. The electronic device according to claim 1, wherein compressing the data using one of the first compression or the second compression comprises: Performing one or more second compression operations of the second compression on the data, the one or more second compression operations compressing data having a second property; And When the one or more compression operations of the second compression do not result in compression of the data because the data does not have the second property, performing one or more first compression operations of the first compression on the data, the one or more first compression operations compressing data having a first property.

5. The electronic device according to claim 4, wherein one or both of the first compression operation and the second compression operation include at least one of pattern matching compression, value matching compression, zero content compression, Lempel-Ziv based compression, Markov compression, and delta compression.

6. The electronic device according to claim 4, wherein when neither the first property nor the second property exists in the data such that neither the first compression nor the second compression results in compressing the data, the data is stored in the first cache memory and the second cache memory in the uncompressed state.

7. The electronic device according to claim 1, wherein after compressing data using one of the first compression or the second compression, the processor is configured to: update metadata associated with the data to indicate that the data is compressed using the respective one of the first compression or the second compression.

8. The electronic device according to claim 1, wherein the processor is further configured to: receive other data in the uncompressed state; compress the other data using a sequence of the second compression and the first compression, wherein the first compression supports variable input formats; after compressing the other data, provide the other data to the first cache memory for storage therein; when retrieving the other data from the first cache memory for storage in the second cache memory, decompress the other data to reverse the effect of the first compression on the other data, thereby restoring the other data to a state where the other data is compressed only using the second compression; and provide the other data compressed using the second compression to the second cache memory for storage therein.

9. The electronic device according to claim 1, wherein the processor receives the data from a memory.

10. The electronic device according to claim 9, wherein the data stored in the memory is compressed using a third compression, and the data is decompressed to restore the data to the uncompressed state before forwarding the data to the processor.

11. The electronic device according to claim 9, wherein the data is compressed using a third compression for communication via one or more communication links between the memory and the processor.

12. A method for compressing data for storage in a cache memory of an electronic device including a cache memory hierarchy and a processor, the cache memory hierarchy including a first cache memory and a second cache memory, the first cache memory being lower in the hierarchy than the second cache memory, the method comprising: receiving, by the processor, data in an uncompressed state; compressing, by the processor, the data by using a second compression and then, when using the second compression does not result in compression of the data, using a first compression; after compressing the data, providing, by the processor, the data to the first cache memory for storage therein; When retrieving the data from the first cache memory for storage in the second cache memory, when using the first compression to compress the data, the processor decompresses the data to reverse the effect of the first compression on the data, thereby restoring the data to the uncompressed state; and the processor provides the data that is compressed using the second compression or in the uncompressed state to the second cache memory for storage therein.

13. The method according to claim 12, wherein: The electronic device further includes a third cache memory in the cache memory hierarchy, and the second cache memory is lower in the hierarchy than the third cache memory; and the method further includes: When retrieving the data from the second cache memory for storage in the third cache memory, when using the second compression to compress the data: The processor decompresses the data to reverse the effect of the second compression on the data, thereby restoring the data to the uncompressed state; and the processor provides the data to the third cache memory for storage therein.

14. The method according to claim 13, wherein: The first compression is a heavier-weight compression, and the second compression is a lighter-weight compression; and Decompressing the data using the lighter-weight compression is faster than decompressing the data using the heavier-weight compression.

15. The method according to claim 12, wherein compressing the data using one of the first compression or the second compression includes: Performing one or more second compression operations of the second compression on the data, the one or more second compression operations compressing data having a second property; and When the one or more compression operations of the second compression do not result in compressing the data because the data does not have the second property, performing one or more first compression operations of the first compression on the data, the one or more first compression operations compressing data having a first property.

16. The method according to claim 15, wherein one or both of the first compression operation and the second compression operation include at least one of pattern matching compression, value matching compression, zero content compression, Lempel-Ziv based compression, Markov compression, and delta compression.

17. The method according to claim 15, wherein when neither the first property nor the second property exists in the data such that neither the first compression nor the second compression results in compressing the data, the method further includes: Storing the data in the uncompressed state in the first cache memory and the second cache memory.

18. The method according to claim 12, wherein the method further includes after compressing the data using one of the first compression or the second compression: The processor updates the metadata associated with the data to indicate that the data is compressed using a respective one of the first compression or the second compression.

19. The method according to claim 12, wherein the method further comprises: receiving, by the processor, other data in the uncompressed state; compressing, by the processor, the other data using a sequence of the second compression and the first compression, the first compression supporting a variable input format; after compressing the other data, providing, by the processor, the other data to the first cache memory for storage therein; when retrieving the other data from the first cache memory for storage in the second cache memory, decompressing the other data to reverse the effect of the first compression on the other data, thereby restoring the other data to a state where the other data is compressed only using the second compression; and providing, by the processor, the other data compressed using the second compression to the second cache memory for storage therein.

20. The method according to claim 12, further comprising: receiving, by the processor, the data from a memory, the data being compressed using a third compression for storage in the memory and being decompressed for transmission to the processor.

Citation Information

Patent Citations

  • glow wire candle

    DE620717C

  • Systems and methods for reducing latency for accessing compressed memory using stratified compressed memory architectures and organization

    US20080055323A1

  • Cache system and a method of operating a cache memory

    US20130311722A1

  • Raid system performance enhancement using compressed data and byte addressable storage devices

    US20170286220A1