Compressed metadata assisted computation
Patent Information
- Application Number
- CN202180065679.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-25
- Filing Date
- 2021-09-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-09-17
Smart Images

Figure CN116194902B_ABST
Abstract
Description
Background Technology
[0001] In addition to the main memory in computing devices, most modern computing devices provide at least one level of cache memory (or "cache"). Generally, a cache is a small, fast-access memory used to store a limited number of copies of data and instructions that will be used to perform various operations (e.g., computations) closer to the functional blocks within the computing device. Caches are typically implemented using high-speed memory circuitry such as static random-access memory (SRAM) integrated circuits and other types of memory circuitry. Compressing data in the cache allows for lower latency access to larger amounts of data, thus further improving the overall performance of the computing system. Attached Figure Description
[0002] This disclosure is illustrated by way of example and not limitation in the figures of the accompanying drawings.
[0003] Figure 1 An implementation scheme for the computing device is shown.
[0004] Figure 2A and Figure 2B An implementation scheme for the redundant address calculation unit is shown.
[0005] Figure 3 Components in a processing unit according to one embodiment are shown.
[0006] Figure 4 Components in a processing unit according to one embodiment are shown.
[0007] Figure 5 The process for performing auxiliary calculations for compressed metadata is illustrated according to one implementation scheme. Detailed Implementation
[0008] The following description illustrates numerous examples of specific details, such as particular systems, components, methods, etc., to provide a good understanding of the implementation schemes. However, it will be apparent to those skilled in the art that at least some implementation schemes can be practiced without these specific details. In other instances, well-known components or methods are not described in detail, or are presented in a simple block diagram format to avoid unnecessarily obscuring the implementation schemes. Therefore, the specific details set forth are merely exemplary. Specific implementations may vary as a result of these exemplary details and are still contemplated within the scope of the implementation schemes.
[0009] Compression of data stored in a cache effectively increases cache capacity and link bandwidth to and from the cache. In one implementation, the process of compressing data into a cache generates compression metadata associated with the compressed data. The compression metadata contains information indicating which values in the compressed data blocks are similar to or identical to each other and which values match a small set of constant values (e.g., 0 or -1).
[0010] In one implementation, when compressed metadata indicates that the data values involved in the operations are identical, cached compressed metadata serves as the basis for deduplication of redundant operations, such as register loads and arithmetic operations. The computational cost of identifying similar or identical data values is already incurred by the compression mechanism; therefore, deduplication is performed with minimal additional computational cost by comparing existing compressed metadata values associated with different addresses. Once a set of identical values is identified, redundant operations involving those identical values are removed, and operations that depend on the removed operations are updated to depend on the remaining equivalent operations. The removal of redundant operations reduces the number of computations and / or load operations performed. This reduces the load on the arithmetic logic unit (ALU) and physical register file, thereby effectively increasing link bandwidth and ultimately leading to increased throughput and reduced power consumption.
[0011] Figure 1 An embodiment of a computing system 100 implementing the aforementioned compressed metadata-assisted computing mechanism is shown. Generally, the computing system 100 is embodied as any of a variety of different types of devices, including but not limited to laptop or desktop computers, mobile devices, servers, etc. The computing system 100 includes multiple components 102-108 that communicate with each other via a bus 101. In the computing system 100, each of the components 102-108 is capable of communicating directly with any of the other components 102-108 via the bus 101 or via one or more other components 102-108. The components 101-108 of the computing system 100 are contained within a single physical housing, such as a laptop or desktop computer chassis or a mobile phone enclosure. In an alternative embodiment, some of the components of the computing system 100 are embodied as peripheral devices, such that the entire computing system 100 does not reside within a single physical housing.
[0012] The computing system 100 also includes user interface devices for receiving or providing information to a user. Specifically, the computing system 100 includes input devices 102, such as a keyboard, mouse, touchscreen, or other devices for receiving information from a user. The computing system 100 displays information to the user via a display 105, such as a monitor, LED display, liquid crystal display, or other output device.
[0013] The computing system 100 also includes a network adapter 107 for transmitting and receiving data via wired or wireless networks. The computing system 100 also includes one or more peripheral devices 108. Peripheral devices 108 may include mass storage devices, location detection devices, sensors, input devices, or other types of devices used by the computing system 100.
[0014] The computing system 100 includes one or more processing units 104, which can operate in parallel in the case of multiple processing units 104. The processing units 104 are configured to receive and execute instructions 109 stored in a memory subsystem 106. In one embodiment, each processing unit 104 includes multiple computing nodes residing on a common integrated circuit substrate. The memory subsystem 106 includes memory devices used by the computing system 100, such as random access memory (RAM) modules, read-only memory (ROM) modules, hard disks, and other non-transitory computer-readable media.
[0015] Some implementations of the computing system 100 may include more than Figure 1 The illustrated embodiments have fewer or more components. For example, some embodiments are implemented without any display 105 or input device 102. Other embodiments have more than one particular type of component; for example, embodiments of computing system 100 may have multiple buses 101, network adapters 107, memory devices 106, etc.
[0016] According to the implementation plan, Figure 2A and Figure 2B The block diagram illustrates the functionality of redundant address calculation logic units 200 and 250. Redundant address calculation logic units 200 and 250 calculate addresses associated with redundant operations indicated by two exemplary types of compressed metadata, C-Pack and Base-Delta-Immediate (BDI), which are cache compression schemes implemented in computing system 100. In addition to the C-Pack and BDI compression schemes discussed herein, other compression schemes are used in alternative implementations, where such compression schemes provide compressed metadata indicating which compressed data values are similar or identical.
[0017] The C-Pack compression scheme uses compact codes to replace frequently occurring data words in representing data in the cache. Frequently occurring words in the compressed data are stored in a dictionary. Common data words (e.g., all zeros) are encoded using a single compact code, rather than being explicitly stored in the dictionary; therefore, dictionary lookups are not used to decode these values. For other values, the C-Pack compression scheme uses compact codes and pointers to entries in the dictionary containing the complete word to represent the value in the cache. Thus, memory is saved when multiple identical data words are represented in the cache using compact references to the same data words in the dictionary. Additionally, partially matching words can be represented in the cache using references to matching portions of words stored in the dictionary or multiple portions of different words in the dictionary.
[0018] Therefore, the similarity between data words in the cache is encoded in the metadata (including code and pointers) used to represent the data words. Specifically, identical code and pointers reference the same constants or dictionary entries, and thus encode the same values. Similarity is determined by comparing metadata values without accessing the actual values in the dictionary.
[0019] refer to Figure 2A The redundant address calculation logic 200 receives C-Pack metadata 201, including code and pointers, and compares the metadata values to identify duplicate address 204 (i.e., the address storing duplicate values) and zero address 206 (i.e., the address storing zero). In one embodiment, the C-Pack metadata 201 includes code representing each 4-byte value, and the calculation logic 200 identifies addresses 204 and 206 based on three possible codes: "00", "01", and "10". Code "00" indicates that the value at the corresponding address is zero. Code "01" indicates that a dictionary index of the 4-byte value is calculated based on the number of previously occurring "10" codes, which is calculated in box 202. Code "10" indicates that the index is directly included in the metadata.
[0020] When the index for the "01" code has been calculated in box 202, that index is compared at box 203 with the explicitly included index corresponding to the "10" code. The set of repeating addresses 204 are the addresses corresponding to the matching indexes found in box 203. Box 205 identifies addresses associated with the "00" encoding in metadata 201; these addresses contain the value zero.
[0021] The codes other than "00", "01", and "10" indicate that the value they represent partially matches one or more values in the dictionary; therefore, a dictionary lookup is performed to identify the matching value. Since dictionary lookups incur a similar performance penalty to normal cache accesses, these codes are ignored by computation logic 200.
[0022] Data blocks stored according to the BDI compression scheme include a base value and a series of Δ values, each Δ value representing a word in the original data block. Each Δ value is also associated with a value indicating whether the base of the Δ is zero or the recorded base value. (See reference) Figure 2B By finding a Δ value with a zero value at box 252 and then comparing the resulting set with a Δ value with a zero base at box 253, redundant address calculation logic 250 identifies a zero address 254 (i.e., the address storing the zero value). The resulting zero address 254 has a zero Δ relative to the zero base and therefore has a zero value.
[0023] In box 255, Δ values with zero Δ and non-zero bases are identified; these addresses 256 have zero Δ relative to the same base and therefore contain matching values. In one embodiment, a similar calculation is performed for Δ values other than zero. Addresses with the same Δ value relative to the same base are identified as matching each other. When a first Δ value equals a second Δ value and a first base value equals a second base value, the first value represented by the first Δ value and the first base value is equal to the second value represented by the second Δ value and the second base value.
[0024] Figure 3 Components in an out-of-order execution pipeline of a processing unit 300 according to one embodiment are shown. This processing unit utilizes matching address information (containing duplicate and / or constant value addresses, such as zero addresses) to reduce the number of redundant data cache accesses performed. Processing unit 300 corresponds to processing unit 104 in computing system 100. In the illustrated configuration, processing unit 300 reduces the number of data cache accesses by identifying addresses containing duplicate data values. Data is read from the cache once, and copied for each duplicate address.
[0025] In processing unit 300, instruction extraction unit 301 retrieves an instruction to be executed (e.g., instruction 109). This instruction specifies the operation to be performed and where to find the data value to which the operation will be performed. Decoding unit 302 converts the instruction into a signal for performing the operation in the processing circuit.
[0026] The register and reorder buffer (ROB) allocation unit 303 facilitates out-of-order instruction execution by allocating registers in the ROB / physical register file 305 for storing operands and the results of specified operations. The allocation unit 303 also updates the register alias table (RAT) 304 to map logical registers to physical registers allocated in the physical register file 305.
[0027] The Arithmetic Logic Unit (ALU) 306 performs logical and arithmetic operations based on the opcode of the received instruction. For load and store operations, the address is calculated in the address calculation unit 307, and the operation is queued in the load / store queue 308. Operations in the load / store queue transfer data between the data cache 309 and the physical register file 305.
[0028] Data cache 309 is a memory device that stores data that may be used by incoming instructions and has access to data in the rest of the memory hierarchy (e.g., main memory 106). Data stored in data cache 309 is compressed when written by compression engine 310. For each data block to be stored in cache 309, compression engine 310 generates compressed metadata representing the data, such as compact compression codes, Δ, and / or other metadata values depending on the compression scheme used. Metadata, rather than the original data, is stored in data cache 309. In an alternative embodiment, some or all of the compressed metadata is stored in another memory device instead of data cache 309. The compressed cached data is also decompressed by compression engine 310 upon being read by reconstructing the original data based on the stored compressed metadata.
[0029] Load / store queue 308 contains a list of load operations, each specifying an address in data cache 309 from which data will be retrieved and loaded into a register in physical register file 305. When a load operation is executed from queue 308, the address 321 associated with the load operation is provided to data cache 309. Compression engine 310 decompresses the data 322 stored at the provided address and returns the decompressed data 322 to load / store queue 308. As specified by the load operation, load / store queue 308 loads the data 322 into a register in physical register file 305.
[0030] The processing unit 300 includes a redundant address calculation unit 311, which receives compressed metadata 323 corresponding to the address 321 being accessed. The compressed metadata 323 is provided by the compression engine 310, which reads the compressed metadata 323 as part of the decompression process.
[0031] In one implementation, the redundancy address calculation unit 311 receives compressed metadata 323, including metadata for the current load operation in the load / store queue 308 and metadata for one or more addresses associated with other load operations in the queue 308. The redundancy address calculation unit 311 compares the values of the metadata 323 to find multiple sets of matching addresses 324, as previously referenced. Figure 2A and Figure 2BAs discussed. Therefore, each request to data cache 309 also results in any matching address 324 being returned along with data 322.
[0032] Each set of two or more duplicate addresses in matching address 324 represents a set of load operations where one or more operations are redundant. Matching address 324 is received in load / store queue 308 by deduplication logic 320, which eliminates redundant load operations. After load / store queue 308 receives data 322 from data cache 309 according to the first arriving load operation in queue 309, the same data 322 is used to complete the eliminated redundant load operations in queue 308 without repeatedly accessing cache 308. Deduplication logic 320 removes redundant operations from queue 308; thus, the number of data cache accesses is reduced.
[0033] In an alternative implementation, when a load operation is processed from queue 308, metadata 323 associated with the load address is saved. Deduplication logic 320 compares the compressed metadata retrieved for each load operation processed from queue 308 with the saved metadata value from a previous load operation. If the compressed metadata used for a subsequent load operation matches any previously saved metadata, this indicates that the data value used for loading was previously retrieved from cache 309, and repeated access to data cache 309 would be redundant. Therefore, the retrieved value is used to complete the subsequent load operation, rather than retrieving the same value from a different address in data cache 309.
[0034] In either case, the same data obtained from a single data cache access is copied to multiple registers in physical register file 305 to complete multiple redundant load operations, since the redundant load operations are for addresses containing the same data value. In an alternative implementation, the redundant data values do not need to be identical, but are sufficiently similar for the executing application. For such applications, data values that differ within a certain tolerance are considered redundant. Therefore, such implementations utilize compression schemes that generate compressed metadata indicating similarity. For example, when using a BDI compression scheme, small Δ values below a threshold can be considered as representing redundant values when referencing the same cardinality.
[0035] In one implementation, matching address 324 also includes a constant value address identified as storing a constant value, such as Figure 2A and Figure 2B The zero addresses 206 and 254 are described in the document. The deduplication logic 320 eliminates unnecessary cache accesses by causing the load / store queue 308 to use the detected constant values instead of retrieving these values from the data cache 309.
[0036] Figure 4 Components in an out-of-order execution pipeline of a processing unit 400 according to one embodiment are shown. This processing unit utilizes matching address information to reduce the number of redundant register writes and arithmetic operations performed. Processing unit 400 corresponds to processing unit 104 in computing system 100. In the illustrated configuration, processing unit 400 reduces the number of arithmetic operations performed in ALU 306 and / or write operations performed in physical register file 305 by identifying addresses containing duplicate data values based on associated compressed metadata values of the duplicate data values.
[0037] Compared to processing unit 300, the data cache 409 in processing unit 400 includes an additional port for receiving address 421 and provides metadata 323 as soon as address 421 is ready, instead of waiting to receive address 321 from load / store queue 308. Redundant address calculation unit 411 determines a matching address based on compressed metadata 323, as previously referenced. Figure 2A and Figure 2B As described. After determining a set of matching addresses, the redundancy address calculation unit 411 provides the matching addresses to the deduplication logic 420 in the allocation unit 403.
[0038] The matching address information includes duplicate addresses, which indicate to the deduplication logic 420 whether two or more load instructions in the pipeline are requesting data from addresses that hold the same value, making one or more load instructions redundant. The deduplication logic 420 eliminates redundant load operations by replacing them with a single load operation targeting a single destination register in the ROB / physical register file 305. Other instructions that depend on the eliminated load operation have their execution dependencies updated to depend on the remaining single load operation. The deduplication logic 420 also instructs the allocation unit 403 to update the RAT 304 so that the alias initially pointing to the register associated with the eliminated load operation is changed to refer to the single physical register loaded in the ROB / physical register file 305. Therefore, fewer load operations are performed and a single register is used to store the data value instead of multiple registers to store the same data value. This reduces the load on the physical register file 305 and cache 409, thereby improving application performance.
[0039] In one implementation, deduplication logic 420 further deduplicates instructions that become duplicates due to redundant load operations. Instructions that depend on the same deduplicated load operation as their operands are identified; any other operands in these instructions also depend on the deduplicated load operation or the same physical register. Duplicate instructions are deduplicated in a manner similar to that of load operations. All instructions except one duplicate instruction are eliminated, resulting in an output register being used to store the output of a single remaining instruction. Any dependent instructions are updated to depend on a single instruction instead of multiple duplicate instructions. This instruction deduplication further reduces pressure on the physical register file and on the operation execution logic (e.g., ALU 306), which has fewer instructions to execute.
[0040] In one implementation, compressed metadata is used to eliminate unnecessary operations by determining when data stored at a certain address matches a specific constant value. For example, for many instructions (such as multiplication or addition), when one operand is known to be zero, the result is a constant or a simpler function of the other operand. Deduplication logic 420 identifies these instructions based on information received from redundant address calculation unit 411 (e.g., zero address 206 or 254) and replaces them with simpler equivalent operations. The replacement operation is determined based on the original instruction and constant value associated with the appropriate replacement function in table 412. For example, the replacement function table indicates that if one of the operands of a multiplication instruction is 1 or 0, the multiplication instruction will be replaced with an identity function or a zero output, respectively. Table 412 is located in dedicated memory, or in another implementation, in physical register file 305, cache 409, or other memory device.
[0041] Operations that initially depended on any eliminated instruction update their dependencies to the corresponding equivalent substitution operation. For example, when performing an addition operation that adds zero to another operand, the result is the other operand. Therefore, the instruction is removed and the dependency on the removed addition operation is updated to depend instead on the instruction that produces the non-zero operand. Similarly, performing a logical AND with zero always results in an output of zero. Therefore, the logical AND operation is removed, and all dependencies on the removed logical AND are replaced with dependencies on the immediate zero.
[0042] In one implementation, components of processing units 300 and 400 are combined. For example, alternative implementations of the processing units include deduplication logic units 320 and 420 to support various mechanisms for deduplicating redundant operations in the pipeline.
[0043] Figure 5A process 500 for eliminating redundancy and constant value operations based on compressed metadata, according to one embodiment, is shown. Process 500 is performed by components of computing system 100 including load / store queue 308, data cache 309 or 409, compression engine 310, redundant address calculation unit 311 or 411, etc.
[0044] Process 500 operates in a processing unit (e.g., processing unit 300 or 400) to perform load and store operations between a data cache (e.g., 309 or 409) and physical register 305, while eliminating redundant operations (e.g., load and arithmetic instructions) and replacing constant value operations with their simpler equivalents. Load / store queue 308 contains a queue of load and / or store operations processed sequentially. At block 501, if the next operation in load / store queue 308 is a store operation, the data to be stored is compressed by compression engine 310 and stored in the data cache, as provided at block 502. For each data value, compression engine 310 generates associated compressed metadata representing that data value. For example, when using a C-Pack compression scheme, compression engine 310 generates compact code that references one or more dictionary entries or frequently used constant values. When using a BDI compression scheme, compression engine 310 generates a Δ value and an indication of whether Δ is relative to a zero-based or non-zero-based system. From box 502, process 500 returns to box 501 to process the next operation in load / store queue 308.
[0045] At box 501, if the next operation in the load / store queue 308 is a load operation, process 500 continues at box 503. At box 503, compression engine 310 reads the compressed metadata associated with the address specified by the load instruction from data cache 409. The compressed metadata 323 is provided to redundant address calculation unit 411. From box 503, process 500 continues at box 505.
[0046] At box 505, compressed metadata 323 is compared with compressed metadata of known constant values (such as 0 and 1). Depending on the compression scheme used, specific metadata values are generated when compressing such constant values.
[0047] For example, when using the C-Pack compression scheme, the redundant address calculation unit 411 maintains a list containing a predetermined subset of compressed codes and their associated constant values (e.g., constant value 0 is represented by the code "00"). The metadata 323 associated with a specified load address is compared with the code representing each constant value to determine whether the load address contains that constant value.
[0048] When using the BDI compression scheme, the redundant address calculation unit 411 maintains a pre-determined subset containing Δ and base values, along with a list of their associated constant values. The value of metadata 323 is compared with the known Δ and base values for each detectable constant value. For example, the constant 0 is encoded with a Δ value of 0 relative to base 0; therefore, the Δ value 0 and base 0 in metadata 323 indicate that the address contains a value of 0. If metadata 323 indicates that the address does not contain a constant value, process 500 continues at block 507.
[0049] At box 507, the compressed metadata 323 of the address is compared with compressed metadata associated with other addresses. In one embodiment, compressed metadata 323 is compared with compressed metadata obtained for a previous load operation. Alternatively, compressed metadata 323 is compared with compressed metadata of a load operation that is still in queue 308 and has not yet been executed. In some embodiments, depending on the compression scheme used, compressed metadata values are equal when the data values represented by the compressed metadata values are equal. In the C-Pack compression scheme, the same values are represented by the same compact code, and in some cases, by pointers to the same dictionary entries. In the BDI compression scheme, the same values are represented by the same Δ value and refer to the same base value. Therefore, box 507 determines whether compressed metadata 323 is equal to other known compressed metadata of other addresses that have been previously loaded or obtained for load operations still in queue 308. In one embodiment, as Figure 2A and Figure 2B As shown, the compressed metadata comparisons in boxes 505 and 507 are performed in parallel.
[0050] If a duplicate address or a constant (e.g., zero) address is not found, the data specified in the load operation is retrieved from data cache 309 or 409 and loaded into a register in physical register file 305. At box 509, data is read from cache 309 or 409 and decompressed by compression engine 310. The decompressed data value is loaded into the destination register specified by the load operation.
[0051] In one implementation, at box 511, compressed metadata 323 of the retrieved data value is stored along with the data value in case the compressed metadata can be reused later (e.g., at box 507) along with matching metadata in a future load operation. From box 511, process 500 returns to box 501 to continue processing the remaining operations in load / store queue 308. In an alternative implementation, process 500 instead continues from box 509 to box 501 without saving the metadata and the retrieved value.
[0052] At block 505, if the compressed metadata 323 retrieved for the load operation matches the constant value metadata, process 500 continues at block 513. At block 513, deduplication logic 420 identifies one or more instructions in the pipeline that depend on the data value of the load operation as an operand. For each of these instructions, deduplication logic 420 performs a lookup in substitution function table 412 to identify an equivalent substitution function. For example, for an addition instruction with two operands, where one operand is determined to be zero based on compressed metadata 323, the substitution function outputs the other operand. In substitution function table 412, the substitution function is associated with the original addition function and the zero constant value. At block 515, deduplication logic 420 replaces the original function with the equivalent function identified from substitution function table 412. Since the constant value has already been identified based on its metadata 323, the original load operation used to read the value from cache 409 is redundant and is eliminated.
[0053] At box 507, if the compressed metadata 323 retrieved for a load operation matches compressed metadata associated with different load operations, process 500 continues at box 517. At box 517, redundant address calculation unit 311 or 411 calculates one or more duplicate addresses based on which addresses are associated with compressed metadata matching compressed metadata 323. Duplicate addresses are addresses in the data cache that store the same value as the data cache address of the load operation, and are therefore redundant.
[0054] The compression metadata values for C-Pack and BDI compression schemes are as previously referenced. Figure 2A and Figure 2B The comparison is performed as described. In an alternative implementation, when the compressed metadata values of other compression schemes reliably indicate whether the original uncompressed values are equal, other compression schemes may be used instead of C-Pack or BDI.
[0055] From block 517, process 500 continues at block 519. In processing unit 300, deduplication logic 320 of load / store queue 308 receives a matching address 324 from redundant address calculation unit 311 and removes one or more redundant load operations from queue 308. In one embodiment, matching address 324 is provided by redundant address calculation unit 311 to deduplication logic 320. Deduplication logic 320 removes any queued load operations from load / store queue 308 that are redundant due to being directed to duplicate addresses. Instead of accessing the cache to perform the removed load operations, load / store queue 308 uses the data value read from the cache when performing the current load operation, thereby eliminating one or more redundant cache accesses.
[0056] In an alternative implementation, deduplication logic 320 eliminates cache accesses of the current load operation by identifying past load operations pointing to duplicate addresses. That is, the metadata values of the current load address and the previously loaded address indicate that these addresses store the same value. Therefore, the data previously retrieved to match the previous load operation is used to serve the current load operation.
[0057] Alternatively, in processing unit 400, deduplication logic 420 reduces the number of data cache accesses by replacing multiple loads of the same data value with a single load of the data value into a single register. Allocation unit 403 updates RAT 304 such that multiple register aliases already associated with the removed load operation are remapped to the single loaded register. From block 519, process 500 continues at block 521.
[0058] At box 521, deduplication logic 320 or 420 updates the dependencies of any operations that initially depended on the operations eliminated at boxes 515 or 519. When the original operation is replaced by an equivalent operation, the dependencies of the original operation are updated to depend on the equivalent operation. For example, for an addition operation where one operand is determined to be zero, the addition operation is replaced by a simpler equivalent operation that outputs the other operand. Thus, any operation that initially depended on the result of the addition operation is updated to depend on the other operand instead.
[0059] When multiple load operations are redundant due to loading the same data value, and are therefore eliminated at box 519, any operations that depend on the destination register of the eliminated load operation are updated so that they instead depend on the destination register loaded by the remaining load operations.
[0060] Elimination operations can cause other operations to become redundant; therefore, at box 523, if additional operations (e.g., instructions) become redundant through previous elimination, process 500 continues at box 525. At box 525, additional redundant operations are eliminated in a manner similar to that described in boxes 513-515 and 519. Any dependencies on these eliminated operations are updated at box 521. Boxes 521-525 are repeated to eliminate redundant operations until no operations are redundant at box 523. From box 523, process 500 continues at box 509 and performs the loading operation.
[0061] A method includes: in response to receiving an instruction to perform a first operation on first data stored in a memory device, obtaining first compressed metadata from the memory device based on an address of the first data; and reducing the number of operations in a set of operations based on the first operation and one or more matching addresses, the one or more matching addresses corresponding to second metadata that matches the first compressed metadata.
[0062] In this method, the first compression metadata includes a first compression code, and the second compression metadata includes a second compression code equal to the first compression code. In response to identifying the first compression code as a compression code within a predetermined subset of compression codes in a set of compression codes used for compressing data in a storage device, a reduction operation is performed.
[0063] In this method, the first compressed metadata includes a first Δ value representing the difference between the first data and the base value, the second compressed metadata includes a second Δ value, and a reduction operation is performed in response to determining that the first Δ value is equal to the second Δ value and that the second Δ value is associated with the same base value as the first Δ value.
[0064] In this method, the memory device is a data cache device. The method also includes generating first compressed metadata when compressing first data in the data cache device, and generating second compressed metadata when compressing second data in the data cache device.
[0065] In this method, the first operation includes a load operation obtained from the load / store queue of the data cache device, and reducing the number of operations in this group of operations includes eliminating one or more queued accesses of the data cache device from the load / store queue.
[0066] The method further includes, in response to determining that the second compressed metadata matches the first compressed metadata, copying the first data to a register associated with each of the one or more matching addresses for each duplicate address.
[0067] The method also includes loading first data into a physical register, and in response to determining that second compressed metadata matches the first compressed metadata, updating the dependencies associated with the one or more matching addresses to depend on the loading of the physical register.
[0068] The method further includes: identifying a constant value corresponding to the second compressed metadata; in response to identifying the constant value, replacing the first operation with an equivalent operation based on the first operation and the identified constant value; and updating operation dependencies that depend on the first operation to depend on the equivalent operation.
[0069] A computing device includes compression logic that, in response to receiving an instruction to perform a first operation on first data stored in a memory device, obtains first compressed metadata from the memory device based on an address of the first data. The computing device also includes address calculation logic coupled to the compression logic, which, in response to identifying second compressed metadata matching the first compressed metadata, determines one or more matching addresses corresponding to the second compressed metadata. The computing device further includes deduplication logic coupled to the address calculation logic, which reduces the number of operations in a set of operations based on the first operation and one or more matching addresses.
[0070] In the computing device, the first compressed metadata includes a first compressed code, and the second compressed metadata includes a second compressed code equal to the first compressed code. The deduplication logic also reduces the number of operations by recognizing the first compressed code as a compressed code within a predetermined subset of compressed codes in a set of compressed codes used to compress data in the memory device, in response to address calculation logic.
[0071] In this computing device, the first compressed metadata includes a first Δ value representing the difference between the first data and a base value, and the second compressed metadata includes a second Δ value. The deduplication logic also reduces the number of operations in response to the address calculation logic determining that the first Δ value is equal to the second Δ value and that the second Δ value is associated with the same base value as the first Δ value.
[0072] In a computing device, the memory device includes a data cache device. The compression logic also generates first compressed metadata when compressing first data in the data cache device, and generates second compressed metadata when compressing second data in the data cache device.
[0073] In the computing device, the first operation includes a load operation obtained from the load / store queue of the data cache device, and the deduplication logic further reduces the number of operations in the group of operations by eliminating one or more queued accesses of the data cache device from the load / store queue.
[0074] In a computing device, for each duplicate address in one or more matching addresses, the deduplication logic, in response to determining that the second compressed metadata matches the first compressed metadata, copies the first data to the register associated with the duplicate address.
[0075] The computing device also includes a load / store queue coupled to address calculation logic, which loads data into physical registers. Deduplication logic, in response to the address calculation logic identifying second compressed metadata that matches the first compressed metadata, updates the operational dependencies associated with one or more matching addresses to be load-dependent on physical registers.
[0076] In the computing device, the redundant address calculation logic identifies a constant value corresponding to the second compressed metadata. The deduplication logic also, in response to identifying the constant value, replaces the first operation with an equivalent operation based on the first operation and the constant value, and updates operation dependencies that depend on the first operation to dependencies that depend on the equivalent operation.
[0077] The computing system includes a memory device and a compression engine coupled to the memory device. The compression engine generates first compressed metadata for compressing first data to be stored in the memory device, and generates second compressed metadata for compressing second data to be stored in the memory device. The computing system also includes redundant address calculation logic coupled to the memory device. In response to receiving an instruction to perform a first operation on the first data stored in the memory device, the redundant address calculation logic obtains the first compressed metadata based on the address of the first data, and in response to determining that the second compressed metadata matches the first compressed metadata, calculates one or more matching addresses for the second data based on the second compressed metadata. The computing system includes deduplication logic coupled to the redundant address calculation logic, which reduces the number of operations in a set of operations based on the first operation and the one or more matching addresses.
[0078] In a computing system, the memory device includes a cache memory device, and the first compressed metadata and the second compressed metadata are stored in the cache memory device.
[0079] The computing system also includes a physical register file coupled to deduplication logic. The deduplication logic reduces the number of operations by reducing the number of register write operations in the physical register file.
[0080] The computing system also includes an arithmetic logic unit coupled with deduplication logic. The deduplication logic reduces the number of operations by reducing the number of computations performed in the arithmetic logic unit.
[0081] As used herein, the term "coupled to" can mean direct coupling or indirect coupling via one or more intermediary components. Any signal provided on the various buses described herein may be time-division multiplexed with other signals and provided on one or more common buses. Additionally, interconnections between circuit components or blocks may be shown as buses or single-signal lines. Each bus in a bus may alternatively be one or more single-signal lines, and each single-signal line in a single-signal line may alternatively be a bus.
[0082] Some implementations may be implemented as a computer program product, which may include instructions stored on a non-transitory computer-readable medium. These instructions may be used to program a general-purpose or special-purpose processor to perform the described operations. Computer-readable media include any mechanism for storing or transmitting information in a machine-readable form (e.g., software, processing application). Non-transitory computer-readable storage media may include, but are not limited to: magnetic storage media (e.g., floppy disks); optical storage media (e.g., CD-ROMs); magneto-optical storage media; read-only memory (ROM); random access memory (RAM); erasable programmable memory (e.g., EPROM and EEPROM); flash memory; or another type of medium suitable for storing electronic instructions.
[0083] Additionally, some implementations can be practiced in a distributed computing environment where computer-readable media are stored on and / or executed by more than one computer system. Furthermore, information transferred between computer systems can be pushed or pulled across the transmission medium connecting the computer systems.
[0084] Typically, the data structures representing computing system 100 and / or portions thereof, carried on a computer-readable storage medium, can be databases or other data structures that can be read by a program and used directly or indirectly for manufacturing hardware including computing system 100. For example, the data structure can be a behavioral-level description or register-transfer-level (RTL) description of hardware functionality in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool, which can synthesize the description to produce a netlist including a list of gates from a synthesis library. The netlist includes gate sets, which also represent the functionality of the hardware including computing system 100. The netlist can then be placed and routed to produce a dataset describing the geometry to be applied to a mask. The mask can then be used in various semiconductor manufacturing steps to produce semiconductor circuits or circuits corresponding to computing system 100. Alternatively, the database on the computer-readable storage medium can be a netlist (with or without a synthesis library), a dataset (as needed), or Graphical Data System (GDS) II data.
[0085] Although the operations of the methods herein are shown and described in a specific order, the order of operations of each method may be changed so that some operations can be performed in reverse order, or that some operations can be performed at least partially concurrently with other operations. In another embodiment, instructions or sub-operations of different operations may be performed intermittently and / or alternately.
[0086] In the foregoing specification, embodiments have been described with reference to specific exemplary embodiments. However, it will be apparent that various modifications and changes can be made to these embodiments without departing from the broader scope of the embodiments set forth in the appended claims. Therefore, this specification and the accompanying drawings should be regarded as illustrative rather than restrictive.
Claims
1. A method comprising: In response to receiving an instruction to load first data into a physical register, first compressed metadata is obtained from the memory device based on the address of the first data stored in the memory device, the first compressed metadata including a compressed representation of the first data stored at the address; The number of operations in the load / store queue is reduced based on the instruction to load the first data into the physical register and one or more addresses in a set of operations, the one or more addresses corresponding to second compressed metadata, wherein the second compressed metadata includes a compressed representation of the second data, and the second compressed metadata matches the first compressed metadata including the compressed representation of the first data within a threshold. as well as The one or more addresses are mapped to the physical registers based on the register alias table associated with the physical registers.
2. The method according to claim 1, wherein: The memory device is a data cache device; and The method further includes: The first compressed metadata is generated when the first data in the data cache device is compressed, and The second compressed metadata is generated when the second data in the data cache device is compressed.
3. The method according to claim 2, wherein: The instruction to load the first data into the physical register is obtained from the load / store queue of the data cache device; and Reducing the number of operations in the set of operations includes eliminating one or more queued accesses to the data cache device from the load / store queue.
4. The method according to claim 2, further comprising: In response to determining that the second compressed metadata matches the first compressed metadata within the threshold, for each of the one or more addresses, the first data is copied to the register associated with each address.
5. The method according to claim 1, further comprising: Load the first data into the physical register; as well as In response to determining that the second compressed metadata matches the first compressed metadata within the threshold, the dependency associated with the one or more addresses is updated to depend on loading the first data into the physical register.
6. The method according to claim 1, further comprising: Identify the constant value corresponding to the second compressed metadata; In response to identifying the constant value, the operation of the group operation in the load / store queue is replaced with an equivalent operation based on the operation and the identified constant value, wherein the operation includes calculation; and Update the operation dependencies that depend on the operation to depend on the equivalent operation.
7. A computing device, comprising: Compression logic is configured to, in response to receiving an instruction to load first data into a physical register, obtain first compressed metadata from the memory device based on the address of the first data stored in the memory device, the first compressed metadata including a compressed representation of the first data stored at the address; Address calculation logic, coupled to the compression logic and configured to determine one or more addresses corresponding to the second compressed metadata in response to identifying the second compressed metadata, wherein the second compressed metadata includes a compressed representation of second data, and the second compressed metadata matches, within a threshold, the first compressed metadata including the compressed representation of the first data; and Deduplication logic, which is coupled to the address calculation logic and configured as follows: The number of operations in the load / store queue is reduced based on the instruction to load the first data into the physical register and one or more addresses in a set of operations; and The one or more addresses are mapped to the physical registers based on the register alias table associated with the physical registers.
8. The computing device according to claim 7, wherein: The first compression metadata includes a first compression code, and the second compression metadata includes a second compression code that is equal to the first compression code; and The deduplication logic is further configured to reduce the number of operations in response to the address calculation logic identifying the first compression code as a compression code in a predetermined subset of compression codes in a set of compression codes for compressing data in the memory device.
9. The computing device according to claim 7, wherein: The first compressed metadata includes a first Δ value representing the difference between the first data and the base value; The second compressed metadata includes a second Δ value; and The deduplication logic is further configured to reduce the number of operations in response to the address calculation logic determining that the first Δ value is equal to the second Δ value and that the second Δ value is associated with the same base value as the first Δ value.
10. The computing device according to claim 7, wherein: The memory device includes a data cache device; and The compression logic is further configured as follows: The first compressed metadata is generated when the first data in the data cache device is compressed, and The second compressed metadata is generated when the second data in the data cache device is compressed.
11. The computing device according to claim 10, wherein: The instruction to load the first data into the physical register is obtained from the load / store queue of the data cache device; and The deduplication logic is further configured to reduce the number of operations in the set of operations by eliminating one or more queued accesses to the data cache device from the load / store queue.
12. The computing device according to claim 7, wherein, The deduplication logic is further configured as follows: For each of the one or more addresses, in response to determining that the second compressed metadata matches the first compressed metadata within the threshold, the first data is copied to the register associated with each address.
13. The computing device according to claim 7, wherein: The load / store queue is coupled to the address calculation logic and configured to load the first data into the physical register, wherein the deduplication logic is configured to: In response to the address calculation logic identifying the second compressed metadata that matches the first compressed metadata within the threshold, the operational dependencies associated with the one or more addresses are updated to depend on the load of the physical register; as well as The number of operations of the group operation in the load / store queue is further reduced based on one or more operation dependencies of the update operation dependencies associated with one or more addresses of the group operation.
14. The computing device according to claim 7, wherein: The address calculation logic is further configured to identify a constant value corresponding to the second compressed metadata; and The deduplication logic is further configured as follows: In response to identifying the constant value, the operation of the group operation in the load / store queue is replaced with an equivalent operation based on the operation and the constant value, and Update the operation dependencies that depend on the operation to depend on the equivalent operation.
15. The computing device according to claim 7, wherein, Identifying that the second compressed metadata matches the first compressed metadata within the threshold includes identifying that the second compressed metadata is equal to the first compressed metadata.
16. The computing device according to claim 7, wherein, To reduce the number of operations in the group operation, the deduplication logic is configured to remove the corresponding load operation of the group operation from the load / store queue for each of the one or more addresses.
17. A computing system, comprising: Memory devices; A compression engine, coupled to the memory device and configured to: Generate first compressed metadata for compressing first data for storage in the memory device, the first compressed metadata including a compressed representation of the first data stored at an address, and Generate second compressed metadata for compressing second data to store in the memory device, the second compressed metadata including a compressed representation of the second data, wherein the second compressed metadata matches the first compressed metadata including the compressed representation of the first data within a threshold; Address calculation logic, which is coupled to the memory device and configured as follows: In response to receiving an instruction to load the first data into a physical register, the first compressed metadata is obtained based on the address of the first data, and In response to determining that the second compressed metadata matches the first compressed metadata within the threshold, one or more addresses of the second data are calculated based on the second compressed metadata; and Deduplication logic, which is coupled to the address calculation logic and configured as follows: The number of operations in the load / store queue is reduced based on the instruction to load the first data into the physical register and one or more addresses in a set of operations; and The one or more addresses are mapped to the physical registers based on the register alias table associated with the physical registers.
18. The computing system according to claim 17, wherein: The memory device includes a cache memory device; and The first compressed metadata and the second compressed metadata are stored in the cache memory device.
19. The computing system of claim 17, further comprising a physical register file coupled to the deduplication logic, wherein the deduplication logic is configured to reduce the number of operations by reducing the number of register write operations in the physical register file.
20. The computing system of claim 17, further comprising an arithmetic logic unit coupled to the deduplication logic, wherein the deduplication logic is configured to reduce the number of operations by reducing the number of computations performed in the arithmetic logic unit.
Citation Information
Patent Citations
Scatter using index array and finite state machine
CN104303142A
Methods and apparatus for multi-load and multi-store vector instructions
CN109992240A