A Parallel Computing Core Caching Method, Device and Medium Based on Multiple Compression Schemes

By adopting a multi-compression scheme and a multi-group association mechanism in the parallel computing core cache, the problems of high concurrent memory access requests and low storage resource utilization are solved, and more efficient cache performance and resource utilization are achieved.

CN119884023BActive Publication Date: 2025-06-10SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510386985.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-10
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing parallel computing core cache scheme is difficult to handle high concurrent memory access requests, and the on-chip storage resource utilization rate is low.

Method used

The parallel computing core cache method based on multiple compression schemes is adopted. By receiving and adapting concurrent memory access requests, combining and requesting, and using a multi-channel group-associated cache construction mechanism in the cache, multiple data compression schemes are stored to optimize compression effects and resource utilization.

Benefits of technology

It improves the adaptability and overall performance of parallel computing core cache, can handle high concurrent memory access requests more effectively, and improves the utilization rate of on-chip storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884023B_ABST
    Figure CN119884023B_ABST
Patent Text Reader

Abstract

This application relates to the field of parallel computing core cache technology, and discloses a parallel computing core cache method, device and medium based on multiple compression schemes. The present invention receives high-concurrency memory access signals of parallel computing cores, performs adaptive arbitration according to cache internal capacity performance parameters to determine the memory access request sequence; sets multiple data compression schemes, determines the characteristics of memory access request data by analyzing memory access data, and optimizes the compression effect according to multiple data compression schemes to determine whether to perform data compression on the memory access data and the compression scheme to be adopted; fully associates the compressed data with the multi-way set-associative mechanism in the cache system, each way of the cache set-associative module stores the compression result of one compression scheme, and data exchange can be performed between multiple set-associative modules to ensure that as the data is continuously refreshed, the most suitable data compression scheme is dynamically used for the data characteristics, so as to improve the utilization efficiency of on-chip cache resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of parallel computing core cache technology, and in particular to a parallel computing core cache based on a multi-compression scheme Background Art

[0002] In modern high-performance computing scenarios, the application of parallel computing cores is becoming more and more extensive, and its requirements for data storage and processing capabilities are becoming increasingly stringent. The traditional parallel computing core cache architecture has gradually exposed many drawbacks when dealing with the large-scale, high-concurrency data reading and writing needs of parallel computing cores. From the perspective of concurrent processing mechanism, given the high concurrency characteristics of parallel computing cores, the cache architecture often needs to synchronously respond to memory read and write requests from multiple threads. However, most of the current parallel computing core cache architectures adopt a processing strategy that decomposes concurrent access into serial access. Although this solution has the advantage of being simple and fast in the implementation process, the negative effect it causes is that it significantly increases the memory access latency of the parallel computing core, which becomes a key performance bottleneck in computing tasks with extremely high performance requirements. In terms of on-chip resource allocation, due to the inherent limitations of on-chip resources, in the resource allocation process of parallel computing cores, computing resources often have a higher tendency to be allocated in priority than storage resources. In this case, how to maximize the storage efficiency of the parallel computing core architecture under the constraints of limited on-chip storage resources has also become a key problem that needs to be overcome.

[0003] At present, cache compression technology is widely studied as an effective means to improve cache performance. However, existing cache compression technologies usually use a single compression scheme to process the entire cache, which cannot fully utilize the characteristics of different data, resulting in limited compression effect and performance improvement. For example, different application data has different characteristics. Some data is suitable for a certain compression method, while other data may be more suitable for other compression methods. Therefore, it is necessary to develop a parallel computing core cache system implementation based on multiple compression schemes to improve the adaptability and overall performance of the parallel computing core cache.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0006] In view of the deficiencies of the prior art, the present invention provides a parallel computing core cache method, device and medium based on a multi-compression scheme, which are used to solve the problems that existing parallel computing core cache schemes are difficult to handle high-concurrency memory access requests and have low utilization rate of on-chip storage resources.

[0007] In some embodiments, the parallel computing core cache method based on a multi-compression scheme includes:

[0008] Step A: Receive n-way concurrent memory access requests from a parallel computing core;

[0009] Step B: Perform adaptive arbitration according to the internal capacity performance parameters of the cache to determine the memory access request sequence;

[0010] First, according to the depth and bit width of the internal storage unit of the read cache, determine the word_sel size in the memory access request according to the bit width of the internal storage unit of the cache, determine the index size in the memory access request according to the depth of the internal storage unit of the cache, and determine the tag size in the memory access request according to the processor bit width; then read the index data in the n-way memory access requests sent from the parallel computing core to the cache, and group and merge the n-way memory access requests according to the index data. The grouping and merging method is to merge the memory access requests with the same index into one group, and the indexes between groups are arranged in order. Through grouping and merging, m groups of n-way concurrent memory access requests are formed; both m and n are positive integers, and m ≤ n;

[0011] Step C: Execute the memory access requests adapted to the multi-cache data compression scheme. The m groups of n-way concurrent memory access requests formed in Step B are sent to the cache. The cache adopts a multi-way set-associative cache construction mechanism. Each way of the cache set-associative module stores a data compression scheme. Determine whether the memory access request hits or misses through tag comparison. If the request hits, trigger the request hit execution mechanism; if the request misses, trigger the request miss execution mechanism; during the execution of the request hit execution mechanism and the request miss execution mechanism, if it is a data read, execute the data compression method stored in the request hit cache set-associative module. If it is a data update, select a compression algorithm according to the input data characteristics, perform data compression, and associate with the storage set-associative module.

[0012] As a further improvement, in Step B, the size determination methods of word_sel, index, and tag are as follows:

[0013] ,

[0014] where 、 、 respectively represent the sizes of word_sel, index, and tag, Width, Depth, and SIZE corerespectively represent the bit width of the cache internal storage unit, the depth of the cache internal storage unit, and the processor bit width.

[0015] As a further improvement, when grouping in step B, if the combined quantity of memory access requests with the same index does not reach n-way, fill in invalid signals to make each group of memory access requests reach n-way. The m groups of n-way concurrent memory access requests formed by grouping all include a TAG signal bit, an Index signal bit, a WORD_SEL signal bit, and a flag bit. The TAG signal bit and the WORD_SEL signal bit directly fill in the TAG signal and the WORD_SEL signal in the n-way concurrent memory access requests sent by the parallel computing core. The Index signal bit fills in the Index signal after grouping, and the memory access request storing the Index signal after grouping has the flag bit set to valid, while the memory access request filled with the invalid signal has the flag bit set to invalid.

[0016] As a further improvement, in step C, a request hit includes a read data request hit, and a request miss includes a read data request miss;

[0017] When a read data request misses, the request miss execution mechanism is as follows: If there is one or more read data requests in the n-way memory access requests sent that do not hit, then trigger the read data request miss execution mechanism, extract the unhit memory access requests row by row from the n-way concurrent memory access requests being executed and send them to the next-level cache; after receiving the request signal from this-level cache, the next-level cache makes a memory access request hit determination. If the next-level cache request hits, it feeds back a response signal to this-level cache. The response signal includes a response address signal and a response data signal. The response data signal first enters the compression / decompression algorithm way determination module for input data feature analysis, completes the selection of the compression algorithm, data compression execution, and group associative module selection according to the input data features, then splits the response address signal into tag, index, and word_sel signals, determines the corresponding cache line in the group associative module according to the index signal, writes the compressed response data into the cache data corresponding cache line of the corresponding group associative module, and writes the split tag data into the cache tag corresponding cache line of the corresponding group associative module;

[0018] When a read data request hits, the request hit execution mechanism is as follows: If there is one or more read data requests in the n-way memory access requests sent that hit, then trigger the read data request hit execution mechanism, obtain the response data from the cache data corresponding cache line of the corresponding group associative module according to the hit group associative module and index data, and send the obtained response data and the group associative module to the compression / decompression algorithm way determination module for processing. The compression / decompression algorithm way determination module determines the data compression algorithm to be executed according to the received group associative module, then completes the reverse decompression operation of the data according to this data compression algorithm, and feeds back the decompressed data to the parallel computing core module.

[0019] As a further improvement, the request hit includes a write data request hit, and the request miss includes a write data request miss;

[0020] When a write data request hits, the request hit execution mechanism is as follows: if one or more write data request hits exist in the valid signals of the n-way memory access request sent, the write data request hit execution mechanism is triggered, and storage data is obtained from the cache line corresponding to the group associative module corresponding to the cache address according to the hit group associative module and index data, and the storage data is sent to the compression and decompression algorithm way judgment module for decompression processing to obtain decompressed storage data, and then the decompressed storage data is combined with the write data sent by the parallel computing core to obtain complete write data, and the complete write data is resent to the compression and decompression algorithm way judgment module for data feature The compression and decompression algorithm way judgment module re-judges the data compression scheme according to the characteristics of the complete written data. If the judgment result is the original data compression scheme, the compression and decompression algorithm way judgment module completes the compression of the complete written data and continues to store the compressed data in the index cache line corresponding to the original group associative module; if the judgment result is a new data compression scheme, the compression and decompression algorithm way judgment module completes the data compression according to the new data compression scheme and writes the compressed data into the index cache line corresponding to the new group associative module, and invalidates the data of the index cache line corresponding to the original group associative module on the other hand;

[0021] When a write data request is missing, the request missing execution mechanism is: if there is a valid signal in the n-way memory access request sent and one or more write data requests are not hit, the write data request missing execution mechanism is triggered. This mechanism does not trigger the data cache action in the current level cache and the compression and decompression algorithm way judgment module action, but directly extracts the missed memory access requests line by line from the concurrently executed n-way memory access requests and sends them to the next level cache.

[0022] As a further improvement, four compression schemes are stored in the compression and decompression algorithm way judgment module as the default data compression scheme. When input data enters, the compression and decompression algorithm way judgment module first performs the default data compression scheme matching work, and selects the corresponding compression scheme according to the data characteristics and compression scheme priority. Each data compression scheme is configured with a group associative storage module associated with it. If a certain data compression algorithm is matched, the compressed data will be stored in the corresponding group associative storage module.

[0023] As a further improvement, the four compression schemes stored in the compression and decompression algorithm way judgment module are: BDI(2,0), BDI(4,0), BDI(4,2) in the BDI compression algorithm, and DIC(8) in the DIC compression algorithm. The default data compression scheme fitting the work process based on these four compression schemes is as follows: The entire cache line data is sliced by 2 bytes to form distributed data of (k + 1) * 2 bytes. Taking the first data 0 as the reference data, the subsequent k data are subtracted from the reference data to obtain difference data 1 to k. Compare whether the difference data 1 to k are all 0. If they are all 0, it conforms to the BDI(2,0) compression algorithm rule, and the 2-byte reference data 0 is stored in the set-associative storage module 0. If the difference data are not all 0, the data slicing is redone; the redone data slicing is based on 4 bytes, and similarly forms distributed data of (k / 2 + 1) * 4 bytes. Again, taking the first data 0 as the reference data, the subsequent k / 2 data are subtracted from the reference data to obtain difference data 1 to k / 2. Compare whether the difference data 1 to k / 2 are all 0. If they are all 0, it conforms to the BDI(4,0) compression algorithm rule, and the 4-byte reference data 0 is stored in the set-associative storage module 1; if the differences are not all 0 but the difference bit width is within 2 bytes, it conforms to the BDI(4,2) compression algorithm rule, and the 4-byte reference data 0 and the k-byte difference data are stored in the set-associative storage module 2; if the difference bit width is greater than 2 bytes, it does not conform to the BDI compression algorithm rule, and the DIC compression algorithm is judged. The judgment process is as follows: The data is sliced by 4 bytes to form distributed data of (k / 2 + 1) * 4 bytes. The distributed data [0, k / 2] is sequentially stored in the data storage dictionary. The data storage dictionary is composed of a content-addressable memory. Only one dictionary index is fed back for the storage of the same data. While the dictionary index is continuously generated, the dictionary depth is judged. When all the dictionary indexes of the distributed data are generated and the depth does not exceed the dictionary depth, it indicates that the repetition rate of the distributed data is good, and the DIC data compression algorithm is adapted. The dictionary index is stored as data in the set-associative storage module 3. If the dictionary index depth generated by the distributed data exceeds the dictionary depth, it indicates that the repetition rate of the distributed data is poor, and the DIC data compression algorithm cannot be used. When the fitting work results of both the BDI compression algorithm and the DIC compression algorithm are fitting, the BDI compression algorithm is preferentially used to compress the data; when the fitting work result of the BDI compression algorithm is not fitting and the fitting work result of the DIC compression algorithm is fitting, the DIC compression algorithm is used to compress the data.

[0024] As a further improvement, the compression and decompression algorithm way judgment module reserves an interface for an extensible compression scheme. When the adaptation results of the default data compression scheme do not fit, the adaptation work of the extensible data compression scheme is started. If there is no extensible data compression scheme or the adaptation result of the extensible data compression scheme is also not suitable, the data is not compressed. The reserved extensible data compression scheme is also associated with a corresponding set-associative storage module. When there is no extensible data compression scheme, the data stored in the corresponding set-associative storage module stores uncompressed cache data.

[0025] In some embodiments, the device includes a processor and a memory storing program instructions. The processor is configured to execute the foregoing parallel computing core cache method based on multiple compression schemes when running the program instructions.

[0026] In some embodiments, the storage medium stores program instructions. When the program instructions are running, they execute the foregoing parallel computing core cache method based on multiple compression schemes.

[0027] The parallel computing core cache method, device, and medium provided by the embodiments of the present disclosure can achieve the following technical effects: The solution of the present invention addresses the problems of the traditional parallel computing core cache scheme, such as difficulty in handling high-concurrency memory access requests and low utilization rate of on-chip storage resources, and proposes a parallel computing core cache system and method based on multiple compression schemes. First, the upper-layer interface of the present invention fully accommodates the high-concurrency memory access signals of the parallel computing core and realizes adaptive arbitration according to the internal capacity performance parameters of the parallel computing core cache to determine the memory access request sequence. Secondly, the present invention fully accommodates multiple data compression schemes. By receiving and analyzing the memory access data in the memory access request sequence, it determines the characteristics of the memory access request data, and based on the multiple accommodated data compression schemes to optimize the compression effect, it determines whether to perform data compression on the memory access data and the compression scheme to be used when performing data compression. Finally, the present invention fully associates the data compressed by multiple compression schemes with the multi-way set-associative mechanism in the classical cache system. Each cache set-associative module stores the compression result of one compression scheme, and sufficient data exchange can be performed between multiple set-associative modules, ensuring that the cache system solution of the present invention can dynamically use the most suitable data compression scheme according to the data characteristics as the data is continuously refreshed, so as to improve the utilization efficiency of on-chip cache resources.

[0028] The above general description and the following description are only exemplary and explanatory and are not used to limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] One or more embodiments are illustrated by way of example in the corresponding drawings, which do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and wherein:

[0030] Figure 1 It is a schematic diagram of the overall architecture of the parallel computing core cache system;

[0031] Figure 2 It is a flowchart of the method described in Embodiment 1;

[0032] Figure 3 It is a flowchart of the memory access request adaptability arbitration strategy;

[0033] Figure 4 It is a flowchart of the memory access request execution mechanism adapted to the multi-cache data compression scheme;

[0034] Figure 5 It is a schematic diagram of the internal architecture of the compression and decompression algorithm way judgment module;

[0035] Figure 6 It is a schematic diagram of the BDI data compression algorithm adapted compression process;

[0036] Figure 7 It is a schematic diagram of the DIC data compression algorithm adapted compression process;

[0037] Figure 8 It is a principle block diagram of the device described in Embodiment 2. Detailed implementation manners

[0038] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to give a thorough understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other instances, well-known structures and devices may be shown in a simplified manner to simplify the drawings.

[0039] The terms "first", "second", etc. in the embodiments of the present disclosure are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.

[0040] Unless otherwise specified, the term "plurality" means two or more.

[0041] In the embodiments of the present disclosure, the character " / " indicates an "or" relationship between the preceding and following objects. For example, A / B means: A or B.

[0042] The term "and / or" is an associative relationship describing an object, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B, these three relationships.

[0043] The term "corresponding" can refer to an associative relationship or a binding relationship. A corresponding to B means that there is an associative relationship or a binding relationship between A and B.

[0044] Embodiment 1

[0045] This embodiment discloses a parallel computing core cache method based on multiple compression schemes. The overall architecture of the parallel computing core cache system is as Figure 1 shown. Among them, the parallel computing core is the upper layer of the computing in the cache architecture of the present invention, responsible for completing specific computing and storage instructions. To ensure upward compatibility of the present invention's solution with the original parallel computing core, the present invention does not change the external interface of the original parallel computing core and fully accepts high-concurrency memory access requests from the parallel computing core.

[0046] The cache architecture of the present invention's solution as a whole adopts a Harvard storage architecture. The L1 cache is divided into an instruction cache and a data cache. The instruction cache is specifically used to process instruction data, and the overall cache structure is read-only and not writable; the data cache is specifically used to process computing data, and the overall cache structure is readable and writable. There is no direct path between the instruction cache and the data cache, but data exchange is performed through the L2 cache. At the same time, in order to further improve the performance of multiple computing cores in the parallel computing core and make full use of the temporal locality and spatial locality characteristics of the cache data, the cache structure of the present invention allocates private instruction caches and data caches for each parallel computing core. The private L1 caches of multiple computing cores also cannot directly perform data exchange and can only transfer data through the L2 cache.

[0047] The cache architecture of the present invention's solution adopts a write-through - non - memory - allocation cache mechanism as a whole. Taking the parallel computing core 1 as an example, the memory access process is specifically described as follows: First, when the program starts running, the parallel computing core 1 sends a fetch instruction signal to the private instruction cache according to the program PC. Since the instruction cache is empty at this time, a cache miss occurs in the private instruction cache of the parallel computing core 1. Therefore, the private instruction cache sends a fetch instruction signal to the L2 cache. After receiving the fetch instruction signal sent from the computing core 1, the L2 cache performs data memory access. Assuming that the data exists in the L2 cache, the L2 cache sends a data response signal to the private instruction cache of the computing core 1. After receiving the data response sent by the L2, the private cache of the computing core 1 stores it in its own cache line and sends a data response signal to the computing core 1, which is the instruction information requested earliest by the computing core. After receiving the instruction information, the computing core 1 enters the decoding - execution and other logics. Assuming that this instruction is also a memory access instruction and it reads data from a certain address, the computing core 1 sends a data request to the private data cache, and the overall execution process is the same as that of reading data from the private instruction cache; assuming that this instruction is a memory access instruction and it writes data to a certain address, the computing core 1 sends a write data request to the private data cache. After receiving the data write request, the corresponding private data cache of the computing core 1 performs a cache hit judgment internally: If the cache hits, the data is written into the corresponding private data cache of the computing core 1, and at the same time, a write data request is sent to the L2 cache. If the cache misses, only a write data request is sent to the L2 cache.

[0048] For the convenience of improving the engineering portability, the present invention's solution integrates various data compression schemes synchronously inside the private L1 cache and L2 cache of each parallel computing core. It does not change the overall computing unit and memory transfer bus of the parallel computing core, but improves data compression and cache performance by changing the internal components of the cache.

[0049] Combined Figure 2 As shown, the embodiments of the present disclosure provide a cache method for a parallel computing core based on multiple compression schemes, including:

[0050] Step A: Receive n - way concurrent memory access requests from the parallel computing core;

[0051] Step B: Perform adaptive arbitration according to the internal capacity and performance parameters of the cache to determine the memory access request sequence;

[0052] Step C: Execute the memory access requests adapted to the multi - cache data compression scheme.

[0053] The mechanism of performing adaptive arbitration according to the internal capacity and performance parameters of the cache in Step B to determine the memory access request sequence is as Figure 3Shown as follows: First, read the depth and bit width of the cache internal storage unit. Determine the word_sel size in the memory access request according to the storage bit width, determine the index size in the memory access request according to the storage depth, and determine the tag size in the memory access request according to the processor bit width. The specific formulas are as follows, where Width, Depth, and SIZE core represent the cache bit width, cache depth, and processor bit width respectively:

[0054] ,

[0055] where , , represent the sizes of word_sel, index, and tag respectively.

[0056] Subsequently, send the n-way memory access requests from the parallel computing core to the cache. Read the index data according to the calibrated index position, and group and merge the n-way memory access requests sent by the parallel computing core according to the index data. The specific grouping and merging scheme is to merge the memory access requests with the same index into one group, and arrange the indexes in order between groups. If the number of memory access requests with the same index does not reach n-way, fill in invalid signals to make each group of memory access requests reach n-way. The m groups of n-way concurrent memory access requests formed by grouping all include a TAG signal bit, an Index signal bit, a WORD_SEL signal bit, and an identification bit. The TAG signal bit and the WORD_SEL signal bit are directly filled with the TAG signal and the WORD_SEL signal in the n-way concurrent memory access requests sent by the parallel computing core, and the Index signal bit is filled with the Index signal after grouping. At the same time, to ensure the stability of the internal communication bit width of the cache, add a 1-bit identification bit to the memory access signals within the group to indicate whether the memory access signal of this path is valid, and maintain the number of concurrent access paths within the group unchanged. For the memory access requests storing the Index signal after grouping, the identification bit is set to valid, and for the memory access requests filled with invalid signals, the identification bit is set to invalid.

[0057] The specific implementation process is as Figure 3 shown in the example. The parallel computing core sends n-way memory access request signals to the cache ( Figure 3In the middle, it is represented by Addr 0 to Addr n. The cache internal adaptability arbitration mechanism first slices the n-way memory access requests according to the cache capacity parameter, dividing them into n-way TAG signals, n-way index signals, and n-way word_sel signals. Subsequently, it compares the values of the n-way index signals, arranges the index values in order into a total of m memory access requests from addr_C0 to addr_Cm. In each memory access request, there are n-way concurrent signals, where the tag signal and the word_sel signal are completely filled with the tag and word_sel signals sent by the parallel computing core, and the index signal uniformly uses the sorted index signal. The identification signal is used to indicate that only the paths with the same index among the n-way concurrent signals are valid. Through grouping and merging, m groups of n-way concurrent memory access requests are formed ( Figure 3 represented by Addr_C0 to Addr_Cn in the middle).

[0058] After the above adaptability arbitration mechanism completes its work, the m times of n-way concurrent memory access requests will be sequentially sent to the specific data units inside the cache in sequence to complete the data reading and writing operations, such as Figure 4 shown.

[0059] First, the cache data module simultaneously reads the cache_tag data of the number of associative groups from the multi-way set-associative cache tag storage according to the common memory access index of the n-way memory access requests sent, that is Figure 4cache_tag0-cache_tagc in the memory. Then compare the TAG data with valid identification bits in the n-way memory access request sent and the cache_tag0-cache_tagc data obtained by memory access to see if they match (tag comparison), and assign the n-bit wide cache hit register corresponding to each group associative module. Specifically, if there is a valid identification bit in a certain group associativity and TAG=cache_tag, then the identification bit in the n-bit wide cache hit register corresponding to the group associativity is pulled high. In this way, the memory access hit / miss situation of the valid signal in the n-way memory access request sent and the group associative position of the cache hit can be obtained. At the same time, since the cache request is divided into read data request and write data request, the overall cache hit situation is divided into four types: read data request hit, read data request miss, write data request hit, and write data request miss. Among them, a read data request hit belongs to data reading, a read data request miss and a write data request hit belong to data updating. When a read data request hits, the input is the address. After judging that a certain cache group-associative module is hit according to the address, the corresponding data decompression scheme stored therein can be executed. When a read data request is missed, the core inputs the address. After there is no data in the cache at this level, a request is sent to the cache at the next level. The cache at the next level responds to the request and gives data. This data needs to be updated to the cache at this level, so the corresponding compression algorithm needs to be selected according to the response data characteristics given by the cache at the next level, and stored in the corresponding group-associative module; when a write data hits, the input is the address and data. The corresponding compression algorithm needs to be selected first according to the hit group-associative module, the data in the cache is read out, and then combined with the write data, and the corresponding compression algorithm is further selected according to the characteristics of the combined data, and stored in the corresponding group-associative module.

[0060] The specific implementation process is described in detail below:

[0061] (1) Read data request missing execution mechanism:

[0062] If there is one or more read data requests among the n memory access requests sent that miss the valid signal, the read data request miss execution mechanism is triggered. The unhit memory access requests will be extracted line by line from the concurrently executed n memory access requests and stored in the MEM_READ_QUEUE (memory read queue) in sequence, and finally the memory access requests will be sent to the next-level cache in sequence. After receiving the request signal from this-level cache, the next-level cache will also perform a memory access request hit judgment. Assuming that the next-level cache request hits, a response signal will be fed back to this-level cache. In the present invention, the response signal fed back by the lower-level cache is divided into a response address signal and a response data signal. The response data signal will first enter the compression and decompression algorithm way judgment module for judgment. This module is a unique module in the present invention's solution and is responsible for functions such as selecting the compression algorithm according to the input data characteristics, executing data compression, and determining the data storage set-associative module according to the selected compression algorithm. After the compression and decompression algorithm way judgment module completes data compression and set-associative module selection based on the input response data signal, the read data request miss execution mechanism will also complete the splitting of the response address signal. This module reuses the adaptive arbitration mechanism in the cache module according to the internal capacity and performance parameters of the cache, and splits the response address signal into tag, index, and word_sel signals according to the same formula. Subsequently, the cache module will determine the set-associative module according to the obtained set-associative selection result, determine the corresponding cache line in the set-associative module according to the index, write the compressed response data into the corresponding cache line of the cache data (cache_data) corresponding to the set-associative module, and write the split tag data into the corresponding cache line of the cache tag (cache_tag) corresponding to the set-associative module. At the same time, to ensure that the parallel computing core can obtain the response data with as low latency as possible, while the response signal fed back by the lower-level cache is being compressed and stored, there will also be a direct path to feedback to the parallel computing core.

[0063] (2) Read data request hit execution mechanism:

[0064] If there is one or more read data requests hitting among the valid signals in the n-way memory access requests sent, the read data request hit execution mechanism is triggered. The response data will be obtained from the corresponding cache line of the cache data (cache_data) corresponding to the hit set-associative module and the sliced index data, and the obtained response data and the set-associative module will be sent to the compression / decompression algorithm way judgment module for processing. Since the data stored in each set-associative module of the cache data (cache_data) is compressed data, in the read data request hit execution mechanism, the compression / decompression algorithm way judgment module is responsible for completing the data decompression operation. Specifically, the compression / decompression algorithm way judgment module is responsible for determining the data compression algorithm to be executed according to the received set-associative module, then completing the reverse decompression operation of the data according to the data compression algorithm, and feeding back the decompressed data to the parallel computing core module.

[0065] (3) Write data request hit execution mechanism:

[0066] If there is one or more write data requests hitting among the valid signals in the n-way memory access requests sent, the write data request hit execution mechanism is triggered. The stored data will be obtained from the corresponding cache line of the cache data (cache_data) corresponding to the hit set-associative module and the sliced index data, and the stored data will be sent to the compression / decompression algorithm way judgment module for decompression processing to obtain the decompressed stored data. Subsequently, the decompressed stored data is combined with the write data sent by the parallel computing core to obtain the complete write data, and the complete write data is sent back to the compression / decompression algorithm way judgment module for data judgment. Since there are data changes in the complete write data compared with the original compressed stored data, the original data compression scheme may not necessarily adapt to the data characteristics after the change, and the compression / decompression algorithm way judgment module needs to re-judge the data compression scheme. If the judgment result is the original data compression scheme, after the compression / decompression algorithm way judgment module completes the compression of the data after the change, it continues to store the compressed data in the corresponding index cache line of the original set-associative module; if the judgment result is a new data compression scheme, on the one hand, the compression / decompression algorithm way judgment module completes the data compression according to the new data compression scheme and writes the compressed data into the corresponding index cache line of the new set-associative module, and on the other hand, it needs to invalidate the data in the corresponding index cache line of the original set-associative module.

[0067] (4) Write data request miss execution mechanism:

[0068] If there is one or more write data requests among the n memory access requests sent that do not hit the valid signal, the write data request miss execution mechanism is triggered. This mechanism does not trigger the actions of the cache data (cache_data) and the compression and decompression algorithm way judgment module in the current-level cache. Instead, it directly extracts the unhit memory access requests line by line from the concurrently executed n memory access requests and stores them in the MEM_WRITE_QUEUE (storage write queue) in sequence. Finally, the memory access requests are sent to the next-level cache in sequence. It should be noted that due to the port width limit of the data request sent from the current-level cache to the next-level cache, if there are data requests waiting to be transmitted in both the MEM_WRITE_QUEUE and the MEM_REQ_QUEUE, the requests in the MEM_READ_QUEUE are preferentially transmitted.

[0069] The compression and decompression algorithm way judgment module is a unique module of the present invention's solution, mainly responsible for functions such as selecting the compression algorithm according to the input data characteristics, executing data compression, and associating with the storage group associative module. As Figure 5 shown, the present invention's solution statistically analyzes the application data characteristics of existing parallel computing cores, comprehensively combines the leading cache data compression solutions in the industry, selects four cases of two data compression solutions as the default compression solutions, and leaves enough extended compression solution interfaces as reserved compression solutions. The overall implementation solution is as follows:

[0070] The present invention's solution first statistically analyzes the compression effects and applicable probabilities of various data compression solutions, and finds that BDI(2,0), BDI(4,0), BDI(4,2) in the BDI compression algorithm and DIC(8) in the DIC compression algorithm have data compression applicable frequencies of up to 11.4%, 15.6%, 10.9% and 9.1% respectively in the tests of various application programs of the parallel computing core, which are much higher than other data compression algorithms. Therefore, they are set as the 4 default data compression solutions of the parallel computing core cache system and method based on multiple compression solutions of the present invention.

[0071] When a cache line containing a large amount of data enters the compression and decompression algorithm way judgment module, the default data compression solution fitting work is first performed, and the fitting work priority of the BDI compression algorithm is higher than that of the DIC compression algorithm. The flowchart of the fitting work of the BDI compression algorithm is as Figure 6As shown in the figure, first slice the entire cache line data by 2 bytes to form distributed data of (k + 1) * 2 bytes. Using the first data 0 as the reference data, subtract the reference data from the subsequent k data to obtain difference data 1 to k. Compare whether the difference data 1 to k are all 0. If they are all 0, it conforms to the BDI(2,0) compression algorithm rule, and only the 2-byte reference data 0 needs to be stored in the set-associative storage module 0 to represent the overall data. Therefore, store the 2-byte reference data 0 in the set-associative storage module 0. If the difference data are not all 0, re-slice the data; the re-sliced data is based on 4 bytes, and similarly form distributed data of (k / 2 + 1) * 4 bytes. Again, use the first data 0 as the reference data, subtract the reference data from the subsequent k / 2 data to obtain difference data 1 to k / 2. Compare whether the difference data 1 to k / 2 are all 0. If they are all 0, it conforms to the BDI(4,0) compression algorithm rule, and only the 4-byte reference data 0 needs to be stored in the set-associative storage module 1 to represent the overall data. Therefore, store the 4-byte reference data 0 in the set-associative storage module 1. If the differences are not all 0 but the difference bit width is within 2 bytes, it conforms to the BDI(4,2) compression algorithm rule, and store the 4-byte reference data 0 and the k-byte difference data in the set-associative storage module 2. If the difference bit width is greater than 2 bytes, it does not conform to the BDI compression algorithm rule, and perform DIC compression algorithm judgment.

[0072] The flowchart of the DIC compression algorithm adaptation work is as Figure 7 shown in the figure. Slice the data by 4 bytes to form distributed data of (k / 2 + 1) * 4 bytes, and sequentially store the distributed data [0, k / 2] in the data storage dictionary. The data storage dictionary is composed of a content-addressable memory (CAM), ensuring that only one dictionary index is fed back for the storage of the same data, thereby realizing the work of converting the actual data into a dictionary index. While continuously generating the dictionary index, perform dictionary depth judgment. When all the dictionary indexes of the distributed data are generated and the depth does not exceed the dictionary depth, it indicates that the repetition rate of the distributed data is good, and the DIC data compression algorithm is adapted. Store the dictionary index as data in the set-associative storage module 3. If the dictionary index depth generated by the distributed data exceeds the dictionary depth, it indicates that the repetition rate of the distributed data is poor, and the DIC data compression algorithm cannot be used. In this embodiment, the dictionary depth is set to 2 bytes to achieve the best data compression effect.

[0073] When the adaptation results of both the BDI compression algorithm and the DIC compression algorithm are in line, the BDI compression algorithm is preferentially used to compress the data; when the adaptation result of the BDI compression algorithm is not in line and the adaptation result of the DIC compression algorithm is in line, the DIC compression algorithm is used to compress the data. When the adaptation results of the default data compression scheme are not in line, the adaptation of the extensible data compression scheme is started. If there is no extensible data compression scheme yet or the adaptation result of the extensible data compression scheme is also not in line, the data is not compressed.

[0074] The present invention configures an associative memory module for each data compression scheme and associates it therewith, and configures an associative memory module for storing uncompressed data at the last stage. If a certain data compression algorithm is in line, the compressed data is stored in the corresponding associative memory module. At the same time, the reserved extensible data compression scheme also has a corresponding associative memory module associated with it. When there is no extensible data compression scheme, the data stored in the corresponding associative memory module will be the same as that in the last-stage associative memory module, storing uncompressed cache data.

[0075] Embodiment 2

[0076] Combined with Figure 8 As shown, the present disclosure provides a parallel computing core cache device 300 based on multiple compression schemes, including a processor 304 and a memory 301. Optionally, the device may further include a communication interface 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can complete communication with each other through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logical instructions in the memory 301 to execute the parallel computing core cache method based on multiple compression schemes in the above embodiments.

[0077] In addition, when the logical instructions in the above-mentioned memory 301 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0078] The memory 301, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the methods in the embodiments of the present disclosure. The processor 304 executes functional applications and data processing by running the program instructions / modules stored in the memory 301, that is, implements the parallel computing core cache method based on multiple compression schemes in the above embodiments.

[0079] The memory 301 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 301 may include a high-speed random access memory and may also include a non-volatile memory.

[0080] Embodiment 3

[0081] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are configured to execute the above-mentioned parallel computing core caching method based on a multi-compression scheme.

[0082] The above-mentioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium.

[0083] The technical solution of the embodiment of the present disclosure may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiment of the present disclosure. The foregoing storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, etc., which are various media that can store program codes, or may also be a transient storage medium.

[0084] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terms used in this application are only for describing the embodiments and do not limit the scope of protection. As used in the description herein, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, or apparatus including the element. Herein, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method parts disclosed in the embodiments, the relevant parts may refer to the description of the method parts.

[0085] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. The skilled person may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. The skilled person can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0086] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms. The units described as separate components can be or can not be physically separated. The components displayed as units can be or can not be physical units, that is, they can be located in one place or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, the various functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

Claims

1. A parallel computing core caching method based on multiple compression schemes, characterized in that: include: Step A, receiving n-way concurrent memory access requests from the parallel computing core; Step B: performing adaptability arbitration according to the internal capacity performance parameters of the cache to determine the memory access request sequence; First, according to the depth and bit width of the internal storage unit of the read cache, the word_sel size in the memory access request is determined according to the bit width of the internal storage unit of the cache, the index size in the memory access request is determined according to the depth of the internal storage unit of the cache, and the tag size in the memory access request is determined according to the bit width of the processor; The size of word_sel, index, and tag is determined as follows: , in , , Respectively represent the size of word_sel, index, tag, Width, Depth and SIZE core They represent the internal storage unit width of the cache, the internal storage unit depth of the cache, and the processor bit width respectively; Then read the index data in the n-way memory access request sent by the parallel computing core to the cache, and group and merge the n-way memory access requests according to the index data. The grouping and merging method is to merge the memory access requests with the same index into one group, and arrange the indexes in order between the groups. Through grouping and merging, m groups of n-way concurrent memory access requests are formed; m and n are both positive integers, and m≤n; Step C, execute the memory access request adapted to the multi-cache data compression scheme, the m groups of n-way concurrent memory access requests formed in step B are sent to the cache, the cache adopts a multi-way group associative cache construction mechanism, each cache group associative module stores a data compression scheme, and determines whether the memory access request is hit or missed by label comparison. If the request is hit, the request hit execution mechanism is triggered, and if the request is missed, the request miss execution mechanism is triggered; during the execution of the request hit execution mechanism and the request miss execution mechanism, if it is data reading, the data compression method stored in the request hit cache group associative module is executed, and if it is data updating, the compression algorithm is selected according to the input data characteristics, data compression is executed, and the storage group associative module is associated.

2. The parallel computing core cache method based on multiple compression schemes according to claim 1, characterized in that: When performing grouping in step B, if the total number of memory access requests with the same index does not reach n-way, an invalid signal is filled in so that each group of memory access requests reaches n-way. The m groups of n-way concurrent memory access requests formed by grouping all include a TAG signal bit, an Index signal bit, a WORD_SEL signal bit and an identification bit. The TAG signal bit and the WORD_SEL signal bit are directly filled with the TAG signal and the WORD_SEL signal in the n-way concurrent memory access request sent by the parallel computing core. The Index signal bit is filled with the Index signal after grouping. The memory access request of the Index signal after grouping is stored, and the identification position is valid. The memory access request filled with the invalid signal has the identification position invalid.

3. The parallel computing core cache method based on multiple compression schemes according to claim 1, characterized in that: In step C, the request hit includes a read data request hit, and the request miss includes a read data request miss; When a read data request is missed, the request missing execution mechanism is as follows: if one or more read data requests do not hit the valid signal in the n-way memory access request sent, the read data request missing execution mechanism is triggered, and the missed memory access requests are extracted row by row from the concurrently executed n-way memory access requests and sent to the next level cache; After receiving the request signal from the cache at this level, the next level cache performs a memory access request hit judgment. If the next level cache request hits, a response signal is fed back to the cache at this level. The response signal includes a response address signal and a response data signal. The response data signal first enters the compression and decompression algorithm way judgment module to perform input data feature analysis, and completes compression algorithm selection, data compression execution, and group associative module selection based on the input data features. Then, the response address signal is split into tag, index, and word_sel signals. The corresponding cache line in the group associative module is determined based on the index signal, and the compressed response data is written into the cache line corresponding to the group associative module corresponding to the cache data, and the split tag data is written into the cache line corresponding to the group associative module corresponding to the cache tag. When a read data request hits, the request hit execution mechanism is: if there is one or more read data request hits in the valid signals of the n-way memory access request sent, the read data request hit execution mechanism is triggered, and the response data is obtained from the cache line corresponding to the group associative module of the cache data according to the hit group associative module and index data, and the obtained response data and the group associative module are sent to the compression and decompression algorithm way judgment module for processing. The compression and decompression algorithm way judgment module determines the data compression algorithm to be executed according to the received group associative module, and then completes the reverse decompression operation of the data according to the data compression algorithm, and feeds the decompressed data back to the parallel computing core module.

4. The parallel computing core cache method based on multiple compression schemes according to claim 1, characterized in that: A request hit includes a write data request hit, and a request miss includes a write data request miss; When a write data request hits, the request hit execution mechanism is as follows: if one or more write data request hits exist in the valid signals of the n-way memory access request sent, the write data request hit execution mechanism is triggered, and storage data is obtained from the cache line corresponding to the group associative module corresponding to the cache address according to the hit group associative module and index data, and the storage data is sent to the compression and decompression algorithm way judgment module for decompression processing to obtain decompressed storage data, and then the decompressed storage data is combined with the write data sent by the parallel computing core to obtain complete write data, and the complete write data is resent to the compression and decompression algorithm way judgment module for data feature The compression and decompression algorithm way judgment module re-judges the data compression scheme according to the characteristics of the complete written data. If the judgment result is the original data compression scheme, the compression and decompression algorithm way judgment module completes the compression of the complete written data and continues to store the compressed data in the index cache line corresponding to the original group associative module; if the judgment result is a new data compression scheme, the compression and decompression algorithm way judgment module completes the data compression according to the new data compression scheme and writes the compressed data into the index cache line corresponding to the new group associative module, and invalidates the data of the index cache line corresponding to the original group associative module on the other hand; When a write data request is missing, the request missing execution mechanism is: if there is a valid signal in the n-way memory access request sent and one or more write data requests are not hit, the write data request missing execution mechanism is triggered. This mechanism does not trigger the data cache action in the current level cache and the compression and decompression algorithm way judgment module action, but directly extracts the missed memory access requests line by line from the concurrently executed n-way memory access requests and sends them to the next level cache.

5. The parallel computing core cache method based on multiple compression schemes according to claim 3 or 4, characterized in that: The compression and decompression algorithm way judgment module stores four compression schemes as the default data compression scheme. When input data enters, the compression and decompression algorithm way judgment module first performs the default data compression scheme matching work, and selects the corresponding compression scheme according to the data characteristics and compression scheme priority. Each data compression scheme is configured with a group associative storage module associated with it. If a data compression algorithm matches, the compressed data will be stored in the corresponding group associative storage module.

6. The parallel computing core cache method based on multiple compression schemes according to claim 5, characterized in that: The four compression schemes stored in the compression and decompression algorithm way judgment module are: BDI(2,0), BDI(4,0), BDI(4,2) in the BDI compression algorithm and DIC(8) in the DIC compression algorithm. The default data compression scheme based on these four compression schemes matches the following workflow: slice the entire cache line data by 2 bytes to form (k+1)*2 bytes of distributed data. The first data 0 is used as the reference data, and the following k data are subtracted from the reference data to obtain the difference data 1 to k. Compare whether the difference data 1 to k are all 0. If they are all 0, it complies with the BDI(2,0) compression algorithm rules. , store the 2-byte reference data 0 in the group associative storage module 0. If the difference data is not all 0, re-slice the data; the re-sliced ​​data is based on 4 bytes, and the distributed data of (k / 2+1)*4 bytes is also formed. The first data 0 is also used as the reference data. The following k / 2 data are subtracted from the reference data to obtain the difference data 1 to k / 2. Compare whether the difference data 1 to k / 2 are all 0. If they are all 0, it complies with the BDI (4,0) compression algorithm rules, and store the 4-byte reference data 0 in the group associative storage module 1; if the difference is not all 0 but the difference bit width is within 2 bytes, it complies with BDI (4,2) compression algorithm rules, the 4-byte reference data 0 and the k-byte difference data are stored in the group associative storage module 2; if the difference bit width is greater than 2 bytes, it does not meet the BDI compression algorithm rules, and the DIC compression algorithm is judged. The judgment process is: slice the data according to 4 bytes to form (k / 2+1)*4 bytes of distributed data, and store the distributed data [0,k / 2] in the data storage dictionary in sequence. The data storage dictionary is composed of content-addressable memory. The storage of the same data will only feedback one dictionary index. While the dictionary index is continuously generated, the dictionary depth judgment is performed. When the dictionary indexes of all distributed data are After the generation is completed and the depth does not exceed the dictionary depth, it means that the repetition rate of the distributed data is good, and the DIC data compression algorithm is adapted, and the dictionary index is stored as data in the group associative storage module 3. If the dictionary index depth generated by the distributed data exceeds the dictionary depth, it means that the repetition rate of the distributed data is poor, and the DIC data compression algorithm cannot be used. When the BDI compression algorithm adaptation work result and the DIC compression algorithm adaptation work result are both consistent, the BDI compression algorithm is used to compress the data; when the BDI compression algorithm adaptation work result is inconsistent and the DIC compression algorithm adaptation work result is consistent, the DIC compression algorithm is used to compress the data.

7. The parallel computing core cache method based on multiple compression schemes according to claim 6 is characterized in that: The compression and decompression algorithm way judgment module reserves an extensible compression scheme interface. When the default data compression scheme adaptation work results are all incompatible, the extensible data compression scheme adaptation work is started. If there is no extensible data compression scheme or the extensible data compression scheme adaptation work results are also incompatible, the data will not be compressed. The reserved extensible data compression scheme also has a corresponding group associative storage module associated with it. When there is no extensible data compression scheme, the data stored in the corresponding group associative storage module stores uncompressed cache data.

8. A parallel computing core cache device based on multiple compression schemes, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the parallel computing core cache method based on multiple compression schemes as described in any one of claims 1 to 7 when running the program instructions.

9. A storage medium storing program instructions, characterized in that: When the program instructions are run, they execute the parallel computing core cache method based on multiple compression schemes as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Scalable application-customized memory compression

    US20190243780A1

  • Data compression storage system and method, processor, and computer storage medium

    WO2021237513A1