Data storage method, data reading method and graphics processor

By using a targeted compression algorithm in the GPU to store data between the GPU and non-GPU memory, the memory capacity bottleneck problem is solved and the data storage and access efficiency is improved.

CN120705075APending Publication Date: 2025-09-26HYGON INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510787347.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The memory capacity of graphics processors has become a bottleneck for performance improvement in high-performance computing and deep learning. Existing memory expansion technologies lead to memory fragmentation and reduced access performance.

Method used

A predetermined target compression algorithm is used to compress the page data to be compressed. The portion of data that is equal to the storable data amount is stored in the graphics processor memory, and the excess portion is stored in the non-graphics processor memory. The storage address is recorded to avoid page reallocation.

Benefits of technology

Improves the data storage and access efficiency of the graphics processor memory, and enables memory capacity expansion without additional overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705075A_ABST
    Figure CN120705075A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data storage method, a data reading method and a graphics processor. The data storage method is applied to the graphics processor, and comprises the following steps: acquiring page data to be compressed; performing compression by using a predetermined target compression algorithm to obtain compressed page data; determining the storable data volume in the memory of the graphics processor according to the target compression rate of the target compression algorithm; in the compressed page data, storing the page data of which the data volume is equal to the storable data volume into a graphics processor memory, storing the page data of which the data volume exceeds the storable data volume into an applied non-graphics processor memory, and recording a storage address; the applied non-graphics processor memory is a memory applied in a processor other than the graphics processor, and the storage capacity is at least determined by the target compression rate. According to the data storage method provided by the embodiment of the invention, the data storage and access efficiency in the memory of the graphics processor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer technology, and in particular to a data storage method, a data reading method, and a graphics processor. Background Art

[0002] With the growing demand for computing resources in the fields of High Performance Computing (HPC) and Deep Learning (DL), the memory capacity of Graphics Processing Unit (GPU) has gradually become a bottleneck restricting performance improvement.

[0003] GPU memory expansion technology addresses memory bottlenecks by increasing physical memory, sharing memory, or optimizing memory management strategies. One existing GPU memory expansion technology reduces the amount of data in GPU memory by compressing data between the GPU's data cache and memory. However, compressed page data can cause memory fragmentation, impacting memory access performance. Against this backdrop, providing a technical solution to improve the data storage and access efficiency of graphics processors has become a pressing technical challenge for those skilled in the art. Summary of the Invention

[0004] Embodiments of the present invention provide a data storage method, a data reading method, and a graphics processor to improve the data storage and access efficiency of the graphics processor.

[0005] In a first aspect, an embodiment of the present invention provides a data storage method, applied to a graphics processor, comprising:

[0006] Get the page data to be compressed;

[0007] Compressing the to-be-compressed page data using a predetermined target compression algorithm to obtain compressed page data; a target compression rate of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory;

[0008] storing, in the compressed page data, page data having a data volume equal to the storable data volume in the graphics processor memory, and storing page data having a data volume exceeding the storable data volume in the applied non-graphics processor memory, and recording the storage addresses of the page data in the applied non-graphics processor memory;

[0009] The applied non-GPU memory is a memory applied for in a processor other than the GPU, and the storage capacity of the applied non-GPU memory is determined based at least on the target compression ratio.

[0010] In a second aspect, an embodiment of the present application provides a data reading method, applied to a graphics processor, comprising:

[0011] Obtaining a data read request, wherein the data read request includes a data read address;

[0012] When it is determined that the page data corresponding to the data read address is compressed page data, and the compressed page data is stored in the graphics processor memory and the applied non-graphics processor memory, based on the data read address, page data having a data volume equal to the data volume storable in the graphics processor memory is obtained from the graphics processor memory, and page data having a data volume exceeding the data volume storable in the applied non-graphics processor memory is obtained based on the storage address in the data read request, thereby obtaining compressed page data; the compressed page data is data stored by the data storage method according to the first aspect; the data volume storable is determined by a target compression ratio of a target compression algorithm adopted by the data storage method;

[0013] Determining a decompression algorithm for the compressed page data using a target compression algorithm, and obtaining decompressed page data;

[0014] Data is read from the decompressed page data.

[0015] In a third aspect, an embodiment of the present application provides a graphics processor, including:

[0016] a compressor configured to obtain page data to be compressed and compress the page data using a predetermined target compression algorithm to obtain compressed page data; a target compression rate of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory;

[0017] a memory controller configured to store, in the compressed page data, page data having a data amount equal to the storable data amount into the graphics processor memory, and store page data having a data amount exceeding the storable data amount into an applied non-graphics processor memory, and record the storage address of the page data in the applied non-graphics processor memory;

[0018] The applied non-GPU memory is a memory applied for in a processor other than the GPU, and the storage capacity of the applied non-GPU memory is determined based at least on the target compression ratio.

[0019] An embodiment of the present invention provides a data storage method, applied to a graphics processor, that first obtains to-be-compressed page data and compresses the to-be-compressed page data using a predetermined target compression algorithm to obtain compressed page data. The target compression ratio of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory. Because the compressed data obtained after the compression of the to-be-compressed page data may not meet the data volume (the amount of data that can be stored) determined by the target compression ratio of the predetermined target compression algorithm, the embodiment of the present invention can, for a compressed page data, store page data of the compressed page data that has a data volume equal to the amount of data that can be stored in the graphics processor memory, and store page data of the compressed page data that has a data volume exceeding the amount of data that can be stored in a previously allocated non-GPU memory, and record the storage address of the page data in the previously allocated non-GPU memory. The previously allocated non-GPU memory is determined based at least on the target compression ratio of the target compression algorithm, so that the compressed page data can be stored completely at one time. This can not only ensure that there are no gaps in the data stored in the GPU memory, but also avoid the GPU using page reallocation to re-apply for GPU memory to store the compressed page data that exceeds the storable data capacity, thereby improving the efficiency of data storage in the GPU memory; when the compressed page data stored in the applied non-GPU memory is needed later, it can be directly obtained based on the recorded storage address. Therefore, on the basis of reducing the capacity occupation of the GPU memory and realizing the expansion of the GPU memory capacity, the additional overhead caused by the GPU using page reallocation to store the compressed page data can be avoided, thereby achieving the purpose of improving the efficiency of data storage in the GPU memory. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0021] Figure 1 This is a flow chart of a data storage method provided by an embodiment of the present invention.

[0022] Figure 2 This is a flow chart of a data reading method provided by an embodiment of the present invention.

[0023] Figure 3 FIG. 4 is a schematic diagram of the structure of a graphics processor provided by an embodiment of the present invention.

[0024] Figure 4 This is a schematic diagram of an implementation process of the data storage method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] As mentioned in the background technology section, as data processing demands continue to increase, the demand for memory capacity on graphics processing units (GPUs) used to assist with data processing operations is also increasing. GPU memory capacity (GPU memory) has gradually become a bottleneck restricting performance improvements. For example, for a currently popular large language model (LLM) with 70 billion parameters, a single model copy requires approximately 1260GB of GPU memory when performing mixed-precision training. However, the memory capacity of a single GPU used for LLM training typically ranges from tens to hundreds of GB. Therefore, multiple GPUs are required for model deployment and training. For example, deploying a single 1260GB LLM copy requires approximately 16 80GB GPUs.

[0026] However, introducing multiple GPUs for training brings many challenges. Although model parallel technologies such as pipeline parallelism and tensor parallelism have successfully solved the problem of model segmentation, these technologies introduce additional high interconnection bandwidth requirements and complex requirements for resource scheduling such as computing and communication.

[0027] Furthermore, there are some factors that make it difficult to expand the physical capacity of GPU memory.

[0028] Since the memory chips in the GPU need to be encapsulated inside the GPU, taking High Bandwidth Memory (HBM) as an example, existing GPUs can only accommodate a certain number of HBM stacks, and the memory capacity of each HBM stack is fixed. It can be seen that the physical capacity of GPU memory is limited by physical space.

[0029] In addition, the demand for GPU physical memory expansion is lower than the demand for memory speed increase. In order to ensure the computing performance of the GPU, the demand for GPU physical memory speed increase is often higher than its expansion demand. Therefore, during the design and development stage of the GPU, the demand for GPU physical memory speed increase rather than expansion demand will be prioritized.

[0030] Memory compression technology can reduce memory usage requirements and has been widely used in CPU systems. By applying compression algorithms to data stored in memory, memory usage is reduced, thereby slowing down CPU performance degradation caused by insufficient memory. Memory compression technology can also be applied to GPUs.

[0031] An existing GPU memory expansion technology reduces data usage in GPU memory by compressing data between the GPU's data cache and memory. However, since data is stored in a pre-set format in GPU memory, compression can leave some memory spaces empty, potentially causing memory fragmentation. To neatly store these compressed pages, the GPU's memory manager uses page reallocation to reclaim memory pages to store the compressed data. However, page reallocation not only consumes a certain amount of bandwidth but also increases memory access latency, impacting data storage and access efficiency.

[0032] To solve the above problems, the present invention provides a data storage method, a data reading method and a graphics processor to improve the data storage and access efficiency in GPU memory.

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0034] Figure 1 This is a flow chart of a data storage method provided by an embodiment of the present invention. Figure 1 The data storage method is applied to a graphics processor and includes the following steps.

[0035] Step S110: Obtain page data to be compressed.

[0036] The data generated during the GPU program execution cannot be fully accommodated in the GPU computing core's internal cache (e.g., L1 cache) and needs to be stored in the GPU memory for subsequent access. The page data to be compressed can be data generated by the GPU computing core that cannot be accommodated in the GPU computing core's cache.

[0037] After obtaining the page data to be compressed, proceed to step S120.

[0038] Step S120: compressing the to-be-compressed page data using a predetermined target compression algorithm to obtain compressed page data.

[0039] The target compression rate of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory.

[0040] The page data to be compressed can be compressed using a predetermined target compression algorithm, thereby reducing the capacity occupied by the page data to be compressed in the graphics processor memory and expanding the memory capacity of the graphics processor.

[0041] In some implementations, to determine a target compression algorithm to improve compression performance of the to-be-compressed page data, the following steps are performed before step S110 is performed.

[0042] Step S100: Acquire debugging page data.

[0043] Before the graphics processor runs a program, it needs to first be tested based on a test data set to analyze the program's performance. The debug page data is the data generated when the graphics processor runs the program based on the test data set. After obtaining the debug page data, step S101 is continued.

[0044] Step S101: compressing the debug page data multiple times using multiple candidate compression algorithms, and determining the compression rate of the debug page data each time using the candidate compression algorithm for each candidate compression algorithm, and obtaining the average compression rate of the debug page data based on the compression rate determined each time.

[0045] Due to the different types of programs running on the GPU, the data generated may vary significantly. Different candidate compression algorithms have different compression performance for different data. Therefore, the debug page data can be compressed multiple times using multiple candidate compression algorithms to obtain the average compression ratio of each algorithm for the debug page data. Subsequently, the appropriate compression algorithm for the program can be selected based on requirements to improve GPU performance.

[0046] The candidate compression algorithms are various commonly used compression algorithms that can perform data compression, such as Huffman coding, LZ77 (a compression algorithm based on a sliding window), DEFLATE (a compression algorithm that combines the LZ77 algorithm and Huffman coding), etc., which are widely used in file compression and data transmission and can completely restore the original data.

[0047] After determining the average compression ratio of the debugging page data under each compression algorithm, proceed to step S102.

[0048] Step S102: selecting a candidate compression algorithm corresponding to the maximum value of the average compression ratios as a target compression algorithm, and taking the average compression ratio of the target compression algorithm as a target compression ratio.

[0049] Selecting the candidate compression algorithm corresponding to the maximum value among the average compression ratios can further reduce the data storage pressure on the GPU memory, thereby improving the memory capacity expansion performance of the GPU. In some other implementations, the target compression algorithm can also be selected based on other performance requirements. For example, the compression algorithm with the fastest compression speed and an average compression ratio that can meet the data storage requirements of the GPU memory can be selected as the target compression algorithm.

[0050] In some implementations, after step S102 is executed, the target compression ratio is saved for use in subsequent steps.

[0051] It is understood that in some implementations, the target compression algorithm can be determined directly. For example, a default compression algorithm in a graphics processor can be used as the target compression algorithm, or a specific compression algorithm can be used as the target compression algorithm. This eliminates the need to perform steps S100 to S102 to determine the target compression algorithm, thereby simplifying the logic of the data storage method and improving the execution efficiency of the data storage method.

[0052] Continue to refer Figure 1 After executing step S120, continue to execute step S130.

[0053] Step S130: storing, in the compressed page data, page data having a data size equal to the storable data size in the GPU memory, and storing, in the non-GPU memory that has been applied for, page data having a data size exceeding the storable data size, and recording the storage address of the page data in the non-GPU memory that has been applied for.

[0054] The applied non-GPU memory is a memory applied for in a processor other than the GPU, and the storage capacity of the applied non-GPU memory is determined based at least on the target compression ratio.

[0055] The GPU memory is global memory on the GPU executing the data storage method, and the allocated non-GPU memory is memory allocated on a processor other than the GPU executing the data storage method, such as CPU memory. In some embodiments, the GPU and the non-GPU have a direct connection to increase the speed of subsequent access to data stored in the GPU memory and the non-GPU memory. For example, the GPU and the non-GPU can be directly connected via a high-bandwidth interconnect technology such as NVLink or CXL to improve the performance of the GPU.

[0056] After the to-be-compressed page data is compressed using a predetermined target compression algorithm, the compressed page data may have different data volumes. That is, if the compressed page data is directly stored in the GPU memory, the data volume of the compressed page data may exceed the data volume that can be stored in the GPU memory, causing the memory manager in the GPU to reallocate pages of the GPU memory, which affects the GPU's storage and access efficiency.

[0057] Furthermore, since each compression algorithm has its own compression ratio, i.e., a predetermined target compression algorithm has its corresponding target compression ratio, after the to-be-compressed page data is compressed according to the target compression algorithm, the amount of data in the compressed page data is at least equal to the storable amount of data; and may even exceed the storable amount of data. It is understood that since the amount of data in the compressed page data is at least equal to the storable amount of data, dividing the compressed page data according to the storable amount of data avoids wasted space.

[0058] Therefore, storing the compressed page data that is equal to the storable data volume in the GPU memory can prevent the GPU memory manager from reallocating pages within the GPU memory, thereby improving data storage and access efficiency. Furthermore, in addition to the compressed page data that is equal to the storable data volume, the compressed page data also includes page data that exceeds the storable data volume. Therefore, this portion of data needs to be stored in the already allocated non-GPU memory to ensure GPU data storage and access performance and avoid the need to re-allocate page memory using page reallocation technology. The compressed page data is then stored as a whole.

[0059] In some embodiments, after executing step S120 to obtain the compressed page data, the data storage method further includes the following step: determining a corresponding actual compression ratio based on the compressed page data. The actual compression ratio can be used to calculate the actual compressed data volume of the compressed page data.

[0060] Specifically, in some embodiments, the actual compression ratio is: the ratio of the uncompressed data volume of the page data to be compressed to the actual compressed data volume of the compressed page data. For example, if the uncompressed data volume of the page data to be compressed is 4KB and the actual compressed data volume of the compressed page data is 2KB, then the actual compression ratio is 4KB / 2KB=2.

[0061] After determining the actual compression ratio corresponding to the compressed page data, the actual compression ratio is compared with the target compression ratio.

[0062] If the actual compression ratio is less than the target compression ratio, it is determined that the actual compressed data amount is greater than the storable data amount, and step S130 is executed, that is, the page data of the compressed page data whose data amount is equal to the storable data amount is stored in the graphics processor memory, and the page data whose data amount exceeds the storable data amount is stored in the applied non-GPU memory, and the storage address of the page data in the applied non-GPU memory is recorded.

[0063] If the actual compression ratio is greater than or equal to the target compression ratio, then the actual compressed data volume is determined to be less than or equal to the storable data volume, indicating that no page data exceeds the storable data volume. The storable data volume of the graphics processor memory satisfies the storage of the compressed page data, and the compressed page data can be directly stored in the graphics processor memory. This indicates that by determining the difference between the actual compression ratio and the target compression ratio, the compressed page data can be accurately and appropriately stored completely, thereby improving data storage efficiency.

[0064] As can be seen, an embodiment of the present invention provides a data storage method, applied to a graphics processor, that first obtains to-be-compressed page data and compresses the to-be-compressed page data using a predetermined target compression algorithm to obtain compressed page data; the target compression ratio of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory. Because the compressed data obtained after the compression of the to-be-compressed page data may not meet the data amount (the amount of data that can be stored) determined by the target compression ratio of the predetermined target compression algorithm, the embodiment of the present invention can, for a compressed page data, on the one hand, store page data of the compressed page data that has a data amount equal to the amount of data that can be stored in the graphics processor memory, and on the other hand, store page data of the compressed page data that has a data amount that exceeds the amount of data that can be stored in a previously applied non-GPU memory, and record the storage address of the page data in the previously applied non-GPU memory. The previously applied non-GPU memory is determined based at least on the target compression ratio of the target compression algorithm, so that the compressed page data can be stored completely at one time. This can not only ensure that there are no gaps in the data stored in the GPU memory, but also avoid the GPU using page reallocation to re-apply for GPU memory to store the compressed page data that exceeds the storable data capacity, thereby improving the efficiency of data storage in the GPU memory; when the compressed page data stored in the applied non-GPU memory is needed later, it can be directly obtained based on the recorded storage address. Therefore, on the basis of reducing the capacity occupation of the GPU memory and realizing the expansion of the GPU memory capacity, the additional overhead caused by the GPU using page reallocation to store the compressed page data can be avoided, thereby achieving the purpose of improving the efficiency of data storage in the GPU memory.

[0065] In some implementations, after determining the target compression algorithm and before executing step S110 , the data storage method further includes the following steps.

[0066] Obtaining the available capacity of the graphics processor memory. The available capacity of the graphics processor memory included in the graphics processor is the maximum memory capacity that can be requested by a program running on the graphics processor, so that in subsequent steps, the program running on the graphics processor can maximize the capacity available to the program.

[0067] After executing the step of obtaining the available capacity of the graphics processor memory, the data storage method further includes: determining a requested capacity on the non-graphics processor memory based on the available capacity, with the requested capacity being the storage capacity of the non-graphics processor memory as the requested non-graphics processor memory; wherein the requested capacity is determined by multiplying the available capacity by a difference between the target compression ratio and a first value.

[0068] The requested capacity of the non-GPU memory is the maximum memory capacity that can be requested by a program running on the GPU. To ensure that the requested non-GPU memory can accommodate all data exceeding the target compression ratio, the requested capacity is calculated by multiplying the available capacity by the difference between the target compression ratio and a first value, where the first value may be 1. This prevents data storage or access exceptions due to insufficient capacity in the non-GPU memory before the GPU memory is fully occupied.

[0069] In some specific implementations, the applied non-GPU memory is further determined based on a pre-configured non-GPU memory capacity threshold to avoid excessive non-GPU memory capacity usage that affects non-GPU performance.

[0070] For example, after determining the target compression algorithm, the target compression ratio of the target compression algorithm is 4, and the actual memory capacity of the graphics processor memory, i.e., the available capacity, is 40 GB. Therefore, a memory space with a capacity of 40 GB is requested for the graphics processor memory, and a memory space with a capacity of 40 GB*(4-1)=120 GB is requested for the non-graphics processor memory (e.g., CPU memory), thereby ensuring normal operation of the program. If, in the above example, the operating system or program pre-configures a non-graphics processor memory capacity threshold of 100 GB, a memory space with a capacity of 100 GB is requested for the non-graphics processor memory, rather than a memory space with a capacity of 120 GB.

[0071] In some specific embodiments, to prevent data stored in the non-GPU memory from being swapped to the hard disk, the requested non-GPU memory is page-locked memory. Since data in the page-locked memory is not swapped to the hard disk, the GPU's access performance to the data in the non-GPU memory is guaranteed.

[0072] In some specific implementations, after the step of using the average compression rate of the target compression algorithm as the target compression rate, the following steps are further performed.

[0073] Saving the target compression ratio; and saving the base address of the applied non-GPU memory, wherein the base address is combined with the storage address to indicate: a storage location of page data exceeding the storable data amount in the applied non-GPU memory.

[0074] After the target compression ratio is saved, when the actual compression ratio needs to be compared with the target compression ratio, the saved target compression ratio can be read, thereby improving the efficiency of the data storage method. Furthermore, after the base address of the requested non-GPU memory is saved, when accessing compressed page data stored in the requested non-GPU memory is needed, fast and accurate memory access can be performed based on the base address.

[0075] In some embodiments, the actual compression ratio and the storage address are recorded in a page table entry for the compressed page data; the page table entry also records a page compression flag for the compressed page data. The page compression flag indicates whether the compressed page data has been compressed; and the actual compression ratio is the ratio of the uncompressed data size of the to-be-compressed page data to the actual compressed data size of the compressed page data.

[0076] By recording the page compression flag, actual compression ratio, and storage address, when subsequently accessing compressed page data stored in the graphics processor memory and the requested non-graphics processor memory, the compressed page data can be read based on the page compression flag, actual compression ratio, and storage address, thereby improving data access performance of the graphics processor.

[0077] In some specific implementations, in the page table entry of the compressed page data, the sizes of the page compression flag, actual compression ratio, and storage address (ie, the offset address in the applied non-processor memory) are 1 bit, 3 bits, and 16 bits, respectively.

[0078] To further improve the data access performance of the graphics processor, in some embodiments, after step S120 and before the step of determining the corresponding actual compression rate based on the compressed page data, the data storage method further includes: storing the compressed page data in the compressed data cache of the graphics processor.

[0079] Compressed data caching can reduce the frequency of GPU / CPU DRAM accesses, improving overall system performance. Since reading compressed page data from the CPU DRAM's allocated non-GPU memory requires address indexing, this involves accessing the TLB (Translation Lookaside Buffer), Page Table, and obtaining and using the base address of the allocated non-GPU memory. By directly caching some compressed page data in the compressed data cache, the access overhead associated with GPU / CPU DRAM accesses can be reduced.

[0080] When a computing core in a graphics processor accesses the compressed data cache, its access speed is faster than that of the graphics processor memory. Therefore, preferentially using the compressed data cache to store the compressed page data can improve the data access performance of the graphics processor. When the compressed data cache cannot accommodate the compressed page data, that is, the compressed page data needs to be stored in the memory, the step of comparing the actual compression rate with the target compression rate is continued, thereby further improving the data storage and reading performance of the graphics processor.

[0081] The embodiment of the present invention also provides a data reading method, which is applied to a graphics processor, referring to Figure 2 , the data reading method includes the following steps.

[0082] Step S201: Acquire a data read request, where the data read request includes a data read address.

[0083] Programs running on the GPU send data read requests through the GPU's computing core to obtain data from the GPU's cache or memory.

[0084] After obtaining the data read request, step S202 is executed.

[0085] Step S202: When it is determined that the page data corresponding to the data read address is compressed page data, and the compressed page data is stored in the graphics processor memory and the requested non-graphics processor memory, based on the data read address, page data having a data volume equal to the data volume storable in the graphics processor memory is obtained from the graphics processor memory, and based on the storage address in the data read request, page data having a data volume exceeding the data volume storable in the requested non-graphics processor memory is obtained from the requested non-graphics processor memory, thereby obtaining compressed page data.

[0086] The compressed page data is the data stored by the data storage method described in any of the aforementioned embodiments; the amount of data that can be stored is determined by the target compression rate of the target compression algorithm adopted by the data storage method.

[0087] If the page data corresponding to the data read address is compressed page data, for example, the page compression flag of the page table entry records that the page data is compressed page data, and the page table entry records the storage address of the requested non-GPU memory, then it means that the compressed page data corresponding to the data read address is stored in the GPU memory and the requested non-GPU memory. In this case, the page data can be read from the GPU memory and the non-GPU memory respectively to form the compressed page data.

[0088] In some specific embodiments, when it is determined that the page table entry for the compressed page data corresponding to the data read address records the storage address of the requested non-GPU memory, the following steps are further performed: obtaining a base address of the requested non-GPU memory; and obtaining a physical address of the non-GPU memory based on the base address and the storage address.

[0089] For example, the storage address of the applied non-GPU memory is the address offset (offset address) at which the compressed page data is stored in the non-GPU memory. Therefore, it is also necessary to obtain the base address of the applied non-GPU memory, and obtain the physical address in the non-GPU memory based on the base address and the storage address. The physical address in the non-GPU memory is the actual physical location where the compressed page data is stored in the non-GPU memory.

[0090] If the page table entry for the compressed page data corresponding to the data read address does not record the storage address of the requested non-GPU memory, it means that the compressed page data corresponding to the data read address is only stored in the GPU memory. Subsequently, it is only necessary to read the compressed page data corresponding to the data read address from the GPU memory.

[0091] In some specific implementations, step S202 includes the following steps:

[0092] In the graphics processor memory, page data having a data volume equal to a storable data volume is obtained according to the data read address; in the applied non-graphics processor memory, page data corresponding to a physical address of the non-graphics processor memory is obtained, where the obtained compressed page data is data exceeding the storable data volume; and compressed page data is obtained based on the page data obtained in the graphics processor memory and the page data obtained in the applied non-graphics processor memory.

[0093] Since the compressed page data is stored separately in the graphics processor memory and the applied non-GPU memory, specifically, among the compressed page data, page data with a data amount equal to the storable data amount is stored in the graphics processor memory, and page data with a data amount exceeding the storable data amount is stored in the applied non-GPU memory; therefore, it is necessary to read from the graphics processor memory and the applied non-GPU memory separately to obtain the complete compressed page data, so that the compressed page data can be subsequently decompressed to obtain the decompressed page data.

[0094] After obtaining the compressed page data, proceed to step S203.

[0095] Step S203: using the target compression algorithm to determine a decompression algorithm for the compressed page data, and obtaining the decompressed page data.

[0096] The target compression algorithm is the compression algorithm used for the compressed page data, and thus a corresponding decompression algorithm may be determined according to the compression method of the target compression algorithm to decompress the compressed page data.

[0097] After obtaining the decompressed page data, proceed to step S204.

[0098] Step S204: reading data from the decompressed page data.

[0099] After the compressed page data is decompressed, the graphics processor can directly read the data and can directly access the data according to program requirements.

[0100] It can be seen that the data reading method provided by the embodiment of the present invention first obtains a data read address of a data read request, and then, when it is determined that the corresponding page data is compressed page data and is stored in the graphics processor memory and the requested non-graphics processor memory, because the compressed page data is compressed using a predetermined target compression algorithm, and in the compressed page data, page data with a data amount equal to the storable data amount is stored in the graphics processor memory, and page data with a data amount exceeding the storable data amount is stored in the requested non-graphics processor memory; therefore, it is necessary to read the corresponding page data from the graphics processor memory and the requested non-graphics processor memory respectively according to the data read address and the storage address in the data read address, and finally obtain the compressed page data, and then decompress the compressed page data using the target compression algorithm to obtain the to-be-compressed page data required to be accessed by the graphics processor, thereby achieving the expansion of the memory capacity of the graphics processor.

[0101] In some implementations, after executing step S201, the method further includes:

[0102] Determining whether the data read address hits the cache of the graphics processor;

[0103] If yes, determining that the page data corresponding to the data read address is not compressed, and reading the data requested by the data read request from the cache hit by the data read address;

[0104] If not, it is determined that the page data corresponding to the data read address is compressed page data.

[0105] The cache of the graphics processor may be a first-level cache (L1 cache) or a second-level cache (L2 cache) or a compressed data cache of the graphics processor.

[0106] When the data read address hits the L1 cache and / or the L2 cache, it indicates that the page data corresponding to the data read address is not compressed, and the page data can be directly accessed in the L1 cache and / or the L2 cache to execute the data read request.

[0107] The compressed data cache can store a certain amount of compressed page data. When the data read address misses the first-level cache and the second-level cache but hits the compressed data cache, it indicates that the compressed page data is stored in the compressed data cache. Therefore, the compressed page data can be quickly read from the compressed data cache, avoiding indexing the graphics processor memory and the requested non-graphics processor memory according to the address of the compressed page data. There is no need to access the TLB and page table of the graphics processor memory, thereby reducing the access overhead of the graphics processor memory and the non-graphics processor memory, thereby improving the data reading efficiency of the graphics processor.

[0108] When the compressed data cache is missed, the physical address of the requested non-GPU memory is calculated using the data read address, storage address, and base address in the manner described above, and the compressed page data is then read according to the requested non-GPU memory physical address and the data read address.

[0109] The embodiment of the present invention also provides a graphics processor. Figure 3 , the graphics processor 1 includes:

[0110] The compressor 11 is configured to obtain page data to be compressed and compress the page data to be compressed using a predetermined target compression algorithm to obtain compressed page data; a target compression rate of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory;

[0111] a memory controller 12 configured to store, in the compressed page data, page data having a data amount equal to the storable data amount in the graphics processor memory, and store page data having a data amount exceeding the storable data amount in an applied non-graphics processor memory, and record the storage address of the page data in the applied non-graphics processor memory;

[0112] The applied non-GPU memory is a memory applied for in a processor other than the GPU, and the storage capacity of the applied non-GPU memory is determined based at least on the target compression ratio.

[0113] Optionally, the memory controller 12 is further configured to, based on the data read address of the acquired data read request, acquire, when it is determined that the page data corresponding to the data read address is compressed page data and that the compressed page data is stored in the graphics processor memory 14 and the requested non-graphics processor memory 21, acquire, from the graphics processor memory 14 based on the data read address, page data having a data volume equal to a data volume that can be stored in the graphics processor memory, and acquire, from the requested non-graphics processor memory 21 based on the storage address in the data read request, page data having a data volume exceeding the data volume that can be stored;

[0114] The compressor 11 is further configured to determine a decompression algorithm for the compressed page data using a target compression algorithm, and obtain decompressed page data;

[0115] The graphics processor 1 further includes:

[0116] The data reading module is used to read data from the decompressed page data.

[0117] Optionally, the graphics processor 1 further includes:

[0118] A compression ratio comparison module is configured to determine an actual compression ratio corresponding to the compressed page data and compare the actual compression ratio with the target compression ratio; the actual compression ratio can be used to calculate an actual compressed data volume of the compressed page data;

[0119] If the actual compression ratio is less than the target compression ratio, determining that the actual compressed data amount is greater than the storable data amount, storing, among the compressed page data, page data having a data amount equal to the storable data amount in the graphics processor memory, storing, among the compressed page data, page data having a data amount exceeding the storable data amount in the requested non-graphics processor memory, and recording the storage addresses of the page data in the requested non-graphics processor memory;

[0120] If the actual compression ratio is greater than or equal to the target compression ratio, it is determined that the actual compressed data amount is less than or equal to the storable data amount, and the compressed page data is stored in the graphics processor memory.

[0121] Optionally, the graphics processor 1 further includes:

[0122] A performance analysis module is configured to obtain debug page data; compress the debug page data multiple times using a plurality of candidate compression algorithms; determine, for each candidate compression algorithm, a compression ratio for each compression of the debug page data using the candidate compression algorithm, and obtain an average compression ratio of the debug page data based on the compression ratio determined each time; select the candidate compression algorithm corresponding to the maximum value among the average compression ratios as a target compression algorithm, and use the average compression ratio of the target compression algorithms as a target compression ratio.

[0123] The performance analysis module can provide a target compression algorithm and a target compression ratio.

[0124] For example, the execution process of the performance analysis module may include:

[0125] First, the performance analysis module will continuously scan the debug page data generated by the GPU calculation and temporarily stored in the GPU memory system during the current performance analysis phase;

[0126] Then, these debugging page data will be continuously transmitted to the compressor 11, which integrates a variety of mainstream compression algorithms. At this time, all compression algorithms are used for compression, and the compression ratio of each compression algorithm is calculated.

[0127] Finally, the compression ratios from multiple iterations are averaged across all compression algorithm dimensions. The compression algorithm with the smallest average compression ratio is selected as the actual compression algorithm for this application, or the target compression algorithm. The corresponding average compression ratio is rounded and written to the target compression ratio register for subsequent use.

[0128] Optionally, the graphics processor 1 further includes:

[0129] a memory request module, configured to obtain an available capacity of the graphics processor memory; determine a requested capacity of the non-graphics processor memory based on the available capacity, as the requested non-graphics processor memory; the requested capacity is the storage capacity of the non-graphics processor memory; wherein the requested capacity is determined by multiplying the available capacity by a difference between the target compression ratio and a first value;

[0130] The target compression rate register 121 is used to store the target compression rate.

[0131] The memory application module can be triggered by the performance analysis module to execute the application of non-GPU memory, that is, the expansion system initialization process.

[0132] The implementation process of the memory application module includes:

[0133] First, detect and obtain the actual parameters of the current host, including:

[0134] 1) Available GPU DRAM capacity;

[0135] 2) Available capacity of CPU DRAM;

[0136] 3) The system environment variable configured by the upper user is the upper threshold of the memory capacity that the memory application module can occupy in the CPU DRAM;

[0137] Next, the actual detected parameter configuration and the data stored in the target compression rate register (target compression rate) are used as input to finalize the final initialization parameters of the current GPU:

[0138] 1) The final target compression ratio;

[0139] 2) The memory size requested on the CPU side.

[0140] For example, if the available capacity of GPU DRAM is 40GB and the final target compression ratio is x3, the memory size requested on the CPU side is 80GB (40GBx(3-1)). Alternatively, if the available capacity of GPU DRAM is 50GB and the final target compression ratio is x4, the memory size requested on the CPU side is 150GB (50GBx(4-1)).

[0141] Afterwards, using the calculated memory size requested on the CPU side, pinned memory is requested on the CPU side. Pinned memory can prevent pages from being swapped to the CPU hard disk. At the same time, the base address (first memory address) of the requested non-GPU memory is recorded in a new register on the GPU: the fallback memory base address register.

[0142] Finally, by writing the specified hardware registers, the compressed data cache, TLB and MMU on the GPU hardware are enabled, and the compressed page data is completely stored using the requested non-GPU memory and GPU memory.

[0143] Optionally, the graphics processor 1 further includes a compressed data cache 13 for storing the compressed page data.

[0144] Optionally, the graphics processor 1 further includes: a base address register 122;

[0145] The base address register 122 is used to store the base address of the requested non-GPU memory. The base address is combined with the storage address to indicate the storage location of the page data exceeding the storable data amount in the requested non-GPU memory.

[0146] As can be seen, in a graphics processor 1 provided by an embodiment of the present invention, the computing core 10 generates to-be-compressed page data, and the compressor 11 can compress the to-be-compressed page data using a predetermined target compression algorithm to obtain compressed page data; the target compression ratio of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory. Because the compressed data obtained after the compression of the to-be-compressed page data may not meet the data volume (the amount of data that can be stored) determined by the target compression ratio of the predetermined target compression algorithm, the embodiment of the present invention can, for a compressed page data, store the page data of the compressed page data that is equal to the amount of data that can be stored in the graphics processor memory 14, and store the page data of the compressed page data that exceeds the amount of data that can be stored in the already applied non-GPU memory 21, and record the storage address of the page data in the already applied non-GPU memory 21. The already applied non-GPU memory 21 is determined based at least on the target compression ratio of the target compression algorithm, so that the compressed page data can be stored completely at one time. Therefore, there are no gaps in the data stored in the GPU memory 21, and the GPU 1 can be prevented from using page reallocation to re-apply for GPU memory to store the compressed page data that exceeds the storable data amount, thereby improving the efficiency of data storage in the GPU memory 14. When the compressed page data stored in the applied non-GPU memory 21 is needed later, it can be directly obtained based on the recorded storage address. Therefore, on the basis of reducing the capacity occupancy of the GPU memory 14 and realizing the expansion of the memory capacity of the GPU 1, the additional overhead caused by the GPU 1 using page reallocation to store the compressed page data can be avoided, thereby achieving the purpose of improving the efficiency of data storage in the GPU memory 14.

[0147] Furthermore, when a computing core in a graphics processor needs to access the compressed page data, a memory controller of the graphics processor first obtains a data read address of a data read request, and then, when it is determined that the corresponding page data is compressed page data and is stored in the graphics processor memory and the applied non-graphics processor memory, since the compressed page data is compressed using a predetermined target compression algorithm, and among the compressed page data, page data with a data amount equal to the storable data amount is stored in the graphics processor memory, and page data with a data amount exceeding the storable data amount is stored in the applied non-graphics processor memory; therefore, it is necessary to read the page data from the graphics processor memory and the applied non-graphics processor memory respectively according to the data read address and the storage address in the data read address, and finally obtain the compressed page data, and then decompress the compressed page data using the target compression algorithm to obtain the to-be-compressed page data that the graphics processor needs to access, thereby achieving the expansion of the memory capacity of the graphics processor.

[0148] To facilitate the description of the functional implementation of the graphics processor provided by the embodiment of the present invention, please refer to Figure 4 , Figure 4 The present invention provides a data storage method according to an embodiment of the present invention.

[0149] Figure 4 In the example, the following is taken to illustrate the compression of 4 pages of data to be compressed: PageA, PageB, PageC, and PageD, with a target compression rate of x4.

[0150] like Figure 4 As shown, each page data to be compressed is compressed by a target compression algorithm predetermined in the compressor to obtain compressed page data. At the same time, each page data to be compressed generates an actual compression ratio:

[0151] The actual compression ratio of PageA after compression: x1;

[0152] The actual compression ratio of PageB after compression is x1.33;

[0153] The actual compression ratio of PageC after compression: x2;

[0154] The actual compression ratio after PageD compression is x4.

[0155] Then, the actual compression ratio of each page to be compressed is compared with the target compression ratio, and each compressed data that is less than the target compression ratio x4 is stored in the GPU memory and the non-GPU memory (i.e. Figure 4 CPU backup memory in ).

[0156] like Figure 4 As shown, the actual compression ratio of PageA, PageB, and PageC is less than the target compression ratio. Therefore, the compressed page data whose data volume is equal to the data volume determined by the target compression ratio (the amount of data that can be stored) is stored in the GPU memory, and the compressed page data whose data volume exceeds the amount of data that can be stored is stored in the CPU backup memory (the non-GPU memory that has been applied for).

[0157] The actual compression ratio of PageD is equal to the target compression ratio, so the compressed page data can be directly stored in the GPU memory.

[0158] The above describes multiple embodiment schemes provided by the embodiments of the present invention. The various optional methods introduced in each embodiment scheme can be combined and cross-referenced with each other without conflict, thereby extending a variety of possible embodiment schemes, which can all be considered as embodiment schemes disclosed and open in the embodiments of the present invention.

[0159] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope defined by the claims.

Claims

1. A data storage method, characterized in that: Applicable to graphics processors, including: Get the page data to be compressed; Compressing the to-be-compressed page data using a predetermined target compression algorithm to obtain compressed page data; a target compression rate of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory; storing, in the compressed page data, page data having a data volume equal to the storable data volume in the graphics processor memory, and storing page data having a data volume exceeding the storable data volume in the applied non-graphics processor memory, and recording the storage addresses of the page data in the applied non-graphics processor memory; The applied non-GPU memory is a memory applied for in a processor other than the GPU, and the storage capacity of the applied non-GPU memory is determined based at least on the target compression ratio.

2. The data storage method according to claim 1, wherein: After the step of obtaining the compressed page data, the data storage method further includes: Determine a corresponding actual compression ratio based on the compressed page data; the actual compression ratio can be used to calculate the actual compressed data volume of the compressed page data; comparing the actual compression ratio with the target compression ratio; If the actual compression ratio is less than the target compression ratio, determining that the actual compressed data amount is greater than the storable data amount, and executing the step of storing, in the compressed page data, page data having a data amount equal to the storable data amount into the graphics processor memory, storing, in addition to the page data having a data amount exceeding the storable data amount into the applied non-graphics processor memory, and recording the storage addresses of the page data in the applied non-graphics processor memory; If the actual compression ratio is greater than or equal to the target compression ratio, it is determined that the actual compressed data amount is less than or equal to the storable data amount, and the compressed page data is stored in the graphics processor memory.

3. The data storage method according to claim 1, wherein: Before the step of obtaining the page data to be compressed, the method further includes: Get debug page data; compressing the debugging page data multiple times using multiple candidate compression algorithms; For each candidate compression algorithm, determining a compression ratio of the debugging page data each time the candidate compression algorithm is used to compress the debugging page data, and obtaining an average compression ratio of the debugging page data based on the compression ratio determined each time; The candidate compression algorithm corresponding to the maximum value of the average compression ratios is selected as the target compression algorithm, and the average compression ratio of the target compression algorithm is used as the target compression ratio.

4. The data storage method according to claim 3, wherein: After the step of selecting the compression algorithm corresponding to the maximum value among the average compression ratios as the target compression algorithm and taking the average compression ratio of the target compression algorithm as the target compression ratio, the method further includes: Obtaining the available capacity of the graphics processor memory; Determining a requested capacity on the non-GPU memory based on the available capacity as the requested non-GPU memory, where the requested capacity is the storage capacity of the non-GPU memory; The requested capacity is determined by multiplying the available capacity by the difference between the target compression rate and the first value.

5. The data storage method according to claim 4, wherein: The requested non-GPU memory is page-locked memory.

6. The data storage method according to claim 4, wherein: After the step of using the average compression rate of the target compression algorithm as the target compression rate, the method further includes: saving the target compression ratio; Furthermore, a base address of the applied non-GPU memory is saved, wherein the base address is combined with the storage address to indicate a storage location of the page data exceeding the storable data amount in the applied non-GPU memory.

7. The data storage method according to claim 2, wherein: The actual compression ratio and the storage address are recorded in a page table entry of the compressed page data; the page table entry also records a page compression mark of the compressed page data; The page compression mark is used to indicate whether the compressed page data has been compressed; the actual compression ratio is the ratio of the uncompressed data volume of the to-be-compressed page data to the actual compressed data volume of the compressed page data.

8. The data storage method according to any one of claims 2 to 6, wherein: After the step of compressing the to-be-compressed page data using a predetermined target compression algorithm to obtain compressed page data, and before the step of determining a corresponding actual compression ratio based on the compressed page data, the method further includes: The compressed page data is stored in a compressed data cache of the graphics processor.

9. A data reading method, characterized in that: Applicable to graphics processors, including: Obtaining a data read request, wherein the data read request includes a data read address; When it is determined that the page data corresponding to the data read address is compressed page data, and the compressed page data is stored in the graphics processor memory and the applied non-graphics processor memory, based on the data read address, page data having a data volume equal to a data volume storable in the graphics processor memory is obtained, and based on the storage address in the data read request, page data having a data volume exceeding the storable data volume is obtained in the applied non-graphics processor memory, thereby obtaining compressed page data; the compressed page data is data stored by the data storage method according to any one of claims 1 to 8; the storable data volume is determined by a target compression ratio of a target compression algorithm adopted by the data storage method; Determining a decompression algorithm for the compressed page data using a target compression algorithm, and obtaining decompressed page data; Data is read from the decompressed page data.

10. The data reading method according to claim 9, wherein: After the step of obtaining the data read request, the method further includes: When it is determined that the page data corresponding to the data read address is compressed page data, determining whether a page table entry of the compressed page data corresponding to the data read address records a storage address of a non-GPU memory that has been applied for; If yes, determining that the compressed page data corresponding to the data read address is stored in the graphics processor memory and the applied non-graphics processor memory; If not, it is determined that the compressed page data corresponding to the data read address is stored in the graphics processor memory.

11. The data reading method according to claim 10, wherein: When it is determined that the page table entry of the compressed page data corresponding to the data read address records the storage location of the non-GPU memory that has been applied for, the method further includes: Obtaining the base address of the requested non-GPU memory; A physical address of the non-GPU memory is obtained based on the base address and the storage address.

12. The data reading method according to claim 11, wherein: The method of acquiring page data having a data amount equal to a storable data amount in the graphics processor memory based on the data read address, and acquiring page data exceeding the storable data amount in the applied non-graphics processor memory based on the storage address in the data read request, to obtain compressed page data, includes: In the graphics processor memory, acquiring page data having a data amount equal to a storable data amount according to the data read address; Acquire, from the applied non-GPU memory, page data corresponding to the physical address of the non-GPU memory, wherein the acquired compressed page data is data exceeding the storable data amount; Compressed page data is obtained based on the page data obtained in the graphics processor memory and the page data obtained in the requested non-graphics processor memory.

13. The data reading method according to claim 12, wherein: After the step of obtaining the data read request, the method further includes: determining whether the data read address hits the cache of the graphics processor; If yes, determining that the page data corresponding to the data read address is not compressed, and reading the data requested by the data read request from the cache hit by the data read address; If not, it is determined that the page data corresponding to the data read address is compressed page data.

14. A graphics processor, characterized in that: include: a compressor configured to obtain page data to be compressed and compress the page data using a predetermined target compression algorithm to obtain compressed page data; a target compression rate of the target compression algorithm is used to determine the amount of data that can be stored in the graphics processor memory; a memory controller configured to store, in the compressed page data, page data having a data amount equal to the storable data amount into the graphics processor memory, and store page data having a data amount exceeding the storable data amount into an applied non-graphics processor memory, and record the storage address of the page data in the applied non-graphics processor memory; The applied non-GPU memory is a memory applied for in a processor other than the GPU, and the storage capacity of the applied non-GPU memory is determined based at least on the target compression ratio.

15. The graphics processor according to claim 14, wherein: The memory controller is further configured to, based on the data read address of the acquired data read request, acquire, when it is determined that the page data corresponding to the data read address is compressed page data and the compressed page data is stored in the graphics processor memory and the requested non-graphics processor memory, acquire, based on the data read address, page data of a data volume equal to a data volume storable in the graphics processor memory, and acquire, based on the storage address in the data read request, page data of a data volume exceeding the data volume storable in the requested non-graphics processor memory; The compressor is further configured to determine a decompression algorithm for the compressed page data using a target compression algorithm, and obtain decompressed page data; The graphics processor further includes: The data reading module is used to read data from the decompressed page data.

16. The graphics processor according to claim 15, wherein: Also includes: A compression ratio comparison module, configured to determine a corresponding actual compression ratio based on the compressed page data; and comparing the actual compression ratio with the target compression ratio; The actual compression ratio can be used to calculate the actual compressed data volume of the compressed page data; If the actual compression ratio is less than the target compression ratio, determining that the actual compressed data amount is greater than the storable data amount, storing, among the compressed page data, page data having a data amount equal to the storable data amount in the graphics processor memory, storing, among the compressed page data, page data having a data amount exceeding the storable data amount in the requested non-graphics processor memory, and recording the storage addresses of the page data in the requested non-graphics processor memory; If the actual compression ratio is greater than or equal to the target compression ratio, it is determined that the actual compressed data amount is less than or equal to the storable data amount, and the compressed page data is stored in the graphics processor memory.

17. The graphics processor according to claim 16, wherein: Also includes: Performance analysis module, used to obtain debugging page data; The debugging page data is compressed multiple times using a plurality of candidate compression algorithms; for each candidate compression algorithm, a compression ratio of the debugging page data compressed each time using the candidate compression algorithm is determined, and an average compression ratio of the debugging page data is obtained based on the compression ratio determined each time; the candidate compression algorithm corresponding to the maximum value of the average compression ratio is selected as the target compression algorithm, and the average compression ratio of the target compression algorithm is used as the target compression ratio.

18. The graphics processor according to claim 17, wherein: Also includes: A memory application module, configured to obtain the available capacity of the graphics processor memory; Determine the requested capacity of the non-GPU memory based on the available capacity as the requested non-GPU memory; The requested capacity is the storage capacity of the non-GPU memory; wherein the requested capacity is determined by multiplying the available capacity by the difference between the target compression rate and the first value; The target compression ratio maintaining module is used to store the target compression ratio.

19. The graphics processor according to any one of claims 15 to 18, wherein: Also includes: The compressed data cache is used to store the compressed page data.

20. The graphics processor according to claim 19, wherein Also includes: Target compression rate register and base address register; The target compression rate register is used to store the target compression rate; The base address register is used to store the base address of the applied non-GPU memory. The base address is combined with the storage address to indicate the storage location of the page data exceeding the storable data amount in the applied non-GPU memory.

Citation Information

Patent Citations

  • Data processing method, device and equipment and readable storage medium

    CN113467958A

  • Adaptive data compression method, system, device and product for database

    CN118838879A

  • Techniques for an efficient fabric attached memory

    US20210133123A1