Data transmission method and device, graphic processing unit, storage medium and terminal equipment
By introducing a decompression synthesis module into the graphics processing unit, the data in the GPU global buffer is dynamically managed, and the problem of high capacity miss rate is solved, which improves GPU performance and improves the supply of GPC bandwidth.
Patent Information
- Application Number
- CN202510122571.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-17
AI Technical Summary
In the prior art, the global buffering of GPUs faces increasing data pressure, resulting in high capacity misses and high cost, making it difficult to effectively solve this problem.
The decompression synthesis module is introduced in the graphics processing unit. By acquiring the compressed data in the global buffer and the updated data in the graphics processor cluster, the decompression, fusion and recompression process is carried out to dynamically manage the data to reduce the capacity miss rate.
By dynamically managing data, the capacity miss rate is reduced, the performance of GPU is improved, and through the compression transmission mechanism, the supply of GPC bandwidth is greatly improved, solving the problem of GPU performance decline.
Smart Images

Figure CN120163701A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of graphics processing technology, and in particular to a data transmission method and apparatus, a graphics processing unit, a storage medium, and a terminal device. Background Art
[0002] With the continuous improvement of process technology and performance requirements of graphics processing units (GPUs), today's GPUs begin to include more and more graphics processor clusters (Graphics Processor Clusters, GPCs) and internal computing units (Process Elements, PEs), which can also be called processing units. In order to provide sufficient data supply to the computing units, the global buffer faces increasing data pressure, more data lines and higher data bandwidth requirements, which brings more complex data exchange structure (crossbar) design to ensure the collaboration between various modules; at the same time, it also brings great challenges to the back-end timing and frequency.
[0003] In the prior art, the miss rate of the overall GPC internal data is reduced by increasing the internal cache capacity of the GPC.
[0004] However, the existing technical solutions will lead to higher GPU costs, and capacity miss is still a technical problem that needs to be solved urgently. Summary of the invention
[0005] This application can reduce the capacity miss rate, realize dynamic management of data during compression and decompression, and improve the performance of GPU.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] In a first aspect, the present application provides a data transmission method for a graphics processing unit, wherein the graphics processing unit includes a global buffer, a graphics processor cluster, and a decompression and synthesis module, and the data transmission method includes: the decompression and synthesis module obtains first compressed data from the global buffer and updates data from the graphics processor cluster; the decompression and synthesis module decompresses the first compressed data and merges the decompressed data with the updated data to obtain merged data; the decompression and synthesis module compresses the merged data and writes it into the global buffer.
[0008] Optionally, the graphics processing unit includes a global buffer in which compressed data and / or uncompressed data is stored. The decompression and composition module obtains first compressed data from the global buffer, including: the decompression and composition module sends a first read request to the global buffer; in response to the first read request, the global buffer returns the first compressed data corresponding to the first read request.
[0009] Optionally, the global buffer returns the first compressed data corresponding to the first read request, including: the global buffer performs a hit test according to the data address in the first read request; if a hit occurs, the global buffer directly returns the first compressed data corresponding to the first read request; if no hit occurs, the global buffer forwards the read request to the dynamic memory controller to obtain the first compressed data corresponding to the first read request from the memory and return it.
[0010] Optionally, the graphics processing unit includes a graphics processor cluster. The decompression and composition module obtains updated data from the graphics processor cluster, including: the decompression and composition module obtains second compressed data of the updated data from the graphics processor cluster; the decompression and composition module decompresses the second compressed data to obtain the updated data.
[0011] Optionally, the graphics processing unit includes a global buffer and a graphics processor cluster. The decompression and composition module obtains updated data from the graphics processor cluster, including: the decompression and composition module obtains the updated data from the graphics processor cluster, and the updated data is the operation result of the graphics processor cluster.
[0012] Optionally, the graphics processor cluster calculates the updated data in the following manner: sending a second read request to the global buffer; in response to the second read request, the global buffer returns the second compressed data corresponding to the second read request; decompressing the second compressed data and performing an operation to obtain the updated data.
[0013] Optionally, the graphics processor cluster includes a processing unit and a fixed function cluster. The processing unit executes shader tasks, and the fixed function cluster executes rasterization tasks, depth test tasks, or color operation tasks. The updated data is the operation result of the processing unit and / or the fixed function cluster.
[0014] Optionally, the decompression and composition module compresses and sends out the fused data, including: the decompression and composition module sends the compressed fused data to the global buffer, where data with the same compression ratio and the same format is stored in the same cache line.
[0015] Optionally, the sizes of the third compressed data obtained by compressing different fused data are the same.
[0016] In a second aspect, the present application discloses a graphics processing unit, which includes: a global buffer for storing compressed data and / or uncompressed data; a graphics processor cluster and a decompression and synthesis module for performing the steps of the data transmission method.
[0017] Optionally, the graphics processing unit further includes: a memory for storing compressed data and / or uncompressed data; a dynamic memory controller for forwarding a first read request to the dynamic memory controller when the global buffer misses.
[0018] In a third aspect, the present application provides a terminal device, which is characterized by including the graphics processing unit described in the second aspect. Alternatively, the terminal device includes a memory and a processor, and a computer program is stored on the memory, and the processor executes the steps of the data transmission method described in the first aspect
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program executes the steps of the data transmission method when being run by a processor.
[0020] Compared with the prior art, the technical solution of the present application has the following beneficial effects:
[0021] In the technical solution of the present application, the decompression and synthesis module in the graphics processing unit obtains the first compressed data and the updated data; decompresses the first compressed data, and fuses the decompressed data with the updated data to obtain fused data; compresses and sends out the fused data. In the technical solution of the present application, by setting a decompression and synthesis module in the graphics processing unit, when some data in the first compressed data is updated, the data can be updated by obtaining the updated data and performing corresponding data fusion, so as to realize the dynamic management of data during the compression and decompression processes, and further improve the performance of the GPU. In addition, in the technical solution of the present application, the interactive data is in a compressed form, which can greatly improve the supply of GPC bandwidth, reduce the capacity miss rate, and implement a lossless compression transmission mechanism, thus effectively solving the problem of the decline in GPU performance caused by the large-scale expansion of the number of PEs and GPCs.
[0022] Furthermore, in the technical solution of the present application, data with the same compression ratio and the same format is stored in the same cache line. By storing multiple data with the same compression ratio in the same cache line, the storage space in the global buffer can be utilized more effectively, thereby greatly improving the capacity of the global buffer and further improving the performance of the GPU. Description of the Drawings
[0023] Figure 1 is a structural diagram of a graphics processing unit provided by an embodiment of the present application;
[0024] Figure 2 is a flowchart of a data transmission method provided by an embodiment of the present application;
[0025] Figure 3 is a schematic diagram of the arrangement of compressed data provided by an embodiment of the present application;
[0026] Figure 4 is a schematic diagram of a data reading process provided by an embodiment of the present application;
[0027] Figure 5 is a schematic diagram of a data writing process provided by an embodiment of the present application;
[0028] Figure 6 is a structural diagram of a data transmission device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0029] As described in the background art, a large number of GPR resources are occupied in the prior art, resulting in a reduction in the number of warps for parallel operation of the GPU and a decrease in GPU performance.
[0030] In the technical solution of the present application, in the case where some data in the first compressed data is updated, the data can be updated by obtaining the updated data and performing corresponding data fusion, so as to realize the dynamic management of data during the compression and decompression processes, and further improve the performance of the GPU. In addition, in the technical solution of the present application, the interactive data is in a compressed form, which can greatly improve the supply of GPC bandwidth, reduce the capacity miss rate, and realize a lossless compression transmission mechanism, thereby effectively solving the problem of the decrease in GPU performance caused by the larger-scale expansion of the number of PEs and GPCs.
[0031] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of the specific embodiments of the present application will be given with reference to the accompanying drawings.
[0032] Refer to Figure 1 , Figure 1 which shows the structure of a graphics processing unit of the present application.
[0033] As Figure 1 shown, the graphics processing unit includes multiple Graphics Processor Clusters (GPCs) 101, a global buffer 102, a Dynamic Memory Controller (DMC) 103, and a memory 104.
[0034] Each graphics processor cluster 102 includes cores, and each core includes a plurality of processing units (Processor Element, PE) 1011. Each processing unit 1011 includes a plurality of general-purpose processing units, texture, and data read / write units, etc. Among them, the processing unit 1011 executes shader tasks.
[0035] Each graphics processor cluster 102 may further include a fixed function cluster (Fix function Cluster, FUC) 1012. Among them, the fixed function cluster 1012 executes rasterization tasks, depth test tasks, or color operation tasks.
[0036] Different from the GPUs in the prior art, in this embodiment, the processing unit 1011, the fixed function cluster FUC 1012, and the global buffer 102 all have compression / decompression units. The compression / decompression unit can compress / decompress data. Correspondingly, the global buffer 102 and the memory 104 may store compressed data and / or uncompressed data.
[0037] Specifically, the processing unit 1011 includes a compression / decompression unit 10111. The compression / decompression unit 10111 can decompress the data read by the processing unit 1011 and compress the operation results processed by the processing unit 1011 and then send them out.
[0038] The fixed function cluster FUC 1012 includes a compression / decompression unit 10121. The compression / decompression unit 10121 can decompress the data read by the fixed function cluster FUC 1012 and compress the operation results processed by the fixed function cluster FUC 1012 and then send them out.
[0039] The global buffer 102 includes a compression / decompression unit 1021. The compression / decompression unit 1021 can compress the data from the GPC and send it to the DMC 103. The compression / decompression unit 1021 can decompress the data from the DMC 103 (i.e., the memory 104).
[0040] It should be noted that the core in this embodiment may also be referred to as a kernel, a compute unit (CU), etc., and this application does not limit this.
[0041] From Figure 1 As can be seen from the architecture of the GPU shown, the data interacted by each module can be compressed data or uncompressed data, and the embodiments of the present invention can manage the change of the compressed form of data.
[0042] Specifically, please refer toFigure 1 and Figure 2 , the data transmission method can be run by the GPU, that is, each step of the data transmission method is executed by the GPU. Specifically, each step of the above method can be executed by a functional module in the GPU (such as Figure 1 the decompression and synthesis module 105 shown). Of course, it can be understood that the data transmission method can also be executed by any other appropriate entity.
[0043] Specifically, the data transmission method can specifically include the following steps:
[0044] Step 201: The decompression and synthesis module 105 obtains the first compressed data from the global buffer 201 and obtains the updated data from the graphics processor cluster 101;
[0045] Step 202: The decompression and synthesis module 105 decompresses the first compressed data and fuses the compressed data with the updated data to obtain the fused data;
[0046] Step 203: The decompression and synthesis module 105 compresses the fused data and sends it out.
[0047] It can be understood that in specific implementation, each step of the above method can be implemented in the form of a software program, and the software program runs in a processor integrated inside the chip or chip module. This method can also be implemented in the form of software combined with hardware, and the present application does not make any restrictions.
[0048] In this embodiment, each step of the above method can be executed during the data compression process, specifically, it can be executed during the data reading process or during the data writing process. The above method can realize the re-synthesis of data when part of the data in the compressed data in the cache is updated during the compressed storage and transmission.
[0049] Taking the execution entity as the decompression and synthesis module 105 as an example, the decompression and synthesis module 105 reads the first compressed data (that is, the data to be updated) and the updated data. The decompression and synthesis module 105 decompresses and expands the first compressed data, and fuses the decompressed data with the updated data to obtain the fused data. The compression / decompression unit 1021 compresses the fused data and writes it into the global buffer 201.
[0050] Furthermore, the global buffer 201 can forward and write the above compressed data into the memory 104 via the DMC 103.
[0051] Further, after the fused data is compressed, the compressed data has a compression flag, which can indicate whether the data is compressed or uncompressed. Through this compression flag, when the global buffer 102 performs cache hit, it can know whether the data stored in the current cache line is in a compressed state or an uncompressed state, so as to know the address coverage range of the current cache line. Through this compression flag, the global buffer 102, the processing unit 1011, and the fixed function cluster FUC1012 can also know whether it is necessary to perform a decompression operation on the data.
[0052] In a non-limiting embodiment, the global buffer 102 stores compressed data, and part of the data of the compressed data is updated. In this case, the decompression synthesis module 105 sends a first read request to the global buffer 102. In response to the first read request, the global buffer 102 returns the first compressed data corresponding to the first read request.
[0053] Further, the first read request includes a data address, which points to the storage space in the global buffer. The global buffer 102 first performs a hit test according to the data address in the first read request. If a hit occurs, indicating that the global buffer 102 stores the corresponding compressed data, the global buffer 102 directly returns the first compressed data corresponding to the first read request.
[0054] If no hit occurs, indicating that the global buffer 102 does not store the corresponding compressed data, the global buffer 102 forwards the read request to the dynamic memory controller 103 to obtain the first compressed data corresponding to the first read request from the memory 104 and return it.
[0055] In a non-limiting embodiment, the updated data can come from the graphics processor cluster 101. Specifically, it can be the processing unit 1011 or the fixed function cluster FUC1012.
[0056] In this embodiment, the decompression synthesis module 105 can obtain the updated data from the graphics processor cluster 101. The updated data is the operation result of the graphics processor cluster 101, specifically, it can be the operation result of the processing unit 1011 and / or the fixed function cluster FUC1012.
[0057] In a specific example, the processing unit 1011 sends a second read request to the global buffer 102. In response to the second read request, the global buffer 102 returns the second compressed data corresponding to the second read request. The compression / decompression unit 10111 decompresses the second compressed data. The processing unit 1011 performs a shader task operation on the decompressed data to obtain the updated data.
[0058] In another specific example, the fixed function cluster FUC1012 sends a second read request to the global buffer 102. In response to the second read request, the global buffer 102 returns the second compressed data corresponding to the second read request. The compression / decompression unit 10121 decompresses the second compressed data. The fixed function cluster FUC1012 performs rasterization tasks, depth test tasks, or color operation tasks on the decompressed data to obtain updated data.
[0059] In a non-limiting embodiment, data with the same compression ratio and the same format is stored in the same cache line. A cache line is the smallest storage unit corresponding to the storage space allocated by the global buffer 102. Wherein, the format of the data refers to the format of the data resource.
[0060] As Figure 3 shown, the compressed data 1 and the compressed data 2 have the same compression ratio and the same format, and the compressed data 1 and the compressed data 2 are stored in the same cache line. The compressed data 3 to the compressed data 6 have the same compression ratio and the same format, and the compressed data 3 to the compressed data 6 are stored in the same cache line.
[0061] Specifically, when the global buffer 102 performs a hit test, it will determine the address range of the cache line according to the arrangement of the compressed data. For example, the address range of the cache line where the compressed data 1 and the compressed data 2 are located may be different from the address range of the cache line where the compressed data 3 to the compressed data 6 are located.
[0062] In the embodiment of the present invention, by storing multiple data with the same compression ratio in the same cache line, the storage space in the global buffer can be utilized more effectively, thereby greatly improving the capacity of the global buffer and further improving the performance of the GPU.
[0063] Please refer to Figure 4 , Figure 4 which shows a data reading process, that is, the process of reading data from the memory.
[0064] In this embodiment, the memory 104 stores compressed data and uncompressed data, which may specifically include compressed data 1, compressed data 2, and compressed data 3, uncompressed data 1 and uncompressed data 2.
[0065] Step 1: The GPU starts running the application program, and then the driver generates a hardware command list for the GPU.
[0066] Step 2: The GPU receives the hardware command list, and then the scheduler starts the GPU task to the local allocator.
[0067] Step 3: The graphics processor cluster locally allocates tasks and then initiates the tasks to the processing unit 1011.
[0068] Step 4: The processing unit 1011 starts to read data and execute the front-end shader tasks.
[0069] Step 5: During the task execution, the processing unit 1011 sends a read request to the global buffer 102 for data extraction (such as instructions, textures, constants, etc.). If the result is hit, the global buffer 102 returns the data to the processing unit 1011 in its original compressed / decompressed state.
[0070] Step 6: If not hit, the read request of the processing unit 1011 will be forwarded to the DMC 103 to obtain data from the memory 104.
[0071] Step 7: The DMC 103 extracts data from the memory 104 based on the metadata information and returns it to the global buffer 102.
[0072] Step 8: The global buffer 102 returns the data to the processing unit 1011 in its original compressed / decompressed state.
[0073] Step 9: Before sending the data to the computing unit, the processing unit 1011 will decode the data into an uncompressed format through the compression / decompression unit 10111.
[0074] Step 10: The processing unit 1011 completes the front-end shader tasks.
[0075] Step 11: The fixed function cluster FUC 1012 starts to work and executes tasks such as rasterization, depth testing, or color operations.
[0076] Step 12: During the task execution, the fixed function cluster FUC 1012 sends a read request to the global buffer 102 for data extraction (such as instructions, textures, constants, etc.). If the result is hit, the global buffer 102 returns the data to the fixed function cluster FUC 1012 in its original compressed / decompressed state.
[0077] Step 13: If not hit, the read request of the fixed function cluster FUC 1012 will be forwarded to the DMC 103 to obtain data from the memory 104.
[0078] Step 14: The DMC 103 extracts data from the memory 104 based on the metadata information and returns it to the global buffer 102.
[0079] Step 15: The global buffer 102 returns the data to the fixed function cluster FUC 1012 in its original compressed / decompressed state.
[0080] Step 16. Before sending the data to the computing unit, the fixed function cluster FUC1012 decodes the data into an uncompressed format through the compression / decompression unit 10121.
[0081] Step 17. The fixed function cluster FUC1012 completes tasks such as rasterization, depth testing, or color operations.
[0082] During Steps 5 and 8, for compressed data, the global buffer 102 can also decompress the data through the compression / decompression unit 1021 and then send the decompressed data to other modules. The operation results obtained in Step 10 and / or Step 17 can be the updated data in the above data transmission method process.
[0083] Please refer to Figure 5 , Figure 5 which shows a data writing process, that is, the process of writing data into the memory.
[0084] Continuing to refer to the foregoing embodiment, in Step 10, the operation result of the processing unit 1011 completing the front-end shader task is obtained. In Step 17, the operation result of the fixed function cluster FUC1012 completing tasks such as rasterization, depth testing, or color operations is obtained. The following writing process is the process of writing the above operation results into the memory 104.
[0085] Taking the writing of the operation result of the processing unit 1011 into the memory 104 as an example, the specific steps are as follows:
[0086] Step 18. The processing unit 1011 compresses the data into a compressed format through the compression / decompression unit 10111.
[0087] Step 19. Send the compressed data to the global buffer 102, and the global buffer 102 allocates a cache line for the compressed data and stores it.
[0088] Step 20. The global buffer 102 can also forward the data to the DMC103, and store the compressed data in the memory 104 through the DMC103.
[0089] Similarly, taking the writing of the operation result of the fixed function cluster FUC1012 into the memory 104 as an example, the specific steps are as follows:
[0090] Step 21. The fixed function cluster FUC1012 compresses the data into a compressed format through the compression / decompression unit 10121.
[0091] Step 22. Send the compressed data to the global buffer 102, and the global buffer 102 allocates a cache line for the compressed data and stores it.
[0092] Step 23: The global buffer 102 can also forward the data to the DMC 103, and store the compressed data in the memory 104 through the DMC 103.
[0093] In addition, the global buffer 102 can also compress and store the data from other modules, or store the compressed data in the memory 104 through the DMC 103.
[0094] For more specific implementation manners of the embodiments of the present application, please refer to the foregoing embodiments, which will not be elaborated herein.
[0095] Please refer to Figure 6 , Figure 6 which shows a data transmission device 60. The data transmission device 60 includes:
[0096] An acquisition module 601, configured to acquire first compressed data and updated data;
[0097] A fusion module 602, configured to decompress the first compressed data, and fuse the decompressed data with the updated data to obtain fused data;
[0098] A compression module 603, configured to compress the fused data and send it out.
[0099] In the embodiments of the present invention, when some data in the first compressed data is updated, the data can be updated by acquiring the updated data and performing corresponding data fusion, so as to realize the dynamic management of data during the compression and decompression processes, and further improve the performance of the GPU. In addition, in this embodiment, the interactive data is in a compressed form, which can greatly improve the supply of the GPC bandwidth, realize a lossless compression transmission mechanism, and thus effectively solve the problem of the performance degradation of the GPU caused by the large-scale expansion of the number of PEs and the number of GPCs.
[0100] Regarding each device and product described in the above embodiments, each module / unit included therein can be a software module / unit, a hardware module / unit, or can be partly a software module / unit and partly a hardware module / unit. For example, for each device and product applied to or integrated into a chip, each module / unit included therein can be implemented in the form of hardware such as circuits. Alternatively, at least some of the modules / units can be implemented in the form of a software program that runs on a processor integrated inside the chip, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a chip module, each module / unit included therein can be implemented in the form of hardware such as circuits. Different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module. Alternatively, at least some of the modules / units can be implemented in the form of a software program that runs on a processor integrated inside the chip module, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a terminal device, each module / unit included therein can be implemented in the form of hardware such as circuits. Different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal device. Alternatively, at least some of the modules / units can be implemented in the form of a software program that runs on a processor integrated inside the terminal device, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits.
[0101] An embodiment of the present application also discloses a storage medium. The storage medium is a computer-readable storage medium, on which a computer program is stored. When the computer program runs, it can execute Figure 1 the steps of the method shown in. The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc. The storage medium can also include a non-volatile memory or a non-transitory memory, etc.
[0102] An embodiment of the present application also discloses a terminal device. The terminal device includes the foregoing graphics processing unit; alternatively, the terminal device includes a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor runs the computer program, it executes the steps of the foregoing instruction compilation method.
[0103] In the embodiments of the present application, "a plurality of" means two or more.
[0104] In the embodiments of the present application, the first, second, etc. descriptions that appear are only for the purpose of indicating and distinguishing the described objects, without any order, and do not represent any special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.
[0105] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner.
[0106] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0107] In several embodiments provided in the present application, it should be understood that the disclosed methods, devices, and systems can be implemented in other ways. For example, the device embodiments described above are only illustrative; for example, the division of the units is only a logical function division, and there can be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0108] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0109] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may be physically included separately for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0110] The above-mentioned integrated unit implemented in the form of a software functional unit may be stored in a computer-readable storage medium. The above-mentioned software functional unit stored in a storage medium includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute some steps of the methods described in various embodiments of the present application.
[0111] Although the present application is disclosed as above, the present application is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. A data transmission method for a graphics processing unit, characterized in that: The graphics processing unit includes a global buffer, a graphics processor cluster and a decompression and synthesis module, and the data transmission method includes: The decompression and synthesis module obtains first compressed data from the global buffer and obtains update data from the graphics processor cluster; The decompression and synthesis module decompresses the first compressed data and merges the decompressed data with the update data to obtain merged data; The decompression and synthesis module compresses the fused data and writes the fused data into the global buffer.
2. The data transmission method according to claim 1, characterized in that: The global buffer stores compressed data and / or uncompressed data, and the decompression and synthesis module obtains the first compressed data from the global buffer, including: The decompression and synthesis module sends a first read request to the global buffer; In response to the first read request, the global buffer returns first compressed data corresponding to the first read request.
3. The data transmission method according to claim 2, characterized in that: The global buffer returns the first compressed data corresponding to the first read request, including: The global buffer performs a hit test according to the data address in the first read request; If a hit occurs, the global buffer directly returns the first compressed data corresponding to the first read request; If there is a miss, the global buffer forwards the read request to the dynamic memory controller to obtain the first compressed data corresponding to the first read request from the memory and return it.
4. The data transmission method according to claim 1, characterized in that: The decompression and synthesis module obtains update data from the graphics processor cluster including: The decompression and synthesis module obtains the second compressed data of the update data from the graphics processor cluster; The decompression and synthesis module decompresses the second compressed data to obtain the updated data.
5. The data transmission method according to claim 1, characterized in that: The decompression and synthesis module obtains update data from the graphics processor cluster including: The decompression and synthesis module directly obtains the update data from the graphics processor cluster, and the update data is the calculation result of the graphics processor cluster.
6. The data transmission method according to claim 5, characterized in that: The graphics processor cluster calculates the update data in the following manner: Sending a second read request to the global buffer; In response to the second read request, the global buffer returns the second compressed data corresponding to the second read request; The decompression and synthesis module decompresses the second compressed data and performs operations to obtain the updated data.
7. The data transmission method according to claim 5, characterized in that: The graphics processor cluster includes a processing unit and a fixed function cluster, the processing unit executes a shader task, the fixed function cluster executes a rasterization task, a depth test task or a color operation task, and the update data is a calculation result of the processing unit and / or the fixed function cluster.
8. The data transmission method according to claim 1, characterized in that: The decompression and synthesis module compresses and sends out the fused data, including: The decompression and synthesis module sends the compressed fused data to the global buffer, wherein data with the same compression rate and the same format are stored in the same cache line.
9. The data transmission method according to any one of claims 1 to 8, characterized in that: The sizes of the third compressed data obtained after compressing different fused data are consistent.
10. A graphics processing unit, characterized in that: include: A global buffer for storing compressed data and / or uncompressed data; Graphics processor cluster; as well as, A decompression and synthesis module, used to execute the steps of the data transmission method described in any one of claims 1 to 9.
11. The graphics processing unit according to claim 10, characterized in that: Also includes: A memory for storing compressed data and / or uncompressed data; The dynamic memory controller is used to forward the first read request to the dynamic memory controller when the global buffer misses.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the data transmission method according to any one of claims 1 to 9 are executed.
13. A terminal device, characterized in that: It comprises a memory and a processor, wherein a computer program is stored in the memory, and is characterized in that the processor executes the steps of the data transmission method described in any one of claims 1 to 9, or comprises a graphics processing unit described in claim 10 or 11.