Method and system for GPU texture compression in runtime
By using GPU for texture compression at runtime, the problem of inefficient CPU compression is solved, efficient and fast texture compression is achieved, and the needs of real-time compression is met.
Patent Information
- Application Number
- CN202510264864.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-20
AI Technical Summary
When the prior art needs to read textures from the outside in real time and compress and save memory, using CPU for texture compression will take up a lot of CPU time and cannot meet the needs of runtime texture compression.
The runtime GPU texture compression method is adopted, and the highly parallel characteristics of the GPU are used to perform compression operations through the calculation shader to achieve efficient texture compression. The specific steps include obtaining the texture file path, loading the texture file data, creating a GPU resource object, creating a calculation shader and a calculation pipeline state, and calling the calculation shader to perform compression operations.
It realizes efficient texture compression, saves memory usage, reduces compression errors, and significantly improves compression speed, meeting the requirements of real-time compression.
Smart Images

Figure CN120182079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer graphics processing, and particularly relates to a method and system for GPU texture compression during runtime. Background Art
[0002] In fields such as game development, virtual reality (VR), augmented reality (AR), and architectural design, compressing textures is a common requirement.
[0003] In the traditional development process, texture compression is usually performed using the CPU during packaging. Although the CPU compression technology has relatively high compression quality, the parallelism of the CPU is low, and the efficiency of performing the compression operation is low. This method is sufficient for textures that already exist and are being used, but in scenarios where textures need to be read from the external in real time, using the CPU for texture compression will cause a large amount of CPU time to be occupied, resulting in lag and unable to meet the requirement of texture compression during runtime.
[0004] Currently, there is still a lack of a texture compression method that can read external textures in real time and compress them to save memory. Summary of the Invention
[0005] In order to solve the problems in the prior art, the present invention aims to provide a method and system for GPU texture compression during runtime, which utilizes the highly parallel characteristics of the GPU to achieve high-efficiency texture compression and meet the requirement of reading external textures in real time and compressing them to save memory.
[0006] To achieve the above object, the present invention provides a method for GPU texture compression during runtime, including the steps of:
[0007] S1: Obtain the texture file path;
[0008] S2: Load the texture file data;
[0009] S3: Create a resource object of the GPU and upload the texture file data to the resource object;
[0010] S4: Create the compute shader of the GPU required for the compression operation and the corresponding compute pipeline state, and bind them to the GPU; the compute shader is written according to a preset compression algorithm;
[0011] S5: Call the compute shader to execute the compression operation and read back the compressed data;
[0012] S6: Save the compressed data obtained.
[0013] As an implementation manner, in the step S1, the method for obtaining the texture file path includes: user input or configuration file.
[0014] As an implementation manner, in the step S2, the file content is read from the texture file using a file reading API.
[0015] As an implementation manner, in the step S4, the compression algorithm includes the steps of:
[0016] S41: Divide the texture in the texture file into 4×4 blocks, calculate the average color value, the maximum color value and the minimum color value of each block and correlate them with each other;
[0017] S42: Pack the data of the average color value into an unsigned integer to obtain compressed color information;
[0018] S43: Convert the maximum color value and the minimum color value into lightness; select a lookup table index and a directory through the lightness; write the index data of the lookup table index and the directory and the compressed color information into a 64-bit binary string according to a preset arrangement rule.
[0019] As an implementation manner, the preset arrangement rule includes:
[0020] Write the compressed color information into the first 24-bit binary string of the 64-bit binary string; write the index data of the lookup table index into the 8-bit binary string after the compressed color information in the 64-bit binary string; write the index data of the directory into the last 32-bit binary string of the 64-bit binary string.
[0021] As an implementation manner, before the step S1, the following steps are further included:
[0022] S7: Write different versions of the compute shader according to the preset compression algorithm for different compression formats in advance;
[0023] In the step S4, select the corresponding compute shader according to the currently used compression format to complete the creation of the compute shader.
[0024] As an implementation manner, before the step S1, the following steps are further included:
[0025] S8: Upload the texture file to a management background;
[0026] S9: Audit the texture file in the management background, and store the texture file in a texture file resource library after the audit passes.
[0027] As an implementation manner, the method for obtaining the texture file path in S1 includes: downloading from the texture file resource library;
[0028] After the user downloads the texture file from the texture file resource library to the local, the following steps are further included:
[0029] Determine whether the data obtained after compression corresponding to the texture file already exists locally. If it exists, end the steps; if not, continue with step S2.
[0030] As an implementation manner, in step S6, the data obtained after compression is written into a local file with the original file name plus the corresponding compression format as the name.
[0031] A runtime GPU texture compression system for implementing the method for runtime GPU texture compression of the present invention includes:
[0032] A file path acquisition module for acquiring the texture file path;
[0033] A texture loading module for loading the texture file data;
[0034] A GPU resource object creation module for creating a resource object of the GPU and uploading the texture file data into the resource object;
[0035] A compute shader and compute pipeline state creation module for creating the compute shader of the GPU required for the compression operation and the corresponding compute pipeline state, and binding them to the GPU;
[0036] A compression execution module for calling the compute shader to execute the compression operation and reading back the data obtained after compression;
[0037] A data saving module for saving the data obtained after compression.
[0038] Due to the adoption of the above technical solutions, the present invention has the following beneficial effects:
[0039] 1. Save memory usage: Directly using uncompressed textures will occupy a large amount of memory. Compressed textures can compress the data to 25% to 30% of the original size, greatly reducing the memory occupancy.
[0040] 2. Lower compression error: By adopting a complex index structure, it can ensure that the color as close as possible to that before compression is obtained during decoding, while maintaining a better visual effect when obtaining a high compression ratio.
[0041] 3. Extremely fast compression speed: The characteristic of using CPU for compression is that it can execute more complex calculation methods and obtain higher-quality compression results. However, the CPU has fewer arithmetic units and slow compression speed, which cannot meet the requirements of real-time compression. The characteristic of GPU is that it has a large number of arithmetic units. And due to the highly parallelized compression algorithm, the operations of each block can be assigned to different arithmetic units for processing, thus achieving a speed increase thousands of times that of CPU compression, so as to meet the requirements of real-time compression.
[0042] In addition, in the compression algorithm, packing the data of the average color value into unsigned integers can reduce memory occupancy and improve data transmission efficiency. By selecting the lookup table index through brightness conversion, this method helps to retain the visual consistency of colors during compression while reducing storage requirements. The selection of brightness is based on the brightness of the color, so as to utilize the characteristic that the human eye is more sensitive to brightness changes than to hue and saturation, and improve the compression quality.
[0043] The compute shader divides the texture into blocks of a fixed size, and each thread independently processes the assigned block without affecting each other. A preset lossy compression algorithm is applied to each block. This compression algorithm allows the GPU to directly decompress and access individual compressed blocks without decompressing the entire texture, thus supporting random access.
[0044] The obtained compressed data is written into a local file with the original file name plus the corresponding compression format as the name, which is convenient for subsequent maintenance and use. Description of the Drawings
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0046] Figure 1 It is a flowchart of the method for GPU texture compression during runtime in the embodiments of the present application.
[0047] Figure 2 It is a schematic diagram of the memory layout after texture compression in the embodiments of the present application. Detailed Embodiments
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0049] Example 1
[0050] Please refer to Figure 1 and Figure 2 , a method for GPU texture compression during runtime in the first embodiment of the present invention, includes the steps:
[0051] S1: Obtain the texture file path;
[0052] The ways to obtain the texture file path include: user input or configuration file.
[0053] S2: Load the texture file data;
[0054] In step S2, use the file reading API to read the file content from the texture file.
[0055] S3: Create a resource object of the GPU and upload the texture file data to the resource object;
[0056] In this embodiment, create a GPU texture object equal to the texture size of the texture file.
[0057] S4: Create the compute shader of the GPU required for the compression operation and the corresponding ComputePipeline State, and bind them to the GPU; the compute shader is written according to a preset compression algorithm;
[0058] In step S4, the compression algorithm includes the steps:
[0059] S41: Divide the texture in the texture file into 4×4 blocks, calculate the average color value, maximum color value, and minimum color value of each block and correlate them with each other;
[0060] In this embodiment, store the average color value, maximum color value, and minimum color value using the float3 structure.
[0061] S42: Pack the data of the average color value into an unsigned integer to obtain compressed color information;
[0062] In this embodiment, Pack the average color value of float3 into uint.
[0063] Packing the data of the average color value into an unsigned integer can reduce memory occupancy and improve data transmission efficiency.
[0064] S43: Convert the maximum color value and the minimum color value into lightness; select a lookup table index (LutIndex) and a directory (Indices) through the lightness; write the index data of the lookup table index and the directory, as well as the compressed color information, into a 64-bit binary string according to a preset layout rule.
[0065] Selecting a lookup table index through lightness conversion helps to preserve the visual consistency of colors during compression while reducing storage requirements. The selection of lightness is based on the brightness of the color, which can take advantage of the fact that the human eye is more sensitive to brightness changes than to hue and saturation, improving the compression quality.
[0066] In this embodiment, the preset layout rule includes:
[0067] Write the compressed color information into the first 24-bit binary string of the 64-bit binary string; write the index data of the lookup table index into the 8-bit binary string after the compressed color information in the 64-bit binary string; write the index data of the directory into the last 32-bit binary string of the 64-bit binary string.
[0068] In this embodiment, the structure of the 64-bit binary string after writing data is as Figure 2 shown. Among them, the compressed color information is distributed in sequence according to RGB classification; among them, red (R) is divided into R0 and R1 segments, each occupying 4 bits; green (G) is divided into G0 and G1 segments, each occupying 4 bits; blue (B) is divided into B0 and B1 segments, each occupying 4 bits. The index data T0, T1, diff, and flip of the lookup table index (LutIndex) occupy 3 bits, 3 bits, 1 bit, and 1 bit respectively; the index data indices of the directory (Indices) occupy 32 bits.
[0069] S5: Call the compute shader to perform the compression operation and read back the compressed data obtained;
[0070] The compute shader divides the texture into blocks of a fixed size, and each thread will independently process the assigned blocks without affecting each other. Apply a preset lossy compression algorithm to each block. This compression algorithm allows the GPU to directly decompress and access a single compressed block without decompressing the entire texture, thus supporting random access.
[0071] S6: Save the compressed data obtained.
[0072] A runtime GPU texture compression system for implementing the method for runtime GPU texture compression described in the present invention in an embodiment of the present invention includes:
[0073] A file path acquisition module for acquiring a texture file path;
[0074] A texture loading module for loading the texture file data;
[0075] A GPU resource object creation module for creating a resource object of the GPU and uploading the texture file data into the resource object;
[0076] A compute shader and compute pipeline state creation module for creating the compute shader of the GPU and the corresponding compute pipeline state required for the compression operation, and binding them to the GPU;
[0077] A compression execution module for calling the compute shader to execute the compression operation and reading back the data obtained after compression;
[0078] A data saving module for saving the data obtained after compression.
[0079] To verify the effects of the embodiments of the present invention, the compression time consumption (unit: ms) of traditional CPU texture compression, GPU texture compression of the present application using an i7,4070 graphics card, and GPU texture compression of the present application using a mobile A14 Bionic is now respectively compared. The results are shown in Table 1.
[0080] Resolution CPU ETC2 (i7-14700K) GPU ETC2 (RTX 4070) GPU ETC2 (A14 Bionic) 2048 12071.91 0.04 0.38 1024 3038.94 0.01 0.25 512 867.51 <0.01 0.13 256 232.49 <0.01 0.03
[0081] Table 1. Comparison table of texture compression time consumption
[0082] As can be seen from Table 1, the texture compression time consumption of the method of the present application is much less than that of traditional CPU texture compression. The time consumption gap is from a thousand times to ten thousand times, showing obvious high efficiency.
[0083] Embodiment 2
[0084] A method for runtime GPU texture compression according to the second embodiment of the present invention has basically the same steps as those of the first embodiment, except that before the step S1, there is also a step:
[0085] S7: Pre-compile different versions of the compute shader according to the preset compression algorithm for different compression formats;
[0086] In the step S4, select the corresponding compute shader according to the currently used compression format to complete the creation of the compute shader. Then, create a buffer of the GPU for saving the compressed texture data, and bind the buffer to the pipeline.
[0087] In the step S5, call the compute shader to execute the compression operation and read back the data obtained after compression through the Dispatch() function; and write the result into the bound buffer.
[0088] In step S6, after compression ends, data is read from the created buffer and copied out of the GPU memory and saved to the texture object for final use.
[0089] Compute shaders supporting different compression formats are pre-written, enabling direct selection and invocation of the compute shader during subsequent creation, thus improving the efficiency and convenience of texture compression operations.
[0090] Embodiment 3
[0091] A method for runtime GPU texture compression according to Embodiment 3 of the present invention has steps substantially the same as those of Embodiment 2, except that before step S1, the following steps are further included:
[0092] S8: Upload the texture file to a management background;
[0093] S9: Review the texture file in the management background, and after the review passes, store the texture file in a texture file resource library.
[0094] The method for obtaining the texture file path in S1 includes: downloading from the texture file resource library.
[0095] Background administrators can also directly configure texture files through the management background.
[0096] The adoption of the texture file resource library provides an original data sample library for users to perform texture compression operations. Additionally, it has the function of uploading texture files, which can supplement texture files according to user needs to better meet user requirements. Background administrators can also directly configure texture files, facilitating the maintenance of the texture file resource library.
[0097] Embodiment 4
[0098] A method for runtime GPU texture compression according to Embodiment 4 of the present invention has steps substantially the same as those of Embodiment 3, except that after the user downloads the texture file from the texture file resource library to the local, the following steps are further included:
[0099] Judge whether the data obtained after compression corresponding to the texture file already exists locally. If it exists, end the steps; if not, continue with step S2.
[0100] The adoption of this step can avoid repeated compression operations on already compressed texture files and improve the overall work efficiency.
[0101] Embodiment 5
[0102] A method for GPU texture compression during runtime according to Embodiment 5 of the present invention has steps substantially the same as those of Embodiment 4, except that in step S6, the data obtained after compression is written into a local file with the original file name plus the corresponding compression format as the name, which is convenient for subsequent maintenance and use.
[0103] It should be noted that the description and drawings of the present invention give preferred embodiments of the present invention. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in this specification. These embodiments are not additional limitations to the content of the present invention. The purpose of providing these embodiments is to make the understanding of the disclosed content of the present invention more thorough and comprehensive. Moreover, the above technical features continue to be combined with each other to form various embodiments not listed above, which are all regarded as the scope described in the specification of the present invention. Further, for those of ordinary skill in the art, improvements or changes can be made according to the above description, and all such improvements and changes should fall within the protection scope of the appended claims of the present invention.
[0104] The present invention has been described in detail with reference to the embodiments accompanied by the drawings. Those of ordinary skill in the art can make various variations of the present invention according to the above description. Therefore, certain details in the embodiments should not constitute limitations to the present invention, and the present invention will take the scope defined by the appended claims as the protection scope of the present invention.
Claims
1. A method for runtime GPU texture compression, comprising the steps of: S1: Get the texture file path; S2: Load the texture file data; S3: Create a GPU resource object and upload the texture file data to the resource object; S4: creating a computing shader and a corresponding computing pipeline state of the GPU required for the compression operation, and binding them to the GPU; the computing shader is written according to a preset compression algorithm; S5: calling the compute shader to perform the compression operation and read back the compressed data; S6: Save the compressed data.
2. The method for runtime GPU texture compression according to claim 1, characterized in that: In the step S1, the way of obtaining the texture file path includes: user input or configuration file.
3. The method for runtime GPU texture compression according to claim 1, characterized in that: In the step S2, the file content is read from the texture file using a file reading API.
4. The method for runtime GPU texture compression according to claim 1, characterized in that: In the step S4, the compression algorithm comprises the steps of: S41: Divide the texture in the texture file into 4×4 blocks, calculate the average color value, the maximum color value and the minimum color value of each block and associate them with each other; S42: Packing the data of the average color value into unsigned integers to obtain compressed color information; S43: converting the maximum color value and the minimum color value into brightness; A lookup table index and a directory are selected according to the brightness; and the index data of the lookup table index and the directory and the compressed color information are written into a 64-bit binary string according to a preset arrangement rule.
5. The method for runtime GPU texture compression according to claim 4, characterized in that: The preset arrangement rules include: Write the compressed color information into the first 24 bits of the 64-bit binary string; write the index data of the lookup table index into the 8-bit binary string after the compressed color information in the 64-bit binary string; write the index data of the directory into the last 32 bits of the 64-bit binary string.
6. The method for runtime GPU texture compression according to claim 4, characterized in that: The S1 step also includes the following steps: S7: Pre-write different versions of the compute shader for different compression formats according to the preset compression algorithm; In the step S4, the corresponding compute shader is selected according to the compression format currently used, and the creation of the compute shader is completed.
7. The method for runtime GPU texture compression according to claim 6, characterized in that: The S1 step also includes the following steps: S8: Upload the texture file to a management backend; S9: reviewing the texture file in the management background, and storing the texture file in a texture file resource library after the review is passed.
8. The method for runtime GPU texture compression according to claim 7, characterized in that: The method of obtaining the texture file path in S1 includes: downloading from the texture file resource library; After the user downloads the texture file from the texture file resource library to the local computer, the process further includes the following steps: It is determined whether the compressed data corresponding to the texture file already exists locally. If yes, the step is ended; if no, the step is continued to step S2.
9. The method for runtime GPU texture compression according to claim 6, characterized in that: In the step S6, the compressed data is written into a local file with the original file name plus the name corresponding to the compression format.
10. A runtime GPU texture compression system, characterized in that: The system is used to implement the runtime GPU texture compression method according to any one of claims 1 to 9, and the runtime GPU texture compression system comprises: File path acquisition module, used to obtain texture file path; A texture loading module, used for loading the texture file data; A GPU resource object creation module, used to create a GPU resource object and upload the texture file data to the resource object; A computing shader and computing pipeline state creation module, used to create the computing shader and corresponding computing pipeline state of the GPU required for the compression operation, and bind them to the GPU; A compression execution module, used for calling the compute shader to execute the compression operation and read back the compressed data; A data storage module stores the data obtained after compression.