Buffer area clearing method and device, electronic equipment and readable storage medium
By dividing the data area and tag area in the GPU video memory and compressing the tag value with the target in the preset relationship, a fast glClear operation is achieved, which solves the problem of low efficiency of glClear operation in GPU rendering, reducing data transmission overhead and time overhead.
Patent Information
- Application Number
- CN202411998419.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
During the GPU rendering process, glClear operation requires a large amount of memory access operations, resulting in large data transmission overhead and time overhead, affecting efficiency and leading to rendering delay.
By dividing the data area and the tag area in the GPU's video memory, using the target compressed tag value in the preset relationship to achieve fast glClear operation, reducing the need for actual data writing.
It realizes fast glClear operations, reduces inventory access, reduces data transmission and time overhead, improves the efficiency of glClear operations, and reduces the latency of GPU rendering.
Smart Images

Figure CN120013742A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a buffer clearing method, device, electronic device and readable storage medium. Background Art
[0002] glClear is a function used in OpenGL (Open Graphics Library) to clear (empty) the data in the frame buffer. The frame buffer is a memory area that stores image data, including the color buffer, depth buffer, template buffer, etc. glClear can be used to clear the data in the frame buffer so that the next image drawn will not be affected by the previous drawing.
[0003] However, during the GPU rendering process, when the GPU executes the glClear operation, it needs to write all the data in the specified buffer in the memory into the same specified data. A large number of memory access operations will generate a large amount of data transmission overhead and time overhead, which not only affects the efficiency of the glClear operation, but also causes GPU rendering delays. Summary of the invention
[0004] In view of the above problems, an embodiment of the present invention is proposed to provide a buffer clearing method that overcomes the above problems or at least partially solves the above problems. Under certain conditions, a fast glClear operation can be implemented, memory access operations can be reduced, and the data transmission overhead and time overhead generated by the glClear operation can be reduced, thereby improving the efficiency of the glClear operation and reducing the delay caused by GPU rendering.
[0005] Correspondingly, the embodiment of the present invention also provides a buffer clearing device, an electronic device, and a computer program product to ensure the implementation and application of the above method.
[0006] In a first aspect, an embodiment of the present invention discloses a buffer clearing method, which is applied to a GPU, wherein the display memory of the GPU includes a data area and a tag area, wherein the data area is used to store data operated by the GPU, and the tag area is used to store a compression tag value of corresponding data in the data area, wherein the compression tag value is used to indicate a compression method of the corresponding data, and the method includes:
[0007] receiving a clear instruction for a target buffer, the clear instruction being used to set data in the target buffer to a target value; the target buffer being an area in the data area;
[0008] In response to the clear instruction, it is queried whether there is a target compression tag value corresponding to the target value in the preset relationship, and if so, the compression tag value of the target buffer in the tag area is set to the target compression tag value.
[0009] In a second aspect, an embodiment of the present invention discloses a buffer clearing device, which is applied to a GPU, wherein the display memory of the GPU includes a data area and a tag area, wherein the data area is used to store data operated by the GPU, and the tag area is used to store a compression tag value of corresponding data in the data area, wherein the compression tag value is used to indicate a compression method of the corresponding data, and the device includes:
[0010] An instruction receiving module, used for receiving a clear instruction for a target buffer, wherein the clear instruction is used for setting the data in the target buffer to a target value; the target buffer is an area in the data area;
[0011] The cache clearing module is used to respond to the clearing instruction and query whether there is a target compression tag value corresponding to the target value in the preset relationship. If so, the compression tag value of the target buffer in the tag area is set to the target compression tag value.
[0012] In a third aspect, an embodiment of the present invention discloses an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of any buffer clearing method described above.
[0013] In a fourth aspect, an embodiment of the present invention discloses a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the buffer clearing method as described in any of the above can be implemented.
[0014] In a fifth aspect, an embodiment of the present invention discloses a computer program product, including a computer program, which, when executed by a processor, performs the steps of any of the aforementioned buffer clearing methods.
[0015] The embodiments of the present invention include the following advantages:
[0016] The buffer clearing method provided by the embodiment of the present invention is applied to a GPU, wherein the GPU has enabled the memory access compression function, and the video memory of the GPU is divided into a data area and a tag area, wherein the data area is used to store the data operated by the GPU, and the tag area is used to store the compression tag value of the corresponding data in the data area, and the compression tag value is used to indicate the compression method of the corresponding data. The embodiment of the present invention realizes fast glClear under the condition that certain conditions are met (the target value of the glClear operation has a corresponding target compression tag value in a preset relationship), and only the compression tag value corresponding to the target buffer needs to be set to the target compression tag value without actually writing the target value into the target buffer, so as to achieve the effect of setting all the data in the target buffer to the target value, thereby reducing the memory access amount required for the glClear operation, thereby reducing the data transmission overhead and time overhead generated by the glClear operation, improving the efficiency of the glClear operation, and reducing the delay generated by GPU rendering. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flow chart of steps of an embodiment of a buffer zone clearing method of the present invention;
[0018] Figure 2 is a flowchart of the steps of a buffer clearing method in an example of the present invention;
[0019] Figure 3 is a structural block diagram of an embodiment of a buffer zone clearing device of the present invention;
[0020] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable when appropriate, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In the embodiments of the present invention, the term "multiple" refers to two or more, and other quantifiers are similar.
[0023] Reference Figure 1 , shows a flowchart of a buffer clearing method embodiment of the present invention, the method is applied to a GPU, the GPU's video memory includes a data area and a tag area, the data area is used to store data operated by the GPU, the tag area is used to store a compression tag value of corresponding data in the data area, the compression tag value is used to indicate the compression method of the corresponding data, the method may include the following steps:
[0024] Step 101, receiving a clear instruction for a target buffer, wherein the clear instruction is used to set the data in the target buffer to a target value; the target buffer is an area in the data area;
[0025] Step 102: In response to the clear instruction, query whether there is a target compression tag value corresponding to the target value in the preset relationship; if so, set the compression tag value of the target buffer in the tag area to the target compression tag value.
[0026] The buffer clearing method provided by the present invention can realize fast glClear under certain conditions. The fast glClear refers to realizing the effect of clearing the target buffer by modifying the compression tag value corresponding to the target buffer without actually writing data to the target buffer, thereby reducing the memory access amount required for the glClear operation, thereby reducing the data transmission overhead and time overhead generated by the glClear operation, improving the efficiency of the glClear operation, and reducing the delay generated by GPU rendering.
[0027] In the embodiment of the present invention, the target buffer refers to a frame buffer, which may be at least one of a color buffer, a depth buffer, or a template buffer.
[0028] The color buffer is a memory area used to store the color information of each pixel during the graphics rendering process. In 3D graphics and 2D graphics processing, the color buffer is an important part of the frame buffer, which records the color information of the final image, which will eventually be sent to the monitor for display. The color buffer can have different color formats and precisions, such as RGB8, RGBA8, RGB10A2, etc. These formats determine the color range and precision that can be stored.
[0029] The depth buffer, also known as the Z buffer, is a memory area used to store the depth (or Z coordinate) information of each pixel in graphics rendering. The depth buffer can be used to help determine the front-to-back order of pixels in a 3D scene, thereby achieving correct occlusion relationships and depth perception.
[0030] The stencil buffer is an important component in graphics rendering. It is similar to the color buffer and the depth buffer and is used to store additional information for each pixel on the screen. The stencil buffer can save an unsigned integer value for each pixel on the screen. The specific meaning of this value depends on the specific application of the program. During the rendering process, this value can be compared with a pre-set reference value, and the color value of the corresponding pixel can be updated based on the comparison result. This comparison process is called stencil testing.
[0031] The depth buffer stores the depth value information of each pixel, while the template buffer stores the template value associated with each pixel. These two buffers usually share the same area and have the same resolution size. Therefore, these two buffers are usually used in combination and are called depth template buffers. By combining depth and template information, an effective way is provided to handle occlusion relationships and achieve advanced graphics effects.
[0032] The embodiment of the present invention is mainly described by taking the color buffer as an example, and the operation processes of other types of buffers are similar and can be referred to each other.
[0033] glClear is a function in OpenGL that clears one or more buffers (such as color buffer, depth buffer, stencil buffer, etc.) to a specified value (such as target value). This function is often used to clear the contents of the screen or a specific rendering target before starting a new rendering pass. The glClear function prototype is as follows:
[0034] void glClear(GLbitfield mask);
[0035] The parameter mask is a bit mask that specifies the buffer to be cleared. It can be a combination of the following values:
[0036] GL_COLOR_BUFFER_BIT : Clear the color buffer.
[0037] GL_DEPTH_BUFFER_BIT: Clear the depth buffer.
[0038] GL_STENCIL_BUFFER_BIT: Clear the stencil buffer.
[0039] Before calling glClear, you need to set the specific value (i.e. target value) of the clearing operation:
[0040] Color buffer: Use glClearColor to set the clear color of the color buffer.
[0041] Depth buffer: Use glClearDepth to set the clear depth value of the depth buffer.
[0042] Stencil buffer: Use glClearStencil to set the clear value of the stencil buffer.
[0043] For example, an example is as follows:
[0044] / / Set the clear color to red
[0045] glClearColor(1.0f,0.0f,0.0f,1.0f);
[0046] / / Set the clear depth value to 1.0
[0047] glClearDepth(1.0);
[0048] / / Set the clear template value to 0
[0049] glClearStencil(0);
[0050] / / Clear the color buffer, depth buffer and stencil buffer
[0051] glClear(GL_COLOR_BUFFER_BIT|GL_DEPTH_BUFFER_BIT|GL_STE NCIL_BUFFER_BIT);
[0052] In the embodiment of the present invention, the clearing instruction for the target buffer refers to an instruction for calling glClear.
[0053] The buffer clearing method provided by the present invention is applied to a GPU, wherein the GPU has enabled a memory access compression function, and the display memory of the GPU is divided into a data area and a tag area, wherein the data area is used to store data operated by the GPU, and the tag area is used to store a compression tag value of corresponding data in the data area, and the compression tag value is used to indicate a compression method of the corresponding data. The memory access compression function is to compress the data used in memory access by using a certain means (such as a certain compression method) to reduce the amount of data.
[0054] It should be noted that the embodiment of the present invention does not limit the size of the data area and the label area. For example, the data area and the label area can be divided according to a preset ratio. For example, the preset ratio is 128:1. Exemplarily, when the kernel driver initializes the GPU, the GPU memory access compression function is turned on, and the video memory is divided into a data area and a label area. Taking 256G of video memory as an example, the first 2G of the video memory space can be used as the label area, and the last 254G as the data area. The physical base address (base) and size (size) corresponding to the label area are filled in the corresponding initialization register for use in calculating the physical address when allocating the compressed label area.
[0055] For data of different formats, when writing into the data area, you can select a suitable compression method according to the data format to compress and write into the data area, and write the corresponding compression tag value into the corresponding tag area. For example, the target buffer is a physical area allocated in the data area, and the tag area corresponds to a compression tag area allocated for the target buffer, which is used to store the compression tag value corresponding to the target buffer, and the compression tag value is used to indicate the compression method corresponding to the compressed data stored in the target buffer. Of course, in a specific implementation, if the data format is not suitable for compression, the original data can also be written into the data area. The embodiment of the present invention does not limit the format of the data operated by the GPU, nor does it limit the compression method used.
[0056] Exemplarily, the data formats supported by the target buffer in the embodiment of the present invention and the compression methods corresponding to different data formats are shown in Table 1.
[0057] Table 1
[0058] Compression mode identifier Compression method 0 No compression 1 Compress in RGBA8 format 2 Compressed in RGB10A2 format 3 Compressed in A2RGB10 format 4 Compressed in D24S8 format 5 Compressed in S8D24 format
[0059] As shown in Table 1, the data formats supported by the embodiment of the present invention may include but are not limited to any of the following: RGBA8, RGB10A2, A2RGB10, D24S8, S8D24. Among them, RGBA8, RGB10A2 and A2RGB10 are data formats (or color formats) that can be used by the color buffer. D24S8 and S8D24 are data formats that can be used by the depth template buffer.
[0060] RGBA8 is a standard texture format that contains red (R), green (G), blue (B), and alpha (A) channels, each with 8 bits. For compression of the RGBA8 format, a variety of compression algorithms can be used, such as ETC2 (Ericsson TextureCompression 2, a texture compression format) or ASTC (Adaptive Scalable Texture Compression, a texture compression format). These compression formats can significantly reduce the storage size of the texture while maintaining image quality.
[0061] RGB10A2 is an HDR (High Dynamic Range) texture format, where the RGB channels each occupy 10 bits and the Alpha channel occupies 2 bits. This format can provide higher color accuracy than the standard 8-bit format. RGB10A2 can be used to store HDR content without losing details. For the compression of RGB10A2, a compression algorithm such as ASTC can be used, which supports the compression of HDR content and can provide different compression ratios and qualities depending on the block size.
[0062] A2RGB10 is another HDR texture format, where the Alpha channel occupies 2 bits and the RGB channels each occupy 10 bits. This format is often used for HDR textures that require an Alpha channel. The compression technology of A2RGB10 is similar to RGB10A2, and compression formats such as ASTC can also be used to reduce storage size.
[0063] S8D24 (or D24S8) is a depth template format, where D24 represents 24 bits of depth channel information and S8 represents 8 bits of template channel information. This format is commonly used in 3D graphics rendering to store depth information and template information for depth testing and template testing. Compression algorithms such as LZ4 can be used for general data compression, but data compression in the S8D24 (or D24S8) format may involve specific graphics hardware and driver support. In some cases, depth and template data can be compressed together with color data to further reduce storage requirements.
[0064] For the target buffer, whether the data stored in the target buffer is compressed and the compression method used can be determined by querying Table 1 according to the data format of the target buffer.
[0065] The embodiment of the present invention pre-establishes a corresponding relationship between preset original data and preset compression label values, which is called a preset relationship. In the preset relationship, each preset original data corresponds to a preset compression label value, and each channel value of each preset original data meets the preset value. See Table 2, which shows a specific schematic diagram of a preset relationship of the present invention.
[0066] Table 2
[0067]
[0068] As shown in Table 2, a preset raw data can be determined by the preset value of each channel, RGB and A (Alpha) represent the channel values of the data format of the color buffer (RGBA8, RGB10A2 and A2RGB10). D and S represent the channel values of the data format of the depth template buffer (D24S8 and S8D24). Each preset raw data in Table 2 corresponds to a preset compression tag value.
[0069] In an example, assume that the data format of data 1 is RGBA8, and the values of the RGB channel and the Alpha channel of data 1 are both 0, such as the color value of data 1 (0,0,0,0). By querying Table 2, the channel values of data 1 match the channel values of the preset original data in the first row of Table 2, and data 1 conforms to the preset original data represented by the first row in Table 2. Therefore, data 1 has a corresponding preset compression tag value of 0011 in Table 2. When the compression tag value is 0011, it can be known that the original data corresponding to the compression tag value 0011 is (0,0,0,0).
[0070] In another example, assume that the data format of data 2 is RGBA8, and the values of the RGB channel and the Alpha channel of data 2 are both 1, such as data 2 is a color value (1,1,1,1). By querying Table 2, data 2 matches the preset original data in row 4 of Table 2, and the corresponding preset compression tag value of data 2 in Table 2 is 1111. When the compression tag value is 1111, it can be known that the corresponding original data is (1,1,1,1).
[0071] Through the preset relationship shown in Table 2, the original data can be quickly obtained according to the preset compression tag value, which can reduce the operation steps of decompressing data and improve data access efficiency.
[0072] Furthermore, the embodiment of the present invention can implement a fast glClear operation using Table 2. Specifically, fast glClear can be implemented when certain conditions are met (the target value has a corresponding target compression tag value in a preset relationship).
[0073] In an optional embodiment of the present invention, the querying whether there is a target compression label value corresponding to the target value in the preset relationship may include:
[0074] Each channel value of the target value is compared with each channel value of each preset original data in the preset relationship. If there is matching preset original data, the preset compression label value corresponding to the matching preset original data is determined to be the target compression label value corresponding to the target value; if there is no matching preset original data, it is determined that the target compression label value corresponding to the target value does not exist in the preset relationship.
[0075] In one example, it is necessary to perform a glClear operation on the target buffer, and the target value of the glClear operation is (0,0,0,0), that is, it is necessary to set all the data in the target buffer to black. According to the embodiment of the present invention, it is only necessary to set the compression tag value corresponding to the target buffer in the tag area to the target compression tag value 0011. Since the original data corresponding to the target compression tag value 0011 is that the values of the RGB channel and the Alpha channel are both 0, the effect of setting all the data in the target buffer to black is achieved.
[0076] Exemplarily, assume that the user program calls the glClear function (a function in OpenGL) and requests that all data in the target buffer (such as the target buffer is the color buffer A) be set to the target value, assuming that the target value is (0,0,0,0). The user-state driver (a driver that implements the OpenGL API) receives the request and queries the preset relationship shown in Table 2. The preset original data in the first row of Table 2 matches each channel of the target value, so the preset compression tag value (0011) corresponding to the target value can be queried as the target compression tag value, that is, the target compression tag value (0011) corresponding to the target value can be queried, so fast glClear can be performed for the color buffer A. The user-state driver calls the GPU to set the compression tag value of the color buffer A in the tag area to the target compression tag value (0011). After setting, the compression tag value of the color buffer A in the tag area is 0011, indicating that all data in the color buffer A is (0,0,0,0), without actually writing the target value (0,0,0,0) into the data area corresponding to the color buffer A. After reading the tag value of 0011, color buffer A is considered to be cleared to (0,0,0,0) without actually clearing it.
[0077] As another example, suppose the user program calls the glClear function again, requesting that all the data in color buffer A be set to the target value (1,1,1,1). The user-state driver receives the request and queries the preset relationship shown in Table 2. It can be found that there is a target compression tag value (1111) corresponding to the target value, so a fast glClear can be performed on color buffer A. The user-state driver calls the GPU to set the compression tag value of color buffer A in the tag area to the target compression tag value (1111). After setting, the compression tag value of color buffer A in the tag area is 1111, indicating that all the data in color buffer A is (1,1,1,1), without actually writing the target value (1,1,1,1) into the data area where color buffer A is located.
[0078] In the case where there is no target compression label value corresponding to the target value in the preset relationship, it means that the target value cannot directly determine the original data through the compression label value, and therefore, the fast glClear method cannot be used.
[0079] The buffer clearing method provided by the present invention can realize fast glClear under the condition that certain conditions are met (the target value has a corresponding target compression tag value in a preset relationship). It only needs to set the compression tag value corresponding to the target buffer to the target compression tag value without actually writing the target value to the target buffer, so as to achieve the effect of setting all the data in the target buffer to the target value, thereby reducing the memory access amount required for the glClear operation, and further reducing the data transmission overhead and time overhead generated by the glClear operation, improving the efficiency of the glClear operation, and reducing the delay caused by GPU rendering.
[0080] In an optional embodiment of the present invention, the method may further include:
[0081] Step S11, receiving a cache allocation request from a user-mode driver, wherein the cache allocation request carries a compression mode identifier corresponding to the data format of the requested buffer; the compression mode identifier is used to indicate a compression mode corresponding to the corresponding data format;
[0082] Step S12, in response to the cache allocation request, allocating the target buffer in the data area, and recording the compression mode identifier in a page table entry of the target buffer;
[0083] Step S13, allocating a compressed label area for the target buffer in the label area; the compressed label area is used to store a compressed label value corresponding to the target buffer;
[0084] Step S14, bind the first physical address of the target buffer corresponding to the data area to the first virtual address, and bind the second physical address of the compressed label area corresponding to the label area to the second virtual address, and return the first virtual address and the second virtual address to the user state driver.
[0085] In a specific implementation, when rendering, the application may need to apply for a target buffer (such as at least one of a color buffer, a depth buffer, and a template buffer). For example, the user program requests the generation of a color buffer by calling an API (Application Programming Interface) provided by OpenGL, such as glgenFramebuffers or glbindFramebuffer. Assuming that the data format of the requested color buffer is RGBA8, the user-mode driver can find out by querying the correspondence between the data format and the compression method shown in Table 1 that the compression method identifier corresponding to the data format RGBA8 is 1, and then sends the cache allocation request and the compression method identifier to the kernel. The kernel allocates a color buffer (such as color buffer A) in the data area, and records the compression method identifier 1 in the page table entry of the color buffer A. In addition, after the color buffer A is allocated in the data area, a corresponding compression tag area is allocated to the color buffer A in the tag area; the compression tag area is used to store the compression tag value corresponding to the color buffer A. Since the data format RGBA8 has a corresponding compression method, the data written to the color buffer A will be stored in a compressed form. Specifically, by querying the compression method recorded in the page table entry and identifying it as 1, the data to be written can be compressed according to the compression method corresponding to the RGBA8 format and written into the color buffer A. Then, the compression tag value corresponding to the color buffer A is calculated based on the written data and the compression method, and written into the compressed tag area corresponding to the tag area of the color buffer A.
[0086] In an example, after the target buffer is allocated in the data area, the physical address of the compressed tag area allocated for the target buffer can be calculated by the following formula:
[0087] (pa >> 7) & size | base (1)
[0088] Where pa is the physical address of the target buffer (called the first physical address), size is the size of the tag area in the video memory, and base is the physical base address of the tag area in the video memory. Pa is the physical address where the compressed tag value corresponding to the target buffer is stored (called the second physical address). ">>" means right shift, "&" means bitwise AND, and "|" means bitwise OR.
[0089] Next, the first physical address corresponding to the target buffer in the data area is bound to the first virtual address (the virtual address of the target buffer), and the second physical address corresponding to the compressed label area in the label area is bound to the second virtual address (the virtual address of the compressed label area corresponding to the target buffer), and the first virtual address and the second virtual address are returned to the user state driver for access by the user state driver.
[0090] In an embodiment of the present invention, whether the data stored in the target buffer is compressed data can be determined by querying Table 1 based on the data format used by the target buffer. For example, for a target buffer, if the data format used by the target buffer is not in Table 1, the corresponding compression mode identifier is 0, indicating that the data stored in the target buffer is uncompressed original data. If the data format used by the target buffer is RGBA8, the corresponding compression mode identifier is 1, indicating that the data stored in the target buffer is compressed data in the RGBA8 format.
[0091] In an optional embodiment of the present invention, setting the compression tag value of the target buffer in the tag area to the target compression tag value may include:
[0092] Step S21: receiving a first command sent by a user mode driver, where the first command carries the second virtual address and the target compression label value;
[0093] Step S22: In response to the first command, query a page table based on the second virtual address to obtain the second physical address;
[0094] Step S23: Based on the second physical address, write the target compressed label value into the label area.
[0095] The first command may include a drawing command or a DMA (Direct Memory Access) copy command. The GPU queries the page table through the second virtual address in the drawing command or the DMA copy command to find the second physical address corresponding to the second virtual address, so that the GPU can operate the area corresponding to the second physical address and write the target compressed tag value into the area corresponding to the second physical address in the tag area.
[0096] DMA is a hardware module in the GPU for high-speed data transmission. The user-state driver sends a DMA copy command to the GPU, which is used to write a compressed label area starting from the second physical address corresponding to the second virtual address in the label area as the target compressed label value. The user-state driver calls the first command to send to the command queue of the GPU, and then the GPU calls the internal DMA module to complete the corresponding operation.
[0097] In one example, the application calls the glClear function to request a glClear operation on the color buffer A, assuming that the target value is (0,0,0,0). The user-state driver receives the request, queries the preset relationship shown in Table 2, and finds that the target value has a corresponding target compression tag value 0011, so a fast glClear can be performed on the color buffer A. Therefore, the user-state driver uses the second virtual address (such as v2) to call the first command, and the first command carries the target compression tag value 0011 and the second virtual address v2. The GPU receives the first command, queries the GPU page table through v2, finds the second physical address m2 corresponding to v2, and writes the target compression tag value into the area corresponding to m2 in the tag area.
[0098] Furthermore, a GPU pipeline clear command may be inserted before the first command, that is, before sending the first call command to the command queue of the GPU, a GPU pipeline clear command is inserted before the first command to ensure data consistency.
[0099] In an optional embodiment of the present invention, the method may further include:
[0100] Step S31: If the target compression label value corresponding to the target value does not exist in the preset relationship, receiving a second command sent by a user-mode driver, where the second command carries a first virtual address, a second virtual address and the target compression label value;
[0101] Step S32: In response to the second command, query a page table based on the first virtual address to obtain the first physical address and a corresponding compression mode identifier;
[0102] Step S33, compressing the target value according to the compression method corresponding to the queried compression method identifier, and writing the compressed data into the data area based on the first physical address;
[0103] Step S34: query a page table based on the second virtual address to obtain the second physical address;
[0104] Step S35: Calculate the compression label value corresponding to the compressed data, and write the calculated compression label value into the label area based on the second physical address.
[0105] In an embodiment of the present invention, when an application calls the glClear function to perform a clear operation on a target buffer, the embodiment of the present invention determines whether a fast glClear operation can be performed on the target buffer based on whether a corresponding target compression tag value exists in a preset relationship for the target value of the glClear operation; if the target value exists in a preset relationship for a corresponding target compression tag value, a fast glClear operation can be performed on the target buffer (such as executing steps S21 to S23); if the target value does not exist in a preset relationship for a corresponding target compression tag value, a fast glClear operation cannot be performed on the target buffer, and the original glClear operation method needs to be used (such as executing steps 31 to S35).
[0106] If the target compression tag value corresponding to the target value does not exist in the preset relationship, it is necessary to write the specific target value into the target buffer to implement the glClear operation.
[0107] Reference Figure 2 , shows a flowchart of a buffer clearing method in an example of the present invention, which may include the following steps:
[0108] A1: Enable GPU memory compression.
[0109] Exemplarily, when the kernel driver initializes the GPU, the GPU memory access compression function is turned on, and the video memory is divided into a data area and a label area. Taking 256G video memory as an example, the first 2G of the video memory space can be used as the label area, and the last 254G as the data area. The physical base address (base) and size (size) corresponding to the label area are filled in the corresponding initialization register for use in calculating the physical address when allocating the compressed label area.
[0110] A2: Apply for the target buffer.
[0111] For example, an application uses functions such as glgenFramebuffers or glbindFramebuffer to request the generation of a color buffer. The user-state driver receives the request and assumes that a color buffer A is requested. Assume that the data format of the requested color buffer A is RGBA8. The user-state driver queries Table 1 to obtain the corresponding compression mode identifier 1, and passes the compression mode identifier to the kernel. The kernel allocates a physical area for the color buffer A in the data area, and writes the compression mode identifier 1 into the page table entry corresponding to the allocated physical area; at the same time, a compressed label area is allocated for the color buffer A in the label area, and the physical position of the compressed label area is calculated according to formula (1). The kernel binds the first physical address of the color buffer A in the data area (such as m1) to the first virtual address (such as v1), and binds the second physical address of the compressed label area corresponding to the color buffer A in the label area (such as m2) to the second virtual address (such as v2), and returns the two virtual addresses (v1 and v2) to the user-state driver for access and use by the user-state driver.
[0112] A3: Call the glClear function to clear the target buffer.
[0113] When an application calls the glClear function to request a glClear operation on color buffer A, it is assumed that the target value set by the application is (0,0,0,0).
[0114] A4: Check whether the target value has a corresponding target compression tag value; if so, execute step A5; if not, execute step A6.
[0115] The user-mode driver queries the preset relationship shown in Table 2. If there is a target compression tag value corresponding to the target value, a fast glClear operation can be performed on color buffer A; if there is no target compression tag value corresponding to the target value, a normal glClear operation is performed on color buffer A.
[0116] A5: Perform a fast glClear operation.
[0117] In this example, the user-state driver queries the preset relationship shown in Table 2. The target compression tag value corresponding to the target value is 0011, and fast glClear can be performed on color buffer A. The user-state driver uses the virtual address v2 to initiate a DMA copy command. The kernel fills m2 and v2 into the GPU page table according to the DMA copy command initiated by the user-state driver. The GPU receives the DMA copy command initiated by the user-state driver, queries the GPU page table through v2, finds the physical address m2 corresponding to v2, and writes the value in the compression tag area corresponding to m2 to 0011.
[0118] A6: Perform a normal glClear operation.
[0119] If there is no target compression tag value corresponding to the target value in the preset relationship, a normal glClear operation is performed on the target buffer.
[0120] It should be noted that in the specific implementation, after the application applies for the target buffer, it can also write target data to the target buffer. The type and value of the target data depend on the specific behavior of the application. For example, the application can call APIs such as glDrawArrays to write target data to the target buffer. After receiving the request, the user-mode driver sends a drawing command to the GPU. If the GPU memory access compression function is turned on, the GPU will write the compressed data to the data area where the target buffer is located, and will write the corresponding compression tag value to the compression tag area corresponding to the target buffer. When the application needs to read the data in the target buffer, the GPU will read the compression tag value corresponding to the target buffer, and decompress the data in the target buffer according to the compression tag value to obtain the decompressed data.
[0121] Furthermore, when writing target data to the target buffer, if the written target data are all preset original data existing in the preset relationship, for example, the application calls the glDrawArrays function to write target values (0,0,0,0) in RGBA8 format to the color buffer A, since the target value is the preset original data in the preset relationship, the preset compression tag corresponding to the preset original data is 0011, and therefore the preset compression tag value 0011 can be written into the compression tag area corresponding to the color buffer A, and the actual value does not need to be written into the color buffer.
[0122] Furthermore, when reading target data from a target buffer, if the compression tag value corresponding to the target buffer is a compression tag value existing in a preset relationship, the decompressed data can be directly obtained through the preset relationship without reading the compressed data in the target buffer and then performing a decompression operation. For example, when it is necessary to read data in a color buffer A, the compression tag value in the compression tag area corresponding to the color buffer A is first read, and the compression tag value read is 0011, which exists in the preset relationship, and the data in the color buffer A can be directly determined to be all (0,0,0,0) according to the preset relationship without performing a decompression operation.
[0123] In a specific implementation, at the beginning of rendering a new frame, the target buffer may be cleared to ensure that the rendering of the new frame is not affected by the content of the previous frame. Alternatively, when switching from one rendering target to another (e.g., in a multi-rendering pass), the target buffer of the new rendering target may need to be cleared to avoid interference from old content. It is understood that the embodiments of the present invention do not limit the scenario and timing of clearing the target buffer.
[0124] It should be noted that if the data format used by the target buffer requested by the application cannot adopt any of the compression methods shown in Table 1, that is, the data in the target buffer is stored as uncompressed original data, then when the application needs to perform a glClear operation on the target buffer, a normal glClear operation should be performed.
[0125] In summary, the buffer clearing method provided by the embodiment of the present invention is applied to a GPU, the GPU has turned on the memory access compression function, and the video memory of the GPU is divided into a data area and a tag area, the data area is used to store the data of the GPU operation, and the tag area is used to store the compression tag value of the corresponding data in the data area, and the compression tag value is used to indicate the compression method of the corresponding data. The embodiment of the present invention realizes fast glClear under the condition that certain conditions are met (the target value of the glClear operation has a corresponding target compression tag value in a preset relationship), and only needs to set the compression tag value corresponding to the target buffer to the target compression tag value without actually writing the target value to the target buffer, so as to achieve the effect of setting all the data in the target buffer to the target value, thereby reducing the memory access amount required for the glClear operation, thereby reducing the data transmission overhead and time overhead generated by the glClear operation, improving the efficiency of the glClear operation, and reducing the delay generated by GPU rendering.
[0126] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0127] Reference Figure 3, shows a structural block diagram of an embodiment of a buffer clearing device of the present invention, the device is applied to a GPU, the display memory of the GPU includes a data area and a tag area, the data area is used to store data operated by the GPU, the tag area is used to store a compression tag value of corresponding data in the data area, the compression tag value is used to indicate a compression method of the corresponding data, the device may include:
[0128] The instruction receiving module 301 is used to receive a clear instruction for a target buffer, wherein the clear instruction is used to set the data in the target buffer to a target value; the target buffer is an area in the data area;
[0129] The cache clearing module 302 is used to query whether there is a target compression tag value corresponding to the target value in the preset relationship in response to the clearing instruction, and if so, set the compression tag value of the target buffer in the tag area to the target compression tag value.
[0130] Optionally, the preset relationship includes a corresponding relationship between preset original data and preset compression label values; each preset original data corresponds to a preset compression label value; and each channel value of each preset original data conforms to a preset value.
[0131] Optionally, the cache clearing module is specifically used to:
[0132] Each channel value of the target value is compared with each channel value of each preset original data in the preset relationship. If there is matching preset original data, the preset compression label value corresponding to the matching preset original data is determined to be the target compression label value corresponding to the target value; if there is no matching preset original data, it is determined that the target compression label value corresponding to the target value does not exist in the preset relationship.
[0133] Optionally, the device further comprises:
[0134] A first receiving module is used to receive a cache allocation request from a user-mode driver, wherein the cache allocation request carries a compression mode identifier corresponding to a data format of a requested buffer; the compression mode identifier is used to indicate a compression mode corresponding to a corresponding data format;
[0135] A first allocation module, configured to allocate the target buffer in the data area in response to the cache allocation request, and record the compression mode identifier in a page table entry of the target buffer;
[0136] A second allocation module is used to allocate a compressed label area to the target buffer in the label area; the compressed label area is used to store a compressed label value corresponding to the target buffer;
[0137] An address binding module is used to bind the first physical address of the target buffer in the data area to the first virtual address, and to bind the second physical address of the compressed label area in the label area to the second virtual address, and return the first virtual address and the second virtual address to the user mode driver.
[0138] Optionally, the cache clearing module includes:
[0139] The first clearing submodule is used to receive a first command sent by a user-mode driver, wherein the first command carries the second virtual address and the target compression label value; in response to the first command, query the page table based on the second virtual address to obtain the second physical address; and write the target compression label value into the label area based on the second physical address.
[0140] Optionally, the cache clearing module includes:
[0141] A second clearing submodule is used to receive a second command sent by a user-mode driver if the target compression label value corresponding to the target value does not exist in the preset relationship, the second command carrying a first virtual address, a second virtual address and the target compression label value; in response to the second command, query a page table based on the first virtual address to obtain the first physical address and a corresponding compression method identifier; compress the target value according to the compression method corresponding to the queried compression method identifier, and write the compressed data to the data area based on the first physical address; query the page table based on the second virtual address to obtain the second physical address; calculate the compression label value corresponding to the compressed data, and write the calculated compression label value to the label area based on the second physical address.
[0142] Optionally, the target buffer includes at least one of a color buffer, a depth buffer or a template buffer.
[0143] The buffer clearing device provided by the embodiment of the present invention is applied to a GPU, the GPU has turned on the memory access compression function, and the video memory of the GPU is divided into a data area and a tag area, the data area is used to store the data operated by the GPU, and the tag area is used to store the compression tag value of the corresponding data in the data area, and the compression tag value is used to indicate the compression method of the corresponding data. The embodiment of the present invention realizes fast glClear under the condition that certain conditions are met (the target value of the glClear operation has a corresponding target compression tag value in a preset relationship), and only needs to set the compression tag value corresponding to the target buffer to the target compression tag value without actually writing the target value to the target buffer, so as to achieve the effect of setting all the data in the target buffer to the target value, thereby reducing the memory access amount required for the glClear operation, thereby reducing the data transmission overhead and time overhead generated by the glClear operation, improving the efficiency of the glClear operation, and reducing the delay generated by GPU rendering.
[0144] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0145] Reference Figure 4 , is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 4 As shown, the electronic device includes: a processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of the buffer clearing method of the aforementioned embodiment.
[0146] An embodiment of the present invention provides a non-transitory computer-readable storage medium. When instructions in the storage medium are executed by a program or a processor of a terminal, the terminal is enabled to perform the steps of the buffer clearing method of the aforementioned embodiment.
[0147] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0148] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0149] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0150] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing terminal device to operate in a predictable manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0151] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0152] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0153] Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A buffer clearing method, characterized in that: Applied to a GPU, the GPU's video memory includes a data area and a tag area, the data area is used to store data operated by the GPU, the tag area is used to store a compression tag value of corresponding data in the data area, the compression tag value is used to indicate a compression method of the corresponding data, the method includes: receiving a clear instruction for a target buffer, the clear instruction being used to set data in the target buffer to a target value; the target buffer being an area in the data area; In response to the clear instruction, it is queried whether there is a target compression tag value corresponding to the target value in the preset relationship, and if so, the compression tag value of the target buffer in the tag area is set to the target compression tag value.
2. The method according to claim 1, characterized in that The preset relationship includes a corresponding relationship between preset original data and preset compression label values; each preset original data corresponds to a preset compression label value; and each channel value of each preset original data conforms to a preset value.
3. The method according to claim 2, characterized in that The querying whether there is a target compression label value corresponding to the target value in the preset relationship includes: Each channel value of the target value is compared with each channel value of each preset original data in the preset relationship. If there is matching preset original data, the preset compression label value corresponding to the matching preset original data is determined to be the target compression label value corresponding to the target value; if there is no matching preset original data, it is determined that the target compression label value corresponding to the target value does not exist in the preset relationship.
4. The method according to claim 1, characterized in that: The method further comprises: Receive a cache allocation request from a user-mode driver, the cache allocation request carrying a compression mode identifier corresponding to a data format of a requested buffer; the compression mode identifier is used to indicate a compression mode corresponding to a corresponding data format; In response to the cache allocation request, allocating the target buffer in the data area, and recording the compression mode identifier in a page table entry of the target buffer; Allocate a compression label area for the target buffer in the label area; the compression label area is used to store the compression label value corresponding to the target buffer; Bind the first physical address of the target buffer corresponding to the data area to the first virtual address, and bind the second physical address of the compressed label area corresponding to the label area to the second virtual address, and return the first virtual address and the second virtual address to the user mode driver.
5. The method according to claim 4, characterized in that The step of setting the compression tag value of the target buffer in the tag area to the target compression tag value comprises: Receive a first command sent by a user mode driver, where the first command carries the second virtual address and the target compression label value; In response to the first command, query a page table based on the second virtual address to obtain the second physical address; Based on the second physical address, the target compressed tag value is written into the tag area.
6. The method according to claim 4, characterized in that The method further comprises: If the target compression label value corresponding to the target value does not exist in the preset relationship, receiving a second command sent by the user mode driver, where the second command carries the first virtual address, the second virtual address and the target compression label value; In response to the second command, query a page table based on the first virtual address to obtain the first physical address and a corresponding compression mode identifier; compressing the target value according to the compression method corresponding to the queried compression method identifier, and writing the compressed data into the data area based on the first physical address; Querying a page table based on the second virtual address to obtain the second physical address; A compression label value corresponding to the compressed data is calculated, and based on the second physical address, the calculated compression label value is written into the label area.
7. The method according to any one of claims 1 to 6, characterized in that: The target buffer includes at least one of a color buffer, a depth buffer, or a template buffer.
8. A buffer zone clearing device, characterized in that: Applied to a GPU, the display memory of the GPU includes a data area and a tag area, the data area is used to store data operated by the GPU, the tag area is used to store a compression tag value of corresponding data in the data area, the compression tag value is used to indicate a compression method of the corresponding data, and the device includes: An instruction receiving module, used for receiving a clear instruction for a target buffer, wherein the clear instruction is used for setting the data in the target buffer to a target value; the target buffer is an area in the data area; The cache clearing module is used to respond to the clearing instruction and query whether there is a target compression tag value corresponding to the target value in the preset relationship. If so, the compression tag value of the target buffer in the tag area is set to the target compression tag value.
9. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of the buffer clearing method according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the buffer clearing method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the buffer clearing method according to any one of claims 1 to 7 are performed.