Rendering system, chip, equipment and rendering method applied to rendering system
By introducing the rendering pipeline manager, the asynchronous operation of geometric processing pipelines in the GPU architecture is realized, which solves the problems of image rendering latency and low efficiency, and improves image rendering efficiency.
Patent Information
- Application Number
- CN202510815636.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
In existing GPU architectures, there are problems of large delays and low efficiency in the image rendering process, especially when multiple geometric processing pipelines need to be run simultaneously, resulting in image rendering delays and inefficiency.
The rendering pipeline manager is introduced to uniformly manage the data segments processed by the geometry processing pipeline, and send rendering instructions to the pixel processing pipeline when there are continuous data segments to realize the asynchronous operation of the geometry processing pipeline.
By running the geometric processing pipeline asynchronously, the delay of image rendering is reduced and the image rendering efficiency is improved.
Smart Images

Figure CN120339483A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of image rendering, and particularly to a rendering system, a chip, a device, and a rendering method applied to the rendering system. Background Art
[0002] A GPU (Graphics Processing Unit) is used to render images displayed on an electronic device.
[0003] In related technologies, there is a GPU architecture: Tile Based Rendering (TBR). In TBR, the screen of an electronic device is divided into multiple tiles, and then each tile is rendered one by one. Specifically, multiple geometry processing pipelines are first used to process the input data stream respectively, and then the processing results are divided into multiple pieces of primitive data based on tile division. Further, a pixel processing pipeline reads and processes these multiple pieces of primitive data to obtain the final rendering result.
[0004] However, in the above related technologies, the pixel processing pipeline needs to wait for multiple geometry processing pipelines to all complete processing before it can process the multiple pieces of primitive data. During this process, the geometry processing pipeline that finishes processing first needs to wait for other geometry processing pipelines to complete before triggering the subsequent rendering process. That is, multiple geometry processing pipelines need to run synchronously. Therefore, the latency of image rendering is relatively large and the efficiency is relatively low. Summary of the Invention
[0005] Embodiments of the present application provide a rendering system, a chip, a device, and a rendering method applied to the rendering system. The technical solutions provided by the embodiments of the present application include the following contents.
[0006] According to one aspect of the embodiments of the present application, a rendering system is provided. The rendering system includes: a rendering pipeline manager, an intermediate buffer, M geometry processing pipelines, and N pixel processing pipelines, where M is an integer greater than 1 and N is an integer greater than or equal to 1; wherein, multiple data segments divided from the input data stream of the rendering system are assigned to the M geometry processing pipelines for processing, and the primitive data obtained after each data segment is processed is stored in the intermediate buffer; The rendering pipeline manager is configured to obtain the numbers and memory information of the data segments processed by the geometry processing pipelines, and the memory information of the data segments is used to determine the memory pages occupied by the primitive data corresponding to the data segments in the intermediate buffer; The rendering pipeline manager is further configured to separately send rendering instructions to the N pixel processing pipelines. The rendering instructions include the numbers of k data segments and memory information, where k is an integer greater than or equal to 1. When k is greater than 1, the numbers of the k data segments are consecutive.
[0007] According to one aspect of the embodiments of the present application, a GPU chip is provided. The GPU chip includes the above-mentioned rendering system.
[0008] According to one aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes a GPU chip, and the GPU chip includes the above-mentioned rendering system.
[0009] According to one aspect of the embodiments of the present application, a rendering method applied to a rendering system is provided. The rendering system includes: a rendering pipeline manager, an intermediate buffer, M geometry processing pipelines, and N pixel processing pipelines, where M is an integer greater than 1 and N is an integer greater than or equal to 1. Among them, multiple data segments divided from the input data stream of the rendering system are assigned to the M geometry processing pipelines for processing, and the primitive data obtained after processing each data segment is stored in the intermediate buffer. The method includes the following steps.
[0010] The rendering pipeline manager obtains the numbers and memory information of the data segments processed by the geometry processing pipelines. The memory information of the data segments is used to determine the memory pages occupied by the primitive data corresponding to the data segments in the intermediate buffer. The rendering pipeline manager separately sends rendering instructions to the N pixel processing pipelines. The rendering instructions include the numbers of k data segments and memory information, where k is an integer greater than or equal to 1. When k is greater than 1, the numbers of the k data segments are consecutive.
[0011] The technical solutions provided by the embodiments of the present application may include the following beneficial effects.
[0012] The rendering system proposed in this application introduces a rendering pipeline manager to uniformly manage the data segments processed by the geometry processing pipeline. Specifically, the rendering pipeline manager obtains the numbers and memory information of the data segments processed by the geometry processing pipeline. After there are k consecutively numbered data segments, it sends rendering instructions to N pixel processing pipelines respectively. That is, as long as there are k consecutively numbered data segments, regardless of the processing progress of the geometry processing pipeline, the rendering pipeline manager will send rendering instructions to indicate rendering. Therefore, each geometry processing pipeline does not need to wait for other geometry processing pipelines to synchronize. Through the rendering pipeline manager, the asynchronous operation of multiple geometry processing pipelines can be achieved (each geometry processing pipeline continuously processes data segments without maintaining synchronization). Therefore, in the embodiments of this application, the latency of image rendering is reduced, thereby improving the efficiency of image rendering. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0014] Figure 1 is a schematic diagram of the rendering system provided in a possible implementation manner of this application; Figure 2 is a schematic diagram of the internal processing flow of the GPU core provided in a possible implementation manner of this application; Figure 3 is a schematic diagram of the rendering system provided in another possible implementation manner of this application; Figure 4 is a schematic diagram of the queue provided in a possible implementation manner of this application; Figure 5 is a schematic diagram showing the non-existence of merged rows in the queue provided in a possible implementation manner of this application; Figure 6 is a schematic diagram showing the existence of merged rows in the queue provided in a possible implementation manner of this application; Figure 7 is a schematic diagram of the queue after the merged rows are removed provided in a possible implementation manner of this application; Figure 8 is a schematic diagram of the queue provided in another possible implementation manner of this application; Figure 9 is a schematic diagram showing the removal of data segments in the queue provided in a possible implementation manner of this application; Figure 10It is a schematic diagram of a rendering system provided in another possible implementation manner of the present application; Figure 11 It is a schematic diagram of a rendering system provided in yet another possible implementation manner of the present application; Figure 12 It is a schematic diagram of a rendering system provided in another possible implementation manner of the present application; Figure 13 It is a flowchart of a rendering method applied to a rendering system provided in a possible implementation manner of the present application; Figure 14 It is a simplified structural block diagram of an electronic device provided in a possible implementation manner of the present application. Detailed implementation manners
[0015] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0016] Before introducing the rendering system proposed by the present application, the following brief description is made on the image rendering technology involved in the related art.
[0017] In the related art, the GPU architecture can be generally divided into two types. One is Immediate Mode Rendering (IMR), and the other is TBR. Among them, TBR is also called tile-based rendering.
[0018] In some embodiments, TBR mainly includes the following parts.
[0019] Chunk rendering. TBR divides the screen of the electronic device into several small rectangular blocks, or called tiles, and each tile is rendered independently. The GPU first converts all geometric data into screen space coordinates, and then classifies and processes these data according to the tiles they cover. For the primitive data in each tile, the GPU will concatenate them and store them in the external memory. Exemplarily, the external memory is the video memory.
[0020] Tile Buffer. When processing each tile, the GPU stores all necessary geometric and texture data in a small cache (usually on-chip storage). This method allows the GPU to reduce access to external memory, thereby improving efficiency. Exemplarily, the tile buffer is the following intermediate buffer.
[0021] Tile-by-tile rendering. The GPU processes each tile in turn until the entire screen is rendered. After the rendering of each tile is completed in its independent tile buffer, the result is written into the final frame buffer.
[0022] Exemplarily, the main advantages of TBR include the following aspects.
[0023] On the one hand, memory bandwidth saving. Since the TBR architecture only needs to access local geometric data and texture data within each tile, it can significantly reduce the read and write operations to external memory, thereby reducing the memory bandwidth requirement.
[0024] On the other hand, energy saving. Reducing the demand for memory bandwidth also means a reduction in power consumption. Therefore, the TBR architecture is widely used on mobile devices.
[0025] On yet another hand, some algorithms can be used to reduce Over Draw to reduce the workload of the fragment shader.
[0026] Exemplarily, the main disadvantages of TBR include the following aspects.
[0027] On the one hand, the process is complex. The process of tiled rendering is relatively complex, including steps such as tiling, storing to the processor, and then reading. Especially when dealing with advanced graphics effects (such as global illumination, complex transparency processing), more hardware and software support may be required.
[0028] On the other hand, there is latency. Since TBR needs to process the screen in tiles, there may be a certain latency in the entire rendering process, especially in scenes that need to be updated frequently.
[0029] On yet another hand, there may be insufficient memory. Since a large amount of data (primitive list) needs to be output to the intermediate buffer (i.e., the above-mentioned external memory), it may lead to Out of Memory (OOM).
[0030] In summary, for TBR, the entire rendering process needs to be disassembled into the following two steps. The entire rendering process includes a geometry processing pipeline and a pixel processing pipeline. Among them, the geometry processing pipeline will perform a series of processing on the data stream, including operations such as vertex shader, tessellator, geometry shader, and viewport transformation. Then, the data result will be divided into many data streams (primitive list) based on tiles and stored in the external memory. After this work is completed, the pixel processing pipeline will read these data streams, perform rasterization and fragment shader, depth and stencil testing, and blending output.
[0031] For a multi-core system, if multiple GPU core units are enabled (each core contains a geometry processing pipeline and a pixel processing pipeline), the following two problems will be encountered.
[0032] First, since the input data stream is serial, if it passes through multiple geometry processing pipelines, a re-ordering process is required before being processed by the pixel processing pipeline. Exemplarily, to solve this problem, a synchronization unit is added at the end of the geometry processing pipeline. However, this may result in performance loss due to mutual waiting.
[0033] Second, due to the existence of the tessellator and the geometry shader, the number of primitives can be inflated several times. Therefore, it is impossible to estimate the storage space for the data stream output by the geometry processing pipeline, so there may be a problem of insufficient storage space.
[0034] Based on the above problems, the present application proposes a rendering system. A rendering pipeline manager is introduced in the rendering system to uniformly manage the data segments processed by the geometry processing pipeline. Specifically, the rendering pipeline manager obtains the numbers and memory information of the data segments processed by the geometry processing pipeline. After there are k consecutive data segments, rendering instructions are sent to each of the N pixel processing pipelines. That is, as long as there are k consecutive data segments, regardless of the processing progress of the geometry processing pipeline, the rendering pipeline manager will send a rendering instruction to indicate rendering. Therefore, each geometry processing pipeline does not need to wait for other geometry processing pipelines to synchronize. Through the rendering pipeline manager, asynchronous operation of multiple geometry processing pipelines can be achieved (each geometry processing pipeline continuously processes data segments without maintaining synchronization). Therefore, in the embodiments of the present application, the latency of image rendering is reduced, thereby improving the efficiency of image rendering. For specific explanations, please refer to the following embodiments.
[0035] Please refer to Figure 1 , which shows a schematic diagram of the rendering system provided in a possible implementation manner of the present application.
[0036] In some embodiments, as Figure 1 shown, the rendering system includes: a rendering pipeline manager 120, an intermediate buffer 110, M geometry processing pipelines 100, and N pixel processing pipelines 130, where M is an integer greater than 1 and N is an integer greater than or equal to 1; among them, multiple data segments divided from the input data stream of the rendering system are assigned to the M geometry processing pipelines 100 for processing, and the primitive data obtained after each data segment is processed is stored in the intermediate buffer 110.
[0037] Exemplarily, the rendering system includes multiple GPU cores. Exemplarily, the internal processing flow of the GPU core is as Figure 2 shown. Exemplarily, as Figure 2 shown, a GPU core includes a geometry processing pipeline 210, a tile divider 220, an intermediate buffer 230, and a pixel processing pipeline 240.
[0038] Exemplarily, a geometry processing pipeline is a module for processing an input data stream or data segments divided based on the input data stream. Exemplarily, the processing of data segments by the geometry processing pipeline includes the following aspects: operations such as vertex shader, tessellator, geometry shader, viewport transformation, etc. Exemplarily, each GPU core includes one or more geometry processing pipelines. The number of geometry processing pipelines in different GPUs may be the same or different. In some embodiments, a GPU core includes M geometry processing pipelines, that is, the number of GPU cores corresponding to M geometry processing pipelines is 1. In some other embodiments, a GPU core includes 1 geometry processing pipeline, that is, the number of GPU cores corresponding to M geometry processing pipelines is M.
[0039] Exemplarily, a pixel processing pipeline is a module for reprocessing the primitive data obtained after being processed by the geometry processing pipeline to obtain a final rendering result. Exemplarily, the processing of primitive data by the pixel processing pipeline includes at least one of the following: performing rasterization and fragment shader, depth and stencil testing, and blending output. Exemplarily, each GPU core includes one or more pixel processing pipelines. The number of pixel processing pipelines in different GPUs may be the same or different. Exemplarily, the pixel processing pipelines can also be distributed outside the GPU. Exemplarily, each pixel processing pipeline is used to process the rendering of multiple tiles on the screen. In some embodiments, a GPU core includes N pixel processing pipelines, that is, the number of GPU cores corresponding to N pixel processing pipelines is 1. In some other embodiments, a GPU core includes 1 pixel processing pipeline, that is, the number of GPU cores corresponding to N pixel processing pipelines is N. In some embodiments, M is equal to N.
[0040] Exemplarily, an intermediate buffer is a storage space for storing the primitive data obtained after each data segment is processed. Exemplarily, the intermediate buffer is distributed in the GPU core or can also be distributed outside the GPU core. Exemplarily, when the intermediate buffer is distributed in the GPU core, each GPU core includes one or more intermediate buffers. Exemplarily, when the intermediate buffer is distributed outside the GPU core, the intermediate buffer is also referred to as external memory or video memory.
[0041] Exemplarily, the primitive data is data related to tile rendering. Exemplarily, the primitive data includes the following aspects of information. The color information of the tile: stored in the tile buffer, used to record the color data of each tile. The depth information: stored in the Depth Buffer, used to record the depth information of each tile. The stencil information: stored in the Stencil Buffer, used to record the stencil information of each tile. Exemplarily, the intermediate buffer includes at least one of the above-mentioned tile buffer, depth buffer, and stencil buffer.
[0042] In some embodiments, the input data stream is divided into multiple data segments. Exemplarily, the multiple data segments are assigned to M geometric processing pipelines for processing. Exemplarily, each data segment is assigned to a geometric processing pipeline for processing. Exemplarily, there are no identical data segments, any two data segments are different from each other, or there is no overlapping part between any two data segments.
[0043] The embodiments of the present application do not limit the way of dividing the input data stream, and multiple data segments are obtained by dividing the input data stream according to a preset division rule. Exemplarily, the input data stream is the original data input to the rendering system for rendering an image. Exemplarily, the input data stream includes a plurality of data related to triangle meshes, such as vertex data of triangle meshes. Exemplarily, the vertex data is data provided for subsequent geometric processing pipelines, including at least one of vertex attributes such as vertex coordinates, texture coordinates, vertex normals, and vertex colors.
[0044] In some embodiments, the rendering pipeline manager is used to obtain the numbers and memory information of the data segments processed by the geometric processing pipelines. The memory information of the data segments is used to determine the memory pages occupied by the primitive data corresponding to the data segments in the intermediate buffer. Exemplarily, as Figure 1 shown, the rendering pipeline manager 120 is used to obtain the numbers and memory information of the data segments processed by the geometric processing pipelines. The memory information of the data segments is used to determine the memory pages occupied by the primitive data corresponding to the data segments in the intermediate buffer.
[0045] Exemplarily, when dividing the input data stream, numbers are respectively assigned to the multiple obtained data segments to obtain the numbers of each data segment. Exemplarily, starting from 0, the numbers of each data segment are sequentially assigned, and the number of each data segment is a non-negative number. Exemplarily, the numbers of adjacent data segments are also consecutive.
[0046] Exemplarily, data segments are allocated to M geometric processing pipelines in a certain order. Exemplarily, the method of allocating data segments to M geometric processing pipelines is as follows. Exemplarily, a third value is set. Starting from the initial data segment, a data segment with a quantity equal to the third value is allocated to the first geometric processing pipeline among the M geometric processing pipelines, a data segment with a quantity equal to the third value is allocated to the second geometric processing pipeline among the M geometric processing pipelines, a data segment with a quantity equal to the third value is allocated to the third geometric processing pipeline among the M geometric processing pipelines, and so on, until each geometric processing pipeline among the M geometric processing pipelines is allocated a data segment with a quantity equal to the third value. In the case where there are still unallocated data segments, starting from the first geometric processing pipeline among the M geometric processing pipelines, data segments with a quantity equal to the third value are continuously allocated until all data segments are allocated. Exemplarily, the third value is 1.
[0047] In some embodiments, the input data stream is simultaneously fed into the geometric processing pipelines of multiple GPU cores. Exemplarily, the entire data stream is segmented by a certain algorithm and processed separately in different geometric processing pipelines. This algorithm can segment the entire data stream according to certain rules, and the distribution of these segments adopts a polling method in the geometric processing pipelines of multiple GPU cores. Exemplarily, a number can be assigned to each data stream segment (a data segment). Exemplarily, if it is a 4-core GPU system (including GPU core 0 (or GPU0), GPU core 1 (or GPU1), GPU core 2 (or GPU2), and GPU core 3 (or GPU3)), data segment 0 (0 is the number of this data segment) will be distributed to the geometric processing pipeline on GPU0, data segment 1 will be distributed to the geometric processing pipeline on GPU1, data segment 2 will be distributed to the geometric processing pipeline on GPU2, data segment 3 will be distributed to the geometric processing pipeline on GPU3, and data segment 4 will be distributed to the geometric processing pipeline on GPU0 again, and so on. Exemplarily, the pixel processing pipeline needs to re-sort the primitve data stream that has been distributed in a scrambled order and processed by the geometric processing pipeline back to the order of the original input data stream.
[0048] This method can ensure that the number of data segments allocated to each geometric processing pipeline is relatively consistent, without significant quantity differences, which is conducive to balancing the number of data segments required by each geometric processing pipeline and achieving parallel processing.
[0049] Exemplarily, each data segment is numbered sequentially according to the order in which it is assigned to the geometric processing pipeline for processing. Exemplarily, each geometric processing pipeline processes the data segments one by one in the order in which the data segments are received. Exemplarily, the geometric processing pipeline first processes the data segments with smaller numbers and then processes the data segments with larger numbers, such as processing the data segments numbered 0, 4, 8, and 12 in sequence. Exemplarily, the geometric processing pipeline processes the allocated data segments in the order of data segment allocation.
[0050] In some embodiments, the memory information of the data segment is used to determine the memory pages occupied by the primitive data corresponding to the data segment in the intermediate buffer. Exemplarily, the memory information of the data segment includes at least one of the number of memory pages, storage address, etc. Exemplarily, during the process of processing the data stream, the geometric processing pipeline continuously applies for and uses memory pages.
[0051] Exemplarily, if the primitive data is continuously occupied in the memory buffer, the memory information includes the number of memory pages. Exemplarily, the position where the primitive data is stored in the intermediate buffer can be determined according to the number of memory pages. Exemplarily, a memory page corresponds to a small storage space in the memory space, and the specific size of this storage space is not limited in this application. Exemplarily, since it is continuously occupied, the storage address of the previously occupied memory pages can be determined according to the number of previously occupied memory pages. Further, the storage address of the memory pages occupied by the current primitive data can be determined according to the number of memory pages occupied by the current primitive data. Exemplarily, whenever a data segment is processed, the geometric processing pipeline sends the number of the data segment and the number of memory pages applied for during the processing of the data segment to the rendering pipeline manager, and the rendering pipeline manager is responsible for collecting the data sent separately by the geometric pipeline processors of all GPU cores.
[0052] Exemplarily, if the primitive data is non - continuously occupied in the memory buffer, also known as out - of - order occupation, the memory information includes the storage address. Exemplarily, the position where the primitive data is stored in the memory buffer can be determined only according to the specific storage address.
[0053] This application determines the storage address by the number of memory pages, which can reduce the data transmission cost and data recording cost (for example, originally K bits were required to record the specific storage address, and in this application, only one bit is required to record the number of memory pages, where K is an integer greater than 1), thereby reducing the amount of data processing during image rendering. Of course, the method of directly recording the specific storage address can improve the accuracy of storage address determination and avoid the occurrence of subsequent storage address calculation errors caused by errors in recording the number of previous memory pages.
[0054] In some embodiments, the rendering pipeline manager is further configured to send rendering instructions to N pixel processing pipelines respectively. The rendering instructions include the numbers of k data segments and memory information, where k is an integer greater than or equal to 1. When k is greater than 1, the numbers of the k data segments are consecutive. Exemplarily, as Figure 1 shown, the rendering pipeline manager 120 is further configured to send rendering instructions to N pixel processing pipelines 130 respectively. The rendering instructions include the numbers of k data segments and memory information, where k is an integer greater than or equal to 1. When k is greater than 1, the numbers of the k data segments are consecutive.
[0055] In some embodiments, since the pixel processing pipeline has the requirement of rendering in order for the input data stream (such as color mixing and depth testing, etc.), the data of the workflow sent to the pixel processing pipeline needs to be in order. Or, it can also be understood that the data segments where the primitive data for rendering sent to the pixel processing pipeline is located need to be consecutive. Therefore, although the data segments processed by different geometry processing pipelines are different and the operating speeds of different geometry processing pipelines are inconsistent, by assigning numbers to each data segment, the pixel processing pipeline can reorder and process them by obtaining the numbers of each data segment.
[0056] In some embodiments, since the operating speeds of each geometry processing pipeline are different, when the rendering pipeline manager triggers the pixel processing pipeline workflow, the data segments in the current queue may not be consecutive. Exemplarily, as Figure 4 shown, the data segments numbered 0, 1, 2, 3, 4, 5, 7, 8 have been completed, but the others have not. Exemplarily, if all of these are packaged and sent to the pixel processing pipeline at this time, it will cause an error. Because the data segments numbered 0, 1, 2, 3, 4, 5, 7, 8 are not a continuous interval, there is a "hole" with the data segment number 6 in the middle, which does not conform to the order-preserving principle. Therefore, when triggering the pixel processing pipeline workflow once (that is, when sending the rendering instructions), consecutive data segment numbers are selected as the pixel processing pipeline workflow (as the numbers of the data segments submitted to the pixel processing pipeline for processing). Exemplarily, in the above example, the data segments numbered 0, 1, 2, 3, 4, 5 can be selected and submitted to the pixel processing pipeline, while the data segments numbered 7 and 8 will not be included in this pixel processing pipeline workflow. Exemplarily, in the above example, k consecutive ones among the data segments numbered 0, 1, 2, 3, 4, 5 can be selected and submitted to the pixel processing pipeline.
[0057] Exemplarily, the rendering pipeline manager obtains the numbers of the data segments processed by the geometry processing pipeline. When the numbers of k data segments are consecutive, the rendering pipeline manager is further configured to send rendering instructions to N pixel processing pipelines respectively. Exemplarily, the k data segments correspond to the workflows of the triggered pixel processing pipelines.
[0058] Exemplarily, the rendering instruction is used to instruct the pixel processing pipeline to perform the rendering process based on the processed primitive data respectively corresponding to the above k data segments. Exemplarily, the rendering pipeline manager sends the rendering instructions to each of the N pixel processing pipelines respectively.
[0059] Exemplarily, for each of the N pixel processing pipelines, based on the rendering instruction, the numbers and memory information of the k data segments are obtained. Exemplarily, each pixel processing pipeline obtains the primitive data respectively corresponding to the k data segments from the corresponding memory pages based on the memory information of the k data segments. Further, based on the primitive data respectively corresponding to the k data segments, the rendering process corresponding to the pixel processing pipeline is performed.
[0060] Exemplarily, the present application does not limit the specific value of k. Exemplarily, k is a preset value. Exemplarily, k is the number of GPU cores included in the rendering system, such as 4. Exemplarily, the numbers of the data segments processed by the geometry processing pipeline obtained by the rendering pipeline manager include numbers 0 to 10, and numbers 0 to 3 are the numbers of the data segments that have been submitted to the N pixel processing pipelines for processing, then numbers 4 to 7 are used as the k data segments indicated by the next sent rendering instruction. That is, the number of data segments indicated by each rendering instruction is the same. Of course, the number of data segments indicated by each rendering instruction may also be different.
[0061] In some embodiments, as Figure 3 shown, multiple data segments divided from the input data stream of the rendering system are assigned to M geometry processing pipelines (including geometry processing pipeline 310) for processing, and the primitive data obtained after processing each data segment is stored in an intermediate buffer (such as intermediate buffer 0 320 corresponding to geometry processing pipeline 310). The rendering pipeline manager 330 is configured to obtain the numbers and memory information of the data segments processed by the geometry processing pipelines (including geometry processing pipeline 310), and the memory information of the data segments is used to determine the memory pages occupied by the primitive data corresponding to the data segments in the intermediate buffer. The rendering pipeline manager 330 is further configured to send rendering instructions to N pixel processing pipelines (including pixel processing pipeline 340) respectively, and the rendering instructions include the numbers and memory information of k data segments, k is an integer greater than or equal to 1, and when k is greater than 1, the numbers of the k data segments are consecutive.
[0062] The rendering system proposed in this application introduces a rendering pipeline manager to uniformly manage the data segments processed by the geometry processing pipeline. Specifically, the rendering pipeline manager obtains the numbers and memory information of the data segments processed by the geometry processing pipeline. After there are k consecutively numbered data segments, it sends rendering instructions to N pixel processing pipelines respectively. That is, as long as there are k consecutively numbered data segments, regardless of the processing progress of the geometry processing pipeline, the rendering pipeline manager will send rendering instructions to indicate rendering. Therefore, each geometry processing pipeline does not need to wait for other geometry processing pipelines to synchronize. Through the rendering pipeline manager, asynchronous operation of multiple geometry processing pipelines can be achieved (each geometry processing pipeline continuously processes data segments without maintaining synchronization). Therefore, the latency of image rendering is reduced in the embodiments of this application, thereby improving the efficiency of image rendering.
[0063] In some embodiments, the rendering pipeline manager is further configured to select one or more data segments with numbers between a first value and a second value from the data segments processed by the geometry processing pipeline as the k data segments.
[0064] In some embodiments, the first value is 1 plus the maximum number of the data segments that have been submitted to the N pixel processing pipelines for processing, and the second value is the maximum number that satisfies the continuity of the numbers of the k data segments.
[0065] Exemplarily, for the data segments that have been submitted to the N pixel processing pipelines for processing, the rendering pipeline manager will not submit the numbers or memory information of these data segments again.
[0066] Exemplarily, if the numbers of the data segments that have been submitted to the N pixel processing pipelines for processing are 0 to 4, then the first value is 1 plus the maximum number 4 of the data segments that have been submitted to the N pixel processing pipelines for processing, that is, the first value is 5.
[0067] Exemplarily, the numbers of the data segments processed by the geometry processing pipeline obtained by the rendering pipeline manager include 0 to 10 and 12 to 13, where the numbers 0 to 4 are the numbers of the data segments that have been submitted to the N pixel processing pipelines for processing. Then the first value is 5, and the second value is 10. The second value is the maximum number that satisfies the continuity of the numbers of the k data segments. Starting from the data segment numbered 5, the data segments numbered 5 to 10 are consecutive processed data segments. Exemplarily, there are 6 consecutive data segments between the numbers 5 and 10, then one or more data segments with numbers between the first value 5 and the second value 10 are selected as the k data segments.
[0068] Exemplarily, when triggering the pixel processing pipeline workflow once, the maximum consecutive data segment number is selected as the pixel processing pipeline workflow. Exemplarily, in the above example, the numbers of the data segments numbered 0, 1, 2, 3, 4, and 5 are selected and submitted to the pixel processing pipeline, while the data segments numbered 7 and 8 will not be included in this pixel processing pipeline workflow.
[0069] Exemplarily, since the data segment numbered 11 has not been processed yet, there is a discontinuity between the data segments numbered 10 and 12. Therefore, even if the data segment numbered 12 is processed, it is not considered as one of the k data segments.
[0070] The technical solution provided by the embodiments of the present application takes the data segment with the most consecutive numbers (also called the maximum consecutive data segment number) from the data segments processed by the geometric processing pipeline by the rendering pipeline manager as the k data segments and sends rendering instructions, without pausing the geometric processing pipeline to wait for other geometric processing pipelines to synchronize, realizing the asynchronous operation of the geometric processing pipeline, which is beneficial to improving the GPU operation performance.
[0071] In addition, it realizes rendering as many consecutive data segments as possible at one time, which is beneficial to accelerating the rendering speed and improving the graphics rendering efficiency. And it does not resend the data segments that have been submitted to the pixel processing pipeline for processing, which is beneficial to avoiding duplicate sending and reducing the cost of data sending and processing.
[0072] In some embodiments, there are M queues set in the rendering pipeline manager, and each queue is used to store the number and memory information of the data segments processed by a geometric processing pipeline.
[0073] Exemplarily, the M queues and the M geometric processing pipelines correspond one by one, and one queue corresponds to one geometric processing pipeline. Exemplarily, taking a system with 4 GPU cores as an example, there are 4 queues in the rendering pipeline manager, which are used to store the numbers and memory information of the processed data segments (such as the numbers of the processed data segments and the memory information of the primitive data corresponding to each numbered data segment) sent by the geometric processing pipelines of 4 GPU cores respectively.
[0074] Exemplarily, each queue includes multiple grids, and one grid is used to store the number of a completed data segment and the memory information of the primitive data corresponding to this data segment (such as the number of memory pages or the so-called page number). Exemplarily, M is 4, and the four queues correspond to four GPU cores. Exemplarily, at a certain moment, the queues in the rendering pipeline manager are as Figure 4As shown. Among them, the gray blocks are the information that has been received at the current moment, while the white blocks are the information that has not been received at the current moment. It can be seen that the operating speeds of different GPU cores are different, that is, the data segments processed by different geometry processing pipelines are different, so the amount of data received by each queue is also different. As Figure 4 shown, in the first queue (that is, queue 410 corresponding to GPU core 0), the data segments processed by the geometry processing pipeline include data segments 0, 4, and 8. As Figure 4 shown, in the second queue (that is, the queue corresponding to GPU core 1), the data segments processed by the geometry processing pipeline include data segments 1 and 5. As Figure 4 shown, in the third queue (that is, the queue corresponding to GPU core 2), the data segment processed by the geometry processing pipeline includes data segment 2. As Figure 4 shown, in the fourth queue (that is, the queue corresponding to GPU core 3), the data segments processed by the geometry processing pipeline include data segments 3 and 7.
[0075] The technical solution provided by the embodiments of the present application uses M queues to separately record the data sent by each of the M geometry processing pipelines, so that the rendering pipeline manager has a relatively clear understanding of the operating conditions of the M geometry processing pipelines (including which data segments are processed and the memory information of the corresponding primitive data). It is beneficial to trigger the execution of subsequent steps of sending rendering instructions and improve the processing efficiency of the rendering pipeline manager.
[0076] In some embodiments, the i-th queue among the M queues is used to store the numbers and memory information of the data segments processed by the i-th geometry processing pipeline among the M geometry processing pipelines, where i is a positive integer less than or equal to M.
[0077] In some embodiments, the rendering pipeline manager is further configured to add the number and memory information of the first data segment to the end of the i-th queue after the i-th geometry processing pipeline finishes processing the first data segment.
[0078] Exemplarily, for the i-th queue, the numbers and memory information of the completed data segments are sequentially added to the i-th queue in the order of completion of the data segments. Exemplarily, the order of the numbers and memory information of the data segments processed in the i-th queue corresponds one-to-one with the order of completion of the data segments.
[0079] In some embodiments, the rendering pipeline manager is further configured to remove the number and memory information of the second data segment from the i-th queue when the second data segment stored in the i-th queue meets the condition of being submitted to the pixel processing pipeline.
[0080] Exemplarily, the condition for being submitted to the pixel processing pipeline is a preset submission condition. Exemplarily, the condition for being submitted to the pixel processing pipeline is a condition for submitting a data segment. Exemplarily, when the second data segment stored in the i-th queue is one of the k data segments, it is considered that the second data segment meets the condition for being submitted to the pixel processing pipeline. Exemplarily, when the second data segment stored in the i-th queue is a data segment numbered between a first value and a second value, it is considered that the second data segment meets the condition for being submitted to the pixel processing pipeline.
[0081] Exemplarily, when the number and memory information of the second data segment stored in the i-th queue are removed from the i-th queue, the number and memory information of the next data segment after the second data segment stored in the i-th queue occupy the storage position corresponding to the number and memory information of the second data segment in the original i-th queue.
[0082] According to the technical solution provided by the embodiment of the present application, when the second data segment will be or has been submitted to the pixel processing pipeline for processing, the rendering instruction already carries the number and memory information of the second data segment. Therefore, the number and memory information of the second data segment are removed from the i-th queue. Under the premise of not affecting the image rendering, it is possible to avoid the number and memory information of the submitted data segment from occupying too much memory space, which is conducive to saving storage space.
[0083] In some embodiments, the rendering pipeline manager is also used to remove the numbers and memory information of the M data segments included in the merged row from the M queues when there is a merged row in the M queues; wherein the merged row includes the number and memory information of a data segment with the smallest number stored in each queue in the M queues, and the numbers of the M data segments included in the merged row are continuous.
[0084] Exemplarily, for the number and memory information of a data segment with the smallest number stored in each of the M queues, if the numbers of a data segment with the smallest number stored in each queue are continuous, then the data segment with the smallest number stored in each of the M queues is used as the data segment in the merged row.
[0085] Exemplarily, M is 4. Exemplarily, queue 1 includes number 0 and number 4, queue 2 includes number 1, queue 3 includes number 2, and queue 4 includes number 3. Then, the number of the smallest numbered data segment in queue 1 is number 0, the number of the smallest numbered data segment in queue 2 is number 1, the number of the smallest numbered data segment in queue 3 is number 2, and the number of the smallest numbered data segment in queue 4 is number 3. Since the numbers of number 0, number 1, number 2, and number 3 are consecutive, number 0, number 1, number 2, and number 3 are used as the numbers in the merged row.
[0086] Exemplarily, as Figure 4 shown, the data segments in the first row of the four queues are all completed and are consecutively numbered. Then, all the elements in the first row are merged, and each element corresponds to the number of a data segment and the memory information of the data segment (including the number of pages).
[0087] Exemplarily, a cumulative value is set for each queue respectively, with an initial value of 0. Exemplarily, the cumulative value of each queue is used to indicate the data segments merged in the queue or the data segments that have been submitted to the pixel processing pipeline (including the number of pages). Exemplarily, when there is an element in all positions in the merged row, the first element of each queue can be popped out (i.e., removed from the queue), and the number of pages is added to the corresponding cumulative value, and so on. In some embodiments, the rendering pipeline manager also records the numbers of the data segments in this row and merges them into the largest consecutive data segment number.
[0088] Exemplarily, as Figure 5 shown, at time t0, there are no processed data segments in queue 510 where GPU core 2 is located. Then, there is no merged row (the processed data segments in queue 510 are lacking in the first row), and no removal is triggered.
[0089] Exemplarily, as Figure 6 shown, at time t2, there are processed data segments in queue 610 where GPU core 2 is located (the processed data segments in queue 610 are included in the first row). Then, there is a merged row, and removal is triggered. Exemplarily, at this time, the cumulative results of the cumulative values of each queue are: GPU0: the number of pages A, GPU1: the number of pages B, GPU2: the number of pages C, GPU3: the number of pages D, and the largest consecutive data segment number is [0:3]. Among them, the number of pages A is used to indicate the number of memory pages of the primitive data corresponding to data segment 0, and so on for the others. A to D are positive integers.
[0090] Exemplarily, based on Figure 6 the merged row, after removing the merged row from the M queues, the result obtained is as Figure 7 shown. Exemplarily, as Figure 7 shown, after data segment 0 and the number of pages A are removed from queue 710 where GPU core 0 is located, the data segment 4 and the number of pages E in the second grid are moved up one grid. The same applies to other queues.
[0091] Exemplarily, as the GPU continues to run, that is, as the geometry processing pipeline continues to run, the resulting queue results are as Figure 8As shown. Assume that at this time, the rendering pipeline manager triggers a pixel processing pipeline workflow, that is, sends a rendering instruction. Then it will first find the largest consecutive data segment number. Since the consecutive number has been found once in the previous round of merging rows, that is, the number of the merging rows is [0:3], and two data segments that can be spliced with the result found in the previous round can still be found in the queue at this time, thus becoming the new largest consecutive data segment. These two data segments are the data segment 4 recorded in the first element 810 of the queue where GPU core 0 is located and the data segment 5 recorded in the first element 820 of the queue where GPU core 1 is located. The merged result is [0:5]. Then, at this time, the rendering pipeline manager accumulates the page numbers of data segments 4 and 5 into their respective cumulative values, and then removes them from their respective queues, obtaining four queues as shown in Figure 9 As shown. The four queues do not include data segment 4 and data segment 5. In some embodiments, the cumulative values of each queue at this time are GPU0: page number A + E, GPU1: page number B + F, GPU2: page number C, GPU3: page number D, and the largest consecutive data segment number is [0:5]. Exemplarily, the rendering pipeline manager will initiate a pixel processing pipeline workflow (that is, send a rendering instruction), and also send the content of this rendering data segment to each pixel processing pipeline, that is, send the largest consecutive data segment number [0:5] to each pixel processing pipeline. The elements here are the grids mentioned above.
[0092] In some embodiments, the rendering pipeline manager is further configured to record the to-be-processed information corresponding to M queues respectively; wherein, for the i-th queue among the M queues, the to-be-processed information corresponding to the i-th queue is used to determine the data segments that have been removed from the i-th queue and not submitted to the pixel processing pipeline.
[0093] Exemplarily, the to-be-processed information includes the above cumulative value and the consecutive data segment number. Exemplarily, the cumulative value of each queue is used to indicate the sum of the page numbers corresponding to the data segments removed from the queue. Exemplarily, the consecutive data segment number is used to indicate the data segments that have been removed from the i-th queue and not submitted to the pixel processing pipeline.
[0094] The technical solution provided by the embodiments of the present application takes into account that the length of each queue is fixed. Therefore, when the merging condition is met, the numbers and memory information of the M data segments included in the merging rows are removed from the M queues, thereby avoiding the situation that there is too much data in the queue and new data cannot be placed, and realizing the reasonable allocation of memory resources. In addition, the to-be-processed information is used to indicate the data segments that have been removed from the i-th queue and not submitted to the pixel processing pipeline, so as to prompt the rendering pipeline manager to submit these data segments, avoiding the situation of missed submission of data segments and ensuring the accuracy of image rendering.
[0095] In some embodiments, the rendering pipeline manager is further configured to, when the memory occupancy rate corresponding to any one of the geometry processing pipelines is greater than or equal to a set threshold, perform the step of separately sending rendering instructions to N pixel processing pipelines.
[0096] Exemplarily, the memory occupancy rate corresponding to a geometry processing pipeline is the ratio of the number of memory pages occupied by the primitive data of the data segment that has been processed by the geometry processing pipeline and has not been submitted to the pixel processing pipeline in the intermediate buffer to the total number of memory pages included in the intermediate buffer. Exemplarily, a certain amount of memory space is pre-allocated for each GPU core where a geometry processing pipeline is located. Exemplarily, the allocated memory space is the intermediate buffer.
[0097] Exemplarily, the memory occupancy rates corresponding to different geometry processing pipelines are different at different processing stages. Exemplarily, when the memory occupancy rate corresponding to one of the geometry processing pipelines is greater than or equal to the set threshold, the step of separately sending rendering instructions to N pixel processing pipelines is performed.
[0098] Exemplarily, the intermediate buffer includes S memory pages, where S is a positive integer, and the set threshold is a%, where a is a positive number. For the rendering pipeline manager, if the number of memory pages occupied by any one of the geometry processing pipelines reaches S×a%, a pixel processing pipeline workflow needs to be initiated, that is, a rendering instruction is sent once.
[0099] Exemplarily, during the processing of the pixel processing pipeline, even if the memory occupancy rate corresponding to any one of the geometry processing pipelines is greater than or equal to the set threshold again, it is necessary to wait until the pixel processing pipeline finishes processing the previously sent rendering instruction, and then the rendering pipeline manager sends the rendering instruction again.
[0100] In some embodiments, when the memory occupancy rate corresponding to any one of the geometry processing pipelines is greater than or equal to the set threshold, the geometry processing pipeline does not stop working. That is to say, sending the rendering instruction and triggering the rendering process of the pixel processing pipeline are also asynchronous with the geometry processing pipeline, and the dependent parts of the two rendering pipelines (including the geometry processing pipeline and the pixel processing pipeline) can be decoupled.
[0101] Considering that if multi-core mutual waiting occurs at a fixed position, a deadlock problem may occur due to insufficient memory space. Specifically, the deadlock problem refers to the situation where if the geometry processing pipeline of a certain GPU core runs out of space before reaching the synchronization point (that is, the memory occupancy rate reaches 100% and it still has not reached the fixed position), other cores will never wait for the signal that the core reaches the synchronization point. The technical solution provided by the embodiments of the present application sends a rendering instruction when the memory occupancy rate corresponding to any geometry processing pipeline is greater than or equal to the set threshold to avoid the occurrence of the deadlock problem. The set threshold proposed by the present application ensures that the rendering instruction is triggered before the memory is exhausted, so the deadlock problem caused by insufficient memory can be avoided.
[0102] In some embodiments, the j-th pixel processing pipeline among the N pixel processing pipelines is configured to obtain the primitive data corresponding to k data segments from the intermediate buffer according to the rendering instruction, render the primitive data corresponding to the k data segments, and after the rendering of the primitive data corresponding to the k data segments is completed, send a completion indication message to the rendering pipeline manager. The completion indication message is used to indicate that the j-th pixel processing pipeline has completed the rendering of the primitive data corresponding to the k data segments, where j is a positive integer less than or equal to N.
[0103] Exemplarily, the completion indication message is used to indicate that the j-th pixel processing pipeline has completed the rendering of the primitive data corresponding to the k data segments. Exemplarily, each pixel processing pipeline sends a completion indication message to the rendering pipeline management. After the N pixel processing pipelines have all sent the completion indication messages, that is, when the rendering pipeline manager receives the completion indication messages sent respectively by the N pixel processing pipelines, it is considered that all the pixel processing pipelines have completed the rendering of the primitive data for the k data segments.
[0104] In some embodiments, the rendering pipeline manager is further configured to release the memory pages occupied by the primitive data corresponding to the k data segments in the intermediate buffer after all the N pixel processing pipelines have completed the rendering of the primitive data corresponding to the k data segments.
[0105] Exemplarily, after the pixel processing pipeline has completed the rendering of the k data segments, it will send back the rendering completion signal (that is, the completion indication message) to the rendering pipeline manager. The rendering pipeline manager will release the memory pages respectively occupied by the primitive data corresponding to the k data segments in the intermediate buffer.
[0106] Exemplarily, as shown in the above embodiments, after the pixel processing pipeline finishes rendering data segments 0 - 5, it will reply the rendering completion signal (i.e., the completion indication information) to the rendering pipeline manager. The rendering pipeline manager will perform a release operation on the memory space of each GPU core. The number of released pages is the cumulative value corresponding to the queues of each core recorded previously, that is, the number of pages of the cumulative removed data segments.
[0107] The technical solution provided by the embodiments of the present application recovers and releases the memory space by releasing the memory pages occupied by the primitive data corresponding to the rendered data segments in the intermediate buffer, and allocates it to the primitive data corresponding to the subsequent data segments to be processed. This not only avoids the occurrence of the OOM problem, but also realizes the dynamic and reasonable allocation and recycling of memory.
[0108] In some embodiments, the rendering system further includes: M tile partitioners.
[0109] In some embodiments, the i-th geometric processing pipeline among the M geometric processing pipelines is used to process the data segments allocated to the i-th geometric processing pipeline among the multiple data segments divided from the input data stream, and obtain the processed data segments, where i is a positive integer less than or equal to M.
[0110] Exemplarily, one geometric processing pipeline corresponds to one tile partitioner. Exemplarily, each GPU core includes one or more tile partitioners.
[0111] In some embodiments, the i-th tile partitioner among the M tile partitioners is used to perform tile partitioning on the processed data segments, obtain the primitive data corresponding to the data segments, and store the primitive data corresponding to the data segments in the intermediate buffer.
[0112] Exemplarily, the tile partitioner is used to divide the screen of the electronic device into multiple tiles. Exemplarily, each tile is a rectangular area. Exemplarily, the area of each tile is the same. Exemplarily, the area of each tile may also be different.
[0113] Exemplarily, the i-th tile partitioner among the M tile partitioners is used to partition the processed data segments according to the tiles, and obtain the primitive data corresponding to the data segments.
[0114] Exemplarily, the tile partitioner also stores the primitive data corresponding to the data segments in the intermediate buffer.
[0115] The technical solution provided by the embodiments of the present application divides the data segments processed by the geometric processing pipeline through the tile partitioner to obtain the primitive data corresponding to the data segments. This is conducive to realizing tile partitioning, and thus realizing the subsequent rendering process based on the pixel processing pipeline.
[0116] In some embodiments, the i-th geometry processing pipeline is further configured to pause processing of new data segments when the number of memory pages occupied by the primitive data corresponding to the data segments that have been processed by the i-th geometry processing pipeline and have not been submitted to the pixel processing pipeline in the intermediate buffer reaches the maximum number of pages allocated to the i-th geometry processing pipeline.
[0117] Exemplarily, when the number of memory pages occupied by the primitive data corresponding to the data segments that have been processed by the i-th geometry processing pipeline and have not been submitted to the pixel processing pipeline in the intermediate buffer reaches the maximum number of pages allocated to the i-th geometry processing pipeline, the processing of new data segments is paused.
[0118] Exemplarily, after reaching the set threshold of the occupancy rate, the pixel rendering workflow is triggered. At this time, the geometry processing pipeline of each core can still continue to work. Whenever a data segment is processed, the number of the data segment and the number of pages it occupies are reported to the rendering pipeline manager. Exemplarily, only when all the pre-allocated pages within a core have been applied for and no pages are released back, the geometry processing pipeline will stop working and wait for pages to be released.
[0119] Exemplarily, if the number of memory pages occupied by the primitive data corresponding to the data segments that have been processed by the i-th geometry processing pipeline and have not been submitted to the pixel processing pipeline in the intermediate buffer is the maximum number of pages allocated to the i-th geometry processing pipeline, that is, when the memory space allocated to the i-th geometry processing pipeline has been fully occupied, the processing of new data segments is paused.
[0120] The present application takes into account that there is no extra memory space to store the primitive data corresponding to the subsequent processed data segments at this time, so the processing of new data segments is paused to avoid the situation of no memory space for storage. This is conducive to ensuring the smooth operation of the image rendering process and reducing additional processing overhead.
[0121] In some embodiments, the rendering system includes multiple GPU cores (or core units), and each GPU core includes at least one geometry processing pipeline and at least one pixel processing pipeline.
[0122] Exemplarily, each GPU core includes one geometry processing pipeline and one pixel processing pipeline. Exemplarily, each GPU core further includes one tile divider and one intermediate buffer.
[0123] Exemplarily, each GPU core includes multiple geometry processing pipelines and multiple pixel processing pipelines. Exemplarily, each GPU core further includes multiple tile dividers and multiple intermediate buffers.
[0124] Exemplarily, the geometry processing pipeline in each GPU core is used to process the data segment allocated to that GPU core. Exemplarily, the geometry processing pipeline in each GPU core not only needs to process the primitive data obtained by its own processing, but also needs to process the primitive data obtained by other GPU cores, so as to achieve tile rendering.
[0125] The technical solution provided by the embodiments of the present application configures at least one geometry processing pipeline and at least one pixel processing pipeline for each GPU core, which is beneficial to realizing parallel processing of at least one geometry processing pipeline and at least one pixel processing pipeline through multiple GPU cores, thereby accelerating the efficiency of image rendering.
[0126] In some embodiments, as Figure 10 shown, it shows a schematic diagram of a rendering system provided in another possible implementation manner of the present application. Exemplarily, as Figure 10 shown, each GPU core 1010 includes a geometry processing pipeline, a tile divider, an intermediate buffer, and a pixel processing pipeline. Exemplarily, the rendering pipeline manager 1020 in the rendering system is located outside the GPU core 1010.
[0127] In some embodiments, as Figure 11 shown, it shows a schematic diagram of a rendering system provided in yet another possible implementation manner of the present application. Exemplarily, as Figure 11 shown, each GPU core 1110 includes a geometry processing pipeline, a tile divider, an intermediate buffer, a pixel processing pipeline, and a rendering pipeline manager. Exemplarily, the rendering pipeline manager is located inside the GPU core 1110. Exemplarily, in each rendering process, in the case of multiple GPU cores and rendering pipeline managers located in different GPU cores, only one of the rendering pipeline managers is called. Exemplarily, regardless of the number of GPU cores, only the rendering pipeline manager in the GPU core 1110 is used to trigger the workflow of the pixel processing pipeline. Of course, even if only the rendering pipeline manager in one GPU core is used to trigger the workflow of the pixel processing pipeline, it is also necessary to trigger the workflow of the pixel processing pipelines in different GPU cores respectively to achieve the purpose of rendering.
[0128] In some embodiments, as Figure 12 shown, it shows a schematic diagram of a rendering system provided in another possible implementation manner of the present application. Exemplarily, as Figure 12As shown, each GPU core 1210 includes a geometry processing pipeline, a tiler, and a pixel processing pipeline. Exemplarily, the rendering pipeline manager 1230 in the rendering system is located outside the GPU core 1210. Exemplarily, the intermediate buffer 1220 in the rendering system is located outside the GPU core 1210. At this time, different GPU cores can share different memory spaces in the same intermediate buffer 1220.
[0129] Exemplarily, when the rendering pipeline manager or the intermediate buffer is located outside the GPU core. In this case, the number of the rendering pipeline manager or the intermediate buffer can also be one or more.
[0130] Hereinafter, the technical solutions provided in the embodiments of the present application will be described in detail by method embodiments. For the content not described in the method embodiments, reference may be made to the above embodiments, which will not be elaborated here.
[0131] Please refer to Figure 13 , which shows a flowchart of a rendering method applied to a rendering system provided in a possible implementation manner of the present application. The execution subject of each step of this method can be Figure 1 the rendering pipeline manager in the rendering system shown. This method may include at least one of the following steps (1310~1320).
[0132] In some embodiments, the rendering system includes: a rendering pipeline manager, an intermediate buffer, M geometry processing pipelines, and N pixel processing pipelines, where M is an integer greater than 1 and N is an integer greater than or equal to 1; among them, multiple data segments divided from the input data stream of the rendering system are assigned to the M geometry processing pipelines for processing, and the primitive data obtained after each data segment is processed is stored in the intermediate buffer.
[0133] Step 1310, the rendering pipeline manager obtains the number and memory information of the data segment processed by the geometry processing pipeline. The memory information of the data segment is used to determine the memory page occupied by the primitive data corresponding to the data segment in the intermediate buffer.
[0134] Step 1320, the rendering pipeline manager sends rendering instructions to the N pixel processing pipelines respectively. The rendering instructions include the numbers and memory information of k data segments, where k is an integer greater than or equal to 1. When k is greater than 1, the numbers of the k data segments are consecutive.
[0135] In some embodiments, the rendering pipeline manager selects one or more data segments with numbers between a first value and a second value from the data segments processed by the geometry processing pipeline as k data segments; wherein, the first value is the maximum number of the data segments that have been submitted to the N pixel processing pipelines plus 1, and the second value is the maximum number that satisfies the condition that the numbers of the k data segments are consecutive.
[0136] In some embodiments, M queues are set in the rendering pipeline manager, and each queue is used to store the number and memory information of the data segments processed by a geometry processing pipeline.
[0137] In some embodiments, the i-th queue among the M queues is used to store the number and memory information of the data segments processed by the i-th geometry processing pipeline among the M geometry processing pipelines, where i is a positive integer less than or equal to M.
[0138] After the rendering pipeline manager finishes processing the first data segment by the i-th geometry processing pipeline, it adds the number and memory information of the first data segment to the end of the i-th queue.
[0139] When the second data segment stored in the i-th queue meets the condition of being submitted to the pixel processing pipeline, the rendering pipeline manager removes the number and memory information of the second data segment from the i-th queue.
[0140] In some embodiments, when there are merged rows in the M queues, the rendering pipeline manager removes the numbers and memory information of the M data segments included in the merged rows from the M queues; wherein, the merged rows include the numbers and memory information of one data segment with the smallest number stored in each of the M queues, and the numbers of the M data segments included in the merged rows are consecutive.
[0141] In some embodiments, the rendering pipeline manager records the to-be-processed information corresponding to each of the M queues; wherein, for the i-th queue among the M queues, the to-be-processed information corresponding to the i-th queue is used to determine the data segments that have been removed from the i-th queue and not submitted to the pixel processing pipeline.
[0142] In some embodiments, when the memory occupancy rate corresponding to any one geometry processing pipeline is greater than or equal to a set threshold, the rendering pipeline manager performs the step of respectively sending rendering instructions to the N pixel processing pipelines.
[0143] Wherein, the memory occupancy rate corresponding to the geometry processing pipeline is the ratio of the number of memory pages occupied by the primitive data corresponding to the data segments that have been processed by the geometry processing pipeline and not submitted to the pixel processing pipeline in the intermediate buffer to the total number of memory pages included in the intermediate buffer.
[0144] In some embodiments, the j-th pixel processing pipeline among the N pixel processing pipelines obtains the primitive data corresponding to k data segments from the intermediate buffer according to the rendering instruction, renders the primitive data corresponding to the k data segments respectively, and after the rendering of the primitive data corresponding to the k data segments is completed, sends a completion indication message to the rendering pipeline manager. The completion indication message is used to indicate that the j-th pixel processing pipeline has completed the rendering of the primitive data corresponding to the k data segments, where j is a positive integer less than or equal to N.
[0145] In some embodiments, after all the N pixel processing pipelines have completed the rendering of the primitive data corresponding to the k data segments, the rendering pipeline manager releases the memory pages occupied by the primitive data corresponding to the k data segments in the intermediate buffer.
[0146] In some embodiments, the rendering system further includes: M tile partitioners.
[0147] In some embodiments, the i-th geometry processing pipeline among the M geometry processing pipelines processes the data segments allocated to the i-th geometry processing pipeline among the multiple data segments obtained by partitioning the input data stream, to obtain processed data segments, where i is a positive integer less than or equal to M.
[0148] In some embodiments, the i-th tile partitioner among the M tile partitioners performs tile partitioning on the processed data segments to obtain the primitive data corresponding to the data segments, and stores the primitive data corresponding to the data segments into the intermediate buffer.
[0149] In some embodiments, when the number of memory pages occupied by the primitive data corresponding to the data segments that have been processed by the i-th geometry processing pipeline and have not been submitted to the pixel processing pipeline in the intermediate buffer reaches the maximum usage number allocated to the i-th geometry processing pipeline, the i-th geometry processing pipeline pauses processing new data segments.
[0150] In some embodiments, the rendering system includes multiple GPU cores, and each GPU core includes at least one geometry processing pipeline and at least one pixel processing pipeline.
[0151] The rendering system proposed in this application introduces a rendering pipeline manager to uniformly manage the data segments processed by the geometry processing pipeline. Specifically, the rendering pipeline manager obtains the numbers and memory information of the data segments processed by the geometry processing pipeline. After there are k consecutively numbered data segments, it sends rendering instructions to N pixel processing pipelines respectively. That is, as long as there are k consecutively numbered data segments, regardless of the processing progress of the geometry processing pipeline, the rendering pipeline manager will send rendering instructions to indicate rendering. Therefore, each geometry processing pipeline does not need to wait for other geometry processing pipelines to synchronize. Through the rendering pipeline manager, the asynchronous operation of multiple geometry processing pipelines can be achieved (each geometry processing pipeline continuously processes data segments without maintaining synchronization). Therefore, the latency of image rendering is reduced in the embodiments of this application, thereby improving the efficiency of image rendering.
[0152] In some embodiments, a GPU chip is further provided. The GPU chip includes the above-mentioned rendering system.
[0153] Please refer to Figure 14 , which is a simplified structural block diagram of an electronic device provided in a possible implementation manner of this application. As Figure 14 shown, the electronic device 1400 includes the above-mentioned GPU chip, and the above-mentioned GPU chip includes the above-mentioned rendering system. The electronic device 1400 can be used to implement the rendering method applied to the rendering system provided in the above embodiments.
[0154] Generally, the electronic device 1400 includes: a processor 1401 and a memory 1402.
[0155] The processor 1401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1401 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1401 may be integrated with a GPU, which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1401 may also include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.
[0156] The memory 1402 may include one or more computer-readable storage media, which may be non-transitory. The memory 1402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1402 are used to store a computer program, which is configured to be executed by one or more processors to implement the rendering method applied to the rendering system described above.
[0157] Those skilled in the art can understand that Figure 14 the structure shown in does not constitute a limitation on the electronic device 1400, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0158] In some embodiments, the electronic device 1400 may be a server, a server cluster, an artificial intelligence computing cluster, a cloud computing cluster, etc. Among them, the artificial intelligence computing cluster may also be simply referred to as an intelligent computing cluster or a smart computing cluster, and this application does not make any limitations in this regard.
[0159] It should be understood that the "plurality" mentioned herein refers to two or more. In addition, the step numbers described herein only exemplarily show a possible execution sequence among steps. In some other embodiments, the above steps may not be executed in the numbered order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the order opposite to the illustration. The embodiments of the present application do not make any limitations in this regard.
[0160] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A rendering system, characterized in that, The rendering system includes: a rendering pipeline manager, an intermediate buffer, M geometry processing pipelines, and N pixel processing pipelines, where M is an integer greater than 1 and N is an integer greater than or equal to 1; among them, multiple data segments divided from the input data stream of the rendering system are assigned to the M geometry processing pipelines for processing, and the primitive data obtained after processing each data segment is stored in the intermediate buffer; The rendering pipeline manager is used to obtain the number and memory information of the data segments processed by the geometry processing pipeline, and the memory information of the data segment is used to determine the memory page occupied by the primitive data corresponding to the data segment in the intermediate buffer; The rendering pipeline manager is further used to send rendering instructions to the N pixel processing pipelines respectively, where the rendering instructions include the numbers and memory information of k data segments, k is an integer greater than or equal to 1, and when k is greater than 1, the numbers of the k data segments are consecutive.
2. The rendering system according to claim 1, wherein, The rendering pipeline manager is further used to select one or more data segments with numbers between a first value and a second value from the data segments processed by the geometry processing pipeline as the k data segments; Among them, the first value is the maximum number of the data segments that have been submitted to the N pixel processing pipelines for processing plus 1, and the second value is the maximum number that satisfies the consecutive numbers of the k data segments.
3. The rendering system according to claim 1, characterized in that, There are M queues set in the rendering pipeline manager, and each queue is used to store the numbers and memory information of the data segments processed by a geometry processing pipeline.
4. The rendering system according to claim 3, wherein The i-th queue among the M queues is used to store the numbers and memory information of the data segments processed by the i-th geometry processing pipeline among the M geometry processing pipelines, and i is a positive integer less than or equal to M; The rendering pipeline manager is further used to add the number and memory information of the first data segment to the end of the i-th queue after the i-th geometry processing pipeline finishes processing the first data segment; The rendering pipeline manager is further used to remove the number and memory information of the second data segment from the i-th queue when the second data segment stored in the i-th queue meets the condition of being submitted to the pixel processing pipeline.
5. The rendering system according to claim 3, wherein, The rendering pipeline manager is further used to remove the numbers and memory information of the M data segments included in the merge row from the M queues when there is a merge row in the M queues; among them, the merge row includes the numbers and memory information of one data segment with the smallest number stored in each of the M queues, and the numbers of the M data segments included in the merge row are consecutive; The rendering pipeline manager is further used to record the pending information corresponding to the M queues respectively; among them, for the i-th queue among the M queues, the pending information corresponding to the i-th queue is used to determine the data segments that have been removed from the i-th queue and have not been submitted to the pixel processing pipeline.
6. The rendering system according to claim 1, wherein: the rendering pipeline manager is further configured to, when the memory occupancy rate corresponding to any one of the geometry processing pipelines is greater than or equal to a set threshold, perform the step of sending rendering instructions to the N pixel processing pipelines respectively; wherein, the memory occupancy rate corresponding to the geometry processing pipeline is the ratio of the number of memory pages occupied by the primitive data corresponding to the data segments that have been processed by the geometry processing pipeline and not submitted to the pixel processing pipeline in the intermediate buffer to the total number of memory pages included in the intermediate buffer.
7. The rendering system according to claim 1, wherein: the j-th pixel processing pipeline among the N pixel processing pipelines is configured to, according to the rendering instruction, obtain the primitive data corresponding to the k data segments respectively from the intermediate buffer, render the primitive data corresponding to the k data segments respectively, and after the primitive data corresponding to the k data segments are rendered, send a completion indication message to the rendering pipeline manager, where the completion indication message is used to indicate that the j-th pixel processing pipeline has completed the rendering of the primitive data corresponding to the k data segments respectively, and j is a positive integer less than or equal to N; the rendering pipeline manager is further configured to, after the N pixel processing pipelines have all completed the rendering of the primitive data corresponding to the k data segments respectively, release the memory pages occupied by the primitive data corresponding to the k data segments respectively in the intermediate buffer.
8. The rendering system according to claim 1, wherein The rendering system further includes: M tile partitioners; the i-th geometry processing pipeline among the M geometry processing pipelines is configured to process the data segments allocated to the i-th geometry processing pipeline among the multiple data segments obtained by partitioning the input data stream, to obtain processed data segments, where i is a positive integer less than or equal to M; the i-th tile partitioner among the M tile partitioners is configured to perform tile partitioning on the processed data segments, to obtain the primitive data corresponding to the data segments, and store the primitive data corresponding to the data segments into the intermediate buffer.
9. The rendering system according to claim 8, wherein: the i-th geometry processing pipeline is further configured to pause processing new data segments when the number of memory pages occupied by the primitive data corresponding to the data segments that have been processed by the i-th geometry processing pipeline and not submitted to the pixel processing pipeline in the intermediate buffer reaches the maximum usage quantity allocated to the i-th geometry processing pipeline.
10. The rendering system according to claim 1, characterized in that, The rendering system includes multiple graphics processing unit (GPU) cores, and each GPU core includes at least one geometry processing pipeline and at least one pixel processing pipeline.
11. A GPU chip, characterized in that, The GPU chip includes the rendering system according to any one of claims 1 to 10.
12. An electronic device, characterized in that, The electronic device includes a GPU chip, and the GPU chip includes the rendering system according to any one of claims 1 to 10.
13. A rendering method applied to a rendering system, characterized in that, The rendering system includes: a rendering pipeline manager, an intermediate buffer, M geometric processing pipelines, and N pixel processing pipelines, where M is an integer greater than 1 and N is an integer greater than or equal to 1; among them, multiple data segments divided from the input data stream of the rendering system are allocated to the M geometric processing pipelines for processing, and the primitive data obtained after each data segment is processed is stored in the intermediate buffer; The method includes: The rendering pipeline manager obtains the numbers and memory information of the data segments processed by the geometric processing pipelines, and the memory information of the data segments is used to determine the memory pages occupied by the primitive data corresponding to the data segments in the intermediate buffer; The rendering pipeline manager sends rendering instructions to the N pixel processing pipelines respectively, and the rendering instructions include the numbers and memory information of k data segments, where k is an integer greater than or equal to 1, and when k is greater than 1, the numbers of the k data segments are consecutive.
Citation Information
Patent Citations
Method of processing graphics primitives, graphics processing system and storage medium
CN112862661A
Parallel processing method and device of graph assembly line and readable storage medium
CN114463160A
Scene rendering method and device, equipment and storage medium
CN115908685A
Buffer area-based pixel rendering order-preserving method and system, and storage medium
CN117689790A
System and method for rendering graphical data
US20030164833A1
Cited By
Graphics processing device and method, graphics processor and computer equipment
CN122089556A