Graphics Processing Unit and Method, Multi-Core Graphics Processing System, Electronic Device and Equipment

By using the geometric processing module and tile division module of the graphics processor in the multi-core graphics processing system, and using the secondary tile-list storage structure, the problem of parallel processing of multi-core GPUs in the tile rendering architecture is solved, and the graphics processing performance is improved.

CN117058288BActive Publication Date: 2025-07-18SUZHOU XIANGDIXIAN COMPUTING TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210490158.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-07
Publication Date
2025-07-18
Estimated Expiration
2042-05-07

AI Technical Summary

Technical Problem

In the tile-based rendering architecture, it is difficult for existing multi-core GPUs to make full use of the advantages of multi-core to perform parallel processing of geometric processing and tile division.

Method used

The graphics processor and multi-core graphics processing system are adopted to process the elements through geometric processing modules and tile division modules, and the secondary tile-list storage structure is used to realize the decentralized storage and centralized management of tile information, and the distribution of digits is carried out in combination with the load balancing strategy to distribute the elements to multiple graphics processors for parallel processing.

Benefits of technology

The parallelization of geometric processing and tile division in multi-core graphics processing system is realized, improving the performance and efficiency of graphics processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058288B_ABST
    Figure CN117058288B_ABST
Patent Text Reader

Abstract

The present disclosure provides a graphics processor, a graphics processing method, a multi-core graphics processing system, an electronic device, and an electronic equipment. The graphics processor includes: a geometry processing module configured to perform geometry processing on the primitives assigned to the present graphics processor; a tile division module configured to perform tile division processing on the primitives assigned to the present graphics processor, save the tile information of each obtained tile to the first tile list of each corresponding tile of the present graphics processor respectively, save the indexes of the first tile lists of each tile to the second tile list of each tile, and save the indexes of the first tile lists corresponding to the multiple graphics processors of the multi-core graphics processing system in the second tile list of the tile.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of GPUs (Graphics Processing Units), and particularly to a graphics processing unit, a multi-core graphics processing system, an electronic device, an electronic equipment, and a graphics processing method. Background Art

[0002] GPUs are widely used in solutions such as personal computers, workstations, servers, embedded systems, and electronic game consoles. Since GPUs are designed with a highly parallel architecture, they have more advantages than general-purpose processors CPU (Central Processing Unit) in large parallel processing algorithms and are very efficient in graphics processing.

[0003] A multi-core graphics processing system (also known as a multi-core GPU) refers to a GPU product that realizes graphics processing functions through multiple GPUs (also known as GPU cores). Each GPU in the multi-core graphics processing system has the same and complete GPU functions.

[0004] The typical process of tile-based GPU graphics processing can be divided into several stages: Host-side (host-side) processing, geometry processing, tile division, rasterization, and pixel processing.

[0005] In existing multi-core GPUs adopting a tile-based rendering architecture, rendering operations before pixel processing need to be completed within the main core. After determining the tile division of the primitives, multiple GPU cores then separately process the tiles allocated to themselves for rasterization and pixel processing.

[0006] The current rendering process is difficult to fully utilize the advantages of multiple cores to achieve geometry processing and tile division through multi-core parallelization. Summary of the Invention

[0007] The object of the present disclosure is to provide a graphics processing unit, a multi-core graphics processing system, an electronic device, an electronic equipment, and a graphics processing method to achieve multi-core parallelization for geometry processing and tile division.

[0008] According to one aspect of the present disclosure, there is provided a graphics processing unit applied to a multi-core graphics processing system. The graphics processing unit at least includes:

[0009] A geometry processing module configured to perform geometry processing on the primitives allocated to this graphics processing unit;

[0010] The tile division module is configured to: perform tile division processing on the primitives allocated to this graphics processor, save the tile information of each obtained tile to the first tile list of each tile corresponding to this graphics processor respectively, save the indexes of the first tile lists of each tile to the second tile list of each tile respectively, and the second tile list of a tile stores the indexes of the first tile lists corresponding to multiple graphics processors in the multi-core graphics processing system for this tile.

[0011] If the graphics processor is the main core in the multi-core graphics processing system, the graphics processor may further include a primitive allocation module, which is configured to: allocate the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system.

[0012] Based on any of the above graphics processor embodiments, allocating the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system, the specific implementation manner may include: grouping the primitives in the image frame according to a predetermined rule, and allocating each group of primitives to multiple graphics processors in the multi-core graphics processing system according to a predetermined load balancing strategy.

[0013] Based on any of the above graphics processor embodiments, the geometry processing module may include a primitive allocation sub-module and multiple geometry processing sub-modules. Among them, the primitive allocation sub-module is configured to: allocate the primitives allocated to this graphics processor to multiple geometry processing sub-modules; the geometry processing sub-module is configured to: perform geometry processing on the primitives allocated to this geometry processing sub-module.

[0014] Based on any of the above graphics processor embodiments, the tile division module may further be configured to:

[0015] Save the indexes of the first tile lists of each tile to the second tile list of each tile respectively.

[0016] According to another aspect of the present disclosure, there is also provided a graphics processing system. This multi-core graphics processing system includes at least two graphics processors, and each graphics processor is configured to:

[0017] Perform geometry processing on the primitives allocated to this graphics processor;

[0018] Perform tile division processing on the primitives allocated to this graphics processor, save the tile information of each obtained tile to the first tile list of each tile corresponding to this graphics processor respectively, save the indexes of the first tile lists of each tile to the second tile list of each tile respectively, and the second tile list of a tile stores the indexes of the first tile lists of the above at least two graphics processors for this tile.

[0019] In at least two graphics processors of a multi-core graphics processing system, one graphics processor is the main core and the other graphics processors are slave cores. Among them, the main core can also be configured to: allocate primitives in an image frame to multiple graphics processors in the multi-core graphics processing system.

[0020] Furthermore, allocating primitives in an image frame to multiple graphics processors in the multi-core graphics processing system, the specific implementation method can include: grouping the primitives in the image frame according to a predetermined rule, and allocating each group of primitives to multiple graphics processors in the multi-core graphics processing system according to a predetermined load balancing strategy.

[0021] Based on any of the above multi-core graphics processing system embodiments, performing geometric processing on the primitives allocated to this graphics processor, the specific implementation method can include: performing geometric processing on the primitives allocated to this graphics processor in parallel. The specific implementation method of parallel geometric processing can refer to the above graphics processor embodiments, and the present disclosure does not limit this.

[0022] Based on any of the above multi-core graphics processing system embodiments, each graphics processor can also be configured to: save the indexes of the first tile list of each tile to the second tile list of each tile respectively.

[0023] According to another aspect of the present disclosure, there is also provided an electronic device, which includes the multi-core graphics processing system described in any of the above embodiments. In some usage scenarios, the product form of this electronic device is a graphics card; in some other usage scenarios, the product form of this electronic device is a CPU motherboard.

[0024] According to another aspect of the present disclosure, there is also provided an electronic equipment, which includes the above electronic device. In some usage scenarios, the product form of this electronic equipment is a portable electronic device, such as a smart phone, a tablet computer, a VR device, etc.; in some usage scenarios, the product form of this electronic equipment is a personal computer, a game console, etc.

[0025] According to another aspect of the present disclosure, there is also provided a graphics processing method, which is applied to a graphics processor in a multi-core graphics processing system. This graphics processing method at least includes the following operations:

[0026] Performing geometric processing on the primitives allocated to this graphics processor;

[0027] Perform tiling processing on the primitives assigned to this graphics processor, and save the tile information of each obtained tile to the first tile list of each corresponding tile of this graphics processor. The indexes of the first tile lists of each tile are respectively saved in the second tile lists of each tile. The second tile list of a tile stores the index of this tile in the first tile lists of multiple graphics processors in the multi-core graphics processing system.

[0028] If this graphics processor is the main core in the multi-core graphics processing system, the above graphics processing method may further include: allocating the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system.

[0029] Furthermore, allocating the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system may specifically include: grouping the primitives in the image frame according to a predetermined rule, and allocating each group of primitives to multiple graphics processors in the multi-core graphics processing system according to a predetermined load balancing strategy.

[0030] Based on any of the above embodiments of the graphics processing method, perform geometric processing on the primitives assigned to this graphics processor. The specific implementation may include: performing geometric processing on the primitives assigned to this graphics processor in parallel.

[0031] Based on any of the above embodiments of the graphics processing method, the graphics processing method may further include: saving the indexes of the first tile lists of each tile to the second tile lists of each tile respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 Schematic diagram of the first tile list and primitive data structure according to an embodiment of the present disclosure;

[0033] Figure 2 Schematic diagram of the second tile list structure provided by an embodiment of the present disclosure;

[0034] Figure 3 Schematic diagram of the parallel geometric processing and tiling process provided by an embodiment of the present disclosure;

[0035] Figure 4 Schematic diagram of the graphics processing system structure based on the multi-core GPU architecture according to an embodiment of the present disclosure;

[0036] Figure 5 Schematic diagram of the graphics processing method flow according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] Before introducing the embodiments of the present disclosure, it should be noted that:

[0038] Some embodiments of the present disclosure are described as processing flows. Although the individual operation steps of the flow may be numbered with sequential step numbers, the operation steps can be implemented in parallel, concurrently, or simultaneously.

[0039] In the embodiments of the present disclosure, terms such as "first" and "second" may be used to describe various features, but these features should not be limited by these terms. These terms are only used to distinguish one feature from another.

[0040] In the embodiments of the present disclosure, the term "and / or" may be used, and "and / or" includes any and all combinations of one or more of the listed associated features.

[0041] It should be understood that when describing the connection relationship or communication relationship between two components, unless it is clearly specified that the two components are directly connected or directly communicate, the connection or communication between the two components can be understood as either direct connection or communication, or indirect connection or communication through an intermediate component.

[0042] In order to make the technical solutions and advantages in the embodiments of the present disclosure clearer and more understandable, the following further details the exemplary embodiments of the present disclosure with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0043] Since tile partitioning requires centralized processing and the corresponding tile information needs to be centrally stored, the GPU operations before rasterization need to be completed within the main core. After determining the tile partitioning of the primitive, multiple GPU cores then process the tiles assigned to themselves respectively. Since the problem of centralized storage of tile information has not been overcome, the prior art does not perform multi-core parallel processing on the geometric processing and tile partitioning processes.

[0044] The purpose of the present disclosure is to provide an implementation solution for multi-core parallel implementation of geometric processing and tile partitioning in a multi-core graphics processing system adopting a tile-based rendering architecture. Among them, the multi-core in the graphics processing system refers to multiple GPUs. A GPU is a processor with computing functions implemented by hardware, which includes components such as computing units and caches, and can be a GPGPU (general-purpose graphics processing unit) or a GPU.

[0045] One embodiment of the present disclosure provides a graphics processor, which is applied to a multi-core graphics processing system and is suitable for a tile-based rendering architecture, such as TBR (Tile Based Render), TBDR (Tile Based Deferred Rendering), etc.

[0046] The above-mentioned graphics processor at least includes a geometry processing module and a tile partitioning module.

[0047] Among them, the geometry processing module is configured to: perform geometry processing on the primitives assigned to this graphics processor.

[0048] The geometry processing module is usually implemented by custom hardware and at least one programmable shader, but the present disclosure does not limit this. In practical applications, the geometry processing function can also be implemented by custom hardware.

[0049] Among them, the tile partitioning module is configured to: perform tile partitioning processing on the primitives assigned to this graphics processor, save the tile information of each obtained tile to the first tile list of each corresponding tile of this graphics processor respectively, save the indexes of the first tile lists of each tile in the second tile list of each tile respectively, and the second tile list of the tile stores the indexes of the first tile lists corresponding to multiple graphics processors in the multi-core graphics processing system.

[0050] The tile partitioning processing includes: determining the primitives covered by each tile. The graphics processor saves the primitive indexes of the primitives covered by each tile determined after the tile partitioning processing to the first tile list corresponding to this graphics processor, so that the module for performing subsequent rendering processes (such as the rasterization module) can read the primitive information of the primitives covered by the tile according to the primitive indexes of the primitives covered by the tile.

[0051] Among them, the primitive index of the primitives covered by each tile determined after the tile partitioning processing is the tile information of each tile obtained after the tile partitioning processing.

[0052] Among them, the first tile list can be but is not limited to being saved in the system memory.

[0053] Among them, the primitive information is generated in the geometry stage, and the primitive information includes the vertex attribute information of the primitive (such as vertex color, vertex coordinates, etc.) and the basic information of the primitive (such as primitive identifier, etc.).

[0054] By way of example and not limitation, such as Figure 1As shown, in some embodiments, the data structure for storing primitive information is a primitive block, and each primitive block includes the primitive information of multiple primitives. Each tile corresponds to a first tile list, and each first tile list includes the primitive indices of the primitives covered by the corresponding tile. The primitive index indicates the storage address of the primitive information of the primitive. For example, the first tile Tile0 covers four primitives, which are: primitive 1-0, primitive 1-1, primitive 2-0, and primitive 3-1. Among them, the primitive information of primitive 1-0 is primitive information 0 in the first primitive block, the primitive information of primitive 1-1 is primitive information 1 in the first primitive block, the primitive information of primitive 2-0 is primitive information 0 in the second primitive block, and the primitive information of primitive 3-1 is primitive information 1 in the third primitive block. Then, the tile list corresponding to the first tile Tile0 includes the primitive index of primitive 1-0 (in this embodiment, the primitive index of primitive 1-0 is the storage address of primitive information 0 in the first primitive block), the primitive index of primitive 1-1 (in this embodiment, the primitive index of primitive 1-1 is the storage address of primitive information 1 in the first primitive block), the primitive index of primitive 2-0 (in this embodiment, the primitive index of primitive 2-0 is the storage address of primitive information 0 in the second primitive block), and the primitive index of primitive 3-1 (in this embodiment, the primitive index of primitive 3-1 is the storage address of primitive information 1 in the third primitive block). Another example, the first tile Tile1 covers three primitives, which are: primitive 1-1, primitive 2-0, and primitive 3-0. Among them, the relevant introductions of primitive 1-1 and primitive 2-0 can refer to the above description, which will not be elaborated here. The primitive information of primitive 3-0 is primitive information 0 in the third primitive block. Then, the first tile list corresponding to the first tile Tile1 includes the primitive index of primitive 1-1, the primitive index of primitive 2-0, and the primitive index of primitive 3-0 (in this embodiment, the primitive index of primitive 3-0 is the storage address of primitive information 0 in the third primitive block).

[0055] In practical applications, the primitive identifier can also be used as the primitive index.

[0056] As can be seen from the above description, the tile information of each tile processed by each graphics processor is separately stored in its corresponding first tile list, which belongs to a dispersed storage method. However, in the rendering stage (such as rasterization) in units of tiles, it is necessary to read the primitive information of all primitives covered by one tile at a time. In order to be able to read the primitive information of all primitives covered by one tile at a time on the premise that the tile information is dispersed, the present disclosure adopts a tile information storage structure of a two-level tile-list. That is, the primitive indices are dispersed and stored in the first tile lists corresponding to the respective graphics processors, and the indices of the primitive indices in each first tile list are centrally stored in the second tile list. Correspondingly, in the rendering stage (such as rasterization) in units of tiles, the indices of the primitive indices can be read from the second tile list, indexed to the first tile list, and then indexed to the corresponding primitive information.

[0057] For simplicity of description, taking the example of dividing into 4 tiles, it is assumed that there are N graphics processors in the graphics processing system participating in geometric processing and tile division, as Figure 2As shown, the second tile list of a tile stores the indexes of the first tile list corresponding to each graphics processor for this tile. For example, in the second tile list of the first tile Tile0, there is a pointer to the first tile list Core0 tile-list0 corresponding to the graphics processor Core0, a pointer to the first tile list Core1 tile-list0 corresponding to the graphics processor Core1, and so on, until the pointer to the first tile list CoreN tile-list0 corresponding to the graphics processor CoreN; in the second tile list of the second tile Tile1, there is a pointer to the first tile list Core0tile-list1 corresponding to the graphics processor Core0, a pointer to the first tile list Core1 tile-list1 corresponding to the graphics processor Core1, and so on, until the pointer to the first tile list CoreN tile-list1 corresponding to the graphics processor CoreN; in the second tile list of the third tile Tile2, there is a pointer to the first tile list Core0 tile-list2 corresponding to the graphics processor Core0, a pointer to the first tile list Core1 tile-list2 corresponding to the graphics processor Core1, and so on, until the pointer to the first tile list CoreN tile-list2 corresponding to the graphics processor CoreN; in the second tile list of the fourth tile Tile3, there is a pointer to the first tile list Core0 tile-list3 corresponding to the graphics processor Core0, a pointer to the first tile list Core1tile-list3 corresponding to the graphics processor Core1, and so on, until the pointer to the first tile list CoreN tile-list3 corresponding to the graphics processor CoreN.

[0058] The above only takes the pointer as an example of the index for illustration. In practical applications, the present disclosure does not limit the specific implementation form of the index.

[0059] The above takes the storage of the indexes of the first tile lists of tiles in the order of the graphics processor numbers in the second tile list as an example for illustration. In practical applications, the indexes of the first tile lists of tiles can also be stored in other orders, and the present disclosure does not limit this.

[0060] With the above-mentioned two-level tile-list storage structure, in the subsequent rendering stage (such as rasterization) in units of tiles, each tile can still be addressed to a unique second tile list, but each second tile list is composed of indexes of the first tile lists at the next level.

[0061] In the embodiments of the present disclosure, it is possible but not limited to that after the main core determines that the tile division modules of each graphics processing unit have completed tile division, the storage address of the second tile list is sent to the subsequent rendering module (such as the rasterization module) in units of tiles, so that the subsequent rendering module reads the indexes of each first tile list from each second tile list block by block according to the storage address, and then reads the primitive indexes of the primitives according to the indexes of each first tile list, and further reads the primitive information according to the indexes of the primitive information of the primitives.

[0062] Based on any of the above embodiments of the graphics processor, the graphics processor may further include a primitive allocation module, which is configured to: allocate the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system. When the graphics processor acts as the main core, its primitive allocation module works; when the graphics processor acts as the slave core, its primitive allocation module does not work.

[0063] It should be noted that the primitive allocation may also be implemented by other modules in the graphics processing system. By way of example and not limitation, the application processor in the graphics processing system may also implement the primitive allocation of each graphics processor.

[0064] The present disclosure does not limit the primitive allocation strategy. In practical applications, the primitive allocation strategy can be formulated according to needs or actual situations. By way of example and not limitation, the primitives can be allocated in a polling manner, or the primitives can be allocated according to a preset load balancing strategy.

[0065] In some embodiments, the primitives in the image frame can be grouped according to a predetermined rule, and each group of primitives can be allocated to multiple graphics processors in the multi-core graphics processing system according to a predetermined load balancing strategy.

[0066] In the embodiments of the present disclosure, the so-called allocation of primitives may include but is not limited to sending the vertex data corresponding to the primitives allocated to a certain graphics processor to the graphics processor.

[0067] In some embodiments, geometric processing can be performed in parallel by multiple geometric pipelines. That is to say, a graphics processor has multiple geometric pipelines. Correspondingly, based on any of the above graphics processor embodiments, the geometric processing module may include a primitive allocation sub-module and multiple geometric processing sub-modules. Among them, the primitive allocation sub-module is configured to allocate the primitives assigned to this graphics processor to multiple geometric processing sub-modules; the geometric processing sub-module is configured to perform geometric processing on the primitives assigned to this geometric processing sub-module.

[0068] Correspondingly, as Figure 3 shown, the main core performs primary allocation of primitives, and each graphics processor (including the main core and slave cores) re-performs primitive allocation on the allocated primitives, allocating the primitives assigned to this graphics processor to multiple geometric pipelines (corresponding to geometric processing sub-modules) in this graphics processor. After multiple geometric pipelines perform geometric processing in parallel, each graphics processor respectively performs tile division on the primitives assigned to this graphics processor.

[0069] Through Figure 3 the parallel geometric processing and tile division shown, the parallel advantages of multi-cores can be fully utilized to improve graphics processing performance.

[0070] Among them, the implementation description of the primitive allocation sub-module can refer to the description of the above primitive allocation module, which will not be elaborated here. In practical applications, the functions of the primitive allocation sub-module can be implemented using existing implementation methods or principles.

[0071] The embodiments of the present disclosure do not limit the specific implementation manner of geometric processing, and existing geometric processing implementation methods can be used.

[0072] Based on any of the above graphics processor embodiments, the second tile list can be pre-configured, and the indexes of each first tile list are fixed. The indexes of each first tile list in the second tile list can also be dynamically changed. Then, the tile division module can also be configured to save the indexes of the tile information of each tile in the first tile list into the tile information of each tile in the second tile list.

[0073] The embodiments of the present disclosure also provide a multi-core graphics processing system based on a tile-based rendering architecture. The graphics processing system includes at least two graphics processors, and the functions implemented by each graphics processor can refer to the description of any of the above graphics processor embodiments.

[0074] In the multi-core graphics processing system provided by the present disclosure in real time, the parallelism advantages of multi-cores are fully utilized in the geometric processing stage and the tile division stage, and performance expansion and improvement are better achieved.

[0075] In the embodiments of the present disclosure, the product form of the graphics processing system may be a SOC (System on Chip) chip.

[0076] The graphics processing system in the embodiments of the present disclosure may be a single-die (wafer) SOC chip or a multi-die interconnected SOC chip.

[0077] Taking one die as an example, the architecture and working principle of the graphics processing system provided by the present disclosure will be described below.

[0078] In Figure 4 In an embodiment shown, the single-die graphics processing system includes multiple GPU cores (GPU Core), and the GPU core is the above-mentioned graphics processor, where one GPU core serves as the main core and other GPU cores serve as slave cores.

[0079] The GPU core is used to process drawing instructions, execute the Pipeline of image rendering according to the drawing instructions, and can also be used to execute other arithmetic instructions. The GPU core further includes: a computing unit, which is used to execute the instructions after shader compilation, belongs to a programmable module, and is composed of a large number of ALUs; a cache (Cache), which is used for caching the data of the GPU core to reduce the access to memory; a rasterization module, which is a fixed stage of the 3D rendering pipeline, and further includes a primitive information calculation module and a pixel information processing module; a tiling module, which performs tiling processing on a frame in the TBR and TBDR GPU architectures; a clipping module, which is a fixed stage of the 3D rendering pipeline and clips the primitives outside the viewing range or the back side that is not displayed; a post-processing module, which is used to perform operations such as scaling, clipping, and rotation on the drawn image; a micro core, which is used for scheduling between various pipeline hardware modules on the GPU core or for task scheduling of multiple GPU cores.

[0080] The GPU core is connected to the on-chip network. Among them, the on-chip network is used for data exchange between various masters and slaves on the graphics processing system. In this embodiment, the on-chip network includes a configuration bus, a data communication network, a communication bus, and so on.

[0081] As Figure 4 shown, the graphics processing system may further include:

[0082] A general-purpose DMA (Direct Memory Access), which is used to perform data transfer between the host side and the memory of the graphics processing system (such as the video card memory). For example, the vertex data of 3D drawing is transferred from the host side to the memory of the graphics processing system through DMA;

[0083] A PCIe controller, an interface for communicating with a host, implements the PCIe protocol, enabling a graphics processing system to be connected to the host via a PCIe interface. Programs such as a graphics API and a graphics card driver are running on the host.

[0084] An application processor is used for scheduling tasks of various modules on the graphics processing system. For example, after the GPU finishes rendering a frame, it notifies the application processor, and then the application processor starts the display controller to display the image drawn by the GPU on the screen.

[0085] A memory controller is used to connect to a memory device and store data on the SOC.

[0086] A display controller controls the output of the frame buffer in memory to a display via a display interface (HDMI, DP, etc.).

[0087] Video decoding can decode the encoded video on the host hard disk into a displayable image.

[0088] Video encoding can encode the original video stream on the host hard disk into a specified format and return it to the host.

[0089] Based on Figure 4 the multi-core graphics processing system architecture shown, in one embodiment, the graphics rendering process is as follows:

[0090] The graphics API of the host (in actual applications, for a mobile graphics processing system, it can also be software on the application processor) sends a drawing instruction to the SOC chip, requesting to render an image frame.

[0091] Among them, the image frame includes at least one object.

[0092] The general DMA transfers the vertex coordinate information of each object in the image frame from the host side to the graphics processing system memory.

[0093] After the computing unit of the main core obtains the above drawing instruction, it decodes the drawing instruction.

[0094] The primitive allocation module of the main core (its function is implemented by the computing unit) distributes the primitives in the image frame to multiple GPU cores including the main core according to a predetermined allocation strategy.

[0095] The primitive allocation sub-module of each GPU core (whose functions are implemented by the computing unit) allocates the primitives assigned to this GPU core to multiple geometric pipelines of this GPU core; each geometric pipeline performs geometric processing in parallel: the vertex shader (whose functions are implemented by the computing unit) obtains the vertex coordinate information corresponding to the primitives assigned to this GPU core from the system memory, and transmits the vertex coordinate information to the geometry shader (whose functions are implemented by the computing unit), and the geometry shader converts the 3D coordinates of the vertex into expanded texture coordinates (i.e., (u, v) coordinates). In addition, the computing unit also performs primitive assembly according to the vertex coordinate information. Among them, the value at the texture coordinate corresponding to the vertex coordinate in the texture map is the vertex color information.

[0096] The vertex coordinate information and vertex texture coordinates of the primitive are saved into the data structure of the primitive in the system memory.

[0097] After the geometric processing, the tile division module of each GPU core performs tile division processing on the primitives in the image frame according to the depth buffer size, and saves the tile division processing result in the first tile list. The first tile list includes the primitive indices of the primitives covered by the tile, as Figure 1 shown, the primitive index indicates the storage address of the primitive information of the primitive.

[0098] When the tile division of all the objects to be displayed in a frame is completed, the rasterization process will be started.

[0099] According to the requirements of deferred rendering, the rasterization module processes each tile one by one. Each time, it reads the index of the first tile list corresponding to the current tile of each GPU core from the second tile list of the current tile, and reads the primitive index (primitive identifier or primitive information storage address) saved in the first tile list according to the index of the first tile list; the rasterization module reads the primitive information of the primitive through the primitive index, and performs pixel coverage testing using the primitive information of the primitive to determine the pixels covered by the primitive, and then performs pixel interpolation calculation and at least one pixel test (by way of example and not limitation, the pixel test may include depth test, stencil test, etc.).

[0100] In the present disclosure, pixel coverage testing, pixel interpolation calculation, and pixel testing are implemented using the prior art.

[0101] After rasterization processing, the information of the output pixels is output.

[0102] When the rasterization of all the first tiles in a frame is completed, then the pixel processing process will be called.

[0103] Specifically, the fragment shader of the GPU core (whose functions are implemented by the computing unit) performs the coloring calculation (such as lighting calculation) of the corresponding pixels for each pixel covered by the primitive in units of tiles.

[0104] An embodiment of the present disclosure further provides an electronic device, which includes the graphics processing system described in any of the above embodiments. In some usage scenarios, the product form of the electronic device is a graphics card; in other usage scenarios, the product form of the electronic device is a CPU motherboard.

[0105] An embodiment of the present disclosure further provides an electronic equipment, which includes the above-mentioned electronic device. In some usage scenarios, the product form of the electronic equipment is a portable electronic device, such as a smart phone, a tablet computer, a VR device, etc.; in some usage scenarios, the product form of the electronic equipment is a personal computer, a game console, a workstation, a server, etc.

[0106] Based on the same inventive concept, an embodiment of the present disclosure further provides a graphics processing method, which is applied to a graphics processor in a multi-core graphics processing system, such as Figure 5 As shown, the method at least includes the following operations:

[0107] Step 501: Geometrically process the primitives allocated to this graphics processor;

[0108] Step 502: Perform tile division processing on the primitives allocated to this graphics processor, save the tile information of each obtained tile to the first tile list of each corresponding tile of this graphics processor respectively, save the indexes of the first tile lists of each tile in the second tile list of each tile respectively, and the second tile list of the tile stores the indexes of the first tile lists of multiple graphics processors in the multi-core graphics processing system where this tile is located.

[0109] If this graphics processor is the main core in the multi-core graphics processing system, the above graphics processing method may further include: allocating the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system.

[0110] Furthermore, the implementation manner of allocating the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system may include: grouping the primitives in the image frame according to a predetermined rule, and allocating each group of primitives to multiple graphics processors in the multi-core graphics processing system according to a predetermined load balancing strategy.

[0111] Based on any of the above embodiments of the graphics processing method, the implementation manner of geometrically processing the primitives allocated to this graphics processor may include: geometrically processing the primitives allocated to this graphics processor in parallel.

[0112] Based on any of the above embodiments of the graphics processing method, the graphics processing method may further include: saving the indexes of the first tile lists of each tile to the second tile list of each tile respectively.

[0113] It should be noted that the above graphic processing method and the above graphics processor are based on the same inventive concept. Therefore, the specific implementation manners of the steps in the method and the noun explanations involved can be referred to the descriptions of the above embodiments, and will not be repeated here.

[0114] Although the preferred embodiments of the present disclosure have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present disclosure.

[0115] Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these modifications and variations.

Claims

1. A graphics processor, applied to a multi-core graphics processing system based on a tile-based rendering architecture, the multi-core graphics processing system including a main graphics processor and a slave graphics processor; The graphics processor is a main graphics processor or a slave graphics processor, and the graphics processor at least includes: A geometry processing module, configured to: perform geometry processing on the primitives assigned to this graphics processor; A tile division module, configured to: perform tile division processing on the primitives assigned to this graphics processor to determine the primitives covered by each tile, save the tile information of each obtained tile to the first tile list of each corresponding tile of this graphics processor respectively, save the indexes of the first tile lists of each tile to the second tile list of each tile respectively, and the second tile list of the tile stores the indexes of the first tile lists corresponding to the multiple graphics processors in the multi-core graphics processing system, so that all the primitive information of the primitives covered by one tile can be read at one time in the subsequent rendering stage on the premise that the tile information is stored dispersedly.

2. The graphics processor according to claim 1, if the graphics processor is the main core in the multi-core graphics processing system, the graphics processor further includes a primitive allocation module, and the primitive allocation module is configured to: allocate the primitives in the image frame to the multiple graphics processors in the multi-core graphics processing system.

3. The graphics processor according to claim 2, wherein the assigning of primitives in an image frame to a plurality of graphics processors in the multi-core graphics processing system comprises: Group the primitives in the image frame according to a predetermined rule, and allocate each group of primitives to the multiple graphics processors in the multi-core graphics processing system according to a predetermined load balancing strategy.

4. The graphics processor according to any one of claims 1 to 3, wherein the geometry processing module includes a primitive allocation sub-module and multiple geometry processing sub-modules; The primitive allocation sub-module is configured to: allocate the primitives assigned to this graphics processor to the multiple geometry processing sub-modules; The geometry processing sub-module is configured to: perform geometry processing on the primitives assigned to this geometry processing sub-module.

5. The graphics processor according to any one of claims 1 to 3, the tile division module is further configured to: Save the indexes of the first tile lists of each tile to the second tile list of each tile respectively.

6. A multi-core graphics processing system, the multi-core graphics processing system is a multi-core graphics processing system based on a tile-based rendering architecture, the graphics processing system includes at least two graphics processors, and the at least two graphics processors include a main graphics processor and a slave graphics processor, and each graphics processor is configured to: Perform geometry processing on the primitives assigned to this graphics processor; Perform tile division processing on the primitives assigned to this graphics processor to determine the primitives covered by each tile, save the tile information of each obtained tile to the first tile list of each corresponding tile of this graphics processor respectively, save the indexes of the first tile lists of each tile to the second tile list of each tile respectively, and the second tile list of the tile stores the indexes of the first tile lists corresponding to the at least two graphics processors, so that all the primitive information of the primitives covered by one tile can be read at one time in the subsequent rendering stage on the premise that the tile information is stored dispersedly.

7. The multi-core graphics processing system according to claim 6, wherein among the at least two graphics processors, one graphics processor is a main core, and the other graphics processors are slave cores; The main core is further configured to: allocate the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system.

8. The multi-core graphics processing system according to claim 7, wherein the step of allocating the primitives in the image frame to a plurality of graphics processors in the multi-core graphics processing system includes: Group the primitives in the image frame according to a predetermined rule, and allocate each group of primitives to multiple graphics processors in the multi-core graphics processing system according to a predetermined load balancing strategy.

9. The multi-core graphics processing system according to any one of claims 6 to 8, wherein the geometric processing of the primitive assigned to the present graphics processor includes: Geometrically process the primitives allocated to this graphics processor in parallel.

10. The multi-core graphics processing system according to any one of claims 6 to 8, wherein each graphics processor is further configured to: respectively save the indexes of the first tile list of each tile to the second tile list of each tile.

11. An electronic device, comprising the multi-core graphics processing system according to any one of claims 6 to 10.

12. An electronic equipment, comprising the electronic device according to claim 11.

13. A graphics processing method is applied to a graphics processor in a multi-core graphics processing system based on a tile-based rendering architecture. The multi-core graphics processing system includes a main graphics processor and a slave graphics processor; The graphics processor is a main graphics processor or a slave graphics processor, and the graphics processing method at least includes: Geometrically process the primitives allocated to this graphics processor. Perform tile division processing on the primitives allocated to this graphics processor to determine the primitives covered by each tile, respectively save the tile information of each obtained tile to the first tile list of each corresponding tile of this graphics processor, and save the indexes of the first tile list of each tile to the second tile list of each tile respectively. The second tile list of the tile stores the indexes of the first tile list of the tile in multiple graphics processors of the multi-core graphics processing system, so that the tile information can be scattered and saved, and all the primitive information covered by one tile can be read at one time in the subsequent rendering stage.

14. The graphics processing method according to claim 13, if this graphics processor is the main core in the multi-core graphics processing system, the method further includes: Allocate the primitives in the image frame to multiple graphics processors in the multi-core graphics processing system.

15. The graphic processing method according to claim 14, wherein the step of allocating the primitive in the image frame to multiple graphics processors in the multi-core graphic processing system includes: Group the primitives in the image frame according to a predetermined rule, and allocate each group of primitives to multiple graphics processors in the multi-core graphics processing system according to a predetermined load balancing strategy.

16. The graphics processing method according to any one of claims 13 to 15, wherein the geometric processing of the primitive assigned to the present graphics processor includes: Geometrically process the primitives allocated to this graphics processor in parallel.

17. The graphics processing method according to any one of claims 13 to 15, the method further includes: Respectively save the indexes of the first tile list of each tile to the second tile list of each tile.

Citation Information

Patent Citations

  • Multi-GPU large-resolution multi-screen graphics block parallel rendering method

    CN107958437A

  • Multi-core geometry processing in a tile based rendering system

    US20090174706A1

  • Graphic processing unit, graphic processing system including the same and rendering method using the same

    US20140333620A1