Graph rendering processing method and system on chip
By using an on-chip network in the graphics rendering system to write graphics rendering requests to shared memory, and allowing the graphics processor and display controller to directly access the shared memory, the problem of inefficient graphics rendering request processing under a heterogeneous architecture is solved, and more efficient graphics rendering processing is achieved.
Patent Information
- Application Number
- CN202510040650.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-16
AI Technical Summary
During the graphics rendering process, the CPU and GPU shared memory and cache under the heterogeneous architecture, resulting in on-chip network consistency checks on requests, increasing access delay and power consumption, and reducing the processing efficiency of graphics rendering requests.
Writing graphics rendering requests to shared memory via an on-chip network, the graphics processor and display controller read and write rendering requests and results directly from shared memory, avoiding consistency checks.
Improves the processing efficiency of graphics rendering requests, reduces delay and power consumption, and enhances processing capabilities.
Smart Images

Figure CN120014132A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a graphics rendering processing method and a system on chip. Background Art
[0002] Currently, a heterogeneous structure of CPU (central processing unit) and GPU (graphics processing unit) is usually used for graphics rendering. In the coupled heterogeneous architecture, the CPU and GPU share memory and cache, and the requests of the CPU and GPU need to access the shared memory through the on-chip network. At the same time, the on-chip network will perform consistency checks on the requests of the CPU and GPU, resulting in a certain access delay and additional power consumption, which leads to low processing efficiency of graphics rendering requests during the graphics rendering process.
[0003] In view of this, how to improve the processing efficiency of graphics rendering requests is an urgent problem to be solved. Summary of the invention
[0004] The present application provides a graphics rendering processing method and a system on chip, which can improve the processing efficiency of graphics rendering requests.
[0005] A first aspect of the present application provides a graphics rendering processing method, which is applied to a system on chip, wherein the system on chip includes a central processing unit, an on-chip network, a graphics processor, a display controller, and a shared memory shared by the central processing unit and the graphics processor. The method includes: the central processing unit writes a graphics rendering request into the shared memory through the on-chip network, and sends description information of the graphics rendering request to the graphics processor through the on-chip network; the graphics processor directly reads the graphics rendering request from the shared memory based on the received description information, and writes the rendering result obtained by processing the graphics rendering request into the shared memory; the display controller directly reads the rendering result from the shared memory, and displays the rendering result through an external display device.
[0006] The technical solution provided by this embodiment of the present application can improve the processing efficiency of graphics rendering requests. Specifically, the central processing unit is connected to the on-chip network, and the shared memory is accessed through the on-chip network. The graphics processor and the display controller can directly access the shared memory, thereby avoiding the consistency check process of the graphics processor and the display controller, and reducing the delay and power consumption caused by the consistency check. It can be seen that through the technical solution provided by this embodiment of the present application, the graphics processor and the display controller can avoid the consistency check of the on-chip network, and by directly accessing the shared memory, the processing efficiency of graphics rendering requests can be improved.
[0007] In a possible implementation, the graphics rendering request includes rendering instructions and rendering data; the central processing unit writes the graphics rendering request into the shared memory through the on-chip network, including: the central processing unit writes the rendering instructions in the graphics rendering request into the instruction queue of the shared memory through the on-chip network, and writes the rendering data in the graphics rendering request into the data queue of the shared memory.
[0008] The technical solution provided in this embodiment of the present application proposes an instruction queue and a data queue in a shared memory to write rendering instructions and rendering data into different queues, thereby performing orderly management of memory data.
[0009] In a possible implementation, the description information is at least used to characterize the storage address of the rendering instruction in the shared memory; the graphics processor directly reads the graphics rendering request from the shared memory based on the received description information, including: the graphics processor directly reads the rendering instruction from the shared memory according to the storage address represented by the description information, and parses the rendering instruction to obtain the storage address of the rendering data in the shared memory; the graphics processor reads the rendering data from the shared memory according to the storage address of the rendering data in the shared memory.
[0010] The technical solution provided in this embodiment of the present application further explains the process of the graphics processor reading the graphics rendering request. The rendering instructions and rendering data with the rendering data address information are read separately, and at the same time, the graphics processor avoids data interaction with the on-chip network, but directly reads the rendering instructions and rendering data in the shared memory, thereby improving the processing efficiency of the graphics rendering request.
[0011] In a possible implementation, the central processing unit sends rendering information of the graphics rendering request to the display controller via the on-chip network, so that the display controller selects an external display device according to the rendering information, and displays the rendering result in the selected external display device according to the rendering information.
[0012] The technical solution provided in this embodiment of the present application further describes the process of the display controller reading rendering information. The display controller obtains the rendering information sent by the central processor through the on-chip network instead of accessing the shared memory through the on-chip network, which improves the data processing efficiency while ensuring the correctness of the rendering.
[0013] In a possible implementation, the graphics processor writes the rendering result obtained by processing the graphics rendering request into the shared memory, including: the graphics processor writes the generated intermediate rendering result into the first result queue of the shared memory during the process of processing the graphics rendering request; after the graphics processor completes the processing of the graphics rendering request, the graphics processor writes the generated final rendering result into the second result queue of the shared memory; accordingly, the display controller directly reads the final rendering result from the second result queue of the shared memory, and displays the final rendering result through an external display device.
[0014] The technical solution provided in this embodiment of the present application further describes the process of the graphics processor performing rendering request processing. The graphics processor writes the rendering result into a pre-planned result queue in the shared memory so that the display controller can read the rendering result.
[0015] In a possible implementation, the system on chip also includes a routing network, and the graphics processor and the display controller access the shared memory through the routing network; the method also includes: the routing network receives a first access request for the shared memory initiated by the graphics processor, and receives a second access request for the shared memory initiated by the display controller, merges the first access request and the second access request into a request data stream, and performs data interaction with the shared memory based on the merged request data stream.
[0016] The technical solution provided in this embodiment of the present application proposes a routing network for a graphics processor and a display controller to access a shared memory. The routing network can pre-merge access requests of the graphics processor and the display controller, reduce the data interface of the shared memory, avoid bandwidth limitations and access conflicts of the shared memory, and thus improve data processing efficiency.
[0017] In a possible implementation, after the routing network receives the first access request and the second access request, the method further includes: the routing network setting respective request identifiers for the first access request and the second access request; after the routing network obtains the request data from the shared memory, based on the request identifier carried in the request data, feeding back the obtained request data to the graphics processor or to the display controller.
[0018] The technical solution provided by this embodiment of the present application is that the routing network adds different identifiers to the requests of the graphics processor and the display controller to identify the corresponding request source based on the identifier carried in the returned data, thereby ensuring the correct transmission of the requested data.
[0019] In a possible implementation, the routing network includes a first data queue of the graphics processor and a second data queue of the display controller; feeding back the acquired request data to the graphics processor or to the display controller includes: if the request data carries a request identifier corresponding to the first access request, writing the request data into the first data queue, and if the request data carries a request identifier corresponding to the second access request, writing the request data into the second data queue.
[0020] The technical solution provided in this embodiment of the present application further describes the queue structure in the routing network, and writes the request data of the graphics processor and the display controller into different data queues so that the graphics processor and the display controller can read the data correctly.
[0021] In a possible implementation, the routing network includes a first request queue of the graphics processor and a second request queue of the display controller, the first request queue is used to store the first access request, and the second request queue is used to store the second access request; wherein the transmission efficiency of the merged request data stream is greater than or equal to the sum of the transmission efficiencies of the first access request and the second access request, and the transmission efficiency of the request data stream is determined by the bit width of the first request queue and the second request queue, and the clock frequency of the first request queue and the second request queue.
[0022] The technical solution provided in this embodiment of the present application proposes that the transmission efficiency of the request data stream is determined by the clock frequency and the bit width, thereby ensuring that the combined request data stream does not reduce the data transmission efficiency.
[0023] In a possible implementation, if the clock frequency of the first request queue or the clock frequency of the second request queue is inconsistent with the clock frequency of the preset output queue, the routing network also includes a first cross-clock unit and a second cross-clock unit; the method also includes: the first cross-clock unit receives access requests in the first request queue, the second cross-clock unit receives access requests in the second request queue, and the first cross-clock unit and the second cross-clock unit output the access requests they respectively receive according to the clock frequency of the preset output queue.
[0024] The technical solution provided in this embodiment of the present application solves the problem of clock frequency asynchrony by constructing a first cross-clock unit and a second cross-clock unit in a routing network.
[0025] A second aspect of the present application provides a system on chip, which includes a central processing unit, an on-chip network, a graphics processor, a display controller, and a shared memory shared by the central processing unit and the graphics processor, wherein: the central processing unit is used to write a graphics rendering request into the shared memory through the on-chip network, and send description information of the graphics rendering request to the graphics processor through the on-chip network; the graphics processor is used to read the graphics rendering request directly from the shared memory based on the received description information, and write the rendering result obtained by processing the graphics rendering request into the shared memory; the display controller is used to read the rendering result directly from the shared memory, and display the rendering result through an external display device. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 A schematic diagram of the steps of a graphics rendering processing method provided in an embodiment of the present application;
[0028] Figure 2 A schematic diagram of a structure for accessing a graphics rendering request in a shared memory provided by one embodiment of the present application;
[0029] Figure 3 A schematic diagram of a structure for accessing rendering results in a shared memory provided in one embodiment of the present application;
[0030] Figure 4 A schematic diagram of a structure of a routing network for data transmission provided in one embodiment of the present application;
[0031] Figure 5 A schematic diagram of the structure of a routing network provided for one embodiment of the present application;
[0032] Figure 6 A schematic diagram of a graphics rendering processing method provided by an embodiment of the present application;
[0033] Figure 7 A schematic diagram of a traditional graphics rendering processing method provided by an embodiment of the present application;
[0034] Figure 8 A schematic diagram of the structure of a system on a chip provided for one embodiment of the present application. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0036] In addition, the descriptions involving "first", "second", etc. in this application are for descriptive purposes only and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more. In addition, the use of "based on" or "according to" means openness and inclusiveness, because the process, steps, calculations or other actions "based on" or "according to" one or more of the conditions or values can be based on additional conditions or beyond the described values in practice.
[0037] The CPU is the computing and control core of the computer system and is mainly used for multi-task management and scheduling. The GPU uses a large number of computing units and an ultra-long pipeline and is mainly used for image processing and parallel computing. The display controller is responsible for managing the data transmission and display process between the computer and the display. Unless otherwise specified, in the subsequent embodiments of this application, the central processing unit is interpreted as the CPU and the graphics processing unit is interpreted as the GPU.
[0038] In a computer graphics system, the CPU issues a graphics rendering request and stores it in a memory unit. The GPU calls the graphics rendering request in the memory unit, performs a series of operations to generate a rendering result and transfers it to the memory unit. The display calls the rendering result in the memory unit to display the image.
[0039] Generally speaking, the combination of CPU and GPU can handle high-performance computing tasks and large amounts of data, especially for graphics rendering tasks. At the same time, CPU and GPU can share memory space to improve data processing efficiency. In the traditional heterogeneous architecture of CPU and GPU, CPU and GPU need to go through an on-chip network when making requests for interaction. The on-chip network will perform consistency checks on each request to ensure data consistency.
[0040] However, the CPU, GPU and display controller communicate through the on-chip network, which involves complex consistency checks that introduce additional access delays, increase energy consumption, and reduce the processing efficiency of graphics rendering requests when performing graphics rendering. In view of this, one or more embodiments of the present application provide a graphics rendering processing method and an on-chip system that can solve the above problems and improve the processing efficiency of graphics rendering requests.
[0041] See also Figure 1 In one embodiment of the present application, a graphics rendering processing method is provided. The method is applied to a system on chip, wherein the system on chip includes a central processing unit, an on-chip network, a graphics processor, a display controller, and a shared memory shared by the central processing unit and the graphics processor. The method may include the following steps:
[0042] S1: The central processing unit writes a graphics rendering request into the shared memory via the on-chip network, and sends description information of the graphics rendering request to the graphics processing unit via the on-chip network.
[0043] The graphics rendering request carries description information of the graphics rendering request, which includes address information of the graphics rendering request stored in the shared memory, and is used by the graphics processor to address and access the shared memory. The central processing unit writes the graphics rendering request to the shared memory through the on-chip network, and the on-chip network performs a consistency check on the graphics rendering request to ensure data consistency. At the same time, the on-chip network sends the description information of the graphics rendering request to the graphics processor.
[0044] S3: The graphics processor directly reads the graphics rendering request from the shared memory based on the received description information, and writes the rendering result obtained by processing the graphics rendering request into the shared memory.
[0045] The above rendering result is the final output image data after the graphics processor performs the graphics rendering process according to the graphics rendering request, which usually includes the color information, depth information and resolution of the pixel. Exemplarily, the above rendering result can be output to the screen of an external display device, or it can be output to a texture for use as an input for post-processing or other rendering processes.
[0046] In this embodiment, the graphics processor directly reads the graphics rendering request from the corresponding position of the shared memory based on the address information represented by the description information, thereby avoiding the interaction between the graphics processor and the on-chip network. Generally speaking, the graphics processor needs to access the shared memory through the on-chip network, and the on-chip network will perform a consistency check on the graphics rendering request read by the graphics processor, causing a certain access delay and additional power consumption. Therefore, the description information and the graphics rendering request are called separately, and the description information is sent to the graphics processor through the on-chip network. Since the data volume of the description information is small, the consistency check will not affect the data processing efficiency. The data volume of the graphics rendering request is large. By directly addressing and calling the graphics processor, it is possible to avoid the large access delay and power consumption caused by the consistency check on the processing of the graphics rendering request, thereby improving the processing efficiency of the graphics rendering request.
[0047] In this embodiment, after the graphics processor obtains the graphics rendering request, it processes it in response to the above-mentioned graphics rendering request, and writes the processed rendering result directly into the shared memory without writing data through the on-chip network, thereby avoiding the consistency check of the rendering result and improving the processing efficiency of the graphics rendering request.
[0048] S5: The display controller directly reads the rendering result from the shared memory, and displays the rendering result through an external display device.
[0049] The display controller reads the rendering result directly from the shared memory, also without passing through the on-chip network, and converts the format of the read rendering result into an output format suitable for the external display device, and the external display device displays the rendering result.
[0050] In view of this, the central processing unit is connected to the on-chip network and the shared memory is accessed through the on-chip network. At the same time, the graphics processor and display controller can avoid the consistency check of the on-chip network and improve the processing efficiency of graphics rendering requests by directly accessing the shared memory.
[0051] In one possible implementation, the display controller displays the rendering result according to the rendering information sent by the central processing unit. The above rendering information may include frame rate, external device identification, etc. Specifically, the central processing unit sends the rendering information of the graphics rendering request to the display controller through the on-chip network, so that the display controller selects an external display device according to the rendering information, and displays the rendering result in the selected external display device according to the rendering information. Exemplarily, the display controller can select the corresponding external display device to display the rendering result according to the external device identification represented by the rendering information, and can also determine the display frame rate and other information of the rendering result, while the external display device that is not selected cannot receive the rendering result.
[0052] In this embodiment, the display controller can obtain rendering information sent by the central processing unit through the on-chip network, thereby ensuring the correctness of rendering while improving data processing efficiency.
[0053] See also Figure 2 In one possible implementation, based on S1, the graphics rendering request includes rendering instructions and rendering data, which are placed in the instruction queue and data queue of the shared memory respectively. The above-mentioned rendering instructions are used to represent the rendering tasks to be executed, and the above-mentioned rendering data represent the graphics data to be processed. The description information carried by the above-mentioned graphics rendering request is used to represent the storage address of the rendering instructions in the shared memory. Specifically, the central processing unit writes the rendering instructions in the graphics rendering request into the instruction queue of the shared memory and writes the rendering data in the graphics rendering request into the data queue of the shared memory through the on-chip network, so as to place the rendering instructions and rendering data into different queues, which is convenient for the graphics processor partition call, thereby effectively managing the memory data.
[0054] Furthermore, in this embodiment, the graphics processor directly reads the rendering instruction from the shared memory based on the storage address represented by the received description information, and parses the rendering instruction to obtain the storage address of the rendering data in the shared memory. The graphics processor reads the rendering data from the shared memory according to the storage address of the rendering data in the shared memory. By directly addressing and calling the shared memory, the graphics processor avoids data interaction with the on-chip network, and directly reads the rendering instruction and rendering data in the shared memory, thereby improving the processing efficiency of the graphics rendering request.
[0055] See also Figure 3 In one possible implementation, based on S3, when the graphics processor writes the rendering result obtained by processing the graphics rendering request into the shared memory, the graphics processor will generate an intermediate rendering result. The above intermediate rendering result can be understood as intermediate data generated in the process of processing the graphics rendering request, and the above intermediate rendering result will also be written into the shared memory.
[0056] Specifically, during the process of processing the graphics rendering request, the graphics processor writes the generated intermediate rendering result into the first result queue of the shared memory; after completing the processing of the graphics rendering request, the graphics processor writes the generated final rendering result into the second result queue of the shared memory; the display controller directly reads the final rendering result from the second result queue of the shared memory, and displays the final rendering result through an external display device.
[0057] In this embodiment, the intermediate rendering results and the final rendering results generated by the graphics processor are placed in different result queues in the shared memory, and the address information of the first result queue and the second result queue is pre-written into the description information and the rendering information by the central processor. The graphics processor writes the rendering result into the designated result queue of the shared memory according to the address information represented by the description information, and the display controller reads the final rendering result under the second result queue according to the address information represented by the rendering information, thereby ensuring the correctness of the rendering.
[0058] See also Figure 4 and Figure 5 In an optional embodiment, the system on chip also includes a routing network, and the routing network is used to provide a data processing and interface channel for the graphics processor and the display controller to access the shared memory.
[0059] Specifically, the routing network receives a first access request for the shared memory initiated by the graphics processor, and receives a second access request for the shared network initiated by the display controller, merges the first access request and the second access request into a request data stream, and performs data interaction with the shared memory based on the merged request data stream.
[0060] The routing network sets respective request identifiers for the first access request and the second access request, respectively, so as to correctly feedback the obtained request data. After the routing network obtains the request data from the shared memory, based on the request identifier carried in the request data, the obtained request data is fed back to the graphics processor or the display controller.
[0061] Exemplarily, a request identifier can be added to the access request by extending the identifier bit. For example, the 2-bit identifier '00' of the original first access request and the second access request is expanded to 3-bit identifiers '100' and '000' to distinguish the access requests, and the corresponding request identifier is also added to the obtained request data to correctly transmit the request data.
[0062] Furthermore, the routing network has a first data queue of the graphics processor and a second data queue of the display controller. The request data can be stored in the corresponding data queue according to the request identifier so that the graphics processor and the display controller can read it correctly. If the request data carries the request identifier corresponding to the first access request, the request data is written into the first data queue. If the request data carries the request identifier corresponding to the second access request, the request data is written into the second data queue.
[0063] Exemplarily, the first access request may be a request for a graphics processor to write a rendering result, the second access request may be a request for a display controller to retrieve a rendering result, and the request data may be a rendering result to be returned. The first access request and the second access request are merged to form a request data stream to interact with the shared memory, the routing network stores the request data returned from the shared memory in the second data queue according to the request identifier, and the display controller calls the request data in the second data queue to implement writing and retrieving of the rendering result.
[0064] In this embodiment, the routing network also has a first request queue of the graphics processor and a second request queue of the display controller, the first request queue is used to store the first access request, and the second request queue is used to store the second access request. It should be noted that the transmission efficiency of the request data stream formed by merging the first access request and the second access request is greater than or equal to the sum of the transmission efficiencies of the first access request and the second access request transmitted separately. The above transmission efficiency is jointly determined by the clock frequency and the bit width, that is, the transmission efficiency of the request data stream is jointly determined by the bit width and the clock frequency of the first request queue and the second request queue.
[0065] For example, if the clock frequency of the first access request and the second access request are both 1 GHz and the bit width is 128 bits, then the merged request data stream requires at least a clock frequency of 2 GHz and a bit width of 256 bits, thereby ensuring that the merged request data stream does not reduce data transmission efficiency.
[0066] In addition, when the clock frequency of the first request queue or the clock frequency of the second request queue is inconsistent with the clock frequency of the preset output queue, the first cross-clock unit and the second cross-clock unit in the routing network perform clock domain conversion to solve the problem of clock asynchrony. Specifically, the first cross-clock unit receives the access request in the first request queue, the second cross-clock unit receives the access request in the second request queue, and the first cross-clock unit and the second cross-clock unit output the access request received by each according to the clock frequency of the preset output queue.
[0067] In view of this, the present embodiment proposes a routing network for a graphics processor and a display controller to access a shared network. The above-mentioned routing network can pre-merge the access requests of the graphics processor and the display controller, reduce the data interface of the shared memory, avoid the bandwidth limitation and access conflict of the shared memory, and thus improve the data processing efficiency of the graphics rendering request. At the same time, the routing network adds different identifiers to the requests of the graphics processor and the display controller to identify the corresponding request source based on the identifier carried in the returned data, and write the request data into different data queues, thereby ensuring the correct transmission of the request data.
[0068] Please refer to Figure 6 The present application provides one or more embodiments of the above-mentioned graphics rendering processing method, and the embodiment is performed according to the following steps:
[0069] Step 1: The CPU issues graphics rendering instructions and passes description information and rendering information to the GPU and display controller.
[0070] Specifically, the CPU writes the graphics rendering instructions into the shared memory through the on-chip network according to the rendering requirements, wherein the graphics rendering instructions include rendering instructions and rendering data, the rendering instructions are written into the instruction queue of the shared memory, and the rendering data are written into the data queue of the shared memory. At the same time, the CPU sends the description information of the graphics rendering instructions to the GPU through the on-chip network, and sends the rendering information to the display controller through the on-chip network.
[0071] Step 2: The GPU accesses the shared memory according to the description information, obtains and processes the rendering data of the graphics rendering instructions, and writes the final rendering result obtained by processing into the shared memory.
[0072] Specifically, the GPU receives the description information, and accesses the shared memory through the routing network according to the address information indicated by the description information, reads the rendering instructions in the shared memory instruction queue, and obtains the rendering data in the data queue according to the storage address of the rendering data indicated by the rendering instructions in the shared memory. Further, the GPU processes the rendering data according to the rendering task indicated by the rendering instructions, and the intermediate rendering results obtained after processing are stored in the first result queue of the shared memory through the routing network, and the final rendering results are stored in the second result queue of the shared memory through the routing network.
[0073] Step 3: The display controller obtains the final rendering result from the shared memory, and transmits the final rendering result to the designated display device for display according to the rendering information.
[0074] Specifically, the display controller receives rendering information from the CPU, and determines the address information of the rendering result, the display device to be displayed, and the specified display frame rate according to the rendering information. At the same time, the display controller accesses the shared memory through the routing network to obtain the final rendering result in the second result queue of the shared memory, and transmits the final rendering result to the display device to be displayed for image display according to the above rendering information.
[0075] In this embodiment, the CPU and GPU have their own independent cache and hardware channels for accessing shared memory, and interact with the shared memory in an asynchronous manner. It should be noted that the consistency of data can be maintained through the GPU driver, and at the same time, the transmission of description information and rendering information can ensure the correctness of the transmitted data.
[0076] Generally, see Figure 7 In traditional graphics rendering methods, the CPU, GPU, and display controller all access shared memory through the on-chip network. Among them, the CPU request and the GPU request will share part of the cache and the on-chip network. The consistency check of the on-chip network will cause a large memory access delay for the GPU and display controller, and will also increase power consumption.
[0077] After the operation of steps 1 to 3, compared with the traditional graphics rendering method, the GPU and the display controller use asynchronous access to the memory with the CPU, and the GPU and the display controller have independent routing networks for accessing the memory area, and the routing network does not need to be checked for consistency, so the graphics processor and the display controller can avoid the access delay caused by the consistency check of the on-chip network when transmitting rendering data and rendering results. At the same time, fewer interface channels on the shared memory can avoid bandwidth limitations and access conflicts of the shared memory, thereby improving the processing efficiency of graphics rendering requests. In addition, setting different data queues in the shared memory can ensure the correct transmission of data.
[0078] See also Figure 8 The present application also provides a system on chip, the system on chip comprising a central processing unit, an on-chip network, a graphics processor, a display controller, and a shared memory shared by the central processing unit and the graphics processor, wherein:
[0079] The central processor 100 is used to write the graphics rendering request into the shared memory through the on-chip network, and send the description information of the graphics rendering request to the graphics processor through the on-chip network.
[0080] The graphics processor 200 is used to directly read the graphics rendering request from the shared memory based on the received description information, and write the rendering result obtained by processing the graphics rendering request into the shared memory.
[0081] The display controller 300 is used to directly read the rendering result from the shared memory and display the rendering result through an external display device.
[0082] Specifically,
[0083] In one embodiment, the central processor 100 is specifically used to send a graphics rendering request through the on-chip network and write it to a specified location in the shared memory, send description information to the graphics processor through the on-chip network to ensure that the graphics processor can correctly call the data in the shared memory, and send rendering information to the display controller through the on-chip network to determine the specific display information and display device.
[0084] In one embodiment, the graphics processor 200 is specifically used to receive description information and obtain rendering instructions of the graphics rendering request according to address information represented by the description information, obtain rendering data of the graphics rendering request according to the address information represented by the rendering instructions, perform data processing on the above rendering data to obtain intermediate rendering results and final rendering results, and write the above intermediate rendering results and final rendering results into the first result queue and the second result queue of the shared memory respectively.
[0085] In one embodiment, the display controller 300 is specifically used to access the second result queue in the shared memory to obtain the final rendering result, and accept rendering information from the central processing unit to determine the external display device to be displayed, and send the above-mentioned final rendering result to the external display device to be displayed for displaying the rendered image.
[0086] In one embodiment, the system on chip further includes: a routing network, the routing network is used for the graphics processor and the display controller to access the shared memory. Specifically, the routing network can store and modulate the requests of the graphics processor and the display controller, merge the modulated requests of the graphics processor and the display controller, and send the merged request data stream to the shared memory for data exchange.
[0087] The above-mentioned external display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0088] The further functional description of each of the above modules is the same as that of the corresponding method embodiments above, and will not be repeated here.
[0089] The systems, modules or devices described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0090] For the convenience of description, the above device is described in various modules according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0091] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods and systems. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware.
[0092] The present application is described with reference to the flowcharts and / or block diagrams of the methods and systems according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0093] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0095] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0096] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0097] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
[0098] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A graphics rendering processing method, characterized in that: The method is applied to a system on chip, the system on chip includes a central processing unit, an on-chip network, a graphics processing unit, a display controller, and a shared memory shared by the central processing unit and the graphics processing unit, and the method includes: The central processing unit writes the graphics rendering request into the shared memory through the on-chip network, and sends the description information of the graphics rendering request to the graphics processor through the on-chip network; The graphics processor directly reads the graphics rendering request from the shared memory based on the received description information, and writes the rendering result obtained by processing the graphics rendering request into the shared memory; The display controller directly reads the rendering result from the shared memory and displays the rendering result through an external display device.
2. The method according to claim 1, characterized in that: The graphics rendering request includes a rendering instruction and rendering data; The central processor writing the graphics rendering request into the shared memory through the on-chip network includes: The central processing unit writes the rendering instruction in the graphics rendering request into the instruction queue of the shared memory and writes the rendering data in the graphics rendering request into the data queue of the shared memory through the on-chip network.
3. The method according to claim 2, characterized in that The description information is at least used to represent the storage address of the rendering instruction in the shared memory; The graphics processor directly reading the graphics rendering request from the shared memory based on the received description information includes: The graphics processor directly reads the rendering instruction from the shared memory according to the storage address represented by the description information, and parses the rendering instruction to obtain the storage address of the rendering data in the shared memory; The graphics processor reads the rendering data from the shared memory according to the storage address of the rendering data in the shared memory.
4. The method according to claim 1, characterized in that: The method further comprises: The central processing unit sends rendering information of the graphics rendering request to the display controller through the on-chip network, so that the display controller selects an external display device according to the rendering information, and displays the rendering result in the selected external display device according to the rendering information.
5. The method according to claim 1 or 4, characterized in that: The graphics processor writes the rendering result obtained by processing the graphics rendering request into the shared memory, including: The graphics processor writes the generated intermediate rendering result into the first result queue of the shared memory during the process of processing the graphics rendering request; After completing the processing of the graphics rendering request, the graphics processor writes the generated final rendering result into the second result queue of the shared memory; Correspondingly, the display controller directly reads the final rendering result from the second result queue of the shared memory, and displays the final rendering result through an external display device.
6. The method according to claim 1, characterized in that The system on chip further includes a routing network, and the graphics processor and the display controller access the shared memory through the routing network; the method further includes: The routing network receives a first access request for the shared memory initiated by the graphics processor, and receives a second access request for the shared memory initiated by the display controller, merges the first access request and the second access request into a request data stream, and performs data interaction with the shared memory based on the merged request data stream.
7. The method according to claim 6, characterized in that After the routing network receives the first access request and the second access request, the method further includes: The routing network sets respective request identifiers for the first access request and the second access request; After the routing network obtains the request data from the shared memory, the routing network feeds back the obtained request data to the graphics processor or to the display controller based on the request identifier carried in the request data.
8. The method according to claim 7, characterized in that The routing network is provided with a first data queue of the graphics processor and a second data queue of the display controller; Feeding back the acquired request data to the graphics processor or to the display controller includes: If the request data carries a request identifier corresponding to the first access request, the request data is written into the first data queue, and if the request data carries a request identifier corresponding to the second access request, the request data is written into the second data queue.
9. The method according to claim 6, characterized in that The routing network is provided with a first request queue of the graphics processor and a second request queue of the display controller, the first request queue is used to store the first access request, and the second request queue is used to store the second access request; Among them, the transmission efficiency of the merged request data stream is greater than or equal to the sum of the transmission efficiencies of the first access request and the second access request, and the transmission efficiency of the request data stream is jointly determined by the bit width of the first request queue and the second request queue, and the clock frequency of the first request queue and the second request queue.
10. The method according to claim 9, characterized in that If the clock frequency of the first request queue or the clock frequency of the second request queue is inconsistent with the clock frequency of the preset output queue, the routing network further includes a first cross-clock unit and a second cross-clock unit; The method further comprises: The first cross-clock unit receives access requests in the first request queue, the second cross-clock unit receives access requests in the second request queue, and the first cross-clock unit and the second cross-clock unit output the access requests they receive according to the clock frequency of the preset output queue.
11. A system on chip, characterized in that: The system on chip includes a central processing unit, an on-chip network, a graphics processing unit, a display controller, and a shared memory shared by the central processing unit and the graphics processing unit, wherein: The central processing unit is used to write the graphics rendering request into the shared memory through the on-chip network, and send the description information of the graphics rendering request to the graphics processor through the on-chip network; The graphics processor is used to directly read the graphics rendering request from the shared memory based on the received description information, and write the rendering result obtained by processing the graphics rendering request into the shared memory; The display controller is used to directly read the rendering result from the shared memory and display the rendering result through an external display device.