Graphics Processor, Chip and Electronic Device
The graphics processor architecture dynamically configures virtual graphics processors by connecting data-instruction dispatchers to graphics processor cores, addressing the inefficiencies in conventional virtualization technologies and enhancing resource utilization and chip design efficiency.
Patent Information
- Application Number
- JP2024562352
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2042-12-23
AI Technical Summary
Conventional graphics processor virtualization technologies struggle to flexibly configure graphics processor cores according to actual needs, leading to inefficient resource utilization.
A graphics processor architecture that includes multiple data-instruction dispatchers and graphics processor cores, allowing for dynamic configuration of virtual graphics processors by connecting dispatchers to cores via a set of data-instruction transmission lines, enabling flexible allocation of cores based on user needs.
This solution allows for flexible configuration of graphics processor cores within virtual graphics processors, improving resource utilization and reducing congestion in chip layout and wiring stages, thereby minimizing chip area and physical layers.
Smart Images

Figure 2025517826000001_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of processors, relates to a graphics processor, and in particular, relates to a graphics processor, a chip, and an electronic device.
Background Art
[0002] A graphics processing unit (GPU), also known as a display core, a video processor, or a display chip, is a microprocessor that specializes in performing arithmetic operations related to images and graphics in personal computers, workstations, game consoles, and some mobile devices (such as tablet PCs, smartphones, etc.). The GPU reduces the dependence on the central processing unit (CPU) of the video card and completes some of the operations of the conventional CPU.
[0003] Generally, the number of graphics processors in an electronic device is somewhat limited. Therefore, in order to more efficiently utilize the limited graphics processor resources and better meet user needs, virtualization technology for graphics processors has emerged. In the virtualization of graphics processors, in order to accept simultaneous use by multiple users, it is required to virtualize an actual graphics processor into multiple virtual graphics processors. Each user uses one virtual graphics processor. Also, each virtual graphics processor can use one or more graphics processor cores. However, in the conventional virtualization technology of graphics processors, the graphics processor cores used by each virtual graphics processor are often fixed, and it is difficult to configure them flexibly according to actual needs.
Summary of the Invention
Problems to be Solved by the Invention
[0004] This application provides a graphics processor, a chip, and an electronic device in which the graphics processor cores included in each virtual graphics processor can be configured according to actual needs.
Means for Solving the Problem
[0005] In a first aspect, an embodiment of the present application provides a graphics processor. The graphics processor includes at least two data-instruction dispatchers and at least two graphics processor cores. Each of the data-instruction dispatchers is connected to at least one of the graphics processor cores. One of the data-instruction dispatchers and one of the graphics processor cores are connected via a set of data-instruction transmission lines. The graphics processor is configured to provide at least one virtual graphics processor. Each of the virtual graphics processors includes one of the data-instruction dispatchers and some or all of the graphics processor cores connected thereto.
[0006] In an implementation manner of the first aspect, the graphics processor is configured to provide n virtual graphics processors based on the received instructions. n is any positive integer less than or equal to N, and N is the number of the graphics processor cores.
[0007] In an implementation manner of the first aspect, the i-th data-instruction dispatcher of the graphics processor is connected to floor(N / n i ) of the graphics processor cores. Both i and n i are positive integers less than or equal to N, and floor is the floor function.
[0008] In an implementation manner of the first aspect, the graphics processor includes N data-instruction dispatchers. Among them, one of the data-instruction dispatchers is connected to N graphics processor cores. Also, m j -mj+1 Each of the data and instruction dispatchers is N / m j connected to m of the graphics processor cores. j and m j+1 are adjacent positive integers divisible by N, where 1 ≤ m j+1 < m j ≤ N.
[0009] In the implementation manner of the first aspect, the graphics processor further includes a data selector. The graphics processor core connected to at least two of the data and instruction dispatchers is connected to the data and instruction dispatcher via the data selector.
[0010] In the implementation manner of the first aspect, the numbers of both the data and instruction dispatcher and the graphics processor core are 8. The connection manner between the data and instruction dispatcher and the graphics processor core includes one 1-to-8 connection, one 1-to-4 connection, two 1-to-2 connections, and four 1-to-1 connections.
[0011] In the implementation manner of the first aspect, the number of physical layers of the graphics processor is configured based on the number of the data and instruction transmission lines.
[0012] In the implementation manner of the first aspect, the data and instruction dispatcher is fully connected to the graphics processor core.
[0013] In the second aspect, the embodiments of the present application provide a chip. The chip includes the graphics processor and input / output pins described in any of the implementation manners of the first aspect of the present application.
[0014] In the second aspect, the embodiments of the present application provide an electronic device. The electronic device includes the graphics processor and a memory described in any of the implementation manners of the first aspect of the present application.
Advantages of the Invention
[0015] The graphics processor provided in the embodiments of this application can provide at least one virtual graphics processor. The graphics processor cores included in each virtual graphics processor can be configured based on actual needs. Therefore, in specific applications, the graphics processor cores included in the virtual graphics processor can be flexibly configured based on actual needs.
[0016] In some embodiments of this application, by optimizing the connection method between the data and instruction dispatcher and the graphics processor core, the number of connections between the data and instruction dispatcher and the graphics processor core can be reduced, and the congestion problem in the chip placement and routing (P&R) stage can be avoided. This is beneficial for reducing the chip area. Also, in some embodiments, the number of layers in the physical layer of the graphics processor is configured based on the number of data and instruction transmission lines. In these embodiments, by optimizing the connection method between the data and instruction dispatcher and the graphics processor core, the number of layers in the physical layer of the graphics processor can be reduced.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0018] Hereinafter, the embodiments of the present application will be described by specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Furthermore, the present application may be implemented or applied by other different specific embodiments. Also, regarding each detailed matter in this specification, various supplements or modifications may be made without departing from the spirit of the present application based on different viewpoints and applications. It should be noted that, without contradiction, the following examples and the features of the examples may be combined with each other.
[0019] In the present application, unless otherwise clearly defined and limited separately, terms such as "attach", "be continuous with", "connect", "fix", etc. should be interpreted in a broad sense. For example, it may be a fixed connection, a removable connection, or an integral one. Also, it may be a mechanical connection or an electrical connection. Furthermore, it may be a direct continuity, an indirect continuity through an intermediate medium, or a communication inside two components or an interaction relationship between two components. Those skilled in the art can interpret the specific meaning of the above terms in the present application according to the specific situation.
[0020] It should be noted that the drawings provided in the following embodiments only roughly illustrate the basic idea of the present application. The drawings only show the assemblies related to the present application, and are not described based on the number, shape, and size of the assemblies actually implemented. The form, number, and ratio of each assembly actually implemented may be arbitrarily changed, and the layout and form of the assemblies may become more complex.
[0021] In the following embodiments, the present application provides a graphics processor. The application scenarios of the graphics processor include, but are not limited to, electronic devices. The electronic devices can be different types of electronic devices such as mobile phones, tablet PCs, personal computers (PCs), personal digital assistants (PDAs), smart watches, netbooks, wearable electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, in-vehicle devices, intelligent cars, smart speakers, robots, smart glasses, etc.
[0022] Referring to FIG. 1 showing a schematic structural diagram of an electronic device 100 according to an embodiment of the present application. The electronic device 100 includes a system processor 110 (which may be, for example, a CPU), a graphics processor 120, a memory 130, and a display 140.
[0023] In a specific operation, the system processor 110 can provide various operations of the user system, including user applications, data processing services, communication services, storage services, game services, or other operations, by starting up and entering the operating system (OS). The graphics processor 120 can provide operations such as graphics processing, rendering services, and correction to the system processor 110. Specifically, referring to FIG. 2, the graphics processor 120 provides operations related to an assembly including a graphics processor core (for example, 121-1, 121-2, ···, 121-k), an instruction configuration processor 122, a crossbar bus 123, etc. Note that k is a positive integer. As can be understood, operations such as graphics processing, rendering services, and correction can be completed by one or more functional modules of the graphics processor 120. For example, one functional module of the graphics processor 120 can appropriately complete one operation. The graphics processor 120 in FIG. 1 may be an independent element connected to the system processor 110 via the communication circuit 150. However, it should be understood that in other examples, the graphics processor 120 may be integrated into the system processor 110.
[0024] Memory 130 may include a random access memory (RAM), a cache memory device, or other volatile memory elements used by system processor 110 or graphics processor 120. Other volatile memory elements used by graphics processor 120 include, for example, caches integrable with graphics processor 120 such as secondary (L2) caches 124-1, 124-2, ···, 124-k in FIG. 2. Further, memory 130 may further include non-volatile memory elements such as a hard disk drive (HDD), a flash memory device, a solid state drive (SSD), or other memory devices for storing an operating system, an application, or other software or firmware used by electronic device 100.
[0025] Electronic devices 100 can communicate with each other via one or more communication links (e.g., one or more network links). For example, the communication link may use metal, glass, optical, air, space, or some other material as a transmission medium. Exemplary communication links include, for example, various communication interfaces and protocols such as Internet Protocol (IP), Ethernet, Universal Serial Bus (USB), Bluetooth (registered trademark), WiFi, or other communication signaling or communication formats (including combinations, improvements, or variants thereof). The communication link may be a direct link, may include intermediate networks, systems, or devices, or may include a logical network link transmitted via multiple physical links.
[0026] The electronic device 100 may include software such as, for example, an operating system, logs, databases, utilities, drivers, network software, user applications, data processing applications, game applications, and other software stored in a computer-readable medium. The software of the electronic device 100 may include one or more platforms controlled by a distributed computing system or cloud computing service. The software of the electronic device 100 may include logical interface elements such as, for example, software-defined interfaces and application programming interfaces (APIs).
[0027] The software of the electronic device 100 can be used to control the operation of the graphics processor 120 to render graphics so as to generate data to be rendered by the graphics processor 120 and output and display it on one or more displays 140.
[0028] The system processor 110, the graphics processor 120, the memory 130, and the display 140 can communicate via a connected communication circuit 150. The exemplary communication circuit 150 may use metal, glass, optical, air, space, or some other material as a transmission medium. Various communication protocols and communication signaling such as, for example, a computer bus (including its combinations or variants) can be used for the communication circuit 150. The communication circuit 150 may be a direct link, may include an intermediate network, system, or device, or may include a logical network link that transmits via a plurality of physical links.
[0029] Figure 2 shows an example of the graphics processor 120 in the embodiment of the present application. As shown in Figure 2, the graphics processor 120 specifically includes a plurality of graphics processor cores 121-1, 121-2, ··· 121-k, an instruction configuration processor 122, a crossbar switch bus 123, and a plurality of L2 caches 124-1, 124-2, 124-3 ···, 124-k. The crossbar switch bus 123 is connected between the graphics processor core 121 and the L2 cache 124 to provide a path for the graphics processor core 121 to access the L2 cache 124 and a path for the L2 cache 124 to return data to the graphics processor core 121. In addition, the L2 cache 124 is further connected to an external memory 130 via a memory interface (MIF).
[0030] In the configuration provided in the embodiment of the present application, the system processor 110 prepares the tasks and data to be executed by the graphics processor 120 and transmits them to the graphics processor core 121 in the form of instruction configurations. Specifically, the instructions issued from the system processor 110 are received by the instruction configuration processor 122, the tasks are analyzed, and directly transmitted to the graphics processor core 121, so that the graphics processor core 121 starts to execute the tasks. Also, for the tasks, they can be transmitted from the instruction configuration processor 122 to the memory 130 via the crossbar switch bus 123, and read from the memory 130 and processed by the graphics processor core 121.
[0031] When the graphics processor core 121 executes a task, the specific process includes the graphics processor core 121 reading and processing external data related to the task from the memory 130 and writing out the data. The graphics processor core 121 performs multi-thread processing that processes a certain amount of data in one instruction. Therefore, in a typical design, in order to reduce the delay of data acquisition and data storage by the graphics processor core 121 and improve the processing efficiency of the graphics processor core 121, an L2 cache 124 is arranged between the graphics processor core 121 and the memory 130, and the L2 cache 124 pre-acquires and caches a large amount of data to reduce the waiting time of the graphics processor core 121.
[0032] In the following, the technical solutions in the embodiments of the present application will be described in detail in combination with the drawings in the embodiments of the present application.
[0033] FIG. 3A shows a schematic structural diagram of a graphics processor 300 according to an embodiment of the present application. As shown in FIG. 3A, the graphics processor 300 includes M data-instruction dispatchers 310-1 to 310-M and N graphics processor clusters 330-1 to 330-N. Both M and N are positive integers greater than or equal to 2. Each graphics processor cluster 330 includes a graphics processing pipeline. In some implementation manners, the numerical values of M and N may be the same. In some other implementation manners, the numerical values of M and N may be different. Each data-instruction dispatcher 310 can be connected to at least one graphics processor cluster 330 and can be connected to a maximum of N graphics processor clusters 330. The connection manner between the data-instruction dispatcher 310 and the graphics processor cluster 330 includes, but is not limited to, a direct connection between the data-instruction dispatcher 310 and the graphics processor cluster 330 or an indirect connection between the data-instruction dispatcher 310 and the graphics processor cluster 330 through a data selector or the like. One data-instruction dispatcher 310 and one graphics processor cluster 330 are connected through a set of data-instruction transmission lines. Each set of data-instruction transmission lines is composed of, for example, 1000 to 2000 data-instruction transmission lines.
[0034] In an embodiment of the present application, the graphics processor 300 is used to provide at least one virtual graphics processor. Each virtual graphics processor includes one data-instruction dispatcher 310 and some or all of the graphics processor cores 330 connected to the data-instruction dispatcher 310. For example, in the graphics processor 300 shown in FIG. 3A, the data-instruction dispatcher 310-1 and the graphics processor cores 330-1 and 330-N connected thereto can provide one virtual graphics processor, and the data-instruction dispatcher 310-2 and the related graphics processor core 330-2 can provide another virtual graphics processor.
[0035] In some implementation manners, the graphics processor 300 can provide a plurality of virtual graphics processors at the same time. Each virtual graphics processor may include a plurality of graphics processor cores 330, but each graphics processor core 330 is only included within one virtual graphics processor. Each user can use one virtual graphics processor. Also, each graphics processor core 330 can be used by only one user at the same time.
[0036] Optionally, FIG. 3B shows a schematic diagram of the connection relationship between the data / instruction dispatcher and the graphics processor core in an embodiment of the present application. Taking the data / instruction dispatcher 310-1 as an example, it is connected to double data rate SDRAM (DDR) via an advanced extensible interface protocol (axi) and an advanced high performance bus (ahb) interface (host interface, HI) 0. The advanced extensible interface protocol is a bus protocol corresponding to a high-performance, high-bandwidth, and low-latency on-chip bus. The advanced extensible interface protocol enables better performance by making the system on chip (SoC) have a smaller area and lower energy loss. The advanced high performance bus is a high-performance bus mainly used for connecting high-performance modules, and mainly includes features such as single clock edge operation, non-three-state execution, and burst transfer support. The graphics processor core 330-1 includes a shader module, a transform feedback (TFB) module, a position primitive assembly (PPA) module, a final primitive assembly (FPA) module, a pixel engine (PE) module, and other modules. The final primitive assembly module is used to write the intermediate result to the memory. The position primitive assembly is used to perform culling operations such as triangle back face culling and zero area culling. The FPA module is used to perform viewport frustum transformation, and the pixel engine module is used to perform operations such as alpha blending for pixels.
[0037] As is apparent from the above description, the graphics processor 300 provided in the embodiment of the present application can provide at least one virtual graphics processor to the user. By virtualizing the graphics processor 300, it becomes possible to share the graphics processor 300 among multiple users.
[0038] According to an embodiment of the present application, the graphics processor 300 is configured to provide n virtual graphics processors based on the received instructions. Here, n is any positive integer less than or equal to N. Further, N is the number of graphics processor cores 330, and may be a numerical value such as 4, 8, 16, etc. Optionally, the n virtual graphics processors include all of the N graphics processor cores 330, but the present application is not limited thereto.
[0039] In an embodiment of the present application, the i-th data-instruction dispatcher of the graphics processor is connected to floor(N / n i ) graphics processor cores. Here, both i and n i are positive integers less than or equal to N, and floor is the floor function. Referring to FIG. 4 here, taking the case of N = 4 as an example, the first data-instruction dispatcher 410-1 of the graphics processor 400 is connected to one graphics processor core 430-1 (n 1 = 3), the second data-instruction dispatcher 410-2 is connected to four graphics processor cores 430-1 to 430-4 (n 2 = 1), the third data-instruction dispatcher 410-3 is connected to two graphics processor cores 430-1 and 430-3 (n 3 = 2), and the fourth data-instruction dispatcher 410-4 is connected to two graphics processor cores 430-2 and 430-4 (n 4= 2). In the embodiments of the present application, the connection method between the data / instruction dispatcher 410 and the graphics processor core 430 includes, but is not limited to, a direct connection between the data / instruction dispatcher 410 and the graphics processor core 430, or an indirect connection between the data / instruction dispatcher 410 and the graphics processor core 430 via a data selector or the like.
[0040] The graphics processor 400 shown in FIG. 4 can be configured to provide one to four virtual graphics processors based on the received instructions. When the graphics processor 400 is configured to provide one virtual graphics processor, the user can use the four graphics processor cores 430-1 to 430-4 via the data / instruction dispatcher 410-2. Also, when the graphics processor 400 is configured to provide two virtual graphics processors, the first user can use the graphics processor cores 430-1 and 430-3 via the data / instruction dispatcher 410-2, and the second user can use the graphics processor cores 430-2 and 430-4 via the data / instruction dispatcher 410-4. Also, when the graphics processor 400 is configured to provide three virtual graphics processors, the first user can use the graphics processor core 430-1 via the data / instruction dispatcher 410-1, the second user can use the graphics processor cores 430-2 and 430-4 via the data / instruction dispatcher 410-2, and the third user can use the graphics processor core 430-3 via the data / instruction dispatcher 410-3. Also, when the graphics processor 400 is configured to provide four virtual graphics processors, the first user can use the graphics processor core 430-1 via the data / instruction dispatcher 410-1, the second user can use the graphics processor core 430-2 via the data / instruction dispatcher 410-2, the third user can use the graphics processor core 430-3 via the data / instruction dispatcher 410-3, and the fourth user can use the graphics processor core 430-4 via the data / instruction dispatcher 410-4.
[0041] As is apparent from the above description, the connection between the data / instruction dispatcher 410 and the graphics processor core 430 in the embodiment of the present application is simplified into one 1-to-4 connection (a total of 1×4 sets of connection lines), two 1-to-2 connections (a total of 2×2 sets of connection lines), and one 1-to-1 connection (a total of 1×1 set of connection lines). Therefore, there are a total of 9 sets of connection lines between the data / instruction dispatcher 410 and the graphics processor core 430 in the embodiment of the present application. Thus, compared with the method in which the data / instruction dispatcher 410 is fully connected to the graphics processor core 430, the connection method provided in the embodiment of the present application has fewer required connection lines, so it is advantageous for avoiding the congestion problem in the P&R stage and reducing the chip area.
[0042] It should be understood that the connection method between the data / instruction dispatcher 410 and the graphics processor core 430 in the case of N = 4 shown in FIG. 4 is only one of the executable methods in the embodiment of the present application, and the present application is not limited thereto. In some implementation manners, the graphics processor core 430 connected to the data / instruction dispatcher 410 may be different from that in FIG. 4. For example, the data / instruction dispatcher 410-1 may be connected to 430-2 instead of 430-1, and the data / instruction dispatcher 410-3 may be connected to 430-2 and 430-4 instead of 430-1 and 430-3. Also, in some other implementation manners, the number of the graphics processor cores 430 connected to the data / instruction dispatcher 410 may be different from that in FIG. 4. For example, the data / instruction dispatcher 410-1 may be connected to two, three, or four graphics processor cores 430, and the data / instruction dispatcher 410-2 may be connected to one, two, or three graphics processor cores 430.
[0043] As a point to be described, in order to improve the utilization efficiency of the cores, all the virtual graphics processors provided by the graphics processor 400 exemplified above use all four graphics processor cores at the same time, but the present application is not limited thereto. For example, when the graphics processor 400 is configured to provide only one virtual graphics processor, the user may use two graphics processor cores 430-1 and 430-3 via the data / instruction dispatcher 410-2. In this case, the remaining two graphics processor cores 430-2 and 430-4 are in an idle state. Also, for example, when the graphics processor 400 is configured to provide two virtual graphics processors, the first user may use one graphics processor core 430-1 via the data / instruction dispatcher 410-1, and the second user may use two graphics processor cores 430-2 and 430-4 via the data / instruction dispatcher 410-4. In this case, the remaining one graphics processor core 430-3 is in an idle state.
[0044] In one embodiment of the present application, the graphics processor includes N data / instruction dispatchers. Among them, one data / instruction dispatcher is connected to N graphics processor cores. Also, m j -m j+1 data / instruction dispatchers are connected to N / m j graphics processor cores. m j and m j+1 are adjacent positive integers divisible by N, and 1 ≤ m j+1 < m j ≤ N. For example, when N = 4, the numerical values of m j and m j+1 include m j+1 = 1 and m j = 2, and m j+1 = 2 and m jTwo types where = 4 are included. Based on this, when the graphics processor provided in the embodiment of the present application includes four data-instruction dispatchers 410, one data-instruction dispatcher 410 is connected to four graphics processor cores 430, and another data-instruction dispatcher 410 is connected to two graphics processor cores 430 (m j+1 = 1 and m j = 2), and the other two data-instruction dispatchers 410 are each connected to one graphics processor core 430 (m j+1 = 2 and m j = 4). In the embodiment of the present application, the connection method between the data-instruction dispatcher 410 and the graphics processor core 430 includes, but is not limited to, the direct connection between the data-instruction dispatcher 410 and the graphics processor core 430, or the indirect connection between the data-instruction dispatcher 410 and the graphics processor core 430 through a data selector or the like.
[0045] Subsequently, the cases of N = 8 and N = 16 are each exemplified to explain the above connection scheme in detail. Here, refer to FIG. 5A. In one example, N = 8, and the graphics processor 500 includes eight data-instruction dispatchers. Among them, one data-instruction dispatcher 510-1 is directly connected to the graphics processor core 530-1 and indirectly connected to the graphics processor cores 530-2 to 530-8 through a data selector. Also, one data-instruction dispatcher 510-5 is indirectly connected to four graphics processor cores 530-5 to 530-8 through a data selector (m j+1 = 1, m j= 2). Also, for the two data-instruction dispatchers 510-3 and 510-7, the data-instruction dispatcher 510-3 is indirectly connected to the two graphics processor cores 530-3 and 530-4 via a data selector, and the data-instruction dispatcher 510-7 is indirectly connected to the two graphics processor cores 530-7 and 530-8 via a data selector (m j+1 = 2, m j = 4). Also, the four data-instruction dispatchers 510-2, 510-4, 510-6, and 510-8 are each indirectly connected to the corresponding graphics processor cores 530-2, 530-4, 530-6, and 530-8 via a data selector (m j+1 = 4, m j = 8).
[0046] The graphics processor 500 shown in FIG. 5A can be configured to provide 1 to 8 virtual graphics processors and can support at least one user (when the graphics processor 500 is configured to provide one virtual graphics processor), and can support up to 8 users (when the graphics processor 500 is configured to provide 8 virtual graphics processors). The graphics processor cores 530 of the graphics processor 500 have 22 possible allocation situations. Specifically, it is shown in Table 1 below. For example, the graphics processor 500 in the 10th case is configured to provide three virtual graphics processors for three users. In this case, two users occupy three graphics processor cores 530, and another user occupies two graphics processor cores 530. Combining with FIG. 5A, the allocation plan of the graphics processor cores 530 may be such that the first user occupies the graphics processor cores 530-1, 530-2, and 530-8, the second user occupies the graphics processor cores 530-5, 530-6, and 530-7, and the third user occupies the graphics processor cores 530-3 and 530-4.
[0047] As a point to be explained, the connection method between the data and instruction dispatcher 510 and the graphics processor core 530 in the embodiments of the present application is not the only one, and in some other implementation methods, other connection methods may be adopted. For example, FIG. 5B shows another connection scheme when N = 8. This scheme has the same effect as the connection scheme shown in FIG. 5A.
[0048]
Table 1
[0049] In the above example, the connection between the data and instruction dispatcher 510 and the graphics processor core 530 is simplified to one 1-to-8 connection (a total of 1×8 sets of connection lines), one 1-to-4 connection (a total of 1×4 sets of connection lines), two 1-to-2 connections (a total of 2×2 sets of connection lines), and four 1-to-1 connections (a total of 4×1 sets of connection lines). Therefore, there are a total of 20 sets of connection lines between the data and instruction dispatcher 510 and the graphics processor core 530 in the embodiments of the present application. Each set of connection lines consists of, for example, 1000 to 2000 data and instruction connection lines. Therefore, compared with the method of fully connecting the data and instruction dispatcher 510 to the graphics processor core 530, the connection method provided in this example has fewer required connection lines, so it is advantageous to avoid the congestion problem in the P&R stage and reduce the chip area.
[0050] Refer to FIG. 5C. In another example, N = 16, and the graphics processor 500 includes 16 data-instruction dispatchers 510. Among them, the first data-instruction dispatcher 510 is directly connected to the first graphics processor core 530 and indirectly connected to the remaining 15 graphics processor cores 530 via a data selector. Also, the ninth data-instruction dispatcher 510 is indirectly connected to eight graphics processor cores 530 via a data selector. Also, the second data-instruction dispatcher 510 is indirectly connected to five graphics processor cores 530 via a data selector. Also, the seventh data-instruction dispatcher 510 is indirectly connected to four graphics processor cores 530 via a data selector. Also, the sixteenth data-instruction dispatcher 510 is indirectly connected to three graphics processor cores 530 via a data selector. Also, the fourth, eleventh, and thirteenth data-instruction dispatchers 510 are indirectly connected to two graphics processor cores 530 via a data selector. Also, the third, fifth, sixth, eighth, tenth, twelfth, fourteenth, and fifteenth data-instruction dispatchers 510 are indirectly connected to one graphics processor core 530 via a data selector. In this example, the connection between the data-instruction dispatcher 510 and the graphics processor core 530 is simplified to one 1-to-16 connection (a total of 1×16 sets of connection lines), one 1-to-8 connection (a total of 1×8 sets of connection lines), one 1-to-5 connection (a total of 1×5 sets of connection lines), one 1-to-4 connection (a total of 1×4 sets of connection lines), one 1-to-3 connection (a total of 1×3 sets of connection lines), three 1-to-2 connections (a total of 3×2 sets of connection lines), and eight 1-to-1 connections (a total of 8×1 sets of connection lines). Therefore, there are a total of 50 sets of connection lines between the data-instruction dispatcher 510 and the graphics processor core 530 in the embodiments of the present application.Therefore, compared with the method in which the data and instruction dispatcher 510 is fully connected to the graphics processor core 530, the connection method provided in this exemplary embodiment requires fewer connection lines, thus avoiding the congestion problem in the P&R stage and being advantageous for reducing the chip area.
[0051] In an embodiment of the present application, the graphics processor 500 may further include a data selector. The graphics processor core 530 connected to at least two data and instruction dispatchers 510 is indirectly connected to the data and instruction dispatcher 510 via a data selector. For example, the graphics processor 500 shown in FIG. 5A includes data selectors 520-2 to 520-8. In the case of the graphics processor core 530 connected to at least two data and instruction dispatchers 510, for example, the graphics processor core 530-5 is indirectly connected to the corresponding data and instruction dispatcher 510 via a data selector. In the embodiment of the present application, the data selector selects at most one data and instruction dispatcher 510 at the same time and connects it to the graphics processor core 530.
[0052] Optionally, in the case of the graphics processor core 530 connected to only one data and instruction dispatcher 510, for example, the graphics processor core 530-1 may be connected to the data and instruction dispatcher without passing through a data selector.
[0053] It should be understood that the data selector in the embodiments of the present application is not limited to any specific device or circuit, and includes any device or circuit capable of selecting and outputting a predetermined one signal from a set of input signals. In an embodiment of the present application, the data-instruction dispatcher 510 is fully connected to the graphics processor core 530. By "fully connected" it means that each data-instruction dispatcher 510 is connected to all the graphics processor cores 530, and each graphics processor core 530 is connected to all the data-instruction dispatchers 510. The connection method between the data-instruction dispatcher 510 and the graphics processor core 530 includes, but is not limited to, a direct connection between the data-instruction dispatcher 510 and the graphics processor core 530, or an indirect connection between the data-instruction dispatcher 510 and the graphics processor core 530 via a data selector or the like. FIG. 6 shows a schematic diagram of the case where the data-instruction dispatcher 510 is fully connected to the graphics processor core 530 when N = 8. In this case, the graphics processor 500 can be configured to provide 1 to 8 virtual graphics processors, and can satisfy all the allocation situations of the processor cores.
[0054] In an embodiment of the present application, the number of physical layers of the graphics processor is configured based on the number of data-instruction transmission lines. Specifically, the number of data-instruction transmission lines included in each physical layer of the graphics processor is limited to a maximum. The fewer the number of data-instruction transmission lines, the fewer the number of physical layers of the graphics processor. Taking the case of N = 8 as an example, when the data-instruction dispatcher is fully connected to the graphics processor core, the number of data-instruction transmission lines is 64 pairs, and the number of physical layers of the graphics processor is configured to be 8 layers. On the other hand, when using the connection method shown in FIG. 5A or FIG. 5B, the number of data-instruction transmission lines is 20 pairs, and the number of physical layers of the graphics processor 500 can be configured to be 4 layers. In this case, by using the connection method shown in FIG. 5A or FIG. 5B, the physical number of layers of the graphics processor 500 can be reduced.
[0055] The present application further provides a chip. FIG. 7 shows a schematic structural diagram of a chip in an embodiment of the present application. The chip includes the graphics processor and input / output pins described in any of the embodiments of the present application.
[0056] The present application further provides an electronic device. The electronic device includes the graphics processor described in any of the embodiments of the present application and a memory communicatively connected to the graphics processor.
[0057] As described above, the graphics processor provided in the embodiments of the present application can provide at least one virtual graphics processor. Through the virtualization of the graphics processor, it becomes possible for multiple users to share the graphics processor jointly. In some embodiments of the present application, by optimizing the connection method between the data and instruction dispatcher and the graphics processor core, the number of connection lines between the data and instruction dispatcher and the graphics processor core can be reduced, thereby avoiding the congestion problem in the chip layout and wiring stages. This is advantageous for reducing the chip area and the number of physical layers of the graphics processor. Therefore, the present application effectively eliminates various drawbacks in the prior art and has high industrial utility value.
[0058] The above embodiments are only illustrative explanations of the principles and effects of the present application and do not limit the present application. Those skilled in the art can supplement or modify the above embodiments on the premise of not departing from the spirit and scope of the present application. Therefore, any equivalent supplements or modifications completed by those skilled in the art without departing from the spirit and technical idea disclosed in the present application are still included in the scope of the claims of the present application.
Description of Reference Numerals
[0059] 100 Electronic device 110 System processor 120 Graphics processor 121-1 to 121-k Graphics Processor Cores 122 Instruction Configuration Processor 123 Crossbar Switch Bus 124-1 to 124-k L2 Caches 130 Memory 140 Display 300 Graphics Processor 310-1 to 310-M Data / Instruction Dispatcher 330-1 to 330-N Graphics Processor Cores 500 Graphics Processor 510-1 to 510-8 Data / Instruction Dispatcher 520-2 to 520-8 Data Selector 530-1 to 530-8 Graphics Processor Cores
Claims
1. A graphics processor, comprising: at least two data-instruction dispatchers and at least two graphics processor cores, each of said data-instruction dispatchers being connected to at least one of said graphics processor cores, and one of said data-instruction dispatchers and one of said graphics processor cores being connected via a set of data-instruction transmission lines; said graphics processor being configured to provide at least one virtual graphics processor, each of said virtual graphics processors comprising one of said data-instruction dispatchers and some or all of said graphics processor cores connected to said one data-instruction dispatcher.
2. The graphics processor according to claim 1, wherein said graphics processor is configured to provide n virtual graphics processors based on received instructions, where n is any positive integer less than or equal to N, and N is the number of said graphics processor cores.
3. The i-th data instruction dispatcher of the graphics processor is connected to floor(N / n i ) graphics processor cores, where both i and n i are positive integers not exceeding N, and floor is the floor function. The graphics processor according to claim 2, characterized in that.
4. The graphics processor includes N of the data-instruction dispatchers, wherein one of the data-instruction dispatchers is connected to N of the graphics processor cores, and m j -m j+1 of the data-instruction dispatchers are connected to N / m j of the graphics processor cores, and m j and m j+1 are adjacent positive integers divisible by N, and 1 ≦ m j+1 < m j ≦ N. The graphics processor according to claim 3, characterized in that.
5. The graphics processor according to claim 4, further comprising a data selector, wherein the graphics processor cores connected to at least two of said data-instruction dispatchers are connected to said data-instruction dispatchers via said data selector.
6. The graphics processor according to claim 4, wherein the number of said data-instruction dispatchers and the number of said graphics processor cores are both eight, and the connection method between said data-instruction dispatchers and said graphics processor cores includes one 1-to-8 connection, one 1-to-4 connection, two 1-to-2 connections, and four 1-to-1 connections.
7. The graphics processor according to claim 1, wherein the number of physical layers of said graphics processor is configured based on the number of said data-instruction transmission lines.
8. The graphics processor according to claim 1, wherein said data-instruction dispatcher is fully connected to said graphics processor core.
9. A chip characterized by including the graphics processor according to any one of claims 1 to 8 and input / output pins.
10. An electronic device characterized by including the graphics processor according to any one of claims 1 to 8 and a memory.
Citation Information
Patent Citations
Dynamic and application-specific virtualized graphics processing
US20190236751A1
Highly parallel virtualized graphics processors
US20220383445A1
Systems and methods for remote graphics processing unit service
US9576332B1