Computer system, graphics processor, and drawing command processing method
By setting up a special command processor in a multi-chip graphics processor, ensuring that the computing tasks are processed within the same chip, solving the cross-chip computing efficiency problem, improving efficiency and reducing latency, optimizing resource utilization, and enhancing the flexibility and scalability of the system.
Patent Information
- Application Number
- CN202410629039.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-05-20
AI Technical Summary
The performance problems of existing graphics processor virtualization technology in cross-chip computing have not been effectively solved, resulting in inefficient computing and increased latency.
By setting the first and second command processors in the multi-chip graphics processor, the computing resources of the first and second chips are managed respectively, ensuring that the computing tasks are only dispatched within the same chip, avoiding cross-chip data exchange.
Improves processing efficiency, reduces system latency, optimizes resource utilization, enhances system flexibility and scalability, and is suitable for high-performance and multi-tasking virtualized computing applications.
Smart Images

Figure CN118350980B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of graphics processing units, and particularly to a command processor in a multi-chip graphics processing unit. Background Art
[0002] Existing graphics processing unit technologies have evolved into multi-chip architectures. A large number of computing units and execution units are arranged in various partitions or clusters for parallel computing. To efficiently utilize and manage computing resources, virtualization technologies have emerged, allowing multiple virtual machines, containers, programs, or tenants in a computer system to share the computing resources in the same graphics processing unit. Virtualization technologies are particularly important in cloud computing, data centers, and artificial intelligence computing scenarios that require multi-tenant environments. There are several common virtualization technologies. For example, the passthrough technology allows a virtual machine to directly control a physical GPU, bypassing the host operating system. The advantage is that it can provide near-native performance, but the resources are exclusive, and it does not support sharing of GPU resources among multiple virtual machines. Another virtualization technology is that a virtual machine calls a specific remoting API to obtain the GPU computing resources on the host. The benefit is that multiple virtual machines can share the GPU, but it will increase additional latency and overhead, and the performance is lower than direct access. Hardware-assisted virtualization technologies, such as Single Root I / O Virtualization (SR-IOV), can split physical resources into multiple independent virtual resources for direct access by different virtual machines. Its advantage is that it can improve performance and resource isolation, but the resource allocation is statically set and lacks flexibility. Some well-known brands have also launched exclusive GPU virtualization management technologies, such as NVIDIA vGPU, which allows the resources of a single GPU to be dynamically allocated to multiple virtual machines. However, the above solutions require expensive dedicated hardware and driver software and are not cost-effective. In addition, none of the above virtualization technologies can effectively solve the problem of cross-chip computing efficiency. If the virtual computing clusters allocated to a computer tenant contain multiple physical computing units that are scattered across multiple chips, a large number of cross-chip bus transmissions will occur during the computing process, seriously affecting the efficiency. Therefore, a technology that can effectively configure virtual resources and avoid efficiency problems remains to be developed. Summary of the Invention
[0003] The present disclosure aims to solve the efficiency bottleneck problem of graphics processing unit virtualization technology. Through the command queue scheduling of the command processor, true multi-chip parallel processing and virtualization management are achieved, thereby fully improving the processing efficiency and system response speed.
[0004] To solve the above technical problems, embodiments of the present disclosure provide a computer system, including a memory, a central processing unit, software, and a graphics processing unit. The graphics processing unit includes a first chip including a first command processor; a second chip including a second command processor; the graphics processing unit further includes a plurality of computing units disposed in a first computing partition and a second computing partition in the first chip, and a third computing partition and a fourth computing partition in the second chip. The first command processor and the second command processor are configured to process drawing commands provided by the computer system to dispatch computing tasks to the computing units. The central processing unit drives the graphics processing unit through the software, such that the first command processor only dispatches computing tasks to the computing units in the first chip, and the second command processor only dispatches computing tasks to the computing units in the second chip.
[0005] In another embodiment, the present disclosure further provides a graphics processing unit that can be disposed in a computer system including a memory, a central processing unit, and software. The graphics processing unit includes a first chip and a second chip. The first chip includes a first command processor, and the second chip includes a second command processor. The graphics processing unit includes a plurality of computing units disposed in a first computing partition and a second computing partition in the first chip, and a third computing partition and a fourth computing partition in the second chip. The first command processor and the second command processor are configured to process drawing commands provided by the computer system to dispatch computing tasks to the computing units. The central processing unit drives the graphics processing unit through the software, such that the first command processor only dispatches computing tasks to the computing units in the first chip, and the second command processor only dispatches computing tasks to the computing units in the second chip.
[0006] In a further embodiment, the present disclosure further provides a method for processing drawing commands, which runs in a computer system including a memory, a central processing unit, a graphics processing unit, and software, wherein the graphics processing unit includes a first chip including a first command processor, and a second chip including a second command processor. The method for processing drawing commands includes: setting the first command processor and the second command processor to be capable of processing drawing commands provided by the computer system to dispatch computing tasks to the computing units; setting the first command processor to only dispatch computing tasks to the computing units in the first chip; and setting the second command processor to only dispatch computing tasks to the computing units in the second chip.
[0007] One of the advantages of the above embodiments is that it can improve processing efficiency and reduce system latency. Through the condition setting of the command processor, any computing task will only be dispatched to the same chip for processing, avoiding the burden of cross-chip data exchange, improving processing efficiency and reducing system latency.
[0008] Another advantage of the above embodiments is that it can optimize resource utilization. The present disclosure classifies the drawing commands and establishes multiple software queues, and then binds them to the command queues that match the conditions, so that the computing cluster resources in all chips can be effectively utilized.
[0009] Another advantage of the above embodiments is that it can increase the flexibility and scalability of the system. The computer system can dynamically bind the command queues according to the type requirements of the graphics operation, and configure the cluster size and quantity in each computing partition, and flexibly adapt to different application requirements on the premise of not generating the burden of cross-chip data exchange.
[0010] In summary, the multi-chip graphics processor command processor system of the present disclosure not only improves performance and efficiency by effectively allocating and managing resources, but also enhances the stability and scalability of the system, and is very suitable for virtualized computing applications that require high performance and multi-task processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present disclosure, and constitute a part of the present disclosure. The illustrative embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure, and do not constitute an improper limitation of the present disclosure. In the drawings:
[0012] Figure 1 is an embodiment of the computer system 100 of the present disclosure.
[0013] Figure 2 is an embodiment of the command processor of the present disclosure.
[0014] Figure 3 is an embodiment of the computing partition of the present disclosure.
[0015] Figure 4 is an embodiment of the operation of the software queue of the present disclosure.
[0016] Figure 5 is a flowchart of the method for processing drawing commands of the present disclosure.
[0017] REFERENCE NUMERALS
[0018] 100: Computer system
[0019] 102: Central processing unit
[0020] 400: Memory
[0021] 106: Software
[0022] 108: Graphics Processing Unit
[0023] 110: First Chip
[0024] 120: Second Chip
[0025] 210: First Command Processor
[0026] 220: Second Command Processor
[0027] 310: First Computing Partition
[0028] 320: Second Computing Partition
[0029] 330: Third Computing Partition
[0030] 340: Fourth Computing Partition
[0031] 201: Command Queue
[0032] 202: Interpretation Module
[0033] 203: Dispatching Module
[0034] 211: First User Pipeline
[0035] 212: Second User Pipeline
[0036] 213: Management Unit
[0037] 215: Firmware
[0038] 221: Third User Pipeline
[0039] 222: Fourth User Pipeline
[0040] 223: Management Unit
[0041] 225: Firmware
[0042] 301: First Computing Cluster
[0043] 302: Second Computing Cluster
[0044] 303: Third Computing Cluster
[0045] 304: Fourth Computing Cluster
[0046] 300: Computing Unit
[0047] 410: First Software Queue
[0048] 420: Second Software Queue
[0049] 430: Third Software Queue
[0050] 440: Fourth Software Queue Detailed Implementation Manner
[0051] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0052] Figure 1 is a computer system 100 of an embodiment of the present disclosure.
[0053] The computer system 100 includes a central processing unit 102, a memory 400, software 106, and a graphics processing unit 108. The graphics processing unit 108 is connected to the central processing unit 102 and the software 106 through a specific bus (not shown). The graphics processing unit 108 in this embodiment further includes a first chip 110 and a second chip 120. The first chip 110 includes a first command processor 210, a first computing partition 310, and a second computing partition 320. Similarly, the second chip 120 includes a second command processor 220, a third computing partition 330, and a fourth computing partition 340.
[0054] The central processing unit 102 is the brain of the computer system 100, which can operate the operating system (OS) and execute the software 106 in the memory 400 to execute most logical operation commands. Modern central processing units usually have multiple cores and multiple threads, and can execute multiple tasks simultaneously to improve computing efficiency and processing speed.
[0055] The memory 106 is coupled to the central processing unit 102 and can store the data required during the operation of the central processing unit 102. For example, the central processing unit 102 can load the software 106 into the memory 400 and then execute its functions. The data during the operation of the central processing unit 102 can also be stored in the central processing unit 102. Generally speaking, the central processing unit 102 can be a random access memory (RAM). In implementation, the present disclosure does not limit the type of the software 106.
[0056] The software 106 is a type of code data stored in a storage device (not shown) of the computer system 100 and is used to guide how the computer system 100 operates. The software 106 referred to in this embodiment particularly includes a driver for the graphics processing unit 108, which can dynamically set the function settings of the graphics processing unit 108. Regarding how the software 106 operates in the computer system 100, it will be described in detail in Figure 4 below.
[0057] The graphics processing unit 108 is a device specifically designed to accelerate image creation for output to a display device and is very important in performing graphics rendering and processing as well as general computing tasks. Because the graphics processing unit 108 has a highly parallel structure, it can also be applied to related calculations in blockchain or artificial intelligence. For ease of understanding, this embodiment is described in a way that the graphics processing unit 108 includes two chips. In implementation, the present disclosure does not limit the number of chips included in the graphics processing unit 108.
[0058] The first chip 110 or the second chip 120 in this embodiment represents a physical die. The first chip 110 mainly includes a first command processor 210 and a plurality of computing units 300 (not shown). According to the physical distribution position in the computer system 100, the computing units 300 can also be roughly classified into a first computing partition 310 and a second computing partition 320. In other words, a specific number of computing units 300 are included in the first computing partition 310, and a specific number of computing units 300 are also included in the second computing partition 320. In this embodiment, the first chip 110 is only divided into two computing partitions, each including the same number of computing units 300. However, it can be understood that in implementation, each chip in the graphics processing unit 108 can be set to more than two computing partitions through software 106, and the number and characteristics of the computing units 300 in each computing partition can also be dynamically adjusted according to requirements. On the other hand, the settings of the third computing partition 330 and the fourth computing partition 340 in the second chip 120 are also based on the same principle and will not be repeated.
[0059] The first chip 110 can be controlled by the central processing unit 102 to change the scheduling and dispatching behavior of drawing commands. For example, the central processing unit 102 can execute the software 106 and send corresponding setting information to the first chip 110 according to the requirements of the computer system 100 for drawing tasks, so that the first chip 110 only dispatches the computing tasks generated after processing the drawing commands to the computing units 300 in the first chip 110. In other words, the computing tasks generated after the first command processor 210 processes the drawing commands will not be sent to the computing units 300 outside the first chip 110. On the other hand, the implementation manner of the second command processor 220 is similar. The computing tasks generated after the second command processor 220 processes the drawing commands will not be sent to the computing units 300 outside the second chip 120.
[0060] In summary, in this embodiment, virtualization technology is implemented through the queue arrangement of the command processor, and the computing tasks are prevented from being assigned to the computing units 300 of different chips. For example, if a computing task requires n computing units 300, through the scheduling of this embodiment, these n computing units 300 must come from the same chip, and there will be no situation across multiple chips. This approach can significantly reduce the resource consumption caused by the cross-chip transmission burden.
[0061] In the graphics processor 108 architecture of this embodiment, various drawing commands from the central processor 102 can be passed through graphics application programming interfaces (APIs) such as DirectX, OpenGL, Vulkan, or computing APIs (such as CUDA, OpenCL). The types of drawing commands can be roughly classified into rendering commands, state setting commands, computing commands, memory operation commands, synchronization commands, query commands, and image encoding commands and image decoding commands.
[0062] The command processor can be implemented as a multi-core and multi-threaded architecture in terms of implementation. Regarding the detailed functions of the command processor, they will be described in Figure 2 below.
[0063] Figure 1 In the graphics processor 108, there are still some bus-related components not shown in the figure, such as local buffers (LBFs) connected in series at the chip input and output ends, or link components (Data Compression Link, DCL) that bridge different chips. Since they are well-known technologies, they will not be described in detail here.
[0064] Figure 2 This is an embodiment of the command processor of the present disclosure. The first command processor 210 mainly includes a management unit 213, a first user pipeline 211, and a second user pipeline 212. The management unit 213 includes firmware 215. The management unit 213 is driven and set by the firmware 215 to control and schedule various working modules in the first command processor 210. In other words, the firmware 215 can drive the first management unit 213 to communicate with the central processor 102, so that the first command processor 210 and the second command processor 220 can process the drawing commands in the computer system 100 and schedule the corresponding computing resources in the graphics processor 108.
[0065] Figure 2The management unit 213 therein mainly functions in command queue selection and binding. In the firmware scheduling, the management unit 213 is responsible for selecting an appropriate command queue and binding it to a specified processing flow. The management unit 213 is also responsible for queue and user pipeline management, including the creation and destruction of command queues, as well as the binding and scheduling operations with different user pipelines. In addition, the management unit 213 is also responsible for responding to the interrupt signals from the host and performing corresponding processing flows, such as reading and writing operations through memory mapped I / O (MMIO), and reading and modifying device configurations and statuses. When necessary, the management unit 213 can also perform context switching, including operations of stopping, saving, and restoring the execution context, to support multitasking and fast switching. In other words, the management unit 213 is a bridge for communication between the graphics processor 108 and the software 106, ensuring that the computing tasks generated by the first command processor 210 and the second command processor 220 can be dispatched to the corresponding computing units 300 for execution in a predetermined manner.
[0066] Since the computer system 100 may need to have multi-user and multi-tasking capabilities, the first command processor 210 in this embodiment has an expandable multi-threaded architecture. For example, the first command processor 210 may include multiple command queues, multiple decoding modules, and multiple dispatching modules. These modules can be grouped into the first user pipeline 211, the second user pipeline 212, etc. under the setting of the management unit 213. A user pipeline is a virtualized concept, generally referring to a set of modules that can independently complete the work required for a single command type, tenant, thread, virtual machine, or container. This set of modules generated by the partitioning concept can physically partition the data processed by different user pipelines and reduce the burden of context switching.
[0067] For example, the first user pipeline 211 in the first command processor 210 includes a command queue 201, a decoding module 202, and a dispatching module 203. The command queue 201 in the first user pipeline 211 is set by the management unit 213 and is associated with a software queue in the memory 400 (refer to Figure 4)Bind to capture and temporarily store the drawing commands in the software queue. The interpretation module 202 interprets the drawing commands one by one from the command queue 201 to generate corresponding computing tasks. The dispatching module 203 is coupled to the interpretation module 202 and is responsible for dispatching the interpreted computing tasks to one or more corresponding computing units 300 in the first computing partition 310 or the second computing partition 320. On the other hand, the second user pipeline 212 also has the same architecture, is controlled by the management unit 213, binds to different software queues, and operates simultaneously with the first user pipeline 211 without interference. Although two groups of user pipelines are shown in the first command processor 210 in this embodiment, it can be understood that, in implementation, the number of user pipelines supported by each command processor is not limited to 2. On the other hand, the second command processor 220 also similarly includes a management unit 223, a third user pipeline 221, and a fourth user pipeline 222. Its operating principle is similar to that of the first command processor 210 and the first user pipeline 211, so it will not be repeated.
[0068] In one embodiment, each user pipeline can be set to specifically process different types of commands, which can ensure that the graphics processor efficiently executes various graphics and computing tasks. For example, the first user pipeline 211 in the first command processor 210 in the first chip 110 can be set to focus on computing commands and memory operation commands. After converting the computing commands and memory operation commands into corresponding computing tasks, the first user pipeline 211 dispatches them to the computing units 300 in the first computing partition 310 or the second computing partition 320 for execution. The user pipelines in the second chip 120, through the setting of the central processing unit 102, can specifically process other types of commands that do not overlap with the scope of responsibility of the first command processor 210, such as image encoding commands and image decoding commands. The third user pipeline 221 and the fourth user pipeline 222 in the second command processor 220 can convert the image encoding commands or image decoding commands into corresponding computing tasks and then dispatch them to the computing units 300 in the third computing partition 330 and the fourth computing partition 340 for processing. On the other hand, the user pipelines in the first command processor 210 and the second command processor 220 are not limited to specifically processing specific types of commands, but can allocate drawing commands through different classification methods. For example, the user pipeline can determine the corresponding software queue to be bound according to various conditions such as task identification codes, resource requirements, tenants, parent programs, execution purposes, and priority levels. In implementation, the types of drawing commands of the graphics processor 108 are not limited to the above, and the binding conditions of the command queues and software queues in the first command processor 210 and the second command processor 220 are not limited to the above.
[0069] In one embodiment, the management unit 213 in the first command processor 210 and the management unit 223 in the second command processor 220 do not operate simultaneously. For example, when the management unit 213 is activated and running, the management unit 223 is in a disabled state. The operation of all user pipelines in the first command processor 210 and the second command processor 220 is coordinated and scheduled by the management unit 213. The advantage of this approach is that through the centralized management of the management unit 213, all user pipelines in the first command processor 210 and the second command processor 220 can function, without wasting idle resources. The management unit 223 only serves as a backup for the management unit 213 and does not interfere with the operation of the management unit 213, which can reduce the situation of resource contention and scheduling interference. The management unit 213 or the management unit 223 mentioned in this embodiment can be a micro control unit (MCU).
[0070] Figure 3 It is an embodiment of the computing partition and computing cluster in the first chip 110 of the present disclosure.
[0071] In this embodiment, the computing units 300 in the first chip 110 are divided into a first computing partition 310 and a second computing partition 320 according to their physical distribution locations. A compute task cluster (CTC) generally refers to a cluster formed by multiple computing units 300 and is a logical set scheduled to complete a computing task. In implementation, various computing tasks may require various different numbers of computing units 300. Therefore, when the management unit 213 allocates computing resources, computing clusters may be logically formed. For example, the first computing partition 310 may include a first computing cluster 301 and a second computing cluster 302. The second computing partition 320 may include a third computing cluster 303 and a fourth computing cluster 304. Each computing cluster is composed of multiple computing units 300 (Compute Unit, CU). In the architecture of a general graphics processor 108, each computing unit 300 can also be regarded as a microprocessor containing several execution units (Execution Units, EU). In this embodiment, since the first computing partition 310 and the second computing partition 320 are both located in the first chip 110, all the computing units 300 therein do not consume bus bandwidth in internal data transmission and communication.
[0072] Whether it is a computing partition formed by physical distribution or a computing cluster formed by logical scheduling, it can be set by either the central processing unit 102 or the management unit 213 and bound to a specific user, thread, or command type to execute the computing tasks dispatched by the first command processor 210 and the second command processor 220. In other words, on the premise of not crossing the multi-chip range, the first command processor 210 of this embodiment can perform flexible scheduling on all computing units 300 in the first chip 110.
[0073] In one embodiment, the compute unit 300 is responsible for executing various graphics and computing tasks. Common computing tasks may include graphics rendering, parallel data processing, video processing, machine learning, physical simulation, etc. Each computing unit also includes multiple execution units (EUs), registers, caches, and basic control logic, etc. In some embodiments of the graphics processor 108, a computing unit may include dozens to hundreds of execution units. The execution unit is a smaller component inside the computing unit and is responsible for executing specific computing commands. Each execution unit can be regarded as a small processor dedicated to executing vector and scalar operations. During graphics rendering or complex computing processes, each execution unit can independently execute a set of specified commands and process specific data elements. For example, the types of commands that an execution unit can process include command execution, data operation, graphics shading, flow control, resource allocation, etc.
[0074] In Figure 3In this case, each of the first computing cluster 301 and the second computing cluster 302 includes a plurality of computing units 300. Generally speaking, each computing unit 300 contains various execution units, that is, its functions can be all-round. However, for specific task requirements, after the central processing unit 102 executes the software 106, according to the operating conditions of the computer system 100 or the requirements of the drawing command operation type, an execution unit mask (EU Mask) can be generated. The execution unit mask allows the computer system 100 to control and manage the active states of various execution units in the computing unit 300 at the software level, such as activating or disabling specific types of execution units therein. For example, in order to set the computing unit 300 to focus on numerical operations or memory type calculation tasks, the central processing unit 102 can set the computing unit 300 in the first computing cluster 301 through the management unit 213, activate the execution units related to numerical operations or memory type calculation tasks therein, and disable the irrelevant execution units. In this way, the computing unit 300 in the first computing cluster 301 can become an optimized execution unit for numerical operations or memory type calculation tasks. The same concept can be applied to the second computing cluster 302, the third computing cluster 303, and the fourth computing cluster 304. The central processing unit 102 can cooperate with the software 106 to apply the corresponding execution unit mask to the computing unit 300 in the computing cluster through the management unit 213 to optimize the operation of various computing tasks. It can be understood that the number and combination of computing units 300 that need to be customized by the execution unit mask in each cluster can vary dynamically according to the requirements of the computing tasks and there is no limitation.
[0075] Figure 4 This is an embodiment of the operation of the command queue of the present disclosure. In actual operation, the computer system 100 of the present disclosure may generate various drawing tasks due to application requirements and needs to be implemented through the graphics processing unit 108. After the software 106 of this embodiment is executed by the central processing unit 102, it can drive the graphics processing unit 108 to configure the task operations of different chips to produce a partitioning effect similar to virtualization technology. In addition, the drawing commands generated during the operation of the computer system 100 can be pre-classified and combined with a configuration method that restricts cross-chip operation to create more optimized operation combinations.
[0076] As Figure 4 shown, during the process of the central processing unit 102 processing the drawing task, a plurality of software queues can be established in the memory 400, such as the first software queue 410, the second software queue 420, the third software queue 430, and the fourth software queue 440.
[0077] The drawing commands processed by the central processing unit 102 can be classified according to types, such as the first type of commands, the second type of commands, etc. In one embodiment, the first type of commands are temporarily stored in the first software queue 410, and the second type of commands are temporarily stored in the second software queue 420, and so on. The first type of commands can be calculation commands and memory operation commands. The second type of commands can be image encoding commands and image decoding commands. It can be understood that in implementation, the classification method of each software queue is not limited to the command type, and it may also be other characteristics. For example, the basis for the central processing unit 102 to classify the drawing commands includes one or more of the following: task identification code, resource requirements, command type. The task identification code can be used to identify the tenant, the parent program, or the execution purpose, so as to distinguish the priority level of resource allocation. The command type includes numerical operation commands, direct memory access commands, and image commands. The resource requirements include memory requirements, the number of computing units required, and the computing priority level.
[0078] To process each software queue in the memory 400, the first chip 110 and the second chip 120 then have to bind the software queue to the command queue in each user pipeline. For the convenience of description, the operation in each command processor is represented in units of user pipelines. For example, the first user pipeline 211 in the first chip 110 is bound to the first software queue 410 in the memory 400 to read and temporarily store the drawing commands therein after being set by the management unit 213 during operation. Similarly, the second user pipeline 212, the third user pipeline 221, and the fourth user pipeline 222 are respectively bound to different software queues in the memory 400. It can be noted that the binding relationship between the software queue and the user pipeline is not processed sequentially, but can be dynamically determined according to the characteristics of each software queue, the chip position corresponding to each user pipeline, the available computing resources in each chip, etc. If the management unit 213 determines that the available computing units 300 in the first chip 110 are suitable for processing the drawing commands in the third software queue 430, then the second user pipeline 212 is bound to the third software queue 430. To achieve the judgment of this binding condition, the central processing unit 102 can cooperate with the software 106 to perform a handshake protocol with the management unit 213 to match the attributes of various software queues with the available hardware resource conditions in the graphics processing unit 108 at any time. For example, in implementation, the central processing unit 102 can send task type binding information to the management unit 213 according to the task type requirements during operation. The management unit 213 can also be driven by the firmware 215 to change the software queue binding settings of each user pipeline after receiving the task type binding information.
[0079] In short, the management unit 213 negotiates with the central processing unit 102 to bind the first user pipeline 211 to the first software queue 410. Thereby, the drawing commands in the first software queue 410 will ultimately generate computing tasks through the first user pipeline 211 and be executed in the first chip 110. That is to say, after the binding relationship is determined, the drawing commands in the first software queue 410 will not be executed by the second chip 120 beyond the scope of the first chip 110. By analogy, after the second software queue 420 is bound to the third user pipeline 221, it will only be processed by the second chip 120 and not by the first chip 110. As for the third software queue 430 being bound to the second user pipeline 212, or the fourth software queue 440 being bound to the fourth user pipeline 222, they are all based on the same principle and will not be repeated here. This is a virtualization partitioning effect achieved through binding. The management unit 213 can instantaneously allocate the computing resources of the graphics processing unit 108 effectively according to the operating conditions of each computing unit 300 in the graphics processing unit 108 without allowing each computing task to cross chips for operation.
[0080] After the first user pipeline 211 generates a computing task according to the drawing command in the first software queue 410, the computing task will be dispatched to the computing partitions in the first chip 110. In this embodiment, there are two computing partitions in the first chip 110. Therefore, the dispatching principle can be Round Robin, that is, use them in turn. And the initially dispatched partition can be the first one starting from the beginning, or the second one calculated from the middle. In a further embodiment, the management unit 213 can first judge which partition contains sufficient available computing clusters. For example, the management unit 213 judges according to the characteristics of the computing task that the number of available computing units 300 in the second computing partition 320 is sufficient to form a computing cluster to process the task, and then commands the first user pipeline 211 to dispatch the computing task to the second computing partition 320 and sets a sufficient number of computing units 300 in the second computing partition 320 to form the corresponding computing cluster to undertake the computing task. By analogy, after the second user pipeline 212 generates a computing task according to the third software queue 430 and finds that the available computing units 300 in the first computing partition 310 are sufficient to process the computing task, it dispatches the computing task to the computing cluster in the first computing partition 310 for processing. As for the third user pipeline 221 and the fourth user pipeline 222 being dispatched to the third computing partition 330 and the fourth computing partition 340 respectively, they are all based on the same principle and will not be repeated here.
[0081] It should be understood that in this embodiment, it is not limited that the computing clusters in each computing partition are statically set or dynamically formed, nor is it limited to the number of computing clusters in each chip or the number of computing units 300 in each computing cluster in the embodiment. The core idea of this embodiment is that when the first command processor 210 and the second command processor 220 dispatch computing tasks, they only dispatch tasks within the same chip and do not dispatch tasks across chips.
[0082] It can be understood that in implementation, more different types of software queues can be derived in the memory 400, and through the protocol configuration between the central processing unit 102 and the management unit 213, each user pipeline is correspondingly bound to the software queues in the memory 400. The software queues can be established in advance before the central processing unit 102 transmits the drawing commands to the graphics processing unit. For example, the central processing unit 102 can establish various software queues in the memory 400 by classifying the drawing commands. The basis for classifying the drawing commands includes one or more of the following: task identification code, resource requirements, and command type. The task identification code generally refers to the code that can be used to identify the tenant, the parent program, or the execution purpose, so as to distinguish the priority levels of resource allocation. The command types include numerical operation commands, direct memory access commands, and image commands. The computing resource requirements include memory requirements, the number of computing unit requirements, and the computing priority levels, etc. In other words, the first command processor 210 and the second command processor 220 are elastic architectures that can be dynamically scheduled according to task characteristics and resource requirements.
[0083] Figure 5 This is the flowchart of the graphics command processing method of the present disclosure. Based on Figures 1 to 4 the system architecture of, the drawing command processing method of the present disclosure can be summarized as the following process.
[0084] In step 501, the drawing processor is driven so that the command processors on each chip only call the computing resources on the same chip.
[0085] This step is performed in the computer system 100 such as Figure 1 . The main purpose is to enable each chip in the graphics processing unit 108 to maintain the same-chip operation starting from when the command processor binds the software queue. This approach can completely avoid the data transmission burden of cross-chip operations. There are several ways to set up the first command processor 210 and the second command processor 220. For example, the central processing unit 102 can execute the software 106 and transmit relevant setting commands to the graphics processing unit 108. After being processed by the graphics processing unit 108, the setting commands can restrict the first command processor 210 in the first chip 110 and the second command processor 220 in the second chip 120 from assigning computing tasks across chips.
[0086] In one embodiment, this setup command may be a core mask. After the central processing unit 102 generates the core mask, the first command processor and the second command processor are initialized. After the first command processor performs the initialization, a masking effect is produced at the dispatch address, so that computing tasks are only dispatched to the computing units in the first chip. Based on the same principle, after the second command processor performs the initialization, computing tasks are only dispatched to the computing units in the second chip.
[0087] In a further embodiment, the management unit 213 in the first command processor 210 may receive the setup command from the central processing unit 102 and change the types of commands that can be processed by the first user pipeline 211 and the second user pipeline 212 accordingly. The management unit 213 may include a firmware 215 for driving the operation of the management unit 213. The design of the firmware 215 enables the operation mode of the management unit 213 to have the flexibility of dynamic adjustment, for example, controlling the computing task dispatching functions of the first command processor 210 and the second command processor 220 according to external setup commands.
[0088] In step 503, the management unit 213 obtains the workload status of each computing cluster in each chip.
[0089] The management unit 213 in the first chip 110 can implement a monitoring function to obtain the workload data of the first partition, the second partition, the third partition, and the fourth partition. More precisely, each partition contains a plurality of computing units 300. In a derivative embodiment, the management unit 213 can obtain the workload data of each computing unit 300 for fine-grained scheduling.
[0090] In step 505, the management unit 213 schedules the binding relationship between the software queue and the command queue according to the workload status. The management unit 213 can communicate with the central processing unit 102 in real time, so that a plurality of command queues 201 in the first chip 110 are correspondingly bound to a plurality of software queues in the memory 400 according to the workload data. Thereby, the command queues 201 in each user pipeline start to extract drawing commands from the bound software queues and process them.
[0091] In step 507, after the first command processor and the second command processor translate the commands in the command queue into computing tasks, they are dispatched to the available computing partitions.
[0092] In one embodiment, the first command processor decodes a first computing task from a first command queue in the first chip, and dynamically forms one or more computing clusters in the first computing partition or the second computing partition according to the task nature of the first computing task, for dispatching and executing the first computing task. Based on the same principle, the second command processor decodes a second computing task from a second command queue in the second chip, and dynamically forms one or more computing clusters in the third computing partition or the fourth computing partition according to the task nature of the second computing task, for dispatching and executing the second computing task.
[0093] On the other hand, the computing clusters in the computing partitions can also be pre-configured static values, and the computing resources required for the computing tasks can be represented by the number of computing clusters. This is a concept of group based scheduling. For example, the first computing task may require five clusters, while the second computing task requires ten clusters. The task nature includes task type, resource requirements, tenant name, or priority level. In general, the task nature can be automatically determined when the central processing unit 102 processes the drawing commands, or be user-customized values.
[0094] The feature of the present disclosure is to use the method of dynamically binding the command processor and the software queue to achieve data partitioning similar to virtualization technology without sacrificing the hardware performance. The command processor in each chip can be implemented in multiple user pipelines, and each user pipeline has the ability to decode, schedule, and dispatch various drawing commands. The embodiments of the present disclosure provide improved technical features in terms of software 106 and firmware 215, enabling the computing unit 300 to be configured through an execution unit mask, and each command processor can be restricted from cross-chip calls through a core mask. In terms of hardware improvement, a workload monitoring function can be implemented in the command processor to facilitate the immediate determination of the dispatching requirements of the computing tasks.
[0095] In a further derived embodiment, under the cooperation of the software 106, the central processing unit 102 can, through the scheduling of the queue, deliberately send the image encoding command and the image decoding command to the computing units 300 in different chips for processing to balance the burden on the chips. For example, the central processing unit 102 can classify the image encoding command and the image decoding command into different software queues, and then enable the user pipelines in different chips to respectively bind the two software queues. Thus, the computing tasks corresponding to the image encoding command and the image decoding command will be executed in different chips.
[0096] In a further derived embodiment, Figure 4Each software queue in the [system] does not necessarily have to be implemented only in the memory 400 of the computer system 100. The computer system 100 can be configured with additional hardware caches, solid-state drives, or utilize the remaining register space in the central processing unit 102 to implement the software queue. On the other hand, the graphics processing unit 108 itself usually has a paired graphics display card memory (not shown), which can also be utilized to implement the software queue.
[0097] In one embodiment, the first chip 201 can be understood as the so-called master die, and the second chip can be understood as the slave die. In the spirit of the distributed command type, the number of chips in the graphics processing unit 108 is not limited to 2.
[0098] One of the advantages of the above embodiment is that it can improve processing efficiency and reduce system latency. Through the condition setting of the command processor, any computing task will only be dispatched to the same chip for processing, avoiding the burden of cross-chip data exchange, improving processing efficiency, and reducing system latency.
[0099] Another advantage of the above embodiment is that it can optimize resource utilization. The present disclosure classifies the drawing commands and establishes them as multiple software queues, and then binds them to the command queues that match the conditions, so that the computing cluster resources in all chips can be effectively utilized.
[0100] Another advantage of the above embodiment is that it can increase the flexibility and scalability of the system. The computer system can dynamically bind the command queues according to the type requirements of the graphics operation, and configure the size and quantity of the clusters in each computing partition, and flexibly adapt to different application requirements without generating the burden of cross-chip data exchange.
[0101] In summary, the multi-chip graphics processor command processor system of the present disclosure not only improves performance and efficiency by effectively allocating and managing resources, but also enhances the stability and scalability of the system, and is very suitable for virtualized computing applications that require high performance and multi-task processing.
[0102] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements that are not explicitly listed, or further includes elements that are inherent to such process, method, article, or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element.
[0103] The embodiments of the present disclosure have been described above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present disclosure, those of ordinary skill in the art can also make many forms without departing from the purpose of the present disclosure and the scope protected by the claims, and all of them fall within the protection scope of the present disclosure.
Claims
1. A computer system, including a memory, a central processing unit, and software, the computer system further includes a graphics processing unit, connecting the memory and the central processing unit, wherein, The graphics processor includes: A first chip including a first command processor; A second chip including a second command processor; A plurality of computing units disposed in a first computing partition and a second computing partition in the first chip, and a third computing partition and a fourth computing partition in the second chip; wherein: The first command processor and the second command processor are configured to process drawing commands provided by the computer system to dispatch computing tasks to the computing units; Wherein, the computer system is characterized in that: The central processing unit drives the graphics processor through the software according to the requirements of the computer system for the drawing task, so that the first command processor only dispatches the computing tasks to the computing units in the first chip, and the second command processor only dispatches the computing tasks to the computing units in the second chip.
2. The computer system according to claim 1, wherein Wherein: The central processing unit classifies the drawing commands and sets them into a plurality of software queues; The basis for the central processing unit to classify the drawing commands includes one or more of the following: task identification code, resource requirements, and command type; The task identification code is used to identify the tenant, the parent program, or the execution purpose, so as to distinguish the priority levels of resource allocation; The command types include numerical operation commands, direct memory access commands, and image commands; and The resource requirements include memory requirements, the number of computing unit requirements, and the computing priority level.
3. The computer system according to claim 2, wherein Wherein: Each of the first command processor and the second command processor includes a plurality of command queues, which are configured to extract the drawing commands from the bound software queues and process them; The central processing unit obtains the workload data of the first computing partition, the second computing partition, the third computing partition, and the fourth computing partition by communicating with the graphics processor; The central processing unit controls the first command processor and the second command processor to bind the plurality of command queues to the plurality of software queues correspondingly according to the workload data by running the software.
4. The computer system according to claim 3, characterized in that, Wherein: The plurality of computing units in the first computing partition, the second computing partition, the third computing partition, and the fourth computing partition further form a plurality of computing clusters respectively; The first command processor decodes a first computing task from the first command queue in the first chip, and dispatches the first computing task to one or more computing clusters located in the first computing partition or the second computing partition according to the task nature of the first computing task; The second command processor decodes a second computing task from the second command queue in the second chip, and dispatches the second computing task to one or more computing clusters located in the third computing partition or the fourth computing partition according to the task nature of the second computing task; and The task nature includes task type, resource requirements, tenant name, or priority level.
5. The computer system according to claim 4, characterized in that, Wherein: The first command processor further includes a first management unit, coupled to the central processing unit, for managing the scheduling and binding of the queues; The second command processor further includes a second management unit, coupled to the central processor, having the same functional structure as the first management unit; When the first management unit is activated and running, the second management unit is in a disabled state, and the operations of the first command processor and the second command processor are scheduled and managed by the first management unit.
6. The computer system according to claim 5, characterized in that, Wherein: After the central processor executes the software, a core mask is generated to initialize the operations of the first command processor and the second command processor; After the first command processor performs the initialization operation, it only assigns the computing tasks to the computing units in the first chip; And After the second command processor performs the initialization operation, it only assigns the computing tasks to the computing units in the second chip.
7. The computer system according to claim 6, characterized in that, Wherein: The first management unit includes firmware for driving the first management unit to communicate with the central processor and maintaining the dynamic binding relationship between the multiple software queues and the multiple command queues.
8. A graphics processor, characterized in that, Set up and operate in a computer system including a memory, a central processor, and software, including: A first chip including a first command processor; A second chip including a second command processor; and Multiple computing units are set in the first computing partition and the second computing partition in the first chip, and the third computing partition and the fourth computing partition in the second chip; wherein: The first command processor and the second command processor are configured to process the drawing commands provided by the computer system to assign computing tasks to the computing units; The graphics processor is characterized in that: The graphics processor is driven by the central processor according to the requirements of the computer system for drawing tasks through the software, so that the first command processor only assigns the computing tasks to the computing units in the first chip, and the second command processor only assigns the computing tasks to the computing units in the second chip.
9. The graphics processor according to claim 8, wherein Wherein: Before the drawing commands are transmitted to the graphics processor, they are classified by the central processor and set as multiple software queues; The basis for classifying the drawing commands includes one or more of the following: task identification code, resource requirements, and command type; The task identification code is used to identify the tenant, the parent program, or the execution purpose in order to distinguish the priority levels of resource allocation; The command types include numerical operation commands, direct memory access commands, and image commands; and The resource requirements include memory requirements, the number of computing unit requirements, and the computing priority levels.
10. The graphics processor according to claim 9, wherein Wherein: Each of the first command processor and the second command processor includes multiple command queues, configured to extract and process the drawing commands from the bound software queues; The graphics processor provides the workload data of the first computing partition, the second computing partition, the third computing partition, and the fourth computing partition to the central processor; And The first command processor and the second command processor are controlled by the central processor to bind the multiple command queues to the multiple software queues correspondingly according to the workload data.
11. The graphics processor according to claim 10, characterized in that, Wherein: The first command processor decodes a first computing task from a first command queue in the first chip, and dynamically forms one or more computing clusters in the first computing partition or the second computing partition according to the task nature of the first computing task, for dispatching and executing the first computing task; The second command processor decodes a second computing task from a second command queue in the second chip, and dynamically forms one or more computing clusters in the third computing partition or the fourth computing partition according to the task nature of the second computing task, for dispatching and executing the second computing task; and The task nature includes task type, resource requirement, tenant name, or priority level.
12. The graphics processor according to claim 11, wherein Wherein: The first command processor further includes a first management unit, coupled to the central processing unit, for managing the scheduling and binding of the queue; The second command processor further includes a second management unit, coupled to the central processing unit, having the same functional structure as the first management unit; When the first management unit is activated and running, the second management unit is in a disabled state, and the operations of the first command processor and the second command processor are scheduled and managed by the first management unit.
13. The graphics processor according to claim 12, wherein, Wherein: The central processing unit generates a core mask after executing the software to initialize the operations of the first command processor and the second command processor; After the first command processor performs the initialization operation, it only dispatches the computing task to the computing units in the first chip; And After the second command processor performs the initialization operation, it only dispatches the computing task to the computing units in the second chip.
14. The graphics processor according to claim 13, wherein Wherein: The first management unit includes firmware for driving the first management unit to communicate with the central processing unit and maintaining the dynamic binding relationship between the multiple software queues and the multiple command queues.
15. A drawing command processing method operates in a computer system including a memory, a central processing unit, a graphics processing unit, and software, wherein, The graphics processing unit at least includes a first chip, a second chip, and multiple computing units. The first chip includes a first command processor, the second chip includes a second command processor, and the multiple computing units are disposed in a first computing partition and a second computing partition in the first chip, and a third computing partition and a fourth computing partition in the second chip; Characterized in that, the drawing command processing method includes: According to the requirements of the computer system for drawing tasks, setting the first command processor and the second command processor to process the drawing commands provided by the computer system to dispatch computing tasks to the computing units; Setting the first command processor to only dispatch the computing tasks to the computing units in the first chip; and Setting the second command processor to only dispatch the computing tasks to the computing units in the second chip.
16. The drawing command processing method according to claim 15, wherein, Further includes: Before transmitting the drawing commands to the graphics processing unit, classifying the drawing commands to establish multiple software queues; wherein: The basis for classifying the drawing commands includes one or more of the following: task identification code, resource requirement, and command type; The task identification code is used to identify the tenant, the parent program, or the execution purpose, so as to distinguish the priority levels of resource allocation; The command types include numerical operation commands, direct memory access commands, and image commands; and The resource requirements include memory requirements, the number of computing units required, and the computing priority levels.
17. The drawing command processing method according to claim 16, wherein each of the first command processor and the second command processor includes a plurality of command queues, characterized in that, The drawing command processing method further includes: Obtaining the workload data of the first computing partition, the second computing partition, the third computing partition, and the fourth computing partition; and Correspondingly binding the multiple command queues to the multiple software queues according to the workload data, so that the multiple command queues are set to extract and process drawing commands from the bound software queues.
18. The drawing command processing method according to claim 17, characterized in that, It further includes: The first command processor decodes a first computing task from the first command queue in the first chip, and dynamically forms one or more computing clusters in the first computing partition or the second computing partition according to the task nature of the first computing task for dispatching and executing the first computing task; And The second command processor decodes a second computing task from the second command queue in the second chip, and dynamically forms one or more computing clusters in the third computing partition or the fourth computing partition according to the task nature of the second computing task for dispatching and executing the second computing task; where: The task nature includes task type, resource requirements, tenant name, or priority level.
19. The drawing command processing method according to claim 18, characterized in that, It further includes: Generating a core mask to initialize the operation of the first command processor and the second command processor; After the first command processor performs the initialization operation, it only dispatches the computing tasks to the computing units in the first chip; And After the second command processor performs the initialization operation, it only dispatches the computing tasks to the computing units in the second chip.
Citation Information
Patent Citations
Dynamic dispatch for workgroup distribution
US20230206382A1