Task execution method, graphics processor, device, medium and product

By using resource pre-allocation instructions in the graphics processor, the resource requirements of a task are estimated and computing unit resources are reserved in advance, thus solving the problem of low task execution efficiency in the prior art and improving data throughput and resource utilization.

CN121900898APending Publication Date: 2026-04-21MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MOORE THREADS TECH CO LTD
Filing Date
2025-12-25
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, graphics processors perform task packaging, resource status query, and resource request sequentially during task execution, resulting in low data throughput, high latency, uneven load on computing units, and low utilization of computing resources due to the need to wait for feedback during the resource request process.

Method used

By using resource pre-request instructions, the amount of resources required for a task can be estimated and the resource information of computing units can be obtained. Target computing units can be identified in advance and resources can be pre-occupied, avoiding waiting for resource status feedback and request feedback, thereby improving resource utilization and task dispatch efficiency.

Benefits of technology

It reduces task execution time overhead, improves the utilization of computing resources and data throughput of computing units, and achieves faster task dispatch and load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121900898A_ABST
    Figure CN121900898A_ABST
Patent Text Reader

Abstract

The invention discloses a task execution method, a graphics processor, equipment, a medium and a product, and relates to the technical field of chips. The method comprises the steps that in response to a received resource pre-application instruction, resource information corresponding to at least two computing units is obtained, and the resource information is used for indicating the computing resource quantity capable of being provided by the computing units; based on a first computing resource applied by the resource pre-application instruction, a resource pre-occupation instruction is sent to a target computing unit in the at least two computing units, and the first computing resource is a computing resource required when the first task is executed and is obtained based on estimation of the first task; in response to the received first task, the first task is distributed to a target computing unit, and the target computing unit is used for executing the first task through a pre-occupied first computing resource. Calculation resources are occupied in advance based on the resource pre-application instruction, the delay of task scheduling improvement by applying for the calculation resources is reduced, and therefore the task execution efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chip technology, and in particular to a task execution method, graphics processor, device, medium, and product. Background Technology

[0002] A Graphics Processing Unit (GPU) is a chip that uses large-scale computing units to process a large number of computational or graphics rendering tasks in parallel. During the execution of computational or graphics rendering tasks, the GPU needs to generate task packages corresponding to the tasks and dispatch the corresponding tasks based on the state of the computing units.

[0003] In related technologies, based on the types and quantities of resources required by the task instances in the task package corresponding to the first task, resource information query requests are sent sequentially to the computing unit. Based on the resource information returned by the computing unit, the target computing unit for executing the first task is determined. A resource request is sent to the target computing unit to occupy the computing resources in the target computing unit used for executing the first task. After receiving feedback from the target computing unit confirming successful resource occupation, the first task is dispatched to the target computing unit.

[0004] However, in the above task execution method, task packaging, resource status query and resource request are executed sequentially and require waiting for feedback from the computing unit, resulting in low data throughput and high latency. Summary of the Invention

[0005] This application provides a task execution method, a graphics processor, a device, a medium, and a product. The technical solutions provided by this application include the following aspects.

[0006] According to one aspect of the embodiments of this application, a task execution method is provided, the method comprising: In response to receiving a resource pre-request instruction, resource information corresponding to at least two computing units is obtained, wherein the resource information is used to indicate the amount of computing resources that the computing units can provide. Based on the first computing resource requested by the resource pre-request instruction, a resource pre-occupancy instruction is sent to the target computing unit among the at least two computing units; the resource pre-occupancy instruction is used to pre-occupy the first computing resource before dispatching the first task; the first computing resource is the computing resource required for the execution of the first task, and is estimated based on the first task; In response to receiving the first task, the first task is dispatched to the target computing unit, which executes the first task using the first computing resources it has pre-occupied.

[0007] According to one aspect of the embodiments of this application, a graphics processor is provided, the graphics processor comprising: A scheduler, in response to receiving a resource pre-request instruction, acquires resource information corresponding to at least two computing units respectively; the resource information is used to indicate the amount of computing resources that the computing units can provide; based on the first computing resources requested by the resource pre-request instruction, sends a resource pre-occupancy instruction to a target computing unit among the at least two computing units; the resource pre-occupancy instruction is used to pre-occupy the first computing resources before dispatching the first task; the first computing resources are the computing resources required for the execution of the first task, and are estimated based on the first task. A dispatcher is configured to dispatch the first task to the target computing unit in response to receiving the first task; the target computing unit is configured to execute the first task using the first computing resources pre-occupied.

[0008] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a graphics processor and a memory; The memory is used to store computer programs; The graphics processor is used to load the computer program to execute the above-described task execution method.

[0009] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the above-described task execution method.

[0010] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium, and a processor reading from the computer-readable storage medium and executing the computer program to implement the above-described task execution method.

[0011] The technical solution provided in this application can bring the following beneficial effects: Based on the resource pre-request instruction, the amount of resources required for the task is estimated and the resource information corresponding to the computing unit is obtained. The target computing unit for executing the task is then determined. On the one hand, the resource information corresponding to the computing unit is obtained in advance before receiving the first task, thereby avoiding waiting for resource status feedback during task scheduling. On the other hand, the resource allocation to the target computing unit is executed in advance or in parallel with receiving the task, avoiding waiting for resource request feedback. This allows for faster task dispatch, reduces task execution time overhead, and improves the utilization rate of computing resources in the computing unit, as well as the data throughput during task execution. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the structure of a computer system provided in an exemplary embodiment of this application; Figure 2 This is a schematic diagram of a task execution flow provided in an exemplary embodiment of this application; Figure 3 This is a flowchart of a task execution method provided in another embodiment of this application; Figure 4 This is a flowchart of a task execution method provided in another embodiment of this application; Figure 5 This is a schematic diagram illustrating the number of examples in the task packages of different tasks provided in an exemplary embodiment of this application; Figure 6 This is a flowchart of a task execution method provided in another embodiment of this application; Figure 7 This is a schematic diagram of a task processing flow provided in an exemplary embodiment of this application; Figure 8 This is a schematic diagram of the structure of a task processing system provided in an exemplary embodiment of this application; Figure 9 This is a structural block diagram of a graphics processor provided in an exemplary embodiment of this application; Figure 10 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application. Detailed Implementation

[0013] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0014] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0016] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user, processor, and computer device data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0018] Figure 1 This is a schematic diagram of the structure of a computer system provided in an exemplary embodiment of this application. The computer system 100 can implement a system architecture that serves as a task execution method. The computer system 100 includes: a computer device 120.

[0019] In some embodiments, the computer device 120 includes a graphics processor and a memory. The memory stores a computer program, and the graphics processor loads the computer program to execute the task execution method provided in this application embodiment.

[0020] In an optional embodiment, the computer device 120 further includes a central processing unit (CPU) for generating resource pre-request instructions corresponding to the first task and generating task execution instructions corresponding to the first task. Optionally, the CPU in the computer device 120 includes a graphics processor driver. Based on the graphics processor driver, the CPU obtains the amount of resources required by a reference task, which is used to indicate a threshold for the amount of resources required by the first task. Based on the amount of resources required by the reference task, the CPU determines the amount of resources for the first computing resource, wherein the amount of resources for the first computing resource is greater than or equal to the actual resource requirement of the first task, thereby generating a resource pre-request instruction.

[0021] The computer device 120 can be an electronic device such as a mobile phone, tablet computer, vehicle terminal (vehicle system), wearable device, personal computer (PC), or a server, such as a physical server or cloud server. This application embodiment does not limit this.

[0022] The task execution method provided in this application embodiment can be executed by a graphics processor in a computer device 120. In some embodiments, the computer device includes hardware devices that need to perform task execution, such as GPUs, multi-core CPUs, Tensor Processing Units (TPUs), and network processors.

[0023] Taking a GPU as an example, in the complete process of GPU processing computing or graphics rendering tasks, the GPU first needs to receive the software instructions corresponding to the first task sent by the GPU driver, and parse the software instructions to obtain the hardware-executable first task. Indicatively, the first task is implemented in the form of a thread bundle or thread block. The resource requirements corresponding to the first task are obtained, including the types and quantities of resources required by the task instance in the task package corresponding to the first task. A resource information query request is sent to the computing unit to obtain the resource information returned by the computing unit. Based on the resource requirements corresponding to the first task and the resource information corresponding to each computing unit, a target computing unit is determined, which is used to execute the first task. A resource request is sent to the target computing unit to occupy the computing resources used to execute the first task. After receiving feedback from the target computing unit confirming successful resource occupation, the first task is dispatched to the target computing unit. The target computing unit executes the relevant task instructions based on the received first task and outputs the execution results to the corresponding storage space.

[0024] For illustrative purposes, please refer to the following: Figure 2 , Figure 2 This is a schematic diagram of a task execution flow provided in an exemplary embodiment of this application, such as... Figure 2 As shown, this task execution method is applied to a graphics processing unit (GPU) in a computer device. The computer device includes a driver unit and a GPU, wherein the GPU includes a command stream processing unit, a packaging unit, a scheduling unit, a dispatching unit, a storage unit, and a computing unit. During task execution, the driver unit first executes step 210, sending a first task execution instruction corresponding to the first task to the command stream processing unit. Based on the first task execution instruction, the command stream processing unit executes step 220, sending the acquired task configuration information corresponding to the first task to the packaging unit. Based on the task configuration information, the packaging unit executes step 230, generating the first task. The first task is implemented as a thread bundle or thread block. After generating the first task, the packaging unit obtains the resource requirements corresponding to the first task and executes step 240, sending the first task and its corresponding resource requirements to the scheduling unit. These resource requirements include the types and quantities of computing resources required by the first task.

[0025] like Figure 2 As shown, based on the resource requirements corresponding to the first task, the scheduling unit executes step 250, sending a resource information query request to the computing unit. It waits for feedback from the computing unit. If the resource information for each computing unit cannot meet the resource requirements of the first task, step 250 is repeated. If the computing unit provides the corresponding resource information, the scheduling unit selects a target computing unit based on the resource information and executes step 260, sending a resource request to the target computing unit. The target computing unit is selected by the scheduling unit from among the computing units, and its resource information meets the resource requirements of the first task. It waits for successful feedback from the target computing unit regarding the resource request. If the target computing unit's resource request fails, step 260 is repeated.

[0026] like Figure 2 As shown, upon successful feedback from the target computing unit regarding its resource request, the dispatch unit receives the first task and relevant information about the target computing unit from the scheduling unit. This relevant information includes the storage space corresponding to the target computing unit. The dispatch unit executes step 270, sending a preload request to the storage unit, and step 280, dispatching the first task to the target computing unit. The preload request instructs the storage unit to perform a preload operation on the storage space corresponding to the target computing unit.

[0027] However, with an increase in the number of computing units and improved task processing efficiency, or with an increase in the amount of tasks to be executed, the requirements for task execution efficiency increase. The above task generation scheme has at least the following problems: 1. Resource requests occur after the first task is received. That is, only after obtaining the resource requirements corresponding to the first task can resource information query requests and resource request requests be sent to the computing units sequentially. Therefore, tasks cannot be dispatched to the computing units after receiving the task but before the resource request is successful. The computing units cannot continuously process tasks in a saturated state, i.e., the load is unbalanced, which reduces task processing performance. 2. The resource request process requires sending at least two requests to the computing units and waiting for feedback from the computing units. This has a large time overhead. The task generation rate cannot match the task processing rate, resulting in idle computing units, low bandwidth utilization, and limiting task execution efficiency.

[0028] To address the aforementioned issues, this application provides a task execution method. This method uses a resource pre-request instruction to estimate the resource requirements of a task and obtain resource information corresponding to the computing unit, thereby determining the target computing unit for task execution. Before receiving the task, resource allocation is performed on the target computing unit. Please refer to... Figure 3 This illustrates a flowchart of a task execution method provided in another embodiment of this application. The method is performed by a computer device (which can be implemented as follows). Figure 1 The method is executed by the computer device 120 shown. The method includes the following steps.

[0029] Step 310: In response to receiving the resource pre-request instruction, obtain the resource information corresponding to at least two computing units respectively.

[0030] The resource pre-request instruction is used to request the first computing resources required to execute the first task. The first computing resources are those needed for the execution of the first task and are estimated based on the first task. In some embodiments, the resource pre-request instruction is generated by the driver and sent to the scheduling module. It instructs the scheduling module to request the first computing resources required for the first task from the computing unit before the driver sends the first task execution instruction corresponding to the first task to the scheduling module. Accordingly, the first computing resources are determined by the type and scale of the first task. The task scale indicates the total amount of data, computation, or events required to complete the first task, and can be represented in ways such as the number of data elements, the number of floating-point operations, or the number of instructions. Indicatively, the driver is a GPU driver.

[0031] Taking a GPU as an example, the types of the first task include, but are not limited to: 1. Graphics rendering tasks: These are used to execute specific stages predefined in the rendering pipeline, each stage having a clear functional definition and input / output configuration. Illustratively, graphics rendering tasks include vertex shading, tessellation, geometry shading, rasterization, pixel shading, etc. 2. General-purpose computing tasks: These are used to execute kernel computing models or functions defined by the administrator, featuring large-scale, parallel execution. Illustratively, general-purpose computing tasks include matrix operations, physics simulation, image or signal processing, data analysis, and sorting. In some embodiments, by updating the corresponding code snippet of the driver, the driver can pre-generate resource pre-request instructions based on the task identifier before the first task is generated or sent to the scheduling module. The aforementioned resource pre-request instructions may include at least one task description information from the type and scale of the first task.

[0032] In some embodiments, the description dimensions of the first computing resource include, but are not limited to, the following features: 1. Resource Type: This refers to the classification of resources required to execute a task based on their functional attributes; that is, different resource types provide different capabilities during task execution. Taking the computing resources corresponding to a task in a GPU as an example, resource types include, but are not limited to: 1.1 Computation type resources are used to provide the ability to perform arithmetic and logical operations during task execution. For example, the computation resources corresponding to computation units such as basic computation units, matrix computation units, and dedicated ray tracing computation units are computation types.

[0033] 1.2 Storage type resources are used to indicate the space for storing instructions and data, and to provide data retention and migration capabilities during task execution. For example, registers for storing local variables are thread-private and fast; shared memory for inter-thread communication is shared by threads and has low latency; caches, including Level 1 (L1) caches, constant caches, texture caches, etc., are used to temporarily store frequently used data loaded from global memory and other storage spaces, and are high-speed and small-capacity; global memory, accessible to all threads, is large in capacity but has low transfer efficiency; etc.

[0034] 1.3 Control and scheduling type resources are used to provide the ability to manage the execution order and control the timing of task execution during task execution. For example, a GPU computing core, also known as a computing unit cluster, includes a thread bundle scheduler, register file, shared memory, etc.; a thread bundle / block scheduling slot is used to indicate the ability of a GPU computing core to manage thread bundles / blocks simultaneously; etc.

[0035] It is worth noting that the different types of computing resources mentioned above are based on the specific capabilities provided during task execution. In other words, the same hardware unit may be implemented as different resource types in different task execution processes.

[0036] 2. Resource Space Size: The resource space size is determined in units of threads and thread bundles / blocks. Illustratively, during the compilation of a GPU device, based on the variable types, array sizes, and control flow declared in the kernel code, the compiler calculates the resource requirements of each thread and thread bundle / block for each storage type and generates corresponding metadata to instruct the driver runtime to match the resource requirements of the corresponding task.

[0037] 3. Immediate Availability: Indicates that within the current scheduling period, the resource requirements of a task do not exceed the amount of idle, immediately allocable resources on the computing unit.

[0038] In some embodiments, the resource pre-request instruction generated by the driver is implemented as a software instruction. Before the resource pre-request instruction is sent to the scheduling module, it first needs to be parsed and converted into a low-level instruction form that can be executed by the hardware. Illustratively, the execution entity that performs the parsing and processing of the resource pre-request instruction is a command stream processor, which is used to convert the software instruction form of the binary command stream into a hardware instruction form. The binary command stream contains instructions describing the task scale and resource requirement information; converting the software instruction form into the hardware instruction form, i.e., parsing the binary command stream, yields a series of control signals used to trigger the scheduling module and subsequent computing units to perform operations. Based on the conversion of the resource pre-request instruction implementation into a hardware instruction form, the resource pre-request instruction is sent to the scheduling module.

[0039] In the embodiments provided in this application, the driver receives a first task execution request from an application and generates a task execution instruction corresponding to the first task. When the driver generates the task execution instruction corresponding to the first task, the driver generates a resource pre-request instruction corresponding to the first task and sends the resource pre-request instruction to the scheduling module.

[0040] In response to receiving a resource pre-request instruction, resource information corresponding to at least two computing units is obtained. This resource information indicates the amount of computing resources that each computing unit can provide. In some embodiments, the resource information includes real-time data on the amount of computing resources currently available to each computing unit. For different resource types, the resource information includes, but is not limited to: 1. For computing type resources, the resource information includes consumed or idle thread bundles / block slots, i.e., the number of thread bundles / blocks that are occupied or can be accommodated. 2. For storage type resources, the resource information includes the number of remaining registers, the size of remaining shared memory, etc.

[0041] The scope of "at least two computing units" includes, but is not limited to, one of the following two cases: at least two computing units are maintained by the scheduling module and are idle computing units; or at least two computing units are all computing units in the computer device. Optionally, the computer device may be a GPU core.

[0042] After receiving a resource pre-request instruction, the scheduling module can obtain resource information corresponding to at least two computing units in ways including but not limited to: 1. The scheduling module sends resource information retrieval requests to at least two computing units. In response, each computing unit queries its current resource usage status and other resource information, and then feeds this real-time resource information back to the scheduling module. The scheduling module then obtains the resource information corresponding to the computing unit.

[0043] 2. The scheduling module internally maintains a resource information table to store resource information corresponding to computing units. This resource information table can be implemented as a data structure, including but not limited to linked lists, circular arrays, bitmaps, etc. Computing units periodically, or upon task completion, report their current resource information to the scheduling module, which then updates the resource information table accordingly. When the scheduling module receives a resource pre-request instruction, it queries the resource information table to obtain the resource information of the corresponding computing unit.

[0044] 3. The scheduling module sets a resource pool consumption counter to record the occupancy status of computing units. Indicatively, when a computing unit has not received a task, all thread bundles / block slots are idle, and the counter value is set to 0. During task execution, the counter value increases accordingly; when a computing unit completes a task, it sends a hardware completion signal to the scheduling module, causing the scheduler to automatically decrease the corresponding counter, thus updating the resource pool consumption counter. When the scheduling module receives a resource pre-request instruction, it queries the resource pool consumption counter to obtain the corresponding computing unit's count data, i.e., it retrieves resource information.

[0045] It is worth noting that the above-mentioned methods for obtaining resource information corresponding to at least two computing units are merely illustrative, and the specific methods for obtaining resource information corresponding to computing units are not limited in the embodiments of this application.

[0046] Step 320: Based on the first computing resource requested by the resource pre-request instruction, send a resource pre-occupancy instruction to the target computing unit in at least two computing units.

[0047] The resource pre-allocation instruction is used to pre-allocate a first computing resource before dispatching the first task. In some embodiments, the implementation of the resource pre-allocation instruction includes, but is not limited to: 1. Send a resource request to the target computing unit and wait for the target computing unit to send a response regarding resource occupancy.

[0048] 2. Before receiving the first task, computing resources for the computing unit are pre-allocated to ensure sufficient pre-allocated resources, forming a logical resource pool. Upon receiving the resource pre-request instruction corresponding to the first task, a secondary request is made from the pre-allocated resources. Illustratively, during kernel runtime, storage resources are pre-allocated through the driver, and the required first computing resources are reclaimed from the pre-allocated storage resources based on the resource pre-occupancy instruction, thus determining the target computing unit.

[0049] It is worth noting that the above-described method of sending resource pre-occupancy instructions to the target computing unit is merely illustrative, and the specific implementation of the resource pre-occupancy instructions in this application embodiment is not limited.

[0050] The target computing unit is one or more computing units among at least two computing units, and the resource information corresponding to the target computing unit satisfies the resource amount of the first computing resource requested by the resource pre-request instruction, either individually or collectively.

[0051] In some embodiments, the resource pre-occupancy process is implemented as follows: When the first computing resource is 1024 bytes of shared memory, based on the resource pre-occupancy instruction sent by the scheduling module to the target computing unit, the target computing unit checks the corresponding shared memory. If there is at least 1024 bytes of free space in the corresponding shared memory (i.e., sufficient resources), the target computing unit obtains the base address corresponding to the 1024 bytes of space in the corresponding shared memory and sends it to the scheduling module. After the resource pre-occupancy is completed, when the scheduling module receives the first task, it dispatches the first task to the target computing unit. When the scheduling module accesses the first computing resource, i.e., the 1024 bytes of storage space in the pre-occupied shared memory, it obtains the storage address corresponding to the first computing resource based on the received base address plus a predetermined offset.

[0052] In the embodiments provided in this application, when the scheduling module sends a resource pre-occupancy instruction to the target computing unit among at least two computing units, it has not received the task execution instruction corresponding to the first task generated by the driver. That is, the scheduling module has completed the operation of sending the resource pre-occupancy instruction to the target computing unit before receiving the first task. Before receiving the first task, the scheduling module needs to perform configuration processing on the first task, including but not limited to the following processing: 1. The driver generates the task execution instructions corresponding to the first task. Indicatively, generating the task execution instructions includes, but is not limited to: receiving the first task execution request sent by the application; allocating storage space for initialization data; configuring registers; and initializing data.

[0053] 2. The driver sends a first task execution instruction to the scheduling module. Indicatively, the first task execution instruction includes the first task and the task configuration information corresponding to the first task.

[0054] 3. Through parsing and processing, the execution instructions of the first task are converted into low-level instruction forms that the hardware can execute.

[0055] 4. Generate the task package for the first task based on the task configuration information corresponding to the first task.

[0056] 5. The scheduling module receives the task package of the first task, enabling task scheduling at the task package level.

[0057] Step 330: In response to receiving the first task, dispatch the first task to the target computing unit.

[0058] The target computing unit executes a first task using pre-occupied first computing resources. If the pre-occupied amount of the first computing resources exceeds the actual resource requirement of the first task, a resource release request is sent to the target computing unit, i.e., releasing redundant computing resources beyond the actual resource requirement from the pre-occupied first computing resources. In some embodiments, after the scheduling module dispatches the first task to the target computing unit, the target computing unit writes a first indication signal to the target address corresponding to the storage module based on the processing status of the first task. Illustratively, when the processing status of the first task is that the task execution is complete, the first indication signal is used to indicate that the execution result of the first task is in place. Based on the first indication signal, the execution result corresponding to the first task is transmitted to the corresponding storage address or discarded. The above-mentioned processing method for the execution result is determined by the content and purpose of the first task. Taking a GPU as an example, the processing method for the execution result includes, but is not limited to, the following: 1. If the first task is an intermediate stage task, the execution result is stored in the GPU's global memory. For example, if the execution result is the output of the geometry shader in the rendering pipeline or the output of a specific layer in the Artificial Intelligence (AI) model, then the execution result is used as the input data for the next task to be processed by the GPU. Therefore, the execution result needs to be stored in global memory so that the next task to be processed can obtain input data from global memory.

[0059] 2. If the first task is the final output task, the execution result is transferred to the host memory. For example, if the execution result is the final result of scientific computing, the final image of image processing, or the result of a database query, then the execution result is used as task input data or the result required by the user, and therefore needs to be stored in the host memory.

[0060] 3. If the primary task is testing, discard the execution result. For example, if the execution result is not a performance metric in a performance benchmark test, or if the execution result is a temporary variable in data preprocessing, the execution result should also be discarded.

[0061] In some embodiments, the specific processing method of the execution result of the first task does not affect the processing status of the first task, i.e., the acquisition of the first indication signal. The scheduling module acquires the first indication signal and, based on the first indication signal, acquires the resource information corresponding to the target computing unit. Optionally, the first indication signal is also used to notify that the execution result corresponding to the first task is ready and can be safely read. Schematically, the acquisition method of the first indication signal includes: the scheduling module continuously polls the status register corresponding to the target computing unit, the status register being used to indicate the idle status of the computing unit. When the corresponding status register indicates that the target computing unit changes from an "occupied" state to an "idle" state, it indicates that the task execution is complete, i.e., the first indication signal is acquired; or the target computing unit sends the first indication signal to the scheduling module.

[0062] It is worth noting that the above-described method of obtaining the first indication signal is merely illustrative, and the specific method of obtaining the first indication signal is not limited in the embodiments of this application.

[0063] In the embodiments provided in this application, upon receiving the first task, the scheduling module has completed the operation of sending a resource pre-occupancy instruction to the target computing unit. Based on the task execution instruction corresponding to the first task, the scheduling module determines that the target computing unit has completed the application for the computing resources required by the first task based on the resource pre-occupancy instruction, and dispatches the first task to the target computing unit. Optionally, when the scheduling module receives the task execution instruction corresponding to the first task, if it does not receive occupancy feedback from the target computing unit, it waits until it receives occupancy feedback; if it has received occupancy feedback, it dispatches the first task to the target computing unit based on the task execution instruction corresponding to the first task.

[0064] In summary, the method provided in this embodiment, based on resource pre-request instructions, estimates the amount of resources required for a task and obtains the resource information corresponding to the computing unit, thereby determining the target computing unit for executing the task. On the one hand, it obtains the resource information corresponding to the computing unit in advance of receiving the first task, thus avoiding waiting for resource status feedback during task scheduling. On the other hand, it executes resource allocation to the target computing unit in advance or in parallel with receiving the task, avoiding waiting for resource request feedback, thereby dispatching tasks more quickly, reducing the time overhead of task execution, improving the utilization rate of computing resources in the computing unit, and the data throughput during task execution.

[0065] Figure 4This is a flowchart of a task execution method provided in another embodiment of this application. The method is performed by a computer device (which can be implemented as follows). Figure 1 The above method is executed by the computer device 120 shown, and is further implemented as steps 410 to 440. It is worth noting that there is no time dependency between steps 410 to 420 and steps 310 to 320, that is, steps 410 to 420 can be executed after steps 310 to 320, or steps 410 to 420 can be executed in parallel with steps 310 to 320.

[0066] Step 410: In response to receiving the first task, generate the task package for the first task.

[0067] The first task is the parsed task. In some embodiments, taking the GPU executing the first task as an example, the process of the scheduling module receiving the first task includes: 1. The driver receives the first task execution request from the application and generates the first task execution instruction corresponding to the first task. The first task execution instruction includes the first task. Illustratively, the application calls the Application Programming Interface (API) in the programming framework. The API is used to manage large-scale lightweight concurrent tasks on the GPU core, including but not limited to thread-level APIs or fiber-level APIs. The pre-processing steps for the driver to generate the first task execution instruction include: 1.1 Allocating storage space for initialization data: The driver obtains a contiguous block of available space in video memory, such as a block of storage space in global memory, through the API, and sends the address of this contiguous available space to the application. Based on this address, the application writes the data it needs to process, such as an image matrix, into the storage space corresponding to that address.

[0068] 1.2 Configuration Registers: The driver configures the rules for GPU task execution by writing initial values ​​to specific registers of the GPU. These specific registers include, but are not limited to: a command queue base address register, used to store the starting physical address of the command queue in video memory. During command fetching, the GPU's instruction fetch unit locates the first command based on the value of the command queue base address register and reads the complete command in offset order. This command includes the first task execution instruction; a doorbell register, used to store a doorbell signal indicating that the GPU has started processing a task; a pipeline status register, used to maintain the status bits of each computing unit in the current GPU task processing pipeline, including idle, busy, suspended, or abnormal, etc. Based on these status bits, the driver can determine the next task operation or interrupt handling; and a page table base address register, used to indicate the initial address of the page table of the global memory management unit. The page table records the mapping relationship between virtual addresses and physical addresses, etc.

[0069] 1.3 Initialization Data: The driver writes the data required during task execution into the storage space allocated in the above steps. Depending on the task type, this data includes, but is not limited to: calculation constants, vertex data, or image textures; for general computing tasks, initialization data also includes writing kernel function parameters, such as pointers and scalar values, into specific parameter memory.

[0070] 2. The driver sends the first task execution instruction to the scheduling module. Illustratively, the driver encapsulates the compiled commands, such as the first task execution instruction, resource addresses (e.g., storage space addresses allocated for initialization data), execution parameters, etc., into a command stream and writes it to the command queue; optionally, the command queue is stored in global memory. The driver sends a doorbell signal to the GPU, such as writing data to a doorbell register, to indicate that the command stream corresponding to the first task to be executed in the command queue is ready.

[0071] 3. Before the first task execution instruction is sent to the scheduling module, it first needs to be parsed and converted into a low-level instruction form that the hardware can execute. Illustratively, the entity that performs the parsing and processing of the first task execution instruction is the command stream processor. The command stream processor parses the instructions describing the task scale and resource requirements—that is, the binary command stream—into a series of control signals used to trigger operations by the scheduling module and subsequent computing units; in other words, it converts the software instruction form into hardware instruction form. Based on the conversion of the first task execution instruction into hardware instruction form, the first task execution instruction is sent to the scheduling module.

[0072] It is worth noting that this embodiment does not limit the timing relationship between the scheduling module receiving the first task and the scheduling module requesting the first computing resource based on the resource pre-request instruction. That is, there is no dependency between the scheduling module receiving the first task and the scheduling module requesting the first computing resource based on the resource pre-request instruction.

[0073] In some embodiments, the scheduling module receives a first task execution instruction corresponding to the first task, and task configuration information corresponding to the first task. Based on the task configuration information, a task package for the first task is generated. The task configuration information is used to indicate the method of task execution, including but not limited to: 1. Resource Requirements: This indicates the first computing resources required by the first task, such as the size of shared memory and the number of registers required by the first task. Optionally, the first task can be implemented as a thread bundle / block.

[0074] 2. Execution Scale: This indicates the correspondence between the first task and the hierarchical organizational structure, which includes grids, thread bundles / blocks, and threads. In other words, it determines the position of the first task within the hierarchical organizational structure.

[0075] 3. Memory parameters: These include the specific values ​​of the function parameters corresponding to the task or pointers used to obtain the memory addresses of the specific values.

[0076] 4. Dependency: Used to indicate the temporal relationship between tasks. For example, the first task needs to wait for the completion of a preceding task before it can be executed, including waiting for specific events or signals. Based on task configuration information, as well as hardware optimization strategies and constraints, the scheduling module determines the optimal task package size, generates the task package for the first task, and schematically divides threads into thread blocks or organizes thread blocks into a grid.

[0077] The scheduling module generates the first task's task package by determining the optimal task package size, enabling task scheduling at the task package level. Compared to finer-grained scheduling levels, this reduces the decision latency of the scheduling module and lowers scheduling overhead. Managing resources on a task package basis avoids the generation of small, discontinuous blocks of free storage space that are sufficient in total but cannot be allocated, improving resource utilization and allocation speed, thereby enhancing the load balancing of computing units. Simultaneously, it receives tasks in advance or in parallel and executes resource allocation to the target computing unit, avoiding waiting for resource request feedback, thus dispatching tasks more quickly and reducing task execution time overhead.

[0078] Step 420: In response to receiving the first task, obtain the actual resource requirements of the first task.

[0079] In response to receiving the task package for the first task, the resource quantity required by the task instances in the task package is obtained to determine the actual resource requirement of the first task. A task instance is typically an independent data element or work item that needs to be executed in parallel within the task package. In some embodiments, the specific implementation of the task instance differs depending on the task type. Taking GPU as an example, schematically, in a graphics rendering task, a task instance is a pixel instance that needs to be drawn; in a general computing task, a task instance is a data element that needs to be processed. Optionally, a task instance corresponds one-to-one with a thread; that is, a task implemented as a thread block contains multiple task instances, where the number of task instances is consistent with the number of threads in the thread block.

[0080] In some embodiments, the first task includes the number of task instances. Optionally, the number of task instances is part of the task configuration information used to indicate the task execution scale and resource requirements. Illustratively, obtaining the actual resource requirements of the first task can be achieved through the following steps: 1. Based on the first task, obtain the unit resource requirement of the task instance. Optionally, the task configuration information includes the unit resource requirement of the task instance.

[0081] 2. Based on the unit resource requirement of a task instance and the number of task instances, calculate the total resource requirement of the task as a whole. For example, the total resource requirement of the first task = the unit resource requirement of the task instance × the number of task instances in the task package of the first task.

[0082] It is worth noting that the method for obtaining the actual resource requirements of the first task is only illustrative, and this embodiment does not limit the specific method for obtaining the actual resource requirements of the first task.

[0083] Improving parallel capacity and throughput involves obtaining the actual resource requirements of the first task, ensuring that the first task only occupies the necessary computing resources, avoiding the waste of computing resources by occupying them without using them, and increasing the number of tasks in each computing unit, thereby improving resource utilization, the parallelism of task execution, and the overall computing throughput.

[0084] It is worth noting that there is no time-series dependency between receiving the first task, generating the task package for the first task, obtaining the actual resource requirements of the first task, receiving the resource pre-request instruction, and pre-occupying the first computing resources. That is, packaging the first task can be executed after resource pre-occupation, and packaging the first task can also be executed in parallel with resource pre-occupation.

[0085] Step 430: Obtain the difference between the resource quantity of the first computing resource and the actual resource requirement.

[0086] Obtain the amount of resources required for a reference task, which is used to indicate a threshold for the amount of resources required for the first task; based on the amount of resources required for the reference task, determine the amount of resources for the first computing resource, which is greater than or equal to the actual resource requirement of the first task.

[0087] In some embodiments, a reference task is pre-set. Illustratively, when the task types of the first tasks are the same, the first tasks can be divided into standardized tasks and non-standardized tasks. The number of task instances and the unit resource requirement of each task instance in the task package of a standardized task are fixed; that is, the total resource requirement of different standardized tasks is the same. The standardized task is the reference task. At least one of the following values ​​is less in the task package of a non-standardized task than in the standardized task: the number of task instances or the unit resource requirement of each task instance. Most of the tasks executed by the computing unit are standardized tasks, and a few are non-standardized tasks. Optionally, the resource pre-request instruction includes the resource quantity of the first computing resource.

[0088] In an optional embodiment, the process of determining the resource quantity of the first computing resource can also be implemented as follows: determining the resource quantity of the first computing resource based on a preset resource amplitude value, wherein the resource amplitude value is preset based on the task type of the first task; illustratively, the resource amplitude value is larger when the task type is a complex task, and smaller when the task type is a simple task.

[0089] Optionally, taking a GPU as an example, the GPU adopts a single-instruction, multi-threaded model, meaning that for the same GPU core, different threads execute the same instruction stream. Therefore, the execution paths of different threads and their corresponding unit resource requirements are fixed before thread execution, thus ensuring that the unit resource requirement of task instances in the task package of the reference task is fixed. When the task types of the first tasks are the same, to simplify the scheduling process and ensure consistency and predictability of thread behavior, the task package of the first task is implemented as a thread block of uniform size and regular resource requirements; that is, the number of task instances and the unit resource requirement of task instances in the task package of the reference task are fixed. Correspondingly, when the task types of the first task are different, the amount of first computing resources required by the first task is usually different. Illustratively, when the first task is a graphics rendering task, the corresponding amount of first computing resources is usually greater than when the first task is a general computing task.

[0090] For illustrative purposes, please refer to the following: Figure 5 , Figure 5 This is a schematic diagram illustrating the number of examples in the task packages of different tasks provided in an exemplary embodiment of this application, such as... Figure 5As shown, the set of tasks to be executed includes N+1 tasks, where the first task 510, the second task 520, the third task 530, the fourth task 540, and the fifth task 550 correspond to task numbers 0, 1, 2, N-1, and N, respectively. Tasks to be executed with task numbers in the range [1, N-1] are all standardized tasks. Therefore, the number of task instances in the task package of standardized tasks and the unit resource requirement of each task instance are fixed, meaning the total resource requirement of tasks to be executed with task numbers in the range [1, N-1] is the same: the number of task instances corresponding to the first task 510 is 6; the number of task instances corresponding to the second task 520, the third task 530, and the fourth task 540 is 8; and the number of task instances corresponding to the fifth task 550 is 7. Standardized tasks are reference tasks. In the task package of non-standardized tasks, at least one of the following values—the number of task instances or the unit resource requirement of each task instance—is less than that of standardized tasks; that is, the first task 510 and the fifth task 550 are non-standardized tasks.

[0091] In some embodiments, the above-mentioned pre-configured resource amount of the first computing resource is implemented in ways including but not limited to: driving to estimate the resource amount of the first computing resource based on the task identifier corresponding to the first task, in order to standardize the number of task instances corresponding to the task and the unit resource requirement of the task instance. Optionally, the task identifier is used to indicate the task type of the first task.

[0092] Since the resource quantity of the first computing resource is pre-configured based on the total resource requirement corresponding to the standardized task, and the resource quantity of the first computing resource is not less than the actual resource requirement of the first task, the resource quantity difference between the resource quantity of the first computing resource and the actual resource requirement is equal to the resource quantity of the first computing resource minus the resource quantity difference between the actual resource requirement.

[0093] Since the first task is typically a standardized task, the resource pre-allocation instruction determines the first set of computing resources required for the first task and ensures that the amount of the first set of computing resources is not less than the actual resource requirement of the first task. This allows for the pre-allocation of the first set of computing resources required for the first task, reducing the time overhead of task execution and improving the utilization rate of computing resources in the computing unit.

[0094] Step 440: Dispatch the first task to the target computing unit and send a resource release request to the target computing unit.

[0095] A resource release request is used to instruct the target computing unit to release computing resources corresponding to the difference in resource quantity.

[0096] In response to receiving the first task, a task package for the first task is generated, wherein the first task is a parsed task; the task package for the first task is dispatched to the target computing unit.

[0097] In some embodiments, the process by which the target computing unit bases a resource release request includes the following steps: 1. Based on the difference between the resource quantity of the first computing resource and the actual resource demand, the scheduling module generates a resource release request. The resource release request includes a unique identifier of the target computing unit and the resource quantity difference.

[0098] 2. The scheduling module sends a resource release request to the target computing unit. Indicatively, the resource release request is sent to the resource management module corresponding to the computing unit group where the target computing unit resides. The resource management module is used to maintain real-time computing resources in the computing unit group, including managing the occupied and idle status of computing resources; responding to the scheduling module's request, it checks, allocates, or releases redundant computing resources corresponding to the request, on a thread block basis; after the task is completed, it automatically reclaims the released computing resources and resets the status of the computing resources to idle, so that they can be used by subsequent tasks.

[0099] 3. The resource management module obtains the resource entries that need to be released corresponding to the first task based on the unique identifier of the target computing unit in the resource release request and the resource quantity difference, including but not limited to the starting address and size of the contiguous register block that meets the resource release requirements, and the pointer to the shared memory block that meets the resource release requirements.

[0100] 4. Based on the resource entries that need to be released corresponding to the first task, release redundant computing resources. Indicatively, the resource management module resets the bitmap status bits of the register file and shared memory region that meet the resource release requirements from "occupied" to "free". It updates the resource counter corresponding to the target computing unit and marks the thread bundle / block slots corresponding to the released computing resources as available.

[0101] 5. After the resource release operation is completed, the resource management module sends a corresponding notification to the scheduling module to indicate the update of available computing resources.

[0102] The scheduling module, upon receiving the first task and based on the occupancy feedback sent by the target computing unit, dispatches the first task to the target computing unit. This dispatch of the first task occurs after successfully securing the first computing resource. In other words, before dispatching the first task, it is necessary to confirm that the computing resources required for the first task have been allocated.

[0103] In some embodiments, the scheduling module includes a packaging unit, a scheduling unit, and a dispatching unit. The packaging unit, based on generating a task package for the first task, sends the first task to the scheduling unit. Upon receiving the packaged first task from the packaging unit, the scheduling unit checks whether the computing resources required for the first task have been fully allocated, based on the number of task instances in the task package. Optionally, if the computing resources required for the first task have been fully allocated, the scheduling unit sends the first task to the dispatching unit, which, upon receiving the first task, sends it to the target computing unit; if the computing resources required for the first task have not been fully allocated, the scheduling unit waits for feedback from the target computing unit regarding resource pre-allocation, or, the scheduling unit sends a resource query request to the computing unit, selects a target computing unit based on the feedback from the resource query request, and then sends a resource request to the target computing unit, waiting for feedback from the target computing unit until the computing resources required for the first task have been fully allocated.

[0104] The scheduling module ensures that the computing resources required for the first task have been fully allocated, thus preventing the first task from being dispatched directly to a computing unit if the required computing resources for the first task have not been fully allocated, which would prevent the first task from being dispatched to a computing unit with insufficient computing resources. This avoids blocking or performance loss caused by allocation failure and improves the robustness of the task execution process.

[0105] In summary, the method provided in this embodiment, based on resource pre-request instructions, estimates the amount of resources required for a task and obtains the resource information corresponding to the computing unit, thereby determining the target computing unit for executing the task. On the one hand, it obtains the resource information corresponding to the computing unit in advance of receiving the first task, thus avoiding waiting for resource status feedback during task scheduling. On the other hand, it executes resource allocation to the target computing unit in advance or in parallel with receiving the task, avoiding waiting for resource request feedback, thereby dispatching tasks more quickly, reducing the time overhead of task execution, improving the utilization rate of computing resources in the computing unit, and the data throughput during task execution.

[0106] The method provided in this embodiment obtains the difference between the resource quantity of the first computing resource and the actual resource demand, and sends a resource release request to the target computing unit. This avoids pre-occupied computing resources exceeding the actual required computing resources, thus preventing waste of computing resources and improving resource utilization. At the same time, it releases redundant computing resources, updates the available resource quantity, and can allocate subsequent tasks to idle computing resources, thereby reducing task dispatch blocking and improving task execution efficiency.

[0107] Figure 6 This is a flowchart of a task execution method provided in another embodiment of this application. The method is performed by a computer device (which can be implemented as follows). Figure 1 The computer device 120 shown is executed, and the above steps 310 to 320 are further implemented as steps 610 to 650.

[0108] Step 610: In response to receiving the resource pre-request instruction, obtain the resource information corresponding to at least two computing units respectively.

[0109] In response to receiving a resource pre-request instruction, a resource information acquisition request is sent to at least two computing units; resource information reported by at least two computing units is received; or, in response to receiving a resource pre-request instruction, the count data of the resource pool consumption counter is obtained to obtain resource information corresponding to at least two computing units respectively, wherein the resource pool consumption counter is updated when at least two computing units have completed their tasks.

[0110] In some embodiments, the scheduling module maintains a resource information table to store resource information corresponding to computing units. This resource information table can be implemented as a data structure, including but not limited to linked lists, circular arrays, bitmaps, etc. Computing units periodically, or upon task completion, report their current resource information to the scheduling module, which then updates the resource information table accordingly. When the scheduling module receives a resource pre-request instruction, it queries the resource information table to obtain the resource information of the corresponding computing unit. Alternatively, the scheduling module sets a resource pool consumption counter to record the occupancy status of computing units. For example, when a computing unit has not received a task, all thread bundles / block slots are idle, and the counter value is set to 0. During task execution, the counter value increases by a corresponding amount; when a computing unit completes a task, it sends a hardware completion signal to the scheduling module, causing the scheduler to automatically decrease the corresponding counter, thereby updating the resource pool consumption counter. When the scheduling module receives a resource pre-request instruction, it queries the resource pool consumption counter to obtain the corresponding computing unit's count data, i.e., retrieves the resource information.

[0111] In other embodiments, the resource pool is implemented as a set of idle computing resources pre-created and maintained by the scheduling module to indicate at least two computing units, i.e., a resource pool. Computing resources are obtained from this set when needed and returned to the set after use, avoiding the overhead of frequent creation and destruction. Obtaining resource information corresponding to at least two computing units through the resource pool includes the following process: 1. The scheduling module knows the total amount of computing resources corresponding to at least two computing units in the resource pool. When the scheduling module obtains idle computing resources from the resource pool, the count data of the resource pool consumption counter increases accordingly; when the scheduling module returns computing resources to the resource pool, the count data of the resource pool consumption counter decreases accordingly. That is, based on the count data of the resource pool consumption counter, the scheduling module obtains the amount of computing resources occupied in the resource pool in real time.

[0112] 2. Based on the total amount of computing resources in the resource pool and the amount of computing resources that are occupied in the resource pool in real time, the scheduling module can obtain the amount of real-time idle computing resources in the resource pool at any time. Illustratively, the amount of real-time idle computing resources = the total amount of computing resources - the amount of computing resources that are occupied in real time.

[0113] 3. Based on the resource pre-request instruction, the scheduling module obtains the data volume of real-time idle computing resources in the resource pool, that is, the resource information corresponding to at least two computing units.

[0114] It is worth noting that the above-mentioned method of obtaining resource information corresponding to at least two computing units is only illustrative, and this embodiment does not limit the specific method of obtaining resource information corresponding to at least two computing units.

[0115] By acquiring resource information in a centralized manner, the scheduling module can compare the available computing resources of at least two computing units from a global perspective. Based on simplifying resource information management, it can make optimal decision-making and select the target computing unit to execute the first task, avoiding uneven load that may be caused by local decision-making. At the same time, when the task is completed, the scheduling module is notified and the count data of the resource pool consumption counter is updated, ensuring that the scheduling module's decision is based on the real-time resource status and improving the success rate of resource allocation.

[0116] Step 620: Based on the first computing resources requested by the resource pre-request instruction, determine the target computing unit from the amount of computing resources that can be provided by at least two computing units.

[0117] The target computing unit can provide computing resources that are greater than or equal to the amount of computing resources of the first computing resource. Based on the resource information corresponding to at least two computing units and the first computing resource requested by the resource pre-request instruction, the target computing unit is determined from the computing resources that the at least two computing units can provide. The resource information corresponding to the at least two computing units indicates the amount of computing resources that each of the at least two computing units can provide, and the computing resources that the target computing unit can provide are greater than or equal to the amount of computing resources of the first computing resource. The target computing unit is one or more of the at least two computing units, and the resource information corresponding to the target computing unit satisfies, individually or collectively, the amount of computing resources requested by the resource pre-request instruction for the first computing resource.

[0118] In some embodiments, among at least two computing units that satisfy the condition that the amount of computing resources that can be provided is greater than or equal to the amount of the first computing resource, the method for determining the target computing unit includes at least one of the following methods: 1. Determining the target computing unit based on resource sufficiency: The scheduling module compares the resource information of at least two computing units and selects the one with the most abundant idle computing resources as the target computing unit. By evenly distributing the resource load, the system avoids saturating a single computing unit and reducing the overall task execution efficiency, thereby improving resource utilization and system throughput.

[0119] 2. Determine the target computing unit based on load conditions: The scheduling module obtains the task queues that are being executed and waiting to be executed on at least two computing units, and selects at least one computing unit with the shortest task queue as the target computing unit, thereby reducing the waiting time for task execution, improving the parallelism of task execution and reducing response latency.

[0120] 3. Determining the target computation unit based on data affinity: The scheduling module obtains historical execution records of tasks and selects the computation unit selected during the last successful execution of the first task or a historical task of similar type or scale as the target computation unit. By reusing the same or similar computation units, the cache hit rate is improved, thereby increasing the task execution success rate.

[0121] It is worth noting that the above-described method for determining the target computing unit is merely illustrative, and this embodiment does not limit the specific method for obtaining and determining the target computing unit. Furthermore, the above method can be implemented independently, or it can be implemented jointly through a weighted comprehensive decision-making process using a scheduling module based on the global load or specific task characteristics of the computing units in the computer device.

[0122] Step 630: Send a resource pre-occupancy instruction to the target computing unit.

[0123] Step 640: In response to receiving occupancy feedback, obtain the storage space corresponding to the target computing unit.

[0124] Receive occupancy feedback sent by the target computing unit, which indicates that the first computing resource has been successfully occupied; in response to the occupancy feedback, acquire the storage space corresponding to the target computing unit.

[0125] Schematic, when computing resources in a target computing unit that satisfy a resource pre-request instruction are successfully occupied, the target computing unit sends an occupation feedback to the scheduling module. In some embodiments, the occupation feedback is implemented in at least one of the following forms: 1. Explicit confirmation signal, using independent communication to provide feedback to the scheduling module. When the resource management module corresponding to the target computing unit receives a resource pre-request instruction, it checks and allocates the computing resources of the target computing unit. Upon successful allocation (i.e., successful occupancy), the resource management module generates a dedicated confirmation signal and sends it to the scheduling module. When the occupancy feedback is implemented as an explicit confirmation signal, the components of the occupancy feedback include, but are not limited to: a success status bit indicating successful allocation, a task identifier indicating the resource pre-request instruction corresponding to the occupancy feedback, and optionally, a resource handle to the allocated shared memory base address.

[0126] 2. Implicit allocation signals, using indirect feedback to save on confirmation steps, thereby improving execution speed. When the scheduling module selects a target computing unit, it assumes sufficient computing resources and successful resource pre-allocation. That is, when the amount of computing resources that the target computing unit can provide is greater than or equal to the amount of resources of the first computing unit, it is indirectly but definitively confirmed that the computing resources in the target computing unit that satisfy the resource pre-allocation instruction are successfully occupied. When the occupancy feedback is implemented as an implicit allocation signal, if the resource pre-allocation fails unexpectedly, when the first task is dispatched to the target computing unit, it may cause hardware abnormalities or task timeouts. In this case, the scheduling module detects the failure of the resource pre-allocation instruction through an error handling mechanism and re-allocates computing resources.

[0127] Schematic, the storage space corresponding to the target computing unit refers to the storage space that needs to be initialized for the target computing unit, used to store data during the execution of the first task. Upon receiving occupancy feedback, the scheduling module identifies the storage space that needs to be initialized. In some embodiments, the scheduling module obtains an initialization instruction based on the task configuration information corresponding to the first task. The initialization instruction includes the type of storage space to be initialized, its starting address, and the initialization method.

[0128] In some embodiments, the possible implementations of the storage space that needs to be initialized include, but are not limited to: shared memory, which is a collaborative workspace for all threads within a thread block and needs to be initialized to zero or loaded with initial input data; local memory, used to store overflowing variables and needs to be initialized to default values; and specific areas in global memory, such as output buffers used to store intermediate results, which need to be cleared before use.

[0129] Step 650: Send a preload request to the storage module.

[0130] The preload request instructs the storage module to perform a preload operation on the storage space corresponding to the target computing unit. Optionally, the storage module is implemented as global memory. In some embodiments, the preload request is used to initialize the storage space corresponding to the target computing unit. Initializing the storage space is used to create a predictable and ready working environment based on the storage space before the task starts execution. The initialization method includes, but is not limited to, setting it to zero, loading initial input data from global memory, or establishing a specific data structure.

[0131] Initialization ensures that there is no residual or uncleared data in the storage space corresponding to the target computing unit, eliminating interference from dirty data and ensuring the correctness and determinism of the calculation results, thereby guaranteeing the correctness of the program and the consistency of the results of each execution. When the initialization method is to load the initial input data, it provides initial values ​​for operations such as multiplication, addition, comparison, and accumulation, avoiding program errors caused by undefined behavior. Furthermore, initialization is performed when pre-allocating resources, that is, before the first task is dispatched to the target computing unit, the storage space corresponding to the target computing unit has been initialized, without the need to pause and wait for the initial input data, reducing memory access latency and improving resource utilization.

[0132] In some embodiments, the process by which the scheduling module sends a preloading request to the storage module to initialize the storage space corresponding to the target computing unit includes, but is not limited to, the following steps: 1. The computing unit receives the preload request sent by the scheduling module and generates a memory transaction request. The memory transaction request includes, but is not limited to: source address, which indicates the virtual address of the initial input data in global memory; destination address, which indicates the storage space address of the target computing unit to which the initial input data needs to be written, including a specific location in registers or shared memory; and data size, which indicates the amount of data to be preloaded.

[0133] 2. The memory management unit receives the preload request and translates the source address from a virtual address to a physical address by querying the page table. The translated source address and the preload request are then sent to the Level 2 (L2) cache for querying. The L2 cache stores shared temporary storage data for at least two computing units. If the translated source address is in the L2 cache (i.e., an L2 cache hit), the initial input data is read directly from the L2 cache; otherwise, the preload request is sent to global memory.

[0134] 3. Based on the preload request, access global memory to obtain the initial input data indicated by the source address. The initial input data read from global memory is sent to the target computing unit, and written to the L2 cache en route for subsequent access.

[0135] 4. Based on the load request and caching strategy, as well as the initial input data, store the initial input data in the L1 cache or write it directly to the Load / Store Unit (LSU).

[0136] 5. Based on the destination address in the load request, the LSU writes the initial input data to the storage space corresponding to the target computing unit. If the storage space corresponding to the target computing unit is a register, the initial input data is written to the thread-specific register file corresponding to the first task; if the storage space corresponding to the target computing unit is shared memory, the initial input data is written to the shared memory area allocated by the thread block corresponding to the first task.

[0137] In summary, the method provided in this embodiment, based on resource pre-request instructions, estimates the amount of resources required for a task and obtains the resource information corresponding to the computing unit, thereby determining the target computing unit for executing the task. On the one hand, it obtains the resource information corresponding to the computing unit in advance of receiving the first task, thus avoiding waiting for resource status feedback during task scheduling. On the other hand, it executes resource allocation to the target computing unit in advance or in parallel with receiving the task, avoiding waiting for resource request feedback, thereby dispatching tasks more quickly, reducing the time overhead of task execution, improving the utilization rate of computing resources in the computing unit, and the data throughput during task execution.

[0138] The method provided in this embodiment allocates the optimal target computing unit to the first task based on the first computing resources requested by the resource pre-request instruction through the scheduling module. That is, the first task can start execution immediately without waiting for idle time, reducing task scheduling delay, improving resource load balancing, and improving resource utilization while improving task processing efficiency.

[0139] For illustrative purposes, please refer to the following: Figure 7 , Figure 7This is a schematic diagram of a task processing flow provided in an exemplary embodiment of this application, such as... Figure 7 As shown, taking GPU execution of the first task as an example, the CPU, including the GPU driver, receives the first task execution request from the application. Based on the first task execution request, step 710 is executed to generate a resource pre-allocation instruction corresponding to the first task, and step 720 is executed to generate a task execution instruction corresponding to the first task. The CPU sends the resource pre-allocation instruction to the GPU, and the command stream processor in the GPU executes step 730 to receive and parse the resource pre-allocation instruction, and sends the resource pre-allocation instruction to the scheduler in the GPU. The scheduler executes step 740 to send a resource pre-occupancy instruction to the target computing unit in at least two computing units based on the resource pre-allocation instruction, and step 750, in response to receiving occupancy feedback, acquires the storage space corresponding to the target computing unit. The occupancy feedback is sent to the scheduler by the target computing unit after it has occupied the storage space based on the resource pre-occupancy instruction.

[0140] like Figure 7 As shown, after the CPU generates the task execution instruction corresponding to the first task, it sends the task execution instruction to the GPU. The command stream processor in the GPU executes step 760, receiving and parsing the task execution instruction. Based on the parsed task execution instruction, the command stream processor sends the task execution instruction to the packetizer in the GPU. The packetizer executes step 770, generating a task packet for the first task in response to receiving the first task. Based on the generated task packet, the packetizer sends the first task to the scheduler in the form of a task packet. The scheduler executes step 780, determining, in response to the occupancy feedback sent by the target computing unit, that the computing resources required for the first task have been allocated. Based on determining that the computing resources required for the first task have been allocated, the scheduler sends the first task to the dispatcher in the GPU. The dispatcher executes step 790, dispatching the first task to the target computing unit and sending a resource release request to the target computing unit.

[0141] For illustrative purposes, please refer to the following: Figure 8 , Figure 8 This is a schematic diagram of the structure of a task processing system provided in an exemplary embodiment of this application, as shown below. Figure 8 As shown, the task processing system is used to execute task execution methods. Optionally, the host 810 includes a CPU and system memory, and the device 820 includes a GPU. The host 810 and the device 820 serve as two collaborative processing units in the task processing system. The host 810 sends a first task execution request and a resource pre-allocation request to the device 820 via an API. The command stream processor 830 receives the first task execution request and the resource pre-allocation request and converts them into first task execution instructions and resource pre-allocation instructions in hardware instruction form, respectively.

[0142] like Figure 8 As shown, the scheduling system 840 receives a resource pre-request instruction sent by the command stream processor 830, obtains resource information corresponding to at least two computing units, and sends a resource pre-occupancy instruction to the target computing unit among the at least two computing units based on the first computing resource requested by the resource pre-request instruction. Optionally, any computing unit in the figure can be used as the target computing unit. For example, computing unit 1 is used as the target computing unit 850. The scheduling system 840 receives the occupancy feedback sent by the target computing unit 850 and obtains the storage space corresponding to the target computing unit 850. Optionally, the storage space corresponding to the target computing unit 850 is implemented as a local register. First, a preload request is sent to the cache 870. If a cache miss occurs, a preload request is sent to the storage system 860. The preload request is used to instruct the storage system 860 to perform a preload operation on the storage space corresponding to the target computing unit 850.

[0143] like Figure 8 As shown, the scheduling system 840 receives the first task execution instruction and the task configuration information corresponding to the first task sent by the command stream processor 830, and generates a task package for the first task based on the task configuration information. In response to receiving the first task, and based on receiving the occupancy feedback sent by the target computing unit 850, the scheduling system 840 dispatches the first task to the target computing unit 850; that is, it dispatches the first task to the target computing unit 850 after determining that the first computing resources required by the first task have been successfully occupied. When dispatching the first task to the target computing unit 850, the scheduling system 840, in response to receiving the first task, obtains the actual resource requirement of the first task; obtains the resource difference between the resource quantity of the first computing resource and the actual resource requirement; dispatches the first task to the target computing unit 850; and sends a resource release request to the target computing unit 850, the resource release request instructing the target computing unit to release the computing resources corresponding to the resource difference.

[0144] The following are embodiments of the graphics processor provided in this application, which can be used to execute the method embodiments of this application. For details not disclosed in the graphics processor embodiments of this application, please refer to the method embodiments of this application.

[0145] Please refer to Figure 9 It illustrates a structural block diagram of a graphics processor provided in an exemplary embodiment of this application, such as... Figure 9 As shown, the graphics processor includes: Scheduler 910 is used to obtain resource information corresponding to at least two computing units in response to receiving a resource pre-request instruction; the resource information is used to indicate the amount of computing resources that the computing units can provide. Scheduler 910 is also used to send a resource pre-occupancy instruction to target computing unit 920 among at least two computing units based on the first computing resources requested by the resource pre-application instruction; the resource pre-occupancy instruction is used to pre-occupy the first computing resources before dispatching the first task; the first computing resources are the computing resources required when the first task is executed, and are estimated based on the first task. Dispatcher 930 is used to dispatch the first task to target computing unit 920 in response to receiving the first task; target computing unit 920 is used to execute the first task using the first computing resources that have been pre-occupied.

[0146] In an optional embodiment, the scheduler 910 is further configured to obtain the actual resource requirements of the first task in response to receiving the first task execution instruction; Scheduler 910 is also used to obtain the difference between the resource quantity of the first computing resource and the actual resource demand. The dispatcher 930 is also used to dispatch the first task to the target computing unit 920, and to send a resource release request to the target computing unit 920, the resource release request being used to instruct the target computing unit 920 to release the computing resources corresponding to the resource difference.

[0147] In an optional embodiment, the scheduler 910 is further configured to obtain the amount of resources required by a reference task, the amount of resources required by the reference task being used to indicate a threshold of the amount of resources required by the first task. Scheduler 910 is also used to determine the amount of resources of the first computing resource based on the amount of resources required by the reference task, wherein the amount of resources of the first computing resource is greater than or equal to the actual resource requirement of the first task.

[0148] In an optional embodiment, the scheduler 910 is further configured to determine a target computing unit 920 from the amount of computing resources that can be provided by at least two computing units based on the first computing resources requested by the resource pre-request instruction, wherein the amount of computing resources that the target computing unit 920 can provide is greater than or equal to the amount of resources of the first computing resource. The scheduler 910 is also used to send resource pre-occupancy instructions to the target computing unit 920.

[0149] In an optional embodiment, the target computing unit 920 is configured to send an occupation feedback to the scheduler 910 based on a resource pre-occupancy instruction, wherein the occupation feedback is used to indicate that the first computing resource has been successfully occupied. The scheduler 910 is also used to, in response to receiving occupancy feedback, acquire the storage space corresponding to the target computing unit 920; and generate a preload request, which is used to instruct a preload operation to be performed on the storage space corresponding to the target computing unit 920. The graphics processor also includes a storage module 940, which, in response to receiving a preload request from the scheduler 910, sends preload data to the storage space corresponding to the target computing unit 920. The preload data is pre-configured by the GPU driver.

[0150] In an optional embodiment, the scheduler 910 is further configured to, in response to receiving the first task, send the first task to the dispatcher 930 based on receiving occupancy feedback sent by the target computing unit 920; The dispatcher 930 is also used to dispatch a first task to the target computing unit 920, wherein the dispatch of the first task is performed after successfully occupying the first computing resource.

[0151] In an optional embodiment, the graphics processor further includes: a packager 950, configured to generate a task package for the first task in response to receiving a first task execution instruction; wherein the scheduler 910 receives the resource pre-request instruction before the packager 950 receives the first task execution instruction, or the scheduler 910 receives the resource pre-request instruction and the packager 950 receives the first task execution instruction synchronously. Packer 950 is also used to send the task packet of the first task to scheduler 910.

[0152] In an optional embodiment, the graphics processor further includes: a command stream processor 960, configured to parse the resource pre-request instruction and the first task execution instruction sent by the GPU driver; The command stream processor 960 is also used to send parsed resource pre-request instructions to the scheduler 910; The command stream processor 960 is also used to send the parsed first task execution instruction to the packer 950.

[0153] In an optional embodiment, the scheduler 910 is further configured to, in response to receiving a task package of the first task, obtain the amount of resources required by the task instance in the task package, and obtain the actual resource requirement of the first task.

[0154] In summary, the graphics processor provided in this embodiment, based on resource pre-request instructions, estimates the amount of resources required for a task and obtains the resource information corresponding to the computing unit, thereby determining the target computing unit for executing the task. On the one hand, it obtains the resource information corresponding to the computing unit in advance of receiving the first task, thus avoiding waiting for resource status feedback during task scheduling. On the other hand, it executes resource allocation to the target computing unit in advance or in parallel with receiving the task, avoiding waiting for resource request feedback, thereby dispatching tasks more quickly, reducing the time overhead of task execution, and improving the utilization rate of computing resources in the computing unit, as well as the data throughput during task execution.

[0155] It should be noted that the graphics processor provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the graphics processor provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0156] Figure 10 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application. Optionally, the computer device 1000 is a terminal.

[0157] The computer device 1000 can be a portable mobile terminal, also referred to as a mobile terminal in this embodiment. Examples include smartphones, tablets, Moving Picture Experts Group Audio Layer III (MP3) players, and Moving Picture Experts Group Audio Layer IV (MP4) players. The computer device 1000 may also be referred to as a user device, portable terminal, or other names.

[0158] Typically, computer device 1000 includes a graphics processor 1001 and a memory 1002.

[0159] The graphics processor 1001 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The graphics processor 1001 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). In some embodiments, the graphics processor 1001 is used to execute the task execution method provided in the embodiments of this application, and the graphics processor 1001 is also used to render and draw the content that the display screen needs to display. In some embodiments, the graphics processor 1001 may also include an AI processor, which is used to handle computational operations related to machine learning.

[0160] The memory 1002 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1002 may also include high-speed random access memory devices and non-volatile storage devices, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one instruction, which is executed by the processor 1002 to implement the task execution methods provided in the various method embodiments of this application.

[0161] In some embodiments, the computer device 1000 may also optionally include: a peripheral device interface 1003 and at least one peripheral device.

[0162] On the other hand, embodiments of this application provide a computer device, which includes a graphics processor and a memory; the memory is used to store computer programs; the graphics processor is used to load the computer programs to execute the task execution method provided in the embodiments of this application above.

[0163] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the task execution method provided in the embodiments of this application above.

[0164] On the other hand, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the task execution method provided in the embodiments of this application described above.

[0165] On the other hand, embodiments of this application provide a computer device including the processor described above. Optionally, the processor is a GPU. The computer device can be at least one of a portable computer, a desktop computer, a server, a server cluster, an AI computing cluster, and a cloud computing cluster. The AI ​​computing cluster can also be simply referred to as an intelligent computing cluster or a smart computing cluster.

[0166] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0167] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0168] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0169] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A task execution method, characterized in that, The method includes: In response to receiving a resource pre-request instruction, resource information corresponding to at least two computing units is obtained; the resource information is used to indicate the amount of computing resources that the computing units can provide. Based on the first computing resource requested by the resource pre-request instruction, a resource pre-occupancy instruction is sent to the target computing unit among the at least two computing units; the resource pre-occupancy instruction is used to pre-occupy the first computing resource before dispatching the first task; the first computing resource is the computing resource required for the execution of the first task, and is estimated based on the first task; In response to receiving the first task, the first task is dispatched to the target computing unit; the target computing unit is used to execute the first task using the first computing resources that have been pre-occupied.

2. The method according to claim 1, characterized in that, The step of dispatching the first task to the target computing unit in response to receiving the first task includes: In response to receiving the first task, obtain the actual resource requirements of the first task; Obtain the difference between the resource quantity of the first computing resource and the actual resource requirement; The first task is dispatched to the target computing unit, and a resource release request is sent to the target computing unit, the resource release request being used to instruct the target computing unit to release the computing resources corresponding to the resource difference.

3. The method according to claim 2, characterized in that, The step of responding to receiving the first task and obtaining the actual resource requirements of the first task includes: In response to receiving the task package of the first task, the resource quantity required by the task instance in the task package is obtained to obtain the actual resource requirement of the first task.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the amount of resources required for a reference task, wherein the amount of resources required for the reference task is used to indicate a threshold for the amount of resources required for the first task; Based on the resource requirements of the reference task, the resource quantity of the first computing resource is determined, wherein the resource quantity of the first computing resource is greater than or equal to the actual resource requirement of the first task.

5. The method according to any one of claims 1 to 4, characterized in that, The response to receiving a resource pre-request instruction, obtaining resource information corresponding to at least two computing units respectively, includes: In response to receiving the resource pre-request instruction, a resource information acquisition request is sent to the at least two computing units; the resource information reported by the at least two computing units is received; or, In response to receiving the resource pre-request instruction, the count data of the resource pool consumption counter is obtained to obtain the resource information corresponding to the at least two computing units respectively, wherein the resource pool consumption counter is updated in real time following the resource occupation or resource release of the at least two computing units.

6. The method according to any one of claims 1 to 5, characterized in that, The step of sending a resource pre-occupancy instruction to a target computing unit among the at least two computing units based on the first computing resource requested by the resource pre-request instruction includes: Based on the first computing resource requested by the resource pre-request instruction, a target computing unit is determined from the amount of computing resources that can be provided by the at least two computing units, wherein the amount of computing resources that the target computing unit can provide is greater than or equal to the amount of the first computing resource; Send the resource pre-occupancy instruction to the target computing unit.

7. The method according to claim 6, characterized in that, After sending the resource pre-occupancy instruction to the target computing unit, the method further includes: Receive occupancy feedback sent by the target computing unit, the occupancy feedback being used to indicate that the first computing resource was successfully occupied; In response to the occupancy feedback, the storage space corresponding to the target computing unit is obtained; A preload request is sent to the storage module, which instructs the storage module to perform a preload operation on the storage space corresponding to the target computing unit.

8. The method according to claim 7, characterized in that, After sending the preload request to the storage module, the process also includes: In response to receiving the first task, and based on receiving the occupancy feedback sent by the target computing unit, the first task is dispatched to the target computing unit, wherein the dispatch of the first task occurs after the first computing resource is successfully occupied.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: In response to receiving the first task, a task package for the first task is generated, wherein the first task is a parsed task; and in response to receiving the resource pre-request instruction, a target computing unit among the at least two computing units is determined; wherein the resource pre-request instruction is received before receiving the first task, or the resource pre-request instruction and the first task are received simultaneously. The task package of the first task is dispatched to the target computing unit.

10. A graphics processor, characterized in that, The graphics processor includes: A scheduler, in response to receiving a resource pre-request instruction, acquires resource information corresponding to at least two computing units respectively; the resource information is used to indicate the amount of computing resources that the computing units can provide; based on the first computing resources requested by the resource pre-request instruction, sends a resource pre-occupancy instruction to a target computing unit among the at least two computing units; the resource pre-occupancy instruction is used to pre-occupy the first computing resources before dispatching the first task; the first computing resources are the computing resources required for the execution of the first task, and are estimated based on the first task. A dispatcher is configured to dispatch the first task to the target computing unit in response to receiving the first task; the target computing unit is configured to execute the first task using the first computing resources pre-occupied.

11. The graphics processor according to claim 10, characterized in that, The scheduler is further configured to, in response to receiving a first task execution instruction, obtain the actual resource requirement of the first task; and obtain the resource difference between the resource quantity of the first computing resource and the actual resource requirement. The dispatcher is further configured to dispatch the first task to the target computing unit and send a resource release request to the target computing unit, the resource release request being used to instruct the target computing unit to release the computing resources corresponding to the resource difference.

12. The graphics processor according to claim 10, characterized in that, The scheduler is further configured to determine a target computing unit from the amount of computing resources that the at least two computing units can provide, based on the first computing resources requested by the resource pre-request instruction, wherein the amount of computing resources that the target computing unit can provide is greater than or equal to the amount of the first computing resource; The scheduler is also used to send the resource pre-occupancy instruction to the target computing unit.

13. The graphics processor according to claim 12, characterized in that, The target computing unit is configured to send an occupation feedback to the scheduler based on the resource pre-occupancy instruction, wherein the occupation feedback is used to indicate that the first computing resource has been successfully occupied. The scheduler is also configured to, in response to receiving the occupancy feedback, acquire the storage space corresponding to the target computing unit; A preload request is generated, which instructs a preload operation to be performed on the storage space corresponding to the target computing unit. The graphics processor also includes: The storage module is used to send preloaded data to the storage space corresponding to the target computing unit in response to receiving a preload request from the scheduler. The preloaded data is pre-configured by the GPU driver.

14. The graphics processor according to claim 13, characterized in that, The scheduler is further configured to, in response to receiving the first task, send the first task to the dispatcher based on receiving the occupancy feedback sent by the target computing unit; The dispatcher is further configured to dispatch the first task to the target computing unit, wherein the dispatch of the first task occurs after the first computing resource has been successfully occupied.

15. The graphics processor according to claim 10, characterized in that, The graphics processor also includes: A packager is configured to generate a task package for the first task in response to receiving the first task execution instruction; wherein the scheduler receives the resource pre-request instruction before the packager receives the first task execution instruction, or the scheduler receives the resource pre-request instruction synchronously with the packager receiving the first task execution instruction; The packer is also used to send the task package of the first task to the scheduler.

16. The graphics processor according to claim 15, characterized in that, The graphics processor also includes: A command stream processor is used to parse the resource pre-request instruction and the first task execution instruction sent by the GPU driver; The command stream processor is also used to send the parsed resource pre-request instruction to the scheduler; The command stream processor is also used to send the parsed first task execution instruction to the packer.

17. The graphics processor according to claim 15, characterized in that, The scheduler is further configured to, in response to receiving the task package of the first task, obtain the amount of resources required by the task instance in the task package, and obtain the actual resource requirement of the first task.

18. A computer device, characterized in that, The computer device includes a graphics processor and memory; The memory is used to store computer programs; The graphics processor is used to load the computer program to execute the task execution method according to any one of claims 1 to 9.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the task execution method as described in any one of claims 1 to 9.

20. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, and a processor reads from and executes the computer program to implement the task execution method as described in any one of claims 1 to 9.