A computing task processing method, device, medium and product

By designing FMA units using hardware description languages, direct hardware scheduling is supported, which solves the problem of hardware scheduling relying on software scheduling. This achieves task balancing of FMA unit resources, improves computing efficiency and resource utilization, and enhances system performance and energy efficiency.

CN120670125BActive Publication Date: 2026-01-23SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511172585.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2026-01-23
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Current hardware scheduling schemes rely on software scheduling and lack direct hardware scheduling capabilities. They do not perform task balancing based on FMA unit resources, which limits the optimization of system energy efficiency and performance.

Method used

The FMA unit is designed using a hardware description language, which supports direct hardware scheduling. It dynamically allocates computing tasks through hardware queues and a parallel scheduler, and determines the target unit for task dispatch based on the resource usage of the FMA unit, thus realizing a hardware-software combined scheduling method.

Benefits of technology

It improves computational efficiency and FMA resource utilization, achieves task balancing based on FMA unit resources, and enhances the overall performance and energy efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670125B_ABST
    Figure CN120670125B_ABST
Patent Text Reader

Abstract

The application discloses a kind of computing task processing method, equipment, medium and product, it is related to integrated circuit technical field.The present scheme designs FMA unit using hardware description language, supports hardware direct scheduling capability;When receiving multiple computing requests, corresponding each computing task is stored to hardware queue, and the scheduling information of each computing task is determined respectively;Scheduling information includes hardware scheduling and software scheduling, i.e.the present scheme supports two kinds of scheduling mode of software and hardware;When the scheduling information of computing task is hardware scheduling, according to the hardware resource use condition of each FMA unit, determine target FMA unit, dispatch computing task to target FMA unit, so that target FMA unit executes computing task, realizes task equalization based on FMA unit resource, improves computing efficiency and FMA resource utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated circuit technology, and in particular to a computing task processing method, device, medium, and product. Background Technology

[0002] Hardware schedulers are crucial in high-performance computing systems, responsible for efficiently managing computing resources to optimize task execution and system throughput. With the proliferation of heterogeneous computing resources, modern hardware schedulers employ pipelined architectures, priority arbitration mechanisms, and intelligent prediction algorithms to adapt to real-time, parallel, and distributed computing needs, and combine hardware acceleration technologies to improve performance.

[0003] However, current hardware scheduling schemes primarily rely on software task dispatch, performing task allocation and scheduling at the software level. Hardware merely serves as a data transmission channel, lacking direct scheduling capabilities at the hardware level, thus limiting scheduling efficiency and response speed. Furthermore, existing schemes use system tasks as the scheduling unit, rather than scheduling based on the actual resource availability of fused multiply-add (FMA) units. This prevents a fundamental balance between the power consumption and performance load of electronic devices, impacting the overall energy efficiency and performance optimization of the system.

[0004] Given the above problems, how to solve the current hardware scheduling problem, which relies on software scheduling, lacks direct hardware scheduling capabilities, and does not perform task balancing based on FMA unit resources, is an urgent issue for technical personnel in this field. Summary of the Invention

[0005] This invention provides a computing task processing method, device, medium, and product to at least solve the problems of current hardware scheduling relying on software scheduling, lacking direct hardware scheduling capabilities, and not performing task balancing based on FMA unit resources.

[0006] This invention provides a method for processing computational tasks, comprising:

[0007] When multiple computing requests are received, the computing data corresponding to each computing request is obtained, and each computing request is combined with the corresponding computing data to generate a computing task;

[0008] Each computing task is stored in a hardware queue, and the scheduling information for each computing task is determined; the scheduling information includes hardware scheduling and software scheduling.

[0009] When the scheduling information for the computing task is hardware scheduling, the target fusion multiply-accumulate computing unit is determined based on the hardware resource usage of each fusion multiply-accumulate computing unit; wherein, the fusion multiply-accumulate computing unit is a computing unit built based on a hardware description language and supports scheduling;

[0010] Establish a mapping relationship between the computation task and the target fusion multiply-accumulate operation unit, and dispatch the computation task to the target fusion multiply-accumulate operation unit so that the target fusion multiply-accumulate operation unit can execute the computation task.

[0011] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described computing task processing methods.

[0012] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described computational task processing methods.

[0013] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any of the above-described computing task processing methods.

[0014] The beneficial effects of this invention are as follows: An FMA unit is designed using a hardware description language, supporting direct hardware scheduling capabilities; when multiple computing requests are received, the corresponding computing tasks are stored in a hardware queue, and the scheduling information for each computing task is determined; the scheduling information includes hardware scheduling and software scheduling, meaning this solution supports both hardware and software scheduling methods; when the scheduling information for a computing task is hardware scheduling, the target FMA unit is determined based on the hardware resource usage of each FMA unit, and the computing task is dispatched to the target FMA unit so that the target FMA unit can execute the computing task, thus achieving task balancing based on FMA unit resources and improving computing efficiency and FMA resource utilization.

[0015] In addition, the present invention also provides a computing task processing device, medium and product, with the same effect as above. Attached Figure Description

[0016] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart of a computational task processing method provided in an embodiment of the present invention;

[0018] Figure 2 This is a diagram illustrating the architecture of a schedulable fused multiply-accumulate operation unit provided in an embodiment of the present invention.

[0019] Figure 3 This is a diagram of the parallel multiply-accumulate scheduling architecture provided in an embodiment of the present invention;

[0020] Figure 4 The scheduling flowchart of the hardware pipeline scheduler provided in the embodiments of the present invention;

[0021] Figure 5 A schematic diagram of a hardware queue structure provided in an embodiment of the present invention;

[0022] Figure 6 This is a timing diagram of the multiply-accumulator scheduling provided in an embodiment of the present invention;

[0023] Figure 7 A flowchart of the hardware scheduling and computing unit provided in an embodiment of the present invention;

[0024] Figure 8 This is a schematic diagram of a computing task processing device provided in an embodiment of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0026] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0027] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] Current hardware scheduling schemes primarily rely on software task dispatch, performing task allocation and scheduling at the software level. Hardware merely serves as a data transmission channel, lacking direct hardware-level scheduling capabilities, thus limiting scheduling efficiency and response speed. Furthermore, existing schemes use system tasks as the scheduling unit, rather than scheduling based on the actual resource availability of FMA units. This fails to fundamentally balance the power consumption and performance load of electronic devices, impacting overall system energy efficiency and performance optimization. To address these issues, this invention provides a computational task processing method. It should be noted that the method provided by this invention can be specifically applied to processors, Systems on Chips (SoCs), or other units with computing and storage functions, depending on the specific implementation.

[0029] Figure 1 This is a flowchart illustrating a computational task processing method provided in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0030] S10: When multiple computing requests are received, obtain the computing data corresponding to each computing request, and combine each computing request with the corresponding computing data to generate a computing task.

[0031] Figure 2 This is a diagram illustrating the architecture of a schedulable fused multiply-accumulate operation unit provided in an embodiment of the present invention. Figure 2 As shown, when multiple computation requests are received, the data acquisition module retrieves the computation data corresponding to each computation request, such as floating-point computation data, from memory or other storage media, and combines each computation request with its corresponding computation data to generate a computation task. It should be noted that this embodiment does not limit the specific method of combining computation requests and computation data; it depends on the specific implementation.

[0032] S11: Store each computing task in a hardware queue and determine the scheduling information for each computing task; the scheduling information includes hardware scheduling and software scheduling.

[0033] It should be noted that, as Figure 2As shown, the parallel scheduling module is a hardware-implemented parallel scheduler, a technology for efficiently managing computing resources. It supports up to 16 Functional Modules (FMAs) and dynamically allocates computing data and tasks to multiple FMA computing units to maximize resource utilization and performance. The parallel scheduler implements scheduling logic through dedicated circuits (such as state machines, pipelines, and lookup tables), offering higher efficiency and lower latency compared to software scheduling. It allocates tasks to idle FMAs, ensuring even load distribution across all computing units and avoiding resource waste; it dynamically adjusts the scheduling order based on task priority to handle resource contention. In this embodiment, the parallel scheduling module employs a pipelined scheduler to divide the scheduling process into multiple stages, improving throughput; a lookup table (LUT) to store scheduling policies and task mapping relationships; and a hardware queue to store tasks to be scheduled, supporting efficient task enqueueing and dequeueing operations.

[0034] Figure 3 This is a diagram of a parallel multiply-accumulate scheduling architecture provided in an embodiment of the present invention. Figure 3 As shown, the parallel scheduler provided by this invention supports both hardware and software scheduling. Therefore, in this embodiment, each computing task is specifically stored in the hardware queue of the parallel scheduling module, and the scheduling information for each computing task is determined accordingly. The scheduling information includes both hardware and software scheduling; therefore, the specific processing method for the current computing task can be determined based on the specific content of the scheduling information. This embodiment does not limit the specific process for determining the scheduling information; it depends on the specific implementation.

[0035] S12: When the scheduling information of the computing task is hardware scheduling, the target fusion multiply-accumulate computing unit is determined according to the hardware resource usage of each fusion multiply-accumulate computing unit; wherein, the fusion multiply-accumulate computing unit is a computing unit built based on a hardware description language and supporting scheduling.

[0036] Furthermore, when the scheduling information for a computing task is confirmed to be hardware scheduling, hardware scheduling for the computing task is executed. Specifically, the target FMA unit is determined based on the hardware resource usage of each FMA unit.

[0037] It's important to note that FMA is a hardware computing unit used to efficiently perform combined multiplication and addition operations, i.e., calculating the expression A×B+C. Its structure mainly consists of a multiplier, an adder, an intermediate result processing module, and a rounding and normalization module. Its working principle comprises six stages: input, multiplication, intermediate processing, addition, rounding and normalization, and output. First, it receives operands A, B, and C. The multiplier calculates the product of A×B. The intermediate result processing module aligns the product with C to retain high precision. The adder adds the two to generate the final result. The rounding and normalization module ensures the result conforms to floating-point standards. Finally, the result is output. FMA avoids rounding errors in intermediate results through fused operations, reducing the number of instructions and computational latency, significantly improving computational accuracy, performance, and energy efficiency.

[0038] It should also be noted that the FMA unit in this embodiment is a computing unit built on a Hardware Description Language (HDL) and supports scheduling. This embodiment does not limit the specific hardware description language used; for example, Verilog can be used.

[0039] Furthermore, this embodiment does not restrict the specific process of determining the target FMA unit based on the hardware resource usage of each FMA unit. However, considering the need to improve resource utilization and computational efficiency, idle FMA units should be selected as the target FMA units.

[0040] S13: Establish a mapping relationship between the computation task and the target fusion multiply-add operation unit, and dispatch the computation task to the target fusion multiply-add operation unit so that the target fusion multiply-add operation unit can execute the computation task.

[0041] Finally, a mapping relationship is established between the computational tasks and the target FMA unit, and the computational tasks are dispatched to the target FMA unit so that the target FMA unit can execute the computational tasks. This embodiment does not limit the specific process of establishing the mapping relationship; it depends on the specific implementation.

[0042] In this embodiment, an FMA unit is designed using a hardware description language, supporting direct hardware scheduling capabilities. When multiple computing requests are received, the corresponding computing tasks are stored in a hardware queue, and the scheduling information for each computing task is determined. The scheduling information includes hardware scheduling and software scheduling, meaning this solution supports both hardware and software scheduling methods. When the scheduling information for a computing task is hardware scheduling, the target FMA unit is determined based on the hardware resource usage of each FMA unit, and the computing task is dispatched to the target FMA unit so that the target FMA unit can execute the computing task. This achieves task balancing based on FMA unit resources, improving computing efficiency and FMA resource utilization.

[0043] Based on the above embodiments, in some embodiments, each computing request is combined with corresponding computing data to generate a computing task, including:

[0044] S101: Obtain the urgency information and task number from the computation request.

[0045] S102: Combine the urgency information, task number, and calculation data corresponding to the calculation request to generate a calculation task.

[0046] To generate a computation task, the urgency information and task number from the computation request are first obtained. It's important to note that the urgency information indicates the urgency level of the computation task. For example, when the urgency information is 2 bits in size, there are four possible urgency levels: 0, 1, 2, and 3, with urgency increasing sequentially. The task number indicates the order in which the computation tasks were received, and its size is typically 6 bits. Finally, the urgency information, task number, and the computation data corresponding to the computation request are combined to generate the computation task. The computation data is typically 64 bits in size. This completes the generation of the computation task, facilitating subsequent task storage and processing.

[0047] Based on the above embodiments, in some embodiments, the scheduling information for each computing task is determined, including:

[0048] S111: Determine the configuration bit information of the preset 32-bit register.

[0049] The lower 16 bits of the configuration information represent software scheduling, and each bit in the lower 16 bits is mapped to the corresponding fused multiply-accumulate operation unit; the higher 16 bits of the configuration information represent hardware scheduling, and each bit in the lower 16 bits is mapped to the corresponding fused multiply-accumulate operation unit.

[0050] S112: Obtain the target configuration bit indicated by the user, and determine the scheduling information of the computing task based on the target configuration bit and the configuration bit information.

[0051] To determine the scheduling information for each computing task, it is first necessary to determine the configuration bit information of the preset 32-bit register of the pipeline scheduler.

[0052] Table 1 Configuration Bit Information Table

[0053]

[0054] As shown in Table 1, the lower 16 bits of the configuration information represent software scheduling, and each bit in the lower 16 bits is mapped to the corresponding FMA unit; the higher 16 bits of the configuration information represent hardware scheduling, and each bit in the lower 16 bits is mapped to the corresponding FMA unit.

[0055] It should be noted that, as shown in the above embodiments, the parallel scheduling module supports a maximum of 16 FMA units. Therefore, the number of FMA units supporting both hardware and software scheduling in Table 1 is also 16. That is, each FMA unit can support both hardware and software scheduling. To determine which scheduling method to use, it is necessary to obtain the target configuration bit indicated by the user and determine the scheduling information of the computing task based on the target configuration bit and the configuration bit information. For example, if the user indicates that the target configuration bit is 8 bits, then software scheduling is confirmed, and the FMA units used subsequently are 8-bit mapped FMA units; if the user indicates that the target configuration bit is 21 bits, then hardware scheduling is confirmed, and the FMA units used subsequently are 21-bit mapped FMA units.

[0056] In this embodiment, scheduling selection is achieved through the configuration bit information of the preset 32-bit register of the pipeline scheduler, which supports both hardware and software scheduling and improves the flexibility of computing task execution.

[0057] Figure 4 The scheduling flowchart of the hardware pipeline scheduler provided in the embodiment of the present invention is shown. Figure 4 As shown, the dashed line represents the hardware scheduling task implementation, and the solid line represents the combined hardware and software scheduling implementation. Therefore, based on the above embodiments, in some embodiments, the target fusion multiply-accumulate operation unit is determined according to the hardware resource usage of each fusion multiply-accumulate operation unit, including:

[0058] S121: Obtain the resource allocation and release status of each integrated multiply-accumulate operation unit.

[0059] S122: Determine the idle status of the corresponding fusion multiply-accumulate operation unit based on the resource allocation and release status.

[0060] S123: Establish the mapping priority of each fused multiply-accumulate operation unit according to each idle status; wherein, the mapping priority of the fused multiply-accumulate operation unit is positively correlated with the corresponding idle status.

[0061] S124: Determine the target fusion multiply-accumulate operation unit in each fusion multiply-accumulate operation unit according to the priority of each mapping.

[0062] The number of target fusion multiply-accumulate operation units is equal to the number of computation tasks.

[0063] After determining to use hardware scheduling for the computing tasks, it is necessary to identify the target FMA units. Specifically, this involves obtaining the resource allocation and release information for each FMA unit. The resource allocation information refers to which tasks or resources are currently allocated to the FMA unit, and the occupancy status of these tasks or resources; while the resource release information refers to which resources are released after the FMA unit completes its tasks, and the current availability status of these resources. Both together reflect the load and resource utilization of the FMA unit.

[0064] Furthermore, by analyzing the resource allocation and release status of the FMA unit, the current idle status of the FMA unit can be determined, i.e., its ability to execute new tasks or resource availability can be assessed. This helps in effective task scheduling and load balancing.

[0065] Subsequently, a mapping priority is established for each FMA unit based on its idle status. It's important to note that the mapping priority of an FMA unit is positively correlated with its idle status; that is, the more idle an FMA unit is, the higher its mapping priority, and the more likely it is to establish a mapping with a computation task. In other words, in practical implementation, computation tasks are preferentially mapped to idle or nearly idle FMAs, and the resource allocation is refreshed. Simultaneously, the task is dispatched to the mapped FMA.

[0066] Finally, the target FMA unit is determined in each FMA unit according to the mapping priority. In this embodiment, there is no limit to the specific number of target FMA units, but it must be ensured that the number of target FMA units is equal to the number of computation tasks. For example, when there are 3 computation tasks, the number of target FMA units is also 3, and these 3 target FMA units should be the highest mapping priority among all available FMA units.

[0067] In this embodiment, by determining the idle status of FMA units, a mapping priority for each FMA unit is established based on each idle status. This facilitates the determination of idle or nearly idle target FMA units for processing computing tasks based on each mapping priority, thereby improving task processing efficiency and resource utilization and achieving load balancing of FMA units.

[0068] Based on the above embodiments, in some embodiments, a mapping relationship is established between the computation task and the target fusion multiply-accumulate operation unit, including:

[0069] S131: Determine the allocation priority of each computing task based on the urgency information of each computing task.

[0070] S132: Based on the mapping priority of each target fusion multiply-accumulate operation unit and the allocation priority of each computing task, establish the mapping relationship between each computing task and each target fusion multiply-accumulate operation unit in a priority order isomorphic manner.

[0071] Each computational task corresponds one-to-one with a multiply-accumulate unit that integrates various objectives.

[0072] To establish the mapping relationship between computational tasks and FMAs, in addition to considering the load of FMA units, the urgency of the computational tasks also needs to be considered. Specifically, firstly, the allocation priority of each computational task is determined based on its urgency information. Then, based on the mapping priority of each target FMA unit and the allocation priority of each computational task, the mapping relationship between each computational task and each target FMA unit is established in a priority-order isomorphic manner.

[0073] It's important to note the isomorphic priority order, meaning each computation task is mapped one-to-one with each target FMA unit according to its priority. For example, given three computation tasks, task 1 has a higher allocation priority than task 2, and task 2 has a higher allocation priority than task 3; and three target FMA units, target FMA unit 1 has a higher mapping priority than target FMA unit 2, and target FMA unit 2 has a higher mapping priority than target FMA unit 3. Therefore, when mapping computation tasks to target FMA units, a mapping relationship is specifically established between computation task 1 and target FMA unit 1, between computation task 2 and target FMA unit 2, and between computation task 3 and target FMA unit 3. This ensures that computation tasks with higher urgency are processed first by idle FMA units, significantly improving task processing efficiency.

[0074] The following is a detailed explanation of the specific process for determining the allocation priority of each computing task:

[0075] Figure 5 This is a schematic diagram of a hardware queue structure provided in an embodiment of the present invention. Figure 5 As shown, the hardware queue employs an enhanced First In, First Out (FIFO) implementation, which involves adding an extra layer of design around the dual-port random access memory (DPRAM) inside the general-purpose FIFO. Specifically, the allocation priority of each computational task is determined based on its urgency, including:

[0076] S141: Based on the urgency information of each computing task, each computing task is divided into multiple task types; among them, the urgency information of each computing task in the same task type is the same.

[0077] S142: Determine whether there is a target task type in each task type with a corresponding number of calculation tasks of 3; if yes, proceed to step S143; if no, proceed to step S145.

[0078] S143: Set the allocation priority of each computation task in the target task type to be higher than the allocation priority of each computation task in other task types; wherein, the allocation priority of each computation task in the target task type is the same.

[0079] S144: Determine the allocation priority of each calculation task in the remaining task types based on the numerical value of the corresponding urgency information.

[0080] S145: Determine the allocation priority of each calculation task in each task type based on the numerical value of the corresponding urgency information.

[0081] First, based on the urgency information of each computational task, the tasks are divided into several task types. As shown in the above embodiment, the urgency information is 2 bits in size, and there are four urgency levels: urgency 0, urgency 1, urgency 2, and urgency 3, with urgency increasing sequentially. Therefore, based on the corresponding urgency information, the computational tasks can also be divided into four task types. It is understood that the urgency information of each computational task within the same task type is the same. Next, it is determined whether there exists a target task type with 3 corresponding computational tasks within each task type; that is, whether the number of computational tasks in each task type is 3.

[0082] If it is confirmed that there are 3 computational tasks in the target task type, then the allocation priority of each computational task in the target task type is set higher than the allocation priority of each computational task in the other task types, regardless of the relative urgency of each computational task in the target task type and the urgency of each computational task in the other task types. It should also be noted that the allocation priority of each computational task in the target task type is the same. For each computational task in the other task types, the allocation priority needs to be determined based on the numerical value of the corresponding urgency information.

[0083] For example, if the number of computation tasks with urgency level 2 reaches 3, then the allocation priority of computation tasks with urgency level 2 is higher than the allocation priority of computation tasks with urgency levels 0, 1, and 3; the allocation priority of the three computation tasks with urgency level 2 is equal. However, the allocation priority of computation tasks with urgency levels 0, 1, and 3 depends on the numerical value of the corresponding urgency information; that is, the allocation priority of computation tasks with urgency level 3 is higher than the allocation priority of computation tasks with urgency level 1, and the allocation priority of computation tasks with urgency level 1 is higher than the allocation priority of computation tasks with urgency level 0.

[0084] If it is confirmed that no task type has 3 computational tasks, then the default FIFO design is followed, and the allocation priority of each computational task in each task type is determined according to the numerical value of the corresponding urgency information. For example, the allocation priority of a computational task with urgency level 3 is higher than that of a computational task with urgency level 2, the allocation priority of a computational task with urgency level 2 is higher than that of a computational task with urgency level 1, and the allocation priority of a computational task with urgency level 1 is higher than that of a computational task with urgency level 0.

[0085] This ensures that computational tasks with the same urgency level reaching the threshold are mapped and processed first, thus improving the processing efficiency of computational tasks.

[0086] It should also be noted that when dividing the computational tasks into multiple task types, there may be multiple task types with a total of 3 computational tasks, meaning there may be multiple target task types; or even all task types with a total of 3 computational tasks, meaning all task types are target task types. To address this situation, based on the above embodiments, in some embodiments, after confirming the existence of a target task type, the following further steps are taken:

[0087] S151: Determine whether there are multiple target task types; if not, proceed to step S143; if yes, proceed to step S152.

[0088] S152: Based on the numerical value of the corresponding urgency information, determine the allocation priority of each calculation task among all target task types, and proceed to step S144.

[0089] Specifically, after confirming the existence of a target task type, the number of target task types is verified to determine if there are multiple target task types. If it is confirmed that there is only one target task type, the process described in the above embodiment can be followed, proceeding to the step of setting the allocation priority of each computing task in the target task type to be higher than the allocation priority of each computing task in the other task types.

[0090] If multiple target task types are confirmed, the allocation priority of each computational task among all target task types is determined based on the numerical value of the corresponding urgency information. Subsequent processes follow the workflow described in the above embodiment, proceeding to the step of determining the allocation priority of each computational task among the remaining task types based on the numerical value of the corresponding urgency information. It is understood that when all task types are target task types, it is unnecessary to proceed to the step of determining the allocation priority of each computational task among the remaining task types based on the numerical value of the corresponding urgency information.

[0091] This ensures the accuracy and rationality of the process for determining the priority of computing tasks.

[0092] As can be seen from the above embodiments, this solution supports both hardware scheduling and software scheduling. Therefore, in some embodiments, when the scheduling information for the computing task is software scheduling, it further includes:

[0093] S161: Obtain a pre-built lookup table; wherein the lookup table contains a mapping between task numbers in the hardware queue and dual-port random access memory addresses.

[0094] S162: Determine the corresponding target dual-port random access memory address in the lookup table based on the task number in the computation request.

[0095] S163: Read the computation task under the target dual-port random access memory address and proceed to step S13.

[0096] First, obtain the pre-built lookup table.

[0097] Table 2 Lookup Table

[0098]

[0099] As shown in Table 2, the lookup table contains the mapping relationship between task numbers in the hardware queue and DPRAM addresses. Since the floating-point data (computation data) in the hardware queue is 32 bits, and a typical FMA calculation requires 96 bits (3 floating-point data entries) encoded with 8 bits, the address span is decimal 12. The software can use the task number to find the specific address of the floating-point data located in the DPRAM within the hardware queue; thus, it can read the computation tasks in the hardware queue in advance.

[0100] Therefore, if the user selects software scheduling, and the software deems a task needing to be processed in advance, it first retrieves a pre-built lookup table and determines the corresponding target DPRAM address in the lookup table based on the task number in the computation request. The software then reads the computation task under the target DPRAM through the DPRAM interface of the hardware queue and proceeds to the step of establishing a mapping relationship between the computation task and the target fused multiply-accumulate unit. In this way, software scheduling of FMA computation tasks is achieved, supporting the priority processing of specified computation tasks.

[0101] Figure 6 This is a timing diagram for multiply-accumulator scheduling provided in an embodiment of the present invention. Based on the above embodiments, in some embodiments, after dispatching the computation task to the target fusion multiply-accumulator unit, the following is further included:

[0102] S171: Obtain the estimated multiply-accumulate clock cycle for the target fusion multiply-accumulate unit to process the computation task, and obtain the current multiply-accumulate clock cycle that the target fusion multiply-accumulate unit has executed in real time.

[0103] S172: Monitor the completion rate of the target fusion multiply-accumulate operation unit in processing computational tasks based on the estimated multiply-accumulate clock cycle and the real-time current multiply-accumulate clock cycle.

[0104] like Figure 6 As shown, computation requests and computation data enter the hardware queue through the bus interface. The hardware queue then sends the task and data to the hardware scheduler / software scheduler (i.e., the parallel scheduler). After a task is dequeued but before scheduling, the resource map refreshes the current resource status. The scheduler obtains the resource usage from the resource map, selects available FMA hardware computing resources, and then dispatches the task to the FMA unit.

[0105] Based on this, in order to determine the completion progress of each computation task, this embodiment can obtain the estimated multiply-accumulate clock cycles for the target FMA unit to process the computation task, and obtain the current multiply-accumulate clock cycles already executed by the target FMA unit in real time. The completion rate of the computation task processed by the target FMA unit is monitored based on the estimated multiply-accumulate clock cycles and the real-time current multiply-accumulate clock cycles. For example, for a computation task, if the estimated multiply-accumulate clock cycles for the target FMA unit to process the task are 32 clock cycles, and the current multiply-accumulate clock cycles are 28 clock cycles, then the completion rate of the computation task is 28 / 30, with 4 cycles of work remaining. During the FMA computation process, the FMA unit notifies the resource mapping module of the current computation completion rate through feedback signals.

[0106] Finally, after the computation task is completed, a feedback signal is sent to the resource mapping module to update the resource status. The resource mapping module updates the resource status of each FMA in the FMA computation resource pool and feeds back the latest status to the scheduler. Through the above steps, the completion rate of each computation task can be clearly determined. The entire process ensures efficient task scheduling and dynamic resource management, thereby realizing the system's parallel computing capabilities.

[0107] After multiple FMA computation tasks are completed, the computation time varies because different FMA units are assigned different computation tasks and data. To monitor the execution status of each computation task in real time and ensure the orderly output of computation results, based on the above embodiments, in some embodiments, before dispatching the computation task to the target fusion multiply-accumulate unit, after establishing the mapping relationship between the computation task and the target fusion multiply-accumulate unit, the following steps are also included:

[0108] S181: Construct a hardware circular queue containing multiple entries.

[0109] S182: Based on the hardware circular queue and the corresponding task number, assign corresponding entries to each computing task.

[0110] S183: Monitor the execution status of each computation task based on the status bits of each entry.

[0111] Specifically, this embodiment also provides a data reordering cache module. This module is a hardware structure for supporting out-of-order execution. Its core function is to ensure that the order in which tasks are submitted is consistent with the order in which execution results are returned, while allowing computational tasks to be executed out of order during the process to improve performance.

[0112] The data reordering cache module is built on a hardware circular buffer. The hardware circular buffer contains multiple entries, each corresponding to a task. The status bits of the entries in the circular buffer include initialization bit, allocation bit, completion bit, and commit bit.

[0113] When a computation task establishes a dispatch relationship with an FMA unit, an entry is assigned to the corresponding computation task based on the task number; this entry is the one before execution. Since the status bits of each entry change during the execution of the computation task, the execution status of each computation task can be monitored based on the status bits of each entry to determine whether the computation task has been assigned, completed, or submitted.

[0114] Furthermore, based on the hardware circular queue, the computation results corresponding to the completed computation tasks can be reordered according to the task number and status bit when dequeued. Specifically, after the target fusion multiply-accumulate unit completes the computation task, it also includes:

[0115] S184: Obtain the calculation results corresponding to each calculation task.

[0116] S185: Controls the blocking movement of the read pointer according to each status bit, and controls the output of each calculation result in sequence according to the task number of the corresponding calculation task.

[0117] Specifically, if a computation task is completed, the status bit of the corresponding entry will be refreshed. At this time, the hardware circular queue will dequeue entries according to the order before execution; if the task is not completed, the circular queue will continue to refresh in a loop until the computation task that should be completed (i.e., the previous queue status bit shows a valid commit bit) completes and outputs its result, and then submits it uniformly according to the input queue order. The following example illustrates this:

[0118] Assume a hardware circular queue has three entries, all initially in the initialized state. Now, computation tasks A, B, and C are submitted sequentially, with their corresponding task numbers indicating the computation task order. When computation task A arrives, entry 0 is allocated, and its corresponding state bit is updated from the initialized state to the allocated state. When computation task B arrives, entry 1 is allocated, and its corresponding state bit is updated from the initialized state to the allocated state. When computation task C arrives, entry 2 is allocated, and its corresponding state bit is updated from the initialized state to the allocated state. During the execution of the three computation tasks, computation task A completes, and the state bit of entry 0 switches from the allocated state to the completed state. Computation task B is not completed, and entry 1 retains the allocated state. Computation task C completes, and the state bit of entry 2 switches from the allocated state to the completed state. At this point, the queue attempts to submit the results sequentially. Because computation task B is not completed, the entire process is blocked at entry 1, and even if computation tasks A and C are completed, they are not submitted. When computation task B completes, the state bit of entry 1 switches from the allocated state to the completed state. The hardware circular queue detects the change, updates the state bits of all entries to the submitted state, and outputs the corresponding computation results sequentially according to the task input order (task number) of computation tasks A, B, and C. Eventually, all entries return to their initial state and are ready to receive new tasks.

[0119] In this embodiment, the status bits of the hardware circular queue are used to control the blocking movement of the read pointer, forcing the calculation results of each task to be output sequentially according to the input order of its corresponding entry in the queue, thus ensuring the orderliness of the calculation results output.

[0120] Based on the above embodiments, in some embodiments, after controlling the output of each calculation result, the following is also included:

[0121] S191: Write each calculation result back to the storage address specified by the corresponding calculation task.

[0122] Specifically, after controlling the output of each calculation result, each calculation result is written back to a specified storage address such as a register or memory to ensure the final storage of the calculation task execution result and the correct data dependency of subsequent tasks. Its role is to maintain data consistency and program correctness.

[0123] Furthermore, in the process of obtaining the computation results corresponding to each computational task, result verification and error checking are necessary to ensure the correctness and consistency of the data. Specifically, verification methods include checksums, hash values, and redundant computation: checksums generate fixed-length values ​​through mathematical operations to compare whether the data is corrupted; hash values ​​use hash functions to generate unique identifiers to ensure that the data has not been tampered with; redundant computation involves calculating the same data multiple times and comparing the results to confirm correctness. If errors are found during verification, error detection codes such as parity checks and cyclic redundancy checks can be used to identify the errors, or error correction codes such as erasure codes can be used to automatically correct errors within a certain range. These methods effectively guarantee the reliability and consistency of data, which is crucial for fields such as high-performance computing, distributed systems, and data storage, improving system stability and data trustworthiness.

[0124] To enable those skilled in the art to better understand this solution, the following description is provided in conjunction with the appendix. Figure 7 The specific process of this plan is explained as follows:

[0125] Figure 7 This is a flowchart illustrating the workflow of the hardware scheduling and computing unit provided in an embodiment of the present invention. Figure 7 As shown, the hardware scheduling computation process is mainly divided into a scheduling phase and a computation phase, with the scheduling phase primarily driven by the parallel scheduler. Specifically, upon receiving a request, the data acquisition module retrieves floating-point data from memory or other storage devices, combines the computation request with the computation data, and stores it in the hardware queue. After entering the hardware queue, a soft / hard scheduling method is selected. Based on the scheduling policy and resource mapping, the computation unit for the load computation task is determined, and task dispatch is performed. After FMA parallel computation, data reordering is performed, followed by data submission and data write-back, thus completing the entire computation process.

[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0127] Figure 8 This is a schematic diagram of a computing task processing device provided in an embodiment of the present invention. Figure 8 As shown, the device includes:

[0128] The acquisition module 10 is used to acquire the computing data corresponding to each computing request when multiple computing requests are received, and to combine each computing request with the corresponding computing data to generate a computing task.

[0129] The first determining module 11 is used to store each computing task in a hardware queue and determine the scheduling information of each computing task; wherein the scheduling information includes hardware scheduling and software scheduling.

[0130] The second determining module 12 is used to determine the target fusion multiply-accumulate operation unit based on the hardware resource usage of each fusion multiply-accumulate operation unit when the scheduling information of the computing task is hardware scheduling; wherein, the fusion multiply-accumulate operation unit is a computing unit built based on a hardware description language and supports scheduling.

[0131] The task dispatch module 13 is used to establish a mapping relationship between the computation task and the target fusion multiply-accumulate operation unit, and dispatch the computation task to the target fusion multiply-accumulate operation unit so that the target fusion multiply-accumulate operation unit can execute the computation task.

[0132] In some embodiments, the acquisition module 10 includes:

[0133] The first acquisition submodule is used to acquire urgency information and task number from the calculation request;

[0134] The combination submodule is used to combine urgency information, task number, and calculation data corresponding to the calculation request to generate a calculation task.

[0135] In some embodiments, the first determining module 11 includes:

[0136] The configuration bit information determination module is used to determine the configuration bit information of a preset 32-bit register; wherein, the lower 16 bits of the configuration bit information represent software scheduling, and each bit in the lower 16 bits is mapped to the corresponding fused multiply-accumulate operation unit; the higher 16 bits of the configuration bit information represent hardware scheduling, and each bit in the lower 16 bits is mapped to the corresponding fused multiply-accumulate operation unit.

[0137] The target configuration bit acquisition module is used to acquire the target configuration bit indicated by the user and determine the scheduling information of the computing task based on the target configuration bit and the configuration bit information.

[0138] In some embodiments, the second determining module 12 includes:

[0139] The resource information acquisition module is used to acquire the resource allocation and release status of each integrated multiply-accumulate operation unit;

[0140] The idle status determination module is used to determine the idle status of the corresponding fused multiply-accumulate operation unit based on the allocation and release status of each resource.

[0141] The mapping priority establishment module is used to establish the mapping priority of each fused multiply-accumulate operation unit according to each idle status; wherein, the mapping priority of the fused multiply-accumulate operation unit is positively correlated with the corresponding idle status;

[0142] The target fusion multiply-accumulate operation unit determination module is used to determine the target fusion multiply-accumulate operation unit in each fusion multiply-accumulate operation unit according to the priority of each mapping.

[0143] The number of target fusion multiply-accumulate operation units is equal to the number of computation tasks.

[0144] In some embodiments, the task dispatch module 13 includes:

[0145] The priority allocation module is used to determine the allocation priority of each computing task based on the urgency information of each computing task.

[0146] The mapping relationship establishment module is used to establish the mapping relationship between each computing task and each target fusion multiply-accumulate operation unit according to the mapping priority of each target fusion multiply-accumulate operation unit and the allocation priority of each computing task in a priority order isomorphic manner.

[0147] Each computational task corresponds one-to-one with a multiply-accumulate unit that integrates various objectives.

[0148] In some embodiments, the priority determination module includes:

[0149] The task type determination module is used to classify each computing task into multiple task types based on the urgency information of each computing task; among them, the urgency information of each computing task in the same task type is the same.

[0150] The first judgment module is used to determine whether there is a target task type with a corresponding number of 3 computational tasks in each task type; if so, the allocation priority of each computational task in the target task type is set higher than the allocation priority of each computational task in other task types; wherein, the allocation priority of each computational task in the target task type is the same; the allocation priority of each computational task in other task types is determined according to the value of the corresponding urgency information; if not, the allocation priority of each computational task in each task type is determined according to the value of the corresponding urgency information.

[0151] In some embodiments, it also includes:

[0152] The second judgment module is used to determine whether there are multiple target task types; if not, it proceeds to the step of setting the allocation priority of each calculation task in the target task type to be higher than the allocation priority of each calculation task in the other task types; if so, it determines the allocation priority of each calculation task in all target task types according to the numerical value of the corresponding urgency information, and proceeds to the step of determining the allocation priority of each calculation task in the other task types according to the numerical value of the corresponding urgency information.

[0153] In some embodiments, it also includes:

[0154] The lookup table acquisition module is used to acquire a pre-built lookup table; the lookup table contains the mapping relationship between task numbers in the hardware queue and dual-port random access memory addresses;

[0155] The storage address determination module is used to determine the corresponding target dual-port random access memory address in a lookup table based on the task number in the computation request.

[0156] The reading module is used to read the computation tasks under the target dual-port random access memory address and proceed to the step of establishing the mapping relationship between the computation tasks and the target fused multiply-accumulate operation unit.

[0157] In some embodiments, it also includes:

[0158] The multiply-accumulate clock cycle acquisition module is used to acquire the estimated multiply-accumulate clock cycle of the target fusion multiply-accumulate operation unit for processing computation tasks, and to acquire the current multiply-accumulate clock cycle that the target fusion multiply-accumulate operation unit has executed in real time.

[0159] The completion monitoring module is used to monitor the completion rate of the computation task processed by the target fusion multiply-accumulate operation unit based on the estimated multiply-accumulate clock cycle and the real-time current multiply-accumulate clock cycle.

[0160] In some embodiments, it also includes:

[0161] The hardware circular queue building module is used to build hardware circular queues containing multiple entries;

[0162] The entry allocation module is used to allocate corresponding entries to each computing task based on the hardware circular queue and the corresponding task number.

[0163] The execution status monitoring module is used to monitor the execution status of each computing task based on the status bits of each entry.

[0164] In some embodiments, it also includes:

[0165] The calculation result acquisition module is used to acquire the calculation results corresponding to each calculation task;

[0166] The calculation result output module is used to control the blocking movement of the read pointer according to each status bit, and to control the output of each calculation result in sequence according to the task number of the corresponding calculation task.

[0167] In some embodiments, it also includes:

[0168] The write-back module is used to write the calculation results back to the storage address specified by the corresponding calculation task.

[0169] For a description of the features in the embodiment corresponding to the computing task processing device, please refer to the relevant description in the embodiment corresponding to the computing task processing method, which will not be repeated here.

[0170] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the computing task processing method.

[0171] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the computing task processing method when it is run.

[0172] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0173] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the computing task processing method.

[0174] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the computing task processing method.

[0175] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0176] The present invention has provided a detailed description of a computing task processing method, device, medium, and product. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of these embodiments are merely for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for processing computational tasks, characterized in that, include: When multiple computing requests are received, the computing data corresponding to each computing request is obtained, and each computing request is combined with the corresponding computing data to generate a computing task; Each computing task is stored in a hardware queue, and scheduling information for each computing task is determined; wherein, the scheduling information includes hardware scheduling and software scheduling. When the scheduling information of the computing task is the hardware scheduling, the target fusion multiply-accumulate computing unit is determined according to the hardware resource usage of each fusion multiply-accumulate computing unit; wherein, the fusion multiply-accumulate computing unit is a computing unit built based on a hardware description language and supports scheduling; A mapping relationship is established between the computation task and the target fusion multiply-accumulate operation unit, and the computation task is dispatched to the target fusion multiply-accumulate operation unit so that the target fusion multiply-accumulate operation unit can execute the computation task; Determine the scheduling information for each of the aforementioned computing tasks, including: The configuration bit information of a preset 32-bit register is determined; wherein, the lower 16 bits of the configuration bit information represent software scheduling, and each bit in the lower 16 bits is mapped to the corresponding fused multiply-accumulate operation unit; the higher 16 bits of the configuration bit information represent hardware scheduling, and each bit in the higher 16 bits is mapped to the corresponding fused multiply-accumulate operation unit. Obtain the target configuration bit indicated by the user, and determine the scheduling information of the computing task based on the target configuration bit and the configuration bit information; The target fused multiply-accumulate unit is determined based on the hardware resource usage of each fused multiply-accumulate unit, including: Obtain the resource allocation and release status of each of the fused multiply-accumulate operation units; Based on the resource allocation and release status, determine the idle status of the corresponding fusion multiply-accumulate operation unit; A mapping priority is established for each of the fused multiply-accumulate operation units based on each of the aforementioned idle conditions; wherein, the mapping priority of the fused multiply-accumulate operation unit is positively correlated with the corresponding idle condition; The target fusion multiply-accumulate operation unit is determined in each of the fusion multiply-accumulate operation units according to the mapping priority. The number of target fusion multiply-accumulate operation units is equal to the number of computation tasks. When the scheduling information for the computing task is the software scheduler, it further includes: Obtain a pre-built lookup table; wherein the lookup table contains a mapping relationship between task numbers in the hardware queue and dual-port random access memory addresses; Based on the task number in the computation request, determine the corresponding target dual-port random access memory address in the lookup table; The computation task under the target dual-port random access memory address is read, and the step of establishing the mapping relationship between the computation task and the target fusion multiply-accumulate operation unit is entered.

2. The computational task processing method according to claim 1, characterized in that, Combining each of the aforementioned computation requests with the corresponding computation data to generate a computation task includes: Obtain the urgency information and task number from the computation request; The urgency information, the task number, and the computation data corresponding to the computation request are combined to generate the computation task.

3. The computational task processing method according to claim 1, characterized in that, Establishing the mapping relationship between the computational task and the target fusion multiply-accumulate operation unit includes: The allocation priority of each computing task is determined based on the urgency information of each computing task; Based on the mapping priority of each target fusion multiply-accumulate operation unit and the allocation priority of each computing task, the mapping relationship between each computing task and each target fusion multiply-accumulate operation unit is established in a priority order isomorphic manner. Each of the aforementioned computational tasks corresponds one-to-one with each of the aforementioned target fusion multiplication and addition operation units.

4. The computational task processing method according to claim 3, characterized in that, The allocation priority of each computing task is determined based on its urgency information, including: Based on the urgency information of each computing task, each computing task is divided into multiple task types; wherein, the urgency information of each computing task within the same task type is the same; Determine whether there exists a target task type among the task types that corresponds to a number of 3 computational tasks; If so, the allocation priority of each computing task in the target task type is set to be higher than the allocation priority of each computing task in the other task types; wherein, the allocation priority of each computing task in the target task type is the same; Based on the numerical value of the corresponding urgency information, the allocation priority of each computing task in the remaining task types is determined; If not, the allocation priority of each computing task in each task type is determined based on the numerical value of the corresponding urgency information.

5. The computational task processing method according to claim 4, characterized in that, After confirming the existence of the target task type, the following is also included: Determine whether the number of the target task types is multiple; If not, proceed to the step of setting the allocation priority of each computing task in the target task type to be higher than the allocation priority of each computing task in the other task types; If so, then based on the numerical value of the corresponding urgency information, determine the allocation priority of each computing task in all target task types, and proceed to the step of determining the allocation priority of each computing task in the remaining task types based on the numerical value of the corresponding urgency information.

6. The computational task processing method according to claim 1, characterized in that, After dispatching the computation task to the target fusion multiply-accumulate unit, the method further includes: The estimated multiply-accumulate clock cycle for the target fusion multiply-accumulate operation unit to process the computation task is obtained, and the current multiply-accumulate clock cycle already executed by the target fusion multiply-accumulate operation unit is obtained in real time. The completion rate of the computation task processed by the target fusion multiply-accumulate operation unit is monitored based on the estimated multiply-accumulate clock cycle and the real-time current multiply-accumulate clock cycle.

7. The computational task processing method according to any one of claims 1 to 6, characterized in that, Before dispatching the computation task to the target fusion multiply-accumulate unit, after establishing the mapping relationship between the computation task and the target fusion multiply-accumulate unit, the method further includes: Construct a hardware circular queue containing multiple entries; Based on the hardware circular queue and the corresponding task number, a corresponding entry is assigned to each computing task; The execution status of each computational task is monitored based on the status bits of each entry.

8. The computational task processing method according to claim 7, characterized in that, After the target fusion multiply-accumulate unit completes the calculation task, the following steps are also included: Obtain the calculation results corresponding to each of the aforementioned calculation tasks; The reading pointer is controlled to move in a blocking manner according to the status bits, and the calculation results are output sequentially according to the task number of the corresponding calculation task.

9. The computational task processing method according to claim 8, characterized in that, After controlling the output of each of the aforementioned calculation results, the following is also included: The calculation results are written back to the storage address specified for the corresponding calculation task.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the computing task processing method as described in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the computing task processing method as described in any one of claims 1 to 9.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the computational task processing method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Task scheduling method and device, equipment and medium

    CN119759543A

  • Hardware instructions to accelerate table-driven mathematical function evaluation

    US20110296146A1