Task processing method of artificial intelligence processor, storage medium and electronic equipment

By integrating task instructions into data packets in an artificial intelligence processor and dynamically allocating according to task priority and computing unit load status, the problem of low task execution efficiency in the prior art is solved, and more efficient and flexible task scheduling and execution are achieved.

CN120066806AActive Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510555198.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

When dynamically adjusting task allocation, existing artificial intelligence processors need to rely on task execution sequence to decode task instructions, resulting in long processing time and low task execution efficiency.

Method used

By obtaining the target tasks in the processor's task queue, integrating task instructions of the same operation type into data packets based on the data dependencies between target tasks, obtaining the task priority and real-time load status of the calculation unit, and allocating the data packet to the target computing unit according to the priority and load status.

Benefits of technology

It reduces processing time, improves task execution efficiency, solves the problem of uneven resource utilization caused by static task allocation, achieves more efficient and flexible task scheduling and execution, and improves cache hit rate, alleviates memory bandwidth bottlenecks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066806A_ABST
    Figure CN120066806A_ABST
Patent Text Reader

Abstract

The invention discloses a task processing method of an artificial intelligence processor, a storage medium and electronic equipment, and relates to the technical field of computers.The method comprises the steps that a target task in a task queue of the processor is obtained; based on the data dependency relationship between the target tasks, integrating task instructions which can be executed by the target tasks and have the same operation type into a data packet; obtaining a task priority of the target task and real-time load states of a plurality of computing units in a processor; and according to the task priority and the real-time load state, distributing the data packet to a target computing unit in the plurality of computing units. Compared with the prior art, the method has the advantages that a data flow triggering mechanism and a dynamic resource allocation strategy are fused, so that the requirement of decoding task instructions one by one according to a task execution sequence is avoided, the processing time is shortened, the task execution efficiency is improved, and meanwhile, the problem of non-uniform resource utilization caused by static task allocation is effectively solved; and more efficient and flexible task scheduling and execution are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method for processing tasks of an artificial intelligence processor, a storage medium, and an electronic device. Background Art

[0002] With the development of artificial intelligence technologies, the requirements for computing performance are getting higher and higher. From simple machine learning algorithms in the early days to complex deep learning models today, such as deep neural networks being widely used in fields such as computer vision, speech recognition, and natural language processing, these applications require powerful computing capabilities to process massive amounts of data and complex operations.

[0003] Currently, in related technologies, when an artificial intelligence processor dynamically adjusts task allocation, it usually needs to rely on the task execution order to decode task instructions, consuming a lot of time and resulting in low task execution efficiency. Summary of the Invention

[0004] The present disclosure provides a method for processing tasks of an artificial intelligence processor, a storage medium, an electronic device, and a program product. Its main purpose is to solve the problem that in related technologies, when an artificial intelligence processor dynamically adjusts task allocation, it usually needs to rely on the task execution order to decode task instructions, consuming a lot of time and resulting in low task execution efficiency.

[0005] In a first aspect, the present application provides a method for processing tasks of an artificial intelligence processor, including: Obtaining a target task in a task queue of a processor; Based on the data dependency relationships between the target tasks, integrating task instructions of the same operation type that the target tasks can execute into a data packet; Obtaining the task priority of the target task and the real-time load status of multiple computing units in the processor; According to the task priority and the real-time load status, allocating the data packet to a target computing unit among the multiple computing units, where the target computing unit is configured to execute a computing task corresponding to the target task based on the data packet.

[0006] In a second aspect, the present application provides a device for processing tasks of an artificial intelligence processor, including: A first obtaining module configured to obtain a target task in a task queue of a processor; An integrating module configured to integrate task instructions of the same operation type that the target tasks can execute into a data packet based on the data dependency relationships between the target tasks; A second obtaining module configured to obtain the task priority of the target task and the real-time load status of multiple computing units in the processor; The allocation module is configured to allocate the data packet to a target computing unit among the multiple computing units according to the task priority and the real-time load status, and the target computing unit is used to execute the computing task corresponding to the target task based on the data packet.

[0007] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method of the first aspect is implemented.

[0008] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the computer program, the method of the first aspect is implemented.

[0009] In a fifth aspect, the present application provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, the method of the first aspect is implemented.

[0010] The task processing method, storage medium, electronic device, and program product of the artificial intelligence processor provided by the present disclosure, wherein the method includes: first, obtaining a target task in the task queue of the processor; integrating task instructions of the same operation type that can be executed by the target task into a data packet based on the data dependency relationship between the target tasks; obtaining the task priority of the target task and the real-time load status of multiple computing units in the processor; and allocating the data packet to a target computing unit among the multiple computing units according to the task priority and the real-time load status, wherein the target computing unit is used to execute the computing task corresponding to the target task based on the data packet. Compared with the current existing technologies, the present application combines the data flow trigger mechanism with the dynamic resource allocation strategy, avoids the need to decode task instructions one by one according to the task execution order, reduces the processing time, improves the task execution efficiency, effectively solves the problem of uneven resource utilization caused by static task allocation, realizes more efficient and flexible task scheduling and execution, and can also significantly improve the cache hit rate and relieve the memory bandwidth bottleneck.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. Description of the Drawings

[0012] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 The figure shows a schematic flowchart of a task processing method for an artificial intelligence processor provided by an embodiment of the present application; Figure 2 The figure shows a schematic structural diagram of an example provided by an embodiment of the present application; Figure 3 The figure shows a schematic flowchart of another task processing method for an artificial intelligence processor provided by an embodiment of the present application; Figure 4 The figure shows a schematic structural diagram of a task processing device for an artificial intelligence processor provided by an embodiment of the present application. Detailed implementation manners

[0014] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0015] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0016] Currently, the field of artificial intelligence is developing rapidly, posing extremely high requirements for computing performance. From early simple machine learning algorithms to today's complex deep learning models, such as deep neural networks being widely used in fields such as computer vision, speech recognition, and natural language processing, these applications require powerful computing capabilities to process massive amounts of data and complex operations. Traditional processor architectures are stretched thin when dealing with artificial intelligence loads, which has prompted the research and development of dedicated processors for artificial intelligence applications. The field of artificial intelligence chips is constantly evolving, gradually evolving from general-purpose graphics processing units (GPUs) to more efficient dedicated artificial intelligence processors, such as Google's Tensor Processing Unit (TPU), etc. In such a development trend, how to better combine processor technology with artificial intelligence applications and improve the performance of artificial intelligence processors in aspects such as data processing, task scheduling, and execution has become an important research direction in the current technical field.

[0017] In the development stage of artificial intelligence, general-purpose CPUs are the main computing resources, but because their architecture is not optimized for artificial intelligence loads, they are less efficient when processing large-scale data and complex neural network operations. With the development of technology, GPUs have become the main platform for deep learning training with their powerful parallel computing capabilities, accelerating high-parallel tasks such as matrix operations through a large number of computing cores. In addition, there are dedicated artificial intelligence chips such as application-specific integrated circuits (ASICs) customized for specific neural network structures or algorithms, which provide efficient computing performance in specific application scenarios. At the same time, traditional data flow processors adopt the design concept of data flow drive combined with control flow, merge data flow and control flow through explicit data transmission requirements, and use data trigger unit computing cores to support parallel execution of various granularities to meet the needs of computing-intensive and data-intensive applications. Data flow processors also use storage structures including instruction cache, direct memory access (DMA) controller and local memory to reduce memory access overhead, and use on-chip interconnect network architecture to achieve efficient communication between multiple high-performance cores and synchronization mechanisms based on atomic instructions to support multi-core synchronization. In terms of task scheduling, traditional strategies usually adopt static task allocation, that is, determining the mapping relationship between tasks and processing units before task execution; in data stream processors, tasks are dynamically executed according to the order in which data is ready. Although this method may have a long data dependency chain for tasks with high real-time requirements, it will increase delays and affect the satisfaction of real-time requirements.

[0018] Traditional artificial intelligence processor architectures have various deficiencies: Due to the generality of their design, when a general-purpose CPU processes artificial intelligence workloads, a large amount of computing resources are idle, resulting in low hardware utilization; Although GPUs perform excellently in parallel computing, they face problems such as high power consumption and high memory bandwidth pressure. Especially in deep learning training, a large amount of data is frequently read from memory, making the memory bandwidth a performance bottleneck; Although ASICs are extremely efficient in specific application scenarios, due to the lack of generality, they need to redesign the chip when facing changes in neural network structures or algorithms, with poor flexibility. Traditional data flow processor architectures also have defects: In terms of hardware utilization, static task allocation easily causes some processing units to be overloaded while other processing units are idle, resulting in resource waste; The random memory access pattern triggered by fine-grained data flow reduces the cache hit rate and increases the pressure on the memory bandwidth; The method relying on static task partitioning has low resource utilization due to unbalanced data dependencies, and the execution time depends on the dynamic data ready order, making it difficult to meet real-time requirements. In addition, explicit data flow programming requires manual management of dependencies, which is difficult to debug and increases the development cost and difficulty. For the task scheduling and execution methods of traditional data flow processors, the static task allocation strategy cannot be dynamically adjusted according to the runtime system state and cannot fully utilize system resources. For example, when one core is overloaded while other cores are idle, tasks cannot be automatically migrated to balance the load. The method of executing tasks depending on the data ready order may cause critical tasks to be delayed in execution and cannot meet the requirements of real-time applications. Traditional methods lack comprehensive consideration of task priorities, data locality, and system resource status, so it is difficult to achieve efficient task scheduling and execution in complex artificial intelligence application scenarios. In summary, these shortcomings limit the effectiveness and adaptability of existing architectures in processing modern artificial intelligence workloads.

[0019] To improve the problem that in related technologies, when an artificial intelligence processor dynamically adjusts task allocation, it usually needs to rely on the task execution order to decode task instructions, which consumes a lot of time and results in low task execution efficiency. This embodiment provides a task processing method for an artificial intelligence processor, as Figure 1 shown, the method includes the following steps: Step 101, obtain a target task in the task queue of the processor.

[0020] Exemplarily, the processor can monitor target tasks whose all input dependencies are satisfied in the task queue, where the target tasks are stored in the task queue after being sorted according to their priorities.

[0021] Step 102, based on the data dependency relationships between the target tasks, integrate task instructions of the same operation type that can be executed by the target tasks into data packets.

[0022] In some examples, the main processor distributes tasks to the task queues of processing units (PUs) through a global scheduling mechanism; subsequently, these tasks are integrated and packaged into data packets, and dispatched to the target PUs according to the requirements of the tasks. Inside the PU, dependency checking is performed through a local task scheduling method, and instructions of the same operation type with ready data are packaged into a data packet.

[0023] Exemplarily, by integrating instructions of the same type into data packets (supporting unified integration of multiple tasks), optimized allocation of resources and improvement of task execution efficiency are achieved. This not only reduces the resource consumption for task startup, but also improves the utilization rate of computing resources by batch-processing similar tasks, ensuring efficient task scheduling and execution.

[0024] Step 103: Obtain the task priority of the target task and the real-time load status of multiple computing units in the processor.

[0025] Exemplarily, the priority of the target task can be calculated based on the criticality score and the estimated execution time, which is used to reflect the importance and urgency of the task. At the same time, by monitoring the load status of multiple computing units in the processor in real time (such as matrix, vector, floating-point, fixed-point units, etc.), including indicators such as computing utilization rate, bus bandwidth utilization rate, and queue depth, the busy degree of each computing unit can be evaluated. By integrating the task priority of the target task and the real-time load status of multiple computing units in the processor, the processor can select the target task with the highest priority and most suitable for the current load condition for allocation, ensuring that critical tasks can be efficiently executed on computing units with sufficient resources, thereby optimizing the overall task processing efficiency and resource utilization rate.

[0026] Step 104: Allocate the data packet to the target computing unit among multiple computing units according to the task priority and the real-time load status.

[0027] Exemplarily, the target computing unit is used to execute the computing task corresponding to the target task based on the data packet. When the target computing unit receives the data packet, it will automatically trigger the corresponding computing module to perform specific operations. After the calculation is completed, the result will be fed back to the data packet generation and dispatch unit, which uses these results to update the task status and guide the subsequent task scheduling process. This ensures that tasks can be executed efficiently and orderly, while optimizing resource allocation and load balancing, especially suitable for complex artificial intelligence computing tasks.

[0028] In some examples, the target task with the highest priority and most suitable for the current load of the processing unit can be dynamically selected according to the current resource status and load situation of the system for allocation and execution, ensuring that each task in the task queue can be processed efficiently and orderly, and optimizing the overall resource utilization and task execution efficiency.

[0029] In some examples, such as Figure 2 A high-performance data-triggered multi-core artificial intelligence processor architecture shown consists of four main parts. First, the General Purpose Processor (GPP) selects common types including ARM, RISC-V, SPARC, etc. Its core responsibility is to run the operating system and execute various control tasks closely related to applications. During the system operation, the GPP undertakes the processing of system-level key tasks, such as promptly responding to interrupt events and precisely controlling various external devices, etc., to ensure the stable operation and efficient cooperation of the entire system. In the artificial intelligence application scenario, the GPP, with the help of specific framework tools, such as TensorFlow XLA, PyTorch Glow, etc., decomposes the artificial intelligence training program into a Data Flow Graph (DFG) according to the unique characteristics of the artificial intelligence application. In this data flow graph, nodes represent specific tasks, and edges represent the dependency relationships between tasks. After completing the task decomposition, the GPP is responsible for reasonably allocating these tasks to each processing unit. The specific operation process is as follows: First, the GPP creates an initialized thread space to build a basic environment for the subsequent task execution; then, it starts the corresponding processing units to enable the tasks to be executed orderly on each PU, thereby achieving the efficient processing of the artificial intelligence training program. Secondly, the Data Triggered Multi-core Accelerator (DTMA) processing unit mainly undertakes the key responsibility of computing acceleration in the system. When the GPP creates a thread, it will reasonably schedule the tasks to the PUs for execution according to the established scheduling strategy, and a task is completed on one PU. During the operation, the PUs interact with other PUs through methods such as mailbox communication and direct memory access transfer to achieve data sharing and collaborative processing.

[0030] Exemplarily, the above PU unit is optimized for the specific computing requirements of artificial intelligence applications and integrates a large number of functional modules for matrix calculation, floating-point operation, fixed-point operation, and trigonometric function operation. These computing modules are all equipped with independent data stream trigger modules, input queues, and output queues, and can orderly receive and process data. The DTMA processing unit is the core computing unit of the data-triggered multi-core artificial intelligence processor, and its architecture design aims to efficiently execute artificial intelligence tasks and support dynamic task scheduling and data stream trigger mechanisms. The main components of the PU may include: computing modules (including matrix compute unit (MCU), vector compute unit (VCU), floating-point multiply-accumulate unit (floating-point MAC), floating-point arithmetic logic unit (floating-point ALU), fixed-point multiply-accumulate unit (fixed-point MAC), fixed-point arithmetic logic unit (fixed-point ALU), etc.), data stream trigger module, input / output queue, register file, memory interface (connected to high bandwidth memory, HBM), data packet generation and dispatch unit, and content-addressable memory (CAM) for quickly identifying and receiving data packets.

[0031] For example, in the processing unit, data packets are transmitted to the target computing unit through the PU bus, and content-addressable memory is used for parallel address comparison to ensure the accurate delivery of data packets. Once a match is successful, the data packet is received and key information such as operation type, task ID, instruction priority, and data length is extracted by parsing its packet header to determine the corresponding computing unit of the data packet, and the payload data is distributed to the input queue of the corresponding module in the order of the instruction array. Each computing module monitors the data readiness status of its input queue. When the conditions are met, it automatically triggers the execution of the corresponding computing operation (such as GEMM or vector addition), writes the result to the output queue, and marks the task ID and instruction ID for subsequent tracking. The computing result is fed back to the data packet generation and dispatch unit for updating the dependency relationship and scheduling logic; if the data required for subsequent tasks is ready, the task is marked as executable. In addition, the system adopts a dynamic priority scheduling strategy, periodically scans the task queue, adjusts the scanning period according to the PU load condition, and selects a suitable computing unit to dispatch data packets based on the real-time load to achieve dynamic load balancing and hardware acceleration. The register file supports fast data access, the high-bandwidth memory is uniformly addressed by the memory management unit to reduce memory access latency, the DMA controller supports direct memory access to efficiently transfer data, and the global interconnection network uses an optical variable topology on-chip interconnection network to support multi-node communication. Peripheral device interfaces such as PCI-E and JTAG provide the ability to connect to external devices.

[0032] Compared with the related technologies, in this embodiment, the target task in the task queue of the processor is first obtained; based on the data dependency relationship between the target tasks, the task instructions of the same operation type that the target tasks can execute are integrated into data packets; the task priority of the target tasks and the real-time load status of multiple computing units in the processor are obtained; according to the task priority and the real-time load status, the data packets are allocated to the target computing unit among the multiple computing units, where the target computing unit is used to execute the computing task corresponding to the target task based on the data packet. By integrating the data flow trigger mechanism and the dynamic resource allocation strategy, the need to decode task instructions one by one according to the task execution order is avoided, the processing time is reduced, the task execution efficiency is improved, and at the same time, the problem of uneven resource utilization caused by static task allocation is effectively solved, realizing more efficient and flexible task scheduling and execution, and can also significantly improve the cache hit rate and relieve the memory bandwidth bottleneck. By intelligently allocating critical path tasks, a more balanced task execution efficiency is achieved, bringing significant beneficial effects in multiple dimensions such as architecture design, storage management, and task scheduling.

[0033] Further, as a refinement and extension of the above embodiments, in order to specifically illustrate the data processing process, optionally, step 102 may specifically include: identifying, from the task instructions of the target task, task instructions that are executable, have the same operation type, and satisfy the data dependency relationship; integrating the identified task instructions to generate the data packet.

[0034] Exemplarily, during the local scheduling (within the PU) process, the data packet generation and dispatch unit is responsible for receiving the calculation results and performing dependency checks, integrating the task instructions into data packets in a specific format, and then dispatching them. When the DTMA processing unit receives the task information presented in the form of an instruction queue generated by the compiler mapping, it will comprehensively scan the task queue during each data packet generation cycle.

[0035] In some examples, if it is found that the data required for a certain task execution is not yet ready, the currently executable instructions of the same type will be integrated into a data packet, and an idle or low-load computing unit will be selected as the target for transmission according to the real-time load status of each computing unit (such as the status of the input queue). After completing the data packet assembly, the data packet is transmitted through the PU bus. At this time, the content-addressable memory hardware in each computing unit will perform a parallel comparison operation on the information transmitted on the bus. Once it identifies that the data is sent to itself, it will execute the receive operation to obtain the corresponding target data. Subsequently, the computing unit is triggered to execute the task through the data flow trigger module. After the calculation is completed, the computing unit feeds back the task ID, instruction ID, and result to the data packet generation and dispatch unit. This unit repeats the above process according to a fixed working cycle, including scanning the instruction queue, sorting out the dependency relationship, organizing the data packet, and dispatching it, so as to ensure that tasks can be executed efficiently and orderly, while optimizing resource utilization and task processing efficiency.

[0036] Optionally, step 103 may specifically include: obtaining the input queue lengths, the current number of tasks being executed, and the computing resource occupancy rates of multiple computing units. Correspondingly, step 104 may specifically include: selecting a target computing unit from multiple computing units according to the task priority of the target task and in combination with the input queue lengths, the current number of tasks being executed, and the computing resource occupancy rates of multiple computing units; and allocating the data packet to the target computing unit.

[0037] For example, in each data-triggered multi-core accelerator processing unit, a dedicated data packet generation and distribution unit can be provided, which is responsible for integrating the instructions of multiple tasks into data packets and distributing them to the PU bus. Its working process starts with a periodic scan of the task queue, checks whether the data dependencies of the tasks are ready, supports parallel scanning of 4 to 8 task queues, and by default, a full scan is performed every 100 working cycles to determine the data packet generation timing. Once the executable tasks are confirmed, instructions of the same type are integrated into a data packet, which includes a packet header with control information and a data part containing specific instructions and data. Subsequently, based on the task priority and the real-time load status of the PU, a suitable target PU is selected and the data packet is efficiently sent out through a star bus structure. This method not only optimizes the task execution process, but also improves the resource utilization rate and task processing efficiency, ensuring that data-intensive tasks can run smoothly under high-load conditions.

[0038] As a refinement of this embodiment, before obtaining the target task in the task queue of the processor, the tasks can be assigned to the task queue in the following ways, but not limited to, such as Figure 3 as shown Figure 3 is a schematic flowchart of a task processing method for an artificial intelligence processor provided by an embodiment of the present disclosure, including: Step 201, based on the tag data of different tasks, determine the task priorities of different tasks.

[0039] For example, corresponding tags can be attached to each task according to the nature and priority of the task to ensure that its execution requirements and dependency conditions can be accurately reflected; then, dynamic priority task scheduling is implemented, and the execution order of the tasks is flexibly adjusted according to the tag information and the current system load situation to ensure that high-priority and critical path tasks are processed first. This process not only improves the flexibility and efficiency of task scheduling, but also optimizes resource utilization, and is particularly suitable for computationally intensive deep learning application scenarios. Through this method, parallel task execution can be effectively managed and the overall computing performance can be improved.

[0040] Optionally, the tag data includes a criticality score and an estimated execution time; correspondingly, step 201 can specifically include: calculating the task priorities of different tasks based on the criticality score and the estimated execution time.

[0041] Exemplarily, the Criticality Score and the Estimated Cycles are two core metrics for evaluating task priorities. The Criticality Score is used to quantify the impact of a task on the overall system performance. It reflects the importance of the task in the data flow graph. A high score indicates that the task is on the critical path, and its delay will directly affect the overall system performance. A low score means that the task can be delayed without significantly affecting the total delay. The Estimated Cycles is the number of execution cycles of a task predicted based on the hardware performance model, which is used to estimate the time resources required for the task to complete. By combining these two metrics, the priority of a task can be calculated through a specific formula to ensure that critical and time-consuming tasks can be executed first, thereby optimizing system resource allocation and reducing the total delay or maximizing the throughput.

[0042] Optionally, the method of this embodiment may specifically further include: obtaining the path relationships of different tasks in the data flow graph; and performing criticality scoring on different tasks according to the impact degree of the path relationships on the processor performance.

[0043] For example, an artificial intelligence training program can be decomposed into a data flow graph through algorithms and toolchains (such as TensorFlow XLA, PyTorch Glow, etc.), where nodes represent tasks and edges represent the dependencies between tasks. Then, task dynamic label injection is performed to attach detailed label information to each task, including task ID, task type, input dependency list, estimated cycles, and criticality score.

[0044] In some examples, the priority calculation formula is: Priority = Criticality Score (60% weight) + Estimated Cycles (40% weight). The Criticality Score, as a core metric for task scheduling and resource allocation, is used to quantify the impact of a task on the overall system performance, and its value usually varies between 0.0 and 1.0. A high score (close to 1.0) means that the task is on the critical path, and its delay will directly affect the overall system performance, such as the output layer of artificial intelligence inference. A low score (close to 0.0) indicates that the task can be delayed without significantly affecting the total delay, such as an offline batch processing task.

[0045] Optionally, the above-mentioned obtaining of the path relationships of different tasks in the data flow graph may specifically further include: dividing different tasks into multiple data processing nodes and multiple data flow paths; and determining the path relationships of different tasks in the data flow graph based on the multiple data processing nodes and multiple data flow paths.

[0046] Exemplarily, first, the data flow graph of the artificial intelligence program is input, such as the computation graph generated by TensorFlow or PyTorch, whose nodes represent tasks (such as matrix multiplication, convolution operation, etc.) and the edges represent data dependency relationships (such as the data flow direction from task A to task B). The first step of the whole process is to construct this data flow graph to ensure that all tasks and their dependency relationships are accurately represented. Then, in the second step, critical path analysis is performed, and the longest path from the starting point to the ending point, that is, the critical path, is calculated using the Critical Path Method (CPM) or dynamic programming algorithm. Furthermore, according to the path relationship, the tasks located on this critical path are given an initial criticality score of 1.0 because they have a direct impact on the overall performance; while the tasks not on the critical path are scored decreasingly according to their distance from the nearest critical path node.

[0047] Optionally, the method of this embodiment may specifically further include: dynamically adjusting the criticality score according to the actual execution situation of different tasks, and updating the estimated execution time of different tasks.

[0048] Exemplarily, label injection is divided into two stages: compile-time static injection and run-time dynamic adjustment to optimize task scheduling. First, in the compile-time static injection stage, based on the data flow graph analysis combined with the hardware model, the label information of each task is initialized, including task ID, task type, input dependency list, estimated cycle, and the initial criticality score, providing a basic basis for subsequent task execution. Then, in the run-time dynamic adjustment stage, the system optimizes the criticality score in real time according to the feedback information of the actual execution to ensure that it can accurately reflect the actual impact and urgency of each task in the current environment.

[0049] Optionally, the above-mentioned dynamically adjusting the criticality score according to the actual execution situation of different tasks and updating the estimated execution time of different tasks may specifically further include: based on the estimated execution time, monitoring the actual execution time, resource occupancy situation of different tasks, and the impact on the overall performance of the processor; adjusting the criticality score according to the monitoring results, and updating the estimated execution time of different tasks.

[0050] Exemplarily, in the compiler static analysis stage, first, the DFG is input, and the task paths crucial to the system performance are identified through critical path analysis; then all task nodes are traversed, and operations such as task ID generation, dependency relationship analysis, estimated execution cycle based on the hardware performance model, determination of the initial criticality score, and provision of hardware optimization tips are completed in sequence, and finally the DFG injected with label information is output. And in the run-time dynamic adjustment stage, when the task execution is completed, the system dynamically updates the label information according to the actual execution situation.

[0051] For example, the task importance score accounts for 60% of the weight and is used to quantify the impact of a task on the overall system performance. In particular, tasks on the critical path will receive higher scores (such as the output layer calculation in AI inference); the estimated execution time accounts for 40% of the weight, and tasks with shorter estimated execution cycles are preferred to reduce queuing delays. The specific formula for calculating the priority is: Priority = α × Criticality Score + β × (1 / Estimated Cycle), where α and β are 0.6 and 0.4 respectively, and these parameters can be dynamically adjusted according to the system state. The calculation of the task importance score is divided into two stages: the static analysis stage during compilation determines the initial score, initializes the label information based on the data flow graph analysis and the hardware model; the dynamic adjustment stage during runtime optimizes the score according to the actual execution feedback, uses the exponential smoothing method to dynamically adjust the criticality score and update the estimated cycle to prevent subsequent task starvation.

[0052] In some examples, the execution data can be recorded by monitoring the actual execution cycles of each task and comparing it with the estimated cycle. If the actual execution cycle of a task exceeds the estimated cycle by 20%, it is considered to have a potential impact on the critical path. Then, positive and negative feedback mechanisms are used to dynamically adjust the criticality score of the task. For tasks on the critical path with increasing delays, their scores will be increased; while for tasks not on the critical path and completed ahead of schedule, their scores will be decreased. Among them, the update of the score uses the exponential smoothing formula as follows: new_score = α * old_score + (1 - α) * adjustment where α is a smoothing factor with a default value of 0.8, which is used to control the adjustment amplitude. The adjustment value is calculated based on the impact of the task on the overall delay: if the task delay causes subsequent tasks to be blocked, the adjustment is set to min(1.0, old_score + 0.1); if the task is completed ahead of schedule and does not affect the dependencies, the adjustment is set to max(0.0, old_score - 0.05). Then the system periodically (e.g., every 1000 cycles) recalculates the global critical path, analyzes and updates the critical path by traversing the entire data flow graph (DFG), ensuring that the criticality score can reflect the actual situation in the latest execution environment, so as to achieve optimal resource allocation and scheduling optimization.

[0053] Exemplarily, first, the highest-priority ready task with satisfied dependencies is selected from the head of the priority queue, and then the processing unit (PU) with the lowest comprehensive load is selected according to the load status table for task allocation. Specifically, the algorithm evaluates the load of each PU based on metrics such as computing utilization, bus bandwidth utilization, and queue depth, and binds the task to the selected PU. At the same time, the load status table is updated to reflect this change. For example, task T100 is assigned to PU2, and the queue depth count of PU2 is increased accordingly. To further optimize the scheduling effect, the algorithm also includes a dynamic feedback mechanism: periodically (such as every 1000 cycles), the weight value is adjusted. If it is found that a certain PU is in an overloaded state for a long time (utilization exceeds 90%), its task allocation weight is reduced to reduce the allocation of new tasks. If the overall throughput of the system decreases, the weight of the criticality score is increased by adjusting the α / β ratio to ensure that tasks on the critical path are processed first, thereby improving the response speed and execution efficiency of the system.

[0054] Step 202: According to the task priorities of different tasks, allocate different tasks to the task queues of at least one processing unit in the processor.

[0055] Optionally, step 202 may specifically include: by monitoring the utilization of at least one processing unit in real time, allocate tasks with task priorities higher than the preset priority threshold to the task queues of the processing units with current utilization lower than the preset utilization threshold.

[0056] Exemplarily, the load-based resource allocation mechanism aims to efficiently allocate tasks to the data-triggered multi-core accelerator processing units according to the real-time load conditions of each processing unit. Each processing unit reports its utilization in real time, including metrics such as the number of instructions executed per cycle and bus bandwidth utilization, so as to allocate high-priority tasks to the processing unit with the lowest current utilization, thereby optimizing resource usage. The computing utilization is completed by counting the active cycle ratio of the computing units (such as matrix units, vector units, etc.) within the processing unit. The specific formula is: utilization = (the number of cycles for computing component 1 to execute instructions + the number of cycles for computing component 2 to execute instructions +... + the number of cycles for computing component n to execute instructions) / (the number of computing components * the total number of cycles). This statistical process is implemented using counters and counter registers in the computing components. The bus bandwidth utilization measures the actual bandwidth occupancy rate of data transmission between the PU and memory or other PUs, and is calculated by the formula bandwidth utilization = (the actual amount of data transmitted) / (the maximum theoretical bandwidth). In addition, monitoring the PU input / output queue depth reflects the number of tasks to be processed to evaluate the backlog situation. The comprehensive load is obtained by weighted summation of the computing utilization, bandwidth utilization, and queue depth.

[0057] In some examples, a global scheduler (such as a GPP or a dedicated hardware module) is responsible for collecting the load metrics of all PUs and maintaining a load status table, based on which dynamic task allocation and load balancing are performed to ensure that the system can operate efficiently according to the real-time load conditions, maximize the utilization of available resources, and reduce task latency.

[0058] In some examples, by adopting a dynamic priority mechanism, combining the criticality score and the estimated execution time of tasks to optimize the task allocation order, tasks on the critical path are ensured to be executed first. At the same time, a load-based resource allocation strategy monitors the utilization rate of processing units in real time, including computing utilization rate, bandwidth utilization rate, and queue depth, and dynamically allocates tasks to the PU with the lowest current load to achieve load balancing within the system. At the local scheduling level, calculations are triggered by data packet flows inside the PU, and the task execution efficiency is further improved according to the task priority and the real-time load status of the PU.

[0059] In some examples, through dynamic label injection technology and runtime parameter adjustment means, dynamic scheduling at the global and local levels is achieved. At the global scheduling level, task allocation is optimized according to the system resource status to ensure that tasks are reasonably distributed among processing units; at the local scheduling level, a data packet flow is used to trigger the calculation process to ensure that tasks on the critical path can be executed efficiently. In addition, the proposed data flow trigger mechanism not only effectively reduces the programming complexity but also greatly reduces the need for manual management of dependencies in explicit data flow programming, thus simplifying the development process and improving the overall execution efficiency of the system.

[0060] Optionally, the method of this embodiment may further specifically include: parallelly comparing the target address of a data packet through a content-addressable memory to determine whether the data packet matches the target computing unit; if the data packet matches the target computing unit and the data packet is in a data-ready state, then trigger the target computing unit to execute the computing task corresponding to the target task based on the data packet.

[0061] Exemplarily, in the task execution and data packet generation process, when the system scans the task queue merged with all task instructions and finds that the data relied on by the task to be executed is not yet ready, it will integrate and package the currently executable instructions of the same type into data packets in a specific format. After completing the assembly of the data packets, they are sent to the PU bus for subsequent processing. Each data packet consists of two parts: a header and a payload. The header contains key control information and identifiers, which are used to guide the classification, integration, and dispatching process of the data packets; the payload carries specific operation instructions and the required data. In addition, a category field is set in the data packet to accurately classify the data packets according to the category of the computing unit (such as matrix computing unit MCU, vector computing unit VCU, etc.), and the data packets are allocated to the corresponding computing units for execution accordingly. Through this mechanism, it is ensured that tasks can be executed efficiently and orderly, while optimizing resource utilization and task processing efficiency.

[0062] In some examples, the PU includes multiple dedicated computing modules optimized for artificial intelligence tasks to achieve efficient task processing. The Matrix Unit supports general matrix multiplication (GEMM), convolution operations, etc., and adopts a parallel computing architecture to support large-scale matrix operations; the Vector Unit focuses on element-wise operations such as addition, multiplication, and reduction operations such as summation, maximum value calculation; the Float Scalar Unit can perform floating-point multiply-add calculations (floating-point MAC) and floating-point arithmetic logic unit (floating-point ALU) operations; the Fixed-point Scalar Unit supports fixed-point multiply-add calculations (fixed-point MAC) and fixed-point arithmetic logic unit (fixed-point ALU) operations. In addition, a special computing unit is provided to support complex operations such as trigonometric function calculations and reduction. The MemoryUnit is responsible for data transfer and synchronization operations, such as Barrier. Each computing module is equipped with an independent data stream trigger module, which receives data through the input queue and passes the calculation results to the next stage via the output queue, ensuring the efficiency and smoothness of data processing. By modularizing multiple computing units (including matrix, vector, floating-point, and fixed-point computing units, etc.), the overall performance is improved, and it can flexibly adapt to diverse artificial intelligence computing needs, enabling the PU to achieve high-efficiency and high-performance computing capabilities when processing complex artificial intelligence tasks.

[0063] Optionally, the method of this embodiment may specifically further include: managing and converting the address spaces of the high-bandwidth memory and the main memory in the processor through the memory interface.

[0064] Exemplarily, the data packet generation and distribution unit adopts a star bus structure to effectively connect all the above-mentioned computing modules, building an efficient data transmission path. At the same time, this unit is also connected to a register file (with a capacity of 128 KB) and on-chip memory high bandwidth memory (High Bandwidth Memory, HBM). In addition, these HBMs are uniformly addressed with the dynamic random access memory (Dynamic Random Access Memory, DDR) on the GPP side through the Memory Management Unit (MMU), thereby realizing the unified management and efficient utilization of memory resources at the system level.

[0065] Exemplarily, the key components and hardware support of the data flow trigger module include CAM hardware, input / output queues, register files, and high bandwidth memory. The CAM hardware is used to achieve fast address matching to ensure that data packets can be accurately delivered to the target computing unit. Each computing module is equipped with an independent input / output queue, which not only decouples the computing process from data transmission but also improves the overall data processing efficiency. The register file provides the ability to quickly access intermediate results, effectively enhancing the data processing speed; while the high bandwidth memory is uniformly addressed through the memory management unit, reducing the memory access latency and further enhancing the data access efficiency.

[0066] Exemplarily, the adoption of high bandwidth memory and unified address management reduces the memory access latency and improves the data throughput capacity, while the register file supports fast data access and efficient memory utilization. In addition, through the hardware optimization method of data trigger comparison, using the content-addressable memory and dedicated queue mechanism, the fast identification and distribution of data packets are realized, effectively reducing the communication overhead, not only enhancing the computing performance of the system but also improving the data processing speed and task execution efficiency.

[0067] Compared with the current existing technologies, in this embodiment, by integrating the data flow trigger mechanism and the dynamic resource allocation strategy, while improving the processing performance of the artificial intelligence scenario, the problem of imbalance between idle and overloaded hardware resources caused by traditional static task partitioning is solved; for the random memory access pattern triggered by fine-grained data flow, an optimized memory access mechanism significantly improves the cache hit rate and alleviates the memory bandwidth bottleneck; through the intelligent allocation strategy for critical path tasks, hotspot concentration is avoided, and a more balanced task execution efficiency is achieved; finally, through the implicit data flow management mechanism, the explicit programming dependence is reduced, the data dependence management process and debugging complexity are simplified, and the contradiction between the core performance and usability of traditional data flow processors is systematically solved from four dimensions: architecture, storage, scheduling, and programming. The processing unit significantly improves the execution efficiency of artificial intelligence tasks and hardware utilization by introducing a data flow trigger mechanism, dynamic priority scheduling, and the design of an efficient computing module. The data flow trigger mechanism realizes the decoupling of computation and data, enabling the system to trigger computations on demand according to actual requirements, improving flexibility and response speed; the dynamic scheduling strategy combines the criticality score of tasks and the real-time load status, optimizes resource allocation, ensures that tasks on the critical path are processed first, and at the same time balances the workload of each processing unit.

[0068] An embodiment of the present application further provides a task processing device for an artificial intelligence processor, as Figure 1 a specific implementation of the method shown in Figure 4 shown, the device includes: a first acquisition module 31, an integration module 32, a second acquisition module 33, and an allocation module 34.

[0069] The first acquisition module 31 is configured to acquire a target task in the task queue of the processor; The integration module 32 is configured to integrate task instructions of the same operation type that can be executed among the target tasks into a data packet based on the data dependence relationship between the target tasks; The second acquisition module 33 is configured to acquire the task priority of the target task and the real-time load status of multiple computing units in the processor; The allocation module 34 is configured to allocate the data packet to a target computing unit among the multiple computing units according to the task priority and the real-time load status, and the target computing unit is used to execute the computing task corresponding to the target task based on the data packet.

[0070] In some examples of this embodiment, the integration module 32 is specifically configured to identify, from the task instructions of the target task, task instructions that are executable, of the same operation type, and satisfy the data dependence relationship; and integrate the identified task instructions to generate the data packet.

[0071] In some examples of this embodiment, the second acquisition module 33 is specifically configured to acquire the input queue lengths, the current number of executing tasks, and the computing resource occupancy rates of the multiple computing units; correspondingly, the allocation module 34 is specifically configured to select the target computing unit from the multiple computing units according to the task priorities of the target tasks and in combination with the input queue lengths, the current number of executing tasks, and the computing resource occupancy rates of the multiple computing units; and allocate the data packet to the target computing unit.

[0072] In some examples of this embodiment, the first acquisition module 31 is specifically configured to determine the task priorities of different tasks based on the tag data of different tasks; and allocate the different tasks to the task queues of at least one processing unit in the processor according to the task priorities of different tasks.

[0073] In some examples of this embodiment, the tag data includes a criticality score and an estimated execution time; correspondingly, the first acquisition module 31 is specifically further configured to calculate the task priorities of the different tasks based on the criticality score and the estimated execution time.

[0074] In some examples of this embodiment, the first acquisition module 31 is specifically further configured to acquire the path relationships of the different tasks in the data flow graph; and perform criticality scoring on the different tasks respectively according to the influence degree of the path relationships on the processor performance.

[0075] In some examples of this embodiment, the first acquisition module 31 is specifically further configured to divide the different tasks into multiple data processing nodes and multiple data flow paths; and determine the path relationships of the different tasks in the data flow graph based on the multiple data processing nodes and the multiple data flow paths.

[0076] In some examples of this embodiment, the first acquisition module 31 is specifically further configured to dynamically adjust the criticality score according to the actual execution conditions of the different tasks, and update the estimated execution time of the different tasks.

[0077] In some examples of this embodiment, the first acquisition module 31 is specifically further configured to monitor the actual execution time, the resource occupancy, and the influence on the overall performance of the processor of the different tasks based on the estimated execution time; adjust the criticality score according to the monitoring results, and update the estimated execution time of the different tasks.

[0078] In some examples of this embodiment, the first acquisition module 31 is further specifically configured to, by monitoring the utilization rate of the at least one processing unit in real time, assign tasks with a task priority higher than a preset priority threshold to the task queue of the processing unit whose current utilization rate is lower than the preset utilization rate threshold.

[0079] In some examples of this embodiment, the allocation module 34 is further specifically configured to, by parallelly comparing the destination address of the data packet through a content-addressable memory, determine whether the data packet matches the target computing unit; if the data packet matches the target computing unit and the data packet is in a data-ready state, trigger the target computing unit to execute the computing task corresponding to the target task based on the data packet.

[0080] In some examples of this embodiment, the allocation module 34 is further specifically configured to, through a memory interface, manage and convert the address spaces of the high-bandwidth memory and the main memory in the processor.

[0081] It should be noted that for other corresponding descriptions of the various functional units involved in the task processing device of the artificial intelligence processor provided in this embodiment, reference can be made to Figure 1 the corresponding descriptions therein, which will not be elaborated here.

[0082] Based on the above methods as shown in Figure 1 and Figure 3 correspondingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above methods as shown in Figure 1 and Figure 3 shown.

[0083] Based on the above methods as shown in Figure 1 and Figure 3 correspondingly, this embodiment further provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above methods as shown in Figure 1 and Figure 3 shown.

[0084] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of this application.

[0085] Based on the above methods as shown in Figure 1 and Figure 3 shown, and Figure 4In the virtual device embodiment shown, to achieve the above object, an embodiment of the present application further provides an electronic device, such as a personal computer or a server. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as shown in Figure 1 and Figure 3 shown above.

[0086] In some embodiments, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.

[0087] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not limit the physical device, and it may include more or fewer components, or combine certain components, or have different component arrangements.

[0088] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of the information processing program and other software and / or programs. The network communication module is used to implement communication between the components inside the storage medium, and communication between other hardware and software in the information processing physical device.

[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the current existing technologies, this embodiment proposes a data-triggered multi-core artificial intelligence processor architecture and a dynamic priority and load task scheduling and execution method based on this architecture, aiming to solve the core pain points of traditional data flow processors in terms of performance, hardware utilization, storage bandwidth pressure, task execution efficiency, and programming complexity. First, by combining data flow triggering and dynamic resource allocation, the processing performance in artificial intelligence application scenarios is improved, and the problem of idle partial processing units and overloaded other processing units caused by static task partitioning is overcome, thereby improving hardware utilization. Secondly, the problems of low cache hit rate and memory bandwidth bottleneck caused by random memory access triggered by fine-grained data flow are optimized, and the pressure on storage bandwidth is improved. In addition, in this embodiment, by optimizing the allocation of tasks on the critical path to avoid hot spots, a more reasonable task allocation strategy is realized, and the task execution efficiency is improved. Finally, by reducing the dependence on explicit data flow programming, the need for manual management of data dependencies is reduced, the debugging process is simplified, and the programming complexity is effectively reduced. The solution of this embodiment not only considers the criticality score and estimated execution time of tasks to optimize resource allocation, but also flexibly schedules tasks according to the real-time load situation to ensure that tasks on the critical path are processed first. In addition, the scheduling parameters can be further optimized by combining reinforcement learning, or the data transfer overhead can be reduced by using 3D stacked memory technology, thereby continuously improving the system performance. This innovation provides an efficient and flexible task execution paradigm for the next-generation heterogeneous computing architecture and lays a solid technical foundation.

[0090] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0091] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A task processing method for an artificial intelligence processor, characterized in that: include: Get the target task in the task queue of the processor; Based on the data dependency relationship between the target tasks, task instructions of the same operation type executable by the target tasks are integrated into a data packet; Acquire the task priority of the target task and the real-time load status of multiple computing units in the processor; According to the task priority and the real-time load status, the data packet is distributed to a target computing unit among the multiple computing units, and the target computing unit is used to execute the computing task corresponding to the target task based on the data packet.

2. The method according to claim 1, characterized in that The step of integrating the task instructions of the same operation type executable by the target tasks into a data packet based on the data dependency relationship between the target tasks includes: From the task instructions of the target task, identifying the task instructions that are executable, have the same operation type, and satisfy the data dependency relationship; The identified task instructions are integrated to generate the data packet.

3. The method according to claim 1, characterized in that Acquiring real-time load status of multiple computing units in the processor includes: Obtaining the input queue lengths, the number of currently executed tasks, and the computing resource occupancy rates of the multiple computing units; The allocating the data packet to a target computing unit among the plurality of computing units according to the task priority and the real-time load status comprises: Selecting the target computing unit from the multiple computing units according to the task priority of the target task and in combination with the input queue lengths, the number of currently executed tasks, and the computing resource occupancy rates of the multiple computing units; The data packets are distributed to the target computing units.

4. The method according to claim 1, characterized in that: Before acquiring the target task in the task queue of the processor, the method further includes: Determine the task priorities of different tasks based on the label data of different tasks; According to the task priorities of different tasks, the different tasks are respectively allocated to the task queue of at least one processing unit in the processor.

5. The method according to claim 4, characterized in that The label data includes a criticality score and an estimated execution time; The step of determining the task priorities of different tasks based on the label data of different tasks includes: Based on the criticality scores and the estimated execution times, task priorities of the different tasks are calculated.

6. The method according to claim 5, characterized in that The method further comprises: Obtaining the path relationship of the different tasks in the data flow graph; According to the influence degree of the path relationship on the processor performance, the different tasks are respectively scored in terms of criticality.

7. The method according to claim 6, characterized in that The obtaining of the path relationship of the different tasks in the data flow graph includes: Dividing the different tasks into a plurality of data processing nodes and a plurality of data flow paths; Based on the multiple data processing nodes and the multiple data flow paths, the path relationship between the different tasks in the data flow graph is determined.

8. The method according to claim 5, characterized in that The method further comprises: According to the actual execution status of the different tasks, the criticality scores are dynamically adjusted, and the estimated execution times of the different tasks are updated.

9. The method according to claim 8, characterized in that The dynamically adjusting the criticality scores according to the actual execution status of the different tasks and updating the estimated execution time of the different tasks include: Based on the estimated execution time, monitoring the actual execution time, resource usage and impact on the overall performance of the processor of the different tasks; The criticality score is adjusted according to the monitoring result, and the estimated execution time of the different tasks is updated.

10. The method according to claim 4, characterized in that The step of allocating the different tasks to the task queues of at least one processing unit in the processor according to the task priorities of the different tasks comprises: By real-time monitoring of the utilization of the at least one processing unit, the task with a task priority higher than a preset priority threshold is allocated to a task queue of a processing unit whose current utilization is lower than a preset utilization threshold.

11. The method according to claim 1, characterized in that: After allocating the data packet to a target computing unit among the plurality of computing units according to the task priority and the real-time load status, the method further includes: Comparing the target address of the data packet in parallel through a content addressable memory to determine whether the data packet matches the target computing unit; If the data packet matches the target computing unit and the data packet is in a data-ready state, the target computing unit is triggered to execute a computing task corresponding to the target task based on the data packet.

12. The method according to claim 1, characterized in that The method further comprises: The address space of the high bandwidth memory and main memory in the processor is managed and converted through the memory interface.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

14. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.

15. A computer program product having a computer program stored thereon, characterized in that: When the computer program product is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Task processing method and system for comprehensive resource management of intelligent monitoring system

    CN117608840A

  • Task resource scheduling method and device, equipment and storage medium

    CN119376890A

  • Distributed task scheduling method and system for heterogeneous tasks based on cloud native, and medium

    CN119759594A

  • Workload scheduling using queues with different priorities

    US20240118920A1

Cited By

  • Distributed computing system and method, electronic device, storage medium and program product

    CN120353557A

  • Task allocation method and device, electronic equipment and storage medium

    CN120523611A

  • Processing method, device and equipment for starting task of multi-core system and storage medium

    CN120596169A

  • Push type loading system based on universal loading accelerator

    CN120596170A

  • Computing task processing method and device, medium and product

    CN120670125A