Resource scheduling system and method based on artificial intelligence chip
Patent Information
- Application Number
- CN202610746272.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-09-25
AI Technical Summary
受限于基于软件采样的传统监测手段(高延迟、低精度及额外算力开销),系统无法精准捕捉芯片内部硬件单元的实时执行状态,导致调度策略与硬件动态能力严重脱节,进而导致系统的效能降低
通过本发明所提供的上述实施例,通过信息交互模块接收用户输入的推理任务及预选调度策略,实现了任务与策略的解耦。用户可根据具体应用场景(如低延迟推理或高吞吐训练)选择最优策略,系统能够灵活响应,从而在保证服务质量的前提下,最大化资源利用效率。芯片监控模块与流处理模块均集成于人工智能芯片内部,形成了紧耦合的监控-执行闭环。芯片监控模块能够基于实时的推理任务负载和预设的调度策略,对流处理模块的计算过程进行细粒度的动态调度与实时监控。这种片上集成设计极大地降低了监控延迟,使得系统能够迅速发现并处理计算瓶颈,保障任务执行的稳定性。流处理模块作为执行单元,专注于推理任务的计算过程。芯片监控模块的介入,使得计算资源的分配不再是静态或盲目的,而是根据任务的实际需求和芯片的实时状态进行动态调配。这有效避免了计算资源的闲置或过载,显著提升了AI芯片的整体算力利用率。集成在芯片内部的监控模块能够精确捕获流处理模块的执行状态,一旦出现异常(如计算错误或死锁),可立即触发相应的保护机制或重新调度策略。这种深度的硬件级监控使得系统的执行行为更具确定性,为上层应用提供了稳定可靠的运行环境。
Smart Images

Figure CN122816784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and for example to a resource scheduling system and method based on an artificial intelligence chip. Background Technology
[0002] With the rapid development of artificial intelligence technology, the parameter scale of large-scale AI (Artificial Intelligence) models, such as large language models and computer vision models, is growing exponentially, placing extremely high demands on the computing power, storage bandwidth, and scheduling efficiency of AI chips. As the core hardware carrier for model operation, the execution state of AI chips directly determines the training and inference efficiency of the model, while the model scheduling strategy is the key link connecting the model requirements with the chip hardware capabilities.
[0003] Model scheduling in related technologies often relies on static configuration at the software level or offline optimization based on historical data. Limited by traditional monitoring methods based on software sampling (high latency, low accuracy, and additional computing power overhead), the system cannot accurately capture the real-time execution status of the internal hardware units of the chip, resulting in a serious disconnect between the scheduling strategy and the dynamic capabilities of the hardware, which in turn leads to a reduction in system performance. Summary of the Invention
[0004] The present invention aims to provide a resource scheduling system, method, electronic device and storage medium based on artificial intelligence chip.
[0005] According to one aspect of the present invention, a resource scheduling system based on an artificial intelligence chip is proposed, comprising: The information interaction module is used to receive user input of inference tasks and pre-selected scheduling strategies; The stream processing module, integrated on the artificial intelligence chip, is used to execute and monitor the computation process of inference tasks in order to generate task execution results; The chip monitoring module, integrated on the artificial intelligence chip, is used to schedule the stream processing module based on inference tasks and pre-selected scheduling strategies.
[0006] According to one aspect of the present invention, a resource scheduling method based on an artificial intelligence chip is proposed, comprising: Receive user input for inference tasks and pre-selected scheduling strategies; Based on inference tasks, pre-selected scheduling strategies, and artificial intelligence chips, the computation process of inference tasks is executed and monitored to generate task execution results.
[0007] According to one aspect of the present invention, an electronic device is provided, comprising: a processor; and a memory storing a computer program, which, when executed by the processor, causes the processor to perform the method described above.
[0008] According to one aspect of the present invention, a non-transitory computer-readable medium is provided, on which readable instructions are stored, which, when executed by a processor, cause the processor to perform the method described above.
[0009] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit the invention.
[0010] Beneficial effects: Through the embodiments provided by this invention, the inference task and pre-selected scheduling strategy are received by the information interaction module, thereby decoupling the task and strategy. Users can select the optimal strategy based on specific application scenarios (such as low-latency inference or high-throughput training), and the system can respond flexibly, maximizing resource utilization efficiency while ensuring service quality. Both the chip monitoring module and the stream processing module are integrated within the AI chip, forming a tightly coupled monitoring-execution closed loop. The chip monitoring module can perform fine-grained dynamic scheduling and real-time monitoring of the stream processing module's computation process based on real-time inference task load and preset scheduling strategies. This on-chip integration design significantly reduces monitoring latency, enabling the system to quickly identify and address computational bottlenecks, ensuring the stability of task execution. The stream processing module, as the execution unit, focuses on the computation process of the inference task. The intervention of the chip monitoring module ensures that the allocation of computing resources is no longer static or blind, but dynamically adjusted according to the actual needs of the task and the real-time status of the chip. This effectively avoids idle or overloaded computing resources and significantly improves the overall computing power utilization of the AI chip. The monitoring module integrated within the chip can accurately capture the execution status of the stream processing module. In the event of an anomaly (such as a calculation error or deadlock), it can immediately trigger the corresponding protection mechanism or rescheduling strategy. This deep hardware-level monitoring makes the system's execution behavior more deterministic, providing a stable and reliable operating environment for upper-layer applications. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without exceeding the scope of protection claimed by the present invention.
[0012] Figure 1 A schematic diagram of the overall system architecture provided in an embodiment of the present invention; Figure 2 A block diagram of a resource scheduling system based on an artificial intelligence chip provided in an embodiment of the present invention; Figure 3 A schematic diagram of the polling scheduling process provided in an embodiment of the present invention; Figure 4 A schematic diagram of a delay-first process provided in an embodiment of the present invention; Figure 5 A schematic diagram of a memory-first process provided for an embodiment of the present invention; Figure 6 A schematic diagram of the data update process provided for embodiments of the present invention; Figure 7 A flowchart of a data update-based resource scheduling method using an artificial intelligence chip, provided for embodiments of the present invention; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0013] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that the invention will be thorough and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.
[0014] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of embodiments of the invention. However, those skilled in the art will recognize that the technical solutions of the invention can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the invention.
[0015] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0016] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0017] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, a first component discussed below could be termed a second component without departing from the teachings of the inventive concept. As used herein, the term "and / or" includes any one and all combinations of one or more of the associated listed items.
[0018] English abbreviations used in the present invention and their corresponding Chinese translations: English abbreviation: AI, Chinese translation: Artificial Intelligence, full English name: Artificial Intelligence; English abbreviation: SP, Chinese translation: stream processor, full English name: Streaming Processor; English abbreviation: PCIe, Chinese translation: Peripheral Component Interconnect Express bus, full English name: PeripheralComponent Interconnect Express English abbreviation: CPU, Chinese translation: Central Processing Unit, full English name: Central Processing Unit; English abbreviation: DSP, Chinese translation: Digital Signal Processor, full English name: Digital SignalProcessor; English abbreviation: ASIC, Chinese translation: Application Specific Integrated Circuit, full English name: Application SpecificIntegrated Circuit; English abbreviation: FPGA, Chinese translation: Field Programmable Gate Array, full English name: Field ProgrammableGate Array; English abbreviation: EISA, Chinese translation: Extended Industry Standard Architecture, full English name: Extended IndustryStandard Architecture; English abbreviation: ROM, Chinese translation: Read Only Memory, full English name: Read Only Memory English abbreviation: RAM, Chinese translation: Random Access Memory, full English name: Random Access Memory English abbreviation: EEPROM, Chinese translation: Electrically Erasable Programmable Read Only Memory, full English name: ElectricallyErasable Programmable Read Only Memory English abbreviation: CD ROM, Chinese translation: compact disc read-only memory, full English name: Compact Disc Read Only Memory.
[0019] Figure 1 is a schematic diagram of the overall system architecture provided by an embodiment of the present invention. Here, HOST refers to the host where the data stream artificial intelligence chip is located. A user completes system startup and sends inference tasks by operating the HOST. The information interaction module can receive inference tasks and a pre-selected scheduling strategy (the user may or may not select a strategy). A stream processing module and a chip monitoring module can be integrated on the data stream AI chip, wherein the stream processing module can specifically be an SP (StreamingProcessor, stream processor) computing unit and an SP monitoring module. The chip monitoring module is responsible for scheduling and managing all SPs, and receives inference tasks and scheduling strategies issued from the HOST. The chip monitoring module includes a scheduler, a sliding window queue, a scheduling strategy and other parts. The SP computing unit in the stream processing module can include a manager, a task queue, a Vector Core, a DDR (Double Data Rate SDRAM) and a TensorCore. The SP monitoring module can monitor the status of each hardware component in the SP computing unit, calculate data such as real-time utilization of the computing unit, memory capacity and bandwidth occupancy, and is responsible for reporting and saving the data to the sliding window queue of the chip monitoring module.
[0020] For specific implementation manners, reference may be made to the following embodiments.
[0021] Figure 2 is a block diagram of a resource scheduling system based on an artificial intelligence chip provided by an embodiment of the present invention. As Figure 2 shows, the system includes: an information interaction module 20, a stream processing module 21 and a chip monitoring module 22.
[0022] the information interaction module 20 is configured to receive an inference task and a pre-selected scheduling strategy input by a user; the stream processing module 21 is integrated on the artificial intelligence chip, and the stream processing module is configured to execute and monitor a calculation process of the inference task to generate a task execution result; the chip monitoring module 22 is integrated on the artificial intelligence chip, and the chip monitoring module is configured to schedule the stream processing module based on the inference task and the pre-selected scheduling strategy.
[0023] In this invention, users can input inference tasks and pre-selected scheduling strategies through the information interaction module 20. The inference task can be a data packet carrying computational requirements, resource constraints, and expected goals. The resource scheduling system of this invention can pre-set multiple scheduling strategies. Users can choose which scheduling strategy to use when inputting the inference task; if the user does not select a strategy, the system can directly use the default strategy.
[0024] The stream processing module 21 and the chip monitoring module 22 are integrated on the AI chip. The chip monitoring module 22 can schedule the stream processing module 21 based on the inference task and a pre-selected scheduling strategy. The scheduled stream processing module 21 can execute the computation process of the inference task and generate the task execution result.
[0025] In some implementations, the pre-selection scheduling strategy can be a round-robin algorithm. The chip monitoring module 22 can number all SPs in the stream processing module, such as 1-10, and schedule them sequentially. Upon receiving a new inference task, the chip monitoring module 22 polls the SPs sequentially by number. If the current SP has sufficient resources, the inference task is placed in that SP. If the current SP lacks resources, the next SP is queried (if SP10 is reached, the next SP is SP1), until sufficient resources are available. The next inference task is queried starting from the SP following the current SP, and so on. If all SPs lack resources, the controller waits until sufficient resources are available for task scheduling. (See reference...) Figure 3 Assuming that the stream processing module 21 contains 4 SPs (SP1-SP4) and their corresponding task allocation and memory usage, the size of the TASK block and the memory required by the code, that is, the maximum virtual memory boundary that a task (process) can have, is defined by the TASK_SIZE parameter. This boundary determines how much memory space the running code can occupy at most.
[0026] When TASK1 arrives, the current SP is SP1. SP1 has enough memory to hold TASK1, so TASK1 is assigned to SP1, and the SP index (stream processor index) is updated to SP2. When TASK2 arrives, SP2 does not have enough memory to hold TASK2, so the SP index points to SP3. After checking SP3, it is confirmed that SP3 has enough space to hold TASK2, so TASK2 is scheduled to SP3, and the SP index is updated to SP4. When TASK3 arrives, SP4 has enough space to hold TASK3, so TASK3 will be scheduled to SP4. This is the round-robin scheduling process.
[0027] This invention decouples tasks and strategies by receiving user-inputted inference tasks and pre-selected scheduling strategies through an information interaction module. Users can choose the optimal strategy based on specific application scenarios (such as low-latency inference or high-throughput training), and the system can respond flexibly, thereby maximizing resource utilization efficiency while ensuring service quality. Both the chip monitoring module and the stream processing module are integrated within the AI chip, forming a tightly coupled monitoring-execution closed loop. The chip monitoring module can perform fine-grained dynamic scheduling and real-time monitoring of the stream processing module's computation process based on real-time inference task load and preset scheduling strategies. This on-chip integration design significantly reduces monitoring latency, enabling the system to quickly identify and address computational bottlenecks, ensuring task execution stability. The stream processing module, as the execution unit, focuses on the computation process of the inference task. The intervention of the chip monitoring module ensures that the allocation of computing resources is no longer static or blind, but dynamically adjusted according to the actual needs of the task and the real-time status of the chip. This effectively avoids idle or overloaded computing resources and significantly improves the overall computing power utilization of the AI chip. The monitoring module integrated within the chip can accurately capture the execution status of the stream processing module. In the event of an anomaly (such as a calculation error or deadlock), it can immediately trigger the corresponding protection mechanism or rescheduling strategy. This deep hardware-level monitoring makes the system's execution behavior more deterministic, providing a stable and reliable operating environment for upper-layer applications.
[0028] According to some embodiments, the stream processing module 21 includes: a computation submodule 211, used to perform the computation process of the inference task to generate the task execution result; and a monitoring submodule 212, used to monitor the computation process of the inference task and report hardware status information in real time.
[0029] In this invention, the computing submodule 211 mainly includes specific computing units such as Tnsor Cores (used to perform Tensor calculations in the AI model) and Vector Cores (used to perform Vector calculations in the AI model), as well as devices such as DDR for storing data. It also has a pre-defined hardware interface to query information such as the utilization rate of the computing hardware and the DDR usage rate, facilitating the execution of the computing process. The monitoring submodule 212 specifically includes a manager, a task queue, and specific monitoring and reporting units. The manager manages the computing submodules in the SP, including hardware initialization, DDR allocation, and data reading and writing functions. It can also receive inference tasks sent by the chip monitoring module 22 for computation. The task queue stores information related to inference tasks issued by the scheduler, i.e., inference and computing task information, such as model / operator addresses, input / output addresses, and inference task IDs, for the manager to read and execute sequentially. The computing monitoring unit monitors the status of each hardware component in the computing submodule, collects data such as the utilization rate of the computing units, memory capacity, and bandwidth usage in real time, and is responsible for reporting and saving this data to the sliding window queue of the chip monitoring module.
[0030] In this invention, while the computation submodule executes the inference task at full speed, the monitoring submodule simultaneously tracks the computation process with fine granularity and reports hardware status information in real time, including computing power utilization and data backlog rate. This real-time feedback of hardware status information enables the system to dynamically perceive the actual load characteristics of the inference task and the chip's instantaneous resource bottlenecks, thereby providing accurate data support for upper-level scheduling strategies. This allows the system to dynamically adjust resource allocation based on real-time feedback (such as elastic scaling or adjusting parallelism), ultimately significantly improving the resource utilization efficiency and operational stability of the AI chip in complex inference scenarios while ensuring low-latency task execution.
[0031] According to some embodiments, the chip monitoring module 22 includes: a sliding window submodule 221, used to acquire and store hardware status information of the stream processing module; and a scheduling submodule 222, used to select a subprocessing unit in the stream processing module based on the inference task, a pre-selected scheduling strategy and the hardware status information, and send the inference task to the subprocessing unit.
[0032] In this invention, the scheduling submodule 222 can receive inference tasks from the information interaction module 20, and then, based on the hardware status information (including computing unit utilization, memory capacity, and bandwidth usage) and pre-selected scheduling strategies collected in the sliding window submodule 221, select one or more SPs, and distribute the inference tasks to the task queues corresponding to the SPs. Different scheduling strategies select SPs in different ways, and the specific selection of one or more SPs needs to be determined according to the scheduling strategy.
[0033] The sliding window submodule 221 is used to collect and analyze the latest hardware status metrics. These metrics are reported in real-time by the monitoring submodule 212 within the stream processing module 21 and may include information such as computing unit utilization, memory capacity, and bandwidth usage. The data stored in the sliding window submodule 221 is the most up-to-date; older data is discarded during the storage process. The stored hardware status information can be provided to the scheduling submodule 222 for scheduling decisions. The sliding window submodule 221 receives reports from the monitoring submodule 212 on computing unit utilization, memory capacity, and bandwidth usage—the hardware status information. In some implementations, data is reported at fixed time thresholds, which can be configured in the driver, such as reporting once every 500ms / 1s. The number of sliding window queue windows is limited, for example, only 10 windows. After storing data for 10 time windows, the data for the 11th time window needs to overwrite the old data.
[0034] The sliding window submodule in this invention continuously acquires and caches the hardware status information of the stream processing module within the most recent time window, providing a basis for the computation of the scheduling submodule. When the scheduling submodule receives an inference task, it no longer relies solely on a static pre-selection scheduling strategy, but instead combines the hardware status trend provided by the sliding window to make a comprehensive judgment, accurately identifying the currently low-load and stable subprocessing unit as the target execution entity. This scheduling logic based on time window smoothing data effectively avoids task misscheduling or frequent migration caused by instantaneous glitches, ensuring that inference tasks are allocated to optimal computing resources while significantly improving the system stability and energy efficiency of the chip under high-load scenarios.
[0035] According to some embodiments, when the pre-selected scheduling strategy is a delay-first algorithm, the scheduling submodule 222 is specifically used to: determine the overall utilization rate of the sub-processing units in the stream processing module based on the preset algorithm, the task information to be scheduled in the hardware status information, the tensor core operation results and the vector core operation results, and sort them according to the overall utilization rate; based on the sorting, allocate the inference task to the corresponding sub-processing unit, and update the sorting based on the updated overall utilization rate after allocation.
[0036] In this invention, the minimum latency of the inference task can be used as the scheduling priority. New inference tasks are scheduled to be computed on the SP with the lightest load to ensure the lowest inference latency, i.e., the latency-first algorithm.
[0037] It can periodically query and analyze data in the sliding window queue, based on Tensor Core / Vector Core utilization metrics (tensor core operation results and vector core operation results) and information on tasks to be scheduled. In some implementations, a single inference computation may involve both Tensor Cores and Vector Cores. When calculating the load, it mainly calculates the utilization of Tensor Cores / Vector Cores simultaneously within the current unit of time, and obtains a comprehensive calculation (not a single metric).
[0038] In some implementations, the overall utilization rate is calculated as (Tensor utilization rate × 70% + Vector utilization rate × 30%). The SPs are then weighted and sorted using a pre-defined algorithm to obtain their overall utilization rate, and then sorted from lowest to highest utilization rate. The pre-defined algorithm is a weighted calculation method for overall utilization rate. Based on the overall utilization rate, the SP with the lowest utilization rate is selected, as a low utilization rate indicates that the SP is relatively idle and has fewer tasks.
[0039] Upon receiving a new inference task, the SP with the lowest overall utilization is selected first. If the SP's DDR (Memory Memory) is sufficient to accommodate the current inference task, the task is scheduled to that SP, the number of tasks pending scheduling in that SP is incremented by 1, the SP's overall utilization is updated, and the SPs are reordered. If the SP with the lowest utilization does not have enough DDR to accommodate the current task, the next lowest utilization SP is selected, and this process continues until the inference task is assigned to an SP. If none of the SPs can accommodate the current inference task, the process waits, updates the SP data, and ensures sufficient resources are available to schedule the current inference task. For the next inference task, the SP with the lowest overall utilization is selected again; here, the SP is the corresponding sub-processing unit.
[0040] refer to Figure 4 , Figure 4 It includes four SPs named SP1-SP4, along with their corresponding task allocation and hardware utilization. The size of the TASK block represents the estimated computational power required by the code. This process is a preliminary assessment in task scheduling or project management, determining in advance how much computational power and how long it will take to complete the code task that is about to run.
[0041] After calculating the overall hardware utilization of SP1-SP4, they are sorted from highest to lowest idle computing power. Specifically: when TASK1 arrives, it is assigned to SP1, which has the most remaining computing power; after adding TASK1, SP1's idle computing power is less than SP4's, so SP1 is moved to SP4. At this point, SP3 has the most remaining computing power. When TASK2 arrives, it is assigned to SP3, which has the most remaining computing power; after adding TASK2, SP3's idle computing power is the lowest, so SP1 is moved to SP2. At this point, SP4 has the most remaining computing power. When TASK3 arrives, TASK2 is assigned to SP4, which has the most remaining computing power. The scheduling is complete. This is the latency-first scheduling process.
[0042] The scheduling submodule of this invention first integrates the task requirements to be scheduled from the hardware status information and combines the real-time calculation results of the tensor core (responsible for matrix operations) and the vector core (responsible for vector operations) to accurately calculate the current overall utilization rate of each sub-processing unit (SP) in the stream processing module. This indicator can comprehensively reflect the true busyness of the SP when handling mixed AI workloads. Subsequently, the module sorts all sub-processing units in ascending or descending order according to the overall utilization rate, and prioritizes the allocation of inference tasks to the sub-processing unit with the lowest current overall utilization rate (i.e., the least idle computing resources), thereby minimizing the queuing time of tasks in the queue. After the task allocation is completed, the system immediately recalculates and updates the overall utilization rate ranking based on the new load situation. This dynamic rolling scheduling mechanism ensures that the system can perceive and avoid local computing hotspots in real time, effectively preventing head-of-queue blocking caused by the overload of a single computing unit. Ultimately, while ensuring extremely low response latency for inference tasks in high-concurrency scenarios, it achieves ultimate load balancing of heterogeneous computing resources within the chip.
[0043] According to some embodiments, when the pre-selected scheduling strategy is a memory-first algorithm, the scheduling submodule 222 is specifically used to: sort the sub-processing units in the stream processing module according to the remaining capacity of the sub-processing units in the hardware status information; and based on the sorting, allocate the inference task to the corresponding sub-processing unit and update the corresponding remaining capacity.
[0044] In this invention, the goal is to prioritize memory usage so that the entire AI chip can accommodate the most inference tasks. By maximizing the utilization of SP's memory, more inference tasks can be accommodated at the same time, thereby improving the parallelism of computation. This is achieved by using a memory-first scheduling strategy.
[0045] The system can periodically query and analyze the data in the sliding window queue, sorting the remaining space of the SP DDR from low to high. Upon receiving a new inference task, it iterates through the SPs from low to high according to their remaining DDR capacity. If the current SP has enough capacity to accommodate the new inference task, it places the task in that SP and updates the remaining capacity of the SP. If none of the SP DDRs are sufficient to accommodate the current task, it waits until the SP data is updated and there is enough DDR to schedule the current inference task. For the next inference task, the system continues to select the SP with the lowest remaining capacity.
[0046] In some implementations, refer to, for example Figure 5 , Figure 5 It contains four SPs named SP1-SP4, along with their corresponding task allocation and memory usage. The size of the TASK block is the amount of memory required for the code. In other words, when creating or scheduling a task (TASK block), all the memory resources required for its execution need to be evaluated and allocated in advance. This memory size determines how much space the task occupies in the chip's memory.
[0047] Sort SP1-SP4 in ascending order of remaining memory. SP1 has the least remaining memory. When TASK1 arrives, SP1, with the least memory space, has enough to accommodate it, so TASK1 is allocated to SP1. When TASK2 arrives, SP1 does not have enough memory, so SP2 is checked. SP2 has enough memory to accommodate TASK2, so TASK2 is moved to SP2, and the remaining memory size of SP2 is updated. When TASK3 arrives, neither SP1 nor SP2 has enough space to accommodate it, so SP4 is checked. SP4 has enough memory to accommodate TASK3, so TASK3 is scheduled to SP4. This is the memory-first scheduling process.
[0048] According to some embodiments, the sliding window submodule 221 includes multiple data blocks for storing multiple sets of hardware indicator data. When the hardware indicator data is updated, the sliding window submodule 221 detects the current position of the head pointer and slides the current position of the head pointer to the right along with the window to generate the next position of the head pointer; the updated hardware indicator data is stored in the data block corresponding to the current position to overwrite the original data in the data block. If the current position corresponds to the position of the last window of the sliding window submodule 221 block, the next position is the position corresponding to the first window of the sliding window submodule 221.
[0049] The sliding window submodule 221 of this invention consists of multiple data blocks that store hardware indicator data. Each data block stores the hardware indicator data of all SPs within a certain time period. A HEAD pointer, i.e., the head of the queue pointer, can be preset to identify the oldest data. Each window stores data within a time window. When the queue is full, new data will overwrite the data in the old time window, thus realizing the writing of new data and the eviction of old data (or empty data).
[0050] When a new SP hardware specification is updated, the sliding window submodule 221 updates the data at the position pointed to by the HEAD pointer and moves the HEAD pointer to the next window. If the queue reaches the end, the next window for HEAD will jump back to the first data block.
[0051] Each time, the indicator data from all windows in the sliding window submodule 221 is read and analyzed. The data update process can be found in [reference needed]. Figure 6 Assume that the time sliding window has five data blocks for storing data.
[0052] The metrics for time 1 are written to the first window block pointed to by HEAD, then HEAD moves to the NEXT position, pointing to the second window block; the metrics for time 2 are written to the second window block pointed to by HEAD, then HEAD moves to the NEXT position, pointing to the third window block; and so on. The metrics for time 5 are written to the fifth window block pointed to by HEAD, then HEAD moves to the NEXT position, pointing to the first window block; the metrics for time 6 are written to the first window block pointed to by HEAD (directly overwriting the metrics for time 1), then HEAD moves to the NEXT position, pointing to the second window block. At this point, the queue contains the metrics for times 6 / 2 / 3 / 4 / 5; the metrics for time 7 are written to the second window block pointed to by HEAD (directly overwriting the metrics for time 2), then HEAD moves to the NEXT position, pointing to the third window block; at this point, the queue contains the metrics for times 6 / 7 / 3 / 4 / 5. Following this process, a sliding window approach is used to achieve minimum-cost updates to hardware performance metrics.
[0053] The submodule in this invention organizes multiple data blocks into a logically closed-loop queue. In dynamic scenarios where hardware indicator data is continuously updated, it accurately locates the storage unit corresponding to the current time window by maintaining the unidirectional sliding trajectory of the head pointer. When the pointer moves to the end of the queue, it automatically wraps back to the head. This end-to-end cyclical overlay mechanism allows the latest generated hardware indicator data to replace the oldest historical data in real time. Thus, without performing high-overhead data movement and memory reallocation operations, it maintains a fixed window length of time-series data samples. This design not only greatly reduces storage bandwidth consumption and system latency but also ensures that the scheduling submodule can make decisions based on the real hardware state trend over the most recent continuous period, effectively filtering out noise data caused by voltage fluctuations or instantaneous interference, and significantly improving the scheduling stability and robustness of the stream processing module under complex operating conditions.
[0054] According to some embodiments, when the pre-selected scheduling strategy is empty, the chip monitoring module 22 performs scheduling based on a polling algorithm.
[0055] In this invention, if the user does not select a scheduling strategy during the input reasoning task stage, i.e. the pre-selected scheduling strategy is empty, the round-robin algorithm is directly used as the scheduling strategy during actual operation.
[0056] When the system lacks optimization preferences for specific hardware states, the polling algorithm maintains a logical pointer to a sub-processing unit and traverses all computing resources in the stream processing module in a fixed order. Upon receiving a new inference task, the scheduling submodule only needs to assign the task to the next sub-processing unit pointed to by the pointer, and then advance (or loop back) the pointer after allocation, without performing complex comparison, calculation, or sorting operations. This scheduling logic simplifies the originally complex decision-making process into a simple sequential addressing, minimizing the computational latency and power consumption of the scheduling submodule itself, while forcibly distributing tasks evenly across various computing units. This eliminates the possibility of localized idle computing power and overheating accumulation at the micro level, providing the most basic, fair, and stable load balancing guarantee for the chip.
[0057] The following describes method embodiments of the present invention, which can be operated in the system embodiments of the present invention. For details not disclosed in the method embodiments of the present invention, please refer to the system embodiments of the present invention.
[0058] Figure 7 A flowchart illustrating a resource scheduling method based on an artificial intelligence chip provided in an embodiment of the present invention. Figure 7 As shown, the resource scheduling method based on artificial intelligence chips includes steps S70 and S71.
[0059] In step S70, the inference task and pre-selected scheduling strategy input by the user are received.
[0060] In step S71, based on the inference task, the pre-selected scheduling strategy, and the artificial intelligence chip, the computation process of the inference task is executed and monitored to generate the task execution result.
[0061] This invention can first initialize the overall system, and then start it up. Figure 1 During the power-on process, the host checks all PCIe (Peripheral Component Interconnect Express) devices on the host, identifies the AI chip via its device ID, loads the corresponding driver, and begins the AI chip startup. Once the AI chip powers on and is activated, it waits for instructions from the host. The host loads the AI chip's initialization program into the AI chip via PCIe and the driver, and the AI chip begins executing the initialization program. The sliding window submodule starts, emptys all indicator data windows in the queue, and begins receiving new data. The scheduling strategy library loads the algorithm corresponding to the preset scheduling strategy and sets the initial algorithm as the default algorithm. The scheduler starts, loads the preset scheduling strategy from the scheduling strategy library, receives inference tasks from the host, reads data from the sliding window submodule, and begins task scheduling based on the preset scheduling strategy and the currently issued computation tasks.
[0062] The SP computing unit (computing submodule) starts up, and components such as the Tensor Core, Vector Core, and DDR power on, complete initialization, and begin reporting their own status information through built-in modules. The SP monitoring module (monitoring submodule) starts up, and the monitoring and reporting module obtains relevant hardware status from the computing unit in real time and reports the data to the sliding window submodule. The task queue is initialized and emptyed; the manager starts polling to check if there are any inference tasks in the task queue. After receiving the signal that the chip initialization is complete, the host sends the scheduling policy to the chip according to the user-configured scheduling policy (or the default scheduling policy if none is available). The chip's scheduler (scheduling submodule) loads the specified scheduling policy according to the configuration sent by the host.
[0063] The host receives inference tasks sent by the application and transmits them to the AI chip via PCIe for inference computation. The scheduler loads collected SP (Service Provider) metric data from the sliding window queue, analyzes and integrates it. Based on the inference task information (inference type, memory usage, etc.), the processed metric data of each SP, and the scheduling policy, the scheduler selects the optimal SP and sends the inference task information to the selected SP. Specifically, it can load the scheduling policy according to the default or user selection, obtain the latest hardware data from the sliding window queue, schedule the TASK to the corresponding SP according to the configured policy, write the model information and input to the SP's corresponding DDR, and add the push task to the SP's task queue.
[0064] The SP monitoring manager polls the task queue in real time. If the queue is not empty, it schedules inference tasks to the compute units for execution in sequence. Based on the received inference tasks, the SP monitoring system sequentially calls the Tensor Cores / Vector Cores to perform inference computations according to the model. After computation is complete, it notifies the controller and returns the inference results. Real-time collection of hardware metrics provides data support for the scheduler. The monitoring submodule polls the SP compute units in real time, reading performance metrics from the compute unit's hardware registers, including Tensor Core / Vector Core utilization, DDR capacity usage, and bandwidth usage. The monitoring submodule then reports the collected metrics to the chip monitoring module.
[0065] The sliding window queue of the chip monitoring module stores hardware metrics reported from the SP monitoring module. If the queue is full, expired metric data will be discarded.
[0066] The function implemented by this method is similar to the function of the system execution provided earlier. Other functions can be found in the previous descriptions and will not be repeated here.
[0067] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 8 As shown, the electronic device 800 of this embodiment may include a memory 801 and a processor 802.
[0068] The memory 801 stores a computer program, which, when executed by the processor 802, causes the processor 802 to perform the method described in the above embodiments.
[0069] The processor 802 and the memory 801 are connected, for example, via a bus.
[0070] Optionally, the electronic device 800 may also include a transceiver. It should be noted that in practical applications, the transceiver is not limited to one, and the structure of the electronic device 800 does not constitute a limitation on the embodiments of the present invention.
[0071] Processor 802 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 802 may also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0072] A bus can include a pathway for transmitting information between the aforementioned components. The bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one thick line is used in the diagram, but this does not imply that there is only one bus or one type of bus.
[0073] The memory 801 can be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or it can be EEPROM (Electrically Erasable Programmable Read Only Memory), CD. ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed discs, laser discs, optical discs, digital universal discs, Blu-ray discs, etc.), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0074] The memory 801 stores application code that executes the present invention, and its execution is controlled by the processor 802. The processor 802 executes the application code stored in the memory 801 to implement the content shown in the foregoing method embodiments.
[0075] Electronic devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Servers can also be included. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0076] The electronic device in this embodiment can be used to execute the method of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
[0077] The present invention also provides a non-transitory computer-readable storage medium having stored computer-readable instructions thereon, which, when executed by a processor, cause the processor to perform the method as described in the above embodiments.
[0078] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a non-transitory computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0079] The embodiments of the present invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, any changes or modifications made by those skilled in the art based on the ideas of the present invention, its specific implementation methods, and its application scope, are all within the scope of protection of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A resource scheduling system based on an artificial intelligence chip, characterized in that, include: The information interaction module is used to receive user input of inference tasks and pre-selected scheduling strategies; A stream processing module, integrated on an artificial intelligence chip, is used to execute and monitor the computation process of the inference task in order to generate task execution results; A chip monitoring module is integrated on the artificial intelligence chip. The chip monitoring module is used to schedule the stream processing module based on the inference task and the pre-selected scheduling strategy.
2. The system according to claim 1, characterized in that, The stream processing module includes: The calculation submodule is used to execute the calculation process of the inference task to generate the task execution result; The monitoring submodule is used to monitor the computation process of the inference task and report hardware status information in real time.
3. The system according to claim 1, characterized in that, The chip monitoring module includes: The sliding window submodule is used to acquire and store the hardware status information of the stream processing module; The scheduling submodule is used to select a subprocessing unit in the stream processing module based on the inference task, the pre-selected scheduling strategy, and the hardware status information, and send the inference task to the subprocessing unit.
4. The system according to claim 3, characterized in that, When the pre-selected scheduling strategy is a delay-first algorithm, the scheduling submodule is specifically used for: Based on the preset algorithm, the task information to be scheduled in the hardware status information, the tensor core operation results, and the vector core operation results, the overall utilization rate of the sub-processing units in the stream processing module is determined, and they are sorted according to the overall utilization rate. Based on the ranking, the reasoning task is assigned to the corresponding sub-processing unit, and the ranking is updated based on the overall utilization rate updated after the assignment.
5. The system according to claim 3, characterized in that, When the pre-selected scheduling strategy is a memory-first algorithm, the scheduling submodule is specifically used for: Based on the remaining capacity of the sub-processing units in the hardware status information, the sub-processing units in the stream processing module are sorted. Based on the sorting, the inference task is assigned to the corresponding sub-processing unit, and the corresponding remaining capacity is updated.
6. The system according to claim 3, characterized in that, The sliding window submodule includes multiple data blocks for storing multiple sets of hardware indicator data.
7. The system according to claim 6, characterized in that, When the hardware indicator data is updated, the sliding window submodule detects the current position of the head pointer and slides the current position of the head pointer to the right along with the window to generate the next position of the head pointer; The updated hardware metrics data is stored in the data block corresponding to the current location, thereby overwriting the original data in the data block.
8. The system according to claim 7, characterized in that, If the current position corresponds to the position of the last window of the sliding window submodule, the next position corresponds to the position of the first window of the sliding window submodule.
9. The system according to claim 1, characterized in that, When the pre-selected scheduling strategy is empty, the chip monitoring module performs scheduling based on a round-robin algorithm.
10. A resource scheduling method based on an artificial intelligence chip, characterized in that, include: Receive user input for inference tasks and pre-selected scheduling strategies; Based on the inference task, the pre-selected scheduling strategy, and the artificial intelligence chip, the computation process of the inference task is executed and monitored to generate task execution results.