NPU Computing Task Scheduling Method, Device and Equipment for Multi-Mode SoC Master Control Chip
By building a two-layer grouping scheduling mechanism and Q-learning task optimization model, combined with Pareto multi-objective optimization algorithm, the problems of unbalanced resource utilization and large task delay in traditional NPU computing task scheduling methods are solved, and more efficient task scheduling and resource management are achieved.
Patent Information
- Application Number
- CN202510217350.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional NPU computing task scheduling methods are difficult to effectively deal with dynamically changing computing loads, resulting in unbalanced computing resource utilization and large task execution delays. Especially in the scenario where multiple computing modes coexist, there are problems of time overlap and resource competition.
By extracting the task feature parameters of the NPU calculation task, generating task descriptors, and establishing a resource state matrix based on the initial scheduling timing table, calculating task overlap, building a Q-learning task optimization model, performing two-layer grouping and sorting tasks, generating task allocation sequences, performing resource load calculations, generating migration cost matrix, performing task reassignment, and finally generating global execution timing through the Pareto multi-objective optimization algorithm for task scheduling.
It achieves a balance between task execution delay and resource utilization, ensures the overall optimization of the scheduling scheme, improves the system's response ability and processing efficiency to burst tasks, reduces the additional overhead of task migration, and optimizes the system's energy efficiency.
Smart Images

Figure CN119718590B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of task scheduling, and particularly relates to a method, device, and equipment for NPU computing task scheduling of a multi-mode SoC master chip. Background Art
[0002] With the rapid development of artificial intelligence technology, the number of NPU computing units integrated in multi-mode SoC master chips is increasing day by day, and the types of computing tasks are becoming increasingly complex and diverse. In practical applications, the NPU needs to process periodic computing tasks and burst computing tasks simultaneously, which poses higher requirements for the scheduling and management of computing resources.
[0003] Traditional NPU computing task scheduling methods mainly adopt static scheduling strategies, which cannot effectively cope with dynamically changing computing loads, resulting in unbalanced utilization of computing resources and relatively large task execution delays. Especially in scenarios where multiple computing modes coexist, there are time overlaps and resource competition problems between NPU computing tasks of different modes, and it is difficult for traditional scheduling methods to achieve reasonable allocation of computing resources while ensuring task execution efficiency. Summary of the Invention
[0004] The main purpose of the present invention is to provide a method, device, and equipment for NPU computing task scheduling of a multi-mode SoC master chip. The present invention achieves a balance between task execution delay and resource utilization rate, ensures the overall optimality of the scheduling scheme, and improves the system's response ability and processing efficiency for burst tasks.
[0005] To achieve the above object, the present invention provides a method for NPU computing task scheduling of a multi-mode SoC master chip, including the following steps:
[0006] Extract the task feature parameters of the NPU computing tasks in the multi-mode SoC master chip, generate a task descriptor according to the task feature parameters, and establish an initial scheduling time sequence table based on the task descriptor;
[0007] Perform partition division on the computing resources of the NPU computing unit to obtain multiple computing resource partitions, and obtain the performance parameters of each computing resource partition, and establish a resource status matrix according to the performance parameters;
[0008] Calculate the task overlap degree based on the initial scheduling time sequence table and the resource status matrix, generate an overlapping time sequence matrix, and construct a Q-learning task optimization model according to the overlapping time sequence matrix;
[0009] Perform two-layer grouping and sorting on the NPU computing tasks according to the Q-learning task optimization model, and generate a task allocation sequence based on the resource status matrix;
[0010] Calculate the resource load of the NPU computing tasks in the task assignment sequence, generate a migration cost matrix, and perform task reallocation according to the migration cost matrix to obtain a task reallocation result;
[0011] Input the task reallocation result into the Pareto multi-objective optimization algorithm to generate a global execution timing sequence of the NPU computing tasks, and perform task scheduling on the NPU computing tasks according to the global execution timing sequence, and output task scheduling execution information.
[0012] The present invention also provides an NPU computing task scheduling device for a multi-mode SoC main control chip, including:
[0013] An extraction module, configured to extract task feature parameters of NPU computing tasks in a multi-mode SoC main control chip, generate a task descriptor according to the task feature parameters, and establish an initial scheduling timing table based on the task descriptor;
[0014] A partitioning module, configured to perform calculation resource partitioning on NPU computing units to obtain multiple calculation resource partitions, obtain performance parameters of each calculation resource partition, and establish a resource status matrix according to the performance parameters;
[0015] A calculation module, configured to calculate a task overlap degree based on the initial scheduling timing table and the resource status matrix, generate an overlapping timing matrix, and construct a Q-learning task optimization model according to the overlapping timing matrix;
[0016] A sorting module, configured to perform two-layer grouping and sorting on NPU computing tasks according to the Q-learning task optimization model, and generate a task assignment sequence based on the resource status matrix;
[0017] A reallocation module, configured to calculate the resource load of the NPU computing tasks in the task assignment sequence, generate a migration cost matrix, and perform task reallocation according to the migration cost matrix to obtain a task reallocation result;
[0018] An output module, configured to input the task reallocation result into the Pareto multi-objective optimization algorithm to generate a global execution timing sequence of the NPU computing tasks, and perform task scheduling on the NPU computing tasks according to the global execution timing sequence, and output task scheduling execution information.
[0019] The present invention also provides a computer device, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0020] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0021] In summary, the technical solution provided by the present invention realizes the classification processing of periodic tasks and burst tasks by constructing a two-layer packet scheduling mechanism, effectively reduces task scheduling conflicts, and improves the utilization efficiency of computing resources. The Q-learning task optimization model is used to dynamically evaluate the task overlap degree, accurately grasp the task execution characteristics, avoid resource allocation conflicts, and reduce task queuing delays. The migration cost matrix is introduced to make task reallocation decisions, comprehensively considering time overhead and energy consumption factors, reducing the additional overhead of task migration, and optimizing the system energy efficiency. The global execution timing is generated based on the Pareto multi-objective optimization algorithm, achieving a balance between task execution delay and resource utilization rate, and ensuring the overall optimality of the scheduling scheme. A method for extracting task feature parameters and constructing a resource status matrix is designed to realize the real-time perception of NPU computing tasks and resource status. Through resource partitioning and dynamic integration strategies, the flexible allocation of computing resources is realized, and the system's response ability and processing efficiency for burst tasks are improved. A complete task scheduling execution information feedback mechanism is established to support the real-time optimization and adjustment of scheduling strategies, enhancing the adaptability and reliability of the system. Description of the Drawings
[0022] Figure 1 is a schematic diagram of the steps of the NPU computing task scheduling method for the multi-mode SoC main control chip in an embodiment of the present invention;
[0023] Figure 2 is a block diagram of the structure of the NPU computing task scheduling device for the multi-mode SoC main control chip in an embodiment of the present invention;
[0024] Figure 3 is a schematic block diagram of the structure of a computer device in an embodiment of the present invention.
[0025] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0026] In order to make the object, technical solution and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0027] Referring to Figure 1 , this embodiment provides a method for scheduling NPU computing tasks of a multi-mode SoC main control chip, including the following steps:
[0028] S1, extracting the task characteristic parameters of the NPU computing task in the multi-mode SoC main control chip, generating a task descriptor according to the task characteristic parameters, and establishing an initial scheduling timing table based on the task descriptor;
[0029] Among them, by collecting features of the NPU computing tasks in the chip, the task feature data such as the execution time, computing resource requirements, task priority, computing mode type, data dependency and task periodicity of each task are obtained. The execution time and resource requirements in the task feature data are numerically standardized to eliminate the differences in different feature dimensions, so that all feature data can be within the same standardized range, which is convenient for subsequent comparison and analysis, and a standardized feature vector is obtained. The standardized feature vector is combined with the priority value of the task for multi-dimensional feature fusion to obtain the task feature parameters. The task feature parameters are input into the feature mapping module for mapping to obtain the task descriptor. The task descriptor is a concentration of task information, including task identifier, computing mode identifier, resource requirement vector, task priority value, expected execution time, and also includes a dependency matrix between tasks. The dependency matrix between tasks can accurately describe the execution order relationship between tasks, ensuring that there will be no task conflict and resource competition during the scheduling process. The NPU computing tasks are classified according to the computing mode identifier in the task descriptor. The NPU computing tasks of the multi-mode SoC chip have different computing modes, such as serial mode, parallel mode or hybrid mode. By classifying the computing modes, a more accurate basis is provided for subsequent scheduling. After the mode classification is completed, the tasks in each mode classification are sorted according to the priority value of the task to obtain the priority sequence of task execution. The time trigger window is calculated based on the task execution priority sequence and the dependency matrix between tasks. A suitable time period is allocated to each task to ensure that the task can be executed on time under the premise of satisfying the dependency. When calculating the time trigger window, the periodicity of the task and the dependency between tasks are considered. The earliest start time and end time of each task are calculated through the scheduling algorithm, and the time resources are reasonably allocated to obtain a time allocation plan including the task start time point and end time point. This plan can effectively prevent time conflicts and resource competition between tasks, while ensuring that tasks can be executed at the scheduled time. The time allocation plan is scheduled and mapped with the task computing resource demand vector to obtain the initial scheduling time sequence table containing the time trigger window and computing resource allocation information of each task.
[0030] S2, partitioning the computing resources of the NPU computing unit to obtain multiple computing resource partitions, and obtaining performance parameters of each computing resource partition, and establishing a resource status matrix according to the performance parameters;
[0031] Specifically, resource statistical analysis is performed on the NPU computing units to obtain computing resource pool data, which includes information such as the number of computing units, the computing power of each computing unit, and the memory bandwidth. The computing power includes the number of floating-point operations that each computing unit can process per second, and the memory bandwidth represents the data transfer speed and bandwidth size. Based on the computing resource pool data, partitioned computing is performed on the NPU computing units. The computing resources of the NPU are divided into several partitions to achieve independent resource management for different computing tasks. The initial partition scheme is roughly divided according to factors such as the number of computing units and computing power. An equilibrium verification calculation is performed on the initial partition scheme to ensure that the load of each computing resource partition can be as balanced as possible, and there will be no situation where some partitions are overloaded with resources while other partitions are idle. Through the equilibrium verification, a processing capacity distribution map of the computing resource partitions is obtained, showing the resource allocation and computing capacity distribution of each partition. Performance parameters are collected for each computing resource partition to obtain partition performance data including the number of computing units, peak computing performance, and memory bandwidth. These performance data can help the scheduling algorithm evaluate the potential and limitations of each computing resource partition, clarify the actual computing power and bandwidth of each partition, and ensure that tasks can be reasonably scheduled according to the resource status of the partitions. The load rate is calculated for the partition performance data. The load rate is an important indicator for evaluating the current usage of each computing resource partition, which is represented by the ratio between the actual load and the maximum load. If the load rate of a partition is too high, it means that the partition is approaching saturation and cannot process more tasks; conversely, a partition with a lower load rate means that there are remaining resources and it is suitable to receive more computing tasks. Through the load rate calculation, the current load status of each computing resource partition is obtained. Based on the load status of each computing resource partition, the remaining computing power is calculated to quantify the computing power that is not currently occupied in each computing resource partition. The remaining computing power value reflects the available processing power of each partition, helps the scheduling algorithm determine which tasks can be assigned to which partitions, and avoids assigning tasks to resource partitions that are already approaching saturation, thereby reducing task execution latency and resource conflicts. According to the remaining computing power values of each computing resource partition, the task queue length is analyzed to determine how many new tasks each partition can currently receive. The longer the queue length, the more tasks the partition has received, and it is necessary to consider whether to increase resources or migrate tasks to other partitions. The determination of the task reception threshold needs to consider not only the remaining computing power of the partition but also the current task queue length. Through the analysis of the task queue length, the task reception threshold of each computing resource partition is obtained. The task reception threshold is fused with the current load status to obtain a resource status matrix, which includes key information such as the load rate, remaining computing power, and task queue length of each computing resource partition.
[0032] S3. Calculate the task overlap degree based on the initial scheduling time sequence table and the resource status matrix, generate an overlapping time sequence matrix, and construct a Q-learning task optimization model according to the overlapping time sequence matrix;
[0033] It should be noted that by performing a time interval analysis on the task time trigger windows in the initial scheduling time sequence table, the start and end times of each task are extracted, obtaining a set that contains the time intervals of all tasks. The time interval of a task represents the execution period of each task and defines the time window required by the task on the resources. By analyzing the set of time intervals, the time overlap situation of any two NPU computing tasks is calculated, obtaining a task pair overlap duration matrix that records the time overlap duration between each pair of tasks. Based on the load rate data in the task pair overlap duration matrix and the resource status matrix, a task overlap score calculation is performed. The load rate reflects the current usage situation of each computing resource partition. Therefore, whether a task can be successfully executed during the overlap period also needs to consider the actual load situation of the resources. The goal of the overlap score calculation is to evaluate the availability of system resources and the possibility of resource conflicts during the task overlap period, thereby providing a reasonable decision-making basis for task scheduling. Through the overlap score calculation, an overlap degree score vector is obtained, which contains the scores of each pair of tasks during the overlap period. The overlap degree score vector is sorted by time interval, and the task pairs with higher scores are given priority, obtaining an overlap time sequence matrix that contains the overlap task identifiers and overlap time sequence information. This matrix shows the time overlap situation between tasks and reflects the competition relationship between different tasks on system resources. Based on the overlap time sequence matrix, the state space of the Q-learning task optimization model is constructed. The Q-learning algorithm optimizes the decision-making strategy according to the state and action of the environment. The state space consists of task state features and resource state features. Task state features include the priority, resource requirements, computing mode, etc. of the task, while resource state features include the load rate, remaining computing power, and available memory bandwidth of each computing resource partition. By fusing these features, a state vector is obtained, representing the overall state of the current system. An action space is constructed based on the state vector. The action space contains task scheduling actions and resource allocation actions. Task scheduling actions refer to how to arrange the execution order and scheduling priority of tasks within a given time window; resource allocation actions are how to reasonably allocate computing resources to each task according to the resource requirements of the tasks and the remaining capabilities of the resources. The selection of these actions directly affects the delay of task execution, the utilization efficiency of resources, and the overall performance of the system. A Q-value update function is designed according to the action set. The Q-value update function is the core of the Q-learning algorithm, and it adjusts the task scheduling strategy by evaluating rewards and punishments. The reward function mainly considers the delay of task execution and resource utilization rate. Delay punishment means that the longer the task execution time, the greater the punishment; while the resource utilization rate reward means that the more effectively the task can utilize computing resources, the higher the overall reward of the system. By comprehensively evaluating the punishment for task delay and the reward for resource utilization rate during the task scheduling process, the scheduling strategy is effectively guided towards the optimal state. Combining the state vector, action set, and reward function, a Q-learning task optimization model is constructed.The goal of this model is to learn how to select optimal task scheduling and resource allocation actions according to the current state of the system through the Q-learning algorithm, so as to maximize the performance of the system. In this process, the Q-learning model continuously obtains feedback from the environment, updates the Q value, optimizes the scheduling strategy, and finally outputs a task stacking scheme, indicating the execution order and required resources of each task in the system.
[0034] S4. Double-layer group sorting is performed on the NPU computing tasks according to the Q-learning task optimization model, and a task allocation sequence is generated based on the resource status matrix;
[0035] Specifically, periodic computing tasks are extracted from the Q-learning model. These tasks have fixed periods and predictable execution times, and have high determinacy during scheduling. By extracting these periodic tasks, the first-layer periodic task set is formed. Based on the load rate data in the resource status matrix, static partition allocation is performed on the periodic task set. Static partition allocation means that the resource requirements of periodic tasks have been predetermined, and these tasks can be reasonably allocated according to the load rate of each computing resource partition. The load rate reflects the current resource usage of each partition. Partitions with lower load rates are preferentially allocated periodic tasks to achieve efficient use of resources. Through this process, an allocation scheme for periodic tasks is generated to ensure that periodic tasks can be reasonably scheduled on appropriate computing resources and avoid excessive resource competition. For the second-layer tasks, namely burst tasks, their characteristics are unpredictable execution durations, large fluctuations in resource requirements, and may occur at any time, requiring flexible scheduling strategies. In the Q-learning model, the characteristics of burst tasks are extracted through state vectors to form the second-layer burst task set. According to the characteristics of these tasks, dynamic partition calculation is performed based on the remaining computing power data in the resource status matrix. The remaining computing power represents the unoccupied processing capacity in the computing resource partition and is the key to determining whether a new burst task can be received. Through dynamic partition calculation, candidate partitions for burst tasks are obtained. The selection of candidate partitions for burst tasks depends on the remaining computing power of the partition and the availability of computing resources. Partitions with higher remaining computing power and sufficient resources are preferentially selected to effectively process burst tasks. Task compatibility analysis is performed based on the candidate partitions for burst tasks and the allocation scheme for periodic tasks to judge the conflicts and competitions in resource allocation between periodic tasks and burst tasks. The execution time of periodic tasks is determined, while burst tasks may occur at any time, resulting in resource competition. Through compatibility analysis, a task resource competition degree matrix is obtained, which reflects the intensity of resource competition between different tasks. This matrix helps the scheduling algorithm evaluate which tasks can share resources within the same time window and which tasks need to be allocated to different computing resource partitions to avoid resource conflicts and task delays. Based on the task resource competition degree matrix, partition selection calculation is performed for burst tasks. According to the resource requirements of the tasks, the remaining computing power of the candidate partitions, and the resource competition degree, the most suitable computing resource partition is selected to execute the burst tasks. The result of partition selection forms the allocation scheme for burst tasks. The allocation scheme for periodic tasks and the allocation scheme for burst tasks are merged to obtain a complete task grouping scheme. During the merging process, scheduling is performed according to factors such as task priorities, resource requirements, and time windows to ensure that all tasks can be executed smoothly on the same time axis and avoid resource conflicts or scheduling delays. Time window allocation calculation is performed on the complete task grouping scheme.According to the priorities and resource requirements of tasks, allocate appropriate execution times for each task to ensure that all tasks are executed within their specified time windows and the execution order among tasks meets the scheduling requirements. After the time window allocation is completed, generate an execution time series for each task. Based on the generated execution time series, perform the allocation of computing resources. Resource allocation needs to consider factors such as the computing power requirements, memory bandwidth requirements, and execution times of each task to achieve the optimal utilization of resources. Serialize the resource occupancy plan of tasks and the execution time series to obtain a task allocation sequence, clarifying the execution order of each task, the resource allocation situation, and the execution times of tasks.
[0036] S5. Calculate the resource load of the NPU computing tasks in the task allocation sequence, generate a migration cost matrix, and perform task reallocation according to the migration cost matrix to obtain the task reallocation result;
[0037] Among them, perform load statistical analysis on the NPU computing tasks in the task allocation sequence, evaluate the task load currently borne by each computing resource partition, and obtain the task load distribution data of each computing resource partition. Based on the task load distribution data, calculate the load balance degree of the computing resource partitions to measure whether the load of each partition is balanced. If the task load of some partitions is much higher than that of other partitions, it will lead to excessive consumption of resources and task delays. The calculation of the load balance degree can help identify the resource partitions with unbalanced loads, so as to determine which partitions need to perform task migration to achieve the purpose of balanced load. Calculate the migration time overhead of the tasks in the unbalanced partitions. The calculation of the migration time overhead refers to calculating the time required to migrate a task from one computing resource partition to another. This process needs to consider factors such as the size of the task, the bandwidth and latency of task migration. The migration time cost of each task is the time overhead required during its process from the source partition to the target partition. By calculating the migration time cost vector, quantify the migration cost of each task. Based on factors such as the computing resources consumed by task migration, the data transfer volume, and the network bandwidth, calculate the task migration energy consumption according to the task migration time cost vector. By calculating the task migration energy consumption, obtain the task migration energy consumption cost vector, which reflects the energy consumed by each task during migration. Fuse the task migration time cost vector and the task migration energy consumption cost vector to obtain the migration cost matrix and evaluate the cost of task migration. Based on the migration cost matrix, select the target partition for task migration for the resource partitions with a load rate lower than the preset threshold. When selecting the target partition during the task migration process, multiple factors need to be considered, including the remaining computing power of the target partition, the load rate, the time and energy consumption of task migration, etc. Partitions with a lower load rate are more suitable for receiving the migrated tasks. By analyzing the migration cost matrix, determine the most suitable target partition for task migration to obtain the task migration plan. The task migration plan includes the target partition for each task migration and its relevant parameters. Reallocate the NPU computing tasks based on the task migration plan. Reallocate the tasks to the computing resource partitions according to the migration plan, so that the task load is balanced, avoiding performance bottlenecks caused by overloading of some partitions, and obtaining the updated task allocation plan, which reflects the latest matching relationship between tasks and resource partitions. Reorganize the resources of the updated task allocation plan to obtain the task reallocation result including task migration information and resource allocation information. The purpose of resource reorganization is to optimize the overall use of computing resources, ensure the most reasonable configuration of resources for each computing resource partition, and at the same time ensure that tasks can be successfully executed under the new allocation plan. The reallocation result includes the migration path, migration time, and migration energy consumption of the tasks, as well as the new task allocation method and resource configuration.
[0038] S6. Input the task reallocation result into the Pareto multi-objective optimization algorithm to generate the global execution timing of the NPU computing tasks, and perform task scheduling on the NPU computing tasks according to the global execution timing, and output the task scheduling execution information.
[0039] Specifically, perform a time overhead analysis on the task execution delays in the task reallocation results to evaluate the delays of each task in its new execution path. Through comprehensive analysis of the task execution time delays, seek an optimal task execution order to minimize the overall task execution time as much as possible, and obtain the delay minimization objective function. At the same time, calculate the utilization rate of the computing resource allocation in the task reallocation results. By analyzing the usage of each computing resource partition, maximize the utilization efficiency of the computing resources to obtain the resource utilization maximization objective function. Combine the delay minimization objective function and the resource utilization maximization objective function to form a multi-objective space. In the multi-objective space, the two objective functions respectively represent the key factors that need to be balanced in task scheduling: on the one hand, it is the task execution delay, and on the other hand, it is the utilization efficiency of the computing resources. During the multi-objective optimization process, handle the potential conflicts between these two objectives. In the multi-objective space, make a reasonable compromise according to the actual requirements of the tasks and the current state of the resources to obtain a constrained optimal solution space. Perform a non-dominated solution search calculation on the multi-objective constraint space. In multi-objective optimization, find solutions that are not inferior to other solutions in all objectives. These solutions are Pareto optimal solutions, representing the optimal task execution timing and resource allocation strategies under the given objectives. Through the non-dominated solution search calculation, obtain a Pareto optimal solution set containing multiple groups of task execution timing and resource allocation strategies. Each group of solutions corresponds to a different task execution order and resource allocation scheme. Verify the dependency relationships according to the task execution timing in the Pareto optimal solution set. There are certain dependency relationships between tasks. For example, some tasks must start execution after other tasks are completed. The process of dependency relationship verification is to ensure that all task dependencies are satisfied during task execution. The result of the dependency relationship verification obtains a subset of execution timings that satisfy the task dependency constraints. These subsets represent the execution orders that meet the inter-task dependency requirements and can ensure that there will be no deadlocks or resource conflicts during the task scheduling process. Based on the subset of execution timings after the dependency relationship verification, perform allocation verification on the computing resources. Confirm that each task can obtain the required computing resources within the specified time, and check whether the resources will be overloaded and whether there is a risk of overload. After the verification is completed, obtain a global execution timing scheme that satisfies the resource constraints. Convert the global execution timing scheme into an execution instruction sequence for NPU computing tasks. The execution instructions include information such as the start time, end time, required resources, and execution order between tasks of each task. Partition and issue the task scheduling instruction set to tasks. Allocate the execution instructions of each task according to the division of the computing resource partitions and generate the execution instruction sequences for each computing resource partition. Distribute all the execution instruction sequences to the corresponding computing resource partitions for execution processing. When each computing resource partition executes its tasks, record and feedback information such as the task execution status, computing performance, and resource utilization rate.
[0040] In one example, task feature parameters of NPU computing tasks in a multi-mode SoC main control chip are extracted, a task descriptor is generated according to the task feature parameters, and an initial scheduling time sequence table is established based on the task descriptor, including:
[0041] Feature collection is performed on the NPU computing tasks in the multi-mode SoC main control chip to obtain task feature data including task execution duration, computing resource demand, task priority value, computing mode type, data dependency relationship, and task periodicity;
[0042] Numerical normalization processing is performed on the execution duration and resource demand in the task feature data to obtain a normalized feature vector, and multi-dimensional feature fusion is performed on the normalized feature vector and the priority value in the task feature data to obtain task feature parameters;
[0043] The task feature parameters are input into a feature mapping module for mapping to obtain a task descriptor including a task identifier, a computing mode identifier, a resource demand vector, a task priority value, an expected execution duration, and a task interdependency relationship matrix;
[0044] The NPU computing tasks are classified by mode according to the computing mode identifier in the task descriptor to obtain a mode classification set, and the tasks in the mode classification set are sorted according to the task priority value to obtain a task execution priority sequence;
[0045] Time-triggered window calculation is performed according to the task execution priority sequence and the task interdependency relationship matrix to obtain a time allocation plan including task start time points and end time points;
[0046] The time allocation plan is scheduled and mapped with the task computing resource demand vector to obtain an initial scheduling time sequence table including the time-triggered windows and computing resource allocation information of each task.
[0047] In this example, feature collection is performed on the NPU computing tasks to extract the specific requirements and resource demands of each task during execution. The task feature data includes the execution duration of the task, the computing resource demand, the task priority value, the computing mode type, the data dependency relationship, and the task periodicity, etc. These features can comprehensively reflect the behavioral characteristics of the task when running in the system. Through the task feature data, the system can understand the computing complexity, execution time requirements, and resource demands of the task. Numerical normalization processing is performed on the execution duration and computing resource demand in the task feature data to make the features of different tasks comparable numerically and avoid the influence of some features on the overall calculation due to too large or too small value ranges. The normalization processing adopts the min-max normalization or Z-score normalization method. Suppose the execution duration of the task is and the computing resource demand is , the standardized eigenvalue and are calculated using the following formula:
[0048] ;
[0049] ;
[0050] where, and are the minimum and maximum values of the execution durations of all tasks respectively, and are the minimum and maximum values of the resource requirements of all tasks respectively. Through the standardization process, it is ensured that the eigenvalue of the execution duration and the resource requirement of all tasks are within the same dimension range, avoiding the influence of too large a numerical range on the scheduling calculation. Multidimensional feature fusion is performed on the standardized eigenvector and the priority value in the task feature data. The priority value of the task is used to determine the priority order of the task during scheduling, taking values such as high, medium, low, etc., and numerical values are assigned. Assuming the priority of the task is , the standardized eigenvector and the priority value are fused into a comprehensive task feature vector by weighted average, and the calculation formula is as follows:
[0051] ;
[0052] where, represents the standardized execution duration vector of the task, represents the standardized resource requirement vector of the task, is the priority value of the task, , and are the weight coefficients of each feature, and these coefficients reflect the importance of different task features in task scheduling. Through multi-dimensional feature fusion, the comprehensive feature of a task is represented as a vector, which can fully reflect the various behavioral features of the task and facilitate subsequent scheduling decisions. The task feature parameters are input into the feature mapping module for mapping to obtain a task descriptor. The task descriptor includes the identifier of the task, as well as the computing mode identifier, resource requirement vector, task priority value, expected execution duration, and the dependency relationship matrix between tasks. Based on the computing mode identifier in the task descriptor, the NPU computing tasks are classified by mode. The tasks are classified according to the computing mode in order to select the most suitable computing resources for each type of task. For example, tasks are divided into two categories: processing-intensive tasks and data-intensive tasks. The former requires more computing resources, while the latter has higher requirements for memory bandwidth. After classifying the tasks, the tasks in the mode classification set are sorted according to the priority of the tasks to obtain the priority sequence of task execution. According to the task execution priority sequence and the dependency relationship matrix between tasks, the time-triggered window calculation is performed. The start and end time points of each task are determined according to the priority and dependency relationship of the tasks. Assume task has a dependency relationship of Dep , task has a start time point of , and an end time point of , then the time allocation scheme is calculated by the following formula:
[0053] ;
[0054] ;
[0055] where, represents the start time point of task , represents the end time point, is the normalized execution duration of task . Through calculation, a time window is allocated for each task to ensure that the tasks are executed in a reasonable order and time arrangement. The time allocation scheme is scheduled and mapped with the task computing resource requirement vector to obtain an initial scheduling time sequence table containing the time-triggered window and computing resource allocation information of each task. Assume task has a resource requirement vector of , then the resource allocation of the task is mapped in the following way:
[0056] ;
[0057] where, It is a scheduling mapping function that calculates the computing resource allocation required for a task based on the task's resource requirements and time window. The initial scheduling time sequence table lists the start time, end time, resource requirements, and allocation of all tasks, providing a detailed execution plan for subsequent scheduling execution.
[0058] In one example, the computing resources of the NPU computing unit are partitioned to obtain multiple computing resource partitions, and the performance parameters of each computing resource partition are obtained. A resource status matrix is established based on the performance parameters, including:
[0059] Perform resource statistical analysis on the NPU computing unit to obtain computing resource pool data including the number of computing units, computing power, and memory bandwidth;
[0060] Based on the computing resource pool data, perform partition calculations on the NPU computing unit to obtain an initial partition scheme for multiple computing resource partitions, and perform balance verification calculations on the initial partition scheme of the computing resource partitions to obtain a processing capacity distribution map of the computing resource partitions;
[0061] Collect performance parameters for each computing resource partition according to the processing capacity distribution map of the computing resource partitions to obtain partition performance data including the number of computing units, peak computing performance, and memory bandwidth;
[0062] Calculate the load rate for the partition performance data to obtain the current load status of each computing resource partition, and calculate the remaining computing power for each computing resource partition based on the current load status to obtain the remaining computing capacity value of the computing resource partition;
[0063] Analyze the task queue length according to the remaining computing capacity value of the computing resource partition to obtain the task reception threshold of each computing resource partition, and fuse the task reception threshold with the current load status to obtain a resource status matrix including load rate, remaining computing power, and task queue length.
[0064] In this example, perform resource statistical analysis on the NPU computing unit to obtain computing resource pool data including the number of computing units, computing power, and memory bandwidth. The number of computing units reflects the total number of physical computing units of the NPU, and the computing power measures the amount of operations that each computing unit can complete per unit time, in FLOPS (floating-point operations per second), while the memory bandwidth determines the data transfer speed between the NPU and the memory. Assume that the NPU has a total of computing units, and the computing power of each computing unit is (in units of FLOPS), and the memory bandwidth is (in units of GB / s). The computing resource pool data is represented as a matrix:
[0065] ;
[0066] Among them, and respectively represent the computing power and memory bandwidth of the th computing unit. Through this resource pool data, the hardware resource situation of the NPU computing unit is evaluated. According to the computing resource pool data, the partition calculation of the NPU computing unit is performed. The computing unit is divided into multiple computing resource partitions to facilitate the subsequent task scheduling and resource allocation. The partitioning strategy is dynamically adjusted according to the computing power and memory bandwidth of the computing unit. Set an initial partitioning scheme, and divide the computing units of the NPU into partitions, and each partition contains several computing units . The computing power and memory bandwidth of each partition are respectively expressed as:
[0067] ;
[0068] ;
[0069] These two formulas respectively calculate the total computing power and total memory bandwidth of partition . The initial partitioning scheme needs to meet the load balance, that is, the computing power and memory bandwidth of each partition are as close as possible. To verify the balance of the initial partitioning scheme, calculate the processing capacity distribution diagram of each partition and observe the load situation among the partitions. Assume that the processing capacity of each partition is the weighted sum of the computing power and memory bandwidth of the partition, and is expressed as:
[0070] ;
[0071] Among them, and are respectively the weight coefficients of the computing power and memory bandwidth, which are used to balance the influence of the two. By drawing the processing capacity distribution diagram, the load situation of each partition is reflected to help adjust the partitioning scheme, so that the computing resource partitions are as balanced as possible, and avoid over-saturation or over-idleness of resources in some partitions. According to the processing capacity distribution diagram of the computing resource partitions, the performance parameters of each partition are collected to obtain the peak computing performance and memory bandwidth of each computing resource partition. The peak computing performance refers to the computing power that can be achieved theoretically when the computing resource partition fully exerts its maximum performance, while the memory bandwidth is the data transmission rate that the partition can withstand. Assume that the peak computing performance of partition is , and its memory bandwidth is , and the performance data of the partition is expressed as:
[0072] ;
[0073] These performance parameters are obtained through hardware testing or simulation. By collecting these performance data, the processing capabilities of each computing resource partition can be understood more precisely. Calculate the load rate of the performance data of the partition to evaluate the current load status of each computing resource partition. The load rate refers to the computing resource partition the degree of current resource utilization, calculated as the ratio of the computing load of the current task to the maximum computing capacity of the partition. Assume the current task in the partition the computing load is , then the load rate is expressed as:
[0074] ;
[0075] wherein, represents the set of tasks being executed in the current partition , is the task corresponding computing load. The higher the load rate, the higher the utilization rate of the resources in this partition, and vice versa, indicating that the resources in this partition have not been fully utilized. Based on the current load status, calculate the remaining computing capacity of each computing resource partition, that is, the additional task computing capacity that the partition can handle under the current task load. The remaining computing capacity The calculation formula is as follows:
[0076] ;
[0077] The remaining computing capacity represents the computing task load that the partition can still accept, which can guide the system to allocate resources reasonably and migrate tasks. Partitions with higher remaining computing power can receive more computing tasks, while partitions with lower remaining computing power require task migration or resource adjustment. Based on the remaining computing capacity of the computing resource partition, analyze the task queue length to obtain the task reception threshold of each computing resource partition. The task reception threshold refers to the maximum number of tasks that each partition can accept, depending on the remaining computing capacity and the current load status of the partition. Assume the task reception threshold of the partition is , then it is expressed as:
[0078] ;
[0079] wherein, is the average computing load of each task, Denotes the floor operation. This threshold represents the number of tasks that a partition can receive within the remaining computing power. The task reception threshold is fused with the current load status to obtain a resource status matrix that includes the load rate, remaining computing power, and task queue length. The resource status matrix Is expressed as:
[0080] ;
[0081] This matrix can comprehensively reflect the current resource usage of each computing resource partition, help the system evaluate the optimization space of resource allocation, and provide a basis for subsequent task scheduling and resource management.
[0082] In one example, the task overlap degree is calculated based on the initial scheduling time sequence table and the resource status matrix, an overlapping time sequence matrix is generated, and a Q-learning task optimization model is constructed according to the overlapping time sequence matrix, including:
[0083] Perform a time interval analysis on the task time trigger window in the initial scheduling time sequence table to obtain a set of time intervals including the start and end times of each task, and calculate the time overlap of any two NPU computing tasks according to the set of time intervals to obtain a task pair overlap duration matrix;
[0084] Based on the task pair overlap duration matrix and the load rate in the resource status matrix, calculate the task overlap score to obtain an overlap degree score vector, and sort the overlap degree score vector according to the time interval to obtain an overlapping time sequence matrix including overlapping task identifiers and overlapping time sequence information;
[0085] Construct a Q-learning state space based on the overlapping time sequence matrix to obtain a state vector including task state features and resource state features;
[0086] Construct an action space for the state vector to obtain an action set including task scheduling actions and resource allocation actions;
[0087] Design a Q-value update function according to the action set to obtain a reward function including task execution delay penalty and resource utilization reward;
[0088] Construct a Q-learning model with the state vector, action set, and reward function to obtain a Q-learning task optimization model for outputting a task stacking scheme.
[0089] In this example, perform a time interval analysis on the task time trigger window in the initial scheduling time sequence table. The initial scheduling time sequence table includes the time trigger window of each NPU computing task, that is, the start time and end time of each task. For each task , its time trigger window is expressed as a pair of timestamps , where is the start time of the task and is the end time of the task. The set of time windows for all tasks is represented as:
[0090] Time Interval Set ;
[0091] where is the total number of tasks. Through time interval analysis, the time overlap duration between any two tasks and is obtained. Assume is the overlap duration between task and , then it is calculated by the following formula:
[0092] ;
[0093] Calculate whether the time windows of two tasks overlap. If they overlap, take the duration of the overlapping part; if not, the overlap duration is zero. By calculating the overlap duration for all task pairs, a task pair overlap duration matrix is obtained, which represents the overlap duration between tasks. The size of matrix is , where each element represents the overlap duration between task and task . Task overlap scoring is performed based on the task pair overlap duration matrix and the load rate in the resource status matrix. Assume the resource status matrix contains the load rate information for each computing resource partition, represented as:
[0094] ;
[0095] where is the load rate of computing resource partition , is the total number of resource partitions. The load rate is calculated by the ratio of the resources occupied by the current task to the total available resources of the resource. Combine the task pair overlap duration matrix and the load rate matrix to calculate the overlap degree score for each pair of tasks. The overlap degree score is represented by the following formula:
[0096] ;
[0097] where is the resource partition where tasks and The load rate. The higher the load rate, the higher the resource utilization rate. Therefore, the penalty value for overlap is greater, and the score is lower. In this way, an overlap degree scoring matrix for each pair of tasks is obtained. , reflecting the impact of the overlap between tasks on computing resources. Sort the overlap degree scoring vector to obtain an overlap time sequence matrix containing the overlap task identifiers and overlap time sequence information. Assume the overlap degree scoring vector is:
[0098] ;
[0099] where represents the overlap degree score between task and task . By sorting the scoring vector, a list of task pairs arranged in the order of overlap degree priority is obtained, that is, the overlap time sequence matrix. Based on the overlap time sequence matrix, construct the state space of Q-learning. The state space is the core part of the Q-learning algorithm and represents the state of the system at any given moment. In the task scheduling scenario, the state space is composed of the state characteristics of tasks and the state characteristics of resources. Task state characteristics include the current execution duration of the task, the execution priority of the task, the remaining execution duration of the task, etc.; resource state characteristics include the load rate of computing resources, the remaining computing power, the memory bandwidth, etc. Combine the state of the task and the state of the resource into a state vector , representing the current state of the system:
[0100] ;
[0101] where represents the state characteristic vector of the task, represents the state characteristic vector of the resource. Based on the state vector, construct the action space. The action space includes task scheduling actions and resource allocation actions. Task scheduling actions determine which task should be executed at the next time point, and resource allocation actions refer to how to allocate computing resources when executing tasks. Assume the action set is , where each action represents a decision on task scheduling or resource allocation. The elements in the action set are expressed as:
[0102] ;
[0103] where is the total number of actions. Based on the action set, design a Q-value update function to guide the optimization of task scheduling and resource allocation. The core idea of the Q-value update function is to update the Q-value through feedback rewards, so that the system can make better decisions in the future. The Q-value update function is defined as:
[0104] ;
[0105] Among them, is the learning rate, is the reward obtained after executing the action is the discount factor, is the maximum Q-value in the next state. The reward is designed according to the penalty for task execution delay and the reward for resource utilization. For example, there is a penalty for task execution delay, and the higher the resource utilization, the greater the reward:
[0106] ;
[0107] Among them, Delay represents the task delay after executing the action Utilization represents the resource utilization after executing the action and
[0108] are the importance coefficients for adjusting the delay and resource utilization. By training the Q-learning model, continuously adjusting the task scheduling and resource allocation strategies, and finally outputting a task optimization plan for task stacking. The Q-learning task optimization model makes optimal scheduling decisions based on the priorities of tasks, the status of resources, and the overlapping relationships between tasks, thereby improving the overall performance of the system.
[0109] In an example, according to the Q-learning task optimization model, the NPU computing tasks are sorted in a two-layer grouping manner, and a task allocation sequence is generated based on the resource status matrix, including:
[0110] Extract the periodic computing tasks according to the Q-learning task optimization model, obtain the first-layer periodic task set, and perform static partition allocation on the first-layer periodic task set based on the load rate in the resource status matrix to obtain the periodic task allocation plan;
[0111] Extract the burst task features from the state vector in the Q-learning task optimization model to obtain the second-layer burst task set, and perform dynamic partition calculation on the second-layer burst task set according to the remaining computing power in the resource status matrix to obtain the burst task candidate partitions;
[0112] Perform partition selection calculation on the burst tasks according to the task resource competition degree matrix to obtain a burst task allocation plan, and merge the periodic task allocation plan and the burst task allocation plan to obtain a complete task grouping plan;
[0113] Perform time window allocation calculation on the complete task grouping plan to obtain the execution time series of each group, and perform computing resource allocation based on the execution time series to obtain the resource occupancy plan of the tasks;
[0114] Serialize the resource occupancy plan and the execution time series to obtain a task allocation sequence containing task execution order and resource allocation information.
[0115] In this example, extract the periodic computing tasks from the Q-learning task optimization model and form the first-layer periodic task set with these tasks. Periodic tasks refer to tasks with a fixed execution period, which are executed regularly within a given time interval. These tasks have relatively stable resource requirements and execution durations. For the periodic task set, let its task set be where each task has the characteristic of periodic execution, and the period is . The resource requirements of periodic tasks are usually fixed, and their execution times can be predicted. Perform static partition allocation on these tasks based on the load rate in the resource status matrix to find a resource partition for each task, so as to make full use of resources and avoid overcrowding. Assume the resource status matrix contains the load rate of each resource partition , and select a suitable partition through the following formula:
[0116] ;
[0117] where represents the load rate of resource partition , and task will be assigned to the resource partition with the minimum load rate. After completing the static partition, obtain the allocation plan of the periodic tasks, and assign a resource partition to each periodic task. Extract the characteristics of the burst tasks based on the state vector in the Q-learning model. Burst tasks refer to tasks that occur within non-periodic time intervals, and their occurrences are random. The set of burst tasks is , and each burst task consumes a large amount of resources when executed. To handle these burst tasks, perform dynamic partition calculation on these tasks based on the remaining computing power in the resource status matrix. The remaining computing power is expressed as:
[0118] ;
[0119] wherein, is the total computing power of the resource partition , and is the total computing resources required for the tasks allocated to this resource partition. The burst tasks will be allocated to the resource partitions with relatively larger remaining computing power to ensure that they can be executed in the shortest time. According to this remaining computing power, the candidate partitions for burst tasks are expressed as:
[0120] ;
[0121] In this way, the burst tasks are dynamically allocated to the resource partitions with more remaining computing power. After the dynamic allocation of burst tasks is completed, task compatibility analysis is performed to determine whether there is resource competition between periodic tasks and burst tasks. Assume that the task resource competition degree matrix represents the resource conflicts between tasks, where each element represents the and resource competition degree between tasks
[0122] ;
[0123] When is relatively high, it indicates that the resource requirements of task and task are relatively similar and the competition is relatively fierce; on the contrary, it indicates that the resource competition is small. Through task compatibility analysis, it is checked whether periodic tasks and burst tasks can share resource partitions. If the resource competition degree is low, these two tasks can coexist in the same resource partition. According to the task resource competition degree matrix, the partition selection calculation of burst tasks is performed to determine the final allocation scheme of burst tasks. Assume that the optimal resource allocation strategy is selected according to the competition degree matrix, then the partition selection formula for burst tasks is:
[0124] ;
[0125] The periodic task allocation scheme and the burst task allocation scheme are merged to obtain a complete task grouping scheme. The time window allocation calculation is performed on the complete task grouping scheme to determine the execution time of each task. By calculating the start and end times of each task, the task execution time sequence is obtained. These tasks are arranged in the scheduling order, and each task is executed on its allocated resource partition. The execution time of the task is expressed by the following formula:
[0126]
[0127] Among them, and respectively represent the start time and end time of task , and represents the time required for task to execute. And the start time of each task must be greater than or equal to the end time of the previous task to avoid time conflicts. After obtaining the task execution time series, calculate the allocation of computing resources based on these time series, and obtain the resource occupancy plan for the tasks. The resource occupancy plan Resource Occupancy includes the computing resources occupied by each task during execution. The goal of computing resource allocation is to maximize the resource utilization rate while avoiding overcrowding of resources. The resource occupancy of tasks is calculated by the following formula:
[0128] ;
[0129] By calculating the resource occupancy of all tasks, obtain the usage of each resource partition during the entire execution process. Combining the resource occupancy plan and the task execution time series, after serialization processing, obtain the task allocation sequence. The task allocation sequence includes the execution order of tasks and computing resource allocation information, ensuring that each task can be completed on time and there are no conflicts in resources.
[0130] Among them, extract periodic computing tasks according to the Q-learning task optimization model to obtain the first-layer periodic task set, and based on the load rate in the resource status matrix, perform static partition allocation on the first-layer periodic task set to obtain the periodic task allocation plan, including:
[0131] Perform periodic feature analysis on the state vector in the Q-learning task optimization model, extract the startup interval time, execution duration, and completion time of each task during consecutive execution cycles, perform numerical statistical processing on the extracted feature parameters to obtain a periodic feature matrix containing the task execution cycle, startup time, and execution duration; perform analysis of variance on the computing tasks based on the startup interval time in the periodic feature matrix, calculate the periodicity index of the tasks according to the results of the analysis of variance and the stability of the execution duration, compare the periodicity index with a preset period threshold to obtain a set of task indices with fixed execution cycles; perform feature space mapping on the task state features in the Q-learning task optimization model according to the set of task indices, extract feature parameters including computing resource requirements, execution priorities, and data dependency relationships, group and organize them according to the task execution cycle to obtain the first-layer periodic task set; perform quantitative calculations on the number of computing units required, memory bandwidth requirements, and computing power requirements for each task in the first-layer periodic task set, normalize the resource requirement parameters in each dimension and construct a feature vector to obtain a standardized task resource requirement vector; perform a correlation analysis between the task resource requirement vector and the current load rate and remaining computing power of each resource partition in the resource state matrix, calculate the matching degree between the task resource requirements and the partition resource supply, use the cosine similarity method to construct a scoring matrix to obtain the task-resource partition fitness matrix; based on the fitness matrix, use the greedy algorithm to perform partition allocation on the first-layer periodic task set, preferentially allocate tasks with high fitness to the corresponding resource partitions, and update the remaining capacity of the resource partitions in real time, complete the allocation of all tasks in the order of task priorities to obtain an initial task allocation plan; perform statistical analysis on the number of tasks, resource occupancy rate, and computing load of each resource partition in the initial task allocation plan, calculate the load balance degree index between partitions, construct a multi-dimensional vector containing task number distribution, resource utilization distribution, and load distribution to obtain the task distribution vector of each resource partition; calculate the load imbalance degree of the current allocation plan according to the task distribution vector, identify resource partitions with too high or too low load, and dynamically adjust the load between partitions through task migration to ensure that the load balance degree of each partition meets the preset threshold requirements to obtain a periodic task allocation plan that meets the load balance constraint.
[0132] Among them, perform burst task feature extraction on the state vector in the Q-learning task optimization model to obtain the second-layer burst task set, and perform dynamic partition calculation on the second-layer burst task set according to the remaining computing power in the resource state matrix to obtain burst task candidate partitions, including:
[0133] Perform temporal feature analysis on the state vectors in the Q-learning task optimization model, calculate the execution frequency and temporal distribution characteristics of each task, perform quantization processing based on the randomness and burstiness indicators of the task arrival pattern, and obtain task temporal feature data including execution regularity degree, temporal distribution density, and burstiness score; conduct clustering analysis based on the task temporal feature data, use the K-means clustering algorithm to classify the temporal features of tasks, calculate the temporal discreteness and fluctuation amplitude of tasks, and perform screening according to the clustering results and a preset burstiness threshold to obtain a set of burst task indices; input the task identifiers in the set of burst task indices into the feature extraction module, extract the feature parameters of resource demand, execution duration, and priority for each burst task, and perform data standardization processing to obtain a second-layer set of burst tasks; conduct statistical analysis on the remaining computing power of each computing resource partition in the resource state matrix, calculate resource capacity indicators including the number of idle computing units, available memory bandwidth, and remaining computing power, and assign weights to the resource indicators to obtain a resource partition score vector; perform preliminary screening on the computing resource partitions according to the resource partition score vector, calculate the resource utilization rate and load balance degree of each partition, set resource capacity thresholds and load balance thresholds, and screen out the resource partitions that meet the requirements for processing burst tasks to obtain a set of candidate partitions; conduct resource demand analysis on the tasks in the second-layer set of burst tasks, calculate the computing density and memory access characteristics of the tasks, construct a task feature vector including computing demand intensity and memory bandwidth demand, and calculate the matching degree with the partition resource characteristics in the set of candidate partitions to obtain a task-partition adaptation matrix; conduct comprehensive evaluation based on the task-partition adaptation matrix and the resource partition score vector, use a weighted scoring method to sort each candidate partition, calculate the affinity score between the task and the partition, generate a priority mapping relationship between the task and the partition, and obtain a partition selection priority table; perform dynamic matching between the partition selection priority table and the real-time resource state, conduct partition screening and sorting according to the task urgency and resource availability, and select appropriate resource partitions in the order of priority to obtain burst task candidate partitions.
[0134] In one example, perform resource load calculation on the NPU computing tasks in the task assignment sequence, generate a migration cost matrix, and perform task reallocation according to the migration cost matrix to obtain the task reallocation result, including:
[0135] Conduct load statistical analysis on the NPU computing tasks in the task assignment sequence, obtain the task load distribution data of each computing resource partition, and calculate the load balance degree of the computing resource partitions based on the task load distribution data to obtain a set of load-unbalanced partitions;
[0136] Calculate the migration time overhead of tasks in the set of resource partitions with unbalanced load to obtain the task migration time cost vector, and calculate the task migration energy consumption based on the task migration time cost vector to obtain the task migration energy consumption cost vector;
[0137] Fuse the task migration time cost vector and the task migration energy consumption cost vector to obtain the migration cost matrix;
[0138] Select the target partition for task migration for the resource partitions with load rates lower than the preset threshold in the migration cost matrix to obtain the task migration plan;
[0139] Reallocate the NPU computing tasks based on the task migration plan to obtain the updated task allocation plan;
[0140] Reorganize the resources of the updated task allocation plan to obtain the task reallocation result including task migration information and resource allocation information.
[0141] In this example, extract the load data of each task on each computing resource partition from the task allocation sequence. The load data includes the occupancy of computing resources, the usage time of resources, etc. Assume that the load data of each computing resource partition is represented as a vector , where represents the load data of the th task on the th resource partition, and the dimension of the load data represents the number of tasks processed by this partition. Calculate the balance of the load distribution data. By calculating the load imbalance degree of each resource partition, judge whether there is a load imbalance in the resource partition. The load balance degree is measured by calculating the difference between the load of each resource partition and the load of other resource partitions. Set the load balance degree calculation formula as:
[0142] ;
[0143] Among them, is the average load of the th resource partition, and the calculation formula is:
[0144] ;
[0145] The smaller the load balance degree , the more balanced the load; on the contrary, it indicates that the load of this resource partition is unbalanced and needs to be migrated. Calculate the migration time of tasks in the set of resource partitions with unbalanced load. The task migration time refers to the time required for a task to migrate from one resource partition to another, which is related to the size of the task, the migration path, and the idle situation of the target resource. The task migration time cost vector Record the migration time of each unbalanced task, where represents the migration time of the th task. Assume that the migration time is related to the task size and the bandwidth of the computing resource. The migration time cost is expressed by the following formula:
[0146] ;
[0147] where, represents the size of task , is the bandwidth of the target resource. Calculate the energy consumption brought by task migration. The task migration energy consumption is related to the task size and the energy efficiency of the computing resource. Assume that the migration energy consumption cost vector is , where represents the migration energy consumption of task . The migration energy consumption is expressed by the following formula:
[0148] ;
[0149] where, is a constant of energy consumption, is the migration time, is the power consumption of the resource. Integrate the time cost and energy consumption cost of task migration to obtain the migration cost matrix , which is expressed as:
[0150] ;
[0151] where, and are weight coefficients used to adjust the importance of time cost and energy consumption cost. After obtaining the migration cost matrix, select the target partition for task migration according to the current load status for the resource partitions with a load rate lower than the preset threshold. Set a load rate threshold , screen the resources with a load rate lower than the threshold in the migration cost matrix, and select a suitable target partition. The task migration plan is expressed as , which is a set containing the target partitions for task migration. Based on the task migration plan, reallocate the NPU computing tasks so that the tasks can be reasonably distributed to each resource partition according to the new load status. The updated plan after task reallocation is expressed as , that is, the new task allocation plan. Reorganize the resources according to the updated task allocation plan so that each resource partition can maximize its computing power. During the reorganization process, optimize the utilization rate and performance of each computing resource based on the task migration information and resource allocation information, and finally obtain the task reallocation result , which includes the task migration information and resource allocation information of each resource partition.
[0152] In one example, the task reallocation result is input into the Pareto multi-objective optimization algorithm to generate the global execution timing of the NPU computing tasks, and the NPU computing tasks are scheduled according to the global execution timing, and the task scheduling execution information is output, including:
[0153] Perform a time overhead analysis on the task execution delay in the task reallocation result to obtain the delay minimization objective function, and calculate the utilization rate of the computing resource allocation in the task reallocation result to obtain the resource utilization maximization objective function;
[0154] Construct a multi-objective space for the delay minimization objective function and the resource utilization maximization objective function to obtain a multi-objective constraint space including task time constraints and resource capacity constraints;
[0155] Perform a non-dominated solution search calculation on the multi-objective constraint space to obtain a Pareto optimal solution set including multiple groups of task execution timings and resource configuration strategies;
[0156] Verify the dependency relationship according to the task execution timing in the Pareto optimal solution set to obtain an execution timing subset that satisfies the task dependency constraints;
[0157] Based on the execution timing subset, perform allocation verification on the computing resources to obtain a global execution timing plan that satisfies the resource constraints, and convert the global execution timing plan into an execution instruction sequence of the NPU computing tasks to obtain a task scheduling instruction set;
[0158] Perform partitioned task distribution on the task scheduling instruction set to obtain the execution instruction sequences of each computing resource partition;
[0159] Distribute the execution instruction sequences to the corresponding computing resource partitions for execution processing to obtain the task scheduling execution information including task execution status, computing performance, and resource utilization rate.
[0160] In this example, the delay factor during task execution is considered. Task delay is affected by multiple factors such as computing resource allocation, task execution order, and task size. Define the task execution delay as , which is jointly determined by the actual time of task execution and the system scheduling time. The calculation formula for task delay is as follows:
[0161] ;
[0162] where represents the execution time of task , and represents the time the task waits to be executed. The execution time Depends on the processing power of the computing resources and the computing requirements of the tasks, while the queue time is related to the position of the task in the resource queue and the scheduling priority of the task. To reduce the total latency of task execution, minimize the execution time and waiting time of the task, and obtain an objective function for latency minimization. Assume the objective function for latency minimization in the task reallocation result is:
[0163] ;
[0164] The objective function represents the sum of the execution latencies of all tasks. By optimizing this objective function, the latency of the tasks can be effectively reduced, thereby improving the overall performance of the system. Calculate the utilization rate of the computing resource allocation in the task reallocation result. It is measured by the ratio of the actual resource usage to the total resource capacity. Assume the total capacity of the th computing resource is , and its actual usage is . The utilization rate of the computing resource
[0165] is expressed as:
[0166] The objective function for maximizing the resource utilization rate of the entire system is expressed as the sum of the utilization rates of all computing resources in the system, and the formula is as follows:
[0167] ;
[0168] Among them, represents the total number of computing resources in the system, represents the actual usage of the th resource, represents the th resource, and and Based on these two objective functions, multi-objective optimization is carried out to construct a multi-objective space, including task time constraints and resource capacity constraints. The task time constraint means that the total time for task execution cannot exceed the maximum execution time preset by the system, and the resource capacity constraint means that the actual usage of each computing resource in the system cannot exceed its total capacity . These constraints are expressed by mathematical formulas as:
[0169] ;
[0170] ;
[0171] According to these constraints, construct a multi-objective constraint space that includes task time constraints and resource capacity constraints. Conduct a non-dominated solution search to obtain a Pareto optimal solution set that includes multiple sets of task execution timings and resource allocation strategies. A non-dominated solution is a solution that cannot be surpassed by other solutions in all objectives. In multi-objective optimization, the Pareto optimal solution set includes all non-dominated solutions, and these solutions achieve the best balance among different objectives. Use various methods for non-dominated solution search, such as the NSGA-II algorithm or other multi-objective evolutionary algorithms. Through these methods, multiple solutions are obtained, and each solution represents a different task scheduling timing and resource allocation strategy. According to the task execution timings in the Pareto optimal solution set, conduct dependency verification. The dependencies between tasks determine the execution order of tasks, ensuring that tasks are executed in the correct order and avoiding execution errors caused by dependency conflicts. Assume task depends on task , then the start time of task must be greater than or equal to the end time of task , that is:
[0172] ;
[0173] During the dependency verification process, check whether each set of solutions satisfies the dependencies of all tasks to obtain a subset of execution timings that satisfy the task dependency constraints. Based on the subset of execution timings, conduct allocation verification for computing resources to ensure that the resources allocated to each task do not exceed the maximum capacity of the resources throughout the execution process. Assume that task uses units of resources during execution, then for each moment , it must be satisfied that:
[0174] ;
[0175] Through verification, ensure that the resource allocation does not exceed its capacity at each moment. Obtain a global execution timing scheme that satisfies all constraints. After converting this scheme into an execution instruction sequence for NPU computing tasks, obtain a task scheduling instruction set. The scheduling instruction set includes the execution timing of each task, the required resources, and the dependencies between tasks. Partition and issue the task scheduling instruction set to obtain the execution instruction sequence for each computing resource partition. These execution instructions are distributed to the corresponding computing resource partitions for processing, and finally obtain task scheduling execution information that includes task execution status, computing performance, and resource utilization. The execution information is used to monitor and adjust the task scheduling strategy in real time to ensure the efficient operation of the system.
[0176] Refer to Figure 2, this embodiment provides an NPU computing task scheduling device for a multi-mode SoC main control chip, including:
[0177] Extraction module 1, configured to extract task feature parameters of NPU computing tasks in the multi-mode SoC main control chip, generate a task descriptor according to the task feature parameters, and establish an initial scheduling time sequence table based on the task descriptor;
[0178] Partitioning module 2, configured to perform calculation resource partitioning on the NPU computing unit to obtain multiple calculation resource partitions, acquire the performance parameters of each calculation resource partition, and establish a resource status matrix according to the performance parameters;
[0179] Calculation module 3, configured to calculate the task overlap degree based on the initial scheduling time sequence table and the resource status matrix, generate an overlapping time sequence matrix, and construct a Q-learning task optimization model according to the overlapping time sequence matrix;
[0180] Sorting module 4, configured to perform two-layer grouping sorting on the NPU computing tasks according to the Q-learning task optimization model, and generate a task allocation sequence based on the resource status matrix;
[0181] Reallocation module 5, configured to calculate the resource load of the NPU computing tasks in the task allocation sequence, generate a migration cost matrix, and perform task reallocation according to the migration cost matrix to obtain a task reallocation result;
[0182] Output module 6, configured to input the task reallocation result into the Pareto multi-objective optimization algorithm, generate the global execution time sequence of the NPU computing tasks, and perform task scheduling on the NPU computing tasks according to the global execution time sequence, and output task scheduling execution information.
[0183] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the description in the above method embodiment, and details are not described herein again.
[0184] Refer to Figure 3 , this embodiment of the present invention also provides a computer device, which may be a server, and its internal structure may be as Figure 3As shown in the figure. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0185] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0186] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0187] Those of ordinary skill in the art can understand that all or part of the processes in the above embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to memory, storage, database, or other media provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0188] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent in such a process, device, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article or method including such an element.
[0189] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A method for scheduling NPU computing tasks of a multi-mode SoC master chip, characterized in that: The following steps are involved: Extracting task characteristic parameters of the NPU computing task in the multi-mode SoC main control chip, generating a task descriptor according to the task characteristic parameters, and establishing an initial scheduling timing table based on the task descriptor; The NPU computing unit is partitioned into computing resource partitions to obtain a plurality of computing resource partitions, and performance parameters of each computing resource partition are obtained, and a resource status matrix is established according to the performance parameters, wherein the resource status matrix includes a load rate, a remaining computing power, and a task queue length; Calculating task overlap based on the initial scheduling time sequence table and the resource state matrix, generating an overlapping time sequence matrix, and constructing a Q-learning task optimization model according to the overlapping time sequence matrix, wherein the state vector in the Q-learning task optimization model includes task state characteristics and resource state characteristics; The NPU computing tasks are grouped and sorted in two layers according to the Q-learning task optimization model, and a task allocation sequence is generated based on the resource state matrix; specifically, the method comprises: extracting periodic computing tasks according to the Q-learning task optimization model to obtain a first-layer periodic task set, and statically partitioning and allocating the first-layer periodic task set based on the load rate in the resource state matrix to obtain a periodic task allocation scheme; extracting burst task features from the state vector in the Q-learning task optimization model to obtain a second-layer burst task set, and dynamically partitioning and calculating the second-layer burst task set based on the remaining computing power in the resource state matrix to obtain a burst task candidate set; Select partitions; perform task compatibility analysis based on the candidate partitions for burst tasks and the periodic task allocation scheme to obtain a task resource competition matrix; perform partition selection calculation on burst tasks according to the task resource competition matrix to obtain a burst task allocation scheme, and merge the periodic task allocation scheme and the burst task allocation scheme to obtain a complete task grouping scheme; perform time window allocation calculation on the complete task grouping scheme to obtain an execution time sequence of each group, and perform computational resource allocation based on the execution time sequence to obtain a resource occupancy scheme for the task; serialize the resource occupancy scheme and the execution time sequence to obtain a task allocation sequence containing task execution order and resource allocation information; Performing resource load calculation on the NPU computing tasks in the task allocation sequence, generating a migration cost matrix, and performing task reallocation according to the migration cost matrix to obtain a task reallocation result; The task reallocation result is input into the Pareto multi-objective optimization algorithm to generate a global execution timing of the NPU computing task, and the NPU computing task is scheduled according to the global execution timing, and task scheduling execution information is output.
2. The NPU computing task scheduling method of the multi-mode SoC master chip according to claim 1 is characterized in that: The extracting task characteristic parameters of the NPU computing task in the multi-mode SoC main control chip, generating a task descriptor according to the task characteristic parameters, and establishing an initial scheduling timing table based on the task descriptor includes: Collect features of NPU computing tasks in multi-mode SoC master chips to obtain task feature data including task execution time, computing resource requirements, task priority value, computing mode type, data dependency, and task periodicity; Performing numerical normalization processing on the task execution time and computing resource requirements in the task feature data to obtain a standardized feature vector, and performing multi-dimensional feature fusion on the standardized feature vector and the task priority value in the task feature data to obtain a task feature parameter; Inputting the task feature parameters into a feature mapping module for mapping to obtain a task descriptor including a task identifier, a computing mode identifier, a task computing resource requirement vector, a task priority value, an expected execution time, and a task dependency matrix; Classifying the NPU computing tasks according to the computing mode identifier in the task descriptor to obtain a mode classification set, and sorting the tasks in the mode classification set according to the task priority values to obtain a task execution priority sequence; Calculate the time trigger window according to the task execution priority sequence and the inter-task dependency matrix to obtain a time allocation scheme including task start time point and end time point; The time allocation scheme is scheduled and mapped with the task computing resource demand vector to obtain an initial scheduling time sequence table including the time trigger window and computing resource allocation information of each task.
3. The NPU computing task scheduling method of the multi-mode SoC master chip according to claim 2 is characterized in that: The NPU computing unit is partitioned into computing resources to obtain a plurality of computing resource partitions, and performance parameters of each computing resource partition are obtained, and a resource state matrix is established according to the performance parameters, including: Perform resource statistics analysis on NPU computing units to obtain computing resource pool data including the number of computing units, computing power, and memory bandwidth; Based on the computing resource pool data, partition calculation is performed on the NPU computing unit to obtain an initial partitioning scheme of multiple computing resource partitions, and a balance verification calculation is performed on the initial partitioning scheme of the computing resource partitions to obtain a processing capacity distribution map of the computing resource partitions; Collecting performance parameters of each computing resource partition according to the processing capacity distribution diagram of the computing resource partition to obtain partition performance data including the number of computing units, peak computing performance and memory bandwidth; Calculating the load rate of the partition performance data to obtain the current load status of each computing resource partition, and calculating the remaining computing power of each computing resource partition based on the current load status to obtain the remaining computing capacity value of the computing resource partition; The task queue length is analyzed according to the remaining computing capacity value of the computing resource partition to obtain the task receiving threshold of each computing resource partition, and the task receiving threshold is fused with the current load state to obtain a resource state matrix.
4. The NPU computing task scheduling method of the multi-mode SoC master chip according to claim 3 is characterized in that: The calculating task overlap based on the initial scheduling time sequence table and the resource status matrix, generating an overlapping time sequence matrix, and constructing a Q-learning task optimization model according to the overlapping time sequence matrix includes: Performing time interval analysis on the task time trigger window in the initial scheduling time sequence table to obtain a time interval set including the start and end time of each task, and performing time overlap calculation on any two NPU computing tasks according to the time interval set to obtain a task pair overlap duration matrix; Based on the task pair overlap duration matrix and the load rate in the resource status matrix, a task overlap score is calculated to obtain an overlap score vector, and the overlap score vector is sorted according to the time interval to obtain an overlap timing matrix including overlap task identifiers and overlap timing information; Constructing a Q-learning state space based on the overlapping time series matrix to obtain a state vector; Constructing an action space for the state vector to obtain an action set including task scheduling actions and resource allocation actions; Designing a Q-value update function according to the action set to obtain a reward function including a task execution delay penalty and a resource utilization rate reward; A Q-learning model is constructed for the state vector, the action set and the reward function to obtain a Q-learning task optimization model for outputting a task stacking solution.
5. The NPU computing task scheduling method of the multi-mode SoC master chip according to claim 1 is characterized in that: The performing resource load calculation on the NPU computing tasks in the task allocation sequence, generating a migration cost matrix, and performing task reallocation according to the migration cost matrix to obtain a task reallocation result includes: Performing load statistics analysis on the NPU computing tasks in the task allocation sequence to obtain task load distribution data of each computing resource partition, and performing load balancing calculation on the computing resource partition based on the task load distribution data to obtain a set of load imbalanced partitions; Calculating the migration time overhead of the tasks in the load imbalance partition set to obtain a task migration time cost vector, and calculating the task migration energy consumption according to the task migration time cost vector to obtain a task migration energy consumption cost vector; The task migration time cost vector and the task migration energy consumption cost vector are cost-fused to obtain a migration cost matrix; Selecting a task migration target partition for a resource partition whose load rate in the migration cost matrix is lower than a preset threshold, and obtaining a task migration plan; Reassign the NPU computing tasks based on the task migration plan to obtain an updated task allocation plan; The updated task allocation scheme is reorganized for resources to obtain a task reallocation result including task migration information and resource allocation information.
6. The NPU computing task scheduling method of the multi-mode SoC master chip according to claim 5 is characterized in that: The task reallocation result is input into the Pareto multi-objective optimization algorithm to generate a global execution timing of the NPU computing task, and the NPU computing task is scheduled according to the global execution timing, and the task scheduling execution information is output, including: Performing a time cost analysis on the task execution delay in the task reallocation result to obtain a delay minimization objective function, and performing a utilization calculation on the computing resource allocation in the task reallocation result to obtain a resource utilization maximization objective function; The delay minimization objective function and the resource utilization maximization objective function are constructed into a multi-objective space to obtain a multi-objective constraint space including task time constraints and resource capacity constraints; Performing non-dominated solution search calculation on the multi-objective constraint space to obtain a Pareto optimal solution set including multiple groups of task execution timings and resource allocation strategies; Verify the dependency relationship according to the task execution timing in the Pareto optimal solution set to obtain a subset of execution timings that meet the task dependency constraints; Based on the execution timing subset, computing resources are allocated and verified to obtain a global execution timing solution that satisfies resource constraints, and the global execution timing solution is converted into an execution instruction sequence of the NPU computing task to obtain a task scheduling instruction set; Distribute the task scheduling instruction set to partition tasks to obtain execution instruction sequences for each computing resource partition; The execution instruction sequence is distributed to the corresponding computing resource partitions for execution processing to obtain task scheduling execution information including task execution status, computing performance and resource utilization.
7. A NPU computing task scheduling device for a multi-mode SoC main control chip, characterized in that: For implementing the steps of the method according to any one of claims 1 to 6, the device comprises: An extraction module is used to extract task characteristic parameters of the NPU computing task in the multi-mode SoC main control chip, generate a task descriptor according to the task characteristic parameters, and establish an initial scheduling timing table based on the task descriptor; A partitioning module is used to partition the computing resources of the NPU computing unit to obtain multiple computing resource partitions, and obtain performance parameters of each computing resource partition, and establish a resource status matrix according to the performance parameters, wherein the resource status matrix includes load rate, remaining computing power and task queue length; A calculation module, used to calculate the task overlap based on the initial scheduling time sequence table and the resource state matrix, generate an overlapping time sequence matrix, and construct a Q-learning task optimization model according to the overlapping time sequence matrix, wherein the state vector in the Q-learning task optimization model includes task state characteristics and resource state characteristics; A sorting module is used to perform double-layer grouping and sorting of NPU computing tasks according to the Q-learning task optimization model, and generate a task allocation sequence based on the resource state matrix; specifically comprising: extracting periodic computing tasks according to the Q-learning task optimization model to obtain a first-layer periodic task set, and statically partitioning and allocating the first-layer periodic task set based on the load rate in the resource state matrix to obtain a periodic task allocation scheme; extracting burst task features from the state vector in the Q-learning task optimization model to obtain a second-layer burst task set, and dynamically partitioning and calculating the second-layer burst task set based on the remaining computing power in the resource state matrix to obtain a burst task set. Issue task candidate partitions; perform task compatibility analysis based on the candidate partitions of burst tasks and the periodic task allocation scheme to obtain a task resource competition matrix; perform partition selection calculation on the burst tasks according to the task resource competition matrix to obtain a burst task allocation scheme, and merge the periodic task allocation scheme and the burst task allocation scheme to obtain a complete task grouping scheme; perform time window allocation calculation on the complete task grouping scheme to obtain an execution time sequence of each group, and perform computing resource allocation based on the execution time sequence to obtain a resource occupancy scheme of the task; serialize the resource occupancy scheme and the execution time sequence to obtain a task allocation sequence containing task execution order and resource allocation information; A reallocation module, used to calculate the resource load of the NPU computing tasks in the task allocation sequence, generate a migration cost matrix, and perform task reallocation according to the migration cost matrix to obtain a task reallocation result; The output module is used to input the task reallocation result into the Pareto multi-objective optimization algorithm, generate the global execution timing of the NPU computing task, schedule the NPU computing task according to the global execution timing, and output the task scheduling execution information.
8. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Heterogeneous network resource allocation method based on reinforcement learning
CN112351433A
Internet of Vehicles calculation unloading method and system based on multi-objective reinforcement learning
CN113961204A