Instruction execution method and device for artificial intelligence chip
Through deep learning-driven feature recognition and resource demand prediction, combined with intelligent scheduling and dynamic micro-instruction set generation, the problem of resource allocation optimization difficulties in AI chips during dynamic and multi-task computing is solved, and the resource utilization efficiency and computing performance are significantly improved.
Patent Information
- Application Number
- CN202510249783.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Existing AI chip designs lack flexibility and adaptability when performing dynamic and multitasking computing, resulting in difficulty in resource allocation optimization and affect execution efficiency and computing performance.
Through deep learning-driven feature recognition and resource demand prediction, a task demand prediction report is generated for computing resources, intelligently schedule computing units and memory resources, dynamically select the optimal instruction execution path, and automatically generate a micro-instruction set, monitor in real time and adjust dynamically to adapt to the current hardware resource configuration.
It significantly improves the resource utilization efficiency of AI chips when performing complex tasks, ensures load balancing of computing units and memory, improves the utilization rate of hardware resources and cache hit rate, and optimizes resource allocation during parallel execution of multi-tasks.
Smart Images

Figure CN120196434A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence chips, and in particular, to an instruction execution method and device for artificial intelligence chips. Background Art
[0002] With the rapid development of artificial intelligence (AI) chips, the computing resource requirements of AI tasks are becoming increasingly complex. Currently, the design of AI chips mainly relies on the optimization of hardware parallelism and resource scheduling to handle complex computing tasks and high-concurrency data streams. However, traditional scheduling strategies are usually based on static hardware configurations, lacking flexibility and adaptability, resulting in difficulties in optimizing resource allocation during the execution of dynamic and multi-task computing, thus affecting execution efficiency and computing performance. In addition, existing resource scheduling methods often do not fully consider the computational dependencies and resource sharing between tasks, and cannot achieve highly refined micro-tuning and dynamic adjustment, further limiting the utilization efficiency of computing resources.
[0003] Existing technologies mainly rely on preset hardware resource configurations and simple scheduling strategies to allocate computing units and memory resources, lacking real-time feedback and optimization of resource requirements during the actual execution process. For example, traditional methods fail to effectively handle dependencies between tasks, resulting in uneven utilization of computing units and memory resources, affecting the performance during multi-task parallel execution. At the same time, existing micro-instruction set generation strategies also lack flexible adaptability and are difficult to dynamically adjust according to feedback during task execution. These deficiencies make it impossible for existing technologies to achieve efficient computing resource scheduling under rapidly changing and complex loads. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an instruction execution method for artificial intelligence chips to solve the problems of suboptimal instruction execution path, inaccurate resource scheduling, and insufficient dynamic micro-instruction set generation in the prior art.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides an instruction execution method for artificial intelligence chips, which includes performing deep learning-driven feature recognition and resource requirement prediction on an input task, generating a requirement prediction report of the task for computing resources by analyzing the computational graph and data dependency relationship of the task;
[0008] Intelligently scheduling computing units and memory resources, dynamically selecting the optimal instruction execution path, and simultaneously optimizing the cache policy;
[0009] Automatically generate a micro-instruction set corresponding to the task according to the computing resource requirements, computing characteristics of the task, and intelligent scheduling results;
[0010] During the execution of the optimal instructions, monitor the execution status of the task in real time, analyze the load of the computing unit, memory access situation, and cache hit rate, and dynamically adjust the micro-instruction set to adapt to the current hardware resource configuration;
[0011] When multiple tasks are executed in parallel, dynamically adjust the execution order of multiple tasks according to the computing load and resource sharing situation of the tasks to optimize resource allocation.
[0012] As a preferred solution of the instruction execution method for an artificial intelligence chip according to the present invention, wherein: the steps for generating a demand prediction report on the computing resources required by the task are as follows.
[0013] Parse the input task through static code analysis and syntax parsing methods to construct a computational graph, where nodes represent computing units and edges represent data dependencies;
[0014] Use a graph neural network to embed the task computational graph and update the features of each node through graph convolution operations;
[0015] Combine the features of a node with the features of its adjacent nodes, and transfer the information of the adjacent nodes to the current node in a weighted manner to update the node features;
[0016] Through the stacking of multiple graph convolution layers, each node gradually accumulates deep features about the task graph and constructs a deep reinforcement learning model;
[0017] Optimize the computing resource requirements of the task through the deep reinforcement learning model. At the same time, during the optimization process, monitor the computing status and resource usage of the task in real time to generate feedback data;
[0018] According to the feedback data, adjust the resource allocation in real time, dynamically optimize the scheduling of computing units, and generate a resource demand prediction report accordingly.
[0019] As a preferred solution of the instruction execution method for an artificial intelligence chip according to the present invention, wherein: the steps for intelligently scheduling computing units and memory resources, dynamically selecting the optimal instruction execution path, and optimizing the cache policy at the same time are as follows.
[0020] Extract the demand information of the task for computing resources and memory resources from the resource demand prediction report;
[0021] Based on the computational graph structure of the task, divide the task into different task units, allocate corresponding computing resources and memory resources to each task unit, and record the computing characteristics of each task unit at the same time;
[0022] Combined with the computing characteristics of the task units, predict the number of computing units and memory bandwidth required for each task unit through machine learning algorithms, and optimize the resource allocation;
[0023] According to the data dependencies and execution order among the task units, analyze the input-output relationships of each task unit to determine the dependency relationships among the task units;
[0024] Construct an instruction execution path graph based on the dependency relationships among the task units, and use a graph search algorithm to calculate the optimal execution path;
[0025] Dynamically adjust the use of the cache hierarchy based on the optimal execution path of the task and the optimized resource allocation.
[0026] As a preferred solution of the instruction execution method for the artificial intelligence chip described in the present invention, wherein: automatically generate a micro-instruction set corresponding to the task according to the computing resource requirements, computing characteristics, and intelligent scheduling results of the task, and the specific steps are as follows.
[0027] Based on the number of computing units, memory bandwidth, and the execution order and dependency relationships among the task units, use a deep reinforcement learning model to predict the micro-instruction set;
[0028] During each task execution process, the deep reinforcement learning model adjusts the generation of the micro-instruction set through multiple trial-and-error learnings;
[0029] At time t, generate a micro-instruction set z t according to the resource requirements and computing characteristics of the task, and optimize and adjust the generation strategy of the micro-instruction set using the deep reinforcement learning model according to the feedback value during the task execution process;
[0030] The deep reinforcement learning model is optimized through the following loss function, and the expression is:
[0031]
[0032] Wherein, is the optimization loss function, T represents the total number of time steps of the task execution, R t represents the actual resource requirements of the task unit at time t, represents the resource requirements predicted by the reinforcement learning model at time t, β t and α t are both weighting factors, F t represents the actual computing characteristics of the task unit at time t, represents the computing characteristics predicted by the model at time t;
[0033] The generation strategy of the micro-instruction set is continuously tried and updated through a deep reinforcement learning model to minimize the task resource requirements and the prediction error of the computing characteristics, thereby generating the micro-instruction set corresponding to the task.
[0034] As a preferred solution of the instruction execution method for the artificial intelligence chip according to the present invention, wherein: at time t, according to the resource requirements and computing characteristics of the task, a micro-instruction set z is generated t , and according to the feedback value during the task execution process, the generation strategy of the micro-instruction set is optimized and adjusted by using a deep reinforcement learning model. The specific steps are as follows.
[0035] During each task execution, the resource requirements and computing characteristics of the task are collected in real time;
[0036] Using a deep reinforcement learning model, a preliminary micro-instruction set is generated according to the resource requirements and computing characteristics collected in real time;
[0037] During the task execution process, monitor the actual resource consumption and computing characteristics, collect the feedback value during the execution process, and use the deep reinforcement learning model to optimize the generation strategy of the micro-instruction set to generate the micro-instruction set z t .
[0038] As a preferred solution of the instruction execution method for the artificial intelligence chip according to the present invention, wherein: during the optimal instruction execution process, the execution state of the task is monitored in real time, the load of the computing unit, the memory access situation and the cache hit rate are analyzed, and the micro-instruction set is dynamically adjusted to adapt to the current hardware resource configuration. The specific steps are as follows.
[0039] Collect the load of the computing unit, memory access and cache hit rate in real time through a hardware monitoring tool;
[0040] Take the load of the computing unit, memory access and cache hit rate as input resources;
[0041] Based on the input resources, train a neural network model through a deep learning method to construct a load analysis model;
[0042] Preprocess and denoise the input resources, and use the load analysis model to diagnose bottlenecks and identify performance problems;
[0043] Based on the bottleneck analysis result, adjust the micro-instruction set through a deep reinforcement learning model;
[0044] Collect the feedback data during the task execution and adjust the micro-instruction set in real time to adapt to the changes in the hardware resources.
[0045] As a preferred solution of the instruction execution method for an artificial intelligence chip according to the present invention, wherein: when multiple tasks are executed in parallel, according to the computational load and resource sharing situation of the tasks, the execution order of multiple tasks is dynamically adjusted to optimize resource allocation. The specific steps are as follows:
[0046] Generate the load priority of each task based on the load of the computing unit, memory access, and cache hit rate;
[0047] Analyze the input-output dependency relationship between tasks through the load priority of each task, and optimize the task execution order through a graph optimization algorithm;
[0048] Evaluate the resource sharing situation among parallel tasks, identify potential resource conflicts, and preferentially allocate resources to the highest-priority tasks;
[0049] Dynamically adjust the task execution order according to the load analysis results and resource conflict evaluation to optimize resource utilization.
[0050] In a second aspect, the present invention provides an instruction execution device for an artificial intelligence chip, including a feature prediction module, a resource scheduling module, an instruction generation module, a status monitoring module, and an order optimization module;
[0051] The feature prediction module is used to perform deep learning-driven feature recognition and resource requirement prediction on the input task, and generate a demand prediction report of the task for computing resources by analyzing the computational graph and data dependency relationship of the task;
[0052] The resource scheduling module is used to intelligently schedule computing unit and memory resources, dynamically select the optimal instruction execution path, and optimize the cache policy at the same time;
[0053] The instruction generation module is used to automatically generate a micro-instruction set corresponding to the task according to the computational resource requirements, computational characteristics of the task, and intelligent scheduling results;
[0054] The status monitoring module is used to monitor the execution status of the task in real time during the optimal instruction execution process, analyze the load of the computing unit, memory access situation, and cache hit rate, and dynamically adjust the micro-instruction set to adapt to the current hardware resource configuration;
[0055] The order optimization module is used to dynamically adjust the execution order of multiple tasks according to the computational load and resource sharing situation of the tasks when multiple tasks are executed in parallel, and optimize resource allocation.
[0056] In a third aspect, the present invention provides a computer device, including a memory and a processor, and the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the instruction execution method for an artificial intelligence chip as described in the first aspect of the present invention is implemented.
[0057] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the instruction execution method for an artificial intelligence chip as described in the first aspect of the present invention is implemented.
[0058] The beneficial effects of the present invention are as follows: By introducing a resource demand prediction and task scheduling method based on deep learning, the present invention significantly improves the resource utilization efficiency of the AI chip when executing complex tasks. Through the deep learning-driven analysis of the task computation graph, it is possible to accurately predict the computing resources and memory bandwidth required for each task, thereby optimizing the allocation of computing resources. At the same time, the dynamically generated micro-instruction set can be adaptively adjusted according to the feedback data monitored in real time to ensure the load balance of the computing unit and the memory, improving the utilization rate of hardware resources and the cache hit rate. When multiple tasks are executed in parallel, the present invention can intelligently adjust the task execution order according to the load priority and resource sharing situation of the tasks, optimize the resource allocation, reduce the resource conflicts between tasks, and further improve the efficiency of parallel computing. Description of the Drawings
[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0060] Figure 1 It is a flowchart of the instruction execution method for an artificial intelligence chip in Embodiment 1.
[0061] Figure 2 It is a module diagram of the instruction execution device for an artificial intelligence chip in Embodiment 1. Detailed Embodiments
[0062] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed embodiments of the present invention in conjunction with the drawings of the specification.
[0063] Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0064] Second, the "one embodiment" or "embodiment" mentioned herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.
[0065] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides an instruction execution method for an artificial intelligence chip, including the following steps:
[0066] S1. Perform deep learning-driven feature recognition and resource requirement prediction on the input task, and generate a demand prediction report of the task for computing resources by analyzing the computational graph and data dependency relationship of the task.
[0067] Furthermore, parse the input task through static code analysis and syntax parsing methods to construct a computational graph, where nodes represent computing units and edges represent data dependency relationships;
[0068] Specifically, nodes include attributes such as computing requirements and memory requirements, and edges describe data transmission and dependency relationships.
[0069] Use a graph neural network to embed the task computational graph, and update the features of each node through graph convolution operations;
[0070] Combine the features of a node with the features of its adjacent nodes, and transmit the information of the adjacent nodes to the current node in a weighted manner to update the node features, thereby gradually enhancing the understanding of each node of the global task graph;
[0071] The update formula for each node is as follows:
[0072]
[0073] Wherein, represents the feature vector of the v-th node at the k + 1-th layer, is the set of adjacent nodes of the v-th node, represents the feature vector of the adjacent node u at the k-th layer, c vu represents the weight coefficient between the node v and the adjacent node u, W (k) represents the weight matrix at the k-th layer, b (k) represents the bias term at the k-th layer, σ is an activation function (such as ReLU), k represents the index of the layer number, v represents the index of the node in the computational graph, and u represents the index of the node adjacent to the node v in the computational graph;
[0074] By stacking multiple graph convolutional layers, each node gradually accumulates deep features of the task graph and constructs a deep reinforcement learning model;
[0075] Among them, the deep features not only reflect the requirements of individual computing units but also contain information from their dependent nodes, enhancing the accuracy of computing resource prediction.
[0076] Preferably, the deep reinforcement learning model can optimize the computing resource requirements of tasks by learning the dynamic relationship between task resource requirements and task execution.
[0077] While optimizing the computing resource requirements of tasks through the deep reinforcement learning model, during the optimization process, the computing status and resource usage of tasks are monitored in real time to generate feedback data;
[0078] Among them, the state in reinforcement learning is the current computing state of the task graph (including the computing load, memory requirements, data transmission bandwidth, etc. of nodes), the action is the scheduling of computing resources (such as the CPU and GPU resources allocated to specific computing units), and the reward is measured based on the effective utilization rate of computing resources and the task completion time.
[0079] Specifically, in the deep reinforcement learning model, the calculation of resource requirements is performed through deep Q-learning (DQN). The specific Q function is defined as:
[0080]
[0081] Among them, Q(s t ,a t ) represents the Q value at time t, given state s t and action a t , that is, the expected return of the task execution result for this state and action pair, is the expectation operator, representing the calculation of the average value of future rewards, r t represents the immediate reward at time t, indicating the immediate return obtained after executing the current action a t , γ represents the discount factor, and its value range is [0,1], represents choosing the action a t+1 that maximizes the Q value at time t + 1, Q(s t+1 ,a t+1 ) represents the Q value obtained at time t + 1 after choosing action a t+1 , that is, the expected return of the task execution result for the next state s t+1 and action a t+1 ;
[0082] Adjust the resource allocation in real time according to the feedback data, dynamically optimize the scheduling of computing units, and generate a resource requirement prediction report accordingly;
[0083] Among them, the resource demand prediction report details the computing requirements, memory requirements, and recommended resource configurations of the tasks.
[0084] Specifically, the key to real-time adjustment of resource allocation through feedback data lies in using a deep reinforcement learning (DQN) model to continuously monitor the task execution status and adjust the resource scheduling strategy based on the state changes. Specifically, by analyzing key metrics such as the computing load, memory requirements, and data transmission bandwidth of each node in the task graph, it is possible to identify in real time which computing units have resource bottlenecks or uneven loads, thereby dynamically optimizing resource allocation. The feedback data is not only used to update the current computing resource configuration but also serves as the input to the reinforcement learning model to continuously optimize the prediction and decision-making process. In each round of adjustment, the deep reinforcement learning (DQN) model will continuously learn the dynamic relationship between task execution and resource requirements, enabling the resource demand prediction report to be generated in real time during task execution and having higher accuracy and practicality.
[0085] Furthermore, assume that there is a computing task that needs to perform image recognition, which contains multiple computing units (such as convolutional operations, pooling operations, fully connected layers, etc.). Each unit has different computing requirements and memory requirements during execution. During the task execution, the computing load and memory requirements of each node are monitored in real time. For example, the convolutional operation has a high computing load, while the pooling operation is relatively light, and the memory requirements are also different. Through the feedback data, it is detected that the computing load of the convolutional operation is too high, which may lead to a resource bottleneck in the GPU computing unit. Based on this monitoring data, the deep reinforcement learning model will decide to increase the GPU resources of the convolutional operation computing unit, optimize the memory bandwidth allocation, and at the same time reduce the resource allocation of the pooling operation. In this way, each node in the task graph will dynamically adjust the resource configuration according to its needs, and the finally generated resource demand prediction report will detail the resource requirements of each computing unit and its optimized configuration, resulting in a significant improvement in the overall task execution efficiency.
[0086] S2. Intelligently schedule computing units and memory resources, dynamically select the optimal instruction execution path, and at the same time optimize the cache policy.
[0087] Furthermore, extract the demand information of the task for computing resources and memory resources from the resource demand prediction report to provide basic data for subsequent task allocation and resource configuration;
[0088] Based on the computing graph structure of the task, divide the task into different task units (for example, matrix operations, data loading, conditional judgments, etc.) and allocate corresponding computing resources and memory resources to each task unit, while recording the computing characteristics (such as computing complexity, memory requirements, etc.) of each task unit;
[0089] For example, matrix operation tasks may require more computing units, while data loading tasks may have higher requirements for memory bandwidth and cache.
[0090] Combined with the computing characteristics of the task units, predict the number of computing units and memory bandwidth required for each task unit through machine learning algorithms, and optimize the resource allocation to achieve the optimal allocation of resources;
[0091] According to the data dependencies and execution order among the task units, analyze the input-output relationships of each task unit to determine the dependency relationships among the task units, which are used to provide a basis for constructing the instruction execution path graph;
[0092] Construct an instruction execution path graph through the dependency relationships among the task units, and use graph search algorithms (such as A* or Dijkstra) to calculate the optimal execution path to optimize the use of computing resources and memory, ensuring the maximization of resource utilization during task execution;
[0093] Specifically, according to the computing dependency relationships and execution order among the task units, construct a task unit dependency graph. Each task unit is represented as a node in the graph, and the data transmission or control dependency relationships among the task units are connected by directed edges.
[0094] It should be noted that if the output of task unit A is the input of task unit B, or the execution of task unit A needs to be carried out after task unit B is completed, then a directed edge from A to B is added in the graph. In this way, the dependency relationships and execution order of the task units can be clearly represented by this directed acyclic graph (DAG). In addition, for each node (task unit), its computing characteristics need to be recorded, such as computing complexity, memory bandwidth requirements, etc. These characteristics help to optimize resource allocation subsequently. The weight of each edge can represent the data transmission delay or the time overhead of execution dependency among the task units, thereby further distinguishing the different execution costs among the task units in the graph. Finally, the task unit dependency graph accurately reflects the execution order, resource requirements, and time constraints of the tasks, laying a foundation for subsequent path optimization.
[0095] Furthermore, after the task unit dependency graph is established, the Dijkstra algorithm is used to calculate the optimal execution path. First, initialize the distance values of each task unit node. The distance of the starting node is set to 0, and the distances of other nodes are set to infinity. Then, start traversing from the starting task unit, and calculate the path cost from the current node to its adjacent nodes one by one. The calculation of the path cost includes not only the computational complexity and resource requirements of the nodes, but also the time required to transfer data through the edges. During the traversal, if it is found that the path cost from the current node to the adjacent node is less than the originally recorded cost, then update the distance value of that adjacent node. In this way, the Dijkstra algorithm continuously searches for and updates the shortest path until the optimal execution path from the starting task unit to all other task units is found. Finally, the calculated optimal execution path is the path that minimizes the computational resource consumption and execution time. This optimal execution path can not only maximize the utilization rate of resources, but also avoid unnecessary delays between task units.
[0096] It should also be noted that when calculating the optimal execution path, it is necessary to always pay attention to the dependency relationships of the task units, ensure that the path selection meets the requirements of data dependencies, and make dynamic adjustments when necessary. In addition, after calculating the optimal execution path, the execution order between task units should be further verified to ensure that all dependency relationships are satisfied and the resource allocation is reasonable.
[0097] Among them, the instruction execution path graph is used to ensure that each optimal execution path can maximize the utilization of computing units and memory resources, and avoid overload and resource conflicts.
[0098] It should also be noted that if during the task execution process, the computing resources (such as computing units or memory) change or the resource occupancy is too high, monitor in real time and dynamically adjust the instruction execution path. For example, switch compute-intensive operations to computing units with lower resource loads, or reorganize the memory access strategy to reduce cache misses.
[0099] Dynamically adjust the use of cache levels according to the optimal execution path of the task and the optimized resource configuration to optimize cache allocation, reduce memory access latency, and improve execution efficiency;
[0100] For example, cache frequently accessed small data in the L1 cache, and cache larger data sets in the L2 or L3 cache. When multiple tasks are in parallel, dynamically adjust the cache level allocation strategy in combination with the resource requirements of the tasks to reduce cache contention and access conflicts.
[0101] S3. Automatically generate the micro-instruction set corresponding to the task according to the computing resource requirements, computing characteristics of the task, and intelligent scheduling results.
[0102] Further, based on the number of computing units, memory bandwidth, and the execution order and dependency relationships among task units, a deep reinforcement learning model is used to predict the micro-instruction set;
[0103] During each task execution, the deep reinforcement learning model (DRL) adjusts the generation of the micro-instruction set through multiple trial-and-error learnings, making the micro-instruction set highly adaptable during execution and capable of optimizing resource utilization and computing characteristics;
[0104] Preferably, using DRL can continuously adjust and optimize the solution under complex task dependencies and resource requirements, making it more in line with the actual task needs. Compared with traditional static scheduling methods, DRL has higher flexibility and adaptability.
[0105] At time t, according to the resource requirements and computing characteristics of the task, a micro-instruction set z is generated t , and according to the feedback value during task execution, the deep reinforcement learning model is used to optimize and adjust the generation strategy of the micro-instruction set;
[0106] Specifically, during each task execution, the resource requirements (such as computing units and memory bandwidth) and computing characteristics (such as computing density and I / O requirements) of the task are collected in real time. Using the deep reinforcement learning model, a preliminary micro-instruction set is generated according to the real-time collected resource requirements and computing characteristics. During task execution, the actual resource consumption and computing characteristics are monitored, the feedback value during the execution process is collected, and the deep reinforcement learning model is used to optimize the generation strategy of the micro-instruction set to generate the micro-instruction set z t .
[0107] The deep reinforcement learning model is optimized through the following loss function, and the expression is:
[0108]
[0109] Among them, is the optimization loss function, representing the error in micro-instruction set generation. t represents the total number of time steps of task execution, and R t represents the actual resource requirements of the task unit at time T, represents the resource requirements predicted by the reinforcement learning model at time T, and β t and α t are both weighting factors, respectively representing weighting the errors at different time steps. F t represents the actual computing characteristics of the task unit at time T, represents the computing characteristics predicted by the model at time t;
[0110] Preferably, by calculating the prediction error F of the actual resource requirements R t and the actual computing characteristics tto optimize the micro-instruction set. Using the weighting factors β t and α t can be flexibly adjusted for different time steps, which provides good support for dynamic optimization and adaptation to different tasks. Specifically, the optimized loss function can be optimized from two aspects: one is by minimizing the resource requirement error (the difference between R t and ), and the other is by minimizing the computational characteristic error (the difference between F t and ). This design is very consistent with the basic idea of the deep reinforcement learning optimization strategy.
[0111] Through the deep reinforcement learning model, continuously try and error and update the generation strategy of the micro-instruction set to minimize the task resource requirements and computational characteristic prediction errors, thereby generating the micro-instruction set corresponding to the task;
[0112] Preferably, through continuous trial and error and adjustment of the micro-instruction set generation strategy, the reinforcement learning model can be continuously optimized, thereby effectively reducing resource waste and conflicts during task execution. Especially when dealing with complex and time-sequentially dependent tasks, this method can significantly improve performance compared with traditional methods.
[0113] Furthermore, each time a task is executed, the micro-instruction set is dynamically adjusted according to the real-time resource requirements and computational characteristics, emphasizing flexibility and adaptability. This enables the micro-instruction set not only to achieve the optimal in terms of resource requirements, but also to avoid excessive computation or resource conflicts during task execution, thereby effectively optimizing the overall performance.
[0114] S4. During the execution of the optimal instructions, monitor the execution status of the task in real time, analyze the load of the computing unit, memory access situation and cache hit rate, and dynamically adjust the micro-instruction set to adapt to the current hardware resource configuration.
[0115] Furthermore, collect the load of the computing unit, memory access and cache hit rate in real time through hardware monitoring tools;
[0116] Among them, the load refers to the amount of tasks carried by the computing unit, the memory access represents the data exchange frequency between the computing unit and the memory, and the cache hit rate reflects the efficiency of data access.
[0117] Take the load of the computing unit, memory access and cache hit rate as input resources, and these input resources directly affect the subsequent load analysis and instruction scheduling;
[0118] Based on the input resources, train a neural network model through deep learning methods to construct a load analysis model;
[0119] Preferably, the deep learning model is applied to automatically learn the complex correlation relationships between computing resources. Especially in the diagnosis of hardware performance bottlenecks and task performance issues, it can greatly improve the accuracy.
[0120] Preprocess and denoise the input resources, and use the load analysis model to diagnose bottlenecks and identify performance issues;
[0121] Specifically, when performing preprocessing and denoising, it is first necessary to collect real-time data such as the load of the computing unit, memory access situation, and cache hit rate, and perform preliminary cleaning and normalization processing. Since the real-time monitoring data may contain noise or incomplete data, signal processing or statistical methods need to be used for denoising. For example, use the mean filtering method based on a sliding window to smooth data fluctuations and remove the noise caused by short-term fluctuations, or use a more complex Denoising Autoencoder model to extract high-quality feature information. In addition, time series analysis methods such as moving average method or wavelet transform can also be applied to detect abnormal fluctuations and eliminate irrelevant interferences. These methods help to improve the accuracy of subsequent analysis and provide more accurate data input for the load analysis model.
[0122] Furthermore, the data after denoising processing will be input into the load analysis model, usually using neural networks (such as Convolutional Neural Network CNN or Recurrent Neural Network RNN) for analysis to capture the complex dependency relationships between the computing unit, memory, and cache. The model will identify situations of excessive load or underutilized resources by learning the patterns in the data, thereby diagnosing possible performance bottlenecks. For example, if the cache hit rate is too low, it may indicate problems with the data access pattern, resulting in a memory bandwidth bottleneck; if the computing unit load is too high, it may be due to unreasonable computing resource configuration. Based on these diagnostic results, the load analysis model can further identify the specific sources of performance issues, such as uneven scheduling of computing units, memory access conflicts, or cache misses, etc., and provide optimization suggestions according to the analysis results. The goal of this process is to improve the utilization efficiency of the overall computing resources through accurate diagnosis, reduce the occurrence of performance bottlenecks, and ultimately optimize the task execution path by real-time adjusting the micro-instruction set.
[0123] Based on the bottleneck analysis results, adjust the micro-instruction set through the deep reinforcement learning model to improve the cache hit rate and reduce the performance loss caused by cache misses;
[0124] Specifically, in the bottleneck analysis stage, real-time obtain the resource consumption data of task execution through hardware monitoring tools, including computing unit load, memory access latency, cache hit rate, etc. These resource consumption data will be used to diagnose the bottlenecks in task execution.
[0125] For example, if it is found that a certain computing unit has a high load or insufficient memory bandwidth, these problems can be identified and the type and location of the bottleneck can be located. Then, these analysis results are used as the basis for adjusting the micro-instruction set.
[0126] Furthermore, during the process of adjusting the micro-instruction set, the deep reinforcement learning model starts to optimize the micro-instruction set according to the bottleneck analysis results and the real-time execution feedback of the task. Specifically, the deep reinforcement learning model uses a reward mechanism to guide the adjustment of the micro-instruction set: if the adjusted micro-instruction set can reduce the bottleneck and improve the resource utilization efficiency, positive feedback is given; otherwise, negative feedback is given. The deep reinforcement learning model continuously optimizes the scheduling strategy of the micro-instruction set according to these feedbacks.
[0127] For example, if the computing unit has a high load, the deep reinforcement learning model will adjust the execution order of the tasks, giving priority to allocating compute-intensive tasks to the units with lighter loads, or adjusting the instruction scheduling method to reduce the pressure on the computing units. If the memory bandwidth becomes a bottleneck, the deep reinforcement learning model may optimize the data transfer strategy to reduce frequent data access, or optimize the cache strategy to make memory access more efficient.
[0128] After each execution, the deep reinforcement learning model will adjust the micro-instruction set again according to the actual running results of the task (such as resource consumption, execution time, system load, etc.). Through continuous trial and error and optimization, the deep reinforcement learning model can gradually adjust the micro-instruction set that best suits the current hardware resources and task characteristics. Eventually, this dynamic adjustment enables the task to be executed efficiently under different hardware environments and loads, minimizing the impact of resource bottlenecks on task performance.
[0129] Collect feedback data during task execution and adjust the micro-instruction set in real time to adapt to changes in hardware resources;
[0130] Preferably, by adjusting the micro-instruction set in real time to adapt to changes in hardware resources, it is ensured that flexible adjustment can be made according to changes in hardware resources. This process avoids waste of hardware resources and low execution efficiency. Especially when the hardware configuration changes, the dynamic adjustment of the micro-instruction set ensures continuous efficient operation.
[0131] S5. When multiple tasks are executed in parallel, dynamically adjust the execution order of multiple tasks according to the computing load and resource sharing situation of the tasks to optimize resource allocation.
[0132] Furthermore, based on the load of the computing unit, memory access, and cache hit rate, generate the load priority of each task;
[0133] Through the load priority of each task, analyze the input-output dependency relationship between tasks, and optimize the task execution order through graph optimization algorithms;
[0134] Preferably, the graph optimization algorithm can ensure that the execution order of tasks best conforms to the availability of resources through multiple rounds of iterative optimization, while minimizing the waiting time between tasks and the idle time of computing resources. For example, through a heuristic search algorithm or a graph search algorithm, calculate the execution time of each task node, and dynamically adjust the order of task execution, so that tasks with higher priorities can start earlier, thereby improving the efficiency of the entire computing process.
[0135] Evaluate the resource sharing situation among parallel tasks, identify potential resource conflicts, and preferentially allocate resources to the highest-priority tasks;
[0136] It should be noted that when parallel tasks are executed, they often need to share certain computing resources or memory, which may cause resource conflicts and reduce the task execution efficiency. Existing technologies usually adopt static allocation strategies, which are prone to uneven resource allocation and performance bottlenecks. This method effectively avoids the "blindness" of resource allocation by real-time evaluating the resource sharing situation among tasks and identifying conflicts, provides more accurate resource scheduling for tasks, and improves resource utilization efficiency.
[0137] Specifically, when evaluating the resource sharing situation among parallel tasks, first analyze the requirements of each task for computing units, memory bandwidth, and caches, and identify possible resource sharing points among tasks, especially conflicts in memory access and cache usage.
[0138] For example, when multiple tasks need to access the same memory area or share caches simultaneously, resource competition will occur, thus affecting the task execution efficiency. Next, based on the input-output dependency relationships among tasks, further determine the conflict situations of these shared resources. For resource conflicts, calculate the resource requirements of tasks and predict the possibility of conflicts. During this process, the load priority of tasks is used as the core basis to ensure that the highest-priority tasks can obtain the required computing resources and memory bandwidth first, thereby avoiding delays caused by insufficient resources. And low-priority tasks will be appropriately delayed or have their resource allocation reduced when resource conflicts are relatively severe. This processing method effectively ensures the timely execution of high-priority tasks, while minimizing resource competition and improving task execution efficiency.
[0139] Dynamically adjust the task execution order according to the load analysis results and resource conflict evaluation to optimize resource utilization.
[0140] This embodiment also provides an instruction execution device for an artificial intelligence chip, including: a feature prediction module, a resource scheduling module, an instruction generation module, a status monitoring module, and a sequence optimization module; The feature prediction module is used to perform feature recognition and resource requirement prediction driven by deep learning on the input task, and generate a demand prediction report of the task for computing resources by analyzing the computational graph and data dependency relationship of the task; The resource scheduling module is used to intelligently schedule computing units and memory resources, dynamically select the optimal instruction execution path, and optimize the cache policy at the same time; The instruction generation module is used to automatically generate a micro-instruction set corresponding to the task according to the computing resource requirements, computing characteristics of the task, and intelligent scheduling results; The status monitoring module is used to monitor the execution status of the task in real time during the execution of the optimal instruction, analyze the load of the computing unit, memory access situation, and cache hit rate, and dynamically adjust the micro-instruction set to adapt to the current hardware resource configuration; The sequence optimization module is used to dynamically adjust the execution sequence of multiple tasks according to the computing load and resource sharing situation of the tasks during multi-task parallel execution, and optimize resource allocation.
[0141] This embodiment also provides a computer device, applicable to the situation of the instruction execution method for an artificial intelligence chip, including: a memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the instruction execution method for an artificial intelligence chip proposed in the above embodiment.
[0142] The computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball, or touchpad set on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0143] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the instruction execution method for an AI chip as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0144] In summary, by introducing a resource demand prediction and task scheduling method based on deep learning, the present invention significantly improves the resource utilization efficiency of an AI chip when executing complex tasks. Through the deep learning-driven analysis of the task computation graph, the computing resources and memory bandwidth required for each task can be accurately predicted, thereby optimizing the allocation of computing resources. At the same time, the dynamically generated micro-instruction set can be adaptively adjusted according to the feedback data monitored in real time to ensure the load balance of the computing unit and the memory, improving the utilization rate of hardware resources and the cache hit rate. When multiple tasks are executed in parallel, the present invention can intelligently adjust the task execution order according to the load priority and resource sharing situation of the tasks, optimize the resource allocation, reduce the resource conflicts between tasks, and further improve the efficiency of parallel computing.
[0145] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for executing instructions for an artificial intelligence chip, characterized in that: include, Perform deep learning-driven feature recognition and resource demand prediction on input tasks, and generate a task demand prediction report on computing resources by analyzing the task's computational graph and data dependencies; Intelligently schedule computing units and memory resources, dynamically select the best instruction execution path, and optimize cache strategies; Automatically generate the microinstruction set corresponding to the task based on the task's computing resource requirements, computing characteristics, and intelligent scheduling results; During the optimal instruction execution process, the execution status of the task is monitored in real time, the load of the computing unit, memory access and cache hit rate are analyzed, and the microinstruction set is dynamically adjusted to adapt to the current hardware resource configuration; When multiple tasks are executed in parallel, the execution order of multiple tasks is dynamically adjusted according to the computing load and resource sharing of the tasks to optimize resource allocation.
2. The instruction execution method for an artificial intelligence chip according to claim 1, characterized in that: The specific steps of generating a task's demand forecast report on computing resources are as follows: Parse the input task through static code analysis and syntax parsing methods to build a computational graph, where nodes represent computational units and edges represent data dependencies; Use graph neural networks to embed the task computation graph and update the features of each node through graph convolution operations; Combine the features of a node with those of its neighboring nodes, pass the information of the neighboring nodes to the current node in a weighted manner, and update the node features; By stacking multiple graph convolutional layers, each node gradually accumulates deep features about the task graph and builds a deep reinforcement learning model; Optimize the computing resource requirements of tasks through deep reinforcement learning models. During the optimization process, monitor the computing status and resource usage of tasks in real time to generate feedback data. Adjust resource allocation in real time based on feedback data, dynamically optimize the scheduling of computing units, and generate resource demand forecast reports based on this.
3. The instruction execution method for an artificial intelligence chip according to claim 2, characterized in that: The intelligent scheduling of computing units and memory resources dynamically selects the optimal instruction execution path and optimizes the cache strategy. The specific steps are as follows: Extract the task's demand information for computing resources and memory resources from the resource demand forecast report; Based on the task computation graph structure, the task is divided into different task units, and corresponding computing resources and memory resources are allocated to each task unit, while the computing characteristics of each task unit are recorded; Combined with the computing characteristics of the task unit, the machine learning algorithm is used to predict the number of computing units and memory bandwidth required for each task unit, and optimize resource allocation; According to the data dependency and execution order between task units, the input-output relationship of each task unit is analyzed to determine the dependency between task units; The instruction execution path graph is constructed through the dependency relationship between task units, and the optimal execution path is calculated using a graph search algorithm; Dynamically adjust the use of cache levels based on the optimal execution path of the task and the optimized resource configuration.
4. The instruction execution method for an artificial intelligence chip according to claim 3, characterized in that: The microinstruction set corresponding to the task is automatically generated according to the computing resource requirements, computing characteristics and intelligent scheduling results of the task. The specific steps are as follows: Based on the number of computing units, memory bandwidth, and the execution order and dependencies between task units, a deep reinforcement learning model is used to predict the microinstruction set. During each task execution, the deep reinforcement learning model adjusts the generation of the microinstruction set through multiple trial-and-error learning; At time t, according to the resource requirements and computing characteristics of the task, a microinstruction set z is generated t , and based on the feedback value during task execution, the deep reinforcement learning model is used to optimize and adjust the microinstruction set generation strategy; The deep reinforcement learning model is optimized by the following loss function, expressed as: in, To optimize the loss function, T represents the total number of time steps for task execution, R t represents the actual resource demand of the task unit at time t, represents the resource demand predicted by the reinforcement learning model at time t, β t and α t are weighting factors, F t represents the actual computing characteristics of the task unit at time t, represents the computational characteristics predicted by the model at time t; Through the deep reinforcement learning model, the microinstruction set generation strategy is continuously tried and updated to minimize the task resource requirements and computing characteristics prediction errors, thereby generating the microinstruction set corresponding to the task.
5. The instruction execution method for an artificial intelligence chip according to claim 4, characterized in that: At time t, according to the resource requirements and computing characteristics of the task, a microinstruction set z is generated. t , and according to the feedback value during the task execution process, the deep reinforcement learning model is used to optimize and adjust the generation strategy of the microinstruction set. The specific steps are as follows: When each task is executed, the resource requirements and computing characteristics of the task are collected in real time; Using deep reinforcement learning models, a preliminary microinstruction set is generated based on the resource requirements and computing characteristics collected in real time; During the task execution, the actual resource consumption and computing characteristics are monitored, feedback values are collected during the execution process, and the generation strategy of the microinstruction set is optimized using the deep reinforcement learning model to generate the microinstruction set z t .
6. The instruction execution method for an artificial intelligence chip according to claim 5, characterized in that: In the optimal instruction execution process, the execution status of the task is monitored in real time, the load of the computing unit, the memory access situation and the cache hit rate are analyzed, and the microinstruction set is dynamically adjusted to adapt to the current hardware resource configuration. The specific steps are as follows: Use hardware monitoring tools to collect the load, memory access, and cache hit rate of computing units in real time; The load, memory access and cache hit rate of the computing unit are taken as input resources; Based on the input resources, the neural network model is trained through deep learning methods to build a load analysis model; Preprocess and denoise input resources, use load analysis models to diagnose bottlenecks and identify performance issues; Based on the bottleneck analysis results, the microinstruction set is adjusted through a deep reinforcement learning model; Collect feedback data during task execution and adjust the microinstruction set in real time to adapt to changes in hardware resources.
7. The instruction execution method for an artificial intelligence chip according to claim 6, characterized in that: When multiple tasks are executed in parallel, the execution order of multiple tasks is dynamically adjusted according to the computing load and resource sharing of the tasks to optimize resource allocation. The specific steps are as follows: Generate the load priority of each task based on the load, memory access and cache hit rate of the computing unit; By analyzing the input-output dependencies between tasks based on the load priority of each task, the task execution order is optimized using graph optimization algorithms. Evaluate resource sharing between parallel tasks, identify potential resource conflicts, and prioritize resource allocation for the highest priority tasks; Based on the load analysis results and resource conflict assessment, the task execution order is dynamically adjusted to optimize resource utilization.
8. An instruction execution device for an artificial intelligence chip, based on the instruction execution method for an artificial intelligence chip according to any one of claims 1 to 7, characterized in that: Including, feature prediction module, resource scheduling module, instruction generation module, state monitoring module and sequence optimization module; The feature prediction module is used to perform deep learning-driven feature recognition and resource demand prediction on the input task, and generate a task demand prediction report on computing resources by analyzing the task's computational graph and data dependency; The resource scheduling module is used to intelligently schedule computing units and memory resources, dynamically select the optimal instruction execution path, and optimize the cache strategy; The instruction generation module is used to automatically generate a microinstruction set corresponding to a task according to the computing resource requirements, computing characteristics and intelligent scheduling results of the task; The state monitoring module is used to monitor the execution state of the task in real time during the optimal instruction execution process, analyze the load, memory access and cache hit rate of the computing unit, and dynamically adjust the microinstruction set to adapt to the current hardware resource configuration; The sequence optimization module is used to dynamically adjust the execution order of multiple tasks and optimize resource allocation according to the computing load and resource sharing of the tasks when multiple tasks are executed in parallel.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the instruction execution method for an artificial intelligence chip described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the instruction execution method for an artificial intelligence chip described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Task scheduling optimization method and system based on equipment state analysis
CN118193169A
Method and system for improving computing power efficiency
CN118550711A
Dynamic workflow optimization system for improved task management efficiency
DE202024101468U1
Cited By
Data distribution method and system
CN120631597A
A data distribution method and system
CN120631597B
Chip, intelligent agent system based on LLM and intelligent agent equipment
CN120973417A
A chip, an LLM-based agent system, and an agent device
CN120973417B
Artificial intelligence acceleration method and system based on heterogeneous hardware
CN121328639A