Instruction optimization scheduling method based on large model and related device
By converting large-scale pre-trained models into calculation task diagrams, determining the scheduling priority of instructions, and optimizing the scheduling strategy based on register pressure state, the resource utilization and computing performance problems of VLIW processors in large-model computing are solved, achieving more efficient resource utilization and model performance improvement.
Patent Information
- Application Number
- CN202510045731.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The existing VLIW processor instruction scheduling strategies are difficult to meet the needs of large-scale pre-trained models for parallel computing and high-speed data processing, resulting in limited resource utilization and computing performance.
By converting the target model into a computing task diagram, the scheduling priority of each instruction is determined, and the instruction scheduling operations are performed based on the register pressure state monitored in real time, dynamic priority scheduling and resource optimization are achieved.
It improves resource utilization efficiency, meets the computing needs of large models, and improves the overall model performance.
Smart Images

Figure CN119473560B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a large model-based instruction optimization scheduling method and related devices. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, deep learning models have performed well in many fields, such as natural language processing, computer vision, and speech recognition. These models, especially large-scale pre-trained models (such as GPT-4, BERT, etc.), contain hundreds of millions to tens of billions of parameters, which poses unprecedented challenges to computing power and hardware resources. Traditional processor architectures are difficult to meet the needs of these models for parallel computing and high-speed data processing.
[0003] Very long instruction word (VLIW) processors improve processing efficiency by packaging multiple operations into one instruction word and executing them in parallel. However, the existing VLIW processor instruction scheduling strategy is usually designed based on computing tasks, which results in the inability to adjust the resource allocation status in time when executing data processing tasks of neural network models (especially large language models) based on the current instruction scheduling strategy, limiting the resource utilization and computing performance of the VLIW processor.
[0004] In addition, there are complex dependencies in neural network models (especially large language models), as well as optimization processing of various computing bottlenecks, and existing instruction scheduling strategies are difficult to meet the data processing requirements of the above scenarios. For example, matrix multiplication and convolution operations in deep learning models are highly parallel, which makes deep learning models usually accompanied by complex dependencies and data reuse requirements, and existing instruction scheduling strategies cannot adjust resource allocation in time to adapt to changes in computing requirements, making it difficult to meet the data processing requirements in deep learning models.
[0005] Therefore, it is urgent to design a technical solution to solve at least one of the above technical problems. Summary of the invention
[0006] In response to the technical problems existing in the prior art, the present application provides an instruction optimization scheduling method based on a large model and related devices, so as to realize instruction scheduling optimization based on a large model, improve resource utilization efficiency, meet the computing requirements of the large model, and improve the overall model performance.
[0007] In a first aspect, an embodiment of the present application provides an instruction optimization scheduling method based on a large model, the method comprising:
[0008] For a target model to be processed, the target model is converted into a computing task graph; wherein the computing task graph includes a plurality of nodes and connecting edges between the plurality of nodes; the plurality of nodes and connecting edges between the plurality of nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model;
[0009] Determining the scheduling priority of each instruction in the target model based on the computing task graph; the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction;
[0010] The instruction scheduling operation for the target model is performed according to the register pressure state monitored in real time and the scheduling priority.
[0011] In a second aspect, the present application provides an instruction optimization scheduling device based on a large model, wherein
[0012] The conversion unit is configured to convert the target model to be processed into a computing task graph; wherein the computing task graph includes a plurality of nodes and connecting edges between the plurality of nodes; the plurality of nodes and connecting edges between the plurality of nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model;
[0013] A determination unit is configured to determine the scheduling priority of each instruction in the target model based on the computing task graph; the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction;
[0014] The execution unit is configured to execute instruction scheduling operations on the target model according to the register pressure state monitored in real time and the scheduling priority.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising:
[0016] at least one processor, memory, and input-output unit;
[0017] The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the large model-based instruction optimization scheduling method of the first aspect.
[0018] In a fourth aspect, a computer-readable storage medium is provided, which includes instructions, and when the instructions are executed on a computer, the computer executes the large model-based instruction optimization scheduling method of the first aspect.
[0019] The beneficial effect of the present application is: providing an instruction optimization scheduling method based on a large model and a related device. In this technical solution, first, for the target model to be processed, the target model is converted into a computing task graph. Among them, the computing task graph contains multiple nodes and connecting edges between multiple nodes; the multiple nodes and connecting edges between multiple nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model. Then, the scheduling priority of each instruction in the target model is determined based on the computing task graph. Among them, the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction. Finally, according to the real-time monitored register pressure state and the scheduling priority, the instruction scheduling operation of the target model is performed.
[0020] The technical solution of the present application realizes dynamic priority scheduling of model instructions through computing task graphs, and further optimizes the scheduling strategy of model instructions based on real-time monitoring of register pressure status, realizes instruction scheduling optimization based on large models, improves resource utilization efficiency, meets the computing requirements of large models, and improves overall model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of an instruction optimization scheduling method based on a large model according to an embodiment of the present application;
[0022] Figure 2 It is a schematic diagram of the principle of an instruction optimization scheduling method based on a large model according to an embodiment of the present application;
[0023] Figure 3 It is a structural schematic diagram of an instruction optimization scheduling method device based on a large model according to an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application;
[0025] Figure 5 It is a structural schematic diagram of a medium device according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0027] With the rapid development of artificial intelligence (AI) technology, deep learning models have performed well in many fields, such as natural language processing, computer vision, and speech recognition. These models, especially large-scale pre-trained models (such as GPT-4, BERT, etc.), contain hundreds of millions to tens of billions of parameters, which poses unprecedented challenges to computing power and hardware resources. Traditional processor architectures are difficult to meet the needs of these models for parallel computing and high-speed data processing.
[0028] In the related art, very long instruction word (VLIW) processors improve processing efficiency by packaging multiple operations into one instruction word and executing them in parallel. However, the existing VLIW processor instruction scheduling strategy is usually designed based on computing tasks, which will result in the inability to adjust the resource allocation status in time when executing data processing tasks of neural network models (especially large language models) based on the current instruction scheduling strategy, which limits the resource utilization and computing performance of the VLIW processor.
[0029] In addition, in the related technologies, there are complex dependencies in neural network models (especially large language models), as well as optimization processing of various computing bottlenecks, and the existing instruction scheduling strategies are difficult to meet the data processing requirements of the above scenarios. For example, the matrix multiplication and convolution operations in deep learning models are highly parallel, which makes deep learning models usually accompanied by complex dependencies and data reuse requirements, and the existing instruction scheduling strategies cannot adjust resource allocation in time to adapt to changes in computing requirements, and it is difficult to meet the data processing requirements in deep learning models.
[0030] In summary, in the related technologies, the existing instruction scheduling strategies usually fail to fully consider the complex dependencies of deep learning models, resulting in the failure to optimize the execution order between instructions, affecting the parallel execution capability of the processor. Secondly, the existing instruction scheduling strategies are mostly static scheduling, which cannot be dynamically adjusted according to the actual computing load and resource utilization during runtime, resulting in low efficiency when processing large models. Thirdly, when allocating computing resources, the existing instruction scheduling strategies fail to fully consider the resource requirements of different instructions, resulting in excessive use or idleness of some computing resources, affecting overall performance. Finally, most of the existing instruction scheduling strategies are general algorithms, which fail to optimize the special computing modes of large deep learning models, and do not take into account the unique computing resources of the processor, and cannot give full play to the performance advantages of VLIW processors.
[0031] Therefore, it is urgent to design a technical solution to solve at least one of the above technical problems.
[0032] In order to solve at least one technical problem in the related art, an embodiment of the present application provides an instruction optimization scheduling method based on a large model and a related device.
[0033] In the technical solution provided by the present application, first, for the target model to be processed, the target model is converted into a computing task graph. The computing task graph includes multiple nodes and connecting edges between the multiple nodes; the multiple nodes and connecting edges between the multiple nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model. Then, the scheduling priority of each instruction in the target model is determined based on the computing task graph. The scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction. Finally, the instruction scheduling operation of the target model is performed according to the real-time monitored register pressure state and the scheduling priority.
[0034] The technical solution of the present application realizes dynamic priority scheduling of model instructions through computing task graphs, and further optimizes the scheduling strategy of model instructions based on real-time monitoring of register pressure status, realizes instruction scheduling optimization based on large models, improves resource utilization efficiency, meets the computing requirements of large models, and improves overall model performance.
[0035] The instruction optimization scheduling method scheme based on the large model provided in the embodiment of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a special device (such as a special terminal device with an instruction optimization scheduling method system based on a large model, etc.). These electronic devices can also be equipped with the chips introduced in the above embodiments. Alternatively, these electronic devices can also be installed with a service program for executing the instruction optimization scheduling method scheme based on the large model.
[0036] Figure 1 A flowchart of an instruction optimization scheduling method based on a large model provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method comprises the following steps:
[0037] 101, for a target model to be processed, converting the target model into a computing task graph;
[0038] 102, determining a scheduling priority of each instruction in the target model based on the computing task graph;
[0039] 103 , executing instruction scheduling operations on the target model according to the register pressure state monitored in real time and the scheduling priority.
[0040] The instructions involved in the embodiments of the present application refer to very long instruction words (VLIW). Here, very long instruction words (VLIW) are an instruction set architecture. In this architecture, the length of an instruction word is very long, and it can specify multiple independent operations at the same time. These operations can be executed in parallel on different functional units of the processor. For example, a VLIW instruction may contain an integer addition operation, a floating-point multiplication operation, and a data loading operation at the same time, and these operations can be performed simultaneously on the integer arithmetic unit, the floating-point arithmetic unit, and the storage unit, thereby improving the parallel processing capability of the processor.
[0041] It is worth noting that the length of VLIW instructions is much longer than traditional instructions. Its format is to combine the encoding of multiple basic operations (such as arithmetic operations, logical operations, data access, etc.) in one long instruction word. These operation fields are independent of each other in the instruction, and each field corresponds to a specific functional unit. For example, a typical VLIW instruction may have a fixed bit segment allocation, one part of the bits is used to specify the operation of the integer operation unit, another part of the bits is used for the operation of the floating-point operation unit, and there are bit segments for controlling storage units, etc.
[0042] The operations in the instructions are explicitly parallel, that is, the compiler or programmer needs to explicitly group operations that can be executed in parallel together to form VLIW instructions. This is different from some other parallel architectures (such as superscalar architectures, whose hardware can dynamically detect and execute parallel operations).
[0043] In a VLIW processor, there are multiple functional units, such as integer unit, floating point unit, load / store unit, etc. After a VLIW instruction is fetched and decoded, each operation field in the instruction is sent to the corresponding functional unit for simultaneous execution. For example, suppose there is a VLIW processor with three functional units: ALU (arithmetic logic unit), FPU (floating point processing unit) and MEM (memory access unit). A VLIW instruction includes an addition operation executed in the ALU, a multiplication operation executed in the FPU, and a data read operation executed in the MEM. These three operations can be performed simultaneously on their respective functional units within one clock cycle, which effectively utilizes the hardware resources of the processor and improves the execution efficiency of instructions.
[0044] In an embodiment of the present application, the computing task graph includes a plurality of nodes and connecting edges between the plurality of nodes. Further optionally, the plurality of nodes and connecting edges between the plurality of nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model.
[0045] In the computing task graph constructed in the embodiment of the present application, nodes and connecting edges are two core components. They have similar setting concepts to the directed acyclic graph (DAG) used in traditional general computing task instruction scheduling, but are further expanded and refined for large model computing tasks.
[0046] Nodes represent specific computing operations in large models, which cover many types. For example, in the common neural network structure of deep learning large models, such as convolution operations in convolutional layers, pooling operations in pooling layers, matrix multiplication and addition in fully connected layers, each independent operation of this type can be abstracted as a node. Taking a simple convolutional neural network (CNN) as an example, an operation of convolving an input feature map with a specific convolution kernel is a node, which clearly identifies the relatively independent execution units in the large model calculation process, facilitating the subsequent analysis of the relationship and dependency between operations.
[0047] The connection edge has two key functions. The first is to reflect the dependency between operations, that is, whether the execution result of a computing operation (corresponding to a node) is a necessary condition for the start of another computing operation. If so, a directed connection edge will be established between the two nodes to indicate the order constraint. The second is to represent the instruction delay, that is, the time interval from the completion of an operation to the start of the next operation that depends on it. This delay information is very important for accurately arranging the instruction scheduling sequence and estimating the overall computing time. For example, after the convolution layer is calculated, its output is used as the input of the pooling layer operation, then there will be a connection edge between the nodes representing the two computing operations, and the delay attribute attached to this edge may be determined by factors such as hardware processing speed and data transmission bandwidth.
[0048] Large models themselves have complex structures and computational processes, often showing multi-level characteristics. For example, in a deep neural network (DNN), starting from the input layer, the data goes through various computational operations in multiple hidden layers in sequence, and finally reaches the output layer to output the results. The computational tasks within each layer and between layers are interrelated and have a sequence, forming a natural hierarchical structure. When constructing a computational task graph, the numerous nodes and the connecting edges between them are arranged in layers according to the actual situation of the instruction scheduling operations corresponding to each computational task in the target model.
[0049] Taking a typical CNN containing multiple convolutional layers, pooling layers and fully connected layers as an example, the input layer data first enters the first convolutional layer for feature extraction. The multiple convolution operation nodes corresponding to this layer and the data-dependent connection edges between them can be classified as the first layer; the feature map after convolution then enters the pooling layer for downsampling operation, and this group of nodes and connection edges corresponding to the pooling operation constitute the second layer; then the pooled result is passed to the subsequent convolutional layer, pooling layer, etc. for further processing, and so on, layer by layer, until the final fully connected layer outputs the classification result. The whole process arranges the nodes and connection edges into multiple different levels in an orderly manner according to the sequence and stage of the computing tasks, intuitively showing the complete computing task hierarchy of the large model from input to output.
[0050] This multi-level arrangement further deepens the presentation and utilization of complex dependencies in large models. In the same level, the connection edges between nodes reflect the close dependence and order relationship between the computing operations in that stage. For example, in a convolution layer, the computing operations of different convolution kernels may have data sharing or execution order requirements. These relationships can be clearly sorted out through the connection edges within the level.
[0051] The connecting edges between different levels reflect the cross-stage dependencies, such as the dependencies between the convolutional layer and the pooling layer, the pooling layer and the subsequent convolutional layer mentioned above, and because each level is located at a different position in the entire large model calculation process, its importance and timeliness to the overall calculation are also different. Based on such a multi-level structure, when optimizing instruction scheduling, it is possible to more clearly determine from a global perspective which operations are core operations on the critical path, which operations can be executed in parallel or appropriately delayed without affecting the overall progress, etc., which helps to more accurately determine the scheduling priority of each instruction, thereby achieving efficient scheduling of large model calculation tasks and improving overall computing efficiency.
[0052] In addition to the hierarchical arrangement based on the traditional computing operation dependencies, the data access relationship related to the hardware-specific computing units is also fully considered to improve this structure. Taking the matrix multiplication computing unit as an example, in large models, matrix multiplication operations are often distributed in different levels. For example, a large number of matrix multiplication operations are involved in the fully connected layer. When expanding resource dependencies based on hardware characteristics, for these nodes involving matrix multiplication, their connection edges not only reflect the conventional calculation order dependencies and instruction delays, but are also additionally related to the data access dependencies at the hardware level.
[0053] For example, after a matrix multiplication operation of a fully connected layer is completed, the result data will be stored in a specific location according to the hardware regulations. If other computing operations (possibly at the same level or subsequent levels) need to use this data, the connection edges between the nodes representing these operations will reflect this dependency based on hardware data access. This integration of hardware characteristics into a multi-level structure makes the computing task graph more consistent with the operation of large models on actual hardware platforms, providing a more comprehensive and accurate basis for further optimizing instruction scheduling, ensuring that instructions can be efficiently executed based on full use of hardware resource advantages, and avoiding performance problems caused by hardware resource limitations or unreasonable data flow.
[0054] In summary, by arranging the nodes and connecting edges in the computing task graph into multiple levels based on the instruction scheduling operations of the target model, not only the computing structure and dependencies of the large model are clearly presented, but also a solid foundation is laid for subsequent instruction scheduling optimization and hardware resource adaptation, helping the large model to achieve efficient and accurate calculations on the corresponding computing platform.
[0055] In the embodiment of the present application, the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph. The scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction.
[0056] From the perspective of the importance of instruction paths, in the large model calculation process constructed by the calculation task graph, different instruction paths have different degrees of influence on the final results. Some instruction paths contain operations that play a key role in the core calculation of the model, such as weight matrix multiplication in the forward propagation process and gradient calculation paths in the back propagation process in deep neural networks. The calculation results of these paths directly affect the output and training effect of the model, so the instructions on the path are of high importance.
[0057] The importance of the instruction path is also reflected in the depth of data dependency. If an instruction is at a deeper position in the dependency chain, that is, it depends on the results of multiple previous instructions, then the execution of this instruction will have a chain reaction on the entire calculation. Once the instruction cannot be executed on time, it may cause multiple subsequent instructions to wait, thus affecting the overall calculation progress. Therefore, instructions on the critical path and with deep data dependencies are usually given higher scheduling priority.
[0058] From the perspective of instruction path timeliness, timeliness mainly involves the impact of the timing of instruction execution on the overall computing efficiency. For instructions that can use hardware resources in a timely manner and will not cause idle resources, their timeliness is higher. For example, when a computing unit (such as a vector computing unit or a matrix multiplication unit) is idle, if the matching instructions (such as vector operations or matrix multiplication operations) can be executed immediately, it can reduce resource waiting time and improve hardware utilization. Therefore, such instructions have advantages in terms of timeliness.
[0059] In addition, the execution results of some instructions are urgently needed by multiple subsequent instructions. Timely execution of these instructions can avoid blocking subsequent instructions. For example, in a pipelined computing process, instructions that are upstream and provide input data for multiple downstream stages have strong timeliness and need to be scheduled first to keep the computing process smooth.
[0060] Exemplarily, after a comprehensive evaluation of the importance and timeliness of the instruction path, the execution order is determined for each instruction. Instructions with high importance and strong timeliness will be placed at the top of the execution order. This is achieved by topological sorting of the computational task graph and other methods. For example, key instructions that have no predecessor dependencies or whose predecessor dependencies have been completed will be prioritized for execution. For instructions with multiple predecessor dependencies, they will be placed in the execution sequence only after all dependent instructions have been executed.
[0061] When considering the execution order, the parallelism between instructions is also taken into account. If there is no data dependency between some instructions and the hardware resources allow, they can be executed at the same time. This can make full use of the parallel computing capabilities of the processor and further improve computing efficiency. For example, in neural network calculations in different branches or independent computing tasks at different levels, these instructions can be carried out in parallel by arranging the execution order reasonably.
[0062] In the embodiments of the present application, different instructions have different requirements and usage methods for register resources. For instructions with high computational complexity and involving a large amount of intermediate data storage (such as complex matrix multiplication operations), more registers need to be allocated to temporarily store data to avoid frequent data read and write operations and improve the operation speed. For simple instructions, only a small number of registers may be needed.
[0063] According to the position and function of the instruction in the computing task graph, as well as its execution order, appropriate register resources are allocated for it. During the scheduling process, it is necessary to ensure that the allocation of register resources can meet the execution requirements of the instructions while avoiding resource waste and conflicts. For example, when two instructions compete for the same set of registers, it is necessary to reasonably arrange the time and method of register usage based on their priority and execution order.
[0064] In summary, by giving priority to scheduling instructions that are on the critical path and have high importance, the waiting time of these instructions can be effectively reduced, and the critical path execution speed of the entire computing task can be accelerated. For example, in the training process of a deep learning model, timely execution of key instructions such as gradient calculation and weight update can make the model converge faster and reduce the training cycle. Reasonable execution order arrangement can make full use of hardware resources. When instructions are scheduled according to timeliness, resource idleness and waiting can be avoided, so that the various computing units and registers of the processor are in working state most of the time. For example, when executing some instructions in parallel, the computing power of multi-core processors can be fully utilized, and at the same time, through reasonable register resource configuration, data read and write conflicts can be reduced and the overall computing throughput can be improved.
[0065] In addition, register resource configuration can be allocated according to the actual needs of instructions to avoid multiple instructions competing for limited register resources at the same time. For example, by allocating specific register groups to instructions of different priorities and types, or adopting a dynamic allocation strategy to adjust according to the execution stage of the instruction and resource usage, register conflicts can be effectively reduced and the stability of instruction execution can be improved. Register resources are allocated according to the position and importance of instructions in the computational task graph, realizing on-demand resource allocation. Sufficient resource support is given to important and computationally complex instructions, and resource allocation is reasonably controlled for simple instructions, thereby improving the overall resource utilization. For example, for a large matrix multiplication instruction, sufficient registers are allocated to store intermediate data and calculation results to ensure that it can complete the calculation efficiently, while for some simple logic judgment instructions, only a small number of necessary registers are allocated.
[0066] This scheduling priority associated with the importance and timeliness of the instruction path, as well as the execution order and register resource configuration, enables instruction scheduling to adapt to different large model computing scenarios. Whether it is model reasoning or training, whether it is processing simple neural networks or complex multi-branch models, the scheduling priority of instructions can be flexibly adjusted according to the specific computing task graph and hardware resource conditions to achieve the best computing effect.
[0067] During the computing process, as the hardware resource status changes (such as changes in register pressure) and the computing task progresses, the scheduling priority can be adjusted dynamically. For example, when register pressure increases, instructions with high register resource requirements can be appropriately postponed, and instructions with low resource requirements can be executed first; when the data of a key instruction is prepared in advance, it can be executed in advance. This dynamic adjustment capability can better cope with various emergencies and real-time changes, and ensure the efficient completion of computing tasks.
[0068] Based on the principles introduced above, in the above 101, by converting the target model into a computing task graph, using nodes to represent the instruction scheduling operations corresponding to each computing task, and using connecting edges to reflect the relationship between them, and arranging them in a hierarchical manner, the complex internal structure of the target model and the dependencies between the computing tasks can be clearly and intuitively presented. For example, in a large deep learning model, the sequence and data transfer dependencies between different layers of neural network calculations (such as the computing tasks corresponding to the convolutional layer, pooling layer, and fully connected layer) can be accurately reflected in the computing task graph. This allows subsequent instruction scheduling to be performed based on this accurate structural information, avoiding unreasonable scheduling due to unclear understanding of the model calculation process, thereby improving the rationality of the overall instruction scheduling.
[0069] In 102, the scheduling priority of each instruction is determined based on the computational task graph, taking full account of the importance and timeliness of the instruction path in the graph. For instructions that are on the critical path and have a greater impact on the timeliness of the overall computational results, higher priority is given and the execution order is prioritized, which can effectively reduce unnecessary waiting time and avoid slowing down the entire model calculation progress due to delays in the execution of key instructions. For example, in the forward propagation process of a deep neural network, the path where the core computational tasks such as matrix multiplication are located is often the critical path. Prioritizing the scheduling of these instructions can speed up the flow of data in the network, thereby improving the computational efficiency of the entire model.
[0070] The scheduling priority includes the register resource configuration of each instruction, which means that register resources are not uniformly allocated to all instructions, but are differentially configured according to factors such as the role and importance of the instruction in the computing task graph. For instructions with high computational complexity, high data dependency, and large impact on the overall calculation, relatively more and more appropriate register resources can be allocated to ensure their smooth and efficient execution; while for relatively simple instructions with less overall impact, fewer register resources are reasonably allocated to avoid resource waste. For example, when performing large-scale matrix multiplication instructions, sufficient registers are allocated to temporarily store intermediate data to prevent frequent data read and write conflicts and improve the computing speed, thereby achieving a good match between register resources and instruction requirements and improving the scientific nature of resource allocation.
[0071] In 103, instruction scheduling operations are performed according to the real-time monitored register pressure status and the established scheduling priority, so that register resources can be dynamically optimized during operation. When it is detected that the register pressure is high, instructions that require less register resources can be scheduled first, or the instructions that have been scheduled to use a large number of registers can be appropriately delayed or adjusted to reduce excessive use and conflicts of register resources; when the register pressure is low, idle resources are fully utilized to schedule more instructions in parallel to improve the throughput of the processor. This dynamic adjustment mechanism ensures that register resources can always be efficiently utilized according to actual conditions, avoiding idle resources or performance bottlenecks caused by resource shortages.
[0072] Whether it is a convolutional neural network, a recurrent neural network or other complex deep learning large model structure, this technical solution can sort out its internal computing task relationship by constructing a computing task graph, and then determine the appropriate scheduling priority and implement dynamic scheduling. This means that it has strong versatility and adaptability, and can cope with the differences in computing tasks, data dependencies, etc. of different large models, flexibly arrange instruction scheduling according to the characteristics of specific models, and ensure that various large models can run efficiently on the corresponding hardware platform.
[0073] During the actual operation of large models, factors such as computing load, data preparation, and hardware resource status may change at any time. With the help of real-time monitoring of register pressure status and dynamically adjustable scheduling priorities, the solution can respond to these changes in a timely manner. For example, if the transmission of some data is delayed at a certain moment, causing the execution conditions of related instructions to change, the scheduling system can quickly adjust the execution order and resource allocation of instructions according to the new situation to ensure that the entire computing process can still be carried out efficiently and stably, reflecting a high degree of flexibility in dealing with various changes in runtime.
[0074] Taking into account the technical effects of the above aspects, the embodiments of the present application enable the large model to realize reasonable and efficient scheduling of instructions, scientific configuration and dynamic optimization of register resources, and flexible adaptation to different situations during the calculation process. Thus, the calculation speed of the large model is significantly improved, the waste of computing resources is reduced, and the output efficiency of the computing results is improved, so that the hardware platform equipped with this technology has more performance advantages when processing large model-related tasks, and enhances its competitiveness in applications in the fields of artificial intelligence, big data, etc., and meets the growing demand for efficient computing of large models.
[0075] As an optional embodiment, in 101, for the target model to be processed, the target model is converted into a computing task graph, such as Figure 2 As shown, it can be implemented as:
[0076] 201. Identify each instruction included in the target model, and identify the topological relationship between each instruction as the dependency relationship between each instruction;
[0077] 202. Construct each instruction in the target model into a corresponding node, and construct a connection edge between each node according to the dependency relationship and / or instruction delay between each instruction;
[0078] 203. Optimizing the dependency relationship according to the hardware device resources deployed by the target model;
[0079] 204. Based on the optimized dependency relationship, optimize the structural relationship between each node and each connection edge to obtain the computing task graph.
[0080] When processing the target model to be processed and converting it into a computational task graph, first, for step 201, a convolutional neural network (CNN) is used as an example of the target model. This CNN model contains various instructions such as convolution operation instructions of the convolution layer, pooling instructions of the pooling layer, matrix multiplication of the fully connected layer, etc. By analyzing the model code or algorithm description, the execution order and data transfer conditions between these different instructions are sorted out to determine the topological relationship between them, and this topological relationship is the dependency relationship between the instructions. For example, the output after the convolution layer calculation is completed will be used as the input of the pooling layer, so there is a dependency relationship between the two instructions, and the execution of the pooling layer depends on the convolution layer to complete the corresponding operation first.
[0081] Next, in step 202, still taking the CNN model as an example, each specific instruction such as convolution operation, pooling operation, matrix multiplication, etc. is constructed into a corresponding node. Then, the connection edge between the nodes is constructed based on the previously identified dependencies and instruction delays. For example, the dependency between the convolution layer and the pooling layer mentioned above will construct a directed connection edge between the node representing the convolution operation and the node representing the pooling operation, and this edge may also be accompanied by instruction delay information. This delay may be determined by factors such as hardware processing speed and data transmission bandwidth. For example, it may take several clock cycles from the output of the convolution layer to the ability of the pooling layer to start processing data. This time period is reflected in the instruction delay attribute of the connection edge.
[0082] Then, in step 203, it is assumed that the target model is to be deployed on a device with a specific hardware architecture, which has an efficient computing unit specifically used for matrix multiplication, as well as limited register resources, etc. When analyzing the dependency, in addition to considering the logical sequence dependency of the instructions themselves, optimization will also be performed in combination with the hardware resource situation. For example, for some instructions that require a large number of matrix multiplication computing units, if the execution order in the original dependency may cause the computing unit to frequently switch tasks or have idle waiting time, then the execution order of the relevant instructions can be adjusted according to the number and performance of the matrix multiplication computing units in the hardware device, so that these instructions can use the hardware resources more efficiently, optimize the dependency between them, and avoid unnecessary hardware resource idleness or use conflicts.
[0083] Finally, in step 204, based on the dependency relationship after hardware resource optimization, the structural relationship between the already constructed nodes and the connecting edges is further improved and adjusted to finally obtain the computing task graph. For example, after optimizing the dependency relationship, it is found that some instructions that can be executed in parallel are actually more suitable for sequential execution after considering the allocation of hardware resources, or the sequence of some instructions needs to be fine-tuned to better adapt to the use rhythm of hardware resources. Then, the connection method and order of the connecting edges between the nodes are changed accordingly, so that the structure of the entire graph is more in line with the requirements of efficient execution on the hardware device. The computing task graph thus obtained can clearly and accurately reflect the computing task structure of the target model in a specific hardware environment and the relationship between the tasks.
[0084] In this way, the final computing task graph has many positive effects. First, the computing task graph can intuitively present the complex computing structure and dependencies between instructions within the target model. Whether the developer is viewing the execution process of the model on a specific hardware, or performing subsequent operations such as instruction scheduling, he can quickly and accurately understand the sequence and mutual constraints of each computing task based on this graph, avoiding problems such as incorrect scheduling or unreasonable resource allocation due to unclear understanding of the model calculation logic. Second, since the dependency is optimized by combining hardware device resources during the construction process, subsequent operations such as instruction scheduling based on the computing task graph can better adapt to the hardware. For example, when performing instruction scheduling to determine the execution order, the characteristics of the hardware's computing unit can be fully utilized, and the execution time of different instructions on different hardware resources can be reasonably arranged to improve the utilization of hardware resources, reduce the performance bottleneck caused by hardware resource restrictions or unreasonable use, thereby improving the computing efficiency of the entire target model on the hardware device, and ensuring that the model can complete the corresponding computing tasks quickly, stably and efficiently, such as model training or reasoning operations can benefit from it.
[0085] In general, these steps help to closely integrate the target model with the actual hardware resources, laying a solid foundation for the subsequent efficient instruction scheduling and optimized computing on the hardware platform.
[0086] Further optionally, in the above step 203, optimizing the dependency relationship according to the hardware device resources deployed by the target model may be implemented as follows:
[0087] Identify the instructions to be optimized in the target model that are associated with the hardware device resources; the hardware device resources include hardware computing units, and the hardware computing units are used to accelerate part or all of the computing tasks performed by the instructions to be optimized; the hardware device resources include at least a VLIW processor; obtain multiple computing operations corresponding to the instructions to be optimized; according to the data access relationship of the hardware computing unit, parse the dependency relationship between the multiple computing operations at the hardware resource level; wherein the dependency relationship at the hardware resource level includes at least: input / output relationship, data storage relationship, reading order, input / output format; based on the dependency relationship at the hardware resource level, expand the local dependency relationship corresponding to the multiple computing operations to obtain the optimized local dependency relationship; integrate the optimized local dependency relationship to obtain the optimized dependency relationship.
[0088] In traditional instruction scheduling scenarios, for general computing tasks, the dependencies between instructions are often analyzed based on register data. For example, by tracking the reading and writing of data in registers, it is determined whether the execution of an instruction depends on the write operation of another instruction to the register. However, when faced with complex tasks such as deep learning large models with unique computing patterns, this dependency analysis based solely on register data is not comprehensive and in-depth enough. Deep learning large models contain many complex computing operations, which have multi-level and intricate dependencies between these operations, and are closely related to the computing resources of the hardware during actual execution. The existing technology usually fails to fully consider these complex dependencies when scheduling instructions, resulting in the execution order between instructions failing to reach the optimal state, thereby limiting the parallel execution capability of the processor and affecting the overall computing efficiency.
[0089] In the embodiment of the present application, the computing task of the large model is represented as a directed acyclic graph (DAG). In this DAG, each node represents a specific computing operation in the large model, such as matrix multiplication, activation function calculation, etc. The edge carries two important information. On the one hand, it reflects the sequential dependency between operations, that is, one operation must be executed after another operation is completed; on the other hand, the edge also represents the instruction delay, that is, the time interval required from the completion of one operation to the start of the next related operation. Through such a DAG construction, the originally obscure and complex multi-level dependencies in the large model can be clearly displayed. For example, in the forward propagation process of a deep neural network, the activation function calculation operation of a certain layer of neurons may depend on the output of the neurons in the previous layer (corresponding to the previous matrix multiplication operation result), and these dependencies can be intuitively represented in the DAG through the connection of nodes and edges. This enables the scheduler to know the intrinsic connection between each computing operation at a glance when performing instruction scheduling, and provides a basic framework for the subsequent reasonable arrangement of the instruction execution order.
[0090] Relying solely on the conventional dependencies presented by DAG is not enough to fully tap the optimization potential of large models on specific hardware platforms, so this solution further expands the dependencies based on hardware-specific computing units. Taking the matrix multiplication computing unit as an example, in deep learning large models, matrix multiplication is a very core and frequently occurring computing operation, and different matrix multiplication operations have specific requirements and associations in the data access method at the hardware level. For example, after a matrix multiplication computing unit performs a multiplication operation, its result data will be stored in a specific cache location. If subsequent related computing operations (which may be the calculation of the next layer of the network, which requires the result of this multiplication as input) want to obtain data efficiently, they need to follow the data access path specified by the hardware. This kind of data dependency at the hardware level is often overlooked in traditional dependency analysis, but this solution takes it into consideration and establishes resource dependencies from the perspective of hardware data access. This means that when scheduling instructions, not only the logical order of the computing operations themselves will be considered, but also the limitations and requirements of hardware resources on data flow and use will be fully taken into account, ensuring that when instructions are executed on the hardware, they can obtain the required data more smoothly and efficiently, avoiding problems such as data waiting and access conflicts that affect execution efficiency.
[0091] Through the above-mentioned in-depth mining and expansion of the dependency relationship of large model instructions combined with hardware characteristics, on the one hand, the execution order of each instruction can be determined more accurately. The scheduler can reasonably arrange which instructions to execute first and which to execute later based on the logical dependency and hardware resource dependency presented by the DAG, avoiding unnecessary waiting or wrong execution order caused by improper dependency processing, making the entire calculation process smoother and more orderly. On the other hand, this optimized dependency processing creates conditions for maximizing the parallel execution capability. Since the dependencies between instructions are clearly sorted out, the scheduler can accurately identify which instructions are not mutually dependent, or which instructions can be executed in parallel through reasonable resource allocation and time scheduling although they have certain dependencies. For example, the calculation operations in different branch networks, if the data they depend on is ready and the hardware resources allow, can be executed in parallel on different computing units at the same time, thereby making full use of the parallel computing resources of the processor, speeding up the calculation speed of the large model, and overall improving the efficiency of instruction scheduling and the performance of the processor when processing large models. In summary, this innovative optimization method for large model dependencies, from the perspective of being more in line with the actual computing characteristics of large models and hardware resource utilization, makes up for the shortcomings of existing technologies in this regard, and provides a strong guarantee for the efficient execution of large models on VLIW processors.
[0092] Based on the above principles, among the many computing instructions in large models (such as convolutional neural networks, deep neural networks, etc.), there are some instructions that are closely related to hardware device resources, and these are the instructions to be optimized. Taking the common large deep learning model as an example, the convolution operation instructions in the convolution layer and the matrix multiplication instructions in the fully connected layer can use the specific computing units in the hardware device to accelerate the calculation during execution. For example, the VLIW processor is equipped with an efficient computing unit specifically for matrix multiplication, so the related instructions involving matrix multiplication in the model belong to this type of instructions to be optimized, because they can use the hardware computing unit to complete part or all of the computing tasks faster. By first identifying these instructions to be optimized, the subsequent optimization based on the dependency relationship of hardware resources can be accurately identified.
[0093] For each instruction to be optimized that has been identified, further disassemble the specific calculation operations they contain. Taking the matrix multiplication instruction as an example, it may involve multiple specific calculation operations such as reading the input matrix data, multiplication and accumulation operations between matrix elements, temporary storage of intermediate results, and output of the final result. These specific calculation operations that constitute the instructions to be optimized are sorted out in order to analyze the relationship between them at the hardware resource level. This is the basic analysis step for the subsequent optimization dependency.
[0094] The hardware computing unit has its own specific data input and output methods as well as internal data storage and reading rules. For example, the hardware computing unit can be the matrix multiplication computing unit and vector computing unit mentioned above. For example, the matrix multiplication computing unit often has a specific cache structure to temporarily store intermediate results, and has strict requirements on the format and reading order of the input matrix data. After the calculation is completed, the result data will be stored in the corresponding cache or register position according to a fixed path for subsequent related operations. Based on these hardware characteristics, the dependencies between the multiple computing operations contained in the instructions to be optimized are analyzed. For example, the output data after a computing operation is completed must be used as the input of another computing operation, and this input must be obtained according to the input format and reading order specified by the hardware computing unit, which forms a dependency relationship at the hardware resource level. For example, in terms of input / output relationships, it is clear which operation's output is the input of which subsequent operation; in terms of data storage relationships, it is known where the calculation result will be stored in the hardware, and how the subsequent operation obtains data from this location; for the reading order, it is necessary to follow the order required by the hardware to read the corresponding data, etc. These are the dependencies between multiple computing operations analyzed from the perspective of hardware resources.
[0095] There are certain local dependencies between the computing operations originally determined by the logic of the large model itself. For example, a certain activation function computing operation depends on the output result of the previous convolution operation, which is a dependency at the computing logic level. Now, combined with the dependencies analyzed from the hardware resource level (such as data storage location, reading order, and other factors that affect the order of operation execution), these local dependencies are expanded. For example, if a computing operation not only needs to wait for the previous operation to complete the output logically, but also needs to wait for the data to be stored in a specific location and meet the corresponding format and order according to the requirements of the hardware computing unit before it can start execution, then this hardware resource-related constraint is included in the dependency, and the original local dependency is enriched and expanded, so as to obtain the optimized local dependency, which is more in line with the actual operation of the hardware and avoids execution conflicts or inefficiencies caused by ignoring hardware resource requirements.
[0096] In the large model, the computing operations corresponding to the instructions to be optimized are distributed in different parts and different levels, which will form multiple local dependencies. After optimizing each local dependency through the above steps, these optimized local dependencies need to be integrated. For example, the relationship between the computing operations involving hardware resource dependencies in different layers is unified, and all the expanded and optimized dependencies are integrated together to finally form a complete optimized dependency that fully considers the hardware device resource situation. In this way, the dependency of the entire large model is no longer limited to the logical level, but fully combines the various constraints and associations at the hardware resource level, laying a solid foundation for the subsequent optimization of instruction execution order based on this, better leveraging the hardware performance advantages, and improving the computing efficiency of the large model on the VLIW processor.
[0097] By optimizing the dependencies in the above steps, the instruction scheduling of large models on specific hardware devices (especially hardware environments including VLIW processors) can be more scientific and reasonable, fully adapt to hardware resources, and achieve efficient computing task execution.
[0098] Further optionally, in the above steps, it is assumed that the computing task types supported by the hardware computing unit include at least one of the following: matrix multiplication operation, activation function calculation, vector operation. The hardware computing unit includes at least one of the following: matrix multiplication calculation unit, activation function calculation unit, vector calculation unit.
[0099] Specifically, in the above steps, given the types of computing tasks supported by the hardware computing unit and the specific units involved, it plays an important role in optimizing dependencies. For matrix multiplication operations, the matrix multiplication computing unit has its own specific rules. For example, the input matrix needs to be passed in according to the specified format, the intermediate results generated by the calculation will be temporarily stored in a specific cache, and subsequent related operations must obtain data according to the storage location and reading order. For example, in the fully connected layer of a deep neural network, a large number of matrix multiplications are involved. When optimizing dependencies, it is necessary to clarify the association between the previous and subsequent operations at the hardware resource level, ensure that data transfer meets the unit requirements, and expand the corresponding local dependencies. In terms of activation function calculation, the activation function calculation unit also has its own characteristics. Taking the ReLU activation function as an example, its input data comes from the results of the previous operation, and the output must be based on the storage and output rules set by the hardware unit for subsequent calculations. When sorting out dependencies, it is necessary to consider its dependence on adjacent computing operations at the hardware resource level such as data access based on these hardware requirements, and improve local dependencies. The vector calculation unit corresponding to vector operations is no exception. It has requirements for the input format of vector data, the order of operations, and the storage of results. In a large model, if a layer has vector operations that intersect with other operations, optimizing the dependencies requires combining the data access relationship of the vector computing unit, accurately grasping its hardware dependencies with surrounding computing operations, expanding and integrating local dependencies, and ultimately optimizing the overall dependencies, allowing the large model to run more efficiently on hardware.
[0100] Further optionally, it is assumed that the computing task type supported by the hardware computing unit is matrix multiplication operation. Based on the above assumption, in the above steps, based on the dependency relationship at the hardware resource level, the local dependency relationship corresponding to multiple computing operations is expanded to obtain the optimized local dependency relationship, which can be implemented as follows:
[0101] Detect at least one node related to the matrix multiplication operation; reset the node position of the at least one node in the computing task graph based on the data storage location corresponding to the matrix multiplication operation and the deployment information of the hardware resources; and / or establish a connection edge between the at least one node and other nodes in the computing task graph.
[0102] When the local dependencies corresponding to multiple computing operations are expanded based on the dependencies at the hardware resource level to obtain optimized local dependencies, when the computing task type supported by the hardware computing unit is matrix multiplication. Specifically, first, detect at least one node related to the matrix multiplication operation. In the constructed computing task graph, the nodes representing the specific computing operation of matrix multiplication will be identified, and these nodes carry the relevant information of the matrix multiplication operation in the entire large model computing process. For example, in a deep neural network, the nodes corresponding to a large number of matrix multiplication operations in the fully connected layer are the focus of attention because they play a key role in promoting the overall calculation. Then, based on the data storage location corresponding to the matrix multiplication operation and the deployment information of the hardware resources, the positions of these nodes in the computing task graph are reset. After the matrix multiplication operation is completed, the result data will be stored in a specific location according to the hardware regulations, such as in a specific cache or register, and this storage location will affect the subsequent related operations to obtain data and the overall calculation order. Then, based on such data storage location and hardware resources (such as the layout of different computing units in the hardware, the setting of data transmission channels and other deployment information), the position of the corresponding node in the computing task graph is adjusted. For example, if the storage location of the result data of a matrix multiplication operation is close to the reading location of the data required for another subsequent key computing operation, in order to make the data transmission more efficient and in line with the optimal logic of hardware resource utilization, the position of the node representing the matrix multiplication operation in the computing task graph will be moved closer to the node corresponding to that key computing operation, so that their sequence and connection relationship are more in line with the actual operation of the hardware.
[0103] At the same time, connection edges will be established between these nodes and other nodes in the computing task graph. Because the result of matrix multiplication is often the input of other subsequent computing operations, from the perspective of hardware resources, not only the logical transmission relationship of data should be considered, but also the requirements of hardware for data access, etc. Therefore, when it is found that the result data of matrix multiplication operation is to be used for other specific computing operations, a connection edge will be established between the nodes representing the matrix multiplication operation and those nodes that need to use its results, so as to clarify the dependency between them under the constraints of hardware resources, and this connection edge may also be accompanied by attribute information such as data transmission time and hardware reading and writing rules. Through such an expansion method, the local dependency relationship is improved to make it more in line with the actual situation at the hardware resource level, thereby obtaining an optimized local dependency relationship, ensuring that the instruction scheduling and calculation of large models on hardware are more scientific, reasonable, efficient and orderly.
[0104] For example, assume that there is a simple neural network model including an input layer, a hidden layer, and an output layer. In the hidden layer, there is a matrix multiplication operation for calculating the weighted sum between neurons, as well as subsequent activation function calculation operations. In terms of hardware resource deployment, there is a dedicated matrix multiplication calculation unit in the processor, which is equipped with a cache to temporarily store the intermediate and final results of the matrix multiplication operation, and this cache is adjacent to the subsequent activation function calculation unit in the hardware architecture, with a short data transmission channel and a high bandwidth, which is designed to speed up the transfer of data from matrix multiplication operations to activation function calculations.
[0105] Originally, in the computational task graph, the nodes representing the hidden layer matrix multiplication operations and the nodes representing the activation function calculations were simply arranged in the conventional logical order, and the distance between the two was relatively far (reflecting only the logical dependency relationship). However, based on the hardware resource situation, since the matrix multiplication operation results are stored in a high-speed cache adjacent to it, in order to allow the subsequent activation function calculation to obtain data more quickly and reduce data transmission delays, the position of the nodes representing the matrix multiplication operation in the computational task graph can be moved closer to the nodes representing the activation function calculations, and they can even be set to be adjacent to each other, thereby reflecting the advantage of nearby data transmission in terms of hardware resources. From a hardware perspective, the process of data from the storage location to the use of the activation function is smoother and more efficient, in line with the logic of efficient hardware utilization, and optimizing local dependencies.
[0106] In another example, in a multi-layer convolutional neural network, there are multiple convolutional layers, pooling layers, and fully connected layers. There are a large number of matrix multiplication operations in the fully connected layer, and the hardware resources are equipped with multiple matrix multiplication computing units and a hierarchical cache structure. After the matrix multiplication operation of one of the fully connected layers is completed, its result will be stored in a specific cache at a certain level, and the subsequent calculation of the next fully connected layer or the related calculation operation of the output layer will require this data. And in terms of hardware architecture, different levels of cache and different computing units have different connection methods and data transmission speeds.
[0107] For example, there is a high-speed data transmission link between the cache storing the result of a matrix multiplication operation of a fully connected layer and the computing unit corresponding to the calculation of the next fully connected layer, which is faster than other paths. Then in the computing task graph, the node position representing the matrix multiplication operation of this fully connected layer will be adjusted according to its storage location and the hardware connection with the computing unit of the next layer, and it will be moved toward the node corresponding to the calculation of the next fully connected layer, so that its layout in the graph can better reflect the advantages of fast data transmission on hardware resources, so that subsequent computing operations that rely on the result of the matrix multiplication can obtain data more efficiently, and the reset node position can better fit the hardware resource deployment, reduce the performance loss caused by unreasonable data transmission paths, optimize the local dependencies in the entire computing process, and thus improve the computing efficiency of the entire CNN model on this hardware platform.
[0108] For example, consider a deep neural network model running in a distributed hardware environment, where different computing nodes are responsible for different computing tasks, and data is exchanged between computing nodes through a high-speed network. Some of the nodes are specifically used to perform matrix multiplication operations, and these nodes are equipped with large-capacity local storage for temporarily storing the results of matrix multiplication, so that subsequent computing operations on other nodes can obtain data. For example, after several computing nodes collaborate to complete the matrix multiplication intensive operations of a certain stage of the model, the results are stored in their respective local storages, and other subsequent nodes will perform related calculations, such as comprehensive summary calculations in subsequent stages. When constructing the computing task graph, the nodes representing matrix multiplication operations and subsequent related calculations that were originally arranged in logical order may be scattered in different areas (only reflecting the logical order). However, based on the deployment of hardware resources, and taking into account the local storage of data on specific computing nodes and the network transmission path for subsequent nodes to obtain data, the position of the node representing the matrix multiplication operation will be adjusted to a position relatively close to those nodes responsible for subsequent related calculations and with more convenient network connections. For example, by adjusting their relative layout in the graph, they can be more in line with the hardware characteristics of efficient data transmission in a distributed environment, reduce network transmission delays, and enable the entire computing process to better utilize the advantages of distributed hardware resources, optimize local dependencies, and ensure that the model runs efficiently in this complex hardware environment.
[0109] Further optionally, it is assumed that the computing task type supported by the hardware computing unit is matrix multiplication operation. Based on the above assumption, in the above steps, based on the dependency relationship at the hardware resource level, the local dependency relationship corresponding to multiple computing operations is expanded to obtain the optimized local dependency relationship, which can be implemented as follows:
[0110] Detect at least one node related to the matrix multiplication operation; reset the node position of the at least one node in the computing task graph based on the data storage location corresponding to the matrix multiplication operation and the deployment information of the hardware resources; and / or establish a connection edge between the at least one node and other nodes in the computing task graph.
[0111] Exemplarily, in the process of optimizing the computational task graph, the nodes closely related to the matrix multiplication operation being performed are first accurately detected. These nodes carry the key information of the matrix multiplication operation in the entire large model calculation process. For example, in the fully connected layer calculation of a deep neural network, after the node corresponding to a matrix multiplication operation is identified, it is further processed based on the information at the hardware resource level. It is known that the result data after the completion of the matrix multiplication operation will be stored in a specific cache location based on the hardware design. The relative layout of this cache location and other components in the hardware and the data transmission path constitute the deployment information of the hardware resources.
[0112] Based on this, if there are other subsequent computing operations (such as specific activation function calculations or matrix multiplication preprocessing of the next layer of the network, etc.) that need to use the result data of the matrix multiplication, then the node position relationship corresponding to the two computing operations will be re-examined in the computing task graph based on the storage location of the result data and the rules for obtaining data from the storage location specified by the hardware (such as the order of data reading, format requirements, and the hardware transmission channel passed through, etc.). The node position that needs to use the result data may be appropriately adjusted toward the node that stores the result data to more intuitively reflect the hardware optimal path for data transmission. At the same time, a data dependency edge with clear hardware resource attributes will definitely be established between the two nodes. This edge not only represents the logical order of calculations, that is, the subsequent computing operations depend on the result of the matrix multiplication operation, but more importantly, it also details the key hardware resource information such as the way to obtain data from the hardware storage location, the hardware rules to be followed, and the possible data transmission delay, thereby completing the precise expansion of the local dependency relationship between the two computing operations from the hardware resource level, so that it can better adapt to the actual hardware operating environment, and effectively ensure the efficiency and stability of the computing process of the entire large model on the hardware platform.
[0113] For example, vector operations are very important when processing computing tasks with vector data structures. For example, in graphics processing, vectors in three-dimensional space are translated, rotated, and scaled, or signal vectors are filtered in signal processing. These vector operations can be efficiently processed by a dedicated vector computing unit. Vector computing units usually have strict requirements on the format of vector data. For example, the length of the data, the element type (such as single-precision floating point or integer), etc. need to comply with hardware regulations. In terms of storage, the vector computing unit may have its own cache structure to store vector data, and the order in which the data is read will vary depending on the type of operation (such as dot product, cross product, etc.). Its input-output relationship is also very clear. For example, the output vector of a vector addition operation will be used as the input vector of the next vector multiplication operation, which forms a dependency at the hardware resource level.
[0114] Take activation function as an example. In neural network, activation function is used to introduce nonlinear factors to neurons, so that neural network can fit various complex functional relationships. Activation functions include ReLU (Rectified Linear Unit), Sigmoid and Tanh. Taking ReLU function as an example, its calculation rule is y = max(0, x). For the input neuron signal x, the output is y.
[0115] The activation function calculation unit may be optimized for a specific activation function. For example, for the ReLU calculation unit, it may have a circuit that quickly determines whether the input is greater than 0. In terms of data storage, its input data usually comes from the output of the previous layer of neurons (such as the result after matrix multiplication) and is stored in a specific register or cache. The output data will provide input for the calculation of the next layer of neurons in the format and storage location specified by the hardware. This input-output relationship and data storage requirements constitute a dependency at the hardware resource level, and can be reflected in the computational task graph through nodes and connecting edges.
[0116] For example, convolution operation is the core computing operation of convolutional neural network (CNN). In the field of image recognition, the convolution layer performs convolution operation on the input image (which can be regarded as a two-dimensional or three-dimensional data matrix) through the convolution kernel to extract the features of the image. For example, a 3-convolution kernel is convolved with a 28-grayscale image to obtain a new feature map. Convolution operation hardware units (such as convolution calculation units) usually have their own storage structure to store convolution kernel data and input image data. The data reading order is carried out according to the rules of convolution operation, for example, starting from the upper left corner of the image, sliding the convolution kernel according to a certain step size for calculation. The format and storage location of the input data (image data) and the output data (feature map data) are strictly regulated, and the output feature map data will be used as the input of the next convolution layer or pooling layer, which forms a dependency relationship at the hardware resource level. It is necessary to accurately construct nodes and connection edges in the computing task graph to reflect these relationships in order to optimize instruction scheduling.
[0117] As an optional embodiment, in 102, determining the scheduling priority of each instruction in the target model based on the computing task graph can be implemented as follows:
[0118] Determine the instruction path corresponding to each instruction in the computing task graph; select a key instruction path from the computing task graph; perform static analysis on the criticality and execution timeliness of the key instruction path to obtain the static priority of each instruction in the key instruction path; perform dynamic analysis on the key instruction path, and adjust the static priority based on the analysis result to obtain the scheduling priority of each instruction in the key instruction path.
[0119] In an embodiment of the present application, the key instruction path includes at least one of the following: a core computing path in a long chain dependency relationship, an instruction path containing preset computing operations, an instruction path whose computing complexity reaches a set condition, an instruction path whose data dependency depth reaches a set threshold, and an instruction path at a preset position in the target model.
[0120] Specifically, in the computing tasks of large models, by constructing a directed acyclic graph (DAG) or other dependency analysis methods, we can find out the instruction paths that have a critical impact on the overall computing completion time, that is, the critical path. For example, in the training process of a deep neural network, the core matrix multiplication operations in forward propagation and backpropagation and the gradient update steps often constitute the critical path. The delay of these operations will directly lead to the extension of the entire model training cycle, because many subsequent calculations depend on their results.
[0121] For complex computing operations, such as large-scale matrix multiplication or high-dimensional convolution operations, they consume more computing resources and time. Once a delay occurs, the overall performance will be greatly affected, so a higher importance weight is given. Operations that are deeply dependent on the results of other instructions will trigger a series of waits and delays if there are problems with the execution of their predecessor instructions, so such instructions are also considered more important. For example, in a multi-layer neural network, the computing operations close to the output layer often depend on the calculation results of the previous layers, and their importance is relatively high. Certain key computing steps are directly related to the output accuracy of the model. For example, in classification tasks, the calculation of the fully connected layer of the last layer plays a decisive role in correct classification, and the priority of these instructions will be increased accordingly.
[0122] For static analysis, before the model calculation task begins, a preliminary priority assessment is performed on each instruction based on prior knowledge of the model structure and algorithm. For example, a basic priority value is assigned to instructions in different areas based on information such as the number of layers in the neural network, the number of neurons in each layer, and the type of operations involved. For instructions in deep networks that involve complex operations, a higher static priority is given.
[0123] For dynamic analysis, first, for register idle time, real-time tracking of the preparation of data required by instructions, and appropriately raising the priority of instructions whose data is ready or about to be ready, because they can be executed immediately or start execution in a short time, reducing the idle time of computing resources. Second, for resource utilization, monitoring the use of resources such as processor computing units and registers, and giving priority to those instructions that can make full use of current idle resources to improve the overall utilization of resources. For example, if a vector computing unit is currently idle, and there is a vector operation instruction waiting to be executed, and its dependency relationship allows, then the priority of this instruction will be increased accordingly. Third, for execution time prediction, historical execution data and current hardware status are combined to predict the execution time of instructions. For instructions with short execution time and no impact on the critical path, priority can be given to quickly complete some small computing tasks and release resources for more important instructions; while for those instructions that may occupy critical resources for a long time and are relatively less important, their priority can be appropriately lowered.
[0124] As an optional embodiment, in 103, the instruction scheduling operation for the target model is performed according to the real-time monitored register pressure state and the scheduling priority, which can be implemented as follows: globally monitoring the target model during the scheduling process to obtain the register monitoring data of the target model; dynamically evaluating the register pressure state based on the register monitoring data; selecting the corresponding target register allocation method from the pre-configured register allocation strategy according to the register pressure state; performing the instruction scheduling operation based on the scheduling priority, and dynamically configuring the corresponding register resources during the instruction scheduling process using the target register allocation method. In the embodiment of the present application, the register monitoring data includes at least: the called register data, the number of free registers, the register allocation frequency, and the register release frequency.
[0125] Specifically, by globally monitoring the target model during the scheduling process to obtain register monitoring data, the use of registers can be fully and meticulously understood. These data include the called register data, the number of free registers, the register allocation frequency, and the register release frequency, providing an accurate basis for subsequent operations. Based on these detailed register monitoring data, the register pressure status is dynamically evaluated, and the tightness or abundance of register resources can be understood in real time. This is like having a dynamic resource map, which allows the scheduling system to clearly know whether the current register resources are facing high pressure and whether there are enough free resources to be used, thereby providing a decision-making basis for reasonable resource allocation. According to the evaluated register pressure status, the corresponding target register allocation method is selected from the pre-configured register allocation strategy, which can make the register resource allocation more flexible and adaptable to the actual situation. When the register pressure is high, a more cautious allocation strategy can be adopted to avoid register conflicts and resource exhaustion; when the pressure is low, the idle resources can be fully utilized to maximize the parallel processing capability. This dynamic allocation method effectively avoids the problem of resource waste or shortage that may be caused by the fixed allocation strategy. Finally, the instruction scheduling operation is performed based on the scheduling priority, and the corresponding register resources are dynamically configured in the instruction scheduling process using the target register allocation method, so that the instruction scheduling and register resource allocation are closely integrated. On the one hand, according to the scheduling priority, key instructions are ensured to be executed first, which improves the efficiency of the key path of the computing task and speeds up the overall computing speed; on the other hand, dynamically configuring register resources ensures that instructions have appropriate resource support during execution, reducing delays caused by resource waiting or conflicts. This series of operations work together to effectively improve the resource utilization and overall computing efficiency of the target model during the computing process, allowing computing tasks to be executed more stably and efficiently.
[0126] In another embodiment, a scheduling algorithm based on a priority queue is used to put all instructions to be executed into a queue according to their priority values. In each scheduling cycle, the scheduler takes the highest priority instruction from the head of the queue for execution. When the priority of an instruction changes (for example, due to data readiness or resource utilization changes), the scheduling algorithm can quickly adjust its position in the queue to ensure that high-priority instructions always get priority execution resources.
[0127] For example, in some cases, a preemptive scheduling strategy is adopted. When a high-priority instruction is ready and the currently executed instruction has a relatively low priority, if preemption is allowed (taking into account factors such as the overhead of resource switching), the execution of the current low-priority instruction can be suspended and resources can be allocated to the high-priority instruction to ensure that the instructions on the critical path can be advanced in time.
[0128] For some situations that are not suitable for frequent preemption, such as when some computing operations have reached a critical stage and the interruption cost is high, non-preemptive scheduling is used to wait for the current instruction to be executed before scheduling high-priority instructions. By reasonably balancing preemptive and non-preemptive scheduling, it can not only ensure the timely execution of key instructions, but also avoid excessive resource switching overhead, thereby improving overall scheduling efficiency.
[0129] During the execution of instructions, feedback information such as execution results and resource usage is continuously collected to optimize the priority evaluation model and scheduling strategy. For example, if it is found that an instruction that was originally considered to have a higher priority does not significantly improve the overall performance during the actual execution process, it may be due to inaccurate estimation of its dependencies or resource requirements. In this case, its priority evaluation parameters can be adjusted accordingly to make subsequent scheduling decisions more accurate, further improving the adaptability of instruction scheduling to the actual operating characteristics of large models and the overall computing efficiency.
[0130] Through the above dynamic regional priority division mechanism, instruction scheduling can more flexibly and efficiently deal with complex situations in the large model calculation process, give full play to the performance advantages of the processor, and compared with the traditional static scheduling mode, it has significant innovation and advantages in improving computing efficiency and resource utilization.
[0131] Further optionally, in the above steps, dynamically evaluating the register pressure state based on the register monitoring data can be implemented as: taking the ratio between the number of used registers and the total number of registers as the register utilization; based on the register allocation frequency, register release frequency, and the register utilization obtained by real-time monitoring, determining the register pressure state in a preset manner.
[0132] In this process, the register utilization is first determined by the ratio of the number of used registers to the total number of registers. This is an indicator that directly measures the degree of occupation of register resources. For example, there are a total of 100 registers, and 60 are currently in use, then the utilization rate is 60%, which can clearly present the general situation of resource usage. Then, the register pressure state is determined according to the preset method based on the real-time monitored register allocation frequency, register release frequency and calculated register utilization. The register allocation frequency reflects the frequency of registers being allocated to instructions for use over a period of time, and the release frequency reflects the speed at which registers are released back to the available state after the instruction is executed. When the allocation frequency is high, the release frequency is low and the utilization rate is also high, it often means that the register pressure is high. Many instructions may be competing for limited register resources, which may easily lead to resource shortage or even insufficient situation. At this time, it is necessary to carefully schedule instructions and reasonably allocate registers to avoid conflicts and waiting, so as to ensure that the computing tasks can be carried out in an orderly manner.
[0133] If the allocation frequency and release frequency are relatively balanced and the utilization rate is at a moderate level, then the register pressure is normal, and instruction scheduling and resource allocation can be carried out at a regular pace to ensure a smooth computing process. If the allocation frequency is low, the release frequency is high, and the utilization rate is low, it means that the register pressure is small and there are more idle register resources. At this time, you can be more aggressive and arrange more instructions for parallel execution in instruction scheduling, making full use of these idle resources to improve computing efficiency and speed up the overall computing speed.
[0134] In this way, the real-time status of register resources can be accurately controlled, and instruction scheduling and register allocation can be flexibly adjusted according to actual pressure, avoiding resource waste or computing jams caused by resource constraints. This greatly improves the resource utilization efficiency and execution efficiency of the entire target model calculation process, ensuring that computing tasks are completed efficiently and stably.
[0135] For example, we first need to determine the key indicator of register utilization, which reflects the degree of register resource occupation, and is reflected by comparing the number of used registers with the total number of registers. For example, it is known that the processor has a certain number of registers. At a certain moment, we check and find that the number of used registers has reached a certain point. If the used registers account for half of the total number of registers, it means that the register utilization is 50%, indicating that half of the register resources have been occupied.
[0136] Next, consider register allocation frequency and release frequency. Register allocation frequency refers to the number of times registers are allocated to instructions for use per unit time. To know this number, you can determine it by counting how many times registers are allocated within a specific period of time, such as the past 10 seconds. If they are allocated 50 times in these 10 seconds, the allocation frequency is 5 times / second. The register release frequency refers to the number of times registers are released after the execution of instructions per unit time. Similarly, you only need to count the number of releases within the corresponding time. For example, if 40 registers are released within 10 seconds, the release frequency is 4 times / second. These two frequencies can show the flow of register resources. If the allocation frequency is much higher than the release frequency, it means that the demand for register resources is very strong and resources may become tight; conversely, if the release frequency is higher than the allocation frequency, it means that register resources are relatively loose.
[0137] Then, the register pressure state is determined according to the preset method. Several different pressure state levels can be defined first, such as: low pressure, medium pressure, and high pressure. For the low pressure state, the judgment standard is set as follows: when the register utilization is lower than a relatively low limit value (such as 30%), and the allocation frequency and release frequency meet the condition that the allocation frequency does not exceed the release frequency, the register pressure is judged to be low at this time. This means that there are relatively many free registers at present, and the speed of register allocation does not exceed the speed of release, and the resources are relatively abundant. For the medium pressure state, when the register utilization is in an intermediate range (such as between 30% and 70%), and the allocation frequency and release frequency are relatively close, it can be judged here by setting the difference between the two in a smaller range, then the register pressure is judged to be in a moderate state. In this case, the use of register resources is relatively balanced and can meet the resource requirements of the current instruction scheduling, but it is still necessary to pay proper attention to the resource usage. For the high pressure state, if the register utilization is higher than a higher limit value (such as 70%), and the allocation frequency is much larger than the release frequency (for example, the allocation frequency is larger than the release frequency plus an appropriate value), then the register pressure is judged to be high. This means that register resources are very tight, and some measures must be taken to alleviate resource pressure, such as adjusting the order of instruction scheduling, giving priority to instructions that require less register resources, and other operations. Finally, during the entire calculation process, it is necessary to constantly monitor the usage, allocation frequency, and release frequency of registers in real time, and continuously and dynamically evaluate the register pressure status based on the previously preset judgment criteria. In this way, the instruction scheduling strategy and register resource allocation method can be flexibly adjusted according to the actual situation, thereby achieving the goal of optimizing computing efficiency.
[0138] Further optionally, in the above steps, according to the register pressure state, the corresponding target register allocation method is selected from the pre-configured register allocation strategy, which can be implemented as follows: if the register pressure state is greater than the set pressure threshold, the first instruction whose register resource usage is lower than the set scheduling threshold is preferentially scheduled, and / or the scheduling timing of the second instruction whose register resource usage is higher than the set scheduling threshold is re-planned; if the register pressure state is less than the set pressure threshold, multiple execution instructions are scheduled in parallel, and the usage time of the register resources is extended to improve the processor throughput and reduce the vacancy rate of the register resources.
[0139] In the process of instruction scheduling and register resource management, dynamic strategies are adopted to optimize performance.
[0140] Exemplarily, when the register pressure state is greater than the set pressure threshold, the system will take two main measures. First, the first instructions whose register resource utilization is lower than the set scheduling threshold are scheduled first. This is because these instructions have relatively small requirements for register resources. When resources are tight, executing them first can avoid instruction waiting and execution delays caused by insufficient registers, so that the calculation process can proceed more smoothly and reduce pauses caused by resource competition. Secondly, for the second instruction whose register resource utilization is higher than the set scheduling threshold, its scheduling timing will be replanned. This may mean postponing the execution of these instructions with large requirements for register resources until the register pressure is relieved, or finding other suitable time windows to arrange their execution, ensuring that register resources can be allocated more reasonably and efficiently under high pressure conditions, avoiding resource exhaustion and a sharp drop in system performance.
[0141] When the register pressure state is less than the set pressure threshold, different strategies are used to further improve performance. At this time, multiple execution instructions will be scheduled in parallel to make full use of the parallel computing capabilities of the processor. Because register resources are relatively abundant, multiple instructions can be executed at the same time without worrying about resource conflicts and overly fierce competition, which can greatly speed up the overall computing speed. In addition, the use time of register resources will be extended, that is, there is no rush to release the data in the register. For example, for some registers that may be occupied by data that will be frequently accessed later, even if the current instruction has been executed, if the subsequent instructions may use these data soon, these registers will not be released temporarily, but the data in them will be retained so that subsequent instructions can quickly access them, reducing the overhead of reloading data, thereby reducing the vacancy rate of register resources, improving resource utilization, and improving processor throughput, so that the entire computing system can play a greater performance advantage when register resources are loose, and achieve more efficient computing task execution.
[0142] As an optional embodiment, after dynamically configuring the corresponding register resources in the instruction scheduling process using the target register allocation method in the above steps, the register resources configured for each instruction in the instruction scheduling process can also be determined. Furthermore, a local optimization strategy is adopted to synchronously schedule multiple instructions using the same register in the instruction scheduling process to reduce the register switching frequency and improve the local utilization of the register.
[0143] For example, after dynamically configuring the corresponding register resources during instruction scheduling using the target register allocation method, the next step is to further determine the specific configuration of the register resources for each instruction. This step allows a clear understanding of the correspondence between each instruction and the register.
[0144] On this basis, local optimization strategies are used to optimize instruction scheduling. In the entire instruction scheduling process, focus on multiple instructions that use the same register, and then schedule them synchronously. The advantage of this is that if different instructions use various registers in a scattered manner, they need to switch frequently between different registers during execution, which will consume a certain amount of time and resources and reduce efficiency. By arranging instructions that use the same register together for synchronous scheduling, the frequency of such register switching can be effectively reduced, so that when executing these instructions, the registers can be used more continuously without switching back and forth between different register resources, thereby improving the local utilization of registers, allowing register resources to be more fully and efficiently used in a local range, which helps to further improve the execution efficiency of the entire computing task, better leverage the advantages of hardware resources, and ensure that the computing process is carried out more smoothly and efficiently.
[0145] As an optional embodiment, after the target register allocation method is used to dynamically configure the corresponding register resources during the instruction scheduling process in the above steps, the dynamic utilization rate of the register resources during the instruction scheduling process can also be predicted through the register allocation prediction model. Furthermore, the minimum interference allocation strategy is adopted to dynamically allocate vacant register resources to different instructions according to the predicted dynamic utilization rate during the instruction scheduling process to reduce the frequency of register renaming and data movement.
[0146] For example, after completing the dynamic configuration of register resources in the instruction scheduling process according to the target register allocation method, in order to make the register resource allocation more scientific and reasonable, the register allocation prediction model will also be used. Through this model, the dynamic utilization of register resources in the instruction scheduling process can be predicted, and the usage and idleness of register resources at different stages can be known in advance. On this basis, the minimum interference allocation strategy is used to carry out subsequent work. In the actual instruction scheduling process, according to the dynamic utilization information obtained by prediction, the register resources that are in an idle state are dynamically allocated for different instructions. The purpose of doing so is to minimize the interference with the registers that have been allocated and are being used, and to avoid frequent renaming of registers and large-scale data movement. Because frequent register renaming and data movement will consume additional time and computing resources, affecting the overall computing efficiency, this prediction-based allocation method with the goal of reducing interference can ensure that the allocation of register resources is more stable and orderly, so that each instruction can smoothly use the appropriate register resources for execution, maximize the efficiency of instruction scheduling and the execution of the entire computing task, and enable the entire computing process to proceed efficiently and stably.
[0147] Further optionally, after the target register allocation method is used to dynamically configure the corresponding register resources during the instruction scheduling process in the above steps, the register release prediction model can be used to predict the register resource usage time of each instruction during the instruction scheduling process. Then, based on the predicted usage time, the register release timing corresponding to each instruction is dynamically configured.
[0148] For example, after dynamically configuring register resources in the instruction scheduling process using the target register allocation method, in order to more efficiently manage register resources, the register release prediction model is further used. Through this model, the register resource usage time of each instruction in the instruction scheduling process is predicted, that is, how long each instruction will occupy the register is estimated in advance.
[0149] Based on the predicted usage time, the release timing of the registers corresponding to each instruction is dynamically configured. The advantage of this is that it can achieve precise control. When the instruction is executed and the registers are no longer needed, the corresponding registers can be released in time, and the registers will not be occupied due to inaccurate judgment. After all, register resources are very valuable. Releasing those registers that are no longer used in time can make them become available free resources as soon as possible, thereby improving the overall utilization of register resources and ensuring that there are enough register resources available for other subsequent instructions, so that the entire instruction scheduling and calculation process can continue more smoothly and efficiently.
[0150] The embodiment of the present application implements dynamic priority scheduling of model instructions through a calculation task graph, and further optimizes the scheduling strategy of model instructions based on the real-time monitoring of register pressure status, implements instruction scheduling optimization based on a large model, improves resource utilization efficiency, meets the computing requirements of the large model, and improves the overall model performance.
[0151] In another embodiment of the present application, a large model-based instruction optimization scheduling device is also provided. Figure 3 The device comprises the following units:
[0152] The conversion unit is configured to convert the target model to be processed into a computing task graph; wherein the computing task graph includes a plurality of nodes and connecting edges between the plurality of nodes; the plurality of nodes and connecting edges between the plurality of nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model;
[0153] A determination unit is configured to determine the scheduling priority of each instruction in the target model based on the computing task graph; the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction;
[0154] The execution unit is configured to execute instruction scheduling operations on the target model according to the register pressure state monitored in real time and the scheduling priority.
[0155] Further optionally, the conversion unit, for a target model to be processed, converts the target model into a computing task graph, and is configured to:
[0156] Identify each instruction included in the target model, and identify the topological relationship between each instruction as the dependency relationship between each instruction;
[0157] Constructing each instruction in the target model into a corresponding node, and constructing connection edges between each node according to the dependency relationship and / or instruction delay between each instruction;
[0158] Optimizing the dependency relationship according to the hardware device resources deployed by the target model;
[0159] Based on the optimized dependency relationship, the structural relationship between each node and each connecting edge is optimized to obtain the computing task graph.
[0160] Further optionally, the conversion unit optimizes the dependency relationship according to the hardware device resources deployed by the target model, and is configured to:
[0161] Identifying instructions to be optimized in the target model that are associated with hardware device resources; the hardware device resources include hardware computing units, and the hardware computing units are used to accelerate part or all of the computing tasks performed by the instructions to be optimized; the hardware device resources include at least a VLIW processor;
[0162] Acquire multiple computing operations corresponding to the instructions to be optimized;
[0163] According to the data access relationship of the hardware computing unit, the dependency relationship between multiple computing operations at the hardware resource level is analyzed; wherein the dependency relationship at the hardware resource level at least includes: input / output relationship, data storage relationship, reading order, and input / output format;
[0164] Based on the dependencies at the hardware resource level, the local dependencies corresponding to multiple computing operations are expanded to obtain optimized local dependencies;
[0165] The optimized local dependencies are integrated to obtain the optimized dependencies.
[0166] Further optionally, the computing task types supported by the hardware computing unit include at least one of the following: matrix multiplication operation, activation function calculation, vector operation; the hardware computing unit includes at least one of the following: a matrix multiplication calculation unit, an activation function calculation unit, and a vector calculation unit.
[0167] Further optionally, the computing task type supported by the hardware computing unit is a matrix multiplication operation; the conversion unit, based on the dependency relationship at the hardware resource level, expands the local dependency relationship corresponding to the multiple computing operations to obtain the optimized local dependency relationship, and is configured as follows:
[0168] Detecting at least one node related to the matrix multiplication operation;
[0169] Based on the data storage location corresponding to the matrix multiplication operation and the deployment information of the hardware resources, resetting the node position of the at least one node in the computing task graph; and / or,
[0170] Establish connection edges between the at least one node and other nodes in the computing task graph.
[0171] Further optionally, the computing task type supported by the hardware computing unit is a matrix multiplication operation; the conversion unit, based on the dependency relationship at the hardware resource level, expands the local dependency relationship corresponding to the multiple computing operations to obtain the optimized local dependency relationship, and is configured as follows:
[0172] Detecting at least one node related to the matrix multiplication operation;
[0173] Based on the data storage location corresponding to the matrix multiplication operation and the deployment information of the hardware resources, resetting the node position of the at least one node in the computing task graph; and / or,
[0174] Establish connection edges between the at least one node and other nodes in the computing task graph.
[0175] Further optionally, the determining unit determines the scheduling priority of each instruction in the target model based on the computing task graph, and is configured to:
[0176] Determine the instruction path corresponding to each instruction in the computing task graph;
[0177] Selecting a key instruction path from the computing task graph; the key instruction path includes at least one of the following: a core computing path in a long chain dependency relationship, an instruction path containing a preset computing operation, an instruction path whose computing complexity reaches a set condition, an instruction path whose data dependency depth reaches a set threshold, and an instruction path at a preset position in the target model;
[0178] Performing static analysis on the criticality and execution timeliness of the key instruction path to obtain the static priority of each instruction in the key instruction path;
[0179] The key instruction path is dynamically analyzed, and the static priority is adjusted based on the analysis result to obtain the scheduling priority of each instruction in the key instruction path.
[0180] Further optionally, the execution unit performs the instruction scheduling operation on the target model according to the register pressure state monitored in real time and the scheduling priority, and is configured to:
[0181] During the scheduling process, the target model is globally monitored to obtain register monitoring data of the target model; the register monitoring data at least includes: called register data, number of free registers, register allocation frequency, and register release frequency;
[0182] dynamically evaluating a register pressure state based on the register monitoring data;
[0183] According to the register pressure state, selecting a corresponding target register allocation mode from a pre-configured register allocation strategy;
[0184] An instruction scheduling operation is performed based on the scheduling priority, and the target register allocation method is used to dynamically configure corresponding register resources during the instruction scheduling process.
[0185] Further optionally, the execution unit dynamically evaluates the register pressure state based on the register monitoring data, and is configured to:
[0186] The ratio between the number of used registers and the total number of registers is taken as the register utilization rate;
[0187] Based on the register allocation frequency, register release frequency, and register utilization rate obtained through real-time monitoring, the register pressure state is determined in a preset manner.
[0188] Further optionally, the execution unit selects a corresponding target register allocation mode from a pre-configured register allocation strategy according to the register pressure state, and is configured to:
[0189] If the register pressure state is greater than a set pressure threshold, firstly scheduling a first instruction whose register resource usage is lower than the set scheduling threshold, and / or re-planning a scheduling timing of a second instruction whose register resource usage is higher than the set scheduling threshold;
[0190] If the register pressure state is less than the set pressure threshold, multiple execution instructions are scheduled in parallel and the use time of register resources is extended to improve processor throughput and reduce the vacancy rate of register resources.
[0191] Further optionally, the execution unit, after dynamically configuring the corresponding register resources in the instruction scheduling process using the target register allocation method, is further configured to:
[0192] Determine register resources to be allocated to each instruction during instruction scheduling;
[0193] A local optimization strategy is adopted to synchronously schedule multiple instructions using the same register during instruction scheduling to reduce the register switching frequency and improve the local utilization of the register.
[0194] Further optionally, the execution unit, after dynamically configuring the corresponding register resources in the instruction scheduling process using the target register allocation method, is further configured to:
[0195] The dynamic utilization of register resources in the instruction scheduling process is predicted through the register allocation prediction model;
[0196] The minimum interference allocation strategy is adopted to dynamically allocate vacant register resources to different instructions according to the predicted dynamic utilization during the instruction scheduling process to reduce the frequency of register renaming and data movement.
[0197] Further optionally, the execution unit, after dynamically configuring the corresponding register resources in the instruction scheduling process using the target register allocation method, is further configured to:
[0198] The register release prediction model is used to predict the register resource usage time of each instruction during instruction scheduling.
[0199] Based on the predicted usage duration, the register release timing corresponding to each instruction is dynamically configured.
[0200] The device can implement various steps in the above method embodiment, which will not be expanded here.
[0201] In an embodiment of the present application, an instruction optimization scheduling device based on a large model is adopted to implement dynamic priority scheduling of model instructions by calculating a task graph, and further optimize the scheduling strategy of model instructions based on a real-time monitored register pressure state, thereby implementing instruction scheduling optimization based on a large model, improving resource utilization efficiency, meeting the computing requirements of the large model, and improving overall model performance.
[0202] See also Figure 4 , Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of the present application. Figure 4 As shown, the embodiment of the present application provides an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, the following steps are implemented: First, for the target model to be processed, the target model is converted into a computing task graph. Among them, the computing task graph contains multiple nodes and connecting edges between multiple nodes; the multiple nodes and connecting edges between multiple nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model. Then, the scheduling priority of each instruction in the target model is determined based on the computing task graph. Among them, the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction. Finally, according to the real-time monitored register pressure state and the scheduling priority, the instruction scheduling operation of the target model is performed.
[0203] See also Figure 5 , Figure 5 A schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present application. Figure 5As shown, this embodiment provides a computer-readable storage medium 600, on which a computer program 611 is stored. When the computer program 611 is executed by a processor, the following steps are implemented: First, for the target model to be processed, the target model is converted into a computing task graph. Among them, the computing task graph contains multiple nodes and connecting edges between multiple nodes; the multiple nodes and connecting edges between multiple nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model. Then, the scheduling priority of each instruction in the target model is determined based on the computing task graph. Among them, the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction. Finally, according to the real-time monitored register pressure state and the scheduling priority, the instruction scheduling operation of the target model is performed.
[0204] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0205] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0206] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0207] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for optimizing instruction scheduling based on a large model, characterized in that: The method comprises: For a target model to be processed, the target model is converted into a computing task graph; wherein the computing task graph includes a plurality of nodes and connecting edges between the plurality of nodes; the plurality of nodes and connecting edges between the plurality of nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model; Determining the scheduling priority of each instruction in the target model based on the computing task graph; the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction; Execute an instruction scheduling operation on the target model according to the register pressure state monitored in real time and the scheduling priority; The determining the scheduling priority of each instruction in the target model based on the computing task graph includes: Determine the instruction path corresponding to each instruction in the computing task graph; Selecting a key instruction path from the computing task graph; the key instruction path includes at least one of the following: a core computing path in a long chain dependency relationship, an instruction path containing a preset computing operation, an instruction path whose computing complexity reaches a set condition, an instruction path whose data dependency depth reaches a set threshold, and an instruction path at a preset position in the target model; Performing static analysis on the criticality and execution timeliness of the key instruction path to obtain the static priority of each instruction in the key instruction path; Dynamically analyzing the key instruction path, and adjusting the static priority based on the analysis result to obtain the scheduling priority of each instruction in the key instruction path; The executing the instruction scheduling operation on the target model according to the real-time monitored register pressure state and the scheduling priority includes: During the scheduling process, the target model is globally monitored to obtain register monitoring data of the target model; the register monitoring data at least includes: called register data, number of free registers, register allocation frequency, and register release frequency; dynamically evaluating a register pressure state based on the register monitoring data; According to the register pressure state, selecting a corresponding target register allocation mode from a pre-configured register allocation strategy; An instruction scheduling operation is performed based on the scheduling priority, and the target register allocation method is used to dynamically configure corresponding register resources during the instruction scheduling process.
2. The instruction optimization scheduling method based on a large model according to claim 1 is characterized in that: For the target model to be processed, converting the target model into a computing task graph includes: Identify each instruction included in the target model, and identify the topological relationship between each instruction as the dependency relationship between each instruction; Constructing each instruction in the target model into a corresponding node, and constructing connection edges between each node according to the dependency relationship and / or instruction delay between each instruction; Optimizing the dependency relationship according to the hardware device resources deployed by the target model; Based on the optimized dependency relationship, the structural relationship between each node and each connecting edge is optimized to obtain the computing task graph.
3. The instruction optimization scheduling method based on a large model according to claim 2 is characterized in that: Optimizing the dependency relationship according to the hardware device resources deployed by the target model includes: Identifying instructions to be optimized in the target model that are associated with hardware device resources; the hardware device resources include hardware computing units, and the hardware computing units are used to accelerate part or all of the computing tasks performed by the instructions to be optimized; the hardware device resources at least include a very long instruction word VLIW processor; Acquire multiple computing operations corresponding to the instructions to be optimized; According to the data access relationship of the hardware computing unit, the dependency relationship between multiple computing operations at the hardware resource level is analyzed; wherein the dependency relationship at the hardware resource level at least includes: input / output relationship, data storage relationship, reading order, and input / output format; Based on the dependencies at the hardware resource level, the local dependencies corresponding to multiple computing operations are expanded to obtain optimized local dependencies; The optimized local dependencies are integrated to obtain the optimized dependencies.
4. The instruction optimization scheduling method based on a large model according to claim 3 is characterized in that: The computing task types supported by the hardware computing unit include at least one of the following: matrix multiplication operation, activation function calculation, vector operation; The hardware computing unit includes at least one of the following: a matrix multiplication computing unit, an activation function computing unit, and a vector computing unit.
5. The instruction optimization scheduling method based on a large model according to claim 4 is characterized in that: The computing task type supported by the hardware computing unit is matrix multiplication operation; The dependency relationship based on the hardware resource level is used to expand the local dependency relationship corresponding to the multiple computing operations to obtain the optimized local dependency relationship, including: Detecting at least one node related to the matrix multiplication operation; Based on the data storage location corresponding to the matrix multiplication operation and the deployment information of the hardware resources, resetting the node position of the at least one node in the computing task graph; and / or, Establish connection edges between the at least one node and other nodes in the computing task graph.
6. The instruction optimization scheduling method based on a large model according to claim 1 is characterized in that: The dynamically evaluating the register pressure state based on the register monitoring data includes: The ratio between the number of used registers and the total number of registers is taken as the register utilization rate; Based on the register allocation frequency, register release frequency, and register utilization rate obtained through real-time monitoring, the register pressure state is determined in a preset manner.
7. The instruction optimization scheduling method based on a large model according to claim 1 is characterized in that: The selecting a corresponding target register allocation mode from a pre-configured register allocation strategy according to the register pressure state includes: If the register pressure state is greater than a set pressure threshold, preferentially scheduling a first instruction whose register resource usage is lower than the set scheduling threshold, and / or re-planning a scheduling timing of a second instruction whose register resource usage is higher than the set scheduling threshold; If the register pressure state is less than the set pressure threshold, multiple execution instructions are scheduled in parallel and the use time of register resources is extended to improve processor throughput and reduce the vacancy rate of register resources.
8. The instruction optimization scheduling method based on a large model according to claim 1 is characterized in that: After dynamically configuring the corresponding register resources in the instruction scheduling process by using the target register allocation method, the method further includes: Determine register resources to be allocated to each instruction during instruction scheduling; A local optimization strategy is adopted to synchronously schedule multiple instructions using the same register during instruction scheduling to reduce the register switching frequency and improve the local utilization of the register.
9. The instruction optimization scheduling method based on a large model according to claim 1 is characterized in that: After dynamically configuring the corresponding register resources in the instruction scheduling process by using the target register allocation method, the method further includes: The dynamic utilization of register resources in the instruction scheduling process is predicted through the register allocation prediction model; The minimum interference allocation strategy is adopted to dynamically allocate vacant register resources to different instructions according to the predicted dynamic utilization during the instruction scheduling process to reduce the frequency of register renaming and data movement.
10. The instruction optimization scheduling method based on a large model according to claim 1, characterized in that: After dynamically configuring the corresponding register resources in the instruction scheduling process by using the target register allocation method, the method further includes: The register release prediction model is used to predict the register resource usage time of each instruction during instruction scheduling. Based on the predicted usage duration, the register release timing corresponding to each instruction is dynamically configured.
11. An instruction optimization scheduling device based on a large model, characterized in that: The device comprises the following units, wherein: The conversion unit is configured to convert the target model to be processed into a computing task graph; wherein the computing task graph includes a plurality of nodes and connecting edges between the plurality of nodes; the plurality of nodes and connecting edges between the plurality of nodes in the computing task graph are arranged into corresponding multiple levels based on the instruction scheduling operations corresponding to each computing task in the target model; A determination unit is configured to determine the scheduling priority of each instruction in the target model based on the computing task graph; the scheduling priority is associated with the importance and timeliness of the instruction path of each instruction in the computing task graph; the scheduling priority includes the execution order of each instruction and the register resource configuration of each instruction; An execution unit configured to execute an instruction scheduling operation on the target model according to the register pressure state monitored in real time and the scheduling priority; The determining unit determines the scheduling priority of each instruction in the target model based on the computing task graph, and is specifically configured as follows: Determine the instruction path corresponding to each instruction in the computing task graph; select a key instruction path from the computing task graph; the key instruction path includes at least one of the following: a core computing path in a long chain dependency relationship, an instruction path containing a preset computing operation, an instruction path whose computing complexity reaches a set condition, an instruction path whose data dependency depth reaches a set threshold, and an instruction path at a preset position in the target model; statically analyze the criticality and execution timeliness of the key instruction path to obtain the static priority of each instruction in the key instruction path; dynamically analyze the key instruction path, and adjust the static priority based on the analysis result to obtain the scheduling priority of each instruction in the key instruction path; The execution unit performs instruction scheduling operations on the target model according to the register pressure state and the scheduling priority monitored in real time, and is specifically configured as follows: During the scheduling process, the target model is globally monitored to obtain register monitoring data of the target model; the register monitoring data includes at least: called register data, the number of free registers, register allocation frequency, and register release frequency; the register pressure state is dynamically evaluated based on the register monitoring data; according to the register pressure state, a corresponding target register allocation method is selected from a pre-configured register allocation strategy; an instruction scheduling operation is performed based on the scheduling priority, and the corresponding register resources are dynamically configured during the instruction scheduling process using the target register allocation method.
12. An electronic device, characterized in that: include: Memory for storing computer software programs; A processor is used to read and execute the computer software program, thereby implementing the large model-based instruction optimization scheduling method described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that: The storage medium stores a computer software program, which, when executed by a processor, implements the large model-based instruction optimization scheduling method as described in any one of claims 1 to 10.
14. A chip, characterized in that: The chip is loaded with a computer software program and / or a hardware unit, and the computer software program and / or the hardware unit are used to implement the large model-based instruction optimization scheduling method as described in any one of claims 1-10.
Citation Information
Patent Citations
Method and device for time-sensitive networking coordinated transfer learning in industrial environment
CN115269134A
Task scheduling method and device, computing system, electronic equipment and storage medium
CN116680063A
Cited By
Intelligent scheduling method and system based on artificial intelligence large model
CN122173229A