Instruction packaging method and device, equipment and storage medium
By recording the transmit clock cycle during the instruction scheduling process and sorting the instructions by the clock cycle, the problem of limited hardware performance improvement in the prior art is solved, and more efficient instruction packet allocation and parallelism improvement are achieved.
Patent Information
- Application Number
- CN202410156947.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2025-08-05
AI Technical Summary
In the ultra-long instruction font architecture, the existing technology fails to effectively improve hardware performance when implementing instruction parallelism, and the existing packaging methods fail to reasonably allocate instruction packages, resulting in limited performance improvement.
By recording and using the transmit clock cycle during the instruction scheduling process, ordering and packaging instructions according to the clock cycle, multiple instruction packets are formed and sent to the hardware execution unit in parallel, avoiding re-analysis of instruction dependencies and hardware resource conflicts.
It improves hardware performance, reduces packaging time, reasonably allocates instruction packages, improves instruction parallelism, and reduces the total execution time.
Smart Images

Figure CN120429012A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of compilers, and particularly to an instruction packing method, device, equipment and storage medium. Background Art
[0002] In fields such as AI (Artificial Intelligence) chips and DSP (Digital Signal Process) chips, the very long instruction word architecture (VLIW) is usually adopted as the chip architecture. The very long instruction word architecture executes an instruction packet in each clock cycle. A single instruction packet is allowed to contain multiple instructions, and the very long instruction word architecture executes multiple instructions in parallel.
[0003] The very long instruction word architecture requires the compiler to pack multiple instructions that can be issued in parallel into one instruction packet. However, when actually implementing instruction parallelism, it is not the case that the higher the performance can be achieved by stuffing more instructions into each instruction packet. All instructions in the same instruction packet need to meet the issue conditions before the instruction packet is issued. Summary of the Invention
[0004] This application provides an instruction packing method, device, equipment and storage medium. The above method provides a new instruction packing method, which can improve the performance of the hardware. The technical solution includes the following contents.
[0005] According to one aspect of this application, an instruction packing method is provided. The method includes the following steps.
[0006] During the execution of instruction scheduling, record the issue clock cycles of the first number of instructions respectively. The instruction scheduling is used to sort the first number of instructions according to the values of the issue clock cycles;
[0007] Pack the instructions with the same issue clock cycle among the first number of instructions to obtain one or more instruction packets. Each instruction packet in the multiple instruction packets is used to be sent to the hardware execution unit according to the corresponding issue clock cycle.
[0008] According to another aspect of this application, an instruction packing device is provided. The device includes the following modules.
[0009] A saving module, configured to record the issue clock cycles of the first number of instructions respectively during the execution of instruction scheduling. The instruction scheduling is used to sort the first number of instructions according to the values of the issue clock cycles;
[0010] A packing module, configured to pack instructions with the same emission clock cycle among the first quantity of instructions, to obtain one or more instruction packets, and each instruction packet in the plurality of instruction packets is configured to be sent to a hardware execution unit according to a corresponding emission clock cycle.
[0011] In an optional embodiment, the saving module is further configured to record the correspondence between the first quantity of instructions and the second quantity of emission clock cycles;
[0012] Wherein, one instruction in the correspondence corresponds to one emission clock cycle, and one emission clock cycle in the correspondence corresponds to one instruction, zero instructions or at least two instructions.
[0013] In an optional embodiment, the packing module is further configured to traverse the second quantity of emission clock cycles one by one, and pack instructions with the same emission clock cycle among the first quantity of instructions, to obtain the plurality of instruction packets.
[0014] In an optional embodiment, the second quantity of emission clock cycles includes a plurality of consecutive emission clock cycles starting from zero; the packing module is further configured to create a first variable; initialize the first variable to zero;
[0015] When the second quantity of emission clock cycles has not been traversed completely and the first emission clock cycle corresponds to an instruction, create a new instruction packet; put the instruction corresponding to the first emission clock cycle into the newly created instruction packet, and the newly created instruction packet is configured to be sent to the hardware execution unit according to the first emission clock cycle, and the first emission clock cycle is the emission clock cycle indicated by the value of the first variable; increment the value of the first variable by one, and continue to traverse the second quantity of emission clock cycles;
[0016] When the second quantity of emission clock cycles has not been traversed completely and the first emission clock cycle does not correspond to an instruction, increment the value of the first variable by one, and continue to traverse the second quantity of emission clock cycles.
[0017] In an optional embodiment, the packing module is further configured to traverse the first quantity of instructions one by one, and pack instructions with the same emission clock cycle among the first quantity of instructions, to obtain the plurality of instruction packets.
[0018] In an optional embodiment, the instruction scheduling is configured to sort the first quantity of instructions in ascending order according to the emission clock cycle, to obtain an instruction sequence, and the emission clock cycle corresponding to the first instruction in the instruction sequence is zero. The packing module is further configured to create a second variable; initialize the second variable to zero;
[0019] When not all the instructions in the instruction sequence are completely packed, obtain the i-th instruction in the instruction sequence;
[0020] When the value of the i-th emission clock cycle is not greater than the value of the second variable, put the i-th instruction into the current instruction packet, where the i-th emission clock cycle is the emission clock cycle corresponding to the i-th instruction; increment i by one, and the initial value of i is one; re-determine whether all the instructions in the instruction sequence have been completely packed;
[0021] When the value of the i-th emission clock cycle is greater than the value of the second variable and there are instructions in the current instruction packet, determine to emit the current instruction packet in the i-th emission clock cycle; create a new instruction packet, increment the value of the second variable by one; re-determine the magnitude relationship between the value of the i-th emission clock cycle and the value of the second variable.
[0022] In an optional embodiment, the packing module is further configured to insert a plurality of delimiters in the instruction sequence, where the delimiters are used to separate instructions with different emission clock cycles; the instruction sequence is a sequence obtained by sorting the first quantity of instructions according to the values of the emission clock cycles through the instruction scheduling;
[0023] Identify the plurality of delimiters in the instruction sequence, and pack the instructions between two adjacent delimiters to obtain the plurality of instruction packets.
[0024] In an optional embodiment, the second quantity of emission clock cycles includes a plurality of consecutive emission clock cycles starting from zero. The packing module is further configured to create a third variable; initialize the third variable to zero;
[0025] When the second quantity of emission clock cycles has not been traversed completely and there is an instruction corresponding to the third emission clock cycle, insert a pseudo-instruction before the first instruction corresponding to the third emission clock cycle; increment the value of the third variable by one, where the third emission clock cycle is the emission clock cycle corresponding to the value of the third variable; continue to traverse the second quantity of emission clock cycles;
[0026] When the second quantity of emission clock cycles has not been traversed completely and there is no instruction corresponding to the third emission clock cycle, increment the value of the third variable by one; continue to traverse the second quantity of emission clock cycles.
[0027] In an optional embodiment, the delimiter is a pseudo-instruction. The packing module is further configured to sequentially check all the instructions in the instruction sequence;
[0028] During the inspection process, if the inspected instruction is the pseudo-instruction, a new instruction packet is created;
[0029] If the inspected instruction is not the pseudo-instruction, the inspected instruction is placed into the newly created instruction packet.
[0030] According to one aspect of the present application, a computer device is provided. The computer device includes: a processor and a memory. The memory stores a computer program, and the computer program is loaded and executed by the processor to implement the instruction packaging method as described above.
[0031] According to another aspect of the present application, a computer-readable storage medium is provided. The storage medium stores a computer program, and the computer program is loaded and executed by the processor to implement the instruction packaging method as described above.
[0032] According to another aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above instruction packaging method.
[0033] The beneficial effects brought by the technical solutions provided by the embodiments of the present application at least include the following.
[0034] By saving the respective issue clock cycles of the first number of instructions during instruction scheduling, and packing the instructions with the same issue clock cycle into the same packet during instruction packaging to obtain one or more instruction packets. Compared with the related art, the packaging method provided by the present application no longer needs to re-analyze and compare information such as instruction dependencies and hardware resources during packaging, which speeds up the packaging speed. Moreover, the packaging method of the present application is no longer limited to the greedy method (related art), and can better distribute the instructions in different instruction packets by virtue of the results obtained during instruction scheduling. The packaging method of the present application generally brings performance improvement. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a schematic diagram of the processing process of an AI compiler provided by the related art;
[0037] Figure 2It is a schematic diagram of the pseudo-code of the table scheduling algorithm provided by the related art;
[0038] Figure 3 It is a schematic diagram of the modulo scheduling algorithm provided by the related art;
[0039] Figure 4 It is a flowchart of the modulo scheduling algorithm provided by the related art;
[0040] Figure 5 It is a flowchart of the instruction packing method provided by the related art;
[0041] Figure 6 It is a schematic diagram of the principle of the instruction packing method provided by an embodiment of the present application;
[0042] Figure 7 It is a flowchart of the instruction packing method provided by an embodiment of the present application;
[0043] Figure 8 It is a schematic diagram of the relationship between instructions, hardware resources, and time resources provided by an embodiment of the present application;
[0044] Figure 9 It is a schematic diagram of the instruction packing result obtained by using the related art;
[0045] Figure 10 It is a schematic diagram of the instruction scheduling result provided by an embodiment of the present application;
[0046] Figure 11 It is a schematic diagram of the instruction packing result obtained by using the instruction packing method of the present application;
[0047] Figure 12 It is a flowchart of the instruction packing method provided by another embodiment of the present application;
[0048] Figure 13 It is a flowchart of the specific instruction packing method provided by an embodiment of the present application;
[0049] Figure 14 It is a flowchart of the instruction packing method provided by another embodiment of the present application;
[0050] Figure 15 It is a flowchart of the specific instruction packing method provided by another embodiment of the present application;
[0051] Figure 16 It is a flowchart of the instruction packing method provided by another embodiment of the present application;
[0052] Figure 17 It is a flowchart of the method for inserting delimiters provided by an embodiment of the present application;
[0053] Figure 18It is a flowchart of a method for packing based on delimiters provided by an embodiment of the present application;
[0054] Figure 19 It is a flowchart of a method for instruction scheduling provided by an embodiment of the present application;
[0055] Figure 20 It is a structural block diagram of an instruction packing device provided by an embodiment of the present application;
[0056] Figure 21 It is a structural block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0057] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0058] First, briefly introduce the nouns involved in the embodiments of the present application.
[0059] In fields such as AI chips and DSP chips, the VLIW architecture is usually adopted as the chip architecture. In order to improve the parallelism of hardware execution, the VLIW architecture usually requires the compiler to pack multiple instructions that can be parallelly issued into an instruction bundle. In the present application, the instruction packing method will be mainly introduced with the AI compiler supporting the AI chip, and the instruction packing method of the present application can also be implemented for compilers such as DSP chips adopting the VLIW architecture.
[0060] AI compiler: The AI compiler usually adopts a classic front-end and back-end structure to connect the model generated by the machine learning framework with the underlying chip. The overall structure of the AI compiler is as Figure 1 shown.
[0061] In Figure 1 , the AI compiler reads the machine learning model 101 generated from machine learning frameworks such as TensorFlow and Pytorch. The AI compiler first performs model parsing 102 on the model 101 in the front end, parsing it into a high-level IR (Intermediate Representation, intermediate expression) (usually a computation graph), and then performs target-independent optimization (optimization independent of the target hardware) 103, such as arithmetic simplification, operator fusion, etc., and outputs the optimized high-level IR. Next, it enters the target-dependent optimization 104 in the back end, such as dedicated instruction mapping, memory allocation, memory access latency hiding, etc., and outputs the device-dependent low-level IR; finally, it enters the code generation stage 105 in the back end to generate the target program 106 that can run on the AI chip.
[0062] The instruction packing method provided by this application is in the code generation stage 105 of the AI compiler.
[0063] In the code generation stage of the AI compiler, the obtained high-level computer language code is divided into multiple basic blocks. A basic block is a linear piece of code. For a basic block, the compiler performs instruction scheduling on the code in the basic block to obtain an instruction sequence. Packing operations are performed on the instructions in the instruction sequence to obtain multiple instruction packets, and then each instruction packet is converted into a continuous segment of binary data with a fixed or variable length, and the converted instruction packets are sent to the hardware (AI chip) for execution. The AI chip processes one instruction packet in parallel in one clock cycle.
[0064] Instruction scheduling: Instruction scheduling can reorder the instruction issuance order, improve the parallelism of instruction execution, and reduce the total duration required for operator execution. Instruction scheduling is subject to various constraints when selecting instructions, such as instruction dependency constraints, hardware resource constraints, register constraints, etc. The main goal of instruction scheduling is to minimize pipeline stalls and seek the optimal solution under these existing constraints.
[0065] Currently, the mainstream instruction scheduling optimization algorithms in the industry include the list scheduling algorithm and the modulo scheduling algorithm.
[0066] The list scheduling algorithm is a way to adjust the instruction order based on the greedy algorithm and heuristic algorithm. This algorithm uses the basic block as the basic interval, and within this range, continuously selects and adjusts the instruction order according to conditions such as instruction dependencies and hardware resource occupancy. Figure 2 The pseudo-code of the list scheduling algorithm provided by the related technology is shown.
[0067] In Figure 2 , Op represents an instruction, D represents the directed acyclic graph of instruction dependencies, D = (N, E), N is the node (Node), representing an instruction; E is the edge, representing the minimum interval between two nodes. The variable Cycle is the simulated hardware clock cycle, initialized to 1, and will increment during the code execution. Whenever Cycle + 1, all instructions that have been completed in the previous Cycle are deleted from the Active list, and each successor node of these instructions in D is checked, and the available instructions are moved into the Ready list. The Ready list contains all instructions that can be used in the current Cycle without affecting the correctness of continuous execution, and is initialized to all children nodes of the parent node in D. The Active list contains all instructions that can actually be issued currently.
[0068] The modulo scheduling algorithm is usually also known as software pipelining. In the code generation process of the AI compiler backend, software pipelining is an important optimization stage. As the basic computational units of AI models, operators have a large number of loop instruction structures, which consume the main chip execution time. Software pipelining optimization is a type of compilation algorithm specifically used to optimize loop instruction structures. The current mainstream software pipelining optimization algorithm in the industry is the modulo scheduling algorithm. Its main algorithm idea is as Figure 3 shown.
[0069] In Figure 3 , assume that the total number of loops of a loop instruction structure is N times. The instruction sequence in its core loop takes T beats for each loop, and T can be equally divided into 3 segments. In Figure 3 's part (A), each loop starts to execute after the previous loop execution is completed. Each loop is represented by I(n), where n ∈ [0, 1, 2,..., N - 2, N - 1]. For one loop, each equally divided segment of code is represented by S(n), Figure 3 each segment of code in is divided into 3 segments, that is, n ∈ [0, 1, 2]. The code after modulo scheduling is as shown in Figure 3 's part (B), where the instructions of the I(n + 1) loop can start to execute in advance without waiting for the I(n) loop to complete. The time difference between the start times of I(n) and I(n + 1) is called the initiation interval II. II is equal to the number of beats occupied by a segment of code in one loop, Figure 3 in part (B) of, II = T / 3. Analyzing the folded loop, it can be found that there is a stable code structure, as shown in Figure 3 's part (C), Figure 3 part (C) of is the code structure after modulo scheduling. It includes three stages, namely the filling stage, including the start part of the code in the first two loops; the core stage, including part of the code in three consecutive loops; and the emptying stage, including the end part of the code in the last two loops.
[0070] For a specific instruction in the loop, assume that in a certain loop, it is at the Kth beat relative to the emission time of the first instruction in this loop. Then its emission time in the folded core stage is T = K mod II, which is exactly the origin of the name of the modulo scheduling. For example, assume that before folding, instruction A is at the 16th beat relative to the emission time of the first instruction in this loop, and II is 10 beats. Then the folded instruction A is emitted at the 6th beat, and the execution time of the core loop is also compressed to 10 beats.
[0071] The basic algorithm flow of modulo scheduling is as shown in Figure 4 shown.
[0072] Step 401, start.
[0073] Step 402, construct a directed acyclic graph. Construct a DAG (directed acyclic graph) according to the input instructions, and the DAG is used to represent the dependency relationship between instructions.
[0074] Step 403, calculate scheduling parameters. For example, the initial value and maximum value of the startup interval, the depth, height, etc. of the nodes in the DAG graph.
[0075] Step 404, sort the instructions. Adjust the scheduling order of the instructions according to the node-related parameters calculated in Step 403.
[0076] Step 405, traverse the startup interval in sequence. Starting from the initial startup interval, sequentially try whether a reasonable scheduling result can be searched until the startup interval is equal to the maximum value.
[0077] Step 406, use the search algorithm to traverse the instructions and calculate the scheduling result. Search in order whether each instruction can be scheduled successfully, that is, there are no resource conflicts and unreasonable dependency relationships between this instruction and the already scheduled instructions.
[0078] Step 407, determine whether to stop the search. Determine whether to stop the search. If the search is stopped, execute Step 408; if the scheduling is not successful, return to execute Step 405.
[0079] Step 408, determine whether the scheduling is successful. If the scheduling is successful, execute Step 409; if the scheduling is not successful, execute Step 410.
[0080] Step 409, adjust the code. If the scheduling is successful, generate a complete code structure according to the scheduling result, including the filling stage, the core stage, and the emptying stage.
[0081] Step 410, end. The scheduling fails, and the original code remains unchanged.
[0082] Through the introduction of the above table scheduling algorithm and modulo scheduling algorithm, it can be found that different instruction scheduling algorithms will display the timing information of instruction emission, and their purpose is to sort the instructions by combining information such as instruction dependencies and hardware resources known to the compiler. This timing information is usually represented by Cycle (clock cycle). During this sorting, different instruction scheduling algorithms will substitute a timing information to describe the emission time of each instruction and the resource information required at each time point, etc.
[0083] In this application, it is hoped to use this timing information to assist in instruction packing, associate instruction packing with instruction scheduling, and thus obtain better instruction packing results. This timing information is the emission clock cycle.
[0084] To facilitate the description of the performance improvement brought by the packaging method provided in this application below, the packaging method provided by the related art is introduced here again. The related art adopts a greedy scheme to pack as many instructions together as possible. The related art is as shown in Figure 5 shown.
[0085] Step 501, start.
[0086] Step 502, create an instruction packet. Create an empty instruction packet and emit the instructions in the basic block in the order after instruction scheduling.
[0087] Step 503, determine whether all instructions have been packed. If there are still instructions not packed, execute Step 504; if all instructions have been packed, execute Step 509.
[0088] Step 504, select an instruction in order.
[0089] Step 505, there is a conflict between the instruction and any instruction in the packet. Compare the current instruction with all the instructions in the current instruction packet to check whether there are register dependency conflicts and hardware resource conflicts between the instruction and any instruction in the packet. If there are other restrictions, check them together. If there are no conflicts, execute Step 508; if there are any conflicts, execute Step 506.
[0090] Step 506, emit the current instruction packet.
[0091] Step 507, create a new instruction packet.
[0092] Step 508, add the instruction to the instruction packet.
[0093] Step 509, end.
[0094] Repeat the above steps until all instructions are emitted.
[0095] It can be found that the packaging method provided by the related art only focuses on information such as instruction dependencies and instruction resources, and will pack instructions together as much as possible. Although the related art can reduce the overall number of instruction packets, it may not necessarily obtain better results in terms of performance.
[0096] Next, the instruction packaging method provided in this application will be introduced.
[0097] Figure 6The figure shows a schematic diagram of the principle of the instruction packing method provided by an exemplary embodiment of the present application. The instruction packing method provided by the embodiment of the present application is executed by the terminal device 601, and the terminal device 601 can be a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc. Optionally, the instruction packing method provided by the embodiment of the present application is executed by a compiler corresponding to a chip adopting the VLIW architecture, such as a compiler corresponding to an AI chip, a compiler corresponding to a DSP chip, etc. The chip adopting the VLIW architecture will execute an instruction packet in each clock cycle, and a plurality of instructions are allowed to be included in one instruction packet, and the chip adopting the VLIW architecture executes multiple instructions in parallel.
[0098] In the instruction packing method provided by the present application, during the execution of instruction scheduling, the emission clock cycles of the first number of instructions are recorded respectively. Instruction scheduling is used to sort the first number of instructions according to the values of the emission clock cycles. For the specific introduction of instruction scheduling, please refer to the above text.
[0099] During the execution of instruction packing, the instructions with the same emission clock cycle among the first number of instructions are packed to obtain one or more instruction packets, and each instruction packet in the plurality of instruction packets is used to be sent to the hardware execution unit according to the corresponding emission clock cycle.
[0100] Combined with reference to Figure 6 , specifically, in the instruction scheduling stage 602, the compiler saves the correspondence 604 between the first number of instructions and the second number of emission clock cycles. In the correspondence 604, one instruction corresponds to one emission clock cycle, and one emission clock cycle corresponds to one instruction, zero instructions or at least two instructions. Schematically, the correspondence 604 is shown in the following table:
[0101] Emission clock cycle 0 Emission clock cycle 1 Emission clock cycle 2 Instruction 1, Instruction 2 Instruction 3
[0102] In the table, emission clock cycle 0 corresponds to instruction 1 and instruction 2, emission clock cycle 1 does not correspond to any instruction, and emission clock cycle 2 corresponds to instruction 3.
[0103] In Figure 6 , in the instruction packing stage 603, the compiler will execute step 605: pack the instructions with the same emission clock cycle among the first number of instructions. Taking the above table as an example, the compiler will construct two instruction packets, namely instruction packet 1 and instruction packet 2. Instruction packet 1 contains instructions 1 and 2, and instruction packet 2 only contains instruction 3. After packing, one or more instruction packets will be obtained, and each instruction packet will be sent to the hardware execution unit according to the corresponding emission clock cycle. For example, instruction packet 1 is sent to the AI chip at emission clock cycle 0. The AI chip includes multiple hardware resources. Instruction 1 is sent to hardware resource b, and instruction 2 is sent to hardware resource c, so as to achieve parallel computing of multiple instructions under the VLIW architecture.
[0104] In summary, the packing method provided by this application directly utilizes the emission clock cycles recorded during the instruction scheduling process, and instructions with the same emission clock cycle are packed into the same instruction packet. This is beneficial to improving the performance of hardware processing, and the specific effect reasoning will be introduced in the following text.
[0105] Figure 7 The flowchart of the instruction packing method provided by an exemplary embodiment of this application is shown. Taking the execution of this method by the Figure 6 compiler in as an example, this method includes:
[0106] Step 720, during the execution of instruction scheduling, record the emission clock cycle of each of the first number of instructions, and the instruction scheduling is used to sort the first number of instructions according to the value of the emission clock cycle;
[0107] Instruction scheduling is used to sort the first number of instructions in ascending order of the emission clock cycle. For example, the instruction sequence after instruction scheduling is Instruction 1 (cycle 0), Instruction 2 (cycle 0), Instruction 3 (cycle 1), Instruction 4 (cycle 2), Instruction 5 (cycle 2), Instruction 6 (cycle 3), where cycle refers to the emission clock cycle.
[0108] The emission clock cycle is the timing information used for sorting in the instruction scheduling algorithm provided by the related technology. Referring to Figure 2 for combination reference, Figure 2 shows a table scheduling algorithm, Figure 2 the cycle in is the emission clock cycle, Figure 2 during the execution of the table scheduling algorithm in, cycle will continuously accumulate. For the detailed introduction of Figure 2 please refer to the above text. Figure 3 and Figure 4 Combined shows a schematic diagram of a modulo scheduling algorithm. Whether it is a table scheduling algorithm, a modulo scheduling algorithm, or other instruction scheduling algorithms, a timing information will be substituted during the sorting process to describe the emission time of each instruction. This timing information is the emission clock cycle utilized by the instruction packing method of this application.
[0109] In one embodiment, record the correspondence between the first number of instructions and the second number of emission clock cycles; wherein, one instruction in the correspondence corresponds to one emission clock cycle, and one emission clock cycle in the correspondence corresponds to one instruction, zero instructions or at least two instructions.
[0110] Schematically, the correspondence is shown in the following table:
[0111]
[0112]
[0113] In the above table, the emission clock cycle 0 corresponds to Instruction 1 and Instruction 2, the emission clock cycle 1 does not correspond to any instruction, and the emission clock cycle 2 corresponds to Instruction 3.
[0114] Step 740: Pack the instructions with the same emission clock cycle among the first quantity of instructions to obtain one or more instruction packets, and each instruction packet in the multiple instruction packets is used to be sent to the hardware execution unit according to the corresponding emission clock cycle.
[0115] Taking the above table as an example, the compiler will construct two instruction packets, corresponding to Instruction Packet 1 and Instruction Packet 2 respectively. Instruction Packet 1 contains Instruction 1 and Instruction 2, and Instruction Packet 2 only contains Instruction 3. After packing, Instruction Packet 1 and Instruction Packet 2 will be obtained, and each instruction packet will be sent to the hardware execution unit according to the corresponding emission clock cycle. For example, Instruction Packet 1 is sent to the AI chip at the emission clock cycle 0. The AI chip includes multiple hardware resources. Instruction 1 is sent to hardware resource b, and Instruction 2 is sent to hardware resource c, so as to implement the parallel computing of multiple instructions under the VLIW architecture.
[0116] In summary, by saving the respective emission clock cycles of the first quantity of instructions during instruction scheduling and packing the instructions with the same emission clock cycle into the same packet during instruction packing to obtain one or more instruction packets, compared with the related art, the packing method provided by this application no longer needs to re-analyze and compare information such as instruction dependencies and hardware resources during packing, thus accelerating the packing speed. Moreover, the packing method of this application is not limited to the greedy method (related art), and can better distribute the instructions in different instruction packets by leveraging the results obtained in the instruction scheduling stage. Generally, the packing method of this application can bring performance improvement.
[0117] Next, the beneficial effects of the instruction packing method provided by this application are theoretically deduced.
[0118] Figure 8 It shows the hardware resources and time resources corresponding to three instructions respectively. Figure 8 In, Instruction A needs to be executed in two clock cycles. Moreover, Instruction A needs to use hardware resource Ob at clock cycle 0 and hardware resource Od at clock cycle 1. Instruction B needs to be executed in two clock cycles. Moreover, Instruction B needs to use hardware resource Oc at clock cycle 0 and hardware resource Ob at clock cycle 1. Instruction C needs to be executed in three clock cycles. Moreover, Instruction C needs to use hardware resource Oa at clock cycle 0, hardware resource Oc at clock cycle 1, and hardware resource Ob at clock cycle 2.
[0119] Suppose there is an instruction sequence where the instructions in the sequence have no dependencies, and the emission order of the instructions in the sequence is ABCA. It should be noted that the two As in the instruction sequence ABCA are two substantially identical instructions. For example, the first A is an "and" instruction, and the second A is also an "and" instruction. The hardware resources and the number of clock cycles required by the two As are the same.
[0120] The instruction packing result obtained by using the related technology is as Figure 9 shown.
[0121] The related technology adopts a greedy algorithm to pack as many instructions as possible into one instruction packet. Figure 9 In it, when the clock cycle 0 arrives, since there are no conflicts among the instructions ABC, they are packed together and sent to the hardware resources. Instruction A is executed using hardware resource 0b at clock cycle 0 and using hardware resource 0d at clock cycle 1. Instruction B is executed using hardware resource 0c at clock cycle 0 and using hardware resource 0b at clock cycle 1. Instruction C is executed using hardware resource 0a at clock cycle 0, using hardware resource 0c at clock cycle 1, and using hardware resource 0b at clock cycle 2.
[0122] The remaining instruction A has travel conflicts at clock cycles 0, 1, and 2 (hardware resource 0b is occupied in each clock cycle) and is not emitted until clock cycle 3 arrives. The remaining instruction A also takes two clock cycles to execute.
[0123] By observing Figure 9 , it can be found that when using the related method to pack the instructions ABCA, the total execution time is 5 clock cycles.
[0124] When using the instruction packing method provided by this application, the correspondence between the first quantity of instructions and the second quantity of emission clock cycles will be saved during instruction scheduling. Figure 10 The correspondence is shown. When the clock cycle 0 arrives, an instruction packet containing instructions A and B will be emitted; when the clock cycle 1 arrives, an instruction packet containing instruction C will be emitted; when the clock cycle 2 arrives, an instruction packet containing another instruction A will be emitted.
[0125] The result of using the instruction packing method provided by this application is as Figure 11 shown.
[0126] According to Figure 10 the scheduling result shown, when the clock cycle 0 arrives, emit instructions A and B; when the clock cycle 1 arrives, emit instruction C; when the clock cycle 2 arrives, emit another instruction A.
[0127] Figure 11 In it, the first instruction A is executed using hardware resource 0b in clock cycle 0 and using hardware resource 0d in clock cycle 1. Instruction B is executed using hardware resource 0c in clock cycle 0 and using hardware resource 0b in clock cycle 1. Instruction C is executed using hardware resource 0a in clock cycle 1, using hardware resource 0c in clock cycle 2, and using hardware resource 0b in clock cycle 3. The second instruction A is executed using hardware resource 0b in clock cycle 2 and using hardware resource 0d in clock cycle 3. Observe Figure 11 , it can be found that there is no conflict in hardware resources in the same clock cycle.
[0128] In Figure 11 In the shown case, the present application adopts the result obtained by instruction scheduling, thereby more reasonably allocating instruction packets, and can make the second instruction A be issued one clock cycle earlier. Compared with the related art, the total execution time of the instruction packing method provided by the present application is reduced by one clock cycle, improving the performance.
[0129] Next, three instruction packing methods are introduced, and one of the three instruction packing methods is applied.
[0130] Figure 12 shows a flowchart of an instruction packing method provided by an exemplary embodiment of the present application. Taking the method being executed by a compiler as an example, the method includes:
[0131] Step 1220, during the execution of instruction scheduling, record the correspondence between the first quantity of instructions and the second quantity of issue clock cycles, where instruction scheduling is used to sort the first quantity of instructions according to the values of the issue clock cycles;
[0132] During the execution of instruction scheduling, save a correspondence table of the first quantity of instructions and the second quantity of issue clock cycles. In the correspondence table, one instruction corresponds to one issue clock cycle, and one issue clock cycle corresponds to one instruction, zero instructions, or at least two instructions.
[0133] Schematically, the correspondence is shown in the following table:
[0134] Emission clock cycle 0 Emission clock cycle 1 Emission clock cycle 2 Instruction 1, Instruction 2 Instruction 3
[0135] In the above table, issue clock cycle 0 corresponds to instruction 1 and instruction 2, issue clock cycle 1 does not correspond to an instruction, and issue clock cycle 2 corresponds to instruction 3.
[0136] Step 1240, traverse the second quantity of issue clock cycles one by one, and pack the instructions with the same issue clock cycle among the first quantity of instructions to obtain multiple instruction packets.
[0137] In one embodiment, the second number of emission clock cycles includes a consecutive plurality of emission clock cycles starting from zero, and traversed in the order starting from zero of the emission clock cycles.
[0138] Figure 13 The flowchart of an instruction packing method provided by an embodiment of the present application is shown.
[0139] Step 1301, start;
[0140] Step 1302, create a first variable;
[0141] Step 1303, initialize the first variable to zero;
[0142] Step 1304, determine whether the traversal of the second number of emission clock cycles is completed;
[0143] Access the comparison table of the first number of instructions and the second number of emission clock cycles in ascending order of the emission clock cycles. Determine whether the traversal of the second number of emission clock cycles is completed. If the traversal is completed, execute step 1309; if the traversal is not completed, execute step 1305.
[0144] Step 1305, determine whether there is an instruction corresponding to the first emission clock cycle;
[0145] The first emission clock cycle is the emission clock cycle indicated by the value of the first variable. For example, if the value of the first variable is 0, the first emission clock cycle is cycle 0; if the value of the first variable is 1, the first emission clock cycle is cycle 1. If there is an instruction corresponding to the first emission clock cycle, execute step 1306; if there is no instruction corresponding to the first emission clock cycle, execute step 1308.
[0146] Step 1306, create a new instruction packet;
[0147] Step 1307, put the instruction corresponding to the first emission clock cycle into the newly created instruction packet;
[0148] The newly created instruction packet is used to be sent to the hardware execution unit according to the first emission clock cycle.
[0149] In one embodiment, if there is a requirement for the instruction order in the instruction packet, the instruction corresponding to the first emission clock cycle is reasonably sorted.
[0150] Step 1308, increment the value of the first variable by one;
[0151] After executing step 1308, return to execute step 1304.
[0152] Step 1309, end.
[0153] In summary, the above embodiment describes a method for executing instruction packaging by iterating through a second number of transmit clock cycles one by one. The above embodiment provides a specific execution strategy. During execution, only the first variable needs to be set to iterate through all transmit clock cycles, making the operation simple. Furthermore, when the value of the first number is greater than the value of the second number, iterating through the second number of transmit clock cycles one by one will achieve a faster packaging speed.
[0154] Figure 14 A flowchart of an instruction packing method provided by an exemplary embodiment of the present application is shown. Taking the method executed by a compiler as an example, the method includes:
[0155] Step 1420: During instruction scheduling, a correspondence between the first number of instructions and the second number of issue clock cycles is recorded, wherein the instruction scheduling is used to sort the first number of instructions according to the number of issue clock cycles.
[0156] During instruction scheduling, a comparison table of a first number of instructions and a second number of issue clock cycles is stored, wherein one instruction corresponds to one issue clock cycle, and one issue clock cycle corresponds to one instruction, zero instruction, or at least two instructions.
[0157] Indicatively, the corresponding relationship is shown in the following table:
[0158] Emission clock cycle 0 Emission clock cycle 1 Emission clock cycle 2 Instruction 1, Instruction 2 Instruction 3
[0159] In the above table, the emission clock cycle 0 corresponds to instruction 1 and instruction 2, the emission clock cycle 1 does not correspond to an instruction, and the emission clock cycle 2 corresponds to instruction 3.
[0160] Step 1440 , traverse the first number of instructions one by one, and pack instructions with the same issue clock cycle in the first number of instructions to obtain multiple instruction packets.
[0161] Figure 15 The flowchart of the instruction packing method provided by one embodiment of the present application is shown. The instruction scheduling is used to sort the first number of instructions in ascending order of the issuance clock cycle to obtain an instruction sequence, wherein the issuance clock cycle corresponding to the first instruction in the instruction sequence is zero.
[0162] Step 1501, start;
[0163] Step 1502, creating a second variable;
[0164] Step 1503, initializing the second variable to zero;
[0165] Step 1504: determine whether all instructions in the instruction sequence have not been packaged;
[0166] According to the instruction order, access the comparison table of the first quantity of instructions and the second quantity of emission clock cycles. Determine whether all the instructions in the instruction sequence have been completely packed. If not all the instructions in the instruction sequence have been completely packed, execute step 1505; if all the instructions in the instruction sequence have been completely packed, then execute step 1512.
[0167] Step 1505, obtain the i-th instruction in the instruction sequence;
[0168] The instruction sequence includes the first quantity of instructions arranged in ascending order of emission clock cycles. The emission clock cycle corresponding to the first instruction in the instruction sequence is zero. The initial value of i is 1.
[0169] Step 1506, determine whether the value of the i-th emission clock cycle is greater than the value of the second variable;
[0170] The i-th emission clock cycle is the emission clock cycle corresponding to the i-th instruction.
[0171] If the value of the i-th emission clock cycle is greater than the value of the second variable, then execute step 1507; if the value of the i-th emission clock cycle is not greater than the value of the second variable, then execute step 1510.
[0172] It can be understood that the first quantity of instructions in the instruction sequence has already been arranged in ascending order of emission clock cycles. Therefore, there is no situation where the value of the i-th emission clock cycle is less than the second variable.
[0173] Step 1507, when there are instructions in the current instruction packet, determine to emit the current instruction packet in the i-th emission clock cycle;
[0174] Step 1508, create a new instruction packet;
[0175] Step 1509, increment the value of the second variable by one;
[0176] After executing step 1509, return to execute step 1506.
[0177] Step 1510, put the i-th instruction into the current instruction packet;
[0178] Step 1511, increment i by one;
[0179] After executing step 1511, return to execute step 1504.
[0180] Step 1512, end.
[0181] In summary, the above embodiments introduce a method for performing instruction packing by traversing the first number of instructions one by one (specifically, traversing the issue clock cycles of the first number of instructions). The above embodiments provide a specific execution idea. During execution, only the second variable needs to be set to traverse all the instructions, and the operation is simple. Moreover, when the value of the first number is less than the value of the second number, traversing the first number of instructions one by one will achieve a faster packing speed.
[0182] Figure 16 The flowchart of an instruction packing method provided by an exemplary embodiment of the present application is shown. Taking the method as being executed by a compiler as an example, the method includes:
[0183] Step 1620, during the execution of instruction scheduling, record the correspondence between the first number of instructions and the second number of issue clock cycles, where instruction scheduling is used to sort the first number of instructions according to the values of the issue clock cycles;
[0184] During the execution of instruction scheduling, save the correspondence table between the first number of instructions and the second number of issue clock cycles. In the correspondence table, one instruction corresponds to one issue clock cycle, and one issue clock cycle corresponds to one instruction, zero instructions, or at least two instructions.
[0185] Schematically, the correspondence is shown in the following table:
[0186] Emission clock cycle 0 Emission clock cycle 1 Emission clock cycle 2 Instruction 1, Instruction 2 Instruction 3
[0187] In the above table, issue clock cycle 0 corresponds to instruction 1 and instruction 2, issue clock cycle 1 does not correspond to any instruction, and issue clock cycle 2 corresponds to instruction 3.
[0188] Step 1640, in the instruction sequence, insert multiple delimiters, where the delimiters are used to separate instructions with different issue clock cycles; the instruction sequence is a sequence obtained by sorting the first number of instructions according to the values of the issue clock cycles through instruction scheduling;
[0189] Schematically, the instruction sequence after inserting the delimiters is: instruction 1 (cycle 0), instruction 2 (cycle 0), instruction 3 (cycle 0), delimiter, instruction 4 (cycle 1), delimiter, instruction 5 (cycle 2), delimiter, instruction 6 (cycle 4), instruction 7 (cycle 4), delimiter, instruction 8 (cycle 5), instruction 9 (cycle 5), where cycle is the issue clock cycle, and there is a situation where there is no instruction corresponding to an issue clock cycle at this time.
[0190] The delimiter can be a pseudo-instruction or a Flag attached to the first instruction.
[0191] Step 1660: Identify multiple delimiters in the instruction sequence, and pack the instructions between two adjacent delimiters to obtain multiple instruction packets.
[0192] Identify the delimiters in the instruction sequence, and pack the instructions between two adjacent delimiters to obtain multiple instruction packets. Thus, the first instruction packet (Instruction 1, Instruction 2, Instruction 3), the second instruction packet (Instruction 4), the third instruction packet (Instruction 5), the fourth instruction packet (Instruction 6, Instruction 7), and the fifth instruction packet (Instruction 8, Instruction 9) are obtained.
[0193] In summary, the above embodiments introduce dividing the instruction packing stage into two independent stages. The first stage is responsible for delimiting the scope of the instruction packets with some delimiters, and the second stage is responsible for identifying these delimiters and packing the instructions. Executing through two independent stages has certain advantages in some special cases. For example, if the compiler needs to implement some other optimizations that are incompatible with the instruction packet format after instruction scheduling and before instruction packing, the solution of the above embodiments can be adopted.
[0194] Specifically for Step 1640 Figure 17 Illustrates a method for inserting delimiters provided by an exemplary embodiment of the present application. The second quantity of emission clock cycles includes a plurality of consecutive emission clock cycles starting from zero, and the delimiter is implemented as a pseudo-instruction.
[0195] Step 1701: Start
[0196] Step 1702: Create a third variable
[0197] Step 1703: Initialize the third variable to zero
[0198] Step 1704: Determine whether the second quantity of emission clock cycles has been traversed
[0199] Access the look-up table of the first quantity of instructions and the second quantity of emission clock cycles in ascending order of the emission clock cycles. Determine whether the second quantity of emission clock cycles has been traversed. If traversed, execute Step 1708; if not traversed, then execute Step 1705.
[0200] Step 1705: Determine whether there is an instruction corresponding to the third emission clock cycle
[0201] The third emission clock cycle is the emission clock cycle corresponding to the value of the third variable. For example, if the value of the third variable is 0, the third emission clock cycle is cycle 0; if the value of the third variable is 1, the third emission clock cycle is cycle 1.
[0202] According to the comparison table of the first quantity of instructions and the second quantity of emission clock cycles, determine whether there is an instruction corresponding to the third emission clock cycle. If there is no instruction corresponding to the third emission cycle, execute step 1707; if there is an instruction corresponding to the third emission cycle, execute step 1706.
[0203] Step 1706, insert a pseudo-instruction before the first instruction corresponding to the third emission clock cycle;
[0204] Step 1707, increment the value of the third variable by one;
[0205] After executing step 1707, return to execute step 1704.
[0206] Step 1708, end.
[0207] In summary, the above embodiments introduce the method of inserting a delimiter. By traversing the second quantity of emission clock cycles one by one, the delimiter is inserted. When executing, only the third variable needs to be set to traverse the second quantity of emission clock cycles, and the operation is simple.
[0208] Regarding step 1660, Figure 18 The flowchart of the packing method provided by an exemplary embodiment of the present application is shown. The delimiter is a pseudo-instruction, and instruction scheduling is used to sort the first quantity of instructions in ascending order of emission clock cycles to obtain an instruction sequence. The emission clock cycle corresponding to the first instruction in the instruction sequence is zero.
[0209] Step 1801, start;
[0210] Step 1802, determine whether all the first quantity of instructions have been traversed;
[0211] If all the first quantity of instructions have been traversed, execute step 1808; if not all the first quantity of instructions have been traversed, execute step 1803.
[0212] In Figure 18 In the specific packing method shown, the valid instructions in the instruction sequence will be sorted in ascending order of emission clock cycles. Check all the instructions in the instruction sequence in order; during the check, if the inspected instruction is a pseudo-instruction, create a new instruction packet; if the inspected instruction is not a pseudo-instruction, put the inspected instruction into the newly created instruction packet.
[0213] Step 1803, check the next instruction;
[0214] Step 1804, determine whether the instruction is a pseudo-instruction;
[0215] If the current instruction is a pseudo-instruction, execute step 1805; if the current instruction is not a pseudo-instruction, execute step 1807.
[0216] Step 1805, transmit the existing instruction packet;
[0217] Step 1806, create a new instruction packet;
[0218] After executing Step 1806, return to execute Step 1802.
[0219] Step 1807, put the instruction into the instruction packet;
[0220] After executing Step 1807, return to execute Step 1802.
[0221] Step 1808, transmit the existing instruction packet;
[0222] When traversing all the instructions of the first quantity is completed, transmit the last instruction packet.
[0223] Step 1809, end.
[0224] In summary, the separator is a pseudo-instruction. The above embodiments introduce a method for identifying pseudo-instructions and performing instruction packaging. The above embodiments will traverse all the instructions in the instruction sequence one by one, and can identify pseudo-instructions, instruction packaging, and instruction packet transmission simultaneously, with relatively high efficiency.
[0225] Figure 19 The flowchart of the instruction scheduling method provided by an exemplary embodiment of the present application is shown. Taking the method being executed by a compiler as an example. The method includes:
[0226] Step 1901, start;
[0227] Step 1902, construct a directed acyclic graph;
[0228] In one embodiment, a directed acyclic graph is constructed according to the input instructions to represent the dependency relationship between instructions. And, priority topological sorting is also performed and some parameters are initialized.
[0229] Step 1903, initialize the variable cycle = 1; the Ready list contains multiple child nodes of the parent nodes in the directed acyclic graph; the Active list is an empty set;
[0230] Step 1904, determine whether the instruction from the Ready list is put into the Active list through two conditions;
[0231] The two conditions include: 1. The instruction can be transmitted in the clock cycle corresponding to the current variable cycle; 2. There is no hardware resource conflict between the instruction and the transmitted instructions.
[0232] If the two conditions are met, put the instruction into the Active list.
[0233] Check the clock cycle information and hardware resource conflict information of the instruction, and put the instructions that meet the emission conditions in the Ready list into the Active list. If there are no instructions that meet the emission conditions, continuously increment the value of the variable cycle until an available instruction is found.
[0234] Step 1905, determine whether the Active list is an empty set;
[0235] If the Active list is not an empty set, execute Step 1906; if the Active list is an empty set, execute Step 1907.
[0236] Step 1906, increment the variable cycle by 1; update the hardware resource information;
[0237] Step 1907, determine whether the Active list contains multiple instructions;
[0238] If the Active list contains multiple instructions, execute Step 1908; if the Active list does not contain multiple instructions, execute Step 1909.
[0239] Step 1908, score the instructions in the Active list and select the most suitable instruction at present;
[0240] Step 1909, adjust the instruction positions;
[0241] If there is only one instruction that can be emitted, directly adjust the order of this instruction; if there are multiple instructions that can be emitted, score each instruction, select the most suitable instruction at present for order adjustment.
[0242] Step 1910, save a comparison table of the variable cycle and instructions;
[0243] Step 1911, put the successor nodes of this instruction in the directed acyclic graph into the Ready list;
[0244] Step 1912, update the hardware resource information;
[0245] Step 1913, determine whether the Active list is an empty set and, the Ready list is an empty set;
[0246] If both the Active list and the Ready list are empty sets, it means that all instructions have been sorted, execute Step 1914; if the Active list and the Ready list are not both empty sets, execute Step 1904.
[0247] Step 1914, end.
[0248] Next, the beneficial effects of the instruction packing method provided by this application will be introduced according to the experimental results.
[0249] Compared with the instruction packing method in the related art, in this application, during instruction scheduling, the launch clock cycles of the first quantity of instructions are directly saved, and during packing, the launch clock cycles are directly used for packing, and the instructions with the same launch clock cycle are packed into one instruction packet. Doing so brings two benefits:
[0250] After the compiler enables instruction table scheduling and instruction modulo scheduling, the performance differences brought by using the packing methods in the related art and the packing method proposed in this application are compared for some common operators. The total execution clock cycles (cycles) are the final execution duration of the operator under hardware, and the unit is cycles. As shown in the following table.
[0251] Operator Related technology This application Cycle reduction Sqrt 21755 18477 15.07% Glu 13143 11710 10.90% 0Atanh 10154 9107 10.31% Eltwise 7778 7019 9.76% Cos 67527 62216 7.87% Tanh 33587 30953 7.84% Sigmoid 6208 5862 5.57% Upsample 54585 54558 0.05%
[0252] As can be seen from the above table, when using the launch clock cycles obtained by instruction scheduling during instruction packing and packing the instructions with the same launch clock cycle into one instruction packet, the total execution duration of the operator can be effectively reduced.
[0253] Figure 20 The structural block diagram of an instruction packing device provided by an exemplary embodiment of this application is shown. The device includes:
[0254] A saving module 2001, configured to record the launch clock cycles of the first quantity of instructions respectively during the execution of instruction scheduling, and the instruction scheduling is used to sort the first quantity of instructions according to the values of the launch clock cycles;
[0255] A packing module 2002, configured to pack the instructions with the same launch clock cycle among the first quantity of instructions to obtain one or more instruction packets, and each instruction packet in the multiple instruction packets is used to be sent to the hardware execution unit according to the corresponding launch clock cycle.
[0256] In an optional embodiment, the saving module 2001 is further configured to record the correspondence between the first quantity of instructions and the second quantity of launch clock cycles; wherein, one instruction in the correspondence corresponds to one launch clock cycle, and one launch clock cycle in the correspondence corresponds to one instruction, zero instructions or at least two instructions.
[0257] In an optional embodiment, the packing module 2002 is further configured to traverse the second quantity of launch clock cycles one by one, and pack the instructions with the same launch clock cycle among the first quantity of instructions to obtain multiple instruction packets.
[0258] In an optional embodiment, the second number of emission clock cycles includes a continuous plurality of emission clock cycles starting from zero. The packing module 2002 is further configured to create a first variable; initialize the first variable to zero; when the second number of emission clock cycles have not been traversed and there is an instruction corresponding to the first emission clock cycle, create a new instruction packet; place the instruction corresponding to the first emission clock cycle into the newly created instruction packet, and the newly created instruction packet is used to be sent to the hardware execution unit according to the first emission clock cycle, and the first emission clock cycle is the emission clock cycle indicated by the value of the first variable; increment the value of the first variable by one, and continue to traverse the second number of emission clock cycles; when the second number of emission clock cycles have not been traversed and there is no instruction corresponding to the first emission clock cycle, increment the value of the first variable by one, and continue to traverse the second number of emission clock cycles.
[0259] In an optional embodiment, the packing module 2002 is further configured to traverse the first number of instructions one by one, and pack the instructions with the same emission clock cycle among the first number of instructions to obtain a plurality of instruction packets.
[0260] In an optional embodiment, the instruction scheduling is used to sort the first number of instructions in ascending order of the emission clock cycle to obtain an instruction sequence, and the emission clock cycle corresponding to the first instruction in the instruction sequence is zero. The packing module 2002 is further configured to create a second variable; initialize the second variable to zero; when not all the instructions in the instruction sequence have been packed, obtain the i-th instruction in the instruction sequence; when the value of the i-th emission clock cycle is not greater than the value of the second variable, place the i-th instruction into the current instruction packet, and the i-th emission clock cycle is the emission clock cycle corresponding to the i-th instruction; increment i by one, and the initial value of i is one; re-determine whether all the instructions in the instruction sequence have been packed; when the value of the i-th emission clock cycle is greater than the value of the second variable and there are instructions in the current instruction packet, determine to emit the current instruction packet in the i-th emission clock cycle; create a new instruction packet, and increment the value of the second variable by one; re-determine the size relationship between the value of the i-th emission clock cycle and the value of the second variable.
[0261] In an optional embodiment, the packing module 2002 is further configured to insert a plurality of delimiters in the instruction sequence, and the delimiters are used to separate instructions with different emission clock cycles; the instruction sequence is a sequence obtained by the instruction scheduling to sort the first number of instructions according to the value of the emission clock cycle;
[0262] Identify a plurality of delimiters in the instruction sequence, and pack the instructions between two adjacent delimiters to obtain a plurality of instruction packets.
[0263] In an optional embodiment, the second quantity of emission clock cycles includes a plurality of consecutive emission clock cycles starting from zero. The packing module 2002 is further configured to create a third variable; initialize the third variable to zero;
[0264] When the second quantity of emission clock cycles has not been traversed completely and there is an instruction corresponding to the third emission clock cycle, insert a pseudo-instruction before the first instruction corresponding to the third emission clock cycle; increment the value of the third variable by one, and the third emission clock cycle is the emission clock cycle corresponding to the value of the third variable; continue to traverse the second quantity of emission clock cycles;
[0265] When the second quantity of emission clock cycles has not been traversed completely and there is no instruction corresponding to the third emission clock cycle, increment the value of the third variable by one; continue to traverse the second quantity of emission clock cycles.
[0266] In an optional embodiment, the delimiter is a pseudo-instruction. The packing module 2002 is further configured to sequentially check all the instructions in the instruction sequence;
[0267] During the checking process, if the checked instruction is a pseudo-instruction, create a new instruction packet;
[0268] If the checked instruction is not a pseudo-instruction, put the checked instruction into the newly created instruction packet.
[0269] In summary, by saving the emission clock cycles of each of the first quantity of instructions during instruction scheduling and packing the instructions with the same emission clock cycle into the same packet during instruction packing to obtain one or more instruction packets, compared with the related art, the packing method provided by this application no longer needs to re-analyze and compare information such as instruction dependencies and hardware resources during packing, which speeds up the packing speed. Moreover, the packing method of this application is not limited to the greedy method (related art), and can better distribute the instructions in different instruction packets by leveraging the results obtained in the instruction scheduling stage. The packing method of this application generally brings performance improvement.
[0270] Figure 21The block diagram of a computer device 2100 provided by an exemplary embodiment of the present application is shown. The computer device 2100 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The computer device 2100 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0271] Generally, the computer device 2100 includes: a processor 2101 and a memory 2102.
[0272] The processor 2101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 2101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 2101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 2101 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 2101 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0273] The memory 2102 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 2102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 2102 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 2101 to implement the instruction packaging method provided in the method embodiments of the present application.
[0274] In some embodiments, the computer device 2100 may further optionally include: a peripheral device interface 2103 and at least one peripheral device. The processor 2101, the memory 2102, and the peripheral device interface 2103 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 2103 through a bus, signal lines, or a circuit board. Exemplarily, the peripheral device may include at least one of: a radio frequency circuit 2104, a display screen 2105, a camera component 2106, an audio circuit 2107, and a power supply 2108.
[0275] The peripheral device interface 2103 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 2101 and the memory 2102. In some embodiments, the processor 2101, the memory 2102, and the peripheral device interface 2103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 2101, the memory 2102, and the peripheral device interface 2103 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0276] The radio frequency circuit 2104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 2104 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 2104 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 2104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 2104 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 2104 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0277] The display screen 2105 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 2105 is a touch display screen, the display screen 2105 also has the ability to collect touch signals on or above the surface of the display screen 2105. The touch signals can be input as control signals to the processor 2101 for processing. At this time, the display screen 2105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 2105, which is disposed on the front panel of the computer device 2100; in other embodiments, there may be at least two display screens 2105, which are respectively disposed on different surfaces of the computer device 2100 or are in a foldable design; in other embodiments, the display screen 2105 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the computer device 2100. Even further, the display screen 2105 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 2105 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0278] The camera module 2106 is used to capture images or videos. Optionally, the camera module 2106 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are respectively any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting functions or other fused shooting functions. In some embodiments, the camera module 2106 may also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0279] The audio circuit 2107 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 2101 for processing, or input to the radio frequency circuit 2104 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the computer device 2100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 2101 or the radio frequency circuit 2104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 2107 may further include a headphone jack.
[0280] The power supply 2108 is used to supply power to each component in the computer device 2100. The power supply 2108 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 2108 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0281] In some embodiments, the computer device 2100 further includes one or more sensors 2109. The one or more sensors 2109 include but are not limited to: an acceleration sensor 2110, a gyroscope sensor 2111, a pressure sensor 2112, an optical sensor 2113, and a proximity sensor 2114.
[0282] The acceleration sensor 2110 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the computer device 2100. For example, the acceleration sensor 2110 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 2101 can control the display screen 2105 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 2110. The acceleration sensor 2110 can also be used for collecting game or user's motion data.
[0283] The gyroscope sensor 2111 can detect the body direction and rotation angle of the computer device 2100. The gyroscope sensor 2111 can cooperate with the acceleration sensor 2110 to collect the 3D actions of the user on the computer device 2100. According to the data collected by the gyroscope sensor 2111, the processor 2101 can achieve the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0284] The pressure sensor 2112 can be disposed on the side frame of the computer device 2100 and / or the lower layer of the display screen 2105. When the pressure sensor 2112 is disposed on the side frame of the computer device 2100, it can detect the holding signal of the user on the computer device 2100, and the processor 2101 can perform left and right hand recognition or quick operation according to the holding signal collected by the pressure sensor 2112. When the pressure sensor 2112 is disposed on the lower layer of the display screen 2105, the processor 2101 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 2105. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0285] The optical sensor 2113 is used to collect the ambient light intensity. In one embodiment, the processor 2101 can control the display brightness of the display screen 2105 according to the ambient light intensity collected by the optical sensor 2113. For example, when the ambient light intensity is high, the display brightness of the display screen 2105 is increased; when the ambient light intensity is low, the display brightness of the display screen 2105 is decreased. In another embodiment, the processor 2101 can also dynamically adjust the shooting parameters of the camera module 2106 according to the ambient light intensity collected by the optical sensor 2113.
[0286] The proximity sensor 2114, also known as the distance sensor, is usually disposed on the front panel of the computer device 2100. The proximity sensor 2114 is used to collect the distance between the user and the front of the computer device 2100. In one embodiment, when the proximity sensor 2114 detects that the distance between the user and the front of the computer device 2100 is gradually decreasing, the processor 2101 controls the display screen 2105 to switch from the lit state to the off state; when the proximity sensor 2114 detects that the distance between the user and the front of the computer device 2100 is gradually increasing, the processor 2101 controls the display screen 2105 to switch from the off state to the lit state.
[0287] Those skilled in the art can understand that Figure 21 the structure shown in does not constitute a limitation on the computer device 2100, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0288] The present application also provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the instruction packaging method provided in the above method embodiment. The present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the instruction packaging method provided in the above method embodiment.
[0289] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0290] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.
[0291] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for packing instructions, characterized in that: The method comprises: During execution of instruction scheduling, recording respective issue clock cycles of a first number of instructions, wherein the instruction scheduling is used to sort the first number of instructions according to the values of the issue clock cycles; Instructions with the same issue clock cycle among the first number of instructions are packaged to obtain one or more instruction packets, each of the multiple instruction packets being sent to a hardware execution unit according to a corresponding issue clock cycle.
2. The method according to claim 1, characterized in that The recording of the issue clock cycles of the first number of instructions includes: Recording a comparison between the first number of instructions and the second number of issue clock cycles; Among them, one instruction in the comparison relationship corresponds to one issuance clock cycle, and one issuance clock cycle in the comparison relationship corresponds to one instruction, zero instruction, or at least two instructions.
3. The method according to claim 2, characterized in that The step of packing instructions with the same transmit clock cycle in the first number of instructions to obtain a plurality of instruction packets includes: The second number of issue clock cycles are traversed one by one, and instructions with the same issue clock cycle among the first number of instructions are packaged to obtain the multiple instruction packets.
4. The method according to claim 3, characterized in that The second number of transmit clock cycles includes a plurality of consecutive transmit clock cycles starting from zero; The step of traversing the second number of issue clock cycles one by one and packing instructions with the same issue clock cycle among the first number of instructions to obtain the multiple instruction packets includes: Creating a first variable; initializing the first variable to zero; If the second number of transmit clock cycles has not been completely traversed and there is an instruction corresponding to the first transmit clock cycle, creating a new instruction packet; placing the instruction corresponding to the first transmit clock cycle into the new instruction packet, and sending the new instruction packet to the hardware execution unit according to the first transmit clock cycle, where the first transmit clock cycle is the transmit clock cycle indicated by the value of the first variable; increasing the value of the first variable by one, and continuing to traverse the second number of transmit clock cycles; When the second number of emission clock cycles has not been completely traversed and no instruction corresponds to the first emission clock cycle, the value of the first variable is increased by one, and the second number of emission clock cycles is continued to be traversed.
5. The method according to claim 2, characterized in that The step of packing instructions with the same transmit clock cycle in the first number of instructions to obtain a plurality of instruction packets includes: The first number of instructions are traversed one by one, and instructions with the same issue clock cycle among the first number of instructions are packaged to obtain the multiple instruction packages.
6. The method according to claim 5, characterized in that The instruction scheduling is used to sort the first number of instructions in ascending order of the emission clock cycles to obtain an instruction sequence, wherein the emission clock cycle corresponding to the first instruction in the instruction sequence is zero; The step of traversing the first number of instructions one by one and packing the instructions with the same transmit clock cycle among the first number of instructions to obtain the multiple instruction packets includes: Creating a second variable; initializing the second variable to zero; If not all instructions in the instruction sequence have been packaged, obtaining the i-th instruction in the instruction sequence; If the value of the i-th transmit clock cycle is not greater than the value of the second variable, placing the i-th instruction into the current instruction packet, where the i-th transmit clock cycle is the transmit clock cycle corresponding to the i-th instruction; incrementing i by one, with the initial value of i being one; and re-determining whether all instructions in the instruction sequence have been packaged; When the value of the i-th emission clock cycle is greater than the value of the second variable and there is an instruction in the current instruction packet, determine to transmit the current instruction packet in the i-th emission clock cycle; create a new instruction packet, increase the value of the second variable by one; and re-determine the size relationship between the value of the i-th emission clock cycle and the value of the second variable.
7. The method according to claim 2, characterized in that The step of packing instructions with the same transmit clock cycle in the first number of instructions to obtain a plurality of instruction packets includes: Inserting a plurality of separators into an instruction sequence, wherein the separators are used to separate instructions having different issue clock cycles; the instruction sequence is a sequence obtained by sorting the first number of instructions according to the values of the issue clock cycles through the instruction scheduling; The multiple delimiters are identified in the instruction sequence, and the instructions in two adjacent delimiters are packaged to obtain the multiple instruction packets.
8. The method according to claim 7, characterized in that The second number of transmit clock cycles includes a plurality of consecutive transmit clock cycles starting from zero; In the instruction sequence, multiple separators are inserted, including: Creating a third variable; initializing the third variable to zero; If the second number of transmit clock cycles has not been completely traversed and there is an instruction corresponding to the third transmit clock cycle, insert a pseudo instruction before the first instruction corresponding to the third transmit clock cycle; increase the value of the third variable by one, and the third transmit clock cycle is the transmit clock cycle corresponding to the value of the third variable; and continue traversing the second number of transmit clock cycles; When the second number of emission clock cycles has not been completely traversed and no instruction corresponds to the third emission clock cycle, the value of the third variable is increased by one; and the second number of emission clock cycles is continued to be traversed.
9. The method according to claim 7, characterized in that The delimiter is a pseudo-instruction; the identifying the multiple delimiters in the instruction sequence, and packing the instructions in two adjacent delimiters to obtain the multiple instruction packets, includes: Checking all instructions in the instruction sequence in order; During the checking process, if the checked instruction is the pseudo instruction, a new instruction packet is created; If the checked instruction is not the pseudo instruction, the checked instruction is placed in a newly created instruction package.
10. A command packaging device, characterized in that: The device comprises: a storage module, configured to record, during execution of instruction scheduling, respective issuance clock cycles of a first number of instructions, wherein the instruction scheduling is configured to sort the first number of instructions according to values of the issuance clock cycles; A packing module is used to pack instructions with the same said emission clock cycle in the first number of instructions to obtain one or more instruction packets, each of the multiple instruction packets is used to be sent to the hardware execution unit according to the corresponding emission clock cycle.
11. The device according to claim 10, characterized in that The storage module is further configured to record a comparison between the first number of instructions and the second number of emission clock cycles; Among them, one instruction in the comparison relationship corresponds to one issuance clock cycle, and one issuance clock cycle in the comparison relationship corresponds to one instruction, zero instruction, or at least two instructions.
12. The device according to claim 11, characterized in that The packing module is further configured to traverse the second number of emission clock cycles one by one, pack the instructions with the same emission clock cycle among the first number of instructions, and obtain the multiple instruction packets.
13. A computer device, characterized in that: The computer device includes: a processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the instruction packing method according to any one of claims 1 to 9.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the instruction packing method according to any one of claims 1 to 9.
15. A computer program product, characterized in that The computer program product stores a computer program, and the computer program is loaded and executed by a processor to implement the instruction packing method according to any one of claims 1 to 9.