Instruction execution method and apparatus based on wait mechanism, and device and storage medium

By using a wait mechanism-based method in instruction execution, matching functional units and generating wait scheduling instructions, the problem of low scheduling efficiency of VLIW or loop structure in the prior art is solved, and higher instruction execution efficiency and program performance are achieved.

WO2025112180A1PCT designated stage expired Publication Date: 2025-06-05SHANGHAI SMARTLOGIC TECHNOLOGY LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/072951
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-01-18
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When the prior art processes ultra-long byte instructions (VLIW) or machine instructions with loop structure, the instruction scheduling efficiency is low, and it is difficult to reasonably allocate functional units to improve the performance of the target program.

Method used

Using the instruction execution method based on the wait mechanism, by obtaining the assigned instructions in the loop structure to be executed, matching the target functional unit according to the instruction type, determining the instruction scheduling result, and generating and executing wait scheduling instructions to optimize the instruction execution efficiency.

Benefits of technology

By evenly allocating the instructions to be allocated to the functional unit and generating wait scheduling instructions, the instruction execution efficiency is significantly improved and the program performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024072951_05062025_PF_FP_ABST
    Figure CN2024072951_05062025_PF_FP_ABST
Patent Text Reader

Abstract

An instruction execution method and apparatus based on a wait mechanism, and a device and a storage medium. The method comprises: acquiring at least one instruction to be allocated in any loop body of a loop structure to be executed (S110); on the basis of the instruction type of each instruction to be allocated, selecting, from at least one functional unit used for executing the instruction(s) to be allocated, a target functional unit matching each instruction to be allocated (S120); on the basis of an instruction delay of each instruction to be allocated and the target functional unit matching each instruction to be allocated, determining an instruction scheduling result (S130); and on the basis of the instruction scheduling result and a wait instruction generation mechanism, generating and executing a wait scheduling instruction (S140).
Need to check novelty before this filing date? Find Prior Art

Description

Instruction execution method, device, equipment and storage medium based on wait mechanism

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311604958.7, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of data processing technology, and for example, to an instruction execution method, apparatus, device, and storage medium based on a wait mechanism. Background Art

[0003] The task of sequencing the operations within a program block or procedure to efficiently utilize processor resources is called instruction scheduling. The execution time of a set of instructions is critically dependent on the order in which they are executed. Instruction scheduling reorders the instructions within a procedure so that as many instructions as possible are executed per cycle, improving instruction execution time.

[0004] Compiler instruction scheduling methods include list scheduling, greedy scheduling, and anti-dependency-breaking scheduling. However, for Very Long Instruction Word (VLIW) instructions or machine instructions with loop structures, these instructions take a long time to execute, making common instruction scheduling methods inefficient. Therefore, how to rationally allocate functional units to improve target program performance and thus increase instruction execution efficiency has become a pressing issue.

[0005] Summary of the Invention

[0006] The present application provides an instruction execution method, apparatus, device and storage medium based on a wait mechanism to improve instruction scheduling efficiency.

[0007] According to one aspect of the present application, a method for executing an instruction based on a wait mechanism is provided, the method comprising:

[0008] Obtain at least one instruction to be allocated in any loop body of a loop structure to be executed;

[0009] selecting, according to the instruction type of each instruction to be assigned, a target functional unit that matches each instruction to be assigned from at least one functional unit for executing the instruction to be assigned;

[0010] Determining an instruction scheduling result according to the instruction delay of each of the instructions to be assigned and based on the target functional units matched by each of the instructions to be assigned;

[0011] According to the instruction scheduling result, based on the wait instruction generation mechanism, a wait scheduling instruction is generated and executed.

[0012] Optionally, according to the instruction type of each instruction to be assigned, selecting a target functional unit that matches each instruction to be assigned from at least one functional unit for executing the instruction to be assigned, includes: determining the functional type of at least one functional unit for executing the instruction to be assigned; according to the instruction type of each instruction to be assigned, based on the functional type of each functional unit, selecting a target functional unit that matches each instruction to be assigned from at least one functional unit.

[0013] Optionally, the instruction scheduling result is determined based on the instruction delay of each of the instructions to be assigned and the target functional unit that matches each of the instructions to be assigned, including: determining the instruction association relationship between each of the instructions to be assigned; and determining the instruction scheduling result based on the instruction association relationship, the target functional unit of each of the instructions to be assigned and the instruction delay of each of the instructions to be assigned.

[0014] Optionally, generating and executing a wait scheduling instruction based on the instruction scheduling result and based on a wait instruction generation mechanism includes: generating a wait instruction for each of the functional units based on the wait instruction generation mechanism according to the instruction scheduling result; generating and executing a wait scheduling instruction based on the wait instruction of each of the functional units.

[0015] According to another aspect of the present application, there is provided an instruction execution device based on a wait mechanism, the device comprising:

[0016] An instruction acquisition module is configured to acquire at least one instruction to be allocated in any loop body of a loop structure to be executed;

[0017] a functional unit matching module configured to select, based on the instruction type of each instruction to be assigned, a target functional unit that matches each instruction to be assigned from at least one functional unit configured to execute the instruction to be assigned;

[0018] a scheduling result determination module configured to determine an instruction scheduling result based on an instruction delay of each of the instructions to be assigned and a target functional unit matched with each of the instructions to be assigned;

[0019] The scheduling instruction generation module is configured to generate and execute a wait scheduling instruction based on the wait instruction generation mechanism according to the instruction scheduling result.

[0020] Optionally, the functional unit matching module includes: a functional type determination unit, configured to determine the functional type of at least one functional unit configured to execute the instruction to be assigned; a functional unit matching unit, configured to select a target functional unit that matches each instruction to be assigned from at least one functional unit based on the instruction type of each instruction to be assigned and the functional type of each functional unit.

[0021] Optionally, the scheduling result determination module includes: an instruction relationship determination unit, configured to determine the instruction association relationship between each of the instructions to be assigned; a scheduling result determination unit, configured to determine the instruction scheduling result based on the instruction association relationship, the target functional unit of each of the instructions to be assigned, and the instruction delay of each of the instructions to be assigned.

[0022] Optionally, the scheduling instruction generation module includes: a wait instruction generation unit, which is configured to generate wait instructions for each of the functional units based on the instruction scheduling results and the wait instruction generation mechanism; and a scheduling instruction generation unit, which is configured to generate and execute wait scheduling instructions based on the wait instructions of each of the functional units.

[0023] According to another aspect of the present application, an electronic device is provided, comprising:

[0024] at least one processor; and

[0025] a memory communicatively connected to the at least one processor; wherein,

[0026] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the instruction execution method based on the wait mechanism described in any embodiment of the present application.

[0027] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the instruction execution method based on the wait mechanism described in any embodiment of the present application when executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] FIG1 is a flowchart of an instruction execution method based on a wait mechanism according to a first embodiment of the present application;

[0029] FIG2 is a flowchart of an instruction execution method based on a wait mechanism according to a second embodiment of the present application;

[0030] FIG3 is a schematic structural diagram of an instruction execution device based on a wait mechanism according to a third embodiment of the present application;

[0031] FIG4 is a schematic structural diagram of an electronic device that implements the instruction execution method based on the wait mechanism according to an embodiment of the present application. DETAILED DESCRIPTION

[0032] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0033] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0034] Example 1

[0035] FIG1 is a flowchart of an instruction execution method based on a wait mechanism provided in Example 1 of the present application. This embodiment is applicable to the case where a machine instruction with a loop structure is reasonably allocated to functional units and instruction scheduling is performed. The method can be executed by an instruction execution device based on a wait mechanism. The instruction execution device based on the wait mechanism can be implemented in the form of hardware and / or software. The instruction execution device based on the wait mechanism can be configured in an electronic device. As shown in FIG1 , the method includes:

[0036] S110 , obtaining at least one instruction to be allocated in any loop body of a loop structure to be executed.

[0037] The loop structure to be executed may be a loop program including at least one loop body; the instruction to be allocated may be the instruction code in the loop body to be allocated to the functional unit.

[0038] S120 . Select, according to the instruction type of each instruction to be assigned, a target functional unit that matches each instruction to be assigned from at least one functional unit for executing the instruction to be assigned.

[0039] Functional units can be hardware execution units that perform different types of operations. Functional units are deployed in processors that execute instructions. For example, a processor that executes VLIW instructions is a VLIW processor. Functional units of different functional types can be deployed within a VLIW processor. For example, functional units of different functional types may include integer units, floating-point units, load-store units, and load-store units.

[0040] Each instruction must ultimately be assigned to a specific functional unit for execution. For example, a processor has four floating-point arithmetic-logic units (FALUs): FALU0, FALU1, FALU2, and FALU3. Ultimately, all multiplication instructions in a program must be assigned to one of these four FALUs for execution. In a VLIW processor, instructions from these four FALUs can be executed concurrently, but instructions from each FALU must be issued sequentially. Therefore, the distribution of multiplication instructions across the four FALUs affects the total execution time of the program, and thus its efficiency.

[0041] Illustratively, according to the instruction type of each instruction to be assigned, a target functional unit whose functional unit configuration matches the instruction type may be selected from at least one functional unit for executing the assigned instruction as the functional unit for executing the corresponding instruction to be assigned.

[0042] In an optional embodiment, according to the instruction type of each instruction to be assigned, a target functional unit matching each instruction to be assigned is selected from at least one functional unit for executing the instruction to be assigned, including: determining the functional type of at least one functional unit for executing the instruction to be assigned; according to the instruction type of each instruction to be assigned, based on the functional type of each functional unit, selecting a target functional unit matching each instruction to be assigned from at least one functional unit.

[0043] Different functional units have different functional types, and functional units of different functional types are used to execute to-be-allocated instructions of different instruction types.

[0044] Exemplarily, the target processor for executing the instruction to be assigned includes functional units A0, A1, B0, B1, C0, and C1. A0 and A1 are functional units of the same functional type and can be used to execute Load type instructions; B0 and B1 are functional units of the same functional type and can be used to execute Index type instructions; C0 and C1 are functional units of the same functional type and can be used to execute Mul type instructions.

[0045] Optionally, based on the instruction type of each instruction to be assigned and the function type of each functional unit, a functional unit with a matching instruction type and function type may be selected as the target functional unit for the corresponding instruction to be assigned. In one clock cycle, one functional unit is bound to only one instruction to be assigned.

[0046] S130 , determining an instruction scheduling result according to the instruction delay of each instruction to be assigned and based on the target functional unit that matches each instruction to be assigned.

[0047] The instruction delay may be the delay time during the execution of the instruction to be assigned, and the instruction delays of different instructions to be assigned may be the same or different.

[0048] Exemplarily, the instruction scheduling result is determined based on the instruction latency of each instruction to be assigned, the target functional unit that matches each instruction to be assigned, and a corresponding scheduling method, such as list scheduling. It should be noted that due to the association and dependency between instructions, the association relationship between instructions also needs to be considered in the process of determining the instruction scheduling result.

[0049] When a compiler is performing instruction scheduling, it needs to rely on a directed acyclic graph (DAG) to analyze the data flow of instructions. DAG is a common internal structure used by compilers to schedule instructions.

[0050] If the DAG formed by all instructions is G = (V, E), where V is a set of vertices (machine instructions, also referred to as instructions, also called nodes in a directed acyclic graph), and E is a set of edges (dependencies between instructions), the total number of nodes in the graph is N(G). V is a set of machine instructions, and each node in the set represents a machine instruction; E is a set of dependency edges between instructions, and each element in E is an ordered pair (edge) (u, v). Where u, v∈V, denoted as u→v, indicates that machine instruction u is the predecessor node of machine instruction v (or v is the successor node of u). If an output operand of u is used as an input operand by v, the edge u→v describes a data dependency. In a directed acyclic graph, the edge set E contains types other than the data dependencies described above, including anti-dependencies, output dependencies, control dependencies, and other types.

[0051] In an optional embodiment, the instruction scheduling result is determined based on the instruction delay of each instruction to be assigned and the target functional unit matching each instruction to be assigned, including: determining the instruction association relationship between each instruction to be assigned; determining the instruction scheduling result based on the instruction association relationship, the target functional unit of each instruction to be assigned and the instruction delay of each instruction to be assigned.

[0052] If there is the following loop structure to be executed: Global G; Loop(N){ Load(Address_0++)->A;(Instruction delay: 4) Load(Address_1++)->B;(Instruction delay: 4) Index(A)->C;(Instruction delay: 2) Index(B)->D;(Instruction delay: 2) Mul(C,G)->E;(Instruction delay: 3) Mul(D,G)->F;(Instruction delay: 3) Store(E,Address_2++); Store(F,Address_3++);}

[0053] The number of loops in the loop body is N. The loop body contains 8 instructions to be allocated, and the variable G is the global data outside the loop body.

[0054] The execution of the Index(A) instruction depends on the predecessor instruction Load(Address_0++). Therefore, the Index(A) instruction and Load(Address_0++) instruction have an associated relationship. The same applies to the other instructions in the loop body, such as Mul(C,G) which depends on Index(A). This example will not be further described.

[0055] For example, if the target processor includes 12 functional units, they are:

[0056] BIU0 / BIU1 / BIU2 / BIU3: Bus Interface Unit, memory access unit, a total of 4, used to execute Load / Store instructions;

[0057] SHU0 / SHU1 / SHU2 / SHU3: Shuffle Unit, data interleaving processing unit, a total of 4, used to execute Index class instructions;

[0058] ALU0 / ALU1 / ALU2 / ALU3: Arithmetic Logic Unit, logical operation unit, a total of 4, used to execute Mul (multiplication) type instructions.

[0059] For a loop body, the allocation results are as follows:

[0060] 1.Load(Address_0++)->A; (allocation functional unit: BIU0)

[0061] 2.Load(Address_1++)->B; (allocation functional unit: BIU1)

[0062] 3. Index (A) -> C; (Assignment functional unit: SHU0)

[0063] 4. Index(B)->D; (Assignment functional unit: SHU1)

[0064] 5.Mul(C,G)->E; (assignment functional unit: ALU0)

[0065] 6.Mul(D,G)->F; (assignment functional unit: ALU1)

[0066] 7.Store(E,Address_2++); (allocation functional unit: BIU2)

[0067] 8.Store(F,Address_3++); (allocation functional unit: BIU3)

[0068] According to the instruction association relationship, the target functional unit of each instruction to be assigned, and the instruction delay of each instruction to be assigned, the generated instruction scheduling results are as follows:

[0069] The serial numbers in the table are the numbers of the instructions to be assigned. Considering instruction latency, the execution of instruction 3 depends on instruction 1, and the latency of instruction 1 is 4 clock cycles. Therefore, instruction 3 must be executed in the 5th clock cycle. The scheduling method for other instructions to be assigned is similar and will not be further described in this embodiment.

[0070] S140 . Generate and execute a wait scheduling instruction according to the instruction scheduling result and based on the wait instruction generation mechanism.

[0071] The wait instruction generation mechanism may be a mechanism for generating a wait instruction for each functional unit.

[0072] In an optional embodiment, according to the instruction scheduling result, based on the wait instruction generation mechanism, a wait scheduling instruction is generated and executed, including: according to the instruction scheduling result, based on the wait instruction generation mechanism, a wait instruction of each functional unit is generated; according to the wait instruction of each functional unit, a wait scheduling instruction is generated and executed.

[0073] For example, continuing the previous example, a wait instruction is generated for each functional unit, and the final wait instruction emission form is as follows:

[0074] BIU0:wait0||BIU1:wait0||BIU2:wait9||BIU3:wait9||SHU0:wait4||SHU1:wait 4||ALU0:wait 6||ALU1:wait 6;

[0075] 1||2||7||8||3||4||5||6; “||” is used to indicate concurrent instructions.

[0076] Traditional instruction scheduling requires generating a scheduling instruction for each clock cycle. This generates many instructions, occupies a large amount of memory space, and reduces instruction generation and execution efficiency. Using wait scheduling instruction generation, only one wait scheduling instruction is generated, resulting in fewer instructions, smaller memory usage, and higher instruction generation and execution efficiency.

[0077] The embodiment of the present application obtains at least one instruction to be assigned in any loop body of a loop structure to be executed; selects a target functional unit that matches each instruction to be assigned from at least one functional unit for executing the instruction to be assigned according to the instruction type of each instruction to be assigned; determines the instruction scheduling result based on the target functional unit that matches each instruction to be assigned according to the instruction delay of each instruction to be assigned; and generates and executes a wait scheduling instruction based on the wait instruction generation mechanism according to the instruction scheduling result. The above technical solution evenly distributes the instructions to be assigned to different functional units so that the total cycle of each functional unit is as balanced as possible, schedules the machine instructions of the assigned functional units and generates wait instructions, and uses wait instructions to improve program performance, thereby improving instruction execution efficiency.

[0078] Example 2

[0079] Figure 2 is a flow chart of an instruction execution method based on a wait mechanism provided in Example 2 of the present application. This embodiment provides an optional example based on the above embodiment.

[0080] As shown in FIG2 , the method includes the following steps:

[0081] S210: Obtain at least one instruction to be allocated in any loop body of the loop structure to be executed.

[0082] S220: Determine a function type of at least one function unit for executing the instruction to be assigned.

[0083] S230 . Selecting a target functional unit that matches each instruction to be assigned from at least one functional unit according to the instruction type of each instruction to be assigned and based on the functional type of each functional unit.

[0084] S240: Determine the instruction association relationship between the instructions to be assigned.

[0085] S250 , determining an instruction scheduling result according to the instruction association relationship, the target functional unit of each instruction to be assigned, and the instruction delay of each instruction to be assigned.

[0086] S260 . Generate wait instructions for each functional unit based on the instruction scheduling result and the wait instruction generation mechanism.

[0087] S270 . Generate and execute a wait scheduling instruction according to the wait instruction of each functional unit.

[0088] In order to illustrate the improved instruction scheduling method of the embodiment of the present application and the role played by the wait scheduling instruction in the instruction scheduling method of the embodiment of the present application, an example is given by comparing it with the traditional scheduling method.

[0089] Take the following loop structure to be executed as an example for explanation: Global G; Loop(N){ Load(Address_0++)->A; (Instruction delay: 4) Load(Address_1++)->B; (Instruction delay: 4) Index(A)->C; (Instruction delay: 2) Index(B)->D; (Instruction delay: 2) Mul(C,G)->E; (Instruction delay: 3) Mul(D,G)->F; (Instruction delay: 3) Store(E,Address_2++); Store(F,Address_3++);}

[0090] The number of loops in the loop body is N. The loop body contains 8 instructions to be allocated, and the variable G is the global data outside the loop body.

[0091] In this embodiment, the target processor is a VLIW architecture that supports the wait mechanism and includes 12 functional units, namely:

[0092] BIU0 / BIU1 / BIU2 / BIU3: Bus Interface Unit, memory access unit, a total of 4, used to execute Load / Store instructions;

[0093] SHU0 / SHU1 / SHU2 / SHU3: Shuffle Unit, data interleaving processing unit, a total of 4, used to execute Index class instructions;

[0094] ALU0 / ALU1 / ALU2 / ALU3: Arithmetic Logic Unit, logical operation unit, a total of 4, used to execute Mul type instructions;

[0095] The traditional functional unit allocation method, considering that the target processor contains a structure of four symmetrical groups of functional units, generally expands the loop four times, and then allocates the instructions in each loop to a fixed group of functional units as much as possible, thereby utilizing the parallel issuance mechanism to improve code execution efficiency. The code after the loop structure to be executed is expanded is as follows: Global G; Loop(N / 4){ 1.Load(Address_0++)->A1; 2.Load(Address_1++)->B1; 3.Index(A1)->C1; 4.Index(B1)->D1; 5.Mul(C1,G)->E1; 6.Mul(D1,G)->F1; 7.Store(E1,Address_2++); 8.Store(F1,Address_3++); Note: The above instructions are assigned to functional units: BIU0, SHU0, ALU0 9.Load(Address_0++)->A2; 10.Load(Address_1++)->B2; 11.Index(A2)->C2; 12.Index(B2)->D2; 13.Mul(C2,G)->E2; 14.Mul(D2,G)->F2; 15.Store(E2,Address_2++); 16.Store(F2,Address_3++); Note: The above instructions are assigned to functional units: BIU1, SHU1, ALU1 17.Load(Address_0++)->A3; 18.Load(Address_1++)->B3; 19.Index(A3)->C3; 20.Index(B3)->D3; 21.Mul(C3,G)->E3; 22.Mul(D3,G)->F3; 23.Store(E3,Address_2++); 24.Store(F3,Address_3++); Note: The above instructions are assigned to functional units: BIU2, SHU2, ALU2 25.Load(Address_0++)->A4; 26.Load(Address_1++)->B4; 27.Index(A4)->C4; 28.Index(B4)->D4; 29.Mul(C4,G)->E4; 30.Mul(D4,G)->F4; 31.Store(E4,Address_2++); 32.Store(F4,Address_3++); Note: The above instructions are assigned to functional units: BIU3, SHU3, ALU3}

[0096] The final instruction scheduling results of the above traditional functional unit allocation scheme are shown in the following table:

[0097] The serial numbers in the table are the numbers of the instructions to be assigned. Under the traditional functional unit allocation scheme, the total execution cycles of the loop unrolled 4 times is 11, and the total execution cycles of the loop body is (N / 4)*11.

[0098] The function unit allocation method of the embodiment of the present application does not perform expansion, but directly allocates all the function units in the original loop body evenly to the available function units. The allocation results are as follows: Global G; Loop(N){ 1.Load(Address_0++)->A; (Allocated function unit: BIU0) 2.Load(Address_1++)->B; (Allocated function unit: BIU1) 3.Index(A)->C; (Allocated function unit: SHU0) 4.Index(B)->D; (Allocated function unit: SHU1) 5.Mul(C,G)->E; (Allocated function unit: ALU0) 6.Mul(D,G)->F; (Allocated function unit: ALU1) 7.Store(E,Address_2++); (Allocated function unit: BIU2) 8.Store(F,Address_3++); (Allocated function unit: BIU3)}

[0099] The final instruction scheduling results of the functional unit allocation scheme of the embodiment of the present application are shown in the following table:

[0100] Based on the above scheduling results, a wait instruction is generated for each functional unit. The final instruction issuance form is:

[0101] BIU0:wait0||BIU1:wait0||BIU2:wait9||BIU3:wait9||SHU0:wait4||SHU1:wait 4||ALU0:wait6||ALU1:wait6;

[0102] 1||2||7||8||3||4||5||6;

[0103] The final total instruction execution cycle is: N+9.

[0104] If N=4, the final execution of the above code is:

[0105] It can be seen that when N=4, the total cycle is 13 beats, which is greater than the 11 beats of the traditional method; however, when N>4, the total cycle (N+9) generated by the scheduling method mentioned in this embodiment will be much smaller than the execution cycle of the traditional scheduling method (N / 4*11).

[0106] Furthermore, traditional scheduling methods require generating a scheduling instruction for each clock cycle. However, this application only needs to generate a wait scheduling instruction for a loop body, and traditional scheduling methods cannot generate wait scheduling instructions using the wait mechanism. In addition, the selection of the loop body in this application is based on the following requirement: it is necessary to ensure that all instructions in the loop body have no data dependencies between different loop iterations to avoid data flow errors.

[0107] Wait is an instruction used in very long byte instruction VLIW processors. The wait instruction involved in this application supports issuance to any functional unit in the VLIW processor. The function of this instruction is to postpone the issuance of instructions on the corresponding functional unit for a specified number of cycles. For example, a VLIW processor includes four functional units: BIU0, BIU1, ALU0, and ALU1. The bus interface unit (BIU) is used to execute memory access instructions, and the arithmetic logic unit (ALU) is used to execute calculation instructions.

[0108] The following is a sample code. There are 9 machine instructions in the sample code, of which instructions 0 to 3 are wait instructions. The function of instruction 0 is to postpone the issuance of the following ALU0 instruction (instruction 6) for 5 cycles, the function of instruction 1 is to postpone the issuance of the following ALU1 instruction (instruction 7) for 6 cycles, and so on.

[0109] Instruction 0: ALU0: wait 5

[0110] Instruction 1: ALU1:wait 6

[0111] Instruction 2: BIU1:wait 7

[0112] Instruction 3: BIU1:wait 8

[0113] Instruction 4: BIU0:load[addr0]->variable0

[0114] Instruction 5: BIU0:load[addr1]->variable1

[0115] Instruction 6: ALU0:variable0+variable1->result0

[0116] Instruction 7: ALU1:variable0-variable1->result1

[0117] Instruction 8: BIU1: store result0->[addr2]

[0118] Instruction 9: BIU1: store result1->[addr3]

[0119] Example 3

[0120] FIG3 is a schematic diagram of the structure of an instruction execution device based on a wait mechanism provided in the third embodiment of the present application. The instruction execution device based on a wait mechanism provided in the embodiment of the present application is applicable to the situation where the functional units of machine instructions with a loop structure are reasonably allocated and instruction scheduling is performed. The instruction execution device based on the wait mechanism can be implemented in the form of hardware and / or software. As shown in FIG3 , the device includes: an instruction acquisition module 301, a functional unit matching module 302, a scheduling result determination module 303 and a scheduling instruction generation module 304. Among them,

[0121] The instruction acquisition module 301 is configured to acquire at least one instruction to be assigned in any loop body of a loop structure to be executed;

[0122] The functional unit matching module 302 is configured to select, based on the instruction type of each instruction to be assigned, a target functional unit that matches each instruction to be assigned from at least one functional unit configured to execute the instruction to be assigned;

[0123] The scheduling result determination module 303 is configured to determine the instruction scheduling result based on the instruction delay of each instruction to be assigned and the target functional unit matched with each instruction to be assigned;

[0124] The scheduling instruction generation module 304 is configured to generate and execute a wait scheduling instruction based on the wait instruction generation mechanism according to the instruction scheduling result.

[0125] The embodiment of the present application obtains at least one instruction to be assigned in any loop body of a loop structure to be executed; selects a target functional unit that matches each instruction to be assigned from at least one functional unit set to execute the instruction to be assigned according to the instruction type of each instruction to be assigned; determines the instruction scheduling result based on the target functional unit that matches each instruction to be assigned according to the instruction delay of each instruction to be assigned; and generates and executes a wait scheduling instruction based on the wait instruction generation mechanism according to the instruction scheduling result. The above technical solution evenly distributes the instructions to be assigned to different functional units so that the total cycle of each functional unit is as balanced as possible, schedules the machine instructions of the assigned functional units and generates wait instructions, and uses wait instructions to improve program performance, thereby improving instruction execution efficiency.

[0126] Optionally, the functional unit matching module 302 includes:

[0127] a function type determining unit configured to determine a function type of at least one function unit configured to execute the instruction to be assigned;

[0128] The functional unit matching unit is configured to select a target functional unit matching each of the instructions to be assigned from at least one functional unit based on the instruction type of each of the instructions to be assigned and the functional type of each of the functional units.

[0129] Optionally, the scheduling result determination module 303 includes:

[0130] An instruction relationship determining unit, configured to determine an instruction association relationship between the instructions to be assigned;

[0131] The scheduling result determining unit is configured to determine the instruction scheduling result according to the instruction association relationship, the target functional unit of each of the instructions to be assigned, and the instruction delay of each of the instructions to be assigned.

[0132] Optionally, the scheduling instruction generation module 304 includes:

[0133] A wait instruction generation unit, configured to generate a wait instruction for each of the functional units based on the wait instruction generation mechanism according to the instruction scheduling result;

[0134] The scheduling instruction generating unit is configured to generate and execute a wait scheduling instruction according to the wait instruction of each functional unit.

[0135] The instruction execution device based on the wait mechanism provided in the embodiment of the present application can execute the instruction execution method based on the wait mechanism provided in any embodiment of the present application, and has the corresponding functional modules and effects of the execution method.

[0136] Example 4

[0137] FIG4 shows a block diagram of an electronic device 40 that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0138] As shown in Figure 4, the electronic device 40 includes at least one processor 41, and a memory connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 41 can perform various appropriate actions and processes according to the computer program stored in the ROM 42 or the computer program loaded from the storage unit 48 into the RAM 43. Various programs and data required for the operation of the electronic device 40 can also be stored in the RAM 43. The processor 41, ROM 42 and RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0139] Multiple components in the electronic device 40 are connected to the I / O interface 45, including an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0140] The processor 41 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors for running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the instruction execution method based on the wait mechanism.

[0141] In some embodiments, the instruction execution method based on the wait mechanism can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the instruction execution method based on the wait mechanism described above can be performed. Alternatively, in other embodiments, the processor 41 can be configured to execute the instruction execution method based on the wait mechanism by any other appropriate means (for example, by means of firmware).

[0142] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] Computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of the present application, computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage medium can include but is not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory RAM, a read-only memory ROM, an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device or any suitable combination of the foregoing.

[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0146] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0147] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited business scalability of traditional physical hosts and virtual private servers (VPS).

[0148] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.

[0149] The above specific implementation methods do not constitute a limitation on the scope of protection of this application.

Claims

1. A method for executing instructions based on a wait mechanism, comprising: Obtain at least one instruction to be allocated in any loop body of the loop structure to be executed; According to the instruction type of each of the instructions to be assigned, selecting a target functional unit matching each of the instructions to be assigned from at least one functional unit for executing the instructions to be assigned; Determining an instruction scheduling result according to the instruction delay of each of the instructions to be assigned and based on the target functional units matched by each of the instructions to be assigned; According to the instruction scheduling result, based on the wait instruction generation mechanism, a wait scheduling instruction is generated and executed.

2. The method according to claim 1, wherein: The step of selecting, according to the instruction type of each instruction to be assigned, a target functional unit matching each instruction to be assigned from at least one functional unit for executing the instruction to be assigned comprises: Determining a function type of at least one function unit for executing the instruction to be assigned; According to the instruction type of each instruction to be allocated and based on the function type of each function unit, a target function unit matching each instruction to be allocated is selected from at least one function unit.

3. The method according to claim 1, wherein: The step of determining the instruction scheduling result according to the instruction delay of each of the instructions to be assigned and based on the target functional unit matched by each of the instructions to be assigned includes: Determining the instruction association relationship between the instructions to be assigned; An instruction scheduling result is determined according to the instruction association relationship, the target functional unit of each of the instructions to be assigned, and the instruction delay of each of the instructions to be assigned.

4. The method according to claim 1, wherein: The step of generating and executing a wait scheduling instruction based on a wait instruction generation mechanism according to the instruction scheduling result includes: Generate a wait instruction for each of the functional units according to the instruction scheduling result and based on the wait instruction generation mechanism; According to the wait instructions of each of the functional units, a wait scheduling instruction is generated and executed.

5. An instruction execution device based on a wait mechanism, comprising: An instruction acquisition module is configured to acquire at least one instruction to be allocated in any loop body of a loop structure to be executed; A functional unit matching module, configured to select, according to the instruction type of each instruction to be assigned, a target functional unit matching each instruction to be assigned from at least one functional unit configured to execute the instruction to be assigned; A scheduling result determination module, configured to determine the instruction scheduling result based on the instruction delay of each of the instructions to be assigned and the target functional unit matched by each of the instructions to be assigned; The scheduling instruction generation module is configured to generate and execute the wait scheduling instruction based on the wait instruction generation mechanism according to the instruction scheduling result.

6. The device according to claim 5, wherein: The functional unit matching module comprises: a function type determination unit, configured to determine a function type of at least one function unit configured to execute the instruction to be assigned; The function unit matching unit is configured to select a target function unit matching each of the instructions to be assigned from at least one function unit based on the instruction type of each of the instructions to be assigned and the function type of each of the function units.

7. The device according to claim 5, wherein: The scheduling result determination module includes: An instruction relationship determination unit, configured to determine an instruction association relationship between the instructions to be assigned; The scheduling result determination unit is configured to determine the instruction scheduling result according to the instruction association relationship, the target functional unit of each of the instructions to be assigned, and the instruction delay of each of the instructions to be assigned.

8. The device according to claim 5, wherein: The scheduling instruction generation module comprises: A wait instruction generating unit, configured to generate a wait instruction for each of the functional units according to the instruction scheduling result and based on the wait instruction generating mechanism; The scheduling instruction generating unit is configured to generate and execute a wait scheduling instruction according to the wait instruction of each of the functional units.

9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the instruction execution method based on the wait mechanism described in any one of claims 1 to 7.

10. A computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a processor to implement the instruction execution method based on the wait mechanism described in any one of claims 1 to 7 when the processor executes the instructions.

Citation Information

Patent Citations

  • Execution control method and device, embedded system, equipment and medium

    CN112698715A

  • Task execution method and device, storage medium and electronic equipment

    CN116107728A

  • Loop task execution method and device, chip and storage medium

    CN116795515A