Block instruction processing method and block instruction processor

CN120051765APending Publication Date: 2025-05-27HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101110.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the instruction fetching stage of existing block instruction processors, the entire block instruction contains multiple instructions, resulting in low instruction execution efficiency and inability to meet complex and huge computing requirements.

Method used

The block instructions are separated into two parts: block header and block body. The block header is used to express dependencies, and the block body is used for specific calculations. By obtaining and parsing the block header in advance, the block body can be pre-scheduled to improve execution efficiency.

Benefits of technology

Through fast fetching and preprocessing of block headers, efficient distribution and parallel execution of block bodies, the execution efficiency of block instructions is significantly improved and the performance of the processor is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051765A_ABST
    Figure CN120051765A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a block instruction processing method and a block instruction processor, and a block instruction comprises a block head and a block body. The method comprises the following steps: acquiring an ith block head; based on the first information indicated by the ith block head, distributing the ith block head to the jth block execution unit; the first information indicated by the ith block head comprises input register information and output register information of the ith block instruction corresponding to the ith block head; obtaining an ith block body through the jth block execution unit based on second information indicated by the ith block head; the ith block corresponds to the block of the ith block instruction; the second information indicated by the ith block head comprises the storage position of the ith block body; i is an integer greater than 1, and j is an integer greater than or equal to 1; and executing N microinstructions included in the ith block, wherein N is an integer greater than or equal to 1. By adopting the embodiment of the invention, the execution efficiency of the block instruction can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

A block instruction processing method and block instruction processor Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method for processing block instructions and a block instruction processor. Background Art

[0002] For a long time, the architectural development of general-purpose single-core central processing units (CPUs) has primarily focused on improving performance by increasing inter-instruction parallelism. Typically, superscalar processor architectures improve performance by issuing multiple instructions per cycle and using hardware logic units to resolve dependencies between instructions after parallelization. However, superscalar processor architectures often consume a large amount of chip area, and their increased power consumption far outweighs the performance gains, ultimately leading to poor processor energy efficiency. To address this, the industry has proposed block instruction processors, which can combine multiple instructions into a single block instruction to execute, effectively improving instruction parallelism and controlling power consumption. However, existing block instruction processors must fetch the entire block instruction before executing it. As mentioned above, since the entire block instruction contains multiple instructions and is large in size, existing block instruction processors consume a significant amount of time fetching the block instruction, resulting in a decrease in overall processor performance and an inability to meet today's increasingly complex and massive computational workloads.

[0003] Summary of the Invention

[0004] The embodiments of the present application provide a method for processing block instructions and a block instruction processor, which can greatly improve the efficiency of instruction execution.

[0005] The processing method of the block instruction provided in the embodiment of the present application can be executed by an electronic device, etc. An electronic device refers to a device that can be abstracted as a computer system, wherein an electronic device that supports the block instruction processing function may also be referred to as a block instruction processing device. The block instruction processing device may be a complete machine of the electronic device, such as: a smart wearable device, a smart phone, a tablet computer, a laptop computer, a desktop computer, a vehicle-mounted computer or a server, etc.; it may also be a system / device composed of multiple complete machines; it may also be a partial device in the electronic device, such as: a chip related to the block instruction processing function, such as a block instruction processor, a system chip (SoC), etc., which is not specifically limited in the embodiment of the present application. Among them, the system chip is also called a system on chip.

[0006] In a first aspect, an embodiment of the present application provides a method for processing a block instruction, wherein the block instruction includes a block header and a block body; the method includes: obtaining the i-th block header; distributing the i-th block header to the j-th block execution unit based on the first information indicated by the i-th block header; the first information indicated by the i-th block header includes the input register information and output register information of the i-th block instruction corresponding to the i-th block header; obtaining the i-th block body through the j-th block execution unit based on the second information indicated by the i-th block header, the i-th block body corresponding to the block body of the i-th block instruction; the second information indicated by the i-th block header includes the storage location of the i-th block body; i is an integer greater than 1, and j is an integer greater than or equal to 1; executing N microinstructions included in the i-th block body, where N is an integer greater than or equal to 1.

[0007] In the prior art, a block instruction processor must fetch the entire block instruction before it can subsequently distribute and execute the block instruction. However, because most block instructions contain multiple instructions and are relatively large, existing block processors consume a significant amount of time during the instruction fetch phase, thereby reducing overall instruction execution efficiency. Through the method provided in the first aspect, embodiments of the present application separate conventional block instructions into two parts: a block header and a block body. Each block instruction consists of a block header and a block body. The block header is primarily used to express the dependencies between block instructions, while the block body is primarily used to express specific calculations. The block body can be composed of multiple microinstructions. Based on this, embodiments of the present application can first continuously and rapidly retrieve the block headers of multiple block instructions and, based on the input register information and output register information indicated by each block header, distribute the block headers sequentially to the corresponding block execution units. This ensures that the block instructions to be executed by different block execution units have almost no dependencies, thereby ensuring reliable parallel execution of multiple block execution units. The block execution units can then quickly retrieve and execute the corresponding block body based on the block body storage location indicated by the block header. In this way, compared with the prior art solution in which the entire block instruction must be retrieved before subsequent distribution and execution can be carried out, resulting in low instruction execution efficiency, the embodiment of the present application is based on a fundamental improvement to the block instruction structure. By obtaining and parsing the block header in advance, the block is pre-scheduled, thereby greatly improving the execution efficiency of the block instruction.

[0008] In one possible implementation, the i-th block header is located in a first storage area, and the i-th block is located in a second storage area; the first storage area stores multiple block headers, and the second storage area stores multiple blocks; the first storage area and the second storage area are pre-divided storage areas in the memory.

[0009] In an embodiment of the present application, the block header and block body of a block instruction can be stored in different areas of memory. In some possible embodiments, two different storage areas can be pre-divided in the memory, one for storing the block headers of multiple block instructions, and the other for storing the block bodies of multiple block instructions. This allows the subsequent processor to efficiently retrieve the block header or block body directly from the corresponding storage area when executing the block instruction, thereby reducing the addressing range.

[0010] In addition, in some possible embodiments, the storage areas of the block header and the block body can be dynamically updated as the program runs, such as reducing the original storage area, or expanding the original storage area when there is insufficient space. This embodiment of the present application does not specifically limit this. For example, the first storage area corresponding to the i-th block header can be a storage area pre-divided in the memory, or the first storage area can also be a storage area obtained by performing a corresponding update on the basis of the pre-divided storage area. For another example, the second storage area corresponding to the i-th block body can be a storage area pre-divided in the memory, or the second storage area can also be a storage area obtained by performing a corresponding update on the basis of the pre-divided storage area.

[0011] Furthermore, the storage locations of multiple block headers in the first storage area can be adjacent to each other, meaning there is no gap or block insertion between the two preceding and succeeding block headers. This ensures the address continuity of the block headers in memory and improves the efficiency of reading block headers. It should be noted that if the disk is damaged, meaning that some of the first storage area may be invalid, the storage locations of multiple block headers can be adjacent to each other in the valid storage areas within the first storage area.

[0012] In one possible implementation, the method further includes: obtaining an i-1th block header, the i-1th block header corresponding to the block header of the i-1th block instruction; obtaining the i-th block header includes: based on third information indicated by the i-1th block header, determining the i-th block instruction executed after the i-1th block instruction and obtaining the i-th block header; the third information indicated by the i-1th block header includes a jump type of the i-1th block instruction.

[0013] In an embodiment of the present application, the block header in a block instruction can also be used to indicate the jump type of the block instruction. Based on this, the embodiment of the present application can quickly determine the block header of the next block instruction to be executed based on the jump type indicated in the block header, thereby continuously and quickly retrieving multiple block headers in sequence. Alternatively, the embodiment of the present application can directly predict and retrieve the next block header based on the jump type indicated by the current block header through a jump prediction method, thereby improving the efficiency of block header retrieval.

[0014] In one possible implementation, the storage location of the i-th block includes the storage locations of the N microinstructions in the i-th block; obtaining the i-th block based on the second information indicated by the i-th block header includes: obtaining the k-th microinstruction from the storage location corresponding to the k-th microinstruction; k is an integer greater than or equal to 1 and less than or equal to N; executing the N microinstructions included in the i-th block includes: executing the k-th microinstruction.

[0015] In this embodiment of the present application, the storage location of each microinstruction in the block can be obtained based on the block storage location indicated by the block header. When the block is executed, the microinstructions are sequentially retrieved from the corresponding storage location and executed. This makes microinstruction fetching faster and more convenient than conventional methods that require fetching the entire block of instructions before performing calculations, thereby effectively improving the execution efficiency of the block instructions.

[0016] In one possible implementation, the i-th block corresponds to N virtual registers, and the N virtual registers correspond one-to-one to the N microinstructions in the i-th block; each virtual register is used to store the execution result obtained after the corresponding microinstruction is executed.

[0017] In the embodiment of the present application, the multiple microinstructions executed sequentially in each block correspond one-to-one to multiple virtual registers. The execution result (i.e., output) of each microinstruction is implicitly written to the corresponding virtual register. In this way, each microinstruction does not require additional fields to express its output. That is, the length of each microinstruction can be fully used to express the opcode and operands, making the entire microinstruction more compact.

[0018] In one possible embodiment, the execution of the kth microinstruction includes: based on the decoding result of the kth microinstruction, determining that the input data of the kth microinstruction includes the execution result of the pth microinstruction among the N microinstructions; based on the relative distance kp between the kth microinstruction and the pth microinstruction, obtaining the execution result of the pth microinstruction from the pth virtual register, and obtaining the execution result of the kth microinstruction based on the execution result of the pth microinstruction; p is an integer greater than 1 or equal to 1 and less than k; and outputting the execution result of the kth microinstruction to the corresponding kth virtual register for storage.

[0019] In an embodiment of the present application, as described above, since a plurality of microinstructions executed in sequence in a block correspond to virtual registers for storing execution results. Therefore, the current microinstruction can obtain the execution result of the preceding microinstruction and complete the corresponding calculation by indexing the virtual register corresponding to the relative distance between the current microinstruction and the preceding microinstruction. Exemplarily, the length of the microinstruction can be 16 bits, of which 10 bits are used to express the opcode, and two 3 bits are used to express two operands (i.e., inputs), of which 3 bits can express a number from 0 to 7, which is equivalent to obtaining the execution result of any one of the 8 microinstructions before the current microinstruction.

[0020] In a possible implementation, the method further includes: if the k-th microinstruction terminates abnormally, submitting the current in-block state of the i-th block instruction to a system register; the system register is used to be accessed by the target program to obtain the in-block state and process the abnormal termination of the k-th microinstruction; wherein the in-block state includes: the storage location of the i-th block header, the storage location of the k-th microinstruction, and the execution result of the microinstruction executed before the k-th microinstruction.

[0021] In an embodiment of the present application, when a microinstruction in a block of instructions terminates abnormally, the current state of the block of instructions can be submitted to a system register so that the corresponding program can access the system register to obtain the state and handle the exception, thereby ensuring efficient and accurate exception handling. The system register can be a register shared between blocks.

[0022] In one possible implementation, the i-th block header is also used to indicate the attributes and types of the i-th block instruction; wherein the attributes include any one or more of the submission strategy, atomicity, visibility, and orderliness of the i-th block instruction, and the types include any one or more of fixed-point, floating-point, custom block, and accelerator call.

[0023] In the embodiment of the present application, the block header can also be used to express the attributes and types of the block instruction. In this way, the present application can quickly obtain a series of information about the block instruction in the first place after the block header is retrieved, so that the block can be executed more efficiently and submitted according to the strategy.

[0024] In a second aspect, an embodiment of the present application provides a block instruction processor, characterized in that the block instruction includes a block header and a block body; the block instruction processor includes a block header instruction fetch unit, a block distribution unit and multiple block execution units; the block header instruction fetch unit is used to obtain the i-th block header; the block distribution unit is used to distribute the i-th block header to the j-th block execution unit based on the first information indicated by the i-th block header; the first information indicated by the i-th block header includes the input register information and output register information of the i-th block instruction corresponding to the i-th block header; the j-th block execution unit is used to obtain the i-th block body based on the second information indicated by the i-th block header; the i-th block body corresponds to the block body of the i-th block instruction; the second information indicated by the i-th block header includes the storage location of the i-th block body; i is an integer greater than 1, and j is an integer greater than or equal to 1; the j-th block execution unit is also used to execute N microinstructions included in the i-th block body, where N is an integer greater than or equal to 1.

[0025] In one possible implementation, the i-th block header is located in a first storage area of ​​a memory, and the i-th block body is located in a second storage area of ​​the memory; a plurality of block headers are stored in the first storage area, and a plurality of blocks are stored in the second storage area; the first storage area and the second storage area are pre-divided storage areas in the memory.

[0026] In a possible implementation, the block header fetch unit is further configured to: obtain an i-1th block header, where the i-1th block header corresponds to a block header of an i-1th block instruction; the block header fetch unit is specifically configured to: determine, based on third information indicated by the i-1th block header, the i-th block instruction executed after the i-1th block instruction and obtain the i-th block header; the third information indicated by the i-1th block header includes a jump type of the i-1th block instruction.

[0027] In one possible embodiment, the storage location of the i-th block includes the storage locations of the N microinstructions in the i-th block; the j-th block execution unit is specifically used to: obtain the k-th microinstruction from the storage location corresponding to the k-th microinstruction; k is an integer greater than or equal to 1 and less than or equal to N; and execute the k-th microinstruction.

[0028] In one possible implementation, the i-th block corresponds to N virtual registers, and the N virtual registers correspond one-to-one to the N microinstructions in the i-th block; each virtual register is used to store the execution result obtained after the corresponding microinstruction is executed.

[0029] In one possible embodiment, the j-th block execution unit is specifically used to: determine, based on the decoding result of the k-th microinstruction, that the input data of the k-th microinstruction includes the execution result of the p-th microinstruction among the N microinstructions; obtain the execution result of the p-th microinstruction from the p-th virtual register based on the relative distance kp between the k-th microinstruction and the p-th microinstruction, and obtain the execution result of the k-th microinstruction based on the execution result of the p-th microinstruction; p is an integer greater than 1 or equal to 1 and less than k; and output the execution result of the k-th microinstruction to the corresponding k-th virtual register for storage.

[0030] In one possible implementation, the j-th block execution unit is further configured to: submit the current in-block state of the ith block instruction to a system register if the k-th microinstruction terminates abnormally; the system register is configured to be accessed by a target program to obtain the in-block state and process the abnormal termination of the k-th microinstruction; wherein the in-block state includes: the storage location of the ith block header, the storage location of the k-th microinstruction, and the execution result of the microinstruction executed before the k-th microinstruction.

[0031] In one possible implementation, the i-th block header is also used to indicate the attributes and types of the i-th block instruction; wherein the attributes include any one or more of the submission strategy, atomicity, visibility, and orderliness of the i-th block instruction, and the types include any one or more of fixed-point, floating-point, custom block, and accelerator call.

[0032] It should be understood that the power supply equipment provided in the second aspect of this application is consistent with the technical solution of the first aspect of this application. Its specific content and beneficial effects can be referred to the power supply equipment provided in the above-mentioned first aspect, and will not be repeated here.

[0033] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the functions involved in the block instruction processing method flow provided in the first aspect above.

[0034] In a fourth aspect, an embodiment of the present application provides a computer program comprising instructions, which, when executed by a computer, enables the computer to execute the functions involved in the block instruction processing method flow provided in the first aspect above.

[0035] In a fifth aspect, an embodiment of the present application provides a chip, which includes a block instruction processor as described in any one of the second aspects above, and is used to implement the functions involved in the process of a block instruction processing method provided in the first aspect above. In one possible design, the chip also includes a memory, which is used to store the program instructions and data necessary for the block instruction processing method, and the block instruction processor is used to call the program code stored in the memory to execute the functions involved in the process of a block instruction processing method provided in the first aspect above. The chip can constitute a chip system, and the chip system can also include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments of the present application or the background technology will be described below.

[0037] FIG1a is a schematic diagram of the organization of a block instruction program provided in an embodiment of the present application.

[0038] FIG1b is a schematic diagram of a block instruction structure provided in an embodiment of the present application.

[0039] FIG2 is a schematic diagram of the structure of a block header provided in an embodiment of the present application.

[0040] FIG3 is a schematic structural diagram of a block provided in an embodiment of the present application.

[0041] FIG4 is a schematic diagram of a block header pointer and a microinstruction pointer provided in an embodiment of the present application.

[0042] FIG5 is a flow chart of a method for processing a block instruction provided in an embodiment of the present application.

[0043] FIG6 is a schematic diagram of a method for obtaining a block header based on a jump type provided in an embodiment of the present application.

[0044] FIG7 is a schematic diagram of an architectural state of a block instruction provided in an embodiment of the present application.

[0045] FIG8 is a schematic diagram of an assembly line of a block head and a block body provided in an embodiment of the present application.

[0046] FIG9 is a schematic diagram of the structure of a block processor provided in an embodiment of the present application.

[0047] FIG10 is a schematic diagram of the structure of another block processor provided in an embodiment of the present application.

[0048] FIG11 is a schematic diagram of the structure of another block processor provided in an embodiment of the present application.

[0049] FIG12 is a schematic diagram of a ring network provided in an embodiment of the present application.

[0050] FIG13 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] The embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0052] The terms "first," "second," "third," and "fourth," as well as "first," "second," "third," and "fourth" in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0053] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0054] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0055] As used in this specification, the terms "component," "module," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0056] First, some terms in this application are explained to facilitate understanding by those skilled in the art.

[0057] (1) Superscalar processor architecture improves processor performance by issuing multiple instructions per beat, and uses hardware logic units to resolve the dependencies between instructions after parallel execution. The algorithm for out-of-order execution of instructions includes Tomasulo's algorithm. The hardware implementation of Tomasulo's algorithm includes reorder buffer, out-of-order issue queue, reservation station, register renaming, branch predication, memory read and write prediction, etc. These modules not only increase the difficulty of hardware design verification, but also consume a lot of chip area, resulting in an increase in power consumption far exceeding the performance improvement, which ultimately makes the processor energy efficiency worse. Based on this, the industry has successively proposed solutions such as very long instruction word (VLIW) architecture and block instruction processor, trying to obtain parallelism between instructions at a higher level and larger granularity to improve processor performance.

[0058] (2) The VLIW architecture relies on the compiler to combine instructions with no dependencies into a single, very long instruction to run, which can improve instruction parallelism to a certain extent. However, because the compiler not only needs to consider the dependencies between instructions but also needs to complete the scheduling problem between instructions, the VLIW processor inevitably increases the complexity of compiler design. Due to the difficulty in designing the compiler complexity, and the incompatibility between the VLIW instruction set and the existing main graph instruction set, which hinders compatibility and restricts the development of the business ecosystem, VLIW processors are not used by mainstream manufacturers.

[0059] (3) Block instruction processors, or instruction block processors (referred to as block processors), can perform parallel operations at the granularity of instruction blocks. Block instructions, or instruction blocks, are composed of multiple instructions. Since the block instruction processor improves the granularity of instruction expression, the CPU hardware can significantly reduce the complexity and power consumption of implementation, thereby improving performance and energy efficiency. Among them, the block instruction is expressed internally as a dependency graph between multiple instructions, while the block instruction is still expressed externally in the traditional jump method. Compared with VLIW, the block instruction processor tends to put multiple dependent instructions into one block, and put multiple non-dependent instructions into different blocks.

[0060] To facilitate understanding of the embodiments of the present application, the following further analyzes and presents the specific technical problems to be solved by this application. As described above, existing block instruction processors can improve instruction parallelism at the granularity of instruction blocks, thereby improving processor performance. However, because an entire block instruction contains multiple instructions and is relatively large in size, existing block instruction processors must fetch the entire block instruction before performing subsequent inter-instruction parallelism analysis and issuance execution. This causes existing block instruction processors to consume a large amount of time during the instruction fetch phase, seriously affecting overall instruction execution efficiency. Therefore, in order to address the problem that current block instruction technology cannot meet actual needs, the technical problems to be solved by this application include the following aspects: separating block instructions into two parts: a block header and a block body, wherein the block header is mainly used to express the dependencies between block instructions, and the block body is used to express the specific calculations of the block instruction. Furthermore, through continuous and rapid instruction fetching and preprocessing of the block header, the dependencies and parallelism between block instructions are obtained, so that the corresponding blocks can be accurately and efficiently distributed to different pipelines for parallel execution, effectively improving the execution efficiency of the block instruction.

[0061] First, the structure of the block instruction provided by this application will be explained.

[0062] Please refer to Figure 1a, which is a schematic diagram of the organization of a block instruction program provided in an embodiment of the present application. As shown in Figure 1a, in the block instruction scenario, the computer program needs to organize the logic of the program into binary codes with block granularity. The block processor then executes the program at the granularity of block instructions by parsing the block instructions. A fixed-length reduced instruction set computer (RISC) instruction or complex instruction set computer (CISC) instruction of a traditional CPU can only complete one computing operation, while a block instruction of a block processor can complete multiple complex computing operations.

[0063] Furthermore, as shown in Figure 1a, the present application designs block instructions into two parts: a block header and a block body (block body / payload). The program needs to separate the block header and the block body and store them in different memory areas. Please also refer to Figure 1b, which is a schematic diagram of a block instruction structure provided by an embodiment of the present application. As shown in Figure 1b, block instruction 0 is composed of block header 0 and block body 0, block instruction 1 is composed of block header 1 and block body 1, block instruction 2 is composed of block header 2 and block body 2, and so on. Among them, block header 0, block header 1, block header 2, block header 3, and block header 4 can be stored in the storage area of ​​the block header, and block body 0, block body 1, block body 2, block body 3, and block body 4 can be stored in the storage area of ​​the block body. Among them, the storage area of ​​the block header and the storage area of ​​the block body can be two fixed storage areas divided out in the memory.

[0064] Specifically, a block header can be used to express the attributes and type of the block instruction, input registers, output registers, a pointer to the block header of the next block instruction, a pointer to the block body of the current block instruction, and so on. Optionally, the compiler often places block headers in adjacent locations in memory. For example, the memory addresses of block headers 0, 1, 2, 3, and 4 shown in Figure 1b can be adjacent to each other, so that subsequent memory addresses can be read continuously, improving the efficiency of retrieval of block headers.

[0065] Specifically, a block can be used to express the specific calculation of a block instruction and can be composed of multiple microinstructions. As shown in Figure 1a, block header 0 may involve more calculation operations and contain more microinstructions, thus occupying more storage space. Optionally, if the calculation operations performed by two block instructions are exactly the same, then the blocks of these two block instructions can be the same. For example, if the calculation operations of block instruction 0 and block instruction 4 are exactly the same, then the block of block instruction 4 can also be block 0. For another example, if the calculation operation of block instruction 4 is the same as part of the calculation operation of block instruction 0, then the block of block instruction 4 can share part of block 0. It should be noted that a group of block headers represents the control and data flow graph of the program, which is the "skeleton" of the program execution. A group of blocks represents the "flesh and blood" of the program.

[0066] Below, the specific structures and corresponding functions of the block header and block body in the block instruction provided by this application will be explained in detail.

[0067] Please refer to Figure 2, which is a schematic diagram of the structure of a block header provided by an embodiment of the present application. As shown in Figure 2, the block header of a block instruction can be a fixed-length RISC instruction, such as a 128-bit RISC instruction, which can include fields such as the type and attributes of the block instruction, input register bitmask, output register bitmask, jump pointer offset, and block pointer offset, as shown below.

[0068] Block instruction types include calculation types such as fixed-point, floating-point, custom blocks, and accelerator calls. Custom blocks allow users to add custom instructions. Alternatively, in this case, a block instruction can actually be hardware such as an accelerator, used to perform operations such as compression and decompression. In this case, there is no actual block in the block instruction, but the input and output can be abstracted into a block header.

[0069] Block instruction attributes: including the block instruction submission strategy, such as the number of times the block instruction is repeated, whether the execution of the block instruction must be atomic / visible / ordered, etc.

[0070] Input register bitmask: Used to indicate the input registers of a block instruction. A block instruction can have up to 32 registers as input. The names (identities, IDs) of these 32 input registers can be expressed in bitmask format.

[0071] Output Register Bitmask: Used to indicate the output registers of a block instruction. A block instruction can have up to 32 registers as output. The IDs of these 32 output registers can be expressed in bitmask format.

[0072] It should be noted that these 32 input / output registers (R0-R31) are all physical registers (i.e., general-purpose registers) shared between blocks. If a block instruction has exactly 32 inputs, then every bit in its 32-bit bit mask is 1. If a block instruction has exactly 16 outputs, then the corresponding 16 bits in its 32-bit bit mask are 1 (for example, if 11110000110000011101111001010001, then the output registers of this block instruction are the 16 general-purpose registers R0, R4, R6, R9, R10, R11, R12, R14, R15, R16, R22, R23, R28, R29, R30, and R31). In other words, the bits in the general-purpose registers used for the block instruction's input / output are 1. Optionally, if there are only 16 general registers in the CPU, each block instruction also has a maximum of 16 input / output registers, and the input / output registers of each block instruction can be indicated by a 16-bit bit mask. This embodiment of the present application does not make specific limitations on this.

[0073] Jump Pointer Offset: This indicates the storage location of the next block header to be executed after the current block completes. Specifically, the block header pointer is expressed as an offset. Adding the jump pointer offset to the current block header pointer yields the next block header pointer, indicating the storage location of the next block header.

[0074] Block pointer offset: This indicates the storage location of the block of instructions in the current block. Specifically, the block pointer is expressed as an offset. Adding the block pointer offset to the previous block pointer yields the current block pointer, which in turn gives the current block's storage location. This is typically the starting location of the current block (i.e., the location of the first microinstruction). The block's end location can be determined based on the block size (i.e., the number of microinstructions).

[0075] Please refer to Figure 3, which is a schematic diagram of the structure of a block provided by an embodiment of the present application. As shown in Figure 3, the block can be composed of a series of microinstructions such as microinstruction 0, microinstruction 1, microinstruction 2, microinstruction 3...microinstruction n, and each microinstruction is executed in sequence. It should be noted that different types of block instructions can define different block formats. Taking the more basic standard block instruction as an example, the microinstruction in the block can be a 16-bit fixed-length RISC instruction. For standard block instructions, each microinstruction in the block can be an instruction with at most two inputs and one output.

[0076] As shown in Figure 3, the input of a block instruction can be read from the output registers of other block instructions through the get instruction in the block body. For example, the first microinstruction (get R0) in Figure 3 reads data written to general register R0 by other block instructions. Another example is the second microinstruction (get R1) in Figure 3 reads data written to general register R1 by other block instructions. Another example is the sixth microinstruction (get R2) in Figure 3 reads data written to general register R2 by other block instructions.

[0077] As shown in Figure 3, the output of a block instruction can be written to the corresponding output register through the set instruction in the block body, that is, written back to one or more of the 32 physical registers mentioned above. In summary, block instructions can access general registers through explicit get / set instructions in microinstructions to obtain the output of other block instructions or output their own execution results. In addition, since the input / output registers of multiple block instructions may overlap, in order to ensure the parallelism of block instructions, as shown in Figure 3, each block instruction can temporarily write the execution result to the shadow register corresponding to the general register when outputting. Subsequent block instructions can also obtain input from the corresponding shadow register through the get instruction. Finally, each block instruction can submit the execution results temporarily stored in the shadow register to the corresponding physical register in sequence.

[0078] Optionally, in addition to the above-mentioned special get instructions and set instructions, the inputs of other microinstructions in the block can mostly come from the outputs of the preceding microinstructions in the current block. In an embodiment of the present application, the execution result of each microinstruction in the block can be implicitly written to its corresponding virtual register. Then, the current microinstruction can use the relative distance between the instructions to index the execution result of the preceding microinstruction temporarily stored in the virtual register as input. In this way, in the microinstruction encoding, each microinstruction does not need to express its own output register, that is, each microinstruction can only express its own opcode and at most two input registers (that is, at most two relative distances). Taking a fixed-length RISC instruction with a microinstruction of 16 bits as an example, its structure can be shown in Table 1 below.

[0079] Table 1

[0080] Bit 15-13Bit 12-10Bit 9-0Link (relative distance)0Link (relative distance)1Opcode (operation code)

[0081] As shown in Table 1 above, taking a 16-bit microinstruction as an example, 10 bits can be used to express the opcode, and two 3-bit bits can be used to express two input registers (i.e., two relative distances). At this time, the relative distance is limited to 1-8, that is, the current microinstruction can only index the execution results of the 1st to 8th microinstructions before the current microinstruction as input. For example, the two inputs of the third microinstruction (add1 2) in Figure 3 come from the execution results of two instructions (i.e., the 1st instruction and the 2nd microinstruction) whose relative distances from the third microinstruction are 1 and 2 respectively. Optionally, if there is only one input, the remaining 6 bits after the opcode can all be used to express one input register, and the relative distance is limited to 1-64.

[0082] For example, the block instructions may be compiled as follows:

[0083] get R0

[0084] get R1

[0085] add T#1 T#2

[0086] const #2

[0087] s1l T#1 T#2

[0088] get R2

[0089] add T#2 T#1

[0090] ld [T#2,#0]

[0091] set T#1R3

[0092] Taking the third microinstruction (addT#1 T#2) in the above block as an example, T#1 represents the virtual register corresponding to the microinstruction with a relative distance of 1 from the current microinstruction (addT#1 T#2), which stores the execution result of the microinstruction get R0. T#2 represents the virtual register corresponding to the microinstruction with a relative distance of 2 from the current microinstruction (addT#1 T#2), which stores the execution result of the microinstruction get R1. Therefore, the actual calculation performed by the third microinstruction (addT#1 T#2) is to add the data in registers R0 and R1.

[0093] As described above, based on the structure of block header and block body, the block instruction provided by this application also defines two levels of program pointers. Please refer to Figure 4, which is a schematic diagram of a block header pointer and microinstruction pointer provided by an embodiment of this application. As shown in Figure 4,

[0094] The block program counter (BPC) records the location of the block header of the currently executing instruction block. By adding the jump pointer offset indicated by the current block header to the current block header pointer, we can obtain the next block header pointer, which indicates the memory location of the next block header to be executed.

[0095] The temporal program counter (TPC) is used to record the location of the currently executing microinstruction. Since a block generally contains multiple microinstructions, the location of each microinstruction in the block can be represented by the TPC.

[0096] In summary, the combination of BPC and TPC can point to the microinstruction currently being executed in the current block of instructions. After each microinstruction is executed, the TPC moves to the next microinstruction. Once all microinstructions in the current block have been executed, the BPC moves to the next block header.

[0097] Based on the detailed description of the block instruction structure in the embodiments corresponding to Figures 1a-4 above, an embodiment of the present application further provides a method for processing a block instruction. Please refer to Figure 5, which is a schematic flow chart of a method for processing a block instruction provided in an embodiment of the present application. This method can be applied to a processor in an electronic device, and specifically to a block instruction processor. As shown in Figure 5, the method may include the following steps S501-S504.

[0098] Step S501: Get the i-th block header.

[0099] Specifically, the block processor sequentially obtains the i-th block header. The i-th block header is the block header of the i-th block instruction, and i is an integer greater than or equal to 1. It should be noted that in the embodiment of the present application, the instruction fetching order of the block headers should be the theoretical execution order of the block instructions during the program execution, that is, the order in which each block instruction finally submits the execution results in sequence. In addition, the serial numbers such as block header 0, block header 1, block header 2, and block header 3 shown in all the figures of the embodiment of the present application only represent the instruction fetching order of the block headers, and have nothing to do with the storage location of the block headers in the memory.

[0100] Optionally, before sequentially obtaining the i-th block header, the block processor also sequentially obtains the i-1th block header, where the i-1th block is the block header for the i-1th block instruction. Based on the third information indicated in the i-1th block header, the block processor can determine the i-th block instruction to be executed after the i-1th block instruction and obtain the corresponding i-th block header. The third indication information indicated by the i-1th block header may include the jump type of the i-1th block instruction. In short, the block processor can quickly determine the next block instruction to be executed based on the jump type indicated by the current block header and obtain the corresponding block header.

[0101] Optionally, in some embodiments of the present application, various jump types indicated in the block header may be as shown in Table 2 below.

[0102] Table 2

[0103]

[0104]

[0105] As shown in Table 2 above, the jump type indicated by the block header may include the above-mentioned deferral (FALL), direct jump (DIRECT), call (CALL), conditional jump (COND), indirect jump (IND), indirect call (INDCALL), return (RET) and concatenation (CONCAT), etc., and may also include any other possible jump type, which is not specifically limited in the embodiments of the present application. As shown in Table 2 above, when the jump type is deferral, direct jump, or call, it is not necessary to calculate through the microinstructions within the block, and the jump can be completed only by parsing the block header. When the jump type is indirect jump, conditional jump, indirect call, and return, it is often necessary to calculate the address of the next block header through the microinstructions within the block. For example, the microinstruction SETBPC: is used to set the BPC absolute address of the next block instruction; for another example, the microinstruction SETBPC.COND: calculates and determines the block header address of the next block instruction among the two possible jump block header addresses.

[0106] Please also refer to FIG. 6 , which is a schematic diagram of a method for obtaining a block header based on a jump type provided in an embodiment of the present application.

[0107] For example, as shown in FIG6 , the jump type indicated in block header 0 is a direct jump (DIRECT). Based on this, the block processor can directly determine that the block instruction pointed to by BNEXT in the block header is the next block instruction to be executed, and obtain the corresponding block header (i.e., block header 1 shown in FIG6 ).

[0108] For example, as shown in FIG6 , the jump type indicated in block header 1 is an indirect jump (IND). As can be seen from Table 2 above, the block processor theoretically needs to first obtain the block body according to the block body address indicated in block header 1, and then calculate the address of the next block instruction to be executed (i.e., the address of the block header) by the microinstruction SETBPC in the block body, thereby obtaining the corresponding block header (i.e., block header 2 shown in FIG6 ). However, in some embodiments of the present application, in order to improve the efficiency of block header instruction fetching, the block processor can only determine the next block instruction to be executed through jump prediction based on the jump type indicated in the current block header 1, thereby directly obtaining block header 2 without having to obtain and execute the corresponding microinstructions.

[0109] For example, as shown in Figure 6 , the jump type indicated in block header 2 is a conditional jump (COND). As can be seen from Table 2 above, the block processor theoretically needs to first obtain the block body based on the block body address indicated in block header 2, and then use the microinstruction SETBPC.COND in the block body to calculate the address of the next block instruction to be executed (i.e., the address of the block header), thereby obtaining the corresponding block header (i.e., block header 3 shown in Figure 6 ). Similarly, to save time and improve efficiency, the block processor can directly predict the next block instruction to be executed based on the jump type indicated in block header 2 through jump prediction, thereby directly obtaining block header 3.

[0110] It should be noted that jump prediction is generally highly accurate. Therefore, when sequentially retrieving block headers, embodiments of the present application can typically use jump prediction to quickly and efficiently retrieve the next block header based on the jump type indicated in the current block header, thereby significantly improving the efficiency of block header instruction fetching and, in turn, the overall execution efficiency of the block instructions. It is understood that if, during subsequent microinstruction execution, the original jump prediction result is incorrect, the block processor can re-retrieve the correct block header based on this calculation result. This will not be further described here.

[0111] Step S502: Distribute the i-th block header to the j-th block execution unit based on the first information indicated by the i-th block header.

[0112] Specifically, the block processor may include multiple block execution units. The block processor may distribute the i-th block header to the corresponding j-th block execution unit based on the first information indicated in the i-th block header. The j-th block execution unit is one of the multiple block execution units included in the block processor, and j is an integer greater than or equal to 1.

[0113] Optionally, the first information indicated by the i-th block header may include the input register information and output register information of the i-th block instruction to which the i-th block header belongs. The input register information and output register information may be the input register Bitmask and output register Bitmask in the block header structure shown in FIG. 2 above.

[0114] It is understandable that the input register information and output register information can reflect the dependency relationship between block instructions. The dependency relationship between the i-th block instruction and other block instructions to be executed on the j-th block execution unit can be relatively large, or even the largest. In other words, the block processor can distribute the block headers to the block execution units with the most dependencies as much as possible based on the input register information and output register information indicated in each block header, so that the block instructions to be executed between different block execution units have almost no dependency, that is, there is almost no data interaction between different block execution units, thereby ensuring the reliable parallel operation of multiple block execution units.

[0115] Step S503 : obtaining the i th block based on the second information indicated by the i th block header through the j th block execution unit.

[0116] Specifically, the j-th block execution unit in the block processor may obtain and execute the i-th block based on the second information indicated by the i-th block header, wherein the i-th block is a block of the i-th block instruction, and the i-th block may include one or more microinstructions.

[0117] The second information indicated by the i-th block header may include the storage location of the i-th block, for example, the block pointer offset in the block header structure shown in FIG. Based on the block pointer offset, the starting location of the block can be determined. Combined with the number of microinstructions in the block and the length of each microinstruction, the storage location of each microinstruction can be determined. In this way, the j-th block execution unit can sequentially retrieve each microinstruction in the i-th block from the corresponding storage location. For example, the k-th microinstruction in the i-th block can be retrieved from the corresponding storage location.

[0118] Optionally, the i-th block header can be located in a first storage area, and the i-th block body can be located in a second storage area. The first storage area and the second storage area can be pre-divided storage areas in the memory. The first storage area can store block headers of multiple block instructions, and the second storage area can store block bodies of multiple block instructions, that is, the block headers and block bodies of block instructions can be stored in different areas of the memory. In addition, the storage locations of the block headers of multiple block instructions in the first storage area can be adjacent to each other, that is, there will be no gap or a block body inserted between the two block headers, thereby ensuring the address continuity of the block headers in the memory and improving the efficiency of reading the block headers.

[0119] Step S504 , executing the N microinstructions included in the i-th block.

[0120] Specifically, the j-th block execution unit in the block processor executes N microinstructions included in the i-th block, where N is an integer greater than or equal to 1. The N microinstructions may be all or part of the microinstructions included in the i-th block. The block processor may execute the N microinstructions sequentially, for example, sequentially executing the k-th microinstruction, the k+1-th microinstruction, and so on of the N microinstructions. Wherein, k is an integer greater than or equal to 1 and less than or equal to N.

[0121] Optionally, based on the description of the embodiment corresponding to FIG. 3 above, the i-th block instruction may correspond to N virtual registers, or in other words, the i-th block may correspond to N virtual registers, and the N virtual registers may correspond one-to-one with the N microinstructions in the i-th block. Thus, the j-th block execution unit executing the k-th microinstruction in the i-th block may include: decoding the k-th microinstruction, and determining, based on the decoding result, that the input data of the k-th microinstruction includes the execution result of the p-th microinstruction among the N microinstructions, i.e., the input of the k-th microinstruction comes from the output of the p-th microinstruction. Where p is an integer greater than or equal to 1 and less than k. Then, the j-th block execution unit may obtain the execution result of the p-th microinstruction from the p-th virtual register among the N virtual registers based on the relative distance kp between the k-th microinstruction and the p-th microinstruction, and calculate the execution result of the k-th microinstruction based on the execution result of the p-th microinstruction as input. Finally, the j-th block execution unit can output the calculated execution result of the k-th microinstruction to the corresponding k-th virtual register for storage so as to be indexed by other subsequent microinstructions.

[0122] Further, please refer to Figure 7, which is a schematic diagram of the architectural state of a block instruction provided in an embodiment of the present application. As shown in Figure 7, the block processor defines two levels of architectural state: the upper level is the shared architectural state between blocks, and the lower level is the private architectural state within the block.

[0123] Inter-block shared architectural state refers to the processor state shared by different instruction blocks, also known as global state. As shown in Figure 7, inter-block shared architectural state primarily includes the state of general registers (R0-R31), BPC, and system registers (BSTATE.EXT).

[0124] The private architectural state within a block is the local state defined within each instruction in the block. As shown in Figure 7, the private architectural state within a block mainly includes the states of the virtual registers and TPC corresponding to the microinstructions within the block.

[0125] It should be noted that the block instruction defines a series of changes to the architectural state. Specifically, when the block instruction is decoded, the state within the block is created. At the end of the execution of the current block instruction, that is, after all microinstructions in the current block have been executed, the state within the block will be released, including clearing the virtual registers corresponding to all microinstructions in the block. Based on this, the block processor in the embodiment of the present application can execute programs and jumps according to the granularity of the block instruction, and maintain accurate execution status through the two levels of architectural state (Global-Local) defined by the block instruction.

[0126] Generally speaking, the execution status of any block instruction can include the following two types:

[0127] (1) Block Commit Success: All microinstructions in the block are executed normally without exception or interruption, and the status in the block is successfully committed, that is, the status in the block is released smoothly.

[0128] (2) Block Exception Terminate: The microinstructions in the block fail to execute normally, and the block state (BSTATE) points to the microinstruction that terminated abnormally and the preceding text of the microinstruction. The block processor can package and copy the block state to the system register so that the corresponding program can access the system register to obtain the block state and handle the exception, thereby ensuring efficient and accurate exception handling. Optionally, when a microinstruction in a block instruction terminates abnormally (for example, a memory access error or an external interrupt is received when performing an addition operation, causing the block instruction to terminate abnormally), the block state of the block instruction includes the TPC of the current abnormally terminated microinstruction and the execution results (T#1-T#8) of multiple microinstructions (generally 8) before the microinstruction. Optionally, it can also include the results that have been written (set) to the shadow register in the current block instruction.

[0129] For example, if the k-th microinstruction in the i-th block terminates abnormally, the current in-block state of the i-th block instruction can be submitted to the system register. The current in-block state of the i-th block instruction can include: the storage location of the i-th block header (i.e., BPC), the storage location of the k-th microinstruction (i.e., TPC), and the execution results of the microinstructions executed before the k-th microinstruction. A subsequent target program (e.g., a program for exception handling) can access the system register to obtain the in-block state and handle the abnormal termination of the k-th microinstruction.

[0130] As mentioned above, regardless of whether a block instruction is submitted normally or terminated abnormally (or abnormally exited), the state within the block will be released (or cleared). If you need to obtain the status of the abnormal exit of the block instruction, you need to access the system register.

[0131] In summary, please refer to Figure 8, which is a pipeline diagram of a block header and a block provided in an embodiment of the present application. As shown in Figure 8, each block header can be a fixed-length RISC instruction, and the block processor can quickly obtain and parse the block header in sequence in the form of a pipeline, thereby obtaining information such as the jump type, input / output registers, and block storage location indicated in the block header. Obviously, the parsing of the block header is much earlier than the calculation of the block. Based on this, the block processor in the embodiment of the present application can rely on the early execution of the block header to pre-allocate the execution resources of the block. This may include: pre-scheduling of the block, allocation of block input and output resources, out-of-order execution and reordering of the block, etc.

[0132] For example, as shown in FIG8 , the block processor sequentially obtains block header 0, block header 1, block header 2, block header 3, etc., and distributes the block headers to the pipelines with the most dependencies based on the information obtained by parsing the block headers. For example, block header 0 is distributed to block pipeline 1 (corresponding to the first block execution unit), block header 1 is distributed to block pipeline 4 (corresponding to the fourth block execution unit), block header 2 is distributed to block pipeline 1 (corresponding to the first block execution unit), block header 3 is distributed to block pipeline 2 (corresponding to the second block execution unit), block header 4 is distributed to block pipeline 1 (corresponding to the first block execution unit), block header 5 is distributed to block pipeline 3 (corresponding to the third block execution unit), and so on. No further details will be given here. Among them, there can be dependencies between blocks 0, 2, and 4 on block pipeline 1, and there can be dependencies between blocks 1, 6, 7, and 8 on block pipeline 4. However, there can be no or minimal dependencies between the blocks on block pipeline 1 and those on block pipeline 4. This facilitates the smooth parallel operation of multiple pipelines. Subsequently, each block pipeline sequentially obtains and executes the block corresponding to the block header. For example, on block pipeline 1, the first block execution unit obtains and executes block 0, block 4 and block 2 in sequence; on block pipeline 2, the second block execution unit obtains and executes block 9, block 3, block 12 and block 1 in sequence; on block pipeline 3, the third block execution unit obtains and executes block 5, block 14 and block 13 in sequence; on block pipeline 4, the fourth block execution unit obtains and executes block 1, block 6, block 7 and block 8 in sequence.

[0133] As shown in Figure 8, in some embodiments of the present application, considering the priority and size of the blocks, out-of-order execution of blocks can be scheduled on each pipeline to improve the efficiency of block instruction execution. For example, on block pipeline 1, block 4 can be executed before block 2. For example, on block pipeline 3, block 14 can be executed before block 13, and so on. This embodiment of the present application does not specifically limit this.

[0134] As shown in Figure 8, the block processor in this embodiment of the present application includes at least one block header pipeline (for continuously acquiring block headers) and multiple block body pipelines (for executing blocks in parallel) when executing block instructions, thereby achieving out-of-order parallel execution of block instructions. In this way, the block processor in this embodiment of the present application can manage block instructions at the granularity of block headers, achieving out-of-order issuance and parallel computing of block instructions, and completing large-scale instruction parallel computing with lower complexity.

[0135] Based on the description of the above-mentioned block instruction processing method embodiment, the embodiment of the present application also provides a block processor (Block Core). Please refer to Figure 9, which is a structural diagram of a block processor provided by the embodiment of the present application. As shown in Figure 9, the block processor 10 may include a block header instruction fetch unit 101, a block distribution unit 102 and multiple block execution units, for example, including a block execution unit 31, a block execution unit 32, a block execution unit 33, etc. Among them, the block header instruction fetch unit 101 is connected to the block distribution unit 102, and the block distribution unit 102 is connected to multiple block execution units such as the block execution unit 31, the block execution unit 32, the block execution unit 33, etc.

[0136] The block header fetch unit 101 is configured to sequentially retrieve the i-th block header from the memory. i is an integer greater than or equal to 1. Optionally, in scenarios such as high-performance computing, in order to further improve the efficiency of fetching block headers, a block header cache may be provided in the block processor 10. In this way, the block header fetch unit 101 may directly sequentially retrieve the i-th block header from the block header cache. Optionally, as described above, the block header fetch unit 101 may determine the address of the next block header (i.e., the location pointed to by the BPC) based on the jump type indicated in the i-1-th block header through jump prediction and retrieve the next block header, thereby improving the efficiency of fetching block headers.

[0137] The block distribution unit 102 is configured to distribute the i-th block header to the j-th block execution unit based on the first information indicated by the i-th block header. The j-th block execution unit is, for example, one of the block execution units 31, 32, or 33 shown in FIG9 . As shown above, the first information may include input register information and output register information of the i-th block instruction to which the i-th block header belongs. The i-th block header has many or even the most dependencies with other block headers previously distributed to the j-th block execution unit, resulting in almost no dependency between the block instructions to be executed by different block execution units, i.e., almost no data interaction between different block execution units, thereby ensuring reliable parallel operation of multiple block execution units.

[0138] The jth block execution unit (e.g., block execution unit 31) is configured to retrieve and execute the i-th block based on the second information indicated by the i-th block header. The second information includes the storage location of the i-th block (i.e., the location indicated by the TPC). Similarly, in scenarios such as high-performance computing, to improve block fetch efficiency, a block cache may be provided in the block processor 10, and the block execution units 31, 32, and 33 may retrieve the i-th block directly from the block cache.

[0139] As described above, the execution results written (set) by each block instruction can be temporarily stored in the shadow register. Finally, each block execution unit in the block processor 10 can submit the execution results of the corresponding block instructions temporarily stored in the shadow register to the corresponding general register in sequence according to the instruction fetch order of the block header.

[0140] In summary, the block processor provided in the embodiment of the present application is a novel CPU implementation based on block instructions. Unlike other block processors, the embodiment of the present application separates block instructions into two parts: a block header and a block body. Therefore, the block processor in the embodiment of the present application has at least two instruction fetch units, namely a block header instruction fetch unit 101 and a block execution unit 31, which are used to fetch block header instructions and block body instructions from the locations pointed to by the BPC and TPC, respectively, thereby greatly improving the overall execution efficiency of block instructions.

[0141] Further, please refer to Figure 10, which is a schematic diagram of the structure of another block processor provided by an embodiment of the present application. As shown in Figure 10, the block processor 10 may include a block header cache, a block header instruction fetch unit 101, a block header decoding unit 103, a block distribution unit 102, a block execution unit 31, a block execution unit 32, a block execution unit 33, a block execution unit 34, and other block execution units, a shared register 104, a memory read / write unit 41, a memory read / write unit 42, a memory read / write unit 43, a memory read / write unit 44, and other memory read / write units, a secondary cache 105, and a bus 106.

[0142] The block header fetch unit 101 is used to sequentially fetch block headers from the block header cache, generally fetching one block header per clock cycle. The specific functions of the block header fetch unit 101 can be found in the description of the embodiment corresponding to FIG5 above, and will not be repeated here.

[0143] The block header decoding unit 103 is configured to decode the block header and send the decoding result to the block distribution unit 102 .

[0144] The block distribution unit 102 is used to discover the dependencies and parallelism between blocks based on the information indicated in the block header (such as input / output register information), and distribute each block header to the corresponding pipeline for execution, that is, distribute the block header to the corresponding block execution unit.

[0145] The block execution unit 31, etc. is used to sequentially obtain and execute each microinstruction in the block corresponding to the current block header according to the information indicated in the block header (such as the storage location of the block, i.e., the block jump pointer offset). As shown in Figure 10, in some possible embodiments, the block execution unit 31, etc. can execute microinstructions according to the classic five-stage pipeline. In addition, if a microinstruction terminates abnormally during the execution process, the block execution unit 31, etc. can submit the block status of the current block instruction to the shared register 104 so that the subsequent corresponding program can access the shared register 104 to obtain the block status and handle the exception. Optionally, a first-level cache can be set in each block execution unit such as the block execution unit 31 for reading microinstructions. As shown in Figure 10, the block execution unit 31, etc. is connected to the memory read and write unit 42, etc. in a one-to-one correspondence. The memory read and write unit 42, etc. is used to submit the execution results of the block instruction executed on the corresponding block execution unit.

[0146] Furthermore, referring to FIG11 , FIG11 is a schematic diagram of the structure of another block processor provided by an embodiment of the present application. As shown in FIG11 , the block processor 10 may include a block header fetch unit 101, a block distribution unit 102, multiple inter-block renaming units (e.g., including an inter-block renaming unit 51, an inter-block renaming unit 52, an inter-block renaming unit 53, and an inter-block renaming unit 54), multiple block out-of-order transmission queues (e.g., including a block out-of-order transmission queue 61, a block out-of-order transmission queue 62, a block out-of-order transmission queue 63, and a block out-of-order transmission queue 64), multiple executable queues (e.g., including an executable queue 71, an executable queue 72, an executable queue 73, and an executable queue 74), multiple block execution units (e.g., including a block execution unit 31, a block execution unit 32, a block execution unit 33, and a block execution unit 34), and a block reordering buffer 107.

[0147] Below, we will take the four parallel execution pipelines shown in Figure 11 as an example to illustrate the main logic of the block processor parsing the block header, which mainly includes the following stages:

[0148] (1) Block header fetch

[0149] As shown in FIG11 , the block header fetch unit 101 sequentially fetches block headers of different block instructions, such as block header 0, block header 1, block header 2, and block header 3, from the block header cache.

[0150] (2) Block header distribution

[0151] As shown in Figure 11, after decoding the block headers, block dispatch unit 102 dispatches block headers 0, 1, 2, and 3, based on register input and output dependencies, to the pipeline with the most dependencies. As shown in Figure 11, block dispatch unit 102 dispatches block headers 0, 3, 4, 8, and 10 to pipeline 1; blocks 1, 5, 9, and 11 to pipeline 2; blocks 2, 6, and 12 to pipeline 3; and blocks 7 and 13 to pipeline 4.

[0152] As shown in Figure 11, there are strong input-output dependencies between the multiple block headers on each pipeline, while there are fewer dependencies between the block headers of different pipelines. For example, on pipeline 1, there are input-output dependencies between block headers 0, 3, 4, 8, and 10. For another example, only the input of block header 8 on pipeline 1 is dependent on the output of block header 1 on pipeline 2. For another example, only the input of block header 1 on pipeline 2 is dependent on the output of block header 0 on pipeline 1, and only the input of block header 7 on pipeline 4 is dependent on the output of block header 2 on pipeline 3. And so on. We will not elaborate on this here.

[0153] Generally, once the block header fetch unit 101 fetches a block header, the block dispatch unit 102 can immediately dispatch it. For example, the block header fetch unit 101 sequentially fetches block header 0, and after decoding, the block dispatch unit 102 can dispatch it to pipeline 1 (corresponding to the block execution unit 31). Subsequently, the block header fetch unit 101 sequentially fetches block header 1, and after decoding, the block dispatch unit 102 can dispatch it to pipeline 2 (corresponding to the block execution unit 32), and so on.

[0154] (3) Inter-block renaming

[0155] As shown in FIG11 , the block header and inter-block renaming unit 51 is used to direct the input and output dependencies of different block headers to the inter-block communication network and physical registers. In other words, the input and output dependencies between blocks are clearly defined on specific general registers, that is, the input of the current block instruction comes from which general register to which the instruction of the block outputs.

[0156] For example, please refer to Figure 12, which is a schematic diagram of a ring network provided in an embodiment of the present application. The interconnection configuration for communication between blocks can be determined by the input and output information indicated in the block header. As described in Figure 2 above, each block instruction has a maximum of 32 inputs and 32 outputs. Figure 12 illustrates a ring network configured between 8 general registers and 8 block execution units. The configuration of this ring network is primarily related to the number of block execution units within the block processor 10.

[0157] As shown in Figure 12, the block execution unit 31 executes the corresponding block instructions (for example, block instruction 0, block instruction 3, block instruction 4, etc.), obtains input from R2 and R6, and outputs the execution results to R4 and R5; the block execution unit 32 executes the corresponding block instructions (for example, block instruction 1, block instruction 5, etc.), obtains input from R4 and R5, and outputs the execution results to R3; the block execution unit 33 executes the corresponding block instructions, obtains input from R2 and R4, and outputs the execution results to R2; the block execution unit 34 executes the corresponding block instructions, obtains input from R0, and outputs the execution results to R0 and R6; the block execution unit 35 executes the corresponding block instructions, obtains input from R2 and R3, and does not output; the block execution unit 36 ​​executes the corresponding block instructions, obtains input from R3 and R6, and outputs the execution results to R1, R4, R7, and so on, which will not be repeated here.

[0158] (4) Block out-of-order emission queue and executable queue

[0159] As shown in FIG11 , the block headers on the corresponding pipeline enter the block out-of-order transmission queue in sequence. For example, block header 0, block header 3, block header 4, etc. enter the block out-of-order transmission queue 61, block header 1, block header 5, block header 9, etc. enter the block out-of-order transmission queue 62, and so on. Details will not be given here.

[0160] As shown in FIG11 , in the block out-of-order emission queue 61, each block header has an input state corresponding to it, which is the arrival status of the input register of each block header. For example, the input of block headers 1 and 3 comes from the output of block header 0. Then, when all microinstructions in block body 0 corresponding to block header 0 are executed (i.e., when block instruction 0 is executed), block body 0 outputs the corresponding execution result. At this time, the input registers of block headers 3 and 1 are both arrived, and block headers 1 and 3 can enter the executable queue 71 and executable queue 72, respectively. For another example, the input of block headers 8 and 9 comes from the output of block header 5. If the block instruction 5 corresponding to block header 5 is large and has not been executed for a long time, the input registers of block headers 8 and 9 will remain in the unreached state, and block headers 8 and 9 will not be able to enter the corresponding executable queue.

[0161] It should be understood that block instructions without dependencies can be directly sent to the executable queue for execution without considering their input status. For example, block header 0 has no input dependencies and can be directly sent to the executable queue 71.

[0162] In addition, the inter-block renaming stage can also detect possible dependency errors in the block dispatch unit 102. For example, in the block dispatch unit 102, the input of block header 4 on pipeline 1 is dependent on the output of block header 3. However, the inter-block renaming unit 51 may find that there is no dependency between block header 4 and block header 3. In this case, block header 3 and block header 4 can be executed out of order. That is, in the block out-of-order emission queue 61, block header 4 can be emitted and executed before block header 3.

[0163] (5) Block execution

[0164] As shown in Figure 11, block execution units 31 and others parse the block headers in their respective executable queues, obtain the corresponding block locations, and begin executing the series of microinstructions in the blocks. For example, block execution unit 31 sequentially parses block headers 0, 3, and so on, then sequentially obtains and executes corresponding blocks 0 and 3; block execution unit 32 sequentially parses block headers 1, 5, and so on, then sequentially obtains and executes corresponding blocks 1 and 5, etc. This will not be further described here.

[0165] It is understandable that, as shown in FIG11 , block instructions are typically executed out of order, and because each block instruction has a different size, some smaller block instructions may have already been executed, but the block instructions preceding them may not have yet begun execution. For example, block instruction 7 corresponding to block header 7 in FIG11 is relatively small, while block instruction 6 corresponding to block header 6 is relatively large. Thus, after block instruction 7 is executed, a large number of microinstructions in block instruction 6 may still have not been executed. However, after the block instructions are executed out of order, they still need to be submitted in order. To this end, as shown in FIG11 , a block reordering buffer 107 is further provided in the block processor 10. The block reordering buffer 107 can be used to maintain the in-order submission of block instructions (in the order in which the block headers were retrieved). As described above, in order to ensure the parallelism of block instructions, the execution results of a block instruction can be output to a shadow register first during execution. Only after all other block instructions preceding the block instruction have been executed can the execution result of the block instruction be submitted to the corresponding general register. For example, as shown in FIG11 , if only block instructions 0, 1, 3, and 5 have been executed, then block instructions 0 and 1 can successfully submit their execution results. Since block instructions 2 and 4 have not yet been executed, block instructions 3 and 5 cannot be submitted temporarily and must wait in the block reordering buffer 107. Only after block instruction 2 is executed can block instructions 2 and 3 be submitted in sequence. Only after block instruction 4 is executed can block instructions 4 and 5 be submitted in sequence.

[0166] As shown in FIG11 , the block reordering buffer 107 can also be used to save the renaming mapping state, that is, to save the dependency relationship between input and output registers between blocks, and the mapping relationship between shadow registers and general registers of each block instruction.

[0167] In summary, the block processor 10 provided in the embodiment of the present application can analyze block header information to discover inter-block dependencies and parallelism, and then launch different blocks onto parallel pipelines for execution. The block header primarily includes block input and output information as well as the jump relationships between blocks. By parsing the block header, the block processor 10 efficiently resolves inter-block dependencies based on the various hardware units shown in Figure 11, and executes block instructions optimistically and concurrently, greatly improving the execution efficiency of block instructions.

[0168] It should be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the block processor 10. In some possible embodiments, the block processor 10 may have more or fewer components than those shown in Figures 9, 10, and 11, or combine certain components, or split certain components, or arrange the components differently. The various components shown in the figures can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits. In addition, the interface connection relationship between the modules illustrated in the embodiment of the present application is only a schematic illustration and does not constitute a structural limitation on the block processor 10. In some possible embodiments, the block processor 10 may also adopt different interface connection methods in the above-mentioned embodiments, or a combination of multiple interface connection methods.

[0169] In summary, the embodiment of the present application provides a block instruction structure, and based on the block instruction structure, provides a block instruction processing method and a corresponding block processor. The embodiment of the present application separates conventional block instructions into two parts: a block header and a block body, that is, each block instruction consists of a block header and a block body. Among them, the block header is mainly used to express the dependency relationship between block instructions, and the block body is mainly used to express specific calculations. The block body can be composed of multiple microinstructions. Based on this, the embodiment of the present application can first continuously and quickly obtain the block headers of multiple block instructions, and based on the input register information and output register information indicated by each block header, distribute the block headers to the corresponding block execution units in sequence, so that the block instructions to be executed between different block execution units have almost no dependency, thereby ensuring the reliable parallelism of multiple block execution units. Then, the block execution unit can quickly obtain the corresponding block based on the block storage location indicated by the block header and execute it. In this way, compared with the prior art solution in which the entire block instruction must be retrieved before subsequent distribution and execution can be carried out, resulting in low instruction execution efficiency, the embodiment of the present application is based on a fundamental improvement to the block instruction structure. By obtaining and parsing the block header in advance, the block is pre-scheduled, which greatly improves the execution efficiency of the block instruction and the performance of the block processor.

[0170] Based on the description of the above method embodiment, the embodiment of the present application also provides an electronic device. Please refer to Figure 13, which is a structural diagram of an electronic device provided by the embodiment of the present application. As shown in Figure 13, the electronic device 110 includes at least a processor 1101, an input device 1102, an output device 1103 and a memory 1104. The electronic device may also include other common components, which are not described in detail here. Among them, the processor 1101, input device 1102, output device 1103 and memory 1104 in the electronic device can be connected via a bus or other means. The electronic device 110 can be a smart wearable device, a smart phone, a tablet computer, a laptop computer, a desktop computer, a vehicle-mounted computer or a server, etc., or it can be a server cluster composed of multiple servers or a cloud computing service center.

[0171] The processor 1101 in the electronic device 110 may be the block processor described in FIG. 9 , FIG. 10 or FIG. 11 .

[0172] The memory 1104 in the electronic device 110 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 1104 can exist independently and be connected to the processor 1101 via a bus. The memory 1104 can also be integrated with the processor 1101.

[0173] A computer-readable storage medium may be stored in the memory 1104 of the electronic device 110. The computer-readable storage medium is used to store a computer program, which includes program instructions. The processor 1101 is used to execute the program instructions stored in the computer-readable storage medium. The processor 1101 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device 110. It is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. In one embodiment, the processor 1101 described in the embodiment of the present application can be used to perform a series of processes of a method for processing block instructions, including: sequentially obtaining an i-th block header; distributing the i-th block header to the j-th block execution unit based on first information indicated by the i-th block header; the first information indicated by the i-th block header includes input register information and output register information of the i-th block instruction to which the i-th block header belongs; obtaining and executing the i-th block body through the j-th block execution unit based on second information indicated by the i-th block header; the i-th block body is the block body of the i-th block instruction, and the i-th block body includes N microinstructions; the second information indicated by the i-th block header includes the storage location of the i-th block body; i is an integer greater than 1, and j and N are integers greater than or equal to 1, etc. For details, please refer to the relevant descriptions in the embodiments corresponding to Figures 1a to 12 above, which will not be repeated here.

[0174] An embodiment of the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium may store a program, and when the program is executed by a processor, the processor can perform part or all of the steps of any one of the above method embodiments.

[0175] An embodiment of the present application also provides a computer program, which includes instructions. When the computer program is executed by a multi-core processor, the processor can execute some or all of the steps of any one of the above method embodiments.

[0176] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0177] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0178] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0179] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0180] In addition, the functional units in the embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0181] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc., specifically a processor in a computer device) to execute all or part of the steps of the above-mentioned methods of each embodiment of the present application. Among them, the aforementioned storage medium may include: U disk, mobile hard disk, magnetic disk, optical disk, read-only memory (Read-Only Memory, abbreviated: ROM) or random access memory (Random Access Memory, abbreviated: RAM) and other media that can store program codes.

[0182] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for processing a block instruction, characterized in that: The block instruction includes a block header and a block body; the method includes: Get the i-th block header; Distribute the i-th block header to the j-th block execution unit based on the first information indicated by the i-th block header; the first information indicated by the i-th block header includes input register information and output register information of the i-th block instruction corresponding to the i-th block header; Obtaining, by the j-th block execution unit, an i-th block based on the second information indicated by the i-th block header, the i-th block corresponding to the block of the i-th block instruction, the second information indicated by the i-th block header including a storage location of the i-th block; i is an integer greater than 1, and j is an integer greater than or equal to 1; Execute N microinstructions included in the i-th block; N is an integer greater than or equal to 1.

2. The method according to claim 1, characterized in that The i-th block header is located in a first storage area, and the i-th block is located in a second storage area; a plurality of block headers are stored in the first storage area, and a plurality of blocks are stored in the second storage area; the first storage area and the second storage area are storage areas pre-divided in the memory.

3. The method according to any one of claims 1-2, characterized in that The method further comprises: Get the i-1th block header, where the i-1th block header corresponds to the block header of the i-1th block instruction; The step of obtaining the i-th block header includes: Based on the third information indicated by the i-1th block header, determine the i-th block instruction executed after the i-1th block instruction and obtain the i-th block header; the third information indicated by the i-1th block header includes a jump type of the i-1th block instruction.

4. The method according to any one of claims 1 to 3, characterized in that The storage location of the i-th block includes the storage locations of the N microinstructions in the i-th block; and obtaining the i-th block based on the second information indicated by the i-th block header includes: Retrieving the k-th microinstruction from a storage location corresponding to the k-th microinstruction; k is an integer greater than or equal to 1 and less than or equal to N; The executing the N microinstructions included in the i-th block includes: executing the k-th microinstruction.

5. The method according to claim 4, characterized in that The i-th block corresponds to N virtual registers, and the N virtual registers correspond one-to-one to the N microinstructions in the i-th block; each virtual register is used to store an execution result obtained after the corresponding microinstruction is executed.

6. The method according to claim 5, characterized in that The executing the kth microinstruction includes: Determining, based on a decoding result of the k-th microinstruction, that input data of the k-th microinstruction includes an execution result of the p-th microinstruction among the N microinstructions; Based on a relative distance kp between the k-th microinstruction and the p-th microinstruction, obtaining an execution result of the p-th microinstruction from the p-th virtual register, and obtaining an execution result of the k-th microinstruction based on the execution result of the p-th microinstruction; p is an integer greater than or equal to 1 and less than k; The execution result of the k-th microinstruction is output to the corresponding k-th virtual register for storage.

7. The method according to any one of claims 4 to 6, characterized in that: The method further comprises: If the k-th microinstruction terminates abnormally, submitting the current intra-block state of the i-th block instruction to a system register; the system register is used to be accessed by the target program to obtain the intra-block state and process the abnormal termination of the k-th microinstruction; The state within the block includes: the storage location of the i-th block header, the storage location of the k-th microinstruction, and the execution result of the microinstruction executed before the k-th microinstruction.

8. The method according to any one of claims 1 to 7, characterized in that The i-th block header is also used to indicate the attributes and types of the i-th block instruction; wherein the attributes include any one or more of the submission strategy, atomicity, visibility, and orderliness of the i-th block instruction, and the types include any one or more of fixed-point, floating-point, custom block, and accelerator call.

9. A block instruction processor, characterized in that The block instruction includes a block header and a block body; the block instruction processor includes a block header instruction fetch unit, a block distribution unit and a plurality of block execution units; The block header fetch unit is used to obtain the i-th block header; The block distribution unit is configured to distribute the i-th block header to the j-th block execution unit based on the first information indicated by the i-th block header; the first information indicated by the i-th block header includes input register information and output register information of the i-th block instruction corresponding to the i-th block header; The j-th block execution unit is configured to obtain an i-th block based on the second information indicated by the i-th block header; the i-th block corresponds to the block of the i-th block instruction; the second information indicated by the i-th block header includes a storage location of the i-th block; i is an integer greater than 1, and j is an integer greater than or equal to 1; The j-th block execution unit is further configured to execute N microinstructions included in the i-th block, where N is an integer greater than or equal to 1.

10. The block instruction processor according to claim 9, characterized in that The i-th block header is located in a first storage area of ​​the memory, and the i-th block is located in a second storage area of ​​the memory; the first storage area stores multiple block headers, and the second storage area stores multiple blocks; the first storage area and the second storage area are pre-divided storage areas in the memory.

11. The block instruction processor according to any one of claims 9 to 10, characterized in that: The block header instruction fetch unit is further configured to: obtain the i-1th block header, where the i-1th block header corresponds to the block header of the i-1th block instruction; The block header instruction fetch unit is specifically configured to: determine the i-th block instruction executed after the i-1-th block instruction and obtain the i-th block header based on the third information indicated by the i-1-th block header; the third information indicated by the i-1-th block header includes a jump type of the i-1-th block instruction.

12. The block instruction processor according to any one of claims 9 to 11, characterized in that: The storage location of the i-th block includes the storage location of each of the N microinstructions in the i-th block; the j-th block execution unit is specifically configured to: Retrieving the k-th microinstruction from a storage location corresponding to the k-th microinstruction; k is an integer greater than or equal to 1 and less than or equal to N; Execute the kth microinstruction.

13. The block instruction processor according to claim 12, characterized in that: The i-th block corresponds to N virtual registers, and the N virtual registers correspond one-to-one to the N microinstructions in the i-th block; each virtual register is used to store an execution result obtained after the corresponding microinstruction is executed.

14. The block instruction processor according to claim 13, wherein: The j-th block execution unit is specifically configured to: Determining, based on a decoding result of the k-th microinstruction, that input data of the k-th microinstruction includes an execution result of the p-th microinstruction among the N microinstructions; Based on a relative distance kp between the k-th microinstruction and the p-th microinstruction, obtaining an execution result of the p-th microinstruction from the p-th virtual register, and obtaining an execution result of the k-th microinstruction based on the execution result of the p-th microinstruction; p is an integer greater than or equal to 1 and less than k; The execution result of the k-th microinstruction is output to the corresponding k-th virtual register for storage.

15. The block instruction processor according to any one of claims 12 to 14, characterized in that: The j-th block execution unit is further configured to: If the k-th microinstruction terminates abnormally, submitting the current intra-block state of the i-th block instruction to a system register; the system register is used to be accessed by the target program to obtain the intra-block state and process the abnormal termination of the k-th microinstruction; The in-block state includes: the storage location of the i-th block header, the storage location of the k-th microinstruction, and the execution result of the microinstruction executed before the k-th microinstruction.

16. The block instruction processor according to any one of claims 9 to 15, characterized in that: The i-th block header is also used to indicate the attributes and types of the i-th block instruction; wherein the attributes include any one or more of the submission strategy, atomicity, visibility, and orderliness of the i-th block instruction, and the types include any one or more of fixed-point, floating-point, custom block, and accelerator call.

17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

18. A chip, characterized in that: The chip includes a memory and a block instruction processor as described in any one of claims 9 to 16 above, and the block instruction processor is coupled to the memory; wherein the block instruction processor is used to call the program code stored in the memory to execute the method described in any one of claims 1 to 8 above.