Programmable Instruction Distribution System

Through the instruction splitting and replay units of the programmable instruction distribution system, the invalid operation problem of GPGPU in branch and loop instruction segments is solved, and execution efficiency and parallel performance are improved.

CN119847610BActive Publication Date: 2025-07-08SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510316224.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

There are a large number of invalid operations when processing branch instruction segments and loop instruction segments, resulting in reduced execution efficiency and parallel performance.

Method used

A programmable instruction distribution system is adopted, including an instruction splitting unit and an instruction replay unit, and by splitting the complete instruction segment into multiple sub-instruction segments and controllingly replaying the distributed instruction segments, the extended RISC-V instruction set is used to achieve rapid distribution.

Benefits of technology

It reduces invalid operations during repeated instructions distribution, and improves the execution efficiency and parallel performance of GPGPU in branch execution and loop execution scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119847610B_ABST
    Figure CN119847610B_ABST
Patent Text Reader

Abstract

The present invention discloses a programmable instruction distribution system, belonging to the technical field of general-purpose graphics processing units. The technical problem to be solved by the present invention is that there are a large number of invalid operations in the GPGPU instruction pipeline when processing branch instruction segments and loop instruction segments, occupying multiple clock cycles and reducing the execution efficiency and parallel performance of the GPGPU. The technical solution is as follows: The system includes an instruction splitting unit and an instruction replay unit; the instruction splitting unit is used to split a complete instruction segment into multiple sub-instruction segments to complete the fast distribution of specific instruction segments; the instruction splitting unit includes an instruction selection switch and a programmable instruction queue; wherein, the instruction selection switch is used to generate corresponding branch instruction segments from the programmable instruction queue according to an instruction mask and distribute the branch instructions; the instruction replay unit is used to record the distributed instruction segments and controllably replay the corresponding instruction segments; the instruction replay unit includes a register redirection logic, an instruction counter, and an instruction recording queue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of general - purpose graphics processors, and more particularly to a programmable instruction distribution system. Background Art

[0002] General - purpose graphics processors (GPGPUs) adopt a single - instruction multiple - thread architecture. The single - instruction multiple - thread architecture packs multiple threads into warps and executes them at the smallest granularity of warps. GPGPUs complete the execution of a warp instruction through a pipeline of scheduling, instruction fetching (instruction buffer), decoding (instruction cache), issuing, execution, and write - back, as shown in the appendix Figure 1 to complete the execution of a warp instruction and achieve highly parallel data computing. Compared with the complex branch management strategy in CPUs, the instruction branch management strategy of GPGPUs is greatly simplified in order to allocate more hardware resources to the computing part. It splits a warp into multiple branch threads for execution through a thread mask, so the branch threads execute the same instruction segment. However, this method requires the execution of a similar instruction segment multiple times, so the branches repeat a series of operations from scheduling to write - back.

[0003] In recent years, artificial intelligence technologies represented by neural networks have developed rapidly, and the computing power requirements of various neural network models have increased exponentially. Matrix operations, as the most basic operators in neural network models, also account for the largest proportion of computing power. To support ultra - large - scale matrix operations, GPGPUs usually first divide a large matrix into blocks and then process each block matrix separately. The block matrix operation instructions have significant regularity: read and write registers with the same label n and repeat specified operations; or read and write registers with label n + m according to a fixed rule and repeat specified operations. However, the instruction distribution strategy of GPGPUs is relatively simple. Even for a completely identical instruction segment, it is still necessary to re - perform scheduling, instruction fetching, and decoding operations. In addition, whether it is a branch instruction or a block matrix operation instruction, due to data dependencies between instructions (the subsequent instruction uses the calculation result of the previous instruction), the scheduler needs to wait for the previous instruction to complete decoding or write - back, as shown in the appendix Figure 1 to allow the subsequent instruction to start execution, which significantly increases the execution latency between instructions. It can be seen that when processing branch instruction segments and loop instruction segments, there are a large number of invalid operations in the GPGPU instruction pipeline, occupying multiple clock cycles and reducing the execution efficiency and parallel performance of GPGPUs. Summary of the Invention

[0004] The technical task of the present invention is to provide a programmable instruction distribution system to solve the problem that when processing branch instruction segments and loop instruction segments, there are a large number of invalid operations in the GPGPU instruction pipeline, occupying multiple clock cycles and reducing the execution efficiency and parallel performance of GPGPUs.

[0005] The technical task of the present invention is achieved in the following manner. A programmable instruction distribution system includes an instruction splitting unit and an instruction replay unit;

[0006] The instruction splitting unit is used to split a complete instruction segment into multiple sub-instruction segments to complete the fast distribution of a specific instruction segment. The instruction splitting unit includes an instruction selection switch and a programmable instruction queue. Among them, the instruction selection switch is used to generate a corresponding branch instruction segment from the programmable instruction queue according to an instruction mask and distribute the branch instruction. The programmable instruction queue is used to store 32 RISC-V instructions;

[0007] The instruction replay unit is used to record the distributed instruction segments and controllably replay the corresponding instruction segments. The instruction replay unit includes a register redirection logic, an instruction counter, and an instruction recording queue. Among them, the instruction counter is used to continuously fetch loop instruction segments from the instruction recording queue according to the number of loops. The register redirection logic is used to replace the register fields in the loop instructions, generate and distribute the loop instructions. The instruction recording queue is used to record 32 RISC-V instructions.

[0008] Preferably, the system also extends the RISC-V instruction set, and defines an instruction fill (MIPT) instruction and an instruction split (MEXP) instruction for the instruction splitting unit, and an instruction replay (ISREPLAY) instruction for the instruction replay unit respectively;

[0009] Among them, the instruction fill instruction includes an inst-half field, an idx field of the instruction fill instruction, and a hi / lo field. The inst-half field is a half instruction represented by an immediate number. The idx field and the hi / lo field of the instruction fill instruction are used to indicate the specific positions where the corresponding instructions are filled into the programmable instruction queue;

[0010] The instruction split instruction includes an idx field of the instruction split instruction, a Mask field, and a mode field of the instruction split instruction. The idx field of the instruction split instruction is used to indicate the instruction to start execution from the instruction queue. The Mask field is used to indicate the instructions to be executed and the instructions to be skipped in the form of a mask. The mode field of the instruction split instruction is used to indicate whether the queue supports looping, that is, to continue executing the head of the queue after reaching the end of the queue;

[0011] The instruction replay instruction includes the mode field, len field, idx field, loop field, and increment field of the instruction replay instruction; the mode field of the instruction replay instruction is used to control whether the instruction replay unit is in the recording mode or the replay mode; in the recording mode, the len field is used to record the length of the instruction, and the idx field of the instruction replay instruction is used to calculate the start position index in the instruction queue; in the replay mode, the idx field of the instruction replay instruction is used to calculate the start position index, and the len field is used to indicate the length of the replay instruction; the loop field is used to indicate the number of times the instruction segment is replayed; the increment field is used for the register correction parameter during replay.

[0012] Preferably, the instruction splitting unit controls the operation of the instruction selection switch and the programmable instruction queue through two dedicated RISC-V instructions, thereby completing the fast distribution of specific instruction segments; specifically as follows:

[0013] When receiving the instruction fill instruction, the instruction selection switch calculates the start position index of the programmable instruction queue according to the control information of the idx field and the hi / lo field of the instruction fill instruction. The formula is as follows: ; then, according to the start position index, the 16-bit immediate number is stored as half an instruction in the corresponding position of the programmable instruction queue;

[0014] When receiving the split instruction, the instruction selection switch calculates the start position index in the programmable instruction queue according to the idx field of the instruction split instruction. The formula is: ; then, using the Mask field as the instruction mask, the instructions in the programmable instruction queue are selected and executed starting from the index position.

[0015] Preferably, a fill status register is set in the instruction selection switch. The fill status register is a 32-bit register, and each bit corresponds to an instruction in the programmable instruction queue;

[0016] Among them, when the fill status register is set to 1, it means that the instruction has been filled completely;

[0017] When the fill status register is set to 0, it means that the instruction has not been filled completely.

[0018] Preferably, before the instruction selection switch fetches an instruction according to the index, it first reads the fill status register, specifically as follows:

[0019] If the instruction to be fetched has been filled completely, it is directly fetched and sent to the emission unit;

[0020] If the instruction to be fetched has not been filled completely, the execution of the corresponding instruction is blocked.

[0021] Preferably, the mode field in the instruction splitting instruction is used to determine whether the instruction selection switch continues to fetch instructions from the head of the programmable instruction queue when the instruction fetch index to be fetched exceeds the queue length. This method allows for the quick generation of a new instruction segment while only replacing a small portion of the instructions in the queue, reducing the instruction filling delay.

[0022] Preferably, the instruction replay unit controls the operation of the instruction counter, the instruction recording queue, and the register redirection logic through a dedicated RISC-V instruction; specifically as follows:

[0023] When receiving an instruction replay instruction, set the instruction counter to the recording mode or the replay mode according to the mode field of the instruction replay instruction;

[0024] In the recording mode, the instruction counter calculates the starting position index according to the idx field of the instruction replay instruction, then saves the subsequent len instructions, and writes them into the instruction recording queue starting from the index in sequence;

[0025] In the replay mode, the instruction counter calculates the starting position index according to the idx field of the instruction replay instruction, fetches len instructions in sequence from the corresponding position in the instruction recording queue, and the len instructions need to be replayed loop times as a loop instruction segment; after the register redirection logic completes one loop of the instruction segment, it corrects the register index of the replay instruction according to the incret field.

[0026] Preferably, a loop counter is set in the instruction counter, and the counter is incremented by 1 after the instruction segment completes a loop until the counter value is equal to the loop value.

[0027] The programmable instruction distribution system of the present invention has the following advantages:

[0028] (1) By deeply analyzing the key steps and operating principles of instruction branching, instruction loop processing, and instruction distribution, the present invention realizes the fast distribution of thread branch instructions and loop operation instructions, thereby reducing the ineffective operations in the repeated distribution process of these two types of instructions and improving the execution efficiency of the GPGPU in scenarios such as branch execution and loop execution;

[0029] (2) The present invention uses the methods of "instruction splitting" and "instruction replay" to perform programmable calls and repeated calls on the same instruction segment, realizing the fast distribution of specific instruction segments; for thread branch programs, by caching the complete branch instruction segment and generating corresponding sub-instruction segments according to the instruction mask, the fast distribution of branch instructions is completed; for loop programs, by recording the executed instruction segments and generating new loop sub-instruction segments through register remapping, the replay distribution of loop instructions is completed;

[0030] (3) The present invention adds a dedicated instruction splitting logic after the decoding logic. Through the instruction queue and instruction mask, it realizes the controllable distribution of the branch instruction segment. After the instruction splitting logic, a dedicated instruction replay logic is added. Through instruction counting and register remapping, it realizes the fast distribution of the loop instruction segment. This design directly generates new distribution instructions after the decoding unit, avoiding the execution delay between instructions caused by repeated scheduling, fetching, and decoding of similar instruction segments, and reducing the instruction cache occupancy to a certain extent, improving the execution efficiency and parallel performance of the GPGPU. Description of the Drawings

[0031] The present invention will be further described below with reference to the drawings.

[0032] Att Figure 1 is a schematic diagram of the instruction pipeline;

[0033] Att Figure 2 is a schematic diagram of the structure of the programmable instruction distribution system;

[0034] Att Figure 3 is a schematic diagram of the extended instruction;

[0035] Att Figure 4 is a detailed architecture diagram of the programmable instruction distribution system;

[0036] Att Figure 5 is a schematic diagram of the kernel program in Embodiment 2. Detailed Embodiments

[0037] The programmable instruction distribution system of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.

[0038] Embodiment 1: As shown in Att Figure 2 , this embodiment provides a programmable instruction distribution system, which includes an instruction splitting unit and an instruction replay unit;

[0039] The instruction splitting unit is used to split the complete instruction segment into multiple sub-instruction segments to complete the fast distribution of specific instruction segments. The instruction splitting unit includes an instruction selection switch and a programmable instruction queue. Among them, the instruction selection switch is used to generate the corresponding branch instruction segment from the programmable instruction queue according to the instruction mask and distribute the branch instructions. The programmable instruction queue is used to store 32 RISC-V instructions. The dotted part in Att Figure 2 is other functional units of the GPGPU instruction pipeline, which together with the present invention form a complete instruction path.

[0040] The instruction replay unit in this embodiment is used to record the distributed instruction segments and controllably replay the corresponding instruction segments; the instruction replay unit includes a register redirection logic, an instruction counter, and an instruction recording queue; among them, the instruction counter is used to continuously fetch the loop instruction segments from the instruction recording queue according to the number of loops; the register redirection logic is used to replace the register fields in the loop instructions, generate and distribute the loop instructions; the instruction recording queue is used to record 32 RISC-V instructions.

[0041] As shown in the appendix Figure 3 As shown, this embodiment also extends the RISC-V instruction set, and respectively defines an instruction padding (MIPT) instruction and an instruction splitting (MEXP) instruction for the instruction splitting unit, and an instruction replay (ISREPLAY) instruction for the instruction replay unit;

[0042] Among them, the instruction padding instruction includes an inst-half field, an idx field of the instruction padding instruction, and a hi / lo field; the inst-half field is a half instruction represented by an immediate number; the idx field and the hi / lo field of the instruction padding instruction are used to indicate the specific position where the corresponding instruction is filled into the programmable instruction queue;

[0043] The instruction splitting instruction includes an idx field of the instruction splitting instruction, a Mask field, and a mode field of the instruction splitting instruction; the idx field of the instruction splitting instruction is used to indicate the instruction to start execution from the instruction queue; the Mask field is used to indicate the instructions to be executed and the instructions to be skipped in the form of a mask; the mode field of the instruction splitting instruction is used to indicate whether the queue supports looping, that is, continue to execute the head of the queue after reaching the end of the queue;

[0044] The instruction replay instruction includes a mode field of the instruction replay instruction, a len field, an idx field of the instruction replay instruction, a loop field, and an incret field; the mode field of the instruction replay instruction is used to control whether the instruction replay unit is in the recording mode or the replay mode; in the recording mode, the len field is used to record the length of the instruction, and the idx field of the instruction replay instruction is used to calculate the start position index in the instruction queue; in the replay mode, the idx field of the instruction replay instruction is used to calculate the start position index, and the len field is used to indicate the length of the replay instruction; the loop field is used to indicate the number of times the instruction segment is replayed; the incret field is used for the register correction parameter during replay.

[0045] As shown in the appendix Figure 4 As shown, the instruction splitting unit in this embodiment controls the operation of the instruction selection switch and the programmable instruction queue through two dedicated RISC-V instructions, and then completes the fast distribution of specific instruction segments; specifically as follows:

[0046] When receiving an instruction filling instruction, the instruction selection switch calculates the starting position index of the programmable instruction queue according to the control information of the idx field and the hi / lo field of the instruction filling instruction. The formula is as follows: ; Then, according to the starting position index, the 16-bit immediate number is stored as half an instruction in the corresponding position of the programmable instruction queue. In addition, to ensure that a 32-bit instruction will not be dispatched until it is completely filled, a filling status register is set in the instruction selection switch. The filling status register is a 32-bit register, and each bit corresponds to an instruction in the programmable instruction queue. When the filling status register is set to 1, it means that the instruction has been completely filled. When the filling status register is set to 0, it means that the instruction has not been completely filled.

[0047] When receiving a disassembling instruction, the instruction selection switch calculates the starting position index in the programmable instruction queue according to the idx field of the instruction disassembling instruction. The formula is: ; Then, using the Mask field as an instruction mask, instructions in the programmable instruction queue are selected and executed starting from the index position. Before the instruction selection switch fetches an instruction according to the index, it first reads the filling status register. Specifically, if the instruction to be fetched has been completely filled, it is directly fetched and sent to the issuing unit. If the instruction to be fetched has not been completely filled, the execution of the corresponding instruction is blocked.

[0048] In this embodiment, the mode field in the instruction disassembling instruction is used to determine whether the instruction selection switch continues to fetch instructions from the head of the programmable instruction queue when the index of the instruction to be fetched exceeds the queue length. This method allows for the quick generation of a new instruction segment on the premise of only replacing a small part of the instructions in the queue, reducing the instruction filling delay.

[0049] In this embodiment, the instruction replay unit controls the operation of the instruction counter, the instruction recording queue, and the register redirection logic through a dedicated RISC-V instruction. Specifically as follows:

[0050] When receiving an instruction replay instruction, the instruction counter is set to the recording mode or the replay mode according to the mode field of the instruction replay instruction.

[0051] In the recording mode, the instruction counter calculates the starting position index according to the idx field of the instruction replay instruction, then saves the subsequent len instructions, and writes them into the instruction recording queue starting from the index in sequence.

[0052] In the replay mode, the instruction counter calculates the starting position index according to the idx field of the instruction to be replayed, and sequentially fetches len instructions from the corresponding position in the instruction record queue. The len instructions, as the loop instruction segment, need to be replayed loop times; after the register redirection logic completes one loop of the instruction segment, it corrects the register index of the replayed instruction according to the incret field; among them, a loop counter is set in the instruction counter, and the counter is incremented by 1 after the instruction segment completes the loop until the counter value is equal to the loop value.

[0053] In addition to the instruction path through the instruction splitting unit and the instruction replay unit as described above, a direct path for other instructions is set in this embodiment. When an instruction does not need to be split or replayed, the instruction is directly distributed as an execution instruction to the subsequent issue unit, execution unit, and write-back unit, reducing the processing latency of such instructions by the distribution strategy. The instructions distributed by the instruction splitting unit will be sent to the instruction replay unit, and through the cooperation of the two units, more complex instruction distribution applications can be achieved.

[0054] Embodiment 2: Taking sparse block matrix multiplication as an example, the specific application and implementation manner of the present invention are introduced. In the matrix multiplication operation of the neural network model, the feature matrix has sparsity (most elements in the matrix are zero and do not affect the result). By skipping the zero elements in the sparse matrix for calculation, the computing throughput of the GPGPU can be significantly improved. Dividing the large-sized matrix into multiple small-sized matrix blocks can make the GPGPU compatible with more application scenarios. To process the sparse matrix, there are a large number of branch program segments in the kernel program of the GPGPU, and to process the block matrix, there are a large number of loop program segments in the program. The kernel program is as shown in the appendix Figure 5 In this scenario, the following applications are carried out in this embodiment:

[0055] In the instruction splitting unit: First, by repeatedly calling the MINP instruction, the complete branch instruction (instruction segment A) is filled into the programmable instruction queue; then, for the sparse matrix block, the MEXP instruction is called, and the Mask mask segment is set to 0011, and the B0 sub-instruction segment is read from the instruction queue. The B0 sub-instruction segment is directly distributed to the execution unit. For the normal matrix block, the MEXP instruction is also called, but the Mask mask is set to 1100, and the B1 sub-instruction segment is read from the instruction queue. The B1 sub-instruction segment is continued to be sent to the instruction replay unit.

[0056] In the instruction replay unit: First, call the ISREPLAY instruction in the recording mode to store the B1 instruction segment into the instruction recording queue; then, call the ISREPLAY instruction in the replay mode, and the instruction counter replays the B1 instruction segment loop times; finally, after replaying an instruction segment once, the register remapping logic corrects the register fields of the B1 instruction segment (index 1 is replaced by n+1, etc.). The instructions of the instruction replay unit are finally distributed to the execution unit to form a complete instruction path with the B0 instruction segment.

[0057] Among them, the B1 sub-instruction segment is the processing instruction segment corresponding to the non-sparse matrix block. The first instruction (LD r1,addrx) and the second instruction (LD r3,addrx) are load instructions. The first instruction (LD r1,addrx) loads the data of matrix 1 from the main memory address addrx into the register r1; the second instruction (LD r3,addrx) loads the data of matrix 2 from the main memory address addrx into the register r3; the third instruction is the matrix multiplication instruction (MYL r5, r1, r3), which performs a matrix multiplication operation on the data in r1 and r3 (matrix 1 multiplied by matrix 2), and stores the result in the register r5; the fourth instruction is the store instruction (SD addrx, r5), which stores the calculation result in r5 into the main memory address addrx.

[0058] The B0 sub-instruction segment is the processing instruction segment corresponding to the sparse matrix block (there are a large number of zero elements in the matrix, and the constant 0 is used as the calculation result). The first instruction LD zero,zero is a load instruction that loads the constant 0; the second instruction SD zero,zero is a store instruction that stores the constant 0 as the calculation result;

[0059] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A programmable instruction distribution system, characterized in that, The system includes an instruction splitting unit and an instruction replaying unit; The instruction splitting unit is used to split a complete instruction segment into multiple sub-instruction segments to complete the fast distribution of a specific instruction segment; the instruction splitting unit includes an instruction selection switch and a programmable instruction queue; wherein, the instruction selection switch is used to generate corresponding branch instruction segments from the programmable instruction queue according to an instruction mask and distribute the branch instructions; the programmable instruction queue is used to store 32 RISC-V instructions; The instruction replaying unit is used to record the distributed instruction segments and controllably replay the corresponding instruction segments; the instruction replaying unit includes a register redirection logic, an instruction counter and an instruction recording queue; wherein, the instruction counter is used to continuously fetch loop instruction segments from the instruction recording queue according to the number of loops; the register redirection logic is used to replace the register fields in the loop instructions to generate and distribute the loop instructions; the instruction recording queue is used to record 32 RISC-V instructions.

2. The programmable instruction distribution system according to claim 1, wherein The system also extends the RISC-V instruction set, and defines an instruction filling instruction and an instruction splitting instruction for the instruction splitting unit and an instruction replaying instruction for the instruction replaying unit respectively; Among them, the instruction filling instruction includes an inst-half field, an idx field of the instruction filling instruction and a hi / lo field; the inst-half field is a half instruction represented by an immediate number; the idx field and the hi / lo field of the instruction filling instruction are used to indicate the specific positions where the corresponding instructions are filled into the programmable instruction queue; The instruction splitting instruction includes an idx field of the instruction splitting instruction, a Mask field and a mode field of the instruction splitting instruction; the idx field of the instruction splitting instruction is used to indicate the instruction starting to be executed from the instruction queue; the Mask field is used to indicate the instructions to be executed and the instructions to be skipped in the form of a mask; the mode field of the instruction splitting instruction is used to indicate whether the queue supports looping, that is, continue to execute the head of the queue after reaching the end of the queue; The instruction replaying instruction includes a mode field of the instruction replaying instruction, a len field, an idx field of the instruction replaying instruction, a loop field and an incret field; the mode field of the instruction replaying instruction is used to control whether the instruction replaying unit is in the recording mode or the replaying mode; in the recording mode, the len field is used to record the length of the instruction, and the idx field of the instruction replaying instruction is used to calculate the starting position index in the instruction queue; in the replaying mode, the idx field of the instruction replaying instruction is used to calculate the starting position index, and the len field is used to indicate the length of the replaying instruction; the loop field is used to indicate the number of times the instruction segment is replayed; the incret field is used for the register correction parameter during replaying.

3. The programmable instruction distribution system according to claim 1 or 2, characterized in that The instruction splitting unit controls the operation of the instruction selection switch and the programmable instruction queue through two dedicated RISC-V instructions, thereby completing the fast distribution of a specific instruction segment; specifically as follows: When receiving an instruction fill instruction, the instruction selection switch calculates the starting position index of the programmable instruction queue according to the control information of the idx field and the hi / lo field of the instruction fill instruction. The formula is as follows: ; Then, according to the starting position index, the 16-bit immediate number is stored as half an instruction in the corresponding position of the programmable instruction queue; When a split instruction is received, the instruction selection switch calculates the starting position index in the programmable instruction queue based on the idx field of the instruction split instruction, and the formula is: ; then, using the Mask field as the instruction mask, select the instructions in the programmable instruction queue starting from the index position for execution.

4. The programmable instruction distribution system according to claim 3, wherein A filling status register is set in the instruction selection switch, and the filling status register is a 32-bit status register, and each bit corresponds to an instruction in the programmable instruction queue; Among them, when the filling status register is set to 1, it indicates that the instruction has been completely filled; When the fill status register is set to 0, it indicates that the instruction has not been filled completely.

5. The programmable instruction distribution system according to claim 4, wherein Before the instruction selection switch fetches an instruction according to the index, it first reads the fill status register, as follows: If the instruction to be fetched has been filled completely, it is directly fetched and sent to the issue unit; If the instruction to be fetched has not been filled completely, the execution of the corresponding instruction is blocked.

6. The programmable instruction distribution system according to claim 5, characterized in that, The mode field in the instruction split instruction is used to determine whether the instruction selection switch continues to fetch instructions from the head of the programmable instruction queue when the index of the instruction to be fetched exceeds the queue length.

7. The programmable instruction distribution system according to claim 6, wherein The instruction replay unit controls the operation of the instruction counter, the instruction record queue, and the register redirection logic through a dedicated RISC-V instruction; as follows: When receiving an instruction replay instruction, set the instruction counter to the record mode or the replay mode according to the mode field of the instruction replay instruction; In the record mode, the instruction counter calculates the starting position index according to the idx field of the instruction replay instruction, then saves the subsequent len instructions, and writes them into the instruction record queue starting from the index in turn; In the replay mode, the instruction counter calculates the starting position index according to the idx field of the instruction replay instruction, fetches len instructions from the corresponding position in the instruction record queue in turn, and the len instructions need to be replayed loop times as the loop instruction segment; after the register redirection logic completes one loop of the instruction segment, it corrects the register index of the replay instruction according to the incret field.

8. The programmable instruction distribution system according to claim 7, wherein, A loop counter is set in the instruction counter, and the counter is incremented after the instruction segment completes a loop until the counter value is equal to the loop value.

Citation Information

Patent Citations

  • Instruction transmitting unit, instruction executing unit, related device and method

    CN114428638A

  • Data access method, natural language reasoning method, processor and computing equipment

    CN119088453A