A processor instruction prefetching and instruction parsing design system
The processor instruction pre-fetch and analysis system addresses cache misses by efficiently managing mixed-width instructions, enhancing performance through optimized pre-fetching and concatenation.
Patent Information
- Application Number
- CN202210384173.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-13
AI Technical Summary
When existing processors process instructions with different bit widths, they cannot effectively parse and splice, resulting in performance degradation, especially when cache miss cannot prefetch instructions in time, affecting processor performance.
A processor instruction prefetching and instruction analysis system is designed, including an instruction storage module, an instruction analysis module and an instruction splicing module. Through effective status flags and counters, the instruction bit width and branch instructions are judged, and the prefetching function is performed when the conditions are met.
It realizes fast and efficient parsing and splicing of 32bit and 16bit instructions, improves processor compatibility and prefetching efficiency, and avoids performance degradation caused by cache miss.
Smart Images

Figure CN114637537B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of integrated circuit instruction prefetching, and particularly to a processor instruction prefetching and instruction parsing design system. Background Art
[0002] Instruction prefetching technology has greatly improved the execution performance of processors. When a cache miss occurs, the performance of the processor may decline because the processor cannot obtain instructions in time. Instruction prefetching technology can pre-store target instructions in the instruction cache in advance to avoid the performance decline of the CPU caused by cache misses. Currently, instruction prefetching is divided into software instruction prefetching and hardware instruction prefetching. For software prefetching, most processors provide prefetch instructions. Software prefetching means that prefetch instructions are compiled by the compiler during compilation to pre-load data in the next-level cache in advance. Since prefetch instructions are used, the prefetch address needs to be calculated, which may cause the processor to not issue prefetch instructions in time, resulting in a decline in the processing ability of the processor. Hardware prefetching can prefetch the data of the next level into the instruction cache in advance based on the hardware circuit, thus avoiding the performance decline caused by cache misses.
[0003] Since the instruction bit widths are different, such as there may be 32-bit instructions and there may also be 16-bit instructions, and there may also be a situation where instructions of different bit widths are stored crosswise. The processor cannot directly use the crosswise stored instructions and needs to first parse and splice the crosswise stored instructions to obtain individual instructions before using them for the processor. Summary of the Invention
[0004] To solve the above technical problems, a processor instruction prefetching and instruction parsing design system of the present invention includes an instruction storage module, an instruction parsing module, an instruction splicing module, and an instruction prefetching module. Among them, the instruction storage module, the instruction parsing module, and the instruction splicing module are in linear series communication. The instruction storage module stores instruction units including inst0, inst1, inst2, and inst3. Whether inst0, inst1, inst2, and inst3 are valid is respectively represented by valid0, valid1, valid2, and valid3. The instruction storage module can continuously store the current instruction according to the register status in the module and the parsing result of the instruction parsing module. The instruction parsing module can parse the bit width and its arrangement of the current instruction, and can also parse whether the current instruction is a branch instruction. The instruction splicing module can splice instructions continuously and without bubbles. The instruction prefetching module can determine whether the address of the current instruction is at the boundary of a conditional instruction cache block. If it is at the boundary of a conditional instruction cache block, then the prefetch function is executed.
[0005] In an embodiment of the present invention, the bit widths of inst0, inst1, inst2, and inst3 in the instruction storage module are all 16 bits, and the bit widths of valid0, valid1, valid2, and valid3 are all 1 bit.
[0006] In an embodiment of the present invention, in the instruction storage module, it is determined whether the current 4 inst registers need to be loaded or cleared based on the states of valid0 - valid3 and the state of Inst_32_buf_afull.
[0007] In an embodiment of the present invention, inst0 - inst3 and valid0 - valid3 both have a clear function and support a flush function.
[0008] In an embodiment of the present invention, the instruction prefetch module can perform a prefetch function based on the address of the current instruction and whether the current instruction is the first branch instruction in the instruction cache block.
[0009] In an embodiment of the present invention, the instruction prefetch module sets a prefetch interval enable configuration register cfg_block_interval_en and a prefetch interval configuration register cfg_block_interval. The prefetch interval enable configuration register cfg_block_interval_en is used to enable the prefetch function for the block interval. When cfg_block_interval_en is invalid, the prefetch module determines whether the current instruction is the first branch jump instruction of the current block for each block.
[0010] In an embodiment of the present invention, the method for determining whether a block is the first branch instruction in the instruction prefetch module is as follows: Assume that the instruction cache block size is 128 x 32 bits. The instruction prefetch module uses a flag register. If the lower 7 bits of the current instruction address are 0, then the flag register is cleared, and a counter counter is used to count each subsequent instruction, with the counter incremented by one each time.
[0011] The above technical solution of the present invention has the following advantages compared with the prior art: The processor instruction prefetch and instruction parsing design system of the present invention uses software prefetch, which means that prefetch instructions are compiled by the compiler at compile time to pre - load data in the next - level cache, and is compatible with 32 - bit instructions and 16 - bit instructions. It can parse different types of batch instructions in a configurable template manner, has strong scalability and compatibility, and the instruction prefetch and instruction parsing are faster and more efficient. Description of the Drawings
[0012] In order to make the content of the present invention easier to be clearly understood, the following further describes the present invention in detail according to specific embodiments of the present invention and in conjunction with the accompanying drawings.
[0013] Figure 1 It is the overall structure diagram of the processor instruction prefetching and instruction parsing design system of the present invention;
[0014] Figure 2 It is the schematic flow diagram of the instruction prefetching of the present invention. Specific Embodiments
[0015] As Figure 1 shown, this embodiment provides a processor instruction prefetching and instruction parsing design system, including an instruction storage module, an instruction parsing module, an instruction splicing module, and an instruction prefetching module. Among them, the instruction storage module, the instruction parsing module, and the instruction splicing module are linearly connected in series for communication. The instruction storage module stores instruction units including inst0, inst1, inst2, and inst3. Whether inst0, inst1, inst2, and inst3 are valid is respectively represented by valid0, valid1, valid2, and valid3. The instruction storage module can continuously store the current instruction according to the register status in the module and the parsing result of the instruction parsing module; the instruction parsing module can parse the bit width and its arrangement of the current instruction, and can parse whether the current instruction is a branch instruction; the instruction splicing module can continuously splice the instructions without bubbles; the instruction prefetching module can determine whether the address of the current instruction is at the boundary of a conditional instruction cache block. If it is at the boundary of a conditional instruction cache block, then the prefetching function is executed.
[0016] Some processor systems support both 32-bit width instructions and 16-bit width instructions at the same time. Among them, the 16-bit width instructions are compressed instructions. The lower 2 bits of the non-compressed instruction op are 2'b11. The lower 2 bits of the compressed instruction op are 2'b00, 2'b01, 2'b10. The data bit width output by the instruction cache is 32 bits. According to the above method, it can be parsed whether the instruction is 32 bits or 16 bits. 32-bit data is read from the instruction cache, and 16-bit instructions and 32-bit instructions may be stored crosswise. The crosswise storage of instructions with different bit widths will cause subsequent modules to be unable to be directly used. Therefore, it is necessary to pre-parse and process the instructions. After parsing and processing, the output 32-bit data only contains one 32-bit instruction or only contains one 16-bit instruction. There are 5 cases for the data output by the instruction cache:
[0017] Instruction cache output high 16 bits Instruction cache output low 16 bits Case 1 High 16 bits of 32-bit instruction Low 16 bits of 32-bit instruction Case 2 16-bit instruction 16-bit instruction Case 3 Low 16 bits of 32-bit instruction 16-bit instruction Case 4 Low 16 bits of 32-bit instruction High 16 bits of 32-bit instruction Case 5 16-bit instruction High 16 bits of 32-bit instruction
[0018] In the instruction storage module, the bit widths of inst0, inst1, inst2, and inst3 are all 16 bits and are used to store instructions. The bit widths of valid0, valid1, valid2, and valid3 are all 1 bit and are used to indicate whether inst0, inst1, inst2, and inst3 are valid respectively. The instruction storage module can accelerate the processing of instructions and enable continuous output of instructions. By the states of valid0 - valid3 and Inst_32_buf_afull, it is judged whether the current 4 inst registers need to be loaded or cleared. Both inst0 - inst3 and valid0 - valid3 have a clear function and support the flushing function.
[0019] The processing procedures of the instruction storage module, the instruction parsing module, and the instruction splicing module are as follows:
[0020] Situation 1: valid0 - valid3 are invalid and Inst_32_buf_afull is invalid. Read the 32 - bit instruction cache instruction into inst1 and inst0. Among them, the high 16 bits of the 32 - bit instruction in the instruction cache are stored in inst1, and the low 16 bits of the 32 - bit instruction in the instruction cache are stored in inst0.
[0021] Situation 2: valid0 - valid1 are valid, valid2 - valid3 are invalid, and Inst_32_buf_afull is invalid. At this time, the instruction parsing module parses the low 2 bits of inst0. If it is 2'b11, then directly splice inst0 and inst1 into a 32 - bit instruction, that is, Inst_32_buf_data = {inst1, inst0}. At the next moment, directly store icache_data[15:0] into inst0 and icache_data[31:16] into inst1. If the subsequent inst_32_buf is full, then the 32 - bit data in the instruction cache will be stored in inst2 and inst3, and the corresponding valid bit will be set to 1. If it is 2'b00 or 2'b01 or 2'b10, then the 16 - bit data of inst0 is used as the low 16 - bit data of the 32 - bit instruction, and the high 16 bits of the 32 - bit instruction are fetched as 16’b0, that is, Inst_32_buf_data = {16’b0, inst0}. At the next moment, load the data in inst1 into inst0. And store icache_data[15:0] into inst1 and icache_data[31:16] into inst2.
[0022] Case 3: valid0 is valid, valid1 - valid3 are invalid, and Inst_32_buf_afull is invalid. The processing method is the same as that in Case 1.
[0023] Case 4: valid0 - valid2 are valid, valid3 is invalid, and Inst_32_buf_afull is invalid. At this time, the instruction parsing module parses the lower 2 bits of inst0. If it is 2'b11, then directly concatenate inst0 and inst1 into a 32-bit instruction, that is, Inst_32_buf_data = {inst1, inst0}. At the next moment, store inst2 into inst0. Directly store icache_data[15:0] into inst1 and icache_data[31:16] into inst2. If it is 2'b00 or 2'b01 or 2'b10, then use the 16-bit data of inst0 as the lower 32 bits of the 32-bit instruction, and the upper 16 bits of the 32-bit instruction are fetched as 0, that is, Inst_32_buf_data = {16'b0, inst0}. At the next moment, load inst1 into inst0 and inst2 into inst1. And store icache_data[15:0] into inst2 and icache_data[31:16] into inst3.
[0024] Case 5: valid0 - valid3 are valid, and Inst_32_buf_afull is invalid. At this time, the instruction parsing module parses the lower 2 bits of inst0. If it is 2'b11, then directly concatenate inst0 and inst1 into a 32-bit instruction, that is, Inst_32_buf_data = {inst1, inst0}. At the next moment, load inst2 into inst0 and load inst3 into inst1. Load inst2 into inst0. And store icache_data[15:0] into inst2 and icache_data[31:16] into inst3.
[0025] Case 6: valid0 - valid3 are invalid, and Inst_32_buf_afull is valid. Directly store icache_data[15:0] into inst0 and icache_data[31:16] into inst1.
[0026] Situation 7: valid1 - valid3 are invalid, valid0 is valid, and Inst_32_buf_afull is valid. Directly store icache_data[15:0] into inst1 and icache_data[31:16] into inst2.
[0027] Situation 8: valid2 - valid3 are invalid, valid0 - valid1 are valid, and Inst_32_buf_afull is valid. Directly store icache_data[15:0] into inst2 and icache_data[31:16] into inst3.
[0028] Situation 9: valid3 is invalid, valid0 - valid2 are valid, and Inst_32_buf_afull is valid. Do nothing.
[0029] As Figure 2 shown, the instruction prefetch module sets the prefetch interval enable configuration register cfg_block_interval_en and sets the prefetch interval configuration register cfg_block_interval. The prefetch interval enable configuration register cfg_block_interval_en is used to enable the prefetch function for the block interval. When cfg_block_interval_en is invalid, the prefetch module judges whether the current instruction is the first branch jump instruction of the current block for each block. If it is the first branch jump instruction, then parse the block where the jump address of this branch jump instruction is located, generate a prefetch address, and the instruction cache executes the prefetch function. The instruction prefetch module only parses the first branch jump instruction of the block, and only the first branch jump instruction of the block executes the above - mentioned prefetch function, and other subsequent branch instructions of the block do not execute the prefetch function. When cfg_block_interval_en is valid, only for the blocks that meet the conditions of the block prefetch interval configuration register, the instruction prefetch module parses the current instruction to judge whether the current instruction is the first branch jump instruction of the current block. If it is the first branch jump instruction, execute the instruction prefetch function, and other subsequent branch instructions of the current block do not execute the prefetch function.
[0030] The method for determining whether a Block is the first branch instruction is as follows: Assume that the instruction cache block size is 128 x 32 bits. The instruction prefetch module uses a flag register. If the lower 7 bits of the current instruction address are 0, then the flag register is cleared. A counter "counter" is used to count each subsequent instruction, incrementing the counter by one each time. When counter is greater than or equal to 0 and less than 127 and the first branch instruction is parsed, the flag register is set to 1. When counter is equal to 127, the flag register is cleared. When the flag register is equal to 1, the branch instruction read is not prefetched. When the flag register is equal to 0, the branch instruction read performs the prefetch function.
[0031] The instruction prefetch module determines whether the address of the current instruction is at the boundary of a cache block that meets the conditions. If it is at the boundary of an instruction cache block that meets the conditions, then the prefetch function is executed. The prefetch interval enable configuration register "cfg_block_interval_en" is used to enable the prefetch function for block intervals. When "cfg_block_interval_en" is invalid, the prefetch module determines the block boundary for each block and executes the prefetch function. When "cfg_block_interval_en" is valid, only the blocks that meet the conditions of the block prefetch interval configuration register perform the expected function. When "cfg_block_interval_en" is valid, for the blocks that meet the "cfg_block_interval" condition, it is determined whether they are at the block boundary based on the value of the boundary range register "cfg_range". If they are at the block boundary, the prefetch function is executed; otherwise, the prefetch function is not executed.
[0032] The method for calculating whether the current instruction is at the block boundary is as follows: Set the boundary range register "cfg_range". The range size between the address of the current instruction that triggers the prefetch and the block boundary can be configured through the configuration register. Assume that the instruction cache block size is 128 x 32 bits. The instruction prefetch module extracts the lower 7 bits of the address where the current instruction is located. If the lower 7 bits of the current instruction address are less than the boundary range register "cfg_range", then it is determined that the current instruction is not at the block boundary, and the instruction cache does not execute the prefetch function. If the lower 7 bits of the current instruction address are greater than or equal to the boundary range register "cfg_range", then it is determined that this address is at the boundary of the instruction cache block. At this time, the instruction prefetch module calculates the prefetch address and sends the prefetch address to the instruction cache to execute the prefetch function.
[0033] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A processor instruction prefetching and instruction parsing design system, characterized in that It includes an instruction storage module, an instruction parsing module, an instruction splicing module, and an instruction prefetching module. Among them, the instruction storage module, the instruction parsing module, and the instruction splicing module communicate linearly in series. The instruction storage module stores instruction units including inst0, inst1, inst2, and inst3. Whether inst0, inst1, inst2, and inst3 are valid is represented by valid0, valid1, valid2, and valid3 respectively. The instruction storage module continuously stores the current instruction according to the register status in the module and the parsing result of the instruction parsing module. The instruction parsing module parses the bit width and its arrangement of the current instruction and determines whether the current instruction is a branch instruction. The instruction splicing module continuously splices the instructions without bubbles. The instruction prefetching module determines whether the address of the current instruction is at the boundary of a conditional instruction cache block. If it is at the boundary of a conditional instruction cache block, then the prefetching function is executed.
2. The processor instruction prefetching and instruction parsing design system according to claim 1, characterized in that: In the instruction storage module, the bit widths of inst0, inst1, inst2, and inst3 are all 16 bits, and the bit widths of valid0, valid1, valid2, and valid3 are all 1 bit.
3. The processor instruction prefetching and instruction parsing design system according to claim 2, characterized in that: In the instruction storage module, it is judged whether the current 4 inst registers need to be loaded or cleared according to the states of valid0 - valid3 and the state of Inst_32_buf_afull.
4. The processor instruction prefetching and instruction parsing design system according to claim 2, characterized in that: Both inst0 - inst3 and valid0 - valid3 have a clearing function and support a flushing function.
5. The processor instruction prefetching and instruction parsing design system according to claim 1, characterized in that: The instruction prefetching module executes the prefetching function according to the address of the current instruction and whether the current instruction is the first branch instruction in the instruction cache block.
6. The processor instruction prefetching and instruction parsing design system according to claim 5, characterized in that: The instruction prefetching module sets the prefetch interval enable configuration register cfg_block_interval_en and the prefetch interval configuration register cfg_block_interval. The prefetch interval enable configuration register cfg_block_interval_en is used to enable the prefetching function of the block interval. When cfg_block_interval_en is invalid, the prefetching module judges whether the current instruction is the first branch jump instruction of the current block for each block.
7. The processor instruction prefetching and instruction parsing design system according to claim 6, characterized in that: The method for the instruction prefetching module to judge whether a block is the first branch instruction is as follows: Assume that the size of the instruction cache block is 128x32 bits. The instruction prefetching module uses a flag register. If the lower 7 bits of the current instruction address are 0, then the flag register is cleared, and a counter counter is used to count each subsequent instruction, with the counter incremented by one each time.
Citation Information
Patent Citations
Method for realizing instruction prefetching in front-end assembly line by advanced pointer method
CN111209043A
Power efficient instruction prefetch mechanism
US20060174090A1