Processor program address buffering method and device
By distinguishing the high-bit and low-bit fields of the program address in a superscalar processor, only writing the high-bit field to the buffer, and optimizing the release of table entries by reordering the buffer, the problems of large hardware resource usage and high power consumption are solved, achieving more efficient resource utilization and reduced power consumption.
Patent Information
- Application Number
- CN202510750466.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-12
AI Technical Summary
In the design of program address buffer in existing superscalar processors, hardware resources are large and power consumption is high. In addition, the increase in the number of pipeline stages leads to a large number of table operations, which increases the chip area and power consumption.
The program address comparison unit is used to distinguish the high-bit field and the low-bit field of the instruction, and only the high-bit field is written into the buffer. The entry release frequency is optimized by reordering the buffer, thereby reducing the number of buffer entries and the access frequency.
The processor's resource overhead and power consumption are reduced while maintaining instruction processing bandwidth and reducing the number of buffer table entries and the number of operations.
Smart Images

Figure CN120631449A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of processor technology, and in particular to a processor program address buffering method and device. Background Art
[0002] Superscalar processor (superscalar CPU): A processor that can execute more than one instruction per clock cycle on average.
[0003] With the evolution of processor technology, today's advanced superscalar processors can now execute hundreds of instructions out of order. As pipelines lengthen and instruction throughput increases, some instructions at different pipeline stages may require the corresponding program address as an operand (for example, branch instructions in the Risc-V instruction set). Since the program address is already obtained during the instruction fetch phase, the program address can be passed back through the pipeline, or a centralized buffer can be used to temporarily store these program addresses in order.
[0004] Modern superscalar processors typically use a centralized buffer to store program addresses. This "program address buffer" temporarily stores the program addresses of these instructions and is removed from the buffer after these instructions are successfully executed and meet the retirement conditions. Because the buffer contains a copy of the program address, when the instruction fetch phase writes to the buffer, the program address's table index in the buffer is simultaneously passed to the pipeline's subsequent stages, where the pipeline reads the program address using this table index. It is generally believed that the number of bits in the buffer's table index is much smaller than the number of bits in the program address width, thus reducing the area overhead associated with pipeline transmission of the entire program address.
[0005] When the instruction fetch unit sends multiple instructions to the processor's back-end module, it will carry the program addresses corresponding to these instructions, assuming that the number of instructions is 1. The unit that temporarily stores program addresses (program address buffer) will write 1 program addresses into its own temporary storage area according to the program order, apply for 1 table entry, and provide the instruction issuance unit with the table entry index of 1 program address in the temporary storage area. At the same time, these indexes will also be recorded in the reorder buffer to prepare for the release of the program address buffer. If the instruction requires the program address as an operand, the instruction issuance unit will use the table entry index to read the program address buffer table entry, and send it to the execution unit for execution together with the instruction opcode and other information. After the execution unit completes the instruction, it notifies the reorder buffer that the instruction is complete. The reorder buffer finds the currently completed instructions according to the program order and releases their corresponding program address buffer table entries. The number of released entries is also 1, and the table entries occupied by them in the reorder buffer are also released.
[0006] It can be seen that the existing solution is relatively simple. Assuming that I instructions need to be processed, there will be a total of I program address buffer table entries written and I program address buffer table entries released. The number of operations is relatively large, and in order to meet the bandwidth of the processor to process instructions simultaneously, the number of program address buffer table entries is usually relatively large. Each table entry needs to store a value with the same bit width as the program address.
[0007] With existing processor technology, if the pipeline is long, the program address corresponding to each instruction will be temporarily stored. This approach has obvious disadvantages in the design iterations of superscalar processors with increasing instruction fetch and decoding widths and increasing pipeline stages:
[0008] 1. As the number of pipeline stages increases, the number of unretired instructions increases. The depth of the program address buffer needs to be designed to the number of instructions when the reorder buffer is fully loaded. The number of table entries implemented in hardware is large, and the area cost is high.
[0009] 2. Assume that the instruction fetch and decode width is W. After W increases, at least W program addresses need to be written to the program address buffer per cycle. The buffer table entries are consumed quickly, which is wasteful.
[0010] 3. Because more hardware resources are used, the above two points will indirectly increase the power consumption of the chip.
[0011] Assuming that I instructions need to be processed, there will be a total of I program address buffer table entry write and I program address buffer table entry release actions, the number of operations is relatively large, increasing the power consumption of the chip.
[0012] In reality, the program addresses of adjacent instructions are often very close together. For example, in the Risc (Reduced Instruction Set Computer) instruction set architecture, instructions generally occupy 2 or 4 bytes, so the program addresses of adjacent instructions often differ by only 1 to 2 bits. However, if each instruction records the entire program address, there will be a large number of table entries recording the same information (because the program is compiled into a single instruction, most of the bit fields of the program address are equal). The program address buffer designed to temporarily store program addresses has a large number of table entries, resulting in a large chip area overhead. If each instruction in the processor's life cycle involves writing and releasing program address buffer table entries, chip power consumption will also increase.
[0013] The present invention proposes a processor program address buffering method and device to solve the problem. Summary of the Invention
[0014] The present invention mainly innovates the storage format and management method of the program address buffer, reducing the resource overhead of the processor for temporarily storing program addresses while maintaining the original processor instruction processing bandwidth, and the frequency of accessing this temporary storage area to reduce chip power consumption, thereby overcoming the problems in the above-mentioned background technology.
[0015] Based on the above technical ideas, the technical solution adopted by the present invention is:
[0016] A processor program address buffer device comprises the following components:
[0017] Instruction fetch unit, program address comparison unit, program address buffer, program address buffer tail value register, instruction issue and execution unit and reorder buffer;
[0018] The instruction fetch unit controls and connects to the program address comparison unit, the program address comparison unit controls and connects to the program address buffer, the program address buffer tail value register, the instruction issuance and execution unit, and the reorder buffer, the program address buffer and the program address buffer tail value register communicate with the program address comparison unit, the instruction issuance controls and connects to the execution unit and the reorder buffer, and the reorder buffer communicates with the program address buffer.
[0019] A processor program address buffering method comprises the following steps:
[0020] S1 instruction fetch and address processing step, which includes the instruction fetch unit output instruction packet link and the program address comparison unit processing flow link;
[0021] S2 instruction emission and execution step, which includes the instruction emission unit processing link and the address splicing and execution link;
[0022] S3 instruction retirement and resource release step, which includes the reorder buffer retirement determination link and the reorder buffer entry release link.
[0023] Further limitation of the above technical solution is the S1 instruction fetch and address processing step, in which the instruction fetch unit outputs the instruction packet, including taking out multiple instructions from the memory in program order, and sending the operation code and complete program address of each instruction to the program address comparison unit.
[0024] A further limitation of the above technical solution is that the processing flow of the program address comparison unit includes address splitting and selection of comparison benchmark. Address splitting includes splitting the complete program address of each instruction into two parts, a high-bit field and a low-bit field. Selecting the comparison benchmark includes, if the program address buffer is empty, using the high-bit field of the first instruction as the initial comparison benchmark, and allocating indexes starting from the head of the free table entry in the program address buffer. If the program address buffer is not empty, using the high-bit field stored in the tail value register as the initial comparison benchmark, and allocating indexes starting from the next position of the current tail table entry in the buffer.
[0025] A further limitation of the above technical solution is that the S1 instruction fetch and address processing step also includes an instruction-by-instruction high-bit comparison and index allocation link. The instruction-by-instruction high-bit comparison and index allocation link includes comparing the high-bit field of each instruction with the current benchmark in sequence according to the instruction program order. If the high-bit field is the same as the benchmark, the table entry index corresponding to the current benchmark is assigned to the instruction. If the high-bit field is different from the benchmark, a new table entry index is assigned to the instruction, and the high-bit field of the instruction is written into the new table entry position of the program address buffer. The current benchmark is updated to the high-bit field of the instruction, and the tail value register is synchronously updated to the high-bit field.
[0026] To further limit the above technical solution, the S2 instruction emission and execution step, in which the instruction emission unit processing link includes receiving the instruction data packet sent by the program address comparison unit, decoding the operation code, and identifying the instruction that requires the program address as an operand.
[0027] A further limitation of the above technical solution is that the address splicing and execution link includes using the table entry index carried by the instruction, reading the corresponding high-bit field from the program address buffer, splicing the high-bit field with the low-bit field of the instruction, restoring it to a complete program address, sending the complete program address as an operand to the execution unit, executing the instruction according to the opcode and operand, and after the instruction is executed, sending an execution completion flag to the reordering buffer.
[0028] A further limitation of the above technical solution is the S3 instruction retirement and resource release step, in which the reordering buffer retirement determination link includes scanning all non-retired instructions in program order, finding a continuous and executed instruction sequence, and extracting the table entry index of the last instruction in the continuous sequence.
[0029] A further limitation of the above technical solution is that the reordering buffer retirement determination link also includes checking the table entry indexes of all non-retired instructions, and recording the table entry index of the last instruction as J. If there is any non-retired instruction with an index ≤ J, the program address buffer table entry will not be released. If the indexes of all non-retired instructions are > J, all table entries with indexes ≤ J in the program address buffer will be released.
[0030] As a further limitation of the above technical solution, the step of releasing the reorder buffer entries includes marking the reorder buffer entries corresponding to the retired instructions as free, clearing the entry occupation flag, and clearing the instruction submission flag.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] 1. The program address comparison unit provides a basis for determining whether to write the program address into the program address buffer, rather than writing all program addresses of all instructions into the program address buffer. Therefore, the number of implementation table entries in the program address buffer can be reduced.
[0033] 2. The program address buffer only temporarily stores the high-bit field of the program address, so the number of implementation table entries of the program address buffer can be further reduced.
[0034] 3. Due to advantage 1, the frequency of writing to the program address buffer becomes lower, which can reduce chip power consumption.
[0035] 4. The reorder buffer can retire multiple instructions at the same time. Each instruction recorded in the reorder buffer records the corresponding program address table entry index, which provides a judgment basis for the reorder buffer to release the program address buffer table entry when retiring instructions. The number of table entries released by the program address buffer is not equal to, but less than or equal to, the number of instructions currently needing to be retired. Therefore, the frequency of releasing table entries in the program address buffer is reduced, which can reduce chip power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 A schematic diagram of the interaction between a program address buffer and other components of a method and apparatus for program address buffering of a processor according to the present invention;
[0038] Figure 2 A flowchart of program address buffer instructions and program address processing of a method and device for program address buffering of a processor according to the present invention;
[0039] Figure 3 A flowchart of a method and device for reordering buffer instructions and program addresses in a processor program address buffer according to the present invention;
[0040] Figure 4The present invention provides a single table entry format diagram in a reordering buffer of a processor program address buffering method and device.
[0041] Among them: 1. Instruction fetch unit; 2. Program address comparison unit; 3. Program address buffer; 4. Program address buffer tail value register; 5. Instruction issue and execution unit; 6. Reorder buffer. DETAILED DESCRIPTION
[0042] The following is combined with Figure 1-Figure 4 The present invention is described in further detail.
[0043] Embodiment 1: This embodiment provides a processor program address buffer device, such as Figure 1-Figure 4 As shown, it includes the following components:
[0044] Instruction fetch unit 1, program address comparison unit 2, program address buffer 3, program address buffer tail value register 4, instruction issue and execution unit 5 and reorder buffer 6;
[0045] The instruction fetch unit 1 is controlled and connected to the program address comparison unit 2, the program address comparison unit 2 is controlled and connected to the program address buffer 3, the program address buffer tail value register 4, the instruction issuance and execution unit 5 and the reorder buffer 6, the program address buffer 3 and the program address buffer tail value register 4 are communicatively connected to the program address comparison unit 2, the instruction issuance is controlled and connected to the execution unit 5 and the reorder buffer 6, and the reorder buffer 6 is communicatively connected to the program address buffer 3.
[0046] A processor program address buffering method comprises the following steps:
[0047] S1 instruction fetch and address processing step, which includes the instruction fetch unit 1 outputting the instruction packet link and the program address comparison unit 2 processing flow link;
[0048] S2 instruction emission and execution step, which includes the instruction emission unit processing link and the address splicing and execution link;
[0049] S3 is an instruction retirement and resource release step, which includes a reorder buffer 6 retirement determination link and a reorder buffer 6 entry release link.
[0050] S1 is an instruction fetch and address processing step. In this step, the instruction fetch unit 1 outputs an instruction packet, which includes fetching multiple instructions from the memory in program order and sending the operation code and complete program address of each instruction to the program address comparison unit 2.
[0051] The processing flow of the program address comparison unit 2 includes address splitting and selecting a comparison benchmark. Address splitting includes splitting the complete program address of each instruction into two parts, a high-bit field and a low-bit field; selecting a comparison benchmark includes if the program address buffer 3 is empty, using the high-bit field of the first instruction as the initial comparison benchmark, and allocating indexes starting from the head of the free table entry in the program address buffer 3; if the program address buffer 3 is not empty, using the high-bit field stored in the tail value register as the initial comparison benchmark, and allocating indexes starting from the next position of the current tail table entry in the buffer.
[0052] S1 instruction fetch and address processing step, which also includes an instruction-by-instruction high-bit comparison and index allocation link. The instruction-by-instruction high-bit comparison and index allocation link includes comparing the high-bit field of each instruction with the current benchmark in sequence according to the instruction program order. If the high-bit field is the same as the benchmark, the table entry index corresponding to the current benchmark is allocated to the instruction. If the high-bit field is different from the benchmark, a new table entry index is allocated to the instruction, and the high-bit field of the instruction is written into the new table entry position of the program address buffer 3. The current benchmark is updated to the high-bit field of the instruction, and the tail value register is synchronously updated to the high-bit field.
[0053] S2 is the instruction emission and execution step. In this step, the instruction emission unit processing link includes receiving the instruction data packet sent by the program address comparison unit 2, decoding the operation code, and identifying the instruction that requires the program address as an operand.
[0054] The address splicing and execution steps include using the table entry index carried by the instruction, reading the corresponding high-bit field from the program address buffer 3, splicing the high-bit field with the low-bit field of the instruction itself, restoring it to the complete program address, sending the complete program address as an operand to the execution unit, executing the instruction according to the opcode and operand, and sending the execution completion flag to the reordering buffer 6 after the instruction is executed.
[0055] S3 instruction retirement and resource release step, in which the retirement determination link of the reorder buffer 6 includes scanning all non-retired instructions in program order, finding a continuous and executed instruction sequence, and extracting the table entry index of the last instruction in the continuous sequence.
[0056] The retirement determination link of the reorder buffer 6 also includes checking the table entry indexes of all non-retired instructions, and recording the table entry index of the last instruction as J. If there is any non-retired instruction with an index ≤ J, the table entry of the program address buffer 3 will not be released. If the indexes of all non-retired instructions are > J, all table entries with indexes ≤ J in the program address buffer 3 will be released.
[0057] The step of releasing the reorder buffer 6 entries includes marking the reorder buffer 6 entries corresponding to the retired instructions as free, clearing the entry occupied flag, and clearing the instruction submitted flag.
[0058] Embodiment 2: This embodiment provides a method and device for processor program address buffering, such as Figure 1-Figure 4 As shown, it includes the following components:
[0059] Instruction fetch unit 1: sends multiple instruction opcodes and their corresponding program addresses arranged in program order to program address comparison unit 2.
[0060] Program address comparison unit 2 selects an initial comparison benchmark based on the valid instruction program address obtained by instruction fetch unit 1 and the empty / full status of the current program address buffer 3 entry provided by program address buffer 3. It then compares the currently received program addresses in sequence and generates program address buffer 3 write requests equal to or less than the number of currently received instructions. It also generates the program address entry index corresponding to the currently received instruction and sends the program address write request to program address buffer 3. It also sends information about the instruction to be executed to instruction issue and execution unit 5, sends information about the currently received valid instruction to reorder buffer 6, and updates program address buffer tail value register 4.
[0061] Program Address Buffer 3: This unit temporarily stores program addresses and consists of multiple entries arranged in first-in, first-out order. Each entry contains the high-order bit field of the program address. Program Address Buffer 3 entries are requested and written by Program Address Comparison Unit 2 and released by Reorder Buffer 6. The Instruction Issue and Execution Unit 5 can read the contents of Program Address Buffer 3 entries but does not modify or release any entries.
[0062] Program address buffer tail value register 4: stores the value of the last entry in the first-in-first-out order in the program address buffer 3 at each moment. Its value is updated by the program address comparison unit 2.
[0063] Instruction Issue and Execution Unit 5: Receives the opcode, low-order program address field, and program address table entry index of a valid instruction. The instruction issue unit decodes the instruction opcode and sends it to the execution unit that can execute the instruction. If the instruction requires a program address as an operand, the high-order program address field is read from the program address buffer 3 using the program address table entry index corresponding to the instruction. This is then concatenated with the received low-order program address field and sent to the execution unit. The execution unit is responsible for executing the instruction according to the opcode and operand and submitting an execution completion flag to the reorder buffer 6.
[0064] Reorder buffer 6: records instructions that have been fetched but not yet retired, their program order, and their corresponding program address table entry indexes, and retires one or more instructions at a time in program order after the execution unit submits the execution completion flag to it out of order.
[0065] The specific working principle is as follows: Take the processing flow of a certain instruction (which requires a program address as an operand) as an example to illustrate:
[0066] The instruction fetch unit 1 fetches the operation codes of multiple instructions from the memory, arranges them in sequence according to the program order, and sends the instruction operation codes and their corresponding program addresses arranged in the program order to the program address comparison unit 2.
[0067] The program address comparison unit 2 receives the operation code and program address of the valid instruction obtained by the instruction fetch unit 1, and selects the value of the program address buffer tail value register 4 or the high bit field of the program address of the first instruction program sequence currently received as the initial comparison benchmark according to the empty or full status of the current program address buffer 3 table entry provided by the program address buffer 3, and compares the currently received program addresses with each other in sequence to generate a program address buffer 3 write request that is less than or equal to the number of instructions currently received (for specific comparison methods, please refer to Figure 2 ). The program address comparison unit 2 generates the program address table entry index corresponding to the currently received instruction (for details on how to assign an index to each instruction, please refer to Figure 2 ), sends the program addresses to be written and their program address table entry indexes to the program address buffer 3; sends the opcode, program address low-bit field and program address table entry index of the currently received valid instruction to the instruction emission and execution unit 5; sends the opcode and program address table entry index of the currently received valid instruction to the reorder buffer 6; writes the program address high-bit field of the last request of the current write request into the program address buffer tail value register 4.
[0068] The program address buffer 3 receives the write request from the program address comparison unit 2 , finds the program address table entry according to the program address table entry index of each write request, and writes the high-bit field of the program address corresponding to the request into the entry.
[0069] After receiving the instruction's opcode, program address low-order field, and program address table entry index from the program address comparison unit 2, the instruction issue and execution unit 5 identifies instructions that require program addresses as operands. Using the program address table entry indexes corresponding to these instructions, the instruction issue and execution unit 5 reads the program address high-order field from the program address buffer 3, concatenates it with the received program address low-order field, and sends the result to the execution unit. The execution unit is responsible for executing the instruction according to the opcode and operand and submitting an execution completion flag to the reorder buffer 6.
[0070] The reorder buffer 6 receives the instruction information from the program address comparison unit 2 before the instruction is issued and executed. Once the instruction information is received, the reorder buffer 6 will register the reorder buffer 6 table entries and write the instruction information according to their program order (see the table entry format for details). Figure 4), after the execution unit submits the execution completion flag to it out of order, it retires one or more instructions at a time in the program order (for details on how to retire and release the program address buffer 3 table entries and the reorder buffer 6 table entries, please refer to Figure 3 ).
[0071] 1. The program address is divided into a high-bit field and a low-bit field. This distinguishes the infrequently changing portion (the high-bit field) from the frequently changing portion (the low-bit field). This provides a basis for determining whether to write to program address buffer 3, rather than writing all instruction program addresses. The patent does not limit how this distinction is made. It can be defined based on program characteristics. For example, if most programs fit within a single page (4096 bytes), the low-bit field can be defined as bits 0 to 11, and the high-bit field can be defined as 12 bits or greater.
[0072] 2. The program address buffer 3 only temporarily stores the high-bit field of the program address.
[0073] 3. When the program address comparison unit 2 receives a program address write request, it filters out program addresses that do not need to be written into the program address buffer 3 by comparing consecutive program addresses.
[0074] 4. Since the program address buffer 3 is a centralized storage resource in the processor, the program address comparison unit 2 needs to generate an entry index (program address entry index) for the high-bit field of the program address corresponding to each instruction in the program address buffer 3 and send it to the instruction issuance and execution unit 5. The instruction issuance and execution unit 5 can read the high-bit field of the program address from the program address buffer 3 using the program address entry index.
[0075] 5. The reorder buffer 6 can retire multiple instructions at the same time. Each instruction recorded in the reorder buffer 6 records the corresponding program address table entry index, which provides a judgment basis for the reorder buffer 6 to release the program address buffer 3 table entries when retiring instructions. The number of table entries released by the program address buffer 3 is not equal to, but less than or equal to, the number of instructions that currently need to be retired.
[0076] Embodiment 3: This embodiment provides a method and device for processor program address buffering, such as Figure 1-Figure 4 As shown, for Example 2, an improved solution is also included:
[0077] 1. The program address buffer tail value register 4 is not essential because it is merely a copy of the high-bit field of the last program address in the first-in-first-out order in the program address buffer 3. Therefore, if the program address buffer 3 is not completely empty, the last program address (high-bit field) in the first-in-first-out order can be directly extracted from the program address buffer 3 as the initial comparison reference.
[0078] 2. The division of the high-bit field and the low-bit field of the program address may be different in different processors and may be adjusted according to the characteristics of the actual program instruction space.
[0079] The above contents are further detailed descriptions of the present invention in conjunction with specific preferred embodiments, so as to facilitate those skilled in the art to understand and apply the present invention. It should not be considered that the specific implementation of the present invention is limited to these descriptions.
Claims
1. A processor program address buffer device, characterized in that: Includes the following components: Instruction fetch unit, program address comparison unit, program address buffer, program address buffer tail value register, instruction issue and execution unit and reorder buffer; The instruction fetch unit controls and connects to the program address comparison unit, the program address comparison unit controls and connects to the program address buffer, the program address buffer tail value register, the instruction issuance and execution unit, and the reorder buffer, the program address buffer and the program address buffer tail value register communicate with the program address comparison unit, the instruction issuance controls and connects to the execution unit and the reorder buffer, and the reorder buffer communicates with the program address buffer.
2. A processor program address buffering method, characterized in that: The following steps are involved: S1 instruction fetch and address processing step, which includes the instruction fetch unit output instruction packet link and the program address comparison unit processing flow link; S2 instruction emission and execution step, which includes the instruction emission unit processing link and the address splicing and execution link; S3 instruction retirement and resource release step, which includes the reorder buffer retirement determination link and the reorder buffer entry release link.
3. A processor program address buffering method according to claim 2, characterized in that: The S1 instruction fetch and address processing step includes the step of outputting an instruction packet by the instruction fetch unit, which includes fetching multiple instructions from the memory in program order and sending the operation code and complete program address of each instruction to the program address comparison unit.
4. A processor program address buffering method according to claim 3, characterized in that: The processing flow of the program address comparison unit includes address splitting and selecting a comparison benchmark. Address splitting includes splitting the complete program address of each instruction into two parts, a high-bit field and a low-bit field. Selecting a comparison benchmark includes, if the program address buffer is empty, using the high-bit field of the first instruction as the initial comparison benchmark, and allocating an index starting from the head of the free table entry in the program address buffer; if the program address buffer is not empty, using the high-bit field stored in the tail value register as the initial comparison benchmark, and allocating an index starting from the next position of the current tail table entry in the buffer.
5. A processor program address buffering method according to claim 4, characterized in that: The S1 instruction fetch and address processing step also includes an instruction-by-instruction high-bit comparison and index allocation link. The instruction-by-instruction high-bit comparison and index allocation link includes comparing the high-bit field of each instruction with the current benchmark in sequence according to the instruction program order. If the high-bit field is the same as the benchmark, the table entry index corresponding to the current benchmark is allocated to the instruction. If the high-bit field is different from the benchmark, a new table entry index is allocated to the instruction, and the high-bit field of the instruction is written into the new table entry position of the program address buffer. The current benchmark is updated to the high-bit field of the instruction, and the tail value register is synchronously updated to the high-bit field.
6. A processor program address buffering method according to claim 5, characterized in that: The S2 instruction emission and execution step includes the instruction emission unit processing link including receiving the instruction data packet sent by the program address comparison unit, decoding the operation code, and identifying the instruction that requires the program address as an operand.
7. A processor program address buffering method according to claim 6, characterized in that: The address splicing and execution link includes using the table entry index carried by the instruction, reading the corresponding high-bit field from the program address buffer, splicing the high-bit field with the low-bit field of the instruction, restoring it to a complete program address, sending the complete program address as an operand to the execution unit, executing the instruction according to the opcode and operand, and sending an execution completion flag to the reordering buffer after the instruction is executed.
8. A processor program address buffering method according to claim 7, characterized in that: The S3 instruction retirement and resource release step, in which the reordering buffer retirement determination link includes scanning all non-retired instructions in program order, finding a continuous and executed instruction sequence, and extracting the table entry index of the last instruction in the continuous sequence.
9. A processor program address buffering method according to claim 7, characterized in that: The reorder buffer retirement determination link also includes checking the table entry indexes of all non-retired instructions, and recording the table entry index of the last instruction as J. If there is any non-retired instruction with an index ≤ J, the program address buffer table entry will not be released. If the indexes of all non-retired instructions are > J, all table entries with indexes ≤ J in the program address buffer will be released.
10. A processor program address buffering method according to claim 7, characterized in that: The step of releasing the reorder buffer table entry includes marking the reorder buffer table entry corresponding to the retired instruction as idle, clearing the entry occupation flag, and clearing the instruction submission flag.
Citation Information
Cited By
Layered reordering buffer device, processor and management method thereof
CN120821502A
A hierarchical reordering buffer device, processor, and management method thereof
CN120821502B