Processor pipeline device, branch instruction processing method and storage medium
By introducing branch heap units into the processor pipeline, cache branch instructions and their execution results, and performing out-of-order or sequential updates, the problem that the processor pipeline takes too long to process predicted error branch instructions is solved, and more efficient branch instruction execution and application execution speed is achieved.
Patent Information
- Application Number
- CN202510495593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing processor pipeline takes too long to process branch instructions that predict errors, resulting in low execution efficiency of branch instructions and application execution speed.
A processor pipeline device is designed, including a front-end processing module and a back-end processing module, and a branch heap unit is introduced into the back-end processing module. The branch heap unit caches branch instructions and their execution results, and updates the front-end branch prediction unit out of order or sequentially as needed.
Through the design of branch heap units, the submission efficiency of branch instructions can be accelerated, the flow time of branch instructions on the pipeline can be reduced, the processor power consumption can be reduced, and the execution efficiency of branch instructions and the execution speed of application can be improved.
Smart Images

Figure CN120010926A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to the field of processor technology, and in particular relates to a processor pipeline device, a branch instruction processing method and a storage medium. Background Art
[0002] The central processing unit (CPU), also commonly referred to as a processor, is the core component of a processor-based system. Processors appear in areas including but not limited to mobile communication devices, intelligent computing terminals, intelligent driving control, and smart home appliances, and are widely used in various computing devices to perform computing tasks for various applications. With the explosive growth of artificial intelligence and the deepening of the concept of the Internet of Things, processors with both high-performance computing and low power consumption have become the cornerstone of the control center sought after in various fields.
[0003] In the face of increasingly complex usage scenarios and real-time control requirements, processors often need to run more complex applications than before. In these complex control applications, branch instructions will account for a higher proportion of the entire program instructions. Branch instructions can change the computer's default behavior of processing instructions in sequence, allowing the program to start processing a different instruction sequence at a branch target address that is different from the next instruction after the branch instruction.
[0004] Most modern processors use an out-of-order emission architecture. After branch instructions and other computational instructions complete the instruction fetch and decoding stages in the processor pipeline, they will be sent out-of-order to the processing units downstream of the pipeline. The out-of-order executed instructions are then recovered and submitted in the order of program execution through the reorder buffer (ROB). In order to improve the execution efficiency of branch instructions, a combination of front-end branch prediction and back-end branch execution and update is usually adopted. The accuracy of the front-end branch prediction depends on the timely feedback of the correct result after the back-end branch instruction is executed. Generally, incorrect prediction of branch instructions requires flushing the processor pipeline and restarting it.
[0005] However, if Figure 1 As shown, if a mispredicted branch instruction appears, it needs to wait for the previous operation instruction to be submitted in the correct program order before it can be submitted and update the front-end predictor. This results in the predictor using the wrong prediction history for subsequent branch predictions during the period before the pipeline flush is completed, thus wasting more time and affecting the efficient execution of branch instructions and the execution speed of the entire application. Summary of the invention
[0006] The present invention provides a processor pipeline device, a branch instruction processing method and a storage medium, aiming to solve the problem that the existing processor pipeline consumes too much time when processing mispredicted branch instructions.
[0007] In order to solve the above technical problems, in a first aspect, the present invention provides a processor pipeline device, the processor pipeline device comprising a front-end processing module and a back-end processing module, wherein: The front-end processing module includes an instruction fetch unit, a branch prediction unit and a decoding unit, wherein the instruction fetch unit is used to read instructions from a memory; the branch prediction unit is used to perform branch prediction according to the instructions to obtain branch instructions corresponding to the instructions; the decoding unit is used to decode the branch instructions and send the decoded branch instructions to the back-end processing module in a disordered order; The back-end processing module includes an out-of-order execution pipeline unit, a branch execution unit and a branch stack unit. The out-of-order execution pipeline unit is used to store the branch instructions sent by the front-end processing module, and send the branch instructions to the branch stack unit in sequence in a queue; the branch execution unit is used to execute the branch instructions; the branch stack unit is used to cache the branch instructions and the execution results generated by the branch execution unit executing the branch instructions, and return the corresponding branch instructions to the branch prediction unit.
[0008] Furthermore, the branch stack unit is also used for: The branch instructions sent by the front-end processing module are obtained, and the branch instructions are cached according to the predicted execution order of the branch instructions. Meanwhile, a corresponding branch heap space index is established for each cached branch instruction.
[0009] Furthermore, the branch stack unit is also used for: According to the branch heap space index, the corresponding branch instruction is read in the cache, and the branch instruction is sent to the branch execution unit for instruction execution.
[0010] Furthermore, the branch stack unit is also used for: The execution result generated by the branch execution unit executing the branch instruction is obtained, and the execution result is cached in the same cache space as the branch instruction according to the branch heap space index corresponding to the branch instruction.
[0011] Furthermore, the branch stack unit is also used for: Determine whether the cached branch instruction has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space.
[0012] Furthermore, the branch stack unit is also used for: Determine whether the cached branch instruction with the oldest cache order has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space.
[0013] In a second aspect, the present invention further provides a branch instruction processing method, the branch instruction processing method is implemented based on the processor pipeline device as described above, and the branch instruction processing method comprises the following steps: The instruction is read through the instruction fetch unit and the branch prediction unit of the front-end processing module, and the branch instruction predicted by the instruction is obtained through branch prediction, and the branch instruction is decoded by the decoding unit, and then the branch instruction is sent to the back-end processing module; The branch instructions sent by the front-end processing module are stored in the out-of-order execution pipeline unit of the back-end processing module, and at the same time, the branch instructions are cached in the branch stack unit according to the predicted execution order; The branch execution unit executes the branch instruction and obtains a corresponding execution result; The execution result corresponding to the branch instruction is cached by the branch stack unit, and the branch instruction with the execution result is returned to the branch prediction unit.
[0014] Furthermore, the execution result corresponding to the branch instruction is cached by the branch stack unit, and the branch instruction with the execution result is returned to the branch prediction unit; Determine whether the cached branch instruction has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space; or: Determine whether the cached branch instruction with the oldest cache order has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space.
[0015] In a third aspect, the present invention also provides a computer device, comprising: a memory, a processor, and a branch instruction processing program stored in the memory and executable on the processor, wherein when the processor executes the branch instruction processing program, the steps in the branch instruction processing method as described in any one of the above embodiments are implemented.
[0016] In a fourth aspect, the present invention further provides a storage medium on which a branch instruction processing program is stored. When the branch instruction processing program is executed by a processor, the steps of the branch instruction processing method described in any one of the above embodiments are implemented.
[0017] The beneficial effect achieved by the present invention is that a processor pipeline device with a branch stack unit structure is proposed, and the branch stack unit of the processor pipeline device can implement reordering buffering for branch instructions. Branch instructions can be executed out of order through the branch stack unit, and can also perform out of order or sequential updates on the front-end branch prediction unit as needed, so that the processor can switch between high performance and high energy efficiency scenarios. At the same time, the branch stack unit can speed up the submission efficiency of branch instructions, reduce the time for branch instructions to flow on the pipeline, and reduce processor power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a logical schematic diagram of the instruction sequence update in the prior art; Figure 2 is a schematic diagram of the structure of a processor pipeline device provided by an embodiment of the present invention; Figure 3 It is a logical schematic diagram of out-of-order update of branch heap units in a processor pipeline device provided by an embodiment of the present invention; Figure 4 It is a logical schematic diagram of sequential updating of branch heap units in a processor pipeline device provided by an embodiment of the present invention; Figure 5 is a flowchart of the steps of the branch instruction processing method provided by an embodiment of the present invention; Figure 6 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0020] Please refer to Figure 2 , Figure 2 1 is a schematic diagram of the structure of a processor pipeline device provided in an embodiment of the present invention. The processor pipeline device 100 includes a front-end processing module 101 and a back-end processing module 102, wherein: The front-end processing module 101 includes an instruction fetch unit 1011, a branch prediction unit 1012 and a decoding unit 1013, wherein the instruction fetch unit 1011 is used to read instructions from a memory; the branch prediction unit 1012 is used to perform branch prediction according to the instructions to obtain branch instructions corresponding to the instructions; the decoding unit 1013 is used to decode the branch instructions and send the decoded branch instructions to the back-end processing module 102 in a disordered order; The back-end processing module 102 includes an out-of-order execution pipeline unit 1021, a branch execution unit 1022 and a branch stack unit 1023. The out-of-order execution pipeline unit 1021 is used to store the branch instructions sent by the front-end processing module 101, and send the branch instructions to the branch stack unit 1023 in sequence in a queue; the branch execution unit 1022 is used to execute the branch instructions; the branch stack unit 1023 is used to cache the branch instructions and the execution results generated by the branch execution unit 1022 executing the branch instructions, and return the corresponding branch instructions to the branch prediction unit 1012.
[0021] Different from the structure of the existing processor pipeline, the processor pipeline device 100 in the embodiment of the present invention is designed with a processor module of a branch file unit 1023, which is used to implement a function similar to a reorder buffer (ROB) in the existing processor pipeline for branch instructions.
[0022] Specifically, in order to realize the function of recording the decoded branch instructions separately in sequence, the branch stack unit 1023 is further used to: The branch instructions sent by the front-end processing module 101 are obtained, and are cached according to the predicted execution order of the branch instructions. Meanwhile, a corresponding branch heap space index is established for each cached branch instruction.
[0023] The function of recording branch instructions in sequence is mainly to select one or more branch instructions from the multi-way transmission instructions of each cycle, apply for space in the branch stack, write the branch instruction into the applied space according to the order of the program, and return the corresponding branch stack space index. It can be understood that the branch stack unit 1023 in the embodiment of the present invention is similar to the reorder buffer, has a certain cache space, and can support the cache of branch instruction data when the cache space is free. Therefore, each time a branch instruction is input, the branch stack unit 1023 is also used to: Determine whether the current pipeline is flushed.
[0024] When the pipeline is flushed, the cache contents of the branch heap unit 1023 are cleared; If pipeline flushing is not involved, data writing to the branch heap unit 1023 is involved, and it is necessary to further determine whether the cache space of the branch heap unit 1023 is free.
[0025] In order to realize the function of reading branch instruction information out of order when the branch instruction is executed, the branch stack unit 1023 is also used for: According to the branch heap space index, the corresponding branch instruction is read in the cache, and the branch instruction is sent to the branch execution unit 1022 for instruction execution.
[0026] In order to implement the function of updating the execution result of the branch instruction out of order after the branch instruction is executed, the branch stack unit 1023 is also used for: The execution result generated by the branch execution unit 1022 executing the branch instruction is obtained, and the execution result is cached in the same cache space as the branch instruction according to the branch heap space index corresponding to the branch instruction.
[0027] In order to implement the function of updating the front-end module out of order according to the execution result of the branch instruction, the branch stack unit 1023 is also used for: Determine whether the cached branch instruction has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit 1012, and the branch instruction and its corresponding execution result are removed from the cache space.
[0028] like Figure 3 As shown, in the out-of-order update mode, as long as the branch stack unit 1023 stores a branch instruction with a written-back branch result, it is read out and updated to the front-end branch prediction unit 1012. After reading, the branch instruction leaves the cache of the branch stack unit 1023.
[0029] In order to realize the function of sequentially updating the front-end module according to the execution results of the branch instructions, the branch stack unit 1023 is also used for: Determine whether the cached branch instruction with the oldest cache order has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit 1012, and the branch instruction and its corresponding execution result are removed from the cache space.
[0030] like Figure 4 As shown, in the sequential update mode, each update needs to find out whether the oldest branch instruction in the branch stack unit 1023 (written into the cache of the branch stack unit in the write order) is written back by the result. If it is written back, it will be read out and updated to the branch prediction unit 1012; otherwise, it will continue to wait until the branch execution unit 1022 writes back the execution result corresponding to the oldest branch instruction.
[0031] It can be understood that the out-of-order update or sequential update implemented by the branch stack unit 1023 in the above functions can be logically switched with each other and will not affect the processing flow of the pipeline. Therefore, during the implementation process, one of them can be used according to needs or the performance of the processor itself, so that the processor can switch between high performance and high energy efficiency scenarios.
[0032] Among the functions that can be implemented by the above-mentioned branch stack unit 1023, the processes involving branch instruction writing, branch result writing, and branch result reading should all be designed with certain logic to determine whether the current pipeline has been flushed. Any pipeline flushing behavior should be regarded as a prediction error of the branch instruction executed by the branch execution unit 1022 within that cycle. Therefore, in order to improve feedback efficiency, when a pipeline flush is detected, the cache of the current branch stack unit 1023 is first cleared so that subsequent branch instructions will no longer be read out to the branch execution unit 1022.
[0033] The beneficial effect achieved by the present invention is that a processor pipeline device with a branch stack unit structure is proposed, and the branch stack unit of the processor pipeline device can implement reordering buffering for branch instructions. Branch instructions can be executed out of order through the branch stack unit, and can also perform out of order or sequential updates on the front-end branch prediction unit as needed, so that the processor can switch between high performance and high energy efficiency scenarios. At the same time, the branch stack unit can speed up the submission efficiency of branch instructions, reduce the time for branch instructions to flow on the pipeline, and reduce processor power consumption.
[0034] The embodiment of the present invention further provides a branch instruction processing method, which is implemented based on the processor pipeline device 100 as described above. Figure 5 , Figure 5 : is a flowchart of the steps of a branch instruction processing method provided by an embodiment of the present invention, wherein the branch instruction processing method comprises the following steps: S201, reading an instruction through an instruction fetch unit and a branch prediction unit of a front-end processing module, obtaining a branch instruction predicted by the instruction through branch prediction, decoding the branch instruction through a decoding unit, and sending the branch instruction to a back-end processing module; S202, storing the branch instruction sent by the front-end processing module through the out-of-order execution pipeline unit of the back-end processing module, and at the same time, caching the branch instruction in the branch stack unit according to the predicted execution order; S203, the branch execution unit executes the branch instruction and obtains a corresponding execution result; S204: Cache the execution result corresponding to the branch instruction through the branch stack unit, and return the branch instruction with the execution result to the branch prediction unit.
[0035] Furthermore, S204, caching the execution result corresponding to the branch instruction through the branch stack unit, and returning the branch instruction with the execution result to the branch prediction unit; S2041, determining whether the cached branch instruction has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space; or: S2042: Determine whether the cached branch instruction with the oldest cache order has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space.
[0036] The branch instruction processing method is implemented based on the processor pipeline device 100 in the above embodiment and can achieve the same technical effect. Please refer to the description in the above embodiment and will not be repeated here.
[0037] The embodiment of the present invention further provides a computer device 300, please refer to Figure 6 , Figure 6300 is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. The computer device 300 includes: a memory 302, a processor 301, and a branch instruction processing program stored in the memory 302 and executable on the processor 301.
[0038] The processor 301 calls the branch instruction processing program stored in the memory 302 to execute the steps of the branch instruction processing method provided in the embodiment of the present invention. Figure 5 , specifically including the following steps: S201, reading an instruction through an instruction fetch unit and a branch prediction unit of a front-end processing module, obtaining a branch instruction predicted by the instruction through branch prediction, decoding the branch instruction through a decoding unit, and sending the branch instruction to a back-end processing module; S202, storing the branch instruction sent by the front-end processing module through the out-of-order execution pipeline unit of the back-end processing module, and at the same time, caching the branch instruction in the branch stack unit according to the predicted execution order; S203, the branch execution unit executes the branch instruction and obtains a corresponding execution result; S204: Cache the execution result corresponding to the branch instruction through the branch stack unit, and return the branch instruction with the execution result to the branch prediction unit.
[0039] Furthermore, S204, caching the execution result corresponding to the branch instruction through the branch stack unit, and returning the branch instruction with the execution result to the branch prediction unit; S2041, determining whether the cached branch instruction has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space; or: S2042: Determine whether the cached branch instruction with the oldest cache order has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space.
[0040] The computer device 300 provided in the embodiment of the present invention can implement the steps in the branch instruction processing method in the above embodiment, and can achieve the same technical effect. Please refer to the description in the above embodiment and will not be repeated here.
[0041] An embodiment of the present invention also provides a storage medium, on which a branch instruction processing program is stored. When the branch instruction processing program is executed by a processor, the various processes and steps in the branch instruction processing method provided by the embodiment of the present invention are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0042] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0043] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0044] The embodiments of the present invention are described above in conjunction with the accompanying drawings. What is disclosed is only the preferred embodiment of the present invention. However, the present invention is not limited to the above-mentioned specific implementation manner. The above-mentioned specific implementation manner is only illustrative rather than restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms and equivalent changes without departing from the scope of protection of the purpose of the present invention and the claims, all of which are within the protection of the present invention.
Claims
1. A processor pipeline device, characterized in that: The processor pipeline device includes a front-end processing module and a back-end processing module, wherein: The front-end processing module includes an instruction fetch unit, a branch prediction unit and a decoding unit, wherein the instruction fetch unit is used to read instructions from a memory; the branch prediction unit is used to perform branch prediction according to the instructions to obtain branch instructions corresponding to the instructions; the decoding unit is used to decode the branch instructions and send the decoded branch instructions to the back-end processing module in a disordered order; The back-end processing module includes an out-of-order execution pipeline unit, a branch execution unit and a branch stack unit. The out-of-order execution pipeline unit is used to store the branch instructions sent by the front-end processing module, and send the branch instructions to the branch stack unit in sequence in a queue; the branch execution unit is used to execute the branch instructions; the branch stack unit is used to cache the branch instructions and the execution results generated by the branch execution unit executing the branch instructions, and return the corresponding branch instructions to the branch prediction unit.
2. The processor pipeline device according to claim 1, characterized in that: The branch stack unit is also used for: The branch instructions sent by the front-end processing module are obtained, and the branch instructions are cached according to the predicted execution order of the branch instructions. Meanwhile, a corresponding branch heap space index is established for each cached branch instruction.
3. The processor pipeline device according to claim 2, characterized in that: The branch stack unit is also used for: According to the branch heap space index, the corresponding branch instruction is read in the cache, and the branch instruction is sent to the branch execution unit for instruction execution.
4. The processor pipeline device according to claim 2, characterized in that: The branch stack unit is also used for: The execution result generated by the branch execution unit executing the branch instruction is obtained, and the execution result is cached in the same cache space as the branch instruction according to the branch heap space index corresponding to the branch instruction.
5. The processor pipeline device according to claim 4, characterized in that: The branch stack unit is also used for: Determine whether the cached branch instruction has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space.
6. The processor pipeline device according to claim 4, characterized in that: The branch stack unit is also used for: Determine whether the cached branch instruction with the oldest cache order has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space.
7. A branch instruction processing method, characterized in that: The branch instruction processing method is implemented based on the processor pipeline device according to any one of claims 1 to 6, and the branch instruction processing method comprises the following steps: The instruction is read through the instruction fetch unit and the branch prediction unit of the front-end processing module, and the branch instruction predicted by the instruction is obtained through branch prediction, and the branch instruction is decoded by the decoding unit, and then the branch instruction is sent to the back-end processing module; The branch instructions sent by the front-end processing module are stored in the out-of-order execution pipeline unit of the back-end processing module, and at the same time, the branch instructions are cached in the branch stack unit according to the predicted execution order; The branch execution unit executes the branch instruction and obtains a corresponding execution result; The execution result corresponding to the branch instruction is cached by the branch stack unit, and the branch instruction with the execution result is returned to the branch prediction unit.
8. The branch instruction processing method according to claim 7, characterized in that: The step of caching the execution result corresponding to the branch instruction through the branch stack unit, and returning the branch instruction with the execution result to the branch prediction unit; Determine whether the cached branch instruction has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space; or: Determine whether the cached branch instruction with the oldest cache order has the corresponding execution result, wherein: If so, the branch instruction with the execution result is read out from the cache space, and the branch instruction is returned to the branch prediction unit, and the branch instruction and the corresponding execution result are removed from the cache space.
9. A computer device, characterized in that: include: A memory, a processor, and a branch instruction processing program stored in the memory and executable on the processor, wherein the processor implements the steps in the branch instruction processing method as described in any one of claims 7 to 8 when executing the branch instruction processing program.
10. A storage medium, characterized in that: The storage medium stores a branch instruction processing program, and when the branch instruction processing program is executed by the processor, the steps in the branch instruction processing method according to any one of claims 7 to 8 are implemented.
Citation Information
Patent Citations
Six-stage pipeline processor based on RISC-V instruction set
CN114721724A
Instruction processing method and device, electronic equipment and computer readable storage medium
CN115080121A
Performing scour recovery using parallel traversal of slice reorder buffer (SROBS)
CN116097215A
Implementation method and device for processor front-end instruction reading queue
CN118444984A
Method and apparatus for a branch instruction pointer table
US5918046A