Instruction fetching method, processor, system-on-chip, and computing device
By receiving branch prediction information in the processor and directly obtaining jump instructions using program counter offset, the problems of latency and dynamic power consumption in the prior art are solved, achieving more efficient and energy-saving instruction fetching, which is suitable for application scenarios with high requirements for real-time performance and battery life.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ALIBABA DAMO (HANGZHOU) TECH CO LTD
- Filing Date
- 2025-08-21
- Publication Date
- 2026-06-04
AI Technical Summary
Existing processors suffer from latency and additional dynamic power consumption when processing discontinuous instructions, especially jump instructions, due to premature address comparisons, which affects processing performance and product battery life.
By receiving branch prediction information, the memory cell where the current jump instruction is located can be determined directly from the cache cell or tightly coupled memory cell, and the instruction can be retrieved based on the offset of the program counter. This avoids sequential address comparison and simultaneous fetching of two branches, reducing timing pressure and dynamic power consumption.
It improves the processor's clock speed and performance in handling non-continuous instructions, enhances battery life, and is suitable for applications with high real-time and battery life requirements.
Smart Images

Figure CN2025116247_04062026_PF_FP_ABST
Abstract
Description
Instruction fetching methods, processors, on-chip systems and computing devices
[0001] This disclosure claims priority to Chinese Patent Application No. 202411709265.9, filed with the China Patent Office on November 26, 2024, entitled “Instruction Fetching Method, Processor, System-on-Chip, and Computing Device”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the technical field of electronic hardware, and in particular to an instruction fetching method, a processor, a system-on-a-chip, and a computing device. Background Technology
[0003] With the development of semiconductor technology and processor architecture technology, modern processors not only pursue higher performance, but also need to meet the requirements of real-time performance and energy efficiency in specific application areas.
[0004] Currently, for high real-time scenarios, processors need to ensure that the processing latency of critical instructions remains within a defined range. If the processor fetches critical instructions from the cache, and if the critical instructions are not in the cache, fetching them from external memory via the Bus Interface Unit (BIU) cannot guarantee stable latency. To address this issue, by adding an independent, tightly coupled memory unit within the processor—independent yet coexisting with the cache—critical instructions can be placed there, achieving critical instruction fetching with stable latency.
[0005] However, if the critical instruction is a non-continuous instruction such as a jump instruction, it may cause a sudden change in the direction of the instruction flow. It is necessary to compare addresses in advance to determine whether the critical instruction is stored in a cache unit or a tightly coupled memory unit at the current jump, and then retrieve the critical instruction from the corresponding memory unit. This method of comparing addresses in advance and then retrieving the instruction will produce a large delay, reduce the processor's clock frequency, and affect the processor's processing performance. Summary of the Invention
[0006] In view of the above, embodiments of this disclosure provide an instruction acquisition method. One or more embodiments of this disclosure also relate to a processor, a system-on-a-chip, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0007] According to a first aspect of the present disclosure, an instruction fetching method is provided, applied to the instruction fetch unit of a processor, the processor further comprising a cache unit, a tightly coupled memory unit, and a program counter, including:
[0008] Receive branch prediction information for the current jump instruction;
[0009] Based on branch prediction information, the current memory unit where the current jump instruction is located is determined from the cache unit and tightly coupled memory unit;
[0010] The current jump instruction is retrieved from the current memory location based on the offset of the current program counter value recorded by the program counter.
[0011] According to a second aspect of the present disclosure, a processor is provided, the processor including an instruction fetch unit, a cache unit, a tightly coupled memory unit, and a program counter;
[0012] The instruction fetch unit is used to receive branch prediction information for the current jump instruction, determine the current memory location of the current jump instruction from the cache unit and the tightly coupled memory unit based on the branch prediction information, and fetch the current jump instruction from the current memory location based on the offset of the current program counter value recorded by the program counter.
[0013] According to a third aspect of the present disclosure, an on-chip system is provided, comprising:
[0014] The control unit and multiple on-chip components, including a processor;
[0015] The control unit is used to control and manage multiple on-chip components, and the processor is used to execute computer programs / instructions, which, when executed by the processor, implement the steps of the above-described instruction acquisition method.
[0016] According to a fourth aspect of the present disclosure, a computing device is provided, comprising:
[0017] Memory and system-on-a-chip;
[0018] The memory is used to store computer programs / instructions, and the system on-chip is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the system on-chip, they implement the steps of the instruction fetching method described above.
[0019] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided that stores a computer program / instructions, which, when executed by a processor, implement the steps of the above-described instruction acquisition method.
[0020] According to a sixth aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the instruction acquisition method described above.
[0021] In one embodiment of this disclosure, during the instruction fetch stage, based on branch prediction information for the current jump instruction, one of the cache units and tightly coupled memory units is directly determined as the current memory unit where the current jump instruction is located. Based on the offset of the current program counter value recorded by the program counter, the current jump instruction is fetched from the current memory unit. Address comparison and instruction fetching are no longer performed sequentially, reducing timing pressure and increasing the processor's clock frequency. At the same time, it effectively avoids the additional dynamic power consumption caused by fetching jump instructions from two branches simultaneously, improving the processor's performance and product battery life when processing non-continuous instructions. It achieves more efficient and energy-saving instruction fetching without sacrificing processing speed, and is suitable for application scenarios with high requirements for real-time performance and product battery life. Attached Figure Description
[0022] Figure 1 is a schematic diagram of the pipeline stage of instruction fetching;
[0023] Figure 2 is a flowchart illustrating a method for obtaining instructions;
[0024] Figure 3 is a flowchart illustrating another method for obtaining instructions;
[0025] Figure 4 is a flowchart of an instruction acquisition method provided in an embodiment of this disclosure;
[0026] Figure 5 is a flowchart illustrating one embodiment of an instruction acquisition method provided in this disclosure;
[0027] Figure 6 is a second schematic flowchart of an instruction acquisition method provided in an embodiment of this disclosure;
[0028] Figure 7 is a comparative schematic diagram of an instruction acquisition method provided in an embodiment of this disclosure;
[0029] Figure 8 is a schematic diagram of the structure of a processor provided in an embodiment of this disclosure;
[0030] Figure 9 is a structural block diagram of a system-on-a-chip provided in an embodiment of this disclosure;
[0031] Figure 10 is a structural block diagram of a computing device provided in an embodiment of this disclosure. Detailed Implementation
[0032] Numerous specific details are set forth in the following description to provide a full understanding of this disclosure. However, this disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this disclosure. Therefore, this disclosure is not limited to the specific implementations disclosed below.
[0033] The terminology used in one or more embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this disclosure. The singular forms “a,” “the,” and “the” as used in one or more embodiments of this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0034] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this disclosure, and similarly, second may also be referred to as first. Depending on the context, the word “if” as used herein may be interpreted as “when”, “in response to a determination”, or “when…”.
[0035] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0036] First, the terms and concepts involved in one or more embodiments of this disclosure will be explained.
[0037] Reduced Instruction Set Computing (RISC): A processor architecture design philosophy that emphasizes using a smaller number of simple instructions to build a processor in order to improve the processor's execution efficiency and speed.
[0038] Reduced Instruction Set Computing (RISC) Processor: A processor designed based on the RISC concept, characterized by a shorter instruction pipeline, fixed instruction format, and simple addressing modes, enabling high-speed processing.
[0039] Complex Instruction Set Computing (CISC): A processor architecture design philosophy that emphasizes the use of a large number of complex instructions to build the processor in order to improve the processor's functionality and flexibility.
[0040] Complex Instruction Set Processor (CISC Processor): A processor designed based on the CISC concept, characterized by having a large number of complex instructions, each of which can perform multiple operations.
[0041] A cache unit is a high-speed, small-capacity random access memory located between the processor and main memory. It temporarily stores frequently accessed data and instructions to reduce the time spent accessing main memory and improve processor efficiency. Cache units are divided into instruction cache (ICache) and data cache (DCache) to optimize instruction and data flows respectively.
[0042] Tightly Coupled Memory (TCM) is a type of random access memory directly connected to the processor. It stores critical data or instructions that require fast access and offers lower access latency than regular cache. Tightly Coupled Memory is divided into Instruction Tightly Coupled Memory (ITCM) and Data Tightly Coupled Memory (DTCM) to optimize instruction and data flows, respectively.
[0043] Static Random Access Memory (SRAM): A type of random access memory that can maintain stable data storage without the need for refresh circuitry, has a relatively fast access speed, and is often used to create caches.
[0044] Instruction Fetch Unit (IFU): A part of the processor responsible for fetching the next instruction to be executed from memory and sending it to the subsequent processing stage.
[0045] Program Counter (PC): A register that stores the address of the next instruction and guides the instruction fetch unit to fetch the next instruction from where.
[0046] Load Store Unit (LSU): Responsible for handling data loading and storage operations, that is, reading data from memory or writing data back to memory.
[0047] Execution Unit (EXU): The execution unit is responsible for executing instructions, including arithmetic and logical operations.
[0048] Arithmetic Logic Unit (ALU): This is the part of the processor that performs basic arithmetic operations (addition, subtraction, etc.) and logical operations (AND, OR, NOT, etc.).
[0049] Bus Interface Unit (BIU): Responsible for communication between the processor and external devices, including data transmission and address sending.
[0050] Branch Unit (BJU): A component unit used to process branch instructions. It is responsible for analyzing branch conditions and predicting the direction of the branch to reduce pipeline stalls caused by waiting for branch instruction results, thereby improving processor execution efficiency.
[0051] Predictor: A component unit used to predict the address of the next instruction. It analyzes historical data and current conditions to predict whether a branch instruction will jump and its target address, thereby reducing pipeline stalls caused by waiting for branch instruction results and improving processor execution efficiency.
[0052] Figure 1 shows a schematic diagram of the pipeline stage for instruction fetching.
[0053] The entire instruction lifecycle, from instruction fetching to execution, typically includes five stages: Instruction Fetch (IF), Instruction Decode (ID), Instruction Execute (EX), Memory Access (MEM), and Memory Access (MEM).
[0054] Instruction fetch phase: The first phase of the instruction cycle, which is the process of reading instructions from memory into the processor.
[0055] Decoding stage: The process of converting the fetched instructions into control signals to determine the specific operation of the instruction and the required operands.
[0056] Execution phase: Perform specific operations based on the decoding results, such as calculations and comparisons.
[0057] Memory access phase: The process that occurs when an instruction needs to access memory (read or write data).
[0058] Write-back phase: The execution result is written back to the register or memory, completing the execution of the instruction.
[0059] The instruction fetch phase of the instruction fetch unit is usually further divided into the following four phases: Program Counter Generation (PCGEN), instruction fetch phase, Instruction Pre-decode (IP) phase, and Instruction Buffer (IB) phase.
[0060] The program generator generation stage includes combinational logic to determine the current memory location of the current jump instruction.
[0061] The instruction fetch phase is responsible for accessing the L1 instruction cache and the predictor record, and for completing the translation from virtual address to physical address.
[0062] The pre-decoding stage is responsible for instruction preprocessing and initiating branch jump requests recorded by the predictor.
[0063] During the caching phase, a branch jump request is initiated for a missed predictor record, as well as a jump request for an indirect branch or function call return. The instructions preprocessed during the pre-decoding phase are encapsulated and cached. Simultaneously, after instruction fusion, up to two instructions are sent to the decoding phase.
[0064] Because the program generator generation stage is logically coupled with other modules, it is not cached by registers and is sometimes not a pipeline stage within the instruction fetch stage.
[0065] Currently, Figure 2 shows a flowchart of an instruction acquisition method, as shown in Figure 2:
[0066] Critical instructions, such as jump instructions, are non-continuous instructions that may cause a sudden change in the direction of the instruction flow.
[0067] During the program counter generation phase: After receiving the combinational logic (instruction retirement unit jump, branch unit jump, pre-decoder predictor jump, cache predictor jump, and program counter increment) for the branch prediction information of the current jump instruction, an address comparison is performed in advance. Based on whether a tightly coupled memory unit is hit, and according to the current program counter value, an access request is initiated to either the cache unit or the tightly coupled memory unit.
[0068] During the instruction fetch phase: the current jump instruction is fetched from the cache unit or the tightly coupled memory unit.
[0069] In the address advance determination scheme shown in Figure 2, the method of comparing addresses in advance and then fetching instructions will cause a large delay in processing timing, reduce the processor's clock frequency, and affect the processor's processing performance.
[0070] To address the aforementioned timing delay issue, Figure 3 illustrates a flowchart of another instruction fetching method, as shown in Figure 3:
[0071] Critical instructions, such as jump instructions, are non-continuous instructions that may cause a sudden change in the direction of the instruction flow.
[0072] During the program counter generation phase: After receiving the combinational logic (instruction retirement unit jump, branch unit jump, pre-decoder predictor jump, cache predictor jump, and program counter increment) for the branch prediction information of the current jump instruction, the current instruction is sequentially executed from the cache unit and tightly coupled memory unit according to the current program counter value. At the same time, the target address is calculated, and the target address is compared with the address range of the cache unit or tightly coupled memory unit to determine the cache unit or tightly coupled memory unit as the valid current memory unit.
[0073] During the instruction fetch phase: the current jump instruction from the valid current memory location is retained.
[0074] In the post-deployment selection scheme shown in Figure 3, although the timing pressure is relatively small, since another branch is executed additionally, even if the tightly coupled memory cell capacity is large, if instructions are fetched simultaneously, it will generate large dynamic power consumption, affecting the processor's battery life and heat dissipation.
[0075] As can be seen, the scheme shown in Figure 2 selects the method of comparing addresses first and then fetching instructions, while the scheme shown in Figure 3 selects the method of comparing addresses while fetching instructions, and then keeping one and discarding the other based on the address comparison results. Neither of these methods can effectively reduce the latency caused by address comparison, avoid additional dynamic power consumption, and reduce the processor's processing performance and product battery life while ensuring the stability of instruction fetching.
[0076] To address the aforementioned problems, this disclosure provides an instruction fetching method. This disclosure also relates to a processor, a system-on-a-chip, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0077] Referring to Figure 4, which shows a flowchart of an instruction fetching method according to an embodiment of the present disclosure, the method is applied to the instruction fetch unit of a processor. The processor also includes a cache unit, a tightly coupled memory unit, and a program counter, and includes the following specific steps:
[0078] Step 402: Receive branch prediction information for the current jump instruction.
[0079] The processor is the core component of a computer system, responsible for executing instructions and processing data, and controlling other parts of the computer system, such as input / output devices, memory, and network interfaces. The processor performs a series of complex operations to complete specific processing tasks, including fetching instructions from storage, decoding instructions, executing instructions, accessing memory, and writing back results. Processors include, but are not limited to, Reduced Instruction Set Processors (RISC) and Complex Instruction Set Processors (CISP).
[0080] The instruction fetch unit is a component unit in the processor responsible for fetching the next instruction to be executed from memory. The instruction fetch unit is used to retrieve instructions from memory during the fetch phase and pass them to subsequent processing phases, such as the decoding phase of the decode unit and the execution phase of the execution unit.
[0081] A cache unit is a storage unit in a processor used to temporarily store data and instructions. The cache unit stores recently accessed data and instructions to reduce the number of times the processor accesses external memory, lower data access latency, and improve processor efficiency. Cache units are typically divided into multiple levels, each with different speeds and capacities. In this embodiment, the cache unit can be an instruction cache unit. For example, in a three-level cache (L1, L2, L3), L1 cache is the fastest but has the smallest capacity, while L3 cache is the slowest but has the largest capacity.
[0082] Tightly coupled memory units are memory units in a processor used to temporarily store critical instructions. Tightly coupled memory units are memory units at the same level as cache units. Compared to cache units, the instructions and data stored in tightly coupled memory units are generally different, such as interrupt handling instructions, real-time operating system kernel instructions, and communication protocol stack instructions. These critical instructions and data are highly sensitive to latency and require rapid response. In this embodiment, the tightly coupled memory unit can be an instruction tightly coupled memory unit.
[0083] The program counter is a register in the processor that stores the address of the next instruction. It instructs the instruction fetch unit to retrieve the next instruction from that address. This instruction address is the program counter value, which is the address of a memory location that points to the location where the next instruction is fetched.
[0084] A jump instruction is a non-continuous instruction used to change the execution order of a program. Jump instructions allow the processor to skip certain instructions or jump to another location in the program to continue execution. Jump instructions are commonly used to implement logic such as conditional branches, loops, subroutine calls, and returns. For example, unconditional jump: `JMP label`, jumps to the label location regardless of the condition. Conditional jump: `JZ label`, jumps to the label location if the zero flag (Z) is 1. Subroutine call: `CALL function`, jumps to the function `function` and saves the return address. Subroutine return: `RET`, returns from the subroutine, restoring to the address of the instruction before the call.
[0085] The current jump instruction is the jump instruction that the processor is currently executing. The processor needs to parse and execute this instruction and update the program counter to point to the new instruction address. The execution of the current jump instruction may affect the program's control flow, resulting in different execution paths.
[0086] Branch prediction information for the current jump instruction predicts whether the current instruction is a jump instruction. This information is the prediction output by the branch prediction unit, including the prediction result (jump or no jump), and may also include the source branch prediction unit. Branch prediction information is generally represented by electrical signals transmitted between components in the processor. The purpose of branch prediction is to reduce pipeline stalls caused by waiting for instruction branch jumps, thereby improving processor execution efficiency.
[0087] Receive branch prediction information for the current jump instruction. One optional method is to receive branch prediction information for the current jump instruction output by the branch prediction unit. The branch prediction unit is a component unit in the processor used to predict the behavior of jump instructions, including but not limited to: instruction retirement unit, branch unit, instruction fetch predictor, and cache predictor.
[0088] For example, a RISC-V processor includes an instruction fetch unit, an instruction cache unit, an instruction tightly coupled memory unit, and a program counter. In a RISC-V processor, n instructions are executed in a pipelined manner. During the fetch stage of the i-th instruction: the instruction fetch unit receives the branch prediction information JumpSignal_i output by the branch prediction unit for the current jump instruction Instruction_i: prediction result: jump occurs; source: branch unit.
[0089] In step 402, branch prediction information for the current jump instruction is received, providing an information basis for subsequently determining the current memory unit where the current jump instruction is located.
[0090] Step 404: Based on the branch prediction information, determine the current memory unit where the current jump instruction is located from the cache unit and the tightly coupled memory unit.
[0091] The current memory location, predicted by branch prediction information, is the memory location where the current jump instruction is stored. It is the fetch object during the fetch phase of the current jump instruction. The current memory location can be either a cache location or a tightly coupled memory location, but not both simultaneously. The current memory location may contain a valid current jump instruction or an invalid current jump instruction.
[0092] Based on branch prediction information, the current storage unit where the current jump instruction is located is determined from the cache unit and tightly coupled storage unit. One possible approach is to determine the current storage unit where the current jump instruction is located from the cache unit and tightly coupled storage unit if the prediction result of the branch prediction information indicates that a jump will occur.
[0093] For example, if the prediction result of the branch prediction information JumpSignal_i is that a jump will occur, the current storage unit Storage Unit_i where the current jump instruction Instruction_i is located is determined to be the instruction tightly coupled storage unit from the instruction cache unit and the instruction tightly coupled storage unit.
[0094] Based on branch prediction information, the current memory location of the current jump instruction is determined from the cache unit and tightly coupled memory unit, providing the basis for fetching the instruction object in the future.
[0095] Step 406: Retrieve the current jump instruction from the current memory location based on the offset of the current program counter value recorded by the program counter.
[0096] The current program counter value is the address of the next instruction (in this embodiment, the current jump instruction) that the program counter updates after the previous instruction has been executed. For example, if the current program counter value is 0x1000, it means that the next instruction to be executed is located at address 0x1000. If the processor executes a jump instruction JMP 0x1010, the program counter will be updated to 0x1010, indicating that the next instruction to be executed is located at address 0x1010.
[0097] The offset of the current program counter is the offset from the address of the instruction preceding the next instruction. This offset is typically fixed, depending on the processor architecture. In most processors, the instruction length is fixed, so the offset is also a fixed value. For example, for a 32-bit processor, each instruction is 4 bytes long, so the offset is 4 bytes. For a 64-bit processor, each instruction may be 8 bytes long or longer.
[0098] Based on the offset of the current program counter value recorded by the program counter, the current jump instruction is retrieved from the current memory unit. One possible approach is to determine the instruction address based on the offset of the current program counter value recorded by the program counter, and retrieve the instruction corresponding to the instruction address from the current memory unit as the current jump instruction.
[0099] For example, based on the offset of the current program counter value 0x1000 recorded by the program counter, the instruction address is determined: 0x1000+4=0x1004. The instruction corresponding to the instruction address (0x1004) is retrieved from the current storage unit Storage Unit_i, which is the instruction tightly coupled storage unit, as the current jump instruction Instruction_i.
[0100] In this embodiment, during the instruction fetch stage, based on branch prediction information for the current jump instruction, one of the cache units and tightly coupled memory units is directly determined as the current memory unit where the current jump instruction is located. Based on the offset of the current program counter value recorded by the program counter, the current jump instruction is fetched from the current memory unit. Address comparison and instruction fetching are no longer performed sequentially, reducing timing pressure and increasing the processor's clock frequency. At the same time, it effectively avoids the additional dynamic power consumption caused by fetching jump instructions from two branches simultaneously, improving the processor's performance and battery life when processing non-continuous instructions. It achieves more efficient and energy-saving instruction fetching without sacrificing processing speed, and is suitable for application scenarios with high requirements for real-time performance and battery life.
[0101] Regarding the issue shown in Figure 3, where the selection scheme after instruction delivery results in significant dynamic power consumption due to simultaneous instruction fetching, affecting the processor's battery life and heat dissipation, targeted optimizations can be made based on scheme 3. This involves directly selecting one of the cache unit or tightly coupled memory unit as the current memory unit, avoiding simultaneous instruction fetching from two memory units and reducing dynamic power consumption.
[0102] Correspondingly, in an optional embodiment of this disclosure, step 404 includes the following specific steps:
[0103] If the prediction result of the branch prediction information is that a jump will occur, the previous memory location where the previous jump instruction was located is determined from the cache unit and the tightly coupled memory unit as the current memory location where the current jump instruction is located.
[0104] The last jump instruction was the most recent jump instruction executed before the current jump instruction.
[0105] The memory location containing the previous jump instruction is the memory location that stores the previous jump instruction. It is the fetch object during the instruction fetch phase of the previous jump instruction. The previous memory location is either a cache location or a tightly coupled memory location, but it cannot be both at the same time. The previous memory location contains a valid previous jump instruction.
[0106] For example, if the prediction result of the branch prediction information JumpSignal_i is that a jump will occur, the previous storage unit where the previous jump instruction Instruction_(i-1) is located is determined from the instruction cache unit and the instruction tightly coupled storage unit: the instruction tightly coupled storage unit is the current storage unit where the current jump instruction Instruction_i is located.
[0107] In this embodiment, a retransmission scheme is adopted to directly determine the previous memory unit where the previous jump instruction is located as the current memory unit where the current jump instruction is located. This avoids fetching instructions from two memory units at the same time, reduces timing pressure, increases the processor's clock frequency, and improves the processor's performance and product battery life when processing non-continuous instructions.
[0108] In an optional embodiment of this disclosure, before determining from the cache unit and tightly coupled storage unit that the previous storage unit where the previous jump instruction was located is the current storage unit where the current jump instruction is located, if the prediction result of the branch prediction information indicates that a jump has occurred, the following specific steps are further included:
[0109] Receive branch prediction information for the initial jump instruction;
[0110] In the case where the branch prediction information records the initial jump instruction and a jump occurs, both the cache unit and the tightly coupled storage unit are determined to be the initial storage unit;
[0111] Based on the offset of the current program counter value recorded by the program counter, the first initial jump instruction is fetched from the cache unit, and based on the offset of the current program counter value recorded by the program counter, the second initial jump instruction is fetched from the tightly coupled memory unit.
[0112] By comparing the current program counter value recorded by the program counter with the address range of the cache unit and the tightly coupled memory unit, the valid initial memory unit is determined from the cache unit and the tightly coupled memory unit;
[0113] The first or second initial jump instruction from the valid initial memory location is retained as the initial jump instruction.
[0114] The initial jump instruction is the jump instruction that the processor executes for the first time.
[0115] Branch prediction information for an initial jump instruction is information predicting whether the initial instruction is a jump instruction. Branch prediction information is the prediction output by the branch prediction unit for the initial jump instruction, including the prediction result (jump or no jump), and may also include the source branch prediction unit. Branch prediction information is generally represented as electrical signals transmitted between component units in the processor.
[0116] The initial memory location is the memory location that may store the initial jump instruction. The initial memory location consists of both a cache location and a tightly coupled memory location. When executing a jump instruction for the first time, since there is no historical data (previous memory location) to refer to, the processor needs to fetch the instruction from both the cache location and the tightly coupled memory location simultaneously to ensure integrity.
[0117] The first initial jump instruction is the initial jump instruction fetched from the cache unit. The first initial jump instruction may be a valid initial jump instruction or an invalid initial jump instruction, which needs to be verified by subsequent address comparison.
[0118] The second initial jump instruction is the initial jump instruction fetched from the tightly coupled memory cell. This second initial jump instruction may be valid or invalid, and needs to be verified during subsequent address comparison.
[0119] The address range of a cache unit is the address interval of the data and instructions stored in the cache unit. The address range of a cache unit is usually fixed and determined by the hardware design. The processor determines whether the first initial jump instruction fetched from the cache unit is a valid initial jump instruction by comparing the current program counter value recorded by the program counter with the address range of the cache unit. The address range of a cache unit is from 0x1000 to 0x1FFF.
[0120] The address range of a tightly coupled memory location is the address interval of the data and instructions stored within that location. This address range is typically fixed and determined by the hardware design. The processor determines whether the second initial jump instruction fetched from the tightly coupled memory location is a valid initial jump instruction by comparing the current program counter value with the address range of the tightly coupled memory location. For example, the address range of a tightly coupled memory location might be from 0x2000 to 0x2FFF.
[0121] The valid initial memory location is the memory location that stores the initial jump instruction. When a jump instruction is executed for the first time, the processor needs to fetch the instruction from the cache and the tightly coupled memory location, and then compare the current program counter value recorded by the program counter with the address range of these two memory locations to determine which memory location stores the valid initial jump instruction. The initial jump instruction in the valid initial memory location will be retained and executed later.
[0122] Based on the offset of the current program counter value recorded by the program counter, the first initial jump instruction is retrieved from the cache unit. One possible approach is to determine the instruction address based on the offset of the current program counter value recorded by the program counter, and retrieve the instruction corresponding to the instruction address from the cache unit as the first initial jump instruction.
[0123] Based on the offset of the current program counter value recorded by the program counter, the second initial jump instruction is retrieved from the tightly coupled memory unit. One possible approach is to determine the instruction address based on the offset of the current program counter value recorded by the program counter, and retrieve the instruction corresponding to the instruction address from the tightly coupled memory unit as the second initial jump instruction.
[0124] For example, in the instruction fetch stage of the first instruction: the fetch unit receives the branch prediction information JumpSignal_1 output by the branch prediction unit for the initial jump instruction Instruction_1: prediction result: jump occurred; source: branch unit. If the branch prediction information JumpSignal_1 records that the initial jump instruction has occurred, it is determined that both the instruction cache unit and the instruction tightly coupled storage unit are the initial storage unit Storage Unit_1. Based on the current program counter value of 0x0000 and the offset 4 recorded by the program counter, the instruction address is determined to be: 0x0000 + 4 = 0x0004. The instruction corresponding to the instruction address (0x0004) is fetched from the instruction cache unit as the first initial jump instruction Instruction_1.1, and the instruction corresponding to the instruction address (0x0004) is fetched from the instruction tightly coupled storage unit as the second initial jump instruction Instruction_1.2. The address range of the instruction cache unit is 0x0000 to 0x0FFF. The address range of the instruction-closed memory location is 0x1000 to 0x1FFF. Comparing the target address 0x0004 with the address range of the instruction cache (0x0000 to 0x0FFF), the target address is found to be within the instruction cache address range. Comparing the target address 0x0004 with the address range of the instruction-closed memory location (0x1000 to 0x1FFF), the target address is found to be outside the instruction-closed memory location range. The instruction cache is determined to be the valid initial memory location. The first initial jump instruction Instruction_1.1 from the valid initial memory location is retained as the initial jump instruction Instruction_1.
[0125] In this embodiment of the disclosure, when a jump instruction is executed for the first time, the instruction is fetched from the cache unit and the tightly coupled memory unit respectively, and the current program counter value recorded by the program counter is compared with the address range of the two memory units to determine the valid initial memory unit. This ensures that a valid initial jump instruction is fetched from the correct memory unit, avoids the additional dynamic power consumption caused by fetching instructions from two memory units at the same time, reduces timing pressure, improves the processor's clock frequency and energy efficiency, and improves the processor's performance and product battery life when processing non-continuous instructions.
[0126] Although targeted optimizations can be made to Scheme 3 by directly selecting one of the cache unit or tightly coupled memory unit as the current memory unit to avoid fetching instructions from two memory units at the same time and reduce dynamic power consumption, the instantaneous power consumption of the processor during the jump will still increase. If the source type of the branch prediction information is relatively stable, a speculative approach can be used to complete the instruction fetching.
[0127] Correspondingly, in an optional embodiment of this disclosure, step 404 includes the following specific steps:
[0128] If the prediction result of the branch prediction information is that a jump will occur, the current storage unit where the current jump instruction is located is determined from the cache unit and the tightly coupled storage unit based on the source type of the branch prediction information.
[0129] The source type of branch prediction information refers to the output source of the branch prediction information, that is, which component unit generates the information predicting whether the current instruction is a jump instruction. Different source types reflect different prediction mechanisms and prediction accuracy, and have different impacts on processor performance and power consumption. The source types of branch prediction information include, but are not limited to: instruction retirement unit, branch unit, predictor (fetch predictor and cache predictor), and program count increment.
[0130] It should be noted that, based on the common practice of using tightly coupled storage units, critical instructions, such as exception vector tables and exception handling functions, are typically stored in these units in common use cases. Jumps into such scenarios are usually triggered by instruction retirement units. Using this method can reduce the time spent entering exceptions and accelerate the processing flow. Normally executing code, on the other hand, generally enters from branch units, predictors, etc., and has lower latency requirements; therefore, the default cache unit is the current storage unit.
[0131] For example, if the prediction result of the branch prediction information JumpSignal_i is that a jump will occur, based on the source type of the branch prediction information JumpSignal_i: instruction retirement unit, the current storage unit Storage Unit_i where the current jump instruction Instruction_i is located is determined to be an instruction tightly coupled storage unit from the instruction cache unit and the instruction tightly coupled storage unit.
[0132] In this embodiment of the disclosure, when the prediction result of the branch prediction information is that a jump will occur, the current memory unit where the current jump instruction is located is directly determined from the cache unit and the tightly coupled memory unit based on the source type of the branch prediction information. This can effectively reduce the instantaneous power consumption of the processor, improve prediction accuracy and processing performance, and effectively avoid the additional dynamic power consumption caused by the simultaneous acquisition of jump instructions by two branches while ensuring performance. This improves the processor's performance and product battery life when processing non-continuous instructions.
[0133] In one optional embodiment of this disclosure, when the prediction result of the branch prediction information indicates that a jump has occurred, the current storage unit where the current jump instruction is located is determined from the cache unit and the tightly coupled storage unit based on the source type of the branch prediction information, including the following specific steps:
[0134] If the prediction result of the branch prediction information is that a jump will occur, and the source type of the branch prediction information is a branch unit, then the cache unit is determined to be the current storage unit where the current jump instruction is located.
[0135] The branch unit is a component unit in the processor used to predict jump instructions. It is responsible for analyzing branch conditions and predicting the direction of the branch to reduce pipeline stalls caused by waiting for jump instruction results, thereby improving processor execution efficiency. The main functions of the branch unit include: Branch condition analysis: Analyzing the condition flags of the jump instruction (such as the zero flag Z, negative flag N, overflow flag V, etc.) to determine whether the branch condition is met. Branch prediction: Predicting whether the instruction will jump based on historical data and current conditions. Common prediction algorithms include static prediction and dynamic prediction. Branch instruction address calculation: Calculating the instruction address of the jump instruction so that the instruction fetch unit can fetch the next instruction from the correct address. Branch history record: Maintaining the Branch History Table (BHT) to record the behavior of past jump instructions for dynamic prediction.
[0136] For example, if the prediction result of the branch prediction information JumpSignal_i is that a jump will occur, based on the source type of the branch prediction information JumpSignal_i: branch unit, the current storage unit Storage Unit_i where the current jump instruction Instruction_i is located is determined to be the instruction cache unit from the instruction cache unit and the instruction tightly coupled storage unit.
[0137] In this embodiment of the disclosure, when the prediction result of the branch prediction information is a jump and the source type of the branch prediction information is a branch unit, the cache unit is directly determined as the current storage unit where the current jump instruction is located. This has low latency requirements, can effectively reduce the instantaneous power consumption of the processor, and improve prediction accuracy and processing performance.
[0138] In one optional embodiment of this disclosure, when the prediction result of the branch prediction information indicates that a jump has occurred, the current storage unit where the current jump instruction is located is determined from the cache unit and the tightly coupled storage unit based on the source type of the branch prediction information, including the following specific steps:
[0139] If the prediction result of the branch prediction information is that a jump will occur, and the source type of the branch prediction information is an instruction retirement unit, then the tightly coupled memory unit is determined to be the current memory unit where the current jump instruction is located.
[0140] The Instruction Retirement Unit (IRU) is a component unit within the processor responsible for verifying instruction execution results and updating the processor's state. The IRU ensures that the execution results of instructions are correctly submitted to the processor's state, including updating register and memory states. After verifying that the instruction execution is error-free, the IRU generates branch prediction information, particularly for jump instructions, indicating whether a jump occurred. The main functions of the IRU include: Verifying instruction execution results: ensuring that instructions have been executed correctly without any exceptions or errors. Updating state: updating register and memory states to reflect the instruction execution results. Generating branch prediction information: for jump instructions, the IRU generates branch prediction information to indicate whether a jump occurred. Exception handling: handling exceptions that may occur during instruction execution.
[0141] For example, if the prediction result of the branch prediction information JumpSignal_i is that a jump will occur, based on the source type of the branch prediction information JumpSignal_i: retired instruction unit, the current storage unit Storage Unit_i where the current jump instruction Instruction_i is located is determined to be an instruction tightly coupled storage unit from the instruction cache unit and the instruction tightly coupled storage unit.
[0142] In this embodiment of the disclosure, when the prediction result of the branch prediction information is that a jump has occurred and the source type of the branch prediction information is an instruction retirement unit, directly determining the tightly coupled storage unit as the current storage unit where the current jump instruction is located can reduce the time to enter the exception, accelerate the processing flow, effectively reduce the instantaneous power consumption of the processor, and improve prediction accuracy and processing performance.
[0143] In one optional embodiment of this disclosure, when the prediction result of the branch prediction information indicates that a jump has occurred, the current storage unit where the current jump instruction is located is determined from the cache unit and the tightly coupled storage unit based on the source type of the branch prediction information, including the following specific steps:
[0144] If the prediction result of the branch prediction information is that a jump will occur, and the source type of the branch prediction information is a predictor, the previous memory unit where the previous jump instruction was located is determined from the cache unit and the tightly coupled memory unit as the current memory unit where the current jump instruction is located.
[0145] The predictor is a component in the processor used to predict the address of the next instruction. By analyzing historical data and current conditions, the predictor predicts whether a jump instruction will execute and its target address, reducing pipeline stalls caused by waiting for jump instruction results and thus improving processor efficiency. The main functions of the predictor include: Static prediction: Predicting based on fixed rules, such as always predicting whether a jump instruction will execute or not. Dynamic prediction: Predicting based on historical data. Common dynamic prediction algorithms include: Local prediction: Considering only the historical behavior of the current branch. Global prediction: Considering the historical behavior of all branches. Hybrid prediction: Combining local and global prediction. Branch Target Buffer (BTB): Used to store the target address of known jump instructions, speeding up the calculation of branch target addresses. Branch History Table: Used to record the behavior of past jump instructions, supporting dynamic prediction.
[0146] In one alternative embodiment of this disclosure, the predictor is at least one of a pre-decoding predictor and a cache predictor.
[0147] The instruction pre-decode predictor is a component in the processor used to predict the address of the next instruction during the instruction pre-decode phase. It analyzes historical data and current conditions to predict whether a jump instruction will execute and its target address, thereby reducing pipeline stalls caused by waiting for jump instruction results and improving processor execution efficiency.
[0148] The Instruction Cache Predictor is a component in the processor used to predict cache access patterns. It analyzes historical data and current conditions to predict whether the next instruction or data will be hit in the cache, thereby optimizing cache usage efficiency, reducing cache misses, and improving processor execution efficiency.
[0149] For example, if the prediction result of the branch prediction information JumpSignal_i is that a jump will occur, based on the source type of the branch prediction information JumpSignal_i: pre-decoder predictor, the previous storage unit where the previous jump instruction Instruction_(i-1) was located is determined from the cache unit and the tightly coupled storage unit as the current storage unit where the current jump instruction Instruction_i is located.
[0150] In this embodiment of the disclosure, when the prediction result of the branch prediction information is a jump and the source type of the branch prediction information is a predictor, the previous memory unit where the previous jump instruction was located is directly determined from the cache unit and the tightly coupled memory unit as the current memory unit where the current jump instruction is located. This has low latency requirements, can effectively reduce the instantaneous power consumption of the processor, and improve prediction accuracy and processing performance.
[0151] Since no address comparison is performed in steps 402 to 406, an incorrect current memory location may be identified, resulting in the fetching of an invalid instruction. Therefore, checks and corrections can be performed during the instruction fetching stage. If an error occurs, the offset of the current program counter value is resent, the current memory location is updated, and the current jump instruction is fetched again.
[0152] In one optional embodiment of this disclosure, after step 406, the following specific steps are further included:
[0153] Compare the current program counter value recorded by the program counter with the address range of the current memory unit to determine whether it falls within the address range;
[0154] If not, update another storage unit from the cache unit and the tightly coupled storage unit to the current storage unit where the current jump instruction is located;
[0155] Retrieve the current jump instruction from the updated current memory location based on the offset of the current program counter value.
[0156] The updated current memory location is another memory location updated after address comparison during the instruction fetch phase. This update process ensures the correctness and stability of instruction fetching and avoids instruction fetching failures caused by incorrect prediction.
[0157] For example, the instruction fetch unit determines that the current storage unit is an instruction-tightly coupled storage unit. It compares the current program counter value recorded by the program counter with the address range of the current storage unit: the current program counter value is 0x2000, while the address range of the instruction-tightly coupled storage unit is 0x1000 to 0x1FFF. The comparison result shows that the current program counter value 0x2000 is not within the address range of the instruction-tightly coupled storage unit, indicating a previous prediction error. Therefore, the instruction cache unit is updated to the current storage unit, Storage Unit_i. Based on the offset of the current program counter value 0x2000, the current jump instruction, Instruction_i, is fetched again from the instruction cache unit.
[0158] In this embodiment of the disclosure, after address comparison is performed during the instruction fetch stage, another storage unit is updated to the current storage unit where the current jump instruction is located. This ensures the correctness and stability of instruction fetching, avoids instruction fetching failure due to incorrect prediction, and ensures high efficiency and stability in application scenarios with high real-time and energy efficiency requirements.
[0159] Referring to one of the above embodiments, Figure 5 shows one of the flowcharts of an instruction acquisition method provided in an embodiment of the present disclosure, as shown in Figure 5:
[0160] Critical instructions, such as jump instructions, are non-continuous instructions that may cause a sudden change in the direction of the instruction flow.
[0161] During the program counter generation phase: After receiving the combinational logic (instruction retirement unit jump, branch unit jump, pre-decoder predictor jump, cache predictor jump, and program counter increment) for the branch prediction information for the initial jump instruction, the current instruction is sequentially executed from the cache unit and tightly coupled memory unit according to the current program counter value. Address comparison is performed, and based on whether the tightly coupled memory unit is hit, an access request is initiated to the cache unit or the tightly coupled memory unit according to the current program counter value.
[0162] During the instruction fetch phase: the initial jump instruction is fetched from the cache unit or tightly coupled memory unit.
[0163] Following this, in the program counter generation phase: after receiving the combinational logic of branch prediction information for the current jump instruction (instruction retirement unit jump, branch unit jump, pre-decoder predictor jump, cache predictor jump, and program counter increment), if the prediction result of the branch prediction information indicates that a jump has occurred, the previous memory location of the last jump instruction is determined from the cache unit and tightly coupled memory unit as the current memory location of the current jump instruction. It is determined that a tightly coupled memory unit has been hit. Based on the current program counter value, an access request is initiated to either the cache unit or the tightly coupled memory unit.
[0164] During the instruction fetch phase: the current jump instruction is fetched from the cache unit or tightly coupled memory unit, and the address is compared to determine whether an error has occurred. If an error occurs, the result is fed back.
[0165] In the retransmission scheme shown in Figure 5, address comparison is bypassed, which reduces the timing pressure of the program counter generation stage and can lead to an increase in processor frequency.
[0166] Referring to the above embodiment, Figure 6 shows a second schematic flowchart of an instruction acquisition method provided in an embodiment of this disclosure, as shown in Figure 6:
[0167] Critical instructions, such as jump instructions, are non-continuous instructions that may cause a sudden change in the direction of the instruction flow.
[0168] During the program counter generation phase: After receiving the combinational logic of branch prediction information for the current jump instruction (instruction retirement unit jump, branch unit jump, pre-decoder predictor jump, cache predictor jump, and program counter increment), if the prediction result of the branch prediction information indicates that a jump has occurred, based on the source type of the branch prediction information (instruction retirement unit, branch unit, pre-decoder predictor, cache predictor), the current memory unit where the current jump instruction is located is determined from the cache unit and tightly coupled memory units. It is then determined whether a tightly coupled memory unit has been hit, and an access request is initiated to either the cache unit or the tightly coupled memory unit.
[0169] During the instruction fetch phase: the current jump instruction is fetched from the cache unit or tightly coupled memory unit, and the address is compared to determine whether an error has occurred. If an error has occurred, the instruction is resent.
[0170] In the speculation scheme shown in Figure 6, the selection of the record / inference selection branch path significantly reduces dynamic power consumption and improves product battery life without performance degradation or only slight degradation. This is the preferred scheme when the speculation success rate is high.
[0171] Figure 7 shows a comparative schematic diagram of an instruction acquisition method provided by an embodiment of this disclosure, relating to the four schemes in Figures 2, 3, 5, and 6:
[0172] The advance address determination scheme shown in Figure 2, during the program counter generation stage:
[0173] Select jump target; address comparison; select current storage unit; send access request.
[0174] During the instruction fetching phase:
[0175] Get the current jump instruction.
[0176] Address comparison is a source of timing pressure. This part of the logic is more complex than other parts, and too many levels of superposition with other logic can affect the processor's clock speed.
[0177] Figure 3 shows the post-issuance selection scheme during the program counter generation stage:
[0178] Select the jump target; compare addresses and issue a request as the current storage unit; send the access request.
[0179] During the instruction fetching phase:
[0180] Get the current jump instruction; select the source based on the address comparison result.
[0181] Address comparison is a source of timing pressure. This part of the logic is more complex than other parts, and excessive stacking with other logic levels can affect the processor's clock speed. Simultaneously, the request to issue to the current memory unit increases power consumption. Accessing both tightly coupled memory units and cache units, but only using instruction data from one of them while discarding the other, leads to wasted power.
[0182] The retransmission scheme shown in Figure 5, during the program counter generation stage:
[0183] Select the jump target; compare addresses and issue a request for the current storage unit; after the first request, issue access requests based on the comparison results.
[0184] During the instruction fetching phase:
[0185] Get the current jump instruction; record the comparison result; select the source based on the address comparison result.
[0186] Address comparison is a source of timing pressure; this part of the logic is more complex than others, and excessive stacking with other logic can affect the processor's clock speed. Simultaneously, it increases power consumption as it handles requests from the current memory unit. Accessing both tightly coupled memory units and cache units but using only one's instruction data while discarding the other leads to wasted power. Recording comparison results relies on the principle of locality; the instruction source after a jump will not change, but even if it does, no error will occur, and the instruction will be resent to select an alternative branch path.
[0187] The speculative scheme shown in Figure 6, during the program counter generation stage:
[0188] Select the jump target; compare addresses and determine the current storage unit based on the source; after the first time, issue access requests based on the comparison results.
[0189] During the instruction fetching phase:
[0190] Get the current jump instruction; record the comparison result; select the source based on the address comparison result.
[0191] Determining the current storage unit based on its source and selecting either a tightly coupled storage unit or a cache unit saves power. However, sending requests to the current storage unit increases power consumption. Accessing both tightly coupled and cache units simultaneously, but using only one type of instruction data while discarding the other, leads to wasted power. Recording comparison results relies on the principle of locality; the instruction source doesn't change after a jump, but even if it does, no error will occur, and the instruction will be resent to select an alternative branch path.
[0192] In summary, the solutions in Figures 5 and 6, combining retransmission and speculative strategies, ensure that performance is not reduced or only slightly reduced while significantly reducing the processor's dynamic power consumption, enabling it to achieve higher performance, longer battery life, and lower heat dissipation requirements in embedded devices and edge computing devices.
[0193] Corresponding to the above method embodiments, this disclosure also provides processor embodiments. FIG8 shows a schematic diagram of the structure of a processor provided in one embodiment of this disclosure. As shown in FIG8, the processor 800 includes an instruction fetch unit 810, a cache unit 820, a tightly coupled memory unit 830, and a program counter 840;
[0194] The instruction fetch unit 810 is used to receive branch prediction information for the current jump instruction, determine the current memory unit where the current jump instruction is located from the cache unit 820 and the tightly coupled memory unit 830 based on the branch prediction information, and fetch the current jump instruction from the current memory unit based on the offset of the current program counter value recorded by the program counter 840.
[0195] In this embodiment of the disclosure, in the processor, the instruction fetch unit, during the instruction fetch stage, directly determines one of the cache units and tightly coupled memory units as the current memory unit where the current jump instruction is located, based on the branch prediction information for the current jump instruction. Based on the offset of the current program counter value recorded by the program counter, the current jump instruction is fetched from the current memory unit. This eliminates the need for sequential address comparison and instruction fetching, reducing timing pressure and increasing the processor's clock frequency. At the same time, it effectively avoids the additional dynamic power consumption caused by fetching jump instructions from two branches simultaneously, improving the processor's performance and battery life when processing non-continuous instructions. It achieves more efficient and energy-saving instruction fetching without sacrificing processing speed, making it suitable for application scenarios with high requirements for real-time performance and battery life.
[0196] The above is an illustrative embodiment of a processor. It should be noted that the technical solution of this processor and the technical solution of the instruction fetching method described above belong to the same concept. Details not described in detail in the processor's technical solution can be found in the description of the technical solution of the instruction fetching method described above.
[0197] Corresponding to the above-described embodiments of the instruction acquisition method, this disclosure also provides embodiments of a system-on-a-chip. Figure 9 shows a schematic diagram of the structure of a system-on-a-chip provided in one embodiment of this disclosure. The system-on-a-chip 900 includes, but is not limited to, a control unit 910 and multiple on-chip components 920, the multiple on-chip components 920 including a processor 9210.
[0198] The control unit 910 is used to control and manage multiple on-chip components 920.
[0199] The processor 9210 is used to execute computer programs / instructions, which, when executed by the processor 9210, implement the steps of the instruction acquisition method described above.
[0200] The System-on-Chip (SoC) 900 is a system-on-a-chip hardware device that integrates a control unit 910 and multiple on-chip components 920. It typically integrates the control unit 910, as well as a central processing unit, a digital signal processor, a processor, and input / output interfaces (I / O), among other on-chip components 920. The SoC 900 can improve performance, reduce power consumption, and decrease physical size, making it suitable for various scenarios such as mobile devices, IoT devices, embedded systems, and edge computing devices. The specific embodiments disclosed herein do not limit the specific implementation of the SoC.
[0201] The control unit 910 is a core component of the system-on-a-chip 900, responsible for instruction decoding, execution flow control, and coordination with multiple on-chip components 920. The control unit 910 controls and manages the workflow of multiple on-chip components 920 by sending control signals, ensuring that each on-chip component can correctly complete its tasks in a predetermined order and logic.
[0202] Multiple on-chip components 920 are various functional hardware components in the system-on-a-chip 900, which work together to complete specific tasks or functions. These multiple on-chip components 920 may include, but are not limited to, a central processing unit, a digital signal processor, a processor, and input / output interfaces.
[0203] The above is an illustrative scheme of a system-on-a-chip (SoC) according to this embodiment. It should be noted that the technical solution of this SoC and the technical solution of the instruction fetching method described above belong to the same concept. Details not described in detail in the SoC technical solution can be found in the description of the instruction fetching method described above.
[0204] Corresponding to the above-described instruction fetching method and system-on-a-chip embodiments, this disclosure also provides embodiments of a computing device. Figure 10 shows a schematic diagram of the structure of a computing device provided in one embodiment of this disclosure. The computing device 1000 includes, but is not limited to, a memory 1010 and a system-on-a-chip 1020.
[0205] The memory 1010 is used to store computer programs / instructions.
[0206] The system-on-chip 1020 is used to execute the following computer program / instruction, which, when executed by the system-on-chip 1020, implements the steps of the above instruction acquisition method.
[0207] The computing device 1000 is an electronic device with computing capabilities, capable of running various software programs to perform specific functions or tasks. The computing device 1000 typically includes a memory 1010 and hardware devices for performing tasks. In embodiments of this disclosure, the computing device 1000 includes, but is not limited to, a system-on-a-chip 1020 and a memory 1010. Specific embodiments of this disclosure do not limit the specific implementation of the computing device.
[0208] The memory 1010 is a component in the computing device 1000 used to store data and computer programs / instructions, and is generally divided into two types: volatile and non-volatile. Volatile memory (such as RAM) loses data after power is turned off and is mainly used for temporary storage of running programs and data; non-volatile memory (such as ROM and flash memory) can retain data after power is turned off and is used to store firmware and persistent data.
[0209] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the instruction acquisition method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the instruction acquisition method described above.
[0210] This disclosure also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, perform the methods as described in any of the foregoing embodiments.
[0211] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the methods in any of the foregoing embodiments.
[0212] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0213] The computer program / instruction includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form.
[0214] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this disclosure are not limited to the described order of actions, because according to the embodiments of this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this disclosure.
[0215] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0216] The preferred embodiments disclosed above are merely illustrative of this disclosure. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments of this disclosure. These embodiments are selected and specifically described in this disclosure to better explain the principles and practical applications of the embodiments of this disclosure, thereby enabling those skilled in the art to better understand and utilize this disclosure. This disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. An instruction fetching method applied to the instruction fetch unit of a processor, the processor further comprising a cache unit, a tightly coupled memory unit, and a program counter, comprising: Receive branch prediction information for the current jump instruction; Based on the branch prediction information, the current storage unit where the current jump instruction is located is determined from the cache unit and the tightly coupled storage unit; The current jump instruction is retrieved from the current storage unit based on the offset of the current program counter value recorded by the program counter.
2. The method according to claim 1, wherein determining the current storage unit where the current jump instruction is located from the cache unit and the tightly coupled storage unit based on the branch prediction information includes: If the prediction result of the branch prediction information is that a jump will occur, the previous storage unit where the previous jump instruction was located is determined from the cache unit and the tightly coupled storage unit as the current storage unit where the current jump instruction is located.
3. The method according to claim 2, further comprising, before determining from the cache unit and the tightly coupled storage unit that the previous storage unit where the previous jump instruction was located is the current storage unit where the current jump instruction is located, when the prediction result of the branch prediction information is that a jump has occurred, the method includes: Receive branch prediction information for the initial jump instruction; If the branch prediction information records that the initial jump instruction has occurred, it is determined that both the cache unit and the tightly coupled storage unit are initial storage units; Based on the offset of the current program count value recorded by the program counter, the first initial jump instruction is retrieved from the cache unit, and based on the offset of the current program count value recorded by the program counter, the second initial jump instruction is retrieved from the tightly coupled storage unit. By comparing the current program counter value recorded by the program counter with the address range of the cache unit and the tightly coupled memory unit, a valid initial memory unit is determined from the cache unit and the tightly coupled memory unit; The first initial jump instruction or the second initial jump instruction from the valid initial storage unit is retained as the initial jump instruction.
4. The method according to claim 1, wherein determining the current storage unit where the current jump instruction is located from the cache unit and the tightly coupled storage unit based on the branch prediction information includes: If the prediction result of the branch prediction information is that a jump will occur, the current storage unit where the current jump instruction is located is determined from the cache unit and the tightly coupled storage unit based on the source type of the branch prediction information.
5. The method according to claim 4, wherein, when the prediction result of the branch prediction information is a jump, determining the current storage unit where the current jump instruction is located from the cache unit and the tightly coupled storage unit based on the source type of the branch prediction information includes: If the prediction result of the branch prediction information is that a jump will occur, and the source type of the branch prediction information is a branch unit, then the cache unit is determined to be the current storage unit where the current jump instruction is located.
6. The method according to claim 4, wherein, when the prediction result of the branch prediction information is a jump, determining the current storage unit where the current jump instruction is located from the cache unit and the tightly coupled storage unit based on the source type of the branch prediction information includes: If the prediction result of the branch prediction information is that a jump will occur, and the source type of the branch prediction information is an instruction retirement unit, then the tightly coupled storage unit is determined to be the current storage unit where the current jump instruction is located.
7. The method according to claim 4, wherein, when the prediction result of the branch prediction information is a jump, determining the current storage unit where the current jump instruction is located from the cache unit and the tightly coupled storage unit based on the source type of the branch prediction information includes: If the prediction result of the branch prediction information is that a jump has occurred, and the source type of the branch prediction information is a predictor, then the previous storage unit where the previous jump instruction was located is determined from the cache unit and the tightly coupled storage unit as the current storage unit where the current jump instruction is located.
8. The method according to claim 7, wherein the predictor is at least one of a pre-decoding predictor and a cache predictor.
9. The method according to any one of claims 1-8, after retrieving the current jump instruction from the current storage unit based on the offset of the current program counter value recorded by the program counter, further comprising: Compare the current program counter value recorded by the program counter with the address range of the current storage unit to determine whether it falls within the address range; If not, update another storage unit from the cache unit and the tightly coupled storage unit to the current storage unit where the current jump instruction is located; Based on the offset of the current program counter value, the current jump instruction is retrieved from the updated current storage unit.
10. A processor, the processor comprising an instruction fetch unit, a cache unit, a tightly coupled memory unit, and a program counter; The instruction fetch unit is configured to receive branch prediction information for the current jump instruction, determine the current storage unit where the current jump instruction is located from the cache unit and the tightly coupled storage unit based on the branch prediction information, and fetch the current jump instruction from the current storage unit based on the offset of the current program counter value recorded by the program counter.
11. A system-on-a-chip, comprising: A control unit and multiple on-chip components, the multiple on-chip components including a processor; The control unit is used to control and manage the plurality of on-chip components, and the processor is used to execute a computer program / instruction that, when executed by the processor, implements the steps of the method according to claims 1 to 9.
12. A computing device, comprising: Memory and system-on-a-chip; The memory is used to store computer programs / instructions, and the system-on-a-chip is used to execute the computer programs / instructions, which, when executed by the system-on-a-chip, implement the steps of the method described in claims 1 to 9.
13. A computer readable storage medium, wherein, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the steps of the method as described in claims 1 to 9.
14. A computer program product, wherein, It includes a computer program that, when executed by a processor, implements the steps of the method as described in claims 1 to 9.