Instruction Tightly-Coupled Memory and Instruction Cache Access Prediction

By using access prediction logic parallel access instructions in the processor, the memory structure that needs to be enabled is predicted based on the location status and program counter value, the power consumption and delay problems during instruction acquisition are solved and system performance is improved.

CN113227970BActive Publication Date: 2025-07-08SIFIVE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980085798.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-28
Filing Date
2019-12-12
Publication Date
2025-07-08
Estimated Expiration
2039-12-12

AI Technical Summary

Technical Problem

When obtaining instructions, the processor needs to access the instruction cache and the instruction is tightly coupled to the memory, resulting in increased power consumption and increased latency, affecting system performance.

Method used

Access prediction logic or predictor is used to access instructions in parallel tightly couple memory and instruction cache, predicting the memory structure that needs to be enabled based on location status and program counter value, reducing unnecessary memory access.

Benefits of technology

By reducing unnecessary memory access, power consumption is reduced and system performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113227970B_ABST
    Figure CN113227970B_ABST
Patent Text Reader

Abstract

Disclosed herein are systems and methods for instruction tightly coupled memory (iTIM) and instruction cache (iCache) access prediction. A processor may use a predictor to enable access to the iTIM or iCache and a specific path (memory structure) based on a location state and a program counter value. The predictor may determine whether to stay in an enabled memory structure, move to and enable a different memory structure, or move to and enable both memory structures. Stay and move predictions may be based on whether a memory structure boundary crossing occurs due to sequential instruction processing, branch or jump instruction processing, branch resolution, and cache miss handling. The program counter and location state indicators may use feedback and be updated in each instruction fetch cycle to determine which memory structure(s) need to be enabled for the next instruction fetch.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Patent Application No. 16 / 553,839, filed on August 28, 2019, which claims priority to U.S. Provisional Patent Application No. 62 / 785,947, filed on December 28, 2018. The entire disclosure of the above applications is hereby incorporated by reference herein. Technical Field

[0003] This disclosure relates to access prediction between an instruction tightly - coupled memory and an instruction cache to fetch instructions. Background Art

[0004] The instruction fetch time between a processor and an off - chip memory system or main memory is typically much slower than the processor execution time. Thus, processors employ instruction caches and instruction tightly - coupled memories to improve system performance. Both types of memories improve latency and reduce power consumption by reducing off - chip memory accesses. However, by having to search both the instruction cache and the instruction tightly - coupled memory for each instruction fetch, the processor uses a large amount of power. Additionally, this can increase latency and reduce system performance. Brief Description of the Drawings

[0005] The present disclosure is best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be emphasized that, as is customary, the various features of the drawings are not drawn to scale. Instead, for clarity, the dimensions of the various features are arbitrarily enlarged or reduced.

[0006] Figure 1 is a block diagram of an example of a processor including access prediction logic for enabling access to one of an instruction cache (iCache) or an instruction tightly - coupled memory (iTIM) in accordance with an embodiment of the present disclosure.

[0007] Figure 2 is a diagram of an example core pipeline of a processor in accordance with an embodiment of the present disclosure.

[0008] Figure 3 is a diagram of an example process and predictor in accordance with an embodiment of the present disclosure.

[0009] Figure 4 is for an example technique for access prediction between an iTIM and an iCache as shown in Figure 3 in accordance with an embodiment of the present disclosure.

[0010] Figure 5 is a diagram of an example process and predictor in accordance with an embodiment of the present disclosure.

[0011] Figure 6A diagram of an example process and branch predictor in accordance with an embodiment of the present disclosure.

[0012] Figure 7 is a diagram of an example technique for access prediction between an iTIM and an iCache as Figure 5 shown. DETAILED DESCRIPTION

[0013] Systems and methods for instruction tightly coupled memory (iTIM) and instruction cache (iCache) access prediction are disclosed herein. The embodiments described herein can be used to eliminate or mitigate the need to access both the iTIM and the iCache when fetching instructions.

[0014] A processor can employ an iTIM and an N-way set associative iCache when fetching instructions to improve system performance. To minimize instruction fetch latency, N paths of the iTIM and the iCache can be accessed in parallel before it is known whether the iTIM or the N-way set associative iCache contains the desired instruction. Power consumption can be reduced by employing position state feedback during instruction fetch. Instead of accessing both N paths of the iTIM and the iCache, the use of feedback can enable the processor to achieve higher performance and / or lower power consumption by accessing one of the iTIM or the iCache (and a specific path).

[0015] The processor can use access prediction logic or a predictor to send an enable signal to access the iTIM or the iCache and a specific path based on a position state and a program counter value. The access prediction logic can determine whether to stay in an enabled memory structure, move to and enable a different memory structure, or move to and enable both memory structures. The stay and move predictions can be based on whether a memory structure boundary crossing occurs due to sequential instruction processing, branch or jump instruction processing, branch resolution, and cache miss handling, where the memory structure is one of the iTIM or the iCache and a specific path. The program counter and the position state indicator can use feedback and be updated during each instruction fetch cycle to determine which (which ones) of the memory structures need to be enabled for the next instruction fetch.

[0016] These and other aspects of the present disclosure are disclosed in the following detailed description, the appended claims, and the drawings.

[0017] As used herein, the term "processor" refers to one or more processors, such as one or more dedicated processors, one or more digital signal processors, one or more microprocessors, one or more controllers, one or more microcontrollers, one or more application processors, one or more central processing units (CPUs), one or more graphics processing units (GPUs), one or more digital signal processors (DSPs), one or more application specific integrated circuits (ASICs), one or more off-the-shelf products, one or more field programmable gate arrays, any other type or combination of integrated circuits, one or more state machines, or any combination thereof.

[0018] The term "circuit" refers to an arrangement of electronic components (e.g., transistors, resistors, capacitors, and / or inductors) configured to implement one or more functions. For example, a circuit may include one or more transistors interconnected to form logic gates that together implement a logical function.

[0019] As used herein, the terms "determine" and "identify" or any variation thereof include selecting, determining, calculating, finding, receiving, ascertaining, establishing, obtaining, or otherwise using one or more of the devices and methods shown and described herein in any manner to identify or determine.

[0020] As used herein, the terms "example", "embodiment", "implementation", "aspect", "feature", or "element" are used as examples, instances, or illustrations. Unless expressly stated otherwise, any example, embodiment, implementation, aspect, feature, or element is independent of every other example, embodiment, implementation, aspect, feature, or element, and may be used in combination with any other example, embodiment, implementation, aspect, feature, or element.

[0021] As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise stated or clear from the context, "X includes A or B" is intended to indicate any natural inclusive arrangement. That is, if X includes A; X includes B; or X includes both A and B, then "X includes A or B" is satisfied in any of the foregoing instances. Further, unless otherwise stated or clear from the context that the singular form is intended, the articles "a" and "an" as used in this application and the appended claims are generally to be construed to mean "one or more".

[0022] In addition, for simplicity of illustration, although the figures and descriptions herein may include a sequence or series of steps or stages, the elements of the methods disclosed herein may occur in various orders or simultaneously. In addition, the elements of the methods disclosed herein may occur with other elements not explicitly presented and described herein. In addition, implementing the methods according to the present disclosure may not require all of the elements of the methods described herein. Although aspects, features, and elements are described herein in specific combinations, each aspect, feature, or element may be used independently or in various combinations with or without other aspects, features, and elements.

[0023] It should be understood that the figures and descriptions of the embodiments have been simplified to illustrate relevant elements for a clear understanding, and many other elements found in a typical processor have been eliminated for clarity purposes. One of ordinary skill in the art can recognize that other elements and / or steps are desirable and / or required in implementing the present disclosure. However, since such elements and steps are not conducive to a better understanding of the present disclosure, no discussion of such elements and steps is provided herein.

[0024] Figure 1 is a block diagram of an example of a processor 1000 having a core 1005, which includes access prediction logic or an access predictor 1010 for accessing one of an instruction cache (iCache) 1015 or an instruction tightly coupled memory (iTIM) 1020 (which may be appropriately referred to as a memory structure or memory structures) according to an embodiment of the present disclosure. In one implementation, the iTIM 1020 may be part of the core 1005. For example, the processor 1000 may be a computing device, a microprocessor, a microcontroller, or an IP core. The processor 1000 may be implemented as an integrated circuit. The processor 1000 may include one or more cores, and each core may include multiple iTIMs and iCaches. In one implementation, the iTIM 1020 may be part of the core 1005. Without departing from the scope of the present disclosure, the claims, or the figures, the access prediction logic 1010 described herein may be appropriately implemented or modified to account for different combinations of cores, iTIMs, and iCaches. The processor 1000 may be configured to decode and execute instructions of an instruction set architecture (ISA) (e.g., the RISC-V instruction set). The processor 1000 may implement a pipelined architecture. In one implementation, the iTIM 1020 may be a low-latency, dedicated RAM having a defined size and configured to have a memory address range. In one implementation, the iCache 1015 may be an N-way set-associative cache with a fixed cache line length, and each cache way has a defined size. The iCache 1015 may have a defined size and be configured to have a memory address range.

[0025] Processor 1000 may include access prediction logic 1010 to improve power consumption when fetching instructions for execution in a pipeline architecture. The access prediction logic 1010 may be updated in each instruction fetch cycle to determine which memory structure(s) to enable. In one implementation, the access prediction logic 1010 may use feedback from a previous instruction fetch cycle to indicate whether to enable the iTIM, the iCache plus which path, or both of these memory structures. The access prediction logic 1010 may process a location state and a program counter value and output an enable signal to enable the iTIM, the iCache plus that path, or both. In one implementation, the access prediction logic 1010 may take into account memory structure out-of-bounds. In one implementation, the access prediction logic 1010 may include sequential instruction processing logic to take into account sequential instruction processing. For example, the sequential instruction processing logic may resolve the consequences of a cross-boundary condition with respect to the current memory structure enabled by the location state. For example, the boundary condition may include a target memory address outside the memory address range of the current memory structure. In one implementation, if an out-of-bounds condition occurs and if the current memory structure is the iTIM, the location state indicator may indicate to activate both the iTIM and the iCache. In one implementation, if an out-of-bounds condition occurs and if the current memory structure is the iCache, a cache prediction algorithm may be used to determine and update the location state. In one implementation, the access prediction logic 1010 may include a branch predictor, a branch history table, a branch target buffer, and / or a return address stack predictor to take into account branch or jump conditions or scenarios that may affect the location state and the program counter value. In one implementation, the access prediction logic 1010 may take into account branch prediction errors that may affect the location state and the program counter value. For example, due to a prediction error, the location state may now be unknown, and the location state indicator may be set to activate the iCache with all paths or both the iTIM and the iCache.

[0026] Figure 2 is a processor 1000 according to an embodiment of the present disclosure, such as Figure 1 the processor 1000, Figure 3 the processor 3000, and Figure 5Diagram of an example core pipeline 2000 of a processor of processor 5000. The core pipeline 2000 may have five pipeline stages for instruction execution, including an instruction fetch stage 2005, an instruction decode stage 2010, an execute instruction stage 2015, a memory read / write stage 2020, and a write-back stage 2025. In some embodiments, the core pipeline 2000 may include fewer or more pipeline stages without departing from the scope of the disclosure, claims, or figures described herein. The term "instruction fetch cycle" as used herein refers to one iteration of the fetch period or stage of instruction execution.

[0027] In the instruction fetch stage 2005, instructions are appropriately fetched from the iTIM, iCache, or memory. The access prediction logic described herein, such as Figure 1 the access prediction logic 1010 of Figure 3 the access prediction logic 3010 of Figure 5 the access prediction logic 5010 of, may be executed, implemented, or processed, at least in part, during the instruction fetch stage 2005. For example, the program counter may be incremented or updated based on instruction type processing, and the location state may be updated based on instruction type processing and hit / miss conditions related to the instruction cache.

[0028] The fetched instructions may then be decoded in the instruction decode stage 2010 and executed in the execute stage 2015. The access prediction logic described herein, such as Figure 1 the access prediction logic 1010 of Figure 3 the access prediction logic 3010 of Figure 5 the access prediction logic 5010 of, may be updated for branch resolution, at least in part, during the execute stage 2015. For example, in the case of a mispredicted branch, the access prediction logic may be updated with a branch misprediction event, a new program counter value, and a new location state. The access prediction logic may then use this updated information to set which memory structures should be enabled. Depending on the instruction type, a read or write to a memory address may occur during the memory read / write stage 2020, and the result may be written to a register during the write-back stage 2025.

[0029] Figure 3 is a diagram of an example process and access prediction logic 3010 in accordance with an embodiment of the present disclosure. Figure 3Shown is a processor 3000 having a core 3005 according to an embodiment of the present disclosure, which includes access prediction logic or an access predictor 3010 for accessing one of an instruction cache (iCache) 3015 or an instruction tightly coupled memory (iTIM) 3020 according to an embodiment of the present disclosure, and can be implemented and configured as described herein. In one implementation, the iTIM 3020 can be part of the core 3005. For example, the processor 3000 can be a computing device, a microprocessor, a microcontroller, or an IP core. The processor 3000 can be implemented as an integrated circuit. The processor 3000 can include one or more cores, and each core can include multiple iTIMs and iCaches as described herein. Without departing from the scope of the present disclosure, the claims, or the drawings, the access prediction logic 3010 described herein can be appropriately implemented or modified for all combinations. The processor 3000 can be configured to decode and execute instructions of an instruction set architecture (ISA) (e.g., the RISC-V instruction set). The processor 3000 can implement, for example Figure 2 the pipeline architecture shown. The processor 3000 can include access prediction logic 3010 to reduce power consumption when fetching instructions for execution in the pipeline architecture.

[0030] The access prediction logic 3010 can include a program counter 3025, a location status indicator 3030, and enable logic 3035. The location status indicator 3030 can set a location status to activate the iTIM, the iCache plus a specific path, the iCache plus all paths when the location status is unknown (referred to herein as the unknown state), or both the iTIM and the iCache when the location status is unknown. The location status indicator 3030 can send the location status to the enable logic 3010. In one implementation, the location status indicator 3030 can use feedback from a previous instruction fetch cycle to set the location status. In one implementation, the location status can be set to the unknown state at initialization due to out-of-bounds from sequential instruction logic processing, due to branch or jump instruction type processing, due to a branch prediction error, or due to a cache miss. In one implementation, the location status indicator 3030 can use branch resolution information to set or update the location status. For example, the branch resolution information can include a branch prediction error event, a new program counter value, and a new location status. In one implementation, the location status indicator 3030 can use branch prediction logic to update the location status. In one implementation, the location status indicator 3030 can use cache hit / miss to update the location status. In one implementation, the location status indicator 3030 can be updated in each instruction fetch cycle.

[0031] The program counter 3025 can retain the memory address of an instruction (also referred to as the program counter value) when fetching and executing the instruction from the memory. The program counter 3025 can include an incrementer, a selector, and a register. When decoding the fetched instruction, the address of the next sequential instruction is formed by adding the byte length of the current instruction to the current program counter value using the incrementer and placing the next sequential instruction in the register. In the case of a branch, the selector selects the address of the target instruction instead of the incremented value and places the target address in the register. For example, the program counter 3025 can be updated with branch resolution information or from branch prediction logic. The program counter 3025 can send the program counter value to the enable logic 3010. In one embodiment, the program counter 3025 can be updated in each instruction fetch cycle.

[0032] The enable logic 3035 processes inputs from the program counter 3025 and the location status indicator 3030, and enables the iTIM, the iCache, and the appropriate path or both in the case of an unknown state. In one embodiment, the access prediction logic 3010 can include sequential instruction processing logic and branch prediction processing as described herein. For example, the access prediction logic 3010 can include a branch predictor, a branch history table, a branch target buffer, and / or a return address stack predictor to account for branch or jump conditions or scenarios that may affect the location status indicator 3030 and the program counter 3025. In one embodiment, the access prediction logic 3010 can account for branch prediction errors that may affect the location status indicator 3030 and the program counter 3025.

[0033] In one embodiment, the activated iTIM 3020 or iCache 3015 and the appropriate path can return the instruction for decoding in the decode instruction stage as Figure 2 shown, and can appropriately update the location status indicator 3030. In this embodiment, in the case of a cache miss, the instruction can be obtained from, for example, the main memory, and the location status indicator 3030 can be appropriately updated. In one embodiment, the location status indicator 3030 can set the location status to unknown, i.e., both the iTIM 3020 and the iCache 3015 can be set for activation.

[0034] In one embodiment, the enable logic 3035 can activate both memory structures when indicating an unknown state, and the appropriate memory structure can be in Figure 2Return the instruction for decoding in the decoding instruction stage shown. In one embodiment, the position status indicator 3030 can be updated appropriately. In this embodiment, in the case of a cache miss, the instruction can be fetched from, for example, the main memory, and the position status indicator 3030 can be updated appropriately.

[0035] Figure 4 is a diagram of an example technique 4000 for access prediction between an iTIM and an iCache as shown Figure 3 in accordance with an embodiment of the present disclosure. The technique includes: providing 4005 the position status of the current instruction cycle; providing 4010 the program counter value; enabling 4015 one or more memory structures based on the position status and the program counter value; returning 4020 the instruction; and updating 4025 the position status and the program counter value. The technique 4000 can be implemented using Figure 1 the processor 1000, Figure 3 the processor 3000, or Figure 5 the processor 5000.

[0036] The technique 4000 includes providing 4005 the position status of the current instruction cycle. In one embodiment, the position status can be known or unknown. In one embodiment, when the position status is known, the position status indicator can indicate the iTIM or the iCache and a specific path. In one embodiment, when the position status is unknown, the position status indicator can indicate the iCache and all paths, or both the iTIM and the iCache and all paths. The iTIM can be, for example, Figure 1 , Figure 3 or Figure 5 the iTIM shown in. The iCache can be, for example, Figure 1 , Figure 3 or Figure 5 the iCache shown in.

[0037] The technique 4000 includes providing 4010 the program counter value.

[0038] The technique 4000 includes enabling 4015 one or more memory structures based on processing the position status and the program counter value. In one embodiment, the enabled memory structure can be the iTIM. In one embodiment, the enabled memory structure can be the iCache and a specific path. In one embodiment, when the position status is unknown, the enabled memory structure can include the iCache with all paths enabled. In one embodiment, when the position status is unknown, the enabled memory structure can include the iTIM and the iCache with all paths enabled.

[0039] Technique 4000 includes a return 4020 instruction. In one embodiment, the instruction can be returned from an enabled and known memory structure. For example, in the case where the location state is known, the instruction can be returned from a specific path in the iTIM or iCache. In one embodiment, the instruction may not be returned from an enabled and known memory structure and may be returned from, for example, main memory or some other memory hierarchy. In this case, the enabled and known memory structure may be the iCache and there is a cache miss. In one embodiment, the instruction can be returned from an enabled memory structure. For example, in the case where the location state is unknown and both the iTIM and the iCache with all paths are enabled, the instruction can be returned from one of the paths in the iTIM or iCache. In one embodiment, the instruction can be returned from a memory that does not include the iTIM and the iCache - for example, main memory or some other memory hierarchy.

[0040] Technique 4000 includes updating 4025 the location state indicator and the program counter. In one embodiment, the location state indicator and the program counter can be updated appropriately in each instruction fetch cycle using feedback regarding the location state and the program counter value. In one embodiment, the location state indicator and the program counter can be updated based on sequential instruction processing as described herein. In one embodiment, the location state indicator and the program counter can be updated based on branch processing described herein. In one embodiment, the location state indicator and the program counter can be updated based on branch resolution described herein. In one embodiment, the location state indicator and the program counter can be updated based on cache hit / miss processing as described herein.

[0041] Figure 5 is a diagram of an example flow and access prediction logic 5010 in accordance with an embodiment of the present disclosure. Figure 5Shown is a processor 5000 having a core 5005 according to an embodiment of the present disclosure, which includes access prediction logic or an access predictor 5010 for accessing one of an instruction cache (iCache) 5015 or an instruction tightly coupled memory (iTIM) 5020, and may be implemented and configured as described herein. In one implementation, the iTIM 5020 may be part of the core 5005. For example, the processor 5000 may be a computing device, a microprocessor, a microcontroller, or an IP core. The processor 5000 may be implemented as an integrated circuit. The processor 5000 may include one or more cores, and each core may include multiple iTIMs and iCaches. Without departing from the scope of the present disclosure, the claims, or the drawings, the access prediction logic 5010 described herein may be appropriately implemented or modified for all combinations. The processor 5000 may be configured to decode and execute instructions of an instruction set architecture (ISA) (e.g., the RISC-V instruction set). The processor 5000 may implement, for example Figure 2 the pipeline architecture shown. The processor 5000 may include access prediction logic 5010 to reduce power consumption when fetching instructions for execution in the pipeline architecture.

[0042] The access prediction logic 5010 may include a program counter 5025, a next program counter logic 5027, a location status indicator 5030, a next location status logic 5033, and an enable logic 5035. The program counter 5025 may be an input to the next program counter logic 5027, the next location status logic 5033, the enable logic 5035, the iTIM 5020, and the iCache 5015. The next program counter logic 5027 may be an input to the program counter 5025. The location status indicator 5030 may be an input to the enable logic 5035 and the next location status logic 5033. The next location status logic 5033 may be an input to the location status indicator 5030. The enable logic 5035 may be an input to the iTIM 5020 and the iCache 5015 including a specific path.

[0043] The location status indicator 5030 may indicate two known states (iTIM state and iCache state) and an unknown state. The location status indicator 5030 may send the location status to the enable logic 5010. In one implementation, the location status indicator 5030 may be updated by the next location status logic 5033.

[0044] The next position state logic OR circuit 5033 can be used as an OR operation for a state machine to implement transitions between three position states - including the iTIM position state, the iCache plus path position state, and the unknown state. In one embodiment, the next position state logic 5033 may be in the unknown state at initialization due to out-of-bounds from sequential instruction logic processing, due to branch or jump instruction type processing, due to branch prediction errors, or due to cache misses. In one embodiment, state transitions can occur based on inputs from the position state indicator 5030, branch resolution, branch prediction logic, sequential instruction logic processing, and cache hit / miss processing. For example, the next position state logic 5033 can use feedback from the position state indicator 5030 from a previous instruction fetch cycle to update the position state.

[0045] In one embodiment, the next position state logic 5033 can use branch resolution information to update the position state. For example, the branch resolution information can include a branch prediction error event that may occur during the execute instruction phase, a new program counter value, and a new position state, as Figure 2 shown. In one embodiment, the next position state logic 5033 can use branch prediction logic to update the position state, as described herein with respect to Figure 6 the description.

[0046] In one embodiment, the next position state logic 5033 can include sequential instruction logic processing to update the position state. For example, the sequential instruction processing logic can determine whether the program counter value has exceeded the address range of the current memory structure.

[0047] Figure 6 is a diagram of an example process and branch prediction logic, circuit, or predictor 6000 according to an embodiment of the present disclosure. In one embodiment, the branch prediction logic 6000 can be as Figure 5Part of the next position state logic 5033 shown in. The program counter 6003 is an input to the branch prediction logic 6000. The branch prediction logic 6000 may include a branch history table 6005, a branch target buffer 6010, a return address stack 6015, and a multiplexer / selector 6020. The branch history table 6005 may store a bit indicating whether a branch was recently taken for each branch instruction. The branch target buffer 6010 may store a source address, a target address of a predicted taken branch, and a position state of the target address. The return address stack 6015 may store a target address and a position state of the target address when an execution procedure call instruction is executed. The return address of the call instruction may be pushed onto the stack, and when the procedure call is completed, the procedure call will return to the target address of the procedure call instruction. When a return instruction is executed, the address outside the return stack is popped, and it is predicted that the return instruction will return to the popped address. The branch prediction logic 6000 may operate and be implemented as a conventional branch predictor that includes a position state. For example, the multiplexer / selector 6020 receives inputs from the branch history table 6005, the branch target buffer 6010, and the return address stack 6015, and selectively outputs branch prediction information, which includes a predicted program counter value, a predicted taken branch, and a target position state. In one embodiment, the output of the branch prediction logic 6000 may be used by the next position state logic 5033 and the next program counter logic 5027 to appropriately update the position state and the program counter value.

[0048] Now also referring to Figure 5 , the program counter 5025 may retain the memory address of an instruction when the instruction is fetched and executed from memory. The program counter 5025 may include an incrementer, a selector, and a register. When decoding the fetched instruction, the byte length of the current instruction is added to the current program counter value using the incrementer and the next sequential instruction is placed in the register to form the address of the next sequential instruction. In the case of a taken branch, the selector selects the address of the target instruction instead of the incremented value and places the target address in the register. The program counter 5025 may be updated by the next program counter logic 5027.

[0049] The next program counter logic 5027 may selectively update the program counter 5025 using information from the program counter 5025, the branch prediction logic 6000 as shown in Figure 6 , and branch resolution information. In one embodiment, the next program counter logic 5027 may use the updated information to update the program counter 5025.

[0050] In one embodiment, the next program counter logic 5027 may update the program counter value using branch resolution information. For example, the branch resolution information may include a branch prediction error event that may occur during the execute instruction phase, a new program counter value, and a new location state, as Figure 2 shown. In one embodiment, the next program counter logic 5027 may update the program counter value using branch prediction logic, as described herein with respect to Figure 6 this.

[0051] The enable logic 5035 processes inputs from the program counter 5025 and the location state indicator 5030, and enables the iTIM for the iTIM state, enables the iCache and the appropriate path for the iCache state, enables the iCache and all paths for the unknown state, or enables both the iTIM and the iCache in the case of the unknown state. In one embodiment, the activated iTIM 5020 or iCache 5015 and the appropriate path may return an instruction for decoding in the decode instruction phase as Figure 2 shown. In one embodiment, in the case of a cache miss, the instruction may be obtained from, for example, main memory, and the next location state logic 5033 may be updated appropriately. In one embodiment, the next location state logic 5033 may be set to unknown.

[0052] In one embodiment, the enable logic 5035 may activate both memory structures when indicating the unknown state and the appropriate memory structure may return an instruction for decoding in the decode instruction phase, as Figure 2 shown. In one embodiment, the next location state logic 5033 may be updated appropriately. In this embodiment, in the case of a cache miss, the instruction may be obtained from, for example, main memory, and the next location state logic 5033 may be updated appropriately.

[0053] Figure 7 is a diagram of an example technique 7000 for access prediction between an iTIM and an iCache as shown in Figure 5 and Figure 6 in accordance with an embodiment of the present disclosure. The technique includes: providing 7005 the location state of the current instruction cycle; providing 7010 the program counter value; enabling 7015 the memory structure based on the location state and the program counter value; returning 7020 the instruction; determining 7025 the location state; and determining 7030 the program counter value. The technique 7000 may be implemented using the processor 1000 of Figure 1 , the processor 3000 of Figure 3 , or the processor 5000 of Figure 5 .

[0054] Technique 7000 includes providing 7005 the position status of the current instruction fetch cycle. In one embodiment, the position status may be known or unknown. In one embodiment, when the position status is known, the position status may indicate the iTIM or iCache and a specific path. In one embodiment, when the position status is unknown, the position status may indicate the iCache and all paths or both the iTIM and the iCache and all paths. The iTIM may be, for example, the iTIM shown in Figure 1 , Figure 3 or Figure 5 . The iCache may be, for example, the iCache shown in Figure 1 , Figure 3 or Figure 5 .

[0055] Technique 7000 includes providing 7010 the program counter value.

[0056] Technique 7000 includes enabling 7015 a memory structure or memory structures based on the processing position status and the program counter value. In one embodiment, the enabled memory structure may be the iTIM. In one embodiment, the enabled memory structure may be the iCache and a specific path. In an embodiment where the position status is unknown, the enabled memory structure may include the iCache with all paths enabled. In an embodiment where the position status is unknown, the enabled memory structure may include the iTIM and the iCache with all paths enabled.

[0057] Technique 7000 includes returning 7020 an instruction. In one embodiment, the instruction may be returned from an enabled and known memory structure. For example, in the case of a known position status, the instruction may be returned from a specific path in the iTIM or iCache. In one embodiment, the instruction may not be returned from an enabled and known memory structure and may be returned from, for example, the main memory or some other memory hierarchy. In this case, the enabled and known memory structure may be the iCache and there is a cache miss. In one embodiment, the instruction may be returned from an enabled memory structure. For example, in the case of an unknown position status, the instruction may be returned from one of the paths in the iTIM or iCache. In one embodiment, the instruction may be returned from a memory that does not include the iTIM and the iCache - for example, the main memory or some other memory hierarchy.

[0058] Technique 7000 includes determining 7025 the next position state. In one embodiment, the position state can be updated in each instruction fetch cycle using feedback regarding the position state. In one embodiment, the position state can be updated based on sequential instruction processing as described herein. In one embodiment, the position state can be updated based on branch processing described herein. In one embodiment, the position state can be updated based on branch resolution described herein. In one embodiment, the position state can be updated based on cache hit / miss processing as described herein. The next position state can be determined based on the provided updates and the position state can be updated.

[0059] Technique 7000 includes determining 7030 the next program counter. In one embodiment, the program counter value can be updated in each instruction fetch cycle. In one embodiment, the program counter value can be updated based on branch processing as described herein. In one embodiment, the program counter value can be updated based on branch resolution as described herein. The next program counter value can be determined based on the provided updates and the program counter value can be updated accordingly.

[0060] Typically, a processor includes: an instruction tightly coupled memory (iTIM); an instruction cache (iCache) having N paths, where N is at least one; and access prediction logic. The access prediction logic is configured to predict which of the iTIM or the iCache and a specific path to fetch an instruction from, enable the predicted iTIM or iCache and specific path based on a location state and a program counter value, and feedback the location state and the program counter value to predict a next location state for a next instruction, wherein the processor is configured to fetch an instruction via the enabled iTIM or iCache and specific path. In one implementation, the access prediction logic is further configured to set the location state to at least one of the iTIM or the iCache and a specific based on at least the program counter value. In one implementation, the access prediction logic is further configured to: when a next program counter value is within an address range of the enabled iTIM or iCache and specific path, set the location state to the currently enabled iTIM or iCache and specific path for the next instruction. In one implementation, the access prediction logic is further configured to: when a next program counter crosses a boundary defined by an address range of the currently enabled iTIM or iCache and specific path, set the location state to an appropriate iTIM or iCache and specific path for the next instruction. In one implementation. The access prediction logic is further configured to: when a next program counter crosses a boundary defined by an address range of the currently enabled iTIM or iCache and specific path, set the location state to an appropriate iTIM and iCache and all N paths for the next instruction. In one implementation, the access prediction logic is further configured to: in the case of a cache path miss, set the location state to the iCache and a different path for the next instruction. In one implementation, the access prediction logic is further configured to: in the case of a cache miss, set the location state to both the iTIM and the iCache and all N paths for the next instruction. In one implementation, the access prediction logic is further configured to: in the case of a mispredicted branch for the next instruction, set the location state to the iCache and all N paths or both the iTIM and the iCache and all N paths. In one implementation, the access prediction logic predicts the location state based on at least branch resolution processing, branch prediction processing, sequential instruction logic processing, cache hit / miss processing, previous program counter values, and previous location states.

[0061] Generally, a method for prediction between memory structures, the method comprising: providing a location state; providing a program counter value; predicting which one of an instruction tightly coupled memory (iTIM) or an instruction cache (iCache) with N paths of a specific path to fetch an instruction from; enabling activation of the predicted iTIM or iCache and the specific path based on the location state and the program counter value; feeding back the location state and the program counter value to predict a next location state for a next instruction; and returning the instruction. In one embodiment, the method further comprises: setting the location state to the iTIM or iCache and at least one of the specific ones based on at least the program counter value. In one embodiment, the method further comprises: when a next program counter value is within an address range of the enabled iTIM or iCache and the specific path, setting the location state to the currently enabled iTIM or iCache and the specific path for the next instruction. In one embodiment, the method further comprises: when a next program counter crosses a boundary defined by an address range of the currently enabled iTIM or iCache and the specific path, setting the location state to an appropriate iTIM or iCache and the specific path for the next instruction. In one embodiment, the method further comprises: when a next program counter crosses a boundary defined by an address range of the currently enabled iTIM or iCache and the specific path, setting the location state to an appropriate iTIM and iCache and all N paths for the next instruction. In one embodiment, the method further comprises: in the case of a cache path miss, setting the location state to the iCache and a different path for the next instruction. In one embodiment, the method further comprises: in the case of a cache miss, setting the location state to the iTIM and iCache and all N paths for the next instruction. In one embodiment, the method further comprises, in the case of a branch misprediction for a next instruction, setting the location state to the iCache and all N paths or both the iTIM and iCache and all N paths. In one embodiment, wherein the prediction is based on at least branch resolution processing, branch prediction processing, sequential instruction logic processing, cache hit / miss processing, previous program counter values, and previous location states.

[0062] Generally, a processor includes: an instruction tightly coupled memory (iTIM), an instruction cache (iCache) having N paths, where N is at least one; a program counter configured to store a program counter value; a location status indicator configured to store a location status; and an enabling circuit configured to enable one of the iTIM or the iCache and a specific path based on the location status and the program counter value, wherein the processor is configured to obtain an instruction through the enabled iTIM or iCache and the specific path. In one embodiment, it further includes a next location status logic configured to set the location status to at least one of the iTIM or the iCache and a specific path based on at least branch resolution processing, branch prediction processing, sequential instruction logic processing, cache hit / miss processing, a previous program counter value, and a previous location status.

[0063] Although some embodiments herein relate to methods, those skilled in the art will understand that they can also be embodied as a system or a computer program product. Thus, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects, which aspects are generally herein referred to as a "processor", "device" or "system". In addition, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon. Any combination of one or more computer-readable media may be utilized. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing parts. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0064] A computer-readable signal medium can include, for example, a propagated data signal in a baseband or as part of a carrier wave, which contains computer-readable program code. Such a propagated signal can take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium that is not a computer-readable storage medium and can communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0065] The program code contained on a computer-readable medium can be transmitted using any suitable medium, including but not limited to CD, DVD, wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0066] Computer program code for performing operations in accordance with aspects of the present invention can be written in any combination of one or more programming languages, including: object-oriented programming languages such as Java, Smalltalk, C++, etc.; and traditional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0067] Aspects of the present invention are described below with reference to the flowchart and / or block diagram of a method, apparatus (system), and computer program product according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions.

[0068] These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing device create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer program instructions can also be stored in a computer-readable medium, which can direct a computer, other programmable data processing device, or other device to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instructions for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0069] Computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0070] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures.

[0071] Although the present disclosure has been described in connection with certain embodiments, it should be understood that the present disclosure is not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications, combinations, and equivalent arrangements included within the scope of the appended claims, which scope is to be accorded the broadest interpretation so as to cover all such modifications and equivalent structures as are permitted under the law.

Claims

1. A processor, comprising: An instruction tightly coupled memory iTIM; An instruction cache iCache having N paths, where N is at least one; and Access prediction logic, the access prediction logic being configured to: Maintain a location status indicator that indicates whether the location status for obtaining an instruction is known or unknown; Predict which one of the iTIM or the iCache and a specific path to obtain an instruction based on the location status indicator and the program counter value; When the location status indicator indicates that the location status is known, enable the predicted iTIM, or the iCache and the specific path; When the location status indicator indicates that the location status is unknown, enable both of the N paths of the iTIM and the iCache for parallel access; and, Update the location status indicator and the program counter value to predict the next location status for the next instruction, wherein the processor is configured to obtain the instruction via the enabled iTIM, the enabled iCache and the specific path, or the enabled iCache and all N paths.

2. The processor according to claim 1, wherein the access prediction logic is further configured to: Set the location status in the location status indicator to at least one of the iTIM or the iCache and the specific path based on at least the program counter value.

3. The processor according to claim 1, wherein the access prediction logic is further configured to: When the next program counter value is within the address range of the enabled iTIM or the enabled iCache and the specific path, set the location status in the location status indicator for the next instruction to the currently enabled iTIM or the currently enabled iCache and the specific path.

4. The processor according to claim 1, wherein the access prediction logic is further configured to: When the next program counter crosses a boundary defined by the address range of the currently enabled iTIM or the currently enabled iCache and the specific path, set the location status in the location status indicator for the next instruction to the appropriate iTIM or iCache and specific path.

5. The processor according to claim 1, wherein the access prediction logic is further configured to: When the next program counter crosses a boundary defined by the address range of the currently enabled iTIM or the currently enabled iCache and the specific path, set the location status in the location status indicator for the next instruction to the appropriate iTIM and iCache and all N paths.

6. The processor according to claim 1, wherein the access prediction logic is further configured to: In the case of a cache path miss, set the location status in the location status indicator for the next instruction to the iCache and a different path.

7. The processor according to claim 1, wherein the access prediction logic is further configured to: In the case of a cache miss, set the position state in the position state indicator to the iTIM, the iCache, and all N paths for the next instruction.

8. The processor according to claim 1, wherein the access prediction logic is further configured to: In the case of a branch misprediction for the next instruction, set the position state in the position state indicator to either the iCache and all N paths or both the iTIM, the iCache, and all N paths.

9. The processor according to claim 1, wherein, The access prediction logic predicts the position state in the position state indicator based on at least branch resolution processing, branch prediction processing, sequential instruction logic processing, cache hit / miss processing, previous program counter value, and previous position state.

10. A method for prediction between memory structures, the method comprising: Providing a position state indicator that indicates whether the position state for obtaining an instruction is known or unknown; Providing a program counter value; Predicting which of an instruction tightly coupled memory iTIM or an instruction cache iCache with N paths to obtain an instruction based on the position state indicator and the program counter value; When the position state indicator indicates that the position state is known, enabling the activation of the predicted iTIM or the iCache and the specific path; When the position state indicator indicates that the position state is unknown, enabling both the iTIM and all N paths of the iCache for parallel access; Updating the position state indicator and the program counter value to predict the next position state for the next instruction; And Returning an instruction via the enabled iTIM, the enabled iCache and the specific path, or the enabled iCache and all N paths.

11. The method according to claim 10, further comprising: Setting the position state to at least one of the iTIM or the iCache and the specific path based on at least the program counter value.

12. The method according to claim 10, further comprising: When the next program counter value is within the address range of the enabled iTIM or the enabled iCache and the specific path, setting the position state in the position state indicator to the currently enabled iTIM or the enabled iCache and the specific path for the next instruction.

13. The method according to claim 10, further comprising: When the next program counter crosses a boundary defined by the address range of the currently enabled iTIM or the currently enabled iCache and the specific path, setting the position state in the position state indicator to the appropriate iTIM or iCache and specific path for the next instruction.

14. The method according to claim 10, further comprising: When the next program counter crosses a boundary defined by the address range of the currently enabled iTIM or the currently enabled iCache and the specific path, set the position state in the position state indicator to the appropriate iTIM and iCache and all N paths for the next instruction.

15. The method according to claim 10, further comprising: In the case of a cache path miss, set the position state in the position state indicator to the iCache and a different path for the next instruction.

16. The method according to claim 10, further comprising: In the case of a cache miss, set the position state in the position state indicator to the iTIM and the iCache and all N paths for the next instruction.

17. The method according to claim 10, further comprising: In the case of a branch misprediction for the next instruction, set the position state in the position state indicator to either the iCache and all N paths or both the iTIM and the iCache and all N paths.

18. The method according to claim 10, wherein The prediction is based on at least branch resolution processing, branch prediction processing, sequential instruction logic processing, cache hit / miss processing, previous program counter value, and previous position state.

19. A processor, comprising: Instruction tightly coupled memory iTIM; An instruction cache iCache having N paths, where N is at least one; A program counter configured to store a program counter value; A position state indicator configured to indicate whether the position state for fetching an instruction is known or unknown; and An enabling circuit configured to: When the position state indicator indicates that the position state is known, enable one of the iTIM or the iCache and a specific path; and When the position state indicator indicates that the position state is unknown, enable both of the iTIM and all N paths of the iCache for parallel access, wherein the processor is configured to fetch an instruction through the enabled iTIM, the enabled iCache and the specific path, or the enabled iCache and all N paths.

20. The processor according to claim 19, further comprising: Next position state logic configured to set the position state in the position state indicator to at least one of the iTIM or the iCache and the specific path based on at least branch resolution processing, branch prediction processing, sequential instruction logic processing, cache hit / miss processing, previous program counter value, and previous position state.

Citation Information

Patent Citations

  • Data cache way prediction

    US20100049912A1