Joint scheduler for high bandwidth multi-shot prefetching

The joint scheduler addresses cache management limitations by co-scheduling prefetch and request accesses, reducing latency and increasing bandwidth through efficient cache utilization and out-of-order execution.

JP2025094927APending Publication Date: 2025-06-25NEXTSILICON LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024217460
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-12
Filing Date
2024-12-12
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

Conventional cache management techniques exhibit limitations in responsiveness and adaptability to varying workloads and data access patterns, affecting data search performance in modern computing systems.

Method used

A joint scheduler is implemented to co-schedule prefetch and request accesses, associating instructions with cache entries and tracking invalidation/eviction, allowing for efficient dispatch of prefetch cycles to ensure data availability in cache during request accesses.

Benefits of technology

This approach reduces memory access latency and increases bandwidth by eliminating redundant cache look-ups and enabling out-of-order execution, enhancing processing unit performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094927000001_ABST
    Figure 2025094927000001_ABST
Patent Text Reader

Abstract

To provide a joint scheduler and method that improve scheduler efficiency to reduce latency and increase execution bandwidth.SOLUTION: A joint scheduler is adapted such that, in response to each hit prefetch access dispatched to respective data relating to respective instructions, the requested accesses relating to respective instructions access respective data in a cache using pointers to respective cache entries storing the respective data, the pointers being associated with the respective instructions, and so as to associate each instruction with a valid indication and a pointer, and in response to each missed prefetch access dispatched for each data relating to each instruction, initiate a read cycle for loading the respective data from a next level memory and cache the respective data in the cache.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In some embodiments, the present invention relates to scheduling memory access in a processing unit, and more specifically, but not limited to, using a joint scheduler to schedule memory access in a processing unit for improving efficiency, reducing latency, and / or increasing bandwidth.

[0002] This application claims the benefit of priority of U.S. Patent Application No. 18 / 537,927, filed on December 13, 2023, the content of which is incorporated herein by reference in its entirety. This application is also related to U.S. Patent Application No. 18 / 770,690, filed on July 12, 2024. The content of the above applications is incorporated herein by reference in its entirety as if fully set forth herein.

Background Art

[0003] In modern computing systems, the performance of a processor highly depends on the effective utilization of cache memory.

[0004] A cache using a high-speed memory device may be deployed close to, and often within, the processor itself to store frequently accessed data and / or instructions, thereby reducing the latency associated with fetching information from main memory, which is significantly slower and has a long access time.

[0005] However, conventional cache management techniques may exhibit inherent limitations in terms of responsiveness and / or adaptability to varying workloads and data access patterns, which can directly affect data search performance.

Summary of the Invention

[0006] The object of the present invention is to provide a method, a system, and a software program product for improving the efficiency of a scheduler of a processing unit to reduce latency and increase execution bandwidth. The above and other objects are achieved by the features of the independent claims. Further implementation forms are apparent from the dependent claims, the description, and the drawings.

[0007] According to a first aspect of the present invention, there is provided a joint scheduler comprising a joint scheduler circuit adapted to dispatch prefetch accesses and request accesses for data related to a plurality of instructions loaded in an execution pipeline of one or more processing circuits. Each prefetch access includes checking whether each respective data is cached in one of a plurality of cache entries of one or more caches, and each request access includes accessing each respective data. The joint scheduler circuit is adapted to: (1) associate each instruction with a valid instruction and a pointer such that, in response to each hit prefetch access dispatched for each respective data related to each one of the plurality of instructions, the request access related to each instruction accesses each respective data in one or more caches using a pointer to each cache entry storing each respective data associated with each instruction; and (2) in response to each miss prefetch access dispatched for each respective data related to each one of the plurality of instructions, start a read cycle for loading each respective data from the next level of memory and cache each respective data in one or more caches.

[0008] According to a second aspect of the present invention, there is provided a method of co-scheduling prefetch accesses and request accesses, including using a joint scheduler circuit adapted to dispatch prefetch accesses and request accesses for data regarding a plurality of instructions loaded into an execution pipeline of one or more processing circuits. Each prefetch access includes checking whether respective data is cached in one of a plurality of cache entries of one or more caches, and each request access includes accessing respective data. The joint scheduler circuit is adapted to: (1) in response to each hit prefetch access dispatched for respective data regarding each one of the plurality of instructions, associate each instruction with a valid indication and a pointer such that the request access regarding each instruction accesses respective data in one or more caches using the pointer to the respective cache entry storing the respective data associated with each instruction; and (2) in response to each miss prefetch access dispatched for respective data regarding each one of the plurality of instructions, start a read cycle to load respective data from the next level of memory and cache respective data in one or more caches.

[0009] According to a third aspect of the present invention, there is provided a joint scheduler comprising a joint scheduler circuit adapted to enhance out-of-order execution of a plurality of instructions loaded into an execution pipeline of one or more processing circuits by tracking invalidation and / or eviction of data stored in each of a plurality of cache entries of one or more caches storing data regarding the plurality of instructions, and dispatching a prefetch cycle to make respective data available in one or more caches during a request access dispatched for one or more of the plurality of instructions by loading the invalidated and / or evicted data regarding one or more of the plurality of instructions.

[0010] In a further implementation form of the first, second, and / or third aspect, the joint scheduler circuit loads each data from the next level of memory after each missed prefetch access, and in response to the successful completion of each read cycle initiated to cache each data in one or more caches, each instruction is adapted to be associated with a valid instruction and a pointer to each cache entry that stores the cached data.

[0011] In a further implementation form of the first, second, and / or third aspect, the joint scheduler circuit loads each data from the next level of memory after each missed prefetch access, and in response to the successful completion of each read cycle initiated to cache each data in one or more caches, another prefetch access is dispatched to update the valid instruction and the pointer to each cache entry that stores the cached data.

[0012] In an optional implementation form of the first, second, and / or third aspect, the joint scheduler circuit tracks the invalidation and / or eviction of data stored in each of a plurality of cache entries, and in response to the eviction of each cache entry that stores the data for each one of a plurality of instructions, each instruction is associated with an invalidation instruction and further adapted to start another prefetch cycle to load the data from the next level of memory into one or more caches.

[0013] In an optional implementation of the first, second, and / or third aspect, the joint scheduler circuit is further adapted to mark each cache entry with an active indication for each prefetch access hit, where the active indication indicates that each cache entry is mapped by a pointer associated with one or more of the instructions, and the active indication is used by the joint scheduler circuit to track invalidation and / or eviction of the data stored in each cache entry.

[0014] In a further implementation of the first, second, and / or third aspect, the joint scheduler circuit is adapted to associate each instruction with a block mark indicating that each prefetch access and / or each request access for each instruction is blocked for dispatch.

[0015] In a further implementation of the first, second, and / or third aspect, the joint scheduler circuit is adapted to remove the block mark associated with each instruction in response to successful completion of the prefetch cycle.

[0016] In a further implementation of the first, second, and / or third aspect, the joint scheduler circuit comprises one or more prefetch ports for dispatching prefetch accesses and one or more request ports for dispatching request accesses. The one or more prefetch ports are separate from and independent of the one or more request ports.

[0017] In a further implementation of the first, second, and / or third aspect, one or more prefetch accesses and one or more request accesses are simultaneously dispatched via one or more independent prefetch ports and one or more request ports, respectively.

[0018] In a further implementation of the first, second, and / or third aspects, for one or more simultaneously dispatched prefetch accesses and one or more request accesses to a common cache entry, in response to completion of the one or more prefetch accesses, at one or more request ports, for the one or more request accesses, a pointer to the common cache entry is directly updated.

[0019] In a further implementation of the first, second, and / or third aspects, the joint scheduler circuit is adapted to dispatch a prefetch cycle in response to eviction and / or invalidation of one or more cache entries storing respective data for each of a plurality of instructions, to load the respective data from a next level of memory into one or more caches, and to associate each instruction with a valid indication and a pointer such that a request access for each instruction accesses the respective data in one or more caches using the pointer to the respective cache entry storing the respective data associated with each instruction.

[0020] In an alternative implementation of the first, second, and / or third aspects, the joint scheduler circuit is further adapted to associate each instruction with a valid indication and a pointer by starting a separate prefetch cycle for each instruction. The joint scheduler circuit is adapted to associate each instruction with a valid indication and a pointer in response to a hit in a separate prefetch cycle.

[0021] Other systems, methods, features, and advantages of the present disclosure will be apparent to or will become apparent to those of ordinary skill in the art upon examination of the following drawings and detailed description. All such additional systems, methods, features, and advantages are intended to be included within this description, be within the scope of the present disclosure, and be protected by the accompanying claims.

[0022] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present invention, but exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. Also, the materials, methods, and examples are illustrative only and not intended to be limiting necessarily.

[0023] The implementation of the methods and / or systems of embodiments of the present invention may include automatically performing or completing selected tasks. Further, depending on the actual instrumentation and facilities of the methods and / or systems of embodiments of the present invention, some selected tasks may be implemented by hardware, software, or firmware, or combinations thereof using an operating system.

[0024] For example, the hardware for performing selected tasks according to embodiments of the present invention may be implemented as a chip or a circuit. Selected tasks according to embodiments of the present invention may be implemented as a plurality of software instructions executed by a computer using any suitable operating system as software. In an exemplary embodiment of the present invention, one or more tasks according to the exemplary embodiments of the methods and / or systems described herein are executed by a data processor such as a computing platform for executing a plurality of instructions. Optionally, the data processor includes volatile memory for storing instructions and / or data, and / or non-volatile storage devices for storing instructions and / or data, such as magnetic hard disks and / or removable media. Optionally, a network connection is also provided. A display, and / or a user input device such as a keyboard or a mouse are also optionally provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In some embodiments of the present invention, the present specification will be described by way of example only with reference to the accompanying drawings. Referring specifically to the drawings in detail here, it is emphasized that the details shown are for illustration and for the purpose of exemplarily considering embodiments of the present invention. In this regard, those skilled in the art will be made clear by the description using the drawings how the embodiments of the present invention can be implemented.

[0026] <Description of the Drawings>

[0027]

Figure 1

[0028]

Figure 2A

Figure 2B

[0029]

Figure 3A

Figure 3B

[0030]

Figure 4A

Figure 4B

Embodiments for Carrying Out the Invention

[0031] In some embodiments of the present invention, it relates to scheduling memory accesses in a processing unit, and more specifically, but not limited to, using a joint scheduler to schedule memory accesses in a processing unit by improving efficiency, reducing latency, and / or increasing bandwidth.

[0032] According to some embodiments of the present invention, there are provided an apparatus, a method, and a computer program product for improving memory accesses dispatched by a processing unit by jointly scheduling prefetch accesses and request accesses.

[0033] Memory latency and access bandwidth are important factors in the overall performance of a processing unit. To reduce latency and memory access bandwidth, a cache is used to hold segments of the memory image.

[0034] Caches used with high-speed memory arrays are typically deployed (L1 cache) close to and at a short distance from important drivers (write operations) or receivers (read operations) of the processing unit, often within the processing unit itself, thereby supporting fast access. However, while fast, caches may have limited capacity and thus may only be able to store relatively small segments of the memory image.

[0035] Therefore, the performance of memory access depends on the cache lookup hit rate. Cache lookup refers to searching for a match between the address of a memory access (read or write operation) and one of the addresses of the data stored in the cache. That is, an address cache typically includes a tag field and an index field. A cache hit means that the address of the accessed data is found in the address cache, and thus the accessed data is cached and stored in the cache.

[0036] In the case of a miss, where the data portion of the memory image required by the read or write operation is not cached, the required data portion must be fetched from the next level of memory, such as main memory, or at least the next higher-level memory, such as an L2 cache, an L3 cache, etc. This can not only increase the latency of the operation due to the long access time imposed by the next level of memory and the possibility of multiple accesses (additional accesses competing with other cache accesses), but also degrade the cache bandwidth.

[0037] Since the cache capacity is limited, when a new memory line (the smallest memory portion cached in the cache) is loaded into the cache, an old memory line must be evicted from the cache to make room for the new memory line. This operation is called replacement because the new memory line replaces the old memory line. When the evicted memory line is accessed again, it must be reloaded into the cache from the higher-level memory.

[0038] To improve the cache hit rate, prefetching may be applied, which means that the prefetch cycle may be started (dispatched) in advance to load data required by one or more instructions into the cache before the "actual" request access is dispatched during the execution of each instruction. In this way, when the actual request access is dispatched, it can be seen that the required data has already been loaded and cached in the cache during the execution of the instruction. In particular, as described above in this specification, during the prefetch access, a complete memory line containing the required data is loaded (cached) into the cache.

[0039] One of the common prefetching algorithms is early cache lookup (hereinafter also referred to as prefetch access in this specification). This is a technique designed to overcome a blocking situation where an instruction may be blocked and its request access (read or write operation) may not be dispatched until one or more conditions are met even if the access address is known.

[0040] Such situations include, for example, an instruction fetch operation dispatched to read data from the instruction cache. Instruction fetch operations are typically performed in order such that each fetch must wait until the previous fetch is complete before each fetch reads data from the instruction cache. If a particular instruction fetch operation is delayed (e.g., due to a cache miss), subsequent instruction fetch operations must be delayed until the delayed fetch operation is complete even if the memory address is known. In another example, a data store operation dispatched to write data to the data cache is an operation that cannot be undone because the old data cannot be restored by being overwritten by the new data. Therefore, the data store operation can be executed only after all previous instructions have been committed even if the memory address is known.

[0041] Since the memory address of the requested access (read operation or write operation) is known, even if the actual request cannot be executed due to the blocking condition, the corresponding prefetch access may be pre-dispatched to the same address and an early lookup may be performed in the cache. In this way, the prefetch access can check whether the memory line containing the data related to the instruction exists in the cache (whether it is cached) without performing the actual requested access (read operation or write operation). In the case of a cache hit, that is, when the requested data is cached, there is no need to do anything further. However, in the case of a cache miss, an early read operation may be dispatched to load the memory line from the upper memory level, and when the early read operation is dispatched, the memory line can be cached in the cache, thereby reducing the latency of the actual request.

[0042] Although the latest processing unit may be able to execute multiple instructions in parallel, it is clear that only a limited number of memory accesses and cache accesses can be dispatched in a given cycle. Therefore, a scheduler is implemented to determine which instruction should be dispatched in each cycle and schedule the instructions accordingly.

[0043] A scheduler may typically include ready logic and a scheduling module. The ready logic is configured to calculate which instructions are eligible to be dispatched. An instruction (an entry in the scheduler) is ready to be dispatched if it is assigned, not in progress in the dispatch pipe, eligible to be dispatched (e.g., the oldest fetch), and not blocked. An instruction may be blocked, for example, if a problem was revealed in a previous dispatch. The blocking is removed (woken up) only after the cause (reason) of the blocking is resolved. For example, a cache miss blocks the dispatch of a given entry until the required memory line is loaded from the next level of memory and cached in the cache.

[0044] A scheduling module that may include one or more dispatch ports may dispatch one or more accesses to ready instruction entries. The scheduling module may employ one or more scheduling mechanisms such as, for example, age based scheduling (selecting the oldest entry first) and / or location based scheduling (selecting the first ready entry in the scheduler). Each scheduled instruction is dispatched, performs a cache lookup, and updates the scheduler at the end of the lookup pipe depending on the completion of the cache lookup, for example, success if the instruction's assignment to the scheduler is released and / or failure in case of a blocking condition.

[0045] According to some embodiments of the present invention, an apparatus, a method, and a computer program product are provided for improving memory accesses dispatched by a processing unit by co-scheduling prefetch accesses and request accesses.

[0046] In particular, the joint scheduler of the processing unit is adapted to schedule the memory access unit, not by dispatching prefetch accesses and data request accesses separately and independently of each other as can be done by an existing scheduler, but by dispatching a prefetch access and a dispatched data request access, or the same instructions that are in correlation with each other, to the memory access unit.

[0047] The joint scheduler may dispatch a prefetch access to look up a cache for data regarding each instruction loaded into the execution pipeline of the processing unit. In the case of a cache hit, i.e., when the requested data is found in the cache, that is, when the data is stored (cached) in one of the cache entries of the cache, the joint scheduler may associate each instruction with a valid indication (e.g., a flag) and a pointer to the cache entry storing the data.

[0048] Also, in the case of a cache miss, i.e., when the requested data is not found in the cache, a read cycle may be initiated to fetch the data from the next level of memory and load the data into the cache. After the completion of the read cycle is successful, the joint scheduler may associate each instruction with a valid indication and a pointer to the cache entry storing the data, as was done in the case of a cache hit. Alternatively, after the completion of the read cycle is successful, the joint scheduler may dispatch another prefetch access that results in a cache hit, which leads to associating each instruction with a valid indication and a pointer to the cache entry storing the data.

[0049] When each instruction is executed, the joint scheduler identifies the valid instructions associated with each instruction, dispatches the required accesses of each instruction using the pointers associated with each instruction, and directly accesses the cache entries mapped by the associated pointers, eliminating the need to search the address cache (tag / index) for the required accesses.

[0050] The joint scheduler may further monitor all of the cache entries pointed to by the pointers associated with the instructions to detect invalidation and / or eviction (removal) of the cache lines stored in these cache entries. If the cache entries mapped by the pointers associated with each instruction are invalidated and / or evicted, the joint scheduler may remove the valid instructions associated with each instruction to indicate that the data for each instruction is no longer cached, i.e., not stored in the cache.

[0051] Furthermore, in response to invalidation and / or eviction, the joint scheduler may dispatch another prefetch access to reload the previously invalidated and / or evicted data and make it available again for subsequent required accesses of each instruction. During another prefetch access, each instruction may be re-associated with a valid instruction and a pointer that maps to the cache entry that stores the reloaded data. Thus, the required accesses of each instruction dispatched by the joint scheduler may directly access the reloaded data stored in the cache entry mapped by the associated pointer using the pointer associated with each instruction.

[0052] Dispatching correlated prefetch accesses and request accesses using a joint scheduler can present significant benefits and advantages compared to existing schedulers.

[0053] First, searching (looking up) the cache's address cache (tag / index fields) and comparing it with the address of the data for each instruction (look-up) can be done only once during the prefetch access dispatched for each instruction whose execution is scheduled. This is because the request access dispatched for each instruction can directly access the cache entry mapped by the pointer associated with each instruction, eliminating the need for another look-up via the cache's tag / index fields. This can significantly reduce the latency of memory and cache accesses and / or increase the access bandwidth.

[0054] Furthermore, the single look-up is performed while each instruction is waiting to be executed, even if the request access for each instruction dispatched during the execution of each instruction does not require a look-up. This can significantly accelerate the execution of instructions and greatly increase the execution performance of the processing unit in terms of time, speed, etc.

[0055] Some existing schedulers can apply a locked prefetch where, if an early prefetch (look-up) results in a cache hit, the hit cache line is not invalidated and / or evicted (replaced), and the pointer to the associated cache entry is locked to be held by the scheduler for use by an actual request access.

[0056] This solution can increase the access bandwidth by eliminating the second lookup by requested access, but there can be significant limitations to this mechanism. First, if all potentially used cache entries of the cache are locked, no additional memory accesses can be provided to request the loading of additional data that is not currently cached. Some schedulers may employ a dedicated mechanism to unlock locks to enable progress. However, such an unlocking mechanism can be very complex and may require extensive hardware resources (e.g., logic circuits, memory cells, etc.) and / or may degrade memory access performance. Additionally, locking cache entries can substantially reduce the cache size, and new data loaded into the cache must replace another entry that may have a higher priority, potentially degrading performance as the cache hit rate decreases.

[0057] On the other hand, the joint scheduler may not lock cache entries that can prevent further data cache blocks, and thus may not degrade the memory access latency and / or bandwidth that can prevent a decrease in execution performance. Rather, the joint scheduler may monitor cache entries mapped by pointers associated with the instructions whose execution is scheduled, track invalidation and / or eviction, and, in the case of invalidation and / or eviction, may reload the data for the instruction.

[0058] Other existing schedulers may utilize a dedicated prefetch buffer so that, in parallel with assigning instructions written to the scheduler to the instructions, the instructions may also be written to the prefetch buffer. The prefetch buffer may be adapted to initiate a prefetch access to look up the address of the data cached in the cache, while the scheduler, in addition to accessing the dedicated prefetch buffer, may issue only the actual request access to look up the cache as if no prefetch had occurred. The prefetch access may be issued only when the data requested by the request access is not found in the prefetch buffer, to send an outstanding request at cache miss. The prefetch access may not be updated in the scheduler.

[0059] This mechanism also presents several major drawbacks and limitations. First, the bandwidth of the prefetch access (lookup) can be significantly poor as potentially each data request is dispatched twice, once during the prefetch access and again during the request access. Further, since the prefetch buffer is merely a duplicate of the scheduler, it consumes additional hardware resources or may reduce in size, which clearly can limit the prefetch capacity and / or cause a drop in the prefetch access. Further, unlike the scheduler, the prefetch buffer may not support age-based scheduling. Also, the prefetch access may be dispatched only once, and if the cache line is evicted after the prefetch access is made to the cache line, no additional prefetch may be dispatched until the actual request access. Thus, the request access has to perform a cache lookup again and may need to access the next level of memory, which can further reduce the memory bandwidth and / or increase the access latency.

[0060] In contrast, when using a joint scheduler, only a single lookup is required during the prefetch access for each instruction before each instruction is executed, while the request access dispatched when each instruction is executed can directly access the mapped cache entry using the pointer associated with each instruction without performing a lookup. By performing the lookup only once, the latency of the request access can be significantly reduced and / or the access bandwidth can be increased. Further, since the joint scheduler does not employ a dedicated buffer, it may not increase the consumption of hardware resources on one hand and may increase the size on the other hand, thereby further reducing the latency of memory access and / or increasing the memory access bandwidth. Additionally, the joint scheduler may employ multiple scheduling algorithms and / or techniques, including other scheduling algorithms such as location-based scheduling in addition to age-based scheduling.

[0061] Furthermore, the joint scheduler can significantly increase the out-of-order execution performance of the processing unit. Since the joint scheduler continuously monitors cache entries to track and identify invalidation events and / or eviction events, the joint scheduler may reload previously invalidated and / or evicted data for instructions in the execution pipeline regardless of their order. Thus, even if data for instructions located further downstream in the pipeline is evicted, this data can be quickly and automatically reused, enabling earlier execution of further downstream instructions, thereby significantly increasing the execution performance of the processing unit as is known in the art.

[0062] Before describing in detail at least one embodiment of the present invention, it is to be understood that the present invention is not necessarily limited to the details of construction and the arrangement of components and / or methods described in the following description and / or shown in the drawings and / or examples in its application. The present invention is capable of other embodiments or of being practiced or carried out in various ways.

[0063] As will be understood by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software aspects and hardware aspects, all of which may generally be referred to as a "circuit," "module," or "system." Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0064] Any combination of one or more computer-readable media may be utilized. A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A more specific list of examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, punch cards, mechanically encoded devices such as raised structures within grooves recording instructions, and any suitable combination of the foregoing, but this list is non-exhaustive. As used herein, a computer-readable storage medium shall not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0065] Computer program code including computer-readable program instructions embodied on a computer-readable medium may be transmitted using any appropriate medium (including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination thereof).

[0066] The computer-readable program instructions described herein can be downloaded from a computer-readable recording medium to respective computing / processing devices, or to an external computer or external storage device, via, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.

[0067] The computer-readable program instructions for carrying out operations of the present invention may be written in any combination of one or more programming languages, for example, assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, etc., or source code or object code written in any combination of one or more programming languages including, for example, object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0068] Computer-readable program instructions may be executed entirely on a user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to implement aspects of the present invention.

[0069] As used herein, aspects of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0070] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions described in the blocks may be performed in an order different from that shown in the drawings. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, as well as combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a special hardware-based system that performs the specified functions or operations, or by a combination of special hardware and computer instructions.

[0071] Referring to the drawings, FIG. 1 is an exemplary process flowchart using a joint scheduler of a processing unit configured to efficiently schedule memory accesses according to some embodiments of the present invention.

[0072] The exemplary process 100 is executed by the joint scheduler of each of one or more processing units to dispatch a prefetch access and a dispatched data request access, or the same instructions that are correlated with each other, to improve the performance of data loading and / or data caching, such as reducing latency, increasing bandwidth, and / or reducing lock time, thereby efficiently scheduling memory accesses in a processing unit having one or more cache memory units.

[0073] The instructions loaded into the execution pipeline of the processing unit may access data stored in a memory accessible to the processing unit, for example, may read and / or write data.

[0074] As is known in the art, to improve performance and reduce latency, the processing unit may include one or more high-speed low-latency cache memories for caching data stored in the memory of the processing unit, for example, an L1 cache and / or an L2 cache, etc. However, while supporting high-speed access, the cache may have limited capacity, so the cache can store only a limited amount of data at any given time.

[0075] When each instruction is executed, the data for a particular instruction may not be cached (not loaded into the cache) instead of being loaded. Therefore, to further improve data access performance, prefetch access to load this data into the cache (caching) may be started in advance, earlier in time, before these instructions are executed, whereby the access time, and thus the execution time of the instructions, can be significantly reduced.

[0076] Also refer to FIGS. 2A and 2B, which are schematic diagrams of an exemplary joint scheduler of a processing unit configured to efficiently schedule memory accesses according to some embodiments of the present invention.

[0077] As shown in FIG. 2A, an exemplary joint scheduler 200, also referred to as a dual dispatch scheduler, of an exemplary processing unit (processor) 202 may include one or more joint scheduler circuits configured and / or adapted to dispatch both prefetch accesses and request accesses for accessing data related to one or more instructions loaded into the execution pipeline 210 of the processing unit 202, such as data reads and / or data writes.

[0078] As is known in the art, the processing unit 202 may include one or more units, circuits, and / or modules, etc. For example, the processing unit 202 may include an input circuit 212 and an output circuit 214 for receiving and outputting data.

[0079] The processing unit 202 may further include a control unit 216 for controlling the operation and / or activities of the processing unit 202, and one or more arithmetic logic units (ALUs) 218 for performing mathematical, arithmetic, and / or logical operations such as addition, subtraction, multiplication, division, shift, logical AND, logical OR, and / or NOT, etc., and may further execute instructions involving mathematical calculations. The ALU 218 may be adapted to operate according to one or more numerical formats such as integers, floating-point numbers, fixed-point numbers, etc.

[0080] The processing unit 202 may include a memory control unit 220 for controlling, managing, and / or processing access to a memory, such as loading, reading, and / or writing data and / or instructions, i.e., program instructions. The memory control unit 220 may be adapted to control one or more memory channels and may include one or more memory control modules known in the art, such as a memory management unit (MMU), etc.

[0081] Furthermore, the processing unit 202 may include one or more caches 222 including an L1 cache having one or more high-speed, low-latency devices for storing (caching) data and / or instructions loaded from a next-level memory 230, thereby making it accessible to the low-latency processing unit 202. The cache 222 is extremely fast, but since its capacity may be limited, it can store only a limited segment of data and / or instructions.

[0082] The next-level memory 230, which can also be controlled by the memory control unit 220, may include, for example, one or more next-level cache units, such as an L2 cache and / or an L3 cache, etc. These are usually slower than the cache 222 but may have a larger capacity. In another example, the next-level memory 230 may include one or more standard memory slow memory arrays (such as system memory and / or application memory, etc.) that can be utilized by one or more volatile memory devices (such as RAM) and / or non-volatile memory devices (such as flash, etc.).

[0083] Note that the processing unit 202 is shown only in an exemplary form. As such, the processing unit 202 may include one or more additional units, circuits, and / or modules, etc. not shown, as known in the art, and / or may lack one or more elements shown in the processing unit 202.

[0084] The processing unit 202 may adopt one or more processor architectures, structures, and / or instruction sets, etc. that support one or more bit widths, such as 32 bits, 56 bits, and / or 64 bits, etc. For example, the processing unit 202 may optionally have a von Neumann architecture, such as a central processing unit (CPU), a multi-core CPU, a data processing unit (DPU), a microcontroller unit (MCU), and / or an accelerated processing unit (ACU), etc. In another example, the processing unit 202 may optionally have a non-von Neumann architecture, such as a graphics processing unit (GPU), a DPU, a field programmable gate array (FPGA), a coarse-grained reconfigurable architecture (CGRA), a neural network accelerator, an intelligence processing unit (IPU), an application-specific integrated circuit (ASIC), a quantum computer, and / or an interconnected computing grid, etc.

[0085] The processing unit 202 may be implemented, structured, and / or deployed according to one or more designs, architectures, and / or implementations. For example, the processing unit 202 may be implemented as a stand-alone device, system, and / or device, etc. In another example, the processing unit 202 may be integrated into one or more higher-level integrated devices. For example, the processing unit 202 may be implemented as an integrated circuit (IC), ASIC, and / or FPGA, etc. in one or more devices with additional elements, such as a computer, server, and / or computing device, etc. In another example, the processing unit 202 may be integrated into one or more higher-level integrated circuits. For example, the processing unit 202 may be implemented as a functional module (such as an IP core, etc.) embedded in one or more integrated components with one or more additional functional elements of the integrated components (such as IC, ASIC, FPGA, CPU, GPU, etc.).

[0086] As shown in FIG. 2B, the exemplary joint scheduler 200 may be adapted to dispatch a plurality of prefetch accesses and data request accesses, also called demand accesses, to access data related to a plurality of instructions loaded into the execution pipeline 210 of the processing unit 202.

[0087] Accordingly, the joint scheduler 200 may include a plurality of instruction entries that each hold a respective control payload for each instruction loaded into the execution pipeline 210 from the time it is assigned until the assignment is released (completed). This control payload is used to determine which of the operations are scheduled.

[0088] Each of the instruction entries may further associate each instruction loaded into the execution pipeline with a respective valid indication and a pointer field including a pointer to a respective cache entry.

[0089] A scheduler entry holds a control payload for each operation from the time it is assigned until the assignment is released (completed). This control payload is used to determine which of the operations are scheduled.

[0090] The valid indication associated with each instruction in each instruction entry of the joint scheduler 200 indicates that the data for each instruction is cached (loaded) in the cache 222, and the pointer field associated with each instruction points to the cache entry in the cache 222 that stores the data for each instruction.

[0091] The valid indication can be utilized using one or more methods, techniques, and / or implementations. For example, each instruction entry may be one or more respective bits (i.e., flags) that are associated, correlated, and / or assigned, and when the bit is set or cleared, it indicates that the data for the instruction of each instruction entry is cached in the cache 222. In another example, the valid indication can be implemented via a pointer field. For example, a pointer field that includes a valid pointer to a valid cache entry indicates that the data for each instruction is cached in the cache 222, while a value and / or pattern that constitutes an invalid pointer that does not point to a valid cache entry (e.g., 0xFFFFFFFF) may indicate that the data for each instruction is not cached in the cache 222.

[0092] For example, the first instruction entry (0) may associate the first instruction (0) with respective valid indication V(0) indicating that data regarding the instruction (0) is cached in cache 222 in the cache entry pointed to by pointer (0) also associated with the instruction (0). In another example, the second instruction entry (1) may associate the second instruction (1) with respective valid indication V(1) indicating that data regarding the instruction (1) is cached in cache 222 in the cache entry pointed to by pointer (1) also associated with the instruction (1). This may be repeated for the Nth instruction entry (N) that associates the Nth instruction (N) with respective valid indication V(N) indicating that data regarding the instruction (N) is cached in cache 222 in the cache entry pointed to by pointer (N) also associated with the instruction (N).

[0093] As is known in the art, a prefetch access is an access (cycle) that starts by accessing an address that maps the requested data to memory in order to load data regarding an instruction whose execution is scheduled by processing unit 202 from the memory. During a prefetch access, a lookup may first be performed to search for the address of the requested data in cache 222. In particular, prior to the execution of an instruction, it is first checked whether the requested data is loaded in one or more caches 222 of processing unit 202, and in case of a miss, i.e., if the requested data is not loaded in cache 222, a prefetch access is started to load the data from the next level of memory into cache 222, i.e., to cache the data.

[0094] A requested access is an access (cycle) that is initiated to load data regarding an instruction currently being executed by the processing unit 202 from the memory. A requested access is typically preceded by a corresponding prefetch access to load the same data from the same address, so the requested data can typically be cached, i.e., stored, in the cache 222.

[0095] The joint scheduler 200 may include one or more prefetch ports for dispatching prefetch accesses and one or more request ports for dispatching requested accesses, and the one or more prefetch ports and the one or more request ports are separate from and independent of each other. As such, the joint scheduler 200 can typically dispatch prefetch accesses and requested accesses regarding different instructions loaded into the execution pipeline 210 simultaneously. Also, if the joint scheduler 200 includes a plurality of prefetch ports, the joint scheduler 200 can dispatch a plurality of prefetch accesses regarding a plurality of different instructions simultaneously. Similarly, if the joint scheduler 200 includes a plurality of request ports, the joint scheduler 200 can dispatch a plurality of requested accesses regarding a plurality of different instructions simultaneously.

[0096] For simplicity, the joint scheduler 200 is described in process 100 as dispatching a single prefetch access and a single request access for a single instruction loaded into the execution pipeline 210. However, the joint scheduler 200 may repeat, extend, and / or scale process 100 to dispatch multiple prefetch accesses and multiple request accesses for multiple instructions loaded into the execution pipeline 210, and thus should not be construed as limited. Further, as described above herein, the joint scheduler 200 may include multiple prefetch ports and / or multiple request ports, and thus may dispatch multiple prefetch accesses and / or request accesses simultaneously.

[0097] Also, for simplicity, the processing unit 202 may include multiple caches 222, but process 100 is described with respect to a single cache 222. However, this should not be construed as limited since the same methodology applied by the joint scheduler 200 is typically equally applicable to and provided for multiple caches 222 that are typically controlled by the memory control unit 220 transparently to the joint scheduler 200.

[0098] As shown in 102, process 100 begins with the joint scheduler 200 traversing one or more instructions scheduled to be loaded into and executed by the execution pipeline 210 of the processing unit 202 and extracting the addresses of data for one or more of the loaded instructions.

[0099] As shown in 104, the joint scheduler 200 may dispatch a prefetch access for each address of each data for each instruction loaded into the execution pipeline 210.

[0100] As shown in 106, in the case of a hit, that is, when each piece of data is cached, that is, loaded and stored in cache 222, process 100 may branch to 108. However, in the case of a miss, that is, when each piece of data is not cached in cache 222, process 100 may branch to 110.

[0101] As shown in 108, in response to the hit prefetch access dispatched for each piece of data, joint scheduler 200 may associate each instruction with a valid instruction and a pointer to each cache entry in cache 222 that stores each piece of data.

[0102] As shown in 110 - 118, in response to the missed prefetch access dispatched for each piece of data, joint scheduler 200 may start a read cycle to load each piece of data from the next level of memory and may cache each piece of data in cache 222.

[0103] As shown in 110, before loading each piece of data from the next level of memory and caching each piece of data in cache 222, joint scheduler 200 may prevent further access to any prefetch access and / or request access regarding each instruction until each piece of data is cached in cache 222 and becomes available.

[0104] For this purpose, joint scheduler 200 may associate each instruction with a block mark indicating that each piece of data regarding each instruction is currently cached and that each prefetch access and / or request access regarding each instruction is blocked from being dispatched until each piece of data is fetched from the next level of memory.

[0105] As shown in 112, the joint scheduler 200 may initiate a read cycle for the next level of memory to fetch each piece of data and store (cache) each piece of data in the cache 222.

[0106] As shown in 114, in response to the successful completion of the read cycle initiated to load each piece of data from the next level of memory and cache each piece of data in the cache 222, the joint scheduler 200 may associate each instruction with a valid instruction and a pointer to each cache entry in the cache 222 that stores the cached each piece of data. This operation is similar to the operation of step 108 executed by the joint scheduler 200 in response to a hit prefetch access.

[0107] Alternatively, as shown in 116, instead of immediately associating each instruction with a valid instruction and a pointer to each cache entry that stores the cached each piece of data, the joint scheduler 200 may branch back to 104 and dispatch (initiate) another prefetch access for each address of each piece of data. Since each piece of data is cached in the cache 222 here, another prefetch access hits at 106, branches to 108, and at 108, the joint scheduler 200 may associate each instruction with a valid instruction and a pointer to each cache entry in the cache 222 that stores the cached each piece of data.

[0108] As shown in 118, in response to the successful completion of the prefetch cycle, the joint scheduler 220 may remove the block mark associated with each instruction, enabling one or more other accesses, typically request accesses, to each instruction for which each piece of data has been loaded and cached in the cache 222.

[0109] As shown in 120, the joint scheduler 220 may dispatch the requested accesses for each instruction.

[0110] In particular, the joint scheduler 220 may check the valid indication associated with each instruction and, accordingly, determine whether each data for each instruction is cached in the cache 222. In response to determining that each data is cached, the joint scheduler 220 may dispatch the requested access using the pointer associated with each instruction and access each data in the indicated cache entry in the cache 222, such as for reading, writing, and / or updating.

[0111] This means that the requested access can be executed without an additional lookup to check whether each data is cached, for example, by traversing the address cache, i.e., the tag / index field of the cache 222, as can be done by existing methods.

[0112] Directly accessing the cache entry mapped by the pointer during the requested access and without an additional lookup can significantly reduce the latency for accessing the data for the currently executing instruction, and thus can significantly improve the performance of the processing unit 202. Further, by using the pointer to access each data in the requested access, the saved time can be used to dispatch and execute significantly more prefetch accesses and / or requested accesses, so that the bandwidth of the data access can be significantly increased, and thus the performance of the processing unit 202 can be further improved.

[0113] Optionally, if one or more prefetch accesses and one or more request accesses to a common cache entry are simultaneously dispatched via respective prefetch ports and request ports, in response to completion of the prefetch access, joint scheduler 200 may directly update, at the request port, pointers to the common cache entry for one or more of the request accesses.

[0114] This means that rather than blocking the request access until each respective data is cached while the corresponding prefetch access is in progress, the request accesses are already in the pipeline, waiting for their turn to be dispatched, and after the prefetch access has completed, joint scheduler 200 can dispatch the request accesses and update their associated pointers.

[0115] Specifically, joint scheduler 200 may identify the cache entry cached by the simultaneous prefetch access, directly update the pointer associated with the corresponding request access, and map the cache entry that caches the data fetched by the simultaneous prefetch access, thereby bypassing the standard mechanism employed for non - simultaneous accesses.

[0116] Optionally, joint scheduler 200 may be adapted to track the invalidation and / or eviction of data stored in each of a plurality of cache entries of cache 222 and, in response thereto, update the valid indication associated with each respective instruction. This means that joint scheduler 200 can monitor the cache entry mapped (indicated) by the pointer associated with each instruction in pipeline 210 and further associated with the valid indication to detect the invalidation and / or eviction of the data cached in the cache entry.

[0117] In response to the eviction of each cache entry that stores respective data for each one of the instructions loaded into execution pipeline 210, joint scheduler 220 may associate each instruction with an invalidation indication indicating that the respective data for the respective instruction is no longer cached within cache 222.

[0118] Joint scheduler 200 may further initiate another prefetch cycle and load respective data from the next level of memory into cache 222. Following another prefetch access that misses, steps 110 - 118 may be repeated, and the valid indication and pointer associated with the instruction may be updated again, indicating that respective data is cached in the cache entry pointed to by the associated pointer.

[0119] Optionally, joint scheduler 220 may be adapted to mark an active indication in each cache entry after each prefetch access hit. An active indication indicating that each cache entry is mapped by a pointer associated with one or more of the instructions loaded into execution pipeline 210 may be used by joint scheduler circuit 220 to track (snoop) the invalidation and / or eviction of data stored in the cache entry marked with the active indication.

[0120] By tracking the invalidation and / or eviction of data stored within cache 222, joint scheduler 220 may significantly enhance the out - of - order execution of one or more of the instructions loaded into pipeline 210.

[0121] To achieve this, the joint scheduler 220 executes steps 110-118 of process 100 in response to detecting invalidation and / or eviction, and dispatches another prefetch cycle to load from the next level of memory 230 the respective invalidated and / or evicted data for one or more of the instructions loaded into pipeline 210, thereby making the respective data available in cache 222 during the requested access dispatched for each instruction.

[0122] As such, data for practically any instruction loaded into pipeline 210 may be available in cache 222 during the requested access for any of the loaded instructions, regardless of the position (order) within pipeline 210, so that the associated data is available in cache 222 during the requested access even when the instructions are executed out of order.

[0123] Refer now to FIGS. 3A and 3B, which are schematic diagrams of an exemplary scheduler of a processing unit configured to schedule memory accesses, particularly memory accesses as known in the art.

[0124] As shown in FIG. 3A, a single dispatch scheduler 310, such as known in the art, may be adapted to dispatch multiple memory accesses to access data for one or more instructions loaded into an execution pipeline, such as execution pipeline 210.

[0125] In particular, single dispatch scheduler 310 may dispatch multiple memory accesses, whether the multiple memory accesses are prefetch accesses or requested accesses, if the multiple memory accesses are not blocked, i.e., are ready for dispatch.

[0126] Each new instruction assigned to the single dispatch scheduler 310 may have its own blocking condition / ready condition (some may be common). Instructions that are "ready to be dispatched" to one or more ports can be scheduled for dispatch.

[0127] Accordingly, the single dispatch scheduler 310 may include ready logic 312 implemented using one or more circuits, components, and / or elements adapted to determine whether each memory access is ready or blocked.

[0128] The single dispatch scheduler 310 may also include an access scheduler 314 that employs one or more scheduling algorithms, such as age-based scheduling, location-based scheduling, etc., to dispatch one or more memory accesses that are ready for dispatch because they do not comply with any blocking conditions via one or more dispatch ports.

[0129] As shown in FIG. 3B, the ready logic 312 may be implemented by monitoring the blocking condition for each memory access individually, regardless of whether it is a prefetch access or a request access, and may generate a ready indication (e.g., a flag) accordingly.

[0130] Now refer to FIGS. 4A and 4B, which are schematic diagrams of an exemplary joint scheduler of a processing unit configured to efficiently schedule memory accesses according to some embodiments of the present invention.

[0131] As shown in FIG. 4A, a joint (dual) dispatch scheduler 410, such as joint scheduler 200, may be adapted to dispatch one or more prefetch accesses and one or more request accesses to data related to one or more instructions loaded into an execution pipeline, such as execution pipeline 210.

[0132] In particular, dual dispatch scheduler 410 may dispatch one or more prefetch accesses and one or more request accesses, each scheduled via one or more respective dispatch ports. That is, prefetch accesses and request accesses are processed differently from each other.

[0133] Dual dispatch scheduler 410 may include ready logic 412 implemented using one or more circuits, components, and / or elements adapted to individually determine block / ready conditions for prefetch accesses and request accesses for each instruction (entry) loaded into execution pipeline 210 and assigned for scheduling by dual dispatch scheduler 410.

[0134] This means that ready logic 412 may correlate between prefetch accesses and request accesses for each instruction, but may individually determine block / ready conditions for each access.

[0135] Dual dispatch scheduler 410 may include a first scheduler, i.e., prefetch access scheduler 414, that employs one or more scheduling algorithms, such as age-based scheduling, location-based scheduling, etc., to dispatch one or more prefetch accesses that are ready for dispatch (ready) because they do not comply with any blocking conditions, via one or more dispatch ports.

[0136] The dual dispatch scheduler 410 may further include a second scheduler, i.e., a request access scheduler 416, which employs one or more scheduling algorithms, such as age-based scheduling, location-based scheduling, etc., to dispatch one or more request accesses that are ready for dispatch without following any blocking conditions via one or more dispatch ports.

[0137] As shown in FIG. 4B, the ready logic 412 may monitor the blocking conditions separately for the prefetch access and the request access regarding the same instruction, and generate respective ready indications (e.g., flags) accordingly. As such, in response to determining that a particular prefetch access regarding a particular instruction is ready for dispatch, the ready logic 412 may issue a prefetch ready indication. Similarly, in response to determining that a particular request access regarding a particular instruction is ready for dispatch, the ready logic 412 may issue a demand ready indication.

[0138] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or to limit the invention to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application, or technical improvements found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

[0139] During the term of the patent that may issue from this application, many related systems, methods, and computer programs are expected to be developed, and the scope of terms such as "processor architecture", "cache architecture", and "scheduling algorithm" is intended to include a priori all such new technologies.

[0140] As used herein, the term "about" refers to ±10%.

[0141] The terms "comprises", "comprising", "includes", "including", "having" and their conjugations mean "including but not limited to". This term encompasses the terms "consisting of" and "consisting essentially of".

[0142] The expression "consisting essentially of" means that a composition or method may include additional ingredients and / or steps, but only to the extent that the additional ingredients and / or steps do not substantially change the basic and novel characteristics of the composition or method recited in the claims.

[0143] As used herein, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" may include a plurality of compounds including mixtures thereof.

[0144] As used herein, the term "exemplary" is used to mean "serving as an example, instance, or illustration." Any embodiment described as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments, and / or does not exclude the incorporation of features from other embodiments.

[0145] As used herein, the term "optionally" is used to mean "provided in some embodiments and not provided in other embodiments." Any particular embodiment of the invention may include a plurality of "optional" features as long as these features do not conflict with each other.

[0146] Throughout this application, various embodiments of the invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an immutable limitation on the scope of the invention. Thus, the description of a range should be considered to specifically disclose all the possible sub-ranges within that range as well as individual numerical values. For example, a description of a range such as from 1 to 6 should be considered to specifically disclose sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., within that range, as well as individual numbers, e.g., 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0147] When a numerical range is recited herein, it is intended to include any recited number (fractional or integral) within the recited range. The phrases "ranging / ranges between" the first recited number and the second recited number and "ranging / ranges" from the first recited number "to" the second recited number are used interchangeably herein to mean including the first recited number, the second recited number, and all fractional and integral numbers therebetween.

[0148] For clarity, it is understood that certain features of the invention described in the context of separate embodiments may be provided in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment may be provided separately, or in any suitable subcombination, or as suitable in other described embodiments of the invention. Specific features described in the context of various embodiments should not be considered essential features of those embodiments, except where the embodiments would not function without those elements.

[0149] Although the invention has been described in connection with its specific embodiments, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.

[0150] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference in their entirety as if each individual publication, patent, or patent application were specifically and individually noted as being incorporated by reference, when it is shown that the individual publication, patent, or patent application is incorporated by reference herein. Also, any citation or identification of a reference in this application should not be construed as an admission that such reference is available as prior art to the present invention. To the extent that section headings are used, they should not necessarily be construed as limiting. Additionally, if there are priority documents for this application, the entireties of those are incorporated by reference herein.

Claims

1. a joint scheduler circuit adapted to dispatch prefetch accesses and request accesses of data for a plurality of instructions loaded into an execution pipeline of at least one processing circuit, each prefetch access including checking whether the respective data is cached in one of a plurality of cache entries of at least one cache, and each request access including accessing the respective data; A joint scheduler comprising: The joint scheduler circuit includes: in response to each prefetch access dispatched to a respective datum for a respective one of the plurality of instructions that hits, the requested access for the respective instruction accesses the respective datum in the at least one cache using a pointer to a respective cache entry that stores the respective datum, the respective instruction being associated with the respective instruction; and in response to each prefetch access dispatched for a respective data item associated with a respective one of said plurality of instructions that misses, initiating a read cycle to load said respective data item from a next level memory and caching said respective data item in said at least one cache; To be adapted, Joint scheduler.

2. the joint scheduler circuit is adapted to load respective data from the next level memory after a respective prefetch access that misses, and to associate, in response to successful completion of a respective read cycle initiated to cache the respective data in the at least one cache, the respective instruction with the valid indication and the pointer to the respective cache entry that stores the respective cached data; The joint scheduler of claim 1 .

3. the joint scheduler circuit is adapted to load the respective data from the next level memory after each prefetch access that misses, and to dispatch another prefetch access in response to successful completion of a respective read cycle initiated to cache the respective data in the at least one cache, and to update the valid indication and the pointer to the respective cache entry storing the respective cached data. The joint scheduler of claim 1 .

4. the joint scheduler circuitry is further adapted to track invalidation and / or eviction of data stored in each of the plurality of cache entries, and, in response to eviction of a respective cache entry storing respective data for a respective one of the plurality of instructions, to associate the respective instruction with an invalidation indication and to initiate another prefetch cycle to load the respective data from the next level memory into the at least one cache. The joint scheduler of claim 1 .

5. the joint scheduler circuitry is further adapted to, for each prefetch access hit, mark a respective cache entry with an active indication, the active indication indicating that the respective cache entry is mapped by a pointer associated with at least one of the instructions, the active indication being used by the joint scheduler circuitry to track invalidation and / or eviction of data stored in the respective cache entry. The joint scheduler of claim 1 .

6. the joint scheduler circuitry is adapted to, for each prefetch access miss, associate with the respective instruction a block mark indicating that each prefetch access and / or each requested access related to the respective instruction is blocked for dispatch; The joint scheduler of claim 1 .

7. the joint scheduler circuitry is adapted to remove the block mark associated with the respective instruction in response to a prefetch cycle completing successfully. The joint scheduler of claim 6.

8. the joint scheduler circuit comprises at least one prefetch port for dispatching the prefetch accesses and at least one request port for dispatching the requested accesses, the at least one prefetch port being separate and independent from the one or more request ports; The joint scheduler of claim 1 .

9. at least one prefetch access and at least one request access are dispatched simultaneously via independent said at least one prefetch port and said at least one request port, respectively; The joint scheduler of claim 8.

10. For at least one simultaneously dispatched prefetch access and at least one request access related to a common cache entry, a pointer to the common cache entry is directly updated for the at least one request access in response to completion of the at least one prefetch access at the at least one request port. The joint scheduler of claim 9.

11. using a joint scheduler circuit adapted to dispatch prefetch accesses and request accesses of data for a plurality of instructions loaded into an execution pipeline of at least one processing circuit, each prefetch access including checking whether the respective data is cached in one of a plurality of cache entries of at least one cache, and each request access including accessing the respective data; 1. A method for jointly scheduling prefetch and request accesses, comprising: The joint scheduler circuit includes: in response to each prefetch access dispatched to a respective datum for a respective one of the plurality of instructions that hits, the requested access for the respective instruction accesses the respective datum in the at least one cache using a pointer to a respective cache entry that stores the respective datum, the respective instruction being associated with the respective instruction; and in response to each prefetch access dispatched for a respective data item associated with a respective one of said plurality of instructions that misses, initiating a read cycle to load said respective data item from a next level memory and caching said respective data item in said at least one cache; To be adapted, method.

12. the joint scheduler circuit is adapted to load respective data from the next level memory after a respective prefetch access that misses, and to associate, in response to successful completion of a respective read cycle initiated to cache the respective data in the at least one cache, the respective instruction with the valid indication and the pointer to the respective cache entry that stores the respective cached data; The method of claim 11.

13. the joint scheduler circuit is adapted to load the respective data from the next level memory after each prefetch access that misses, and to invoke another prefetch access in response to successful completion of a respective read cycle initiated to cache the respective data in the at least one cache, and to update the valid indication and the pointer to the respective cache entry storing the respective cached data. The method of claim 11.

14. the joint scheduler circuitry is further adapted to track invalidation and / or eviction of data stored in each of the plurality of cache entries, and, in response to eviction of a respective cache entry storing respective data for a respective one of the plurality of instructions, to associate the respective instruction with an invalidation indication and to initiate another prefetch cycle to load the respective data from the next level memory into the at least one cache. The method of claim 11.

15. the joint scheduler circuitry is further adapted to, for each prefetch access hit, mark a respective cache entry with an active indication, the active indication indicating that the respective cache entry is mapped by a pointer associated with at least one of the instructions, the active indication being used by the joint scheduler circuitry to track invalidation and / or eviction of data stored in the respective cache entry. The method of claim 11.

16. the joint scheduler circuitry is adapted to, for each prefetch access miss, associate with the respective instruction a block mark indicating that each prefetch access and / or each requested access related to the respective instruction is blocked for dispatch; The method of claim 11.

17. the joint scheduler circuitry is adapted to remove the block mark associated with the respective instruction in response to a prefetch cycle completing successfully.

17. The method of claim 16.

18. out-of-order execution of a plurality of instructions loaded into an execution pipeline of at least one processing circuit; tracking invalidation and / or eviction of data stored in each of a plurality of cache entries of at least one cache that stores data relating to the plurality of instructions; dispatching a prefetch cycle to load the invalidated and / or evicted data for at least one of the plurality of instructions, thereby making the respective data available in the at least one cache during a requested access dispatched for the at least one instruction; a joint scheduler circuit adapted to augment the Joint scheduler.

19. The joint scheduler, in response to evicting and / or invalidating at least one cache entry that stores respective data relating to a respective one of the plurality of instructions, initiating said prefetch cycle to load said respective data from a next level memory into said at least one cache; associating the respective instructions with a validity indication and a pointer such that the requested access for the respective instructions accesses the respective datum in the at least one cache using a pointer to a respective cache entry storing the respective datum, the pointer being associated with the respective instruction; adapted to dispatch the prefetch cycle by 20. The joint scheduler of claim 18.

20. the joint scheduler circuitry is further adapted to associate the respective instruction with the valid indication and the pointer by initiating a separate prefetch cycle for the respective instruction; the joint scheduler circuitry is adapted to associate the respective instruction with the valid indication and the pointer in response to a hit of the other prefetch cycle; 20. The joint scheduler of claim 19.