Joint scheduler for high bandwidth multi-sample prefetching

By using a joint scheduler at the processing unit, the limitations of traditional cache management technology in terms of responsiveness and adaptability are solved, more efficient memory access is achieved, reducing latency and increasing bandwidth, and significantly improving the execution performance of the processing unit.

CN120144179APending Publication Date: 2025-06-13NEXTSILICON LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411825237.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-12
Filing Date
2024-12-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional cache management technology has limitations in responsiveness and adaptability, which affects data retrieval performance.

Method used

The combined scheduler is adopted to improve the memory access efficiency of the processing unit by sending prefetch access and demand access. The scheduler includes a joint scheduler circuit for scheduling memory access at the processing unit, increasing efficiency, reducing latency and increasing bandwidth.

Benefits of technology

By directly accessing the data in the cache, the memory access latency and bandwidth are reduced, and the execution performance of the processing unit is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144179A_ABST
    Figure CN120144179A_ABST
Patent Text Reader

Abstract

The invention provides a joint scheduler for high bandwidth multi-sample prefetching. The joint scheduler is adapted to send a prefetch access and a demand access of data related to a plurality of instructions loaded in an execution pipeline of processing circuitry. Each prefetch access includes checking whether a respective data is cached in a cache entry, and each demand access includes accessing the respective data. The joint scheduler is adapted to, in response to each hit prefetch access sent for a respective data associated with a respective instruction, associate the respective instruction with a valid indication and a pointer to a respective cache entry storing the respective data, causing the demanded access associated with the respective instruction to access the respective data in the cache using the associated pointer; and in response to each missed prefetch access sent for the respective data associated with the respective instruction, initiating a read cycle to load the respective data from a next level memory and cache the respective data in the cache.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 537,927, filed on December 13, 2023, the entire content of which is incorporated herein by reference.

[0003] This application is also related to U.S. Patent Application No. 18 / 770,690, filed on July 12, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0004] In some embodiments, the present invention relates to scheduling memory accesses at a processing unit, and more specifically but not exclusively to using a combined scheduler to schedule memory accesses at a processing unit with improved efficiency, reduced latency, and / or increased bandwidth. Background Art

[0005] In modern computing systems, the performance of a processor depends to a large extent on the effective utilization of cache memory.

[0006] A cache employing a high-speed memory device can be deployed in close proximity to a processing unit, typically within the processor itself, to store frequently accessed data and / or instructions, thereby reducing the latency associated with retrieving information from the main memory, which is significantly slower and has a longer access time.

[0007] However, traditional cache management techniques may exhibit inherent limitations in terms of responsiveness and / or adaptability to different workloads and data access patterns, which can directly affect data retrieval performance. Summary of the Invention

[0008] The object of the present invention is to provide methods, systems, and software program products for improving the efficiency of a processing unit scheduler to reduce latency and increase execution bandwidth. The above object and other objects are achieved by the features of the independent claims. Further embodiments can be seen from the dependent claims, the description, and the drawings.

[0009] According to a first aspect of the present invention, there is provided a combined scheduler comprising: a combined scheduler circuit adapted to send prefetch accesses and demand accesses for data associated with a plurality of instructions, the plurality of instructions being loaded in an execution pipeline of one or more processing circuits. Each prefetch access includes checking whether a corresponding data is cached in one of a plurality of cache entries of one or more caches, and each demand access includes accessing the corresponding data. The combined scheduler circuit is adapted to: (1) in response to each hit prefetch access sent for a corresponding data associated with a corresponding instruction of the plurality of instructions, associate the corresponding instruction with a valid indication and a pointer pointing to a corresponding cache entry that stores the corresponding data, such that the demand access associated with the corresponding instruction uses the associated pointer to access the corresponding data in the one or more caches; and (2) in response to each miss prefetch access sent for the corresponding data associated with the corresponding instruction of the plurality of instructions, initiate a read cycle to load the corresponding data from a next-level memory and cache the corresponding data in the one or more caches.

[0010] According to a second aspect of the present invention, there is provided a method for combined scheduling of prefetch accesses and demand accesses, comprising: using a combined scheduler circuit adapted to send prefetch accesses and demand accesses for data associated with a plurality of instructions, the plurality of instructions being loaded in an execution pipeline of one or more processing circuits, each prefetch access including checking whether a corresponding data is cached in one of a plurality of cache entries of one or more caches, and each demand access including accessing the corresponding data. The combined scheduler circuit is adapted to: (1) in response to each hit prefetch access sent for a corresponding data associated with a corresponding instruction of the plurality of instructions, associate the corresponding instruction with a valid indication and a pointer pointing to a corresponding cache entry that stores the corresponding data, such that the demand access associated with the corresponding instruction uses the associated pointer to access the corresponding data in the one or more caches; and (2) in response to each miss prefetch access sent for the corresponding data associated with the corresponding instruction of the plurality of instructions, initiate a read cycle to load the corresponding data from a next-level memory and cache the corresponding data in the one or more caches.

[0011] According to a third aspect of the present invention, there is provided a combined scheduler, comprising: a combined scheduler circuit adapted to enhance the out-of-order execution of a plurality of instructions loaded in an execution pipeline of one or more processing circuits by: tracking the invalidation and / or eviction of data stored in each of a plurality of cache entries of one or more caches, the at least one cache storing data related to the plurality of instructions; and sending a prefetch cycle to load invalidated data and / or evicted data related to one or more of the plurality of instructions, so that the corresponding data is available in the one or more caches during a demand access sent for one or more instructions.

[0012] In a further implementation of the first aspect, the second aspect, and / or the third aspect, the combined scheduler circuit is adapted to: in response to the successful completion of a corresponding read cycle, the corresponding read cycle initiating to load the corresponding data from the next-level memory and caching the corresponding data in the one or more caches after a corresponding miss prefetch access, associate the corresponding instruction with the valid indication and the pointer pointing to the corresponding cache entry, the corresponding cache entry storing the cached corresponding data.

[0013] In a further implementation of the first aspect, the second aspect, and / or the third aspect, the combined scheduler circuit is adapted to: in response to the successful completion of a corresponding read cycle, the corresponding read cycle initiating to load the corresponding data from the next-level memory and caching the corresponding data in the one or more caches after a corresponding miss prefetch access, send another prefetch access to update the valid indication and the pointer pointing to the corresponding cache entry, the corresponding cache entry storing the cached corresponding data.

[0014] In an alternative implementation of the first aspect, the second aspect, and / or the third aspect, the combined scheduler circuit is further adapted to: track the invalidation and / or eviction of data stored in each of the plurality of cache entries, and in response to the eviction of the corresponding cache entry, the corresponding cache entry storing the corresponding data related to the corresponding instruction of the plurality of instructions, associate the corresponding instruction with an invalid indication, and initiate another prefetch cycle to load the corresponding data from the next-level memory into the one or more caches.

[0015] In alternative embodiments of the first, second, and / or third aspects, the combined scheduler circuit is further adapted to: for each hit prefetch access, mark the corresponding cache entry with an activity indication that indicates that the corresponding cache entry is mapped by a pointer associated with one or more of the plurality of instructions, and the combined scheduler uses the activity indication to track invalidation and / or eviction of data stored in the corresponding cache entry.

[0016] In further embodiments of the first, second, and / or third aspects, the combined scheduler circuit is adapted to: for each miss prefetch access, associate the corresponding instruction with a block marker that indicates that each prefetch access and / or demand access associated with the corresponding instruction is blocked from being sent.

[0017] In further embodiments of the first, second, and / or third aspects, the combined scheduler circuit is adapted to: in response to successful completion of the prefetch cycle, delete the block marker associated with the corresponding instruction.

[0018] In further embodiments of the first, second, and / or third aspects, the combined scheduler circuit includes one or more prefetch ports for sending the prefetch accesses and one or more demand ports for sending the demand accesses, and the one or more prefetch ports are separate and independent from the one or more demand ports.

[0019] In further embodiments of the first, second, and / or third aspects, the one or more prefetch accesses and the one or more demand accesses are sent simultaneously by the respective one or more prefetch ports and the one or more demand ports.

[0020] In further embodiments of the first, second, and / or third aspects, for the one or more prefetch accesses and the one or more demand accesses associated with a common cache entry that are sent simultaneously, in response to completion of the one or more prefetch accesses, a pointer to the common cache entry is directly updated in the one or more demand ports for the one or more demand accesses.

[0021] In a further implementation form of the first aspect, the second aspect, and / or the third aspect, the combined scheduler is adapted to: send the prefetch cycle in response to the eviction and / or invalidation of one or more cache entries storing the corresponding data, the corresponding data being related to a corresponding instruction of the plurality of instructions, by: initiating the prefetch cycle to load the corresponding data from a lower-level memory into the one or more caches; and associating the corresponding instruction with a valid indication and a pointer to a corresponding cache entry that stores the corresponding data, such that a demand access related to the corresponding instruction uses the associated pointer to access the corresponding data in the one or more caches.

[0022] In an alternative implementation form of the first aspect, the second aspect, and / or the third aspect, the combined scheduler circuit is further adapted to: associate the corresponding instruction with the valid indication and the pointer by initiating another prefetch cycle. Wherein the combined scheduler circuit is adapted to: in response to a hit in the another prefetch cycle, associate the corresponding instruction with the valid indication and the pointer.

[0023] Those skilled in the art will become or will have become clear about other systems, methods, features, and advantages of the present disclosure after reading the following drawings and detailed description. All these additional systems, methods, features, and advantages should be included in this description, fall within the scope of the present disclosure, and be protected by the accompanying claims.

[0024] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification (including definitions) shall prevail. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.

[0025] The implementation of the method and / or system of the embodiments of the present invention may involve automatically performing or completing selected tasks. In addition, for the actual instruments and devices according to the embodiments of the present invention, an operating system can be used to implement several selected tasks through hardware, software, firmware, or a combination thereof.

[0026] For example, the hardware for performing a selected task according to an embodiment of the present invention can be implemented as a chip or a circuit. As software, the selected task according to an embodiment of the present invention can be implemented as a plurality of software instructions executed by a computer using any suitable operating system. In an exemplary embodiment of the present invention, one or more tasks according to an exemplary embodiment of the method and / or system described herein are executed by a data processor (e.g., a computing platform for executing a plurality of instructions). Optionally, the data processor includes a volatile memory for storing instructions and / or data and / or a non-volatile memory for storing instructions and / or data (e.g., a magnetic hard disk and / or a removable medium). Optionally, a network connection is also provided. Optionally, a display and / or a user input device (e.g., a keyboard or a mouse) are also provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Some embodiments of the present invention are described herein only by way of example and with reference to the accompanying drawings. Now, with specific reference to the accompanying drawings for detailed elaboration, it should be emphasized that these details are shown by way of example and are used to discuss the embodiments of the present invention illustratively. In this regard, the description in conjunction with the accompanying drawings enables those skilled in the art to clearly understand how to practice the embodiments of the present invention.

[0028] In the drawings:

[0029] Figure 1 is a flowchart of an exemplary process of a joint scheduler using a processing unit configured to efficiently schedule memory access according to some embodiments of the present invention;

[0030] Figure 2A and Figure 2B is a schematic diagram of an exemplary joint scheduler of a processing unit configured to efficiently schedule memory access according to some embodiments of the present invention;

[0031] Figure 3A and Figure 3B is a schematic diagram of an exemplary scheduler of a processing unit configured to schedule memory access; and

[0032] Figure 4A and Figure 4B is a schematic diagram of an exemplary joint scheduler of a processing unit configured to efficiently schedule memory access according to some embodiments of the present invention. DETAILED DESCRIPTION

[0033] In some embodiments, the present invention relates to scheduling memory access at a processing unit, and more specifically but not exclusively to using a joint scheduler to schedule memory access at the processing unit with increased efficiency, reduced latency, and / or increased bandwidth.

[0034] According to some embodiments of the present invention, there are provided an apparatus, a method, and a computer program product for improving memory accesses sent by a processing unit by jointly scheduling prefetch accesses and demand accesses.

[0035] Memory latency and access bandwidth are key factors in the overall performance of a processing unit. To reduce latency and memory access bandwidth, caches are used to hold fragments of the memory image.

[0036] Caches using high-speed memory arrays are typically deployed very close and in a short distance to the key drivers (write operations) or receivers (read operations) of the processing unit, usually within the processing unit itself (L1 cache), to support fast access. However, although very fast, the capacity of caches may be limited, which may restrict them to storing relatively small fragments of the memory image.

[0037] Therefore, the performance of memory access depends on the hit rate of cache lookups. A cache lookup refers to finding a match between the address of a memory access (read or write operation) and one of the data addresses stored in the cache (i.e., the address cache typically includes tag and index fields). A cache hit means that the address of the accessed data is found in the address cache, and thus the accessed data is cached and stored in the cache.

[0038] If a miss occurs, i.e., the data portion of the memory image required for a read or write operation is not cached, the required data portion must be fetched from the next lower-level memory (e.g., main memory) or from the next higher-level memory (e.g., L2 cache, L3 cache, etc.). This may increase the latency of the operation and impair the cache bandwidth, as the next lower-level memory requires a longer access time and the access may be performed multiple times (extra accesses conflict with other cache accesses).

[0039] Due to the limited cache capacity, once a new memory line (the smallest memory portion cached in the cache) is loaded into the cache, an old memory line must be evicted from the cache to make room for the new memory line. This operation is called replacement, as the new memory line will replace the old memory line. Once the evicted memory line is accessed again, it must be reloaded into the cache from a higher memory hierarchy.

[0040] To increase the cache hit rate, prefetching can be applied, which means that prefetch cycles can be initiated (scheduled) in advance to load the data required by one or more instructions into the cache before the "actual" demand access is scheduled during the execution of the corresponding instruction. Thus, once the actual demand access is sent, during the execution of the instruction, it will find that the required data has been loaded and cached in the cache. Specifically, as described above, during a prefetch access, the complete memory line containing the required data is loaded into the cache.

[0041] A popular prefetching algorithm is early cache lookup (also referred to hereinafter as prefetch access), a technique designed to overcome blocking situations where instructions may be blocked and their demand access (read or write operations) may not be sent even though the access address is already known until one or more conditions are met.

[0042] For example, such situations may include that instruction fetch operations are scheduled to read data from the instruction cache, and these operations are typically executed in sequence, so each fetch must wait until the previous fetch has completed before reading its data from the instruction cache. If a particular instruction fetch operation is delayed (e.g., due to a cache miss), subsequent instruction fetch operations must be delayed until the delayed fetch operation is completed, even though their memory addresses are known. In another example, a data store operation that writes data to the data cache is an irreversible operation because the old data is overwritten by the new data and thus cannot be recovered. Therefore, the data store operation can only be executed after all previous instructions have been committed, even though its memory address is known.

[0043] Since the memory address of the demand access (read or write operation) is known, even if the actual request cannot be executed due to a blocking situation, the corresponding prefetch access can be sent in advance to the same address to perform an early lookup in the cache. In this way, the prefetch access can check whether the memory line containing the data related to the instruction exists in the cache (cache) without performing the actual demand access (read or write operation). If there is a cache hit, it means that the requested data has been cached and thus no other operation can be performed. However, if there is a cache miss, an early read operation can be sent to load the memory line from a higher memory level and cache it in the cache, thereby reducing the latency of the actual demand after the actual demand is sent.

[0044] Although modern processing units may be capable of executing multiple instructions in parallel, it is obvious that since only a limited number of memory accesses and cache accesses can be sent in a given cycle, a scheduler is implemented to determine which instructions should be sent in each cycle and schedule the instructions accordingly.

[0045] The scheduler generally includes a ready logic and a scheduling module. The ready logic is configured to calculate which instructions are eligible to be sent. An instruction (an entry in the scheduler) is ready to be sent if it has been allocated, is not being processed in the send pipeline, is eligible to be sent (e.g., the earliest fetch), and is not blocked. For example, an instruction may be blocked if the previous send of the instruction indicates a problem. The block is removed (woken up) only after the source (reason) of the block is resolved. For example, a cache miss will block the send of a given entry until the required memory line is loaded from the next level of memory and cached in the cache.

[0046] The scheduling module may include one or more send ports and may send accesses to one or more ready instruction entries. The scheduling module may employ one or more scheduling mechanisms, such as age-based scheduling (selecting the oldest entry first), location-based scheduling (selecting the first ready entry in the scheduler), etc. Each scheduled instruction is sent, a cache lookup is performed, and the scheduler at the end of the lookup pipeline is updated according to its completion, e.g., successfully when the instruction is released from the scheduler, unsuccessfully in case of a block, etc.

[0047] According to some embodiments of the present invention, there are provided an apparatus, a method, and a computer program product for improving memory accesses sent by a processing unit by jointly scheduling prefetch accesses and demand accesses.

[0048] Specifically, a joint scheduler of a processing unit may be adapted to schedule a memory access unit by sending prefetch accesses and data demand accesses, or to send the same instructions in association with each other, rather than sending prefetch accesses and data demand accesses separately and independently as in existing schedulers.

[0049] The joint scheduler may send a prefetch access to look up in the cache data related to a corresponding instruction loaded in the execution pipeline of the processing unit. In the case of a cache hit, when the requested data is found in the cache, i.e., the data is stored (cached) in one of the multiple cache entries of the cache, the joint scheduler may associate the corresponding instruction with a valid indication (e.g., a flag) and a pointer to the cache entry storing the data.

[0050] In addition, if a cache miss occurs, i.e., the requested data is not found in the cache, a read cycle can be initiated to fetch the data from the next level of memory and load it into the cache. After successfully completing the read cycle, the combined scheduler can associate the corresponding instruction with a valid indication and a pointer to the cache entry that stores the data upon cache hit. Alternatively, after successfully completing the read cycle, the combined scheduler may send another prefetch access, which will also result in a cache hit, thereby associating the corresponding instruction with a valid indication and a pointer to the cache entry that stores the data.

[0051] When executing the corresponding instruction, the combined scheduler can identify the valid indication associated with the corresponding instruction and can use the pointer associated with the corresponding instruction to send a demand access for the corresponding instruction to directly access the cache entry mapped by the associated pointer, thereby eliminating the need for the demand access to search the address cache (tag / index).

[0052] The combined scheduler can also monitor all cache entries pointed to by the pointers associated with the instructions to facilitate detection of whether the cache lines stored in these cache entries are invalid and / or evicted (removed). If a cache entry mapped by a pointer associated with a corresponding instruction is invalid and / or evicted, the combined scheduler can remove the valid indication associated with the corresponding instruction to indicate that the data related to the corresponding instruction is no longer cached, i.e., not stored in the cache.

[0053] In addition, in response to invalidation and / or eviction, the combined scheduler can send another prefetch access to reload the previously invalidated and / or evicted data and make it available again for a subsequent demand access of the corresponding instruction. Since during the another prefetch access, the corresponding instruction can be associated again with a valid indication and a pointer that maps the cache entry storing the reloaded data. Therefore, the demand access for the corresponding instruction sent by the combined scheduler can directly access the reloaded data using the pointer associated with the corresponding instruction, and the reloaded data is stored in the cache entry mapped by the associated pointer.

[0054] Using the combined scheduler to send prefetch accesses and demand accesses related to each other may have greater benefits and advantages than existing schedulers.

[0055] First, during a prefetch access scheduled for each instruction to be executed, searching (looking up) the address cache (tag / index field) of the cache for comparison (lookup) with the data address associated with a corresponding instruction may be done only once because the demand access sent for a corresponding instruction can directly access the cache entry mapped by the pointer associated with the corresponding instruction, thus eliminating the need for another lookup through the tag / index field of the cache. This can significantly reduce memory access and cache access latency and / or increase access bandwidth.

[0056] In addition, a single lookup is performed while the corresponding instruction is still waiting to be executed, and the demand access for the corresponding instruction sent during the execution of the corresponding instruction may not require a lookup. This can significantly speed up instruction execution, thus significantly improving the execution performance of the processing unit in terms of time, speed, etc.

[0057] Some existing schedulers may apply lock prefetching, in which case if an early prefetch (lookup) results in a cache hit, the hit cache line will be locked so that it is not allowed to be invalidated and / or evicted (replaced), and a pointer pointing to the relevant cache entry will be retained in the scheduler for use by the actual demand access.

[0058] Although this solution may increase access bandwidth by eliminating the second lookup through the demand access, this mechanism may have significant limitations. First, if all cache entries that may be used in the cache are locked, additional memory accesses requesting to load additional data that is not currently cached cannot be satisfied. Some schedulers may adopt a dedicated mechanism to release the lock to facilitate progress. However, this lock release mechanism may be very complex, requiring a large amount of hardware resources (such as multiple logic circuits, multiple memory units, etc.), and / or reducing memory access performance. In addition, locking cache entries may harm performance because it effectively reduces the cache size, and new data loaded into the cache must replace another entry that may have a higher priority, thus reducing the cache hit rate.

[0059] On the other hand, the combined scheduler does not lock cache entries, which can prevent cache blocking for further data and thus does not reduce memory access latency and / or bandwidth, preventing a decline in execution performance. Instead, the combined scheduler monitors the cache entries mapped by the pointer associated with the instruction scheduled to be executed to track invalidation and / or eviction, and can reload the data associated with the instruction in the case of invalidation and / or eviction.

[0060] Other existing schedulers can use a dedicated prefetch buffer. In this way, while instructions are being allocated to the scheduler that is being written, these instructions can also be written to a prefetch buffer. The prefetch buffer can be adapted to initiate a prefetch access that will search (look up) data addresses cached in the cache, and the scheduler, in addition to accessing the dedicated prefetch buffer, can only issue actual demand accesses that may search (look up) the cache as if no prefetch had been performed. A prefetch access can only be issued when the data requested by a demand access is not found in the prefetch buffer to send an outstanding request in case of a cache miss. The prefetch access may not be updated in the scheduler.

[0061] There are also some major drawbacks and limitations to this mechanism. First, the prefetch access (look up) bandwidth can be very poor because each data request may be dispatched twice, once during the prefetch access and once during the demand access. In addition, the prefetch buffer is merely a copy of the scheduler and may thus consume additional hardware resources or be reduced in size, which will obviously limit the prefetch capacity and / or result in fewer prefetch accesses. In addition, unlike the scheduler, the prefetch buffer may not support age-based scheduling. In addition, the prefetch access may only be dispatched once, and if the cache line is evicted after a prefetch access to the cache line, no other prefetch will be sent before the actual demand access. Therefore, the demand access must perform the cache lookup again and may need to access the next level of memory, which may further reduce the memory bandwidth and / or increase the access latency.

[0062] In contrast, using the combined scheduler only requires one lookup during a prefetch access related to the corresponding instruction before executing the corresponding instruction, and the demand access sent when executing the corresponding instruction can directly access the mapped cache entry using the pointer associated with the corresponding instruction without performing a lookup. Performing only one lookup can significantly reduce the demand access latency and / or increase the access bandwidth. In addition, the combined scheduler does not employ a dedicated buffer, so on the one hand, it does not increase the consumption of hardware resources, and on the other hand, it can be increased in size, thereby further reducing the memory access latency and / or increasing the memory access bandwidth. In addition, the combined scheduler can employ a variety of scheduling algorithms and / or techniques, including age-based scheduling and other scheduling algorithms such as location-based scheduling, etc.

[0063] In addition, the combined scheduler can significantly improve the out-of-order execution performance of the processing unit. Since it continuously monitors the cache entries to track and identify invalidation and / or eviction events, the combined scheduler can reload previously invalidated and / or evicted data associated with the instructions in the execution pipeline, regardless of their order. Thus, even if the data associated with an instruction located further downstream in the pipeline is evicted, this data will be quickly and automatically available again, enabling the execution of more downstream instructions earlier, and thus significantly increasing out-of-order execution, which, as is well known in the art, can significantly improve the execution performance of the processing unit.

[0064] Before explaining in detail at least one embodiment of the present invention, it is to be understood that the application of the present invention is not necessarily limited to the details of the construction and component arrangements and / or methods set forth in the following description and / or shown in the drawings and / or embodiments. The present invention is capable of other embodiments or of being practiced or carried out in various ways.

[0065] Those skilled in the art will recognize that aspects of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which are generally referred to herein as "circuits", "modules", or "systems". In addition, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied therein.

[0066] Any combination of one or more computer-readable media may be used. A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device (such as a punch card or raised structures in a groove having instructions recorded thereon), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0067] The computer program code containing computer-readable program instructions is contained on a computer-readable medium and can be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, etc., or any suitable combination of the foregoing.

[0068] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.

[0069] The computer-readable program instructions for performing the operations of the present invention can be written in any combination of one or more programming languages, such as assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (e.g., Smalltalk, C++ etc.) and conventional procedural programming languages (e.g., the "C" programming language or similar programming languages).

[0070] The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuits in order to perform aspects of the present invention.

[0071] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0072] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or sometimes may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a special-purpose hardware system that performs the specified functions or acts, or by a combination of special-purpose hardware and computer instructions.

[0073] Referring to the accompanying drawings, Figure 1 is a flowchart of an exemplary process of a joint scheduler using a processing unit configured to efficiently schedule memory accesses according to some embodiments of the present invention.

[0074] An exemplary process 100 may be executed by a joint scheduler of each of one or more processing units to schedule memory accesses, and efficiently schedule memory accesses in a processing unit having one or more cache memory units by sending prefetch accesses and data demand accesses or sending the same instructions in association with each other, thereby improving the performance of data loading and / or data caching, for example, reducing latency, increasing bandwidth, reducing lock time, etc.

[0075] Multiple instructions loaded into an execution pipeline of a processing unit may access (e.g., read and / or write) data stored in a memory accessible by the processing unit.

[0076] As is well known in the art, to improve performance and reduce latency, the processing unit may be equipped with one or more high-speed low-latency cache memories, such as L1 cache, L2 cache, etc., for caching data stored in the memory of the processing unit. However, although supporting high-speed access, the capacity of the cache may be limited, so only a limited amount of data can be stored at any given time.

[0077] To further improve data access performance, since data related to a specific instruction may not be cached (loaded into the cache), instead of loading the data when executing the corresponding instruction, a prefetch access can be initiated in advance so as to load this data into the cache (the cache) before executing these instructions, thereby significantly reducing the access time and further reducing the execution time of the instructions.

[0078] See also Figure 2A and Figure 2B , Figure 2A and Figure 2B are schematic diagrams of an exemplary combined scheduler of a processing unit configured to efficiently schedule memory accesses according to some embodiments of the present invention.

[0079] As Figure 2A shown, an exemplary combined scheduler 200 (alternatively designated as a dual scheduler) of an exemplary processing unit (processor) 202 may include one or more combined scheduler circuits configured to and / or adapted to send both prefetch accesses and demand accesses to access data related to one or more instructions, the plurality of instructions being loaded in an execution pipeline 210 of the processing unit 202, e.g., data reads, data writes, etc.

[0080] The processing unit 202 may include one or more units, circuits, modules, etc. known in the art. For example, the processing unit 202 may include an input circuit 212 and an output circuit 214 for receiving and outputting data.

[0081] The processing unit 202 may further include a control unit 216 for controlling the operation and / or activities of the processing unit 202 and one or more arithmetic logic units (ALUs) 218, e.g., addition, subtraction, multiplication, division, shift, AND, OR, NOT, and / or etc., and the processing unit 202 may further execute instructions involving mathematical calculations. The ALU 218 may be adapted to operate according to one or more digital formats (e.g., integer, floating point, fixed point, and / or etc.).

[0082] The processing unit 202 may include a memory control unit 220 for controlling, managing, and / or processing accesses to memory, e.g., loading, reading, writing of data and / or instructions (i.e., program instructions). The memory control unit 220 may be adapted to control one or more memory channels and may include one or more memory control modules known in the art, e.g., a memory management unit (MMU), etc.

[0083] In addition, the processing unit 202 may include one or more caches 222, such as an L1 cache, which includes one or more high-speed low-latency devices for storing (caching) data and / or instructions loaded from the next-level memory 230, and thus the processing unit 202 can access this data and / or instructions with low latency. Although extremely fast, the capacity of the cache 222 may be limited, and thus only a limited fragment of data and / or instructions can be stored.

[0084] The next-level memory 230 may also be controlled by the memory control unit 220. For example, the next-level memory 230 may include one or more next-level cache units, such as an L2 cache, an L3 cache, etc., which are generally slower than the cache 222 but have a larger capacity. In another example, the next-level memory 230 may include one or more standard memory low-speed memory arrays, such as a system memory, an application memory, etc., which may be used by one or more volatile memory devices (such as RAM) and / or non-volatile memory devices (such as Flash, etc.).

[0085] It should be noted that the processing unit 202 is shown only in an exemplary form. Thus, the processing unit 202 may include one or more additional units, circuits, modules, etc. not shown, and / or lack one or more elements shown for the processing unit 202, as is known in the art.

[0086] The processing unit 202 may adopt one or more processor architectures, structures, and / or instruction sets and / or the like that support one or more bit widths, such as 32 bits, 56 bits, 64 bits, and / or the like. For example, the processing unit 202 may optionally have a von Neumann architecture, such as a Central Processing Unit (CPU), a multi-core CPU, a Data Processing Unit (DPU), a Microcontroller Unit (MCU), an Accelerated Processing Unit (ACU), etc. In another example, the processing unit 202 may optionally have a non-von Neumann architecture, such as a Graphical Processing Unit (GPU), a DPU, a Field-Programmable Gate Array (FPGA), a Coarse-Grained Reconfigurable Architecture (CGRA), a neural network accelerator, an Intelligence Processing Unit (IPU), an Application-Specific Integrated Circuit (ASIC), a quantum computer, an interconnected computing grid, etc.

[0087] The processing unit 202 may be implemented, constructed, and / or deployed according to one or more designs, structures, and / or implementation manners. For example, the processing unit 202 may be implemented as an independent device, system, apparatus, etc. In another example, the processing unit 202 may be integrated in one or more higher-level integrated devices. For example, the processing unit 202 may be implemented as an integrated circuit (IC), an ASIC, an FPGA, etc. in one or more devices, such as a computer, a server, a computing device, etc. including additional elements. In another example, the processing unit 202 may be integrated in one or more higher-level integrated circuits. For example, the processing unit 202 may be implemented as a functional module (such as an IP core, etc.) embedded in one or more integrated components (such as an IC, an ASIC, an FPGA, a CPU, a GPU, etc.), and the integrated components include one or more additional functional elements of the integrated components.

[0088] As Figure 2BAs shown, the exemplary joint scheduler 200 may be adapted to send a plurality of prefetch accesses and data demand accesses (alternatively designated as demand accesses) for accessing data related to a plurality of instructions loaded in the execution pipeline 210 of the processing unit 202.

[0089] Accordingly, the joint scheduler 200 may include a plurality of instruction entries, each instruction entry storing a related control payload for each instruction loaded in the execution pipeline 210 from allocation to release (completion). The control payload is used to determine which operations are scheduled.

[0090] Each instruction entry may also associate each of the plurality of instructions loaded in the execution pipeline with a corresponding valid indication and a pointer field, the pointer field including a pointer to a corresponding cache entry.

[0091] A scheduler entry stores the related control payload for each operation from its allocation to release (completion). This control payload is used to determine which operations are scheduled.

[0092] The valid indication associated with each instruction in a corresponding instruction entry of the joint scheduler 200 indicates that the data related to the corresponding instruction is cached (loaded) in the cache 222, and the pointer field associated with the corresponding instruction points to the cache entry in the cache 222 that stores the data related to the corresponding instruction.

[0093] One or more methods, techniques, and / or embodiments may be used to utilize the valid indication. For example, each instruction entry may be associated, correlated, and / or assigned one or more corresponding bits (i.e., flags) that, when set or cleared, indicate that the data related to the instruction of the corresponding instruction entry is cached in the cache 222. In another example, the valid indication may be implemented via the pointer field. For example, a pointer field including a valid pointer pointing to a valid cache entry indicates that the data related to the corresponding instruction is cached in the cache 222, while a value, pattern, etc. (e.g., 0xFFFFFFFF) constituting an invalid pointer that does not point to a valid cache entry may indicate that the data related to the corresponding instruction is not cached in the cache 222.

[0094] For example, a first instruction entry (0) may associate a first instruction (0) with a corresponding valid indication V(0), where the corresponding valid indication indicates that data related to the instruction (0) is cached in the cache 222 of a cache entry, and the cache entry pointed to by the pointer (0) is also associated with the instruction (0). In another example, a second instruction entry (1) may associate a second instruction (1) with a corresponding valid indication V(1), where the corresponding valid indication indicates that data related to the instruction (1) is cached in the cache 222 of a cache entry, and the cache entry pointed to by the pointer (1) is also associated with the instruction (1). This can be repeated until an Nth instruction entry (N) may associate an Nth instruction (N) with a corresponding valid indication V(N), where the corresponding valid indication indicates that data related to the instruction (N) is cached in the cache 222 of a cache entry, and the cache entry pointed to by the pointer (N) is also associated with the instruction (N).

[0095] As is well known in the art, a prefetch access is an access (cycle) that starts loading data related to an instruction that the processing unit 202 plans to execute by accessing the memory to map the requested data address. During the prefetch access, a lookup may first be performed to search for the requested data address in the cache 222. Specifically, the prefetch access starts before instruction execution to first check whether the requested data has been loaded into one or more caches 222 of the processing unit 202, and if there is a miss, i.e., the requested data is not loaded into the cache 222, then the data is loaded from a lower-level memory into the cache 222, i.e., the data is cached.

[0096] A demand access is an access (cycle) that is started to load data related to an instruction that the processing unit 202 is currently executing from the memory. Since the demand access typically precedes a corresponding prefetch access to load the same data from the same address, the requested data can usually be cached, i.e., stored in the cache 222.

[0097] The combined scheduler 200 may include one or more prefetch ports for sending prefetch accesses and one or more demand ports for sending demand accesses, and these ports are separate and independent of each other. Therefore, the combined scheduler 200 can send prefetch accesses and demand accesses simultaneously, which are typically related to different instructions loaded into the execution pipeline 210. In addition, assuming that the combined scheduler 200 includes multiple prefetch ports, the combined scheduler 200 can send multiple prefetch accesses related to multiple different instructions simultaneously. Similarly, assuming that the combined scheduler 200 includes multiple demand ports, the combined scheduler 200 can send multiple demand accesses related to multiple different instructions simultaneously.

[0098] For simplicity, the joint scheduler 200 is described in process 100 to send a single prefetch access and a single demand access, both of which are related to a single instruction loaded into the execution pipeline 210. However, this should not be construed as a limitation, as the joint scheduler 200 may repeat, extend, and / or scale process 100 to send multiple prefetch accesses and multiple demand accesses, where the multiple prefetch accesses and the multiple demand accesses are related to multiple instructions loaded into the execution pipeline 210. Additionally, as previously described, the joint scheduler 200 may include multiple prefetch ports and / or multiple demand ports and thus may send multiple prefetch accesses and / or multiple demand accesses simultaneously.

[0099] Also for simplicity, although the processing unit 202 may include multiple caches 222, process 100 is described for a single cache 222. However, this should not be construed as a limitation, as the same method applied by the joint scheduler 200 may be similarly applied to and serve multiple caches 222, where the multiple caches 222 are controlled by a memory control unit 220 and generally control the joint scheduler 200 transparently.

[0100] As shown in step 102, process 100 begins by traversing one or more instructions loaded into an execution pipeline 210 of a processing unit 202 and scheduled for execution by the processing unit 202 from the joint scheduler 200 and extracting data addresses related to one or more of the loaded instructions.

[0101] As shown in step 104, the joint scheduler 200 may send a prefetch access to a corresponding address of a corresponding instruction, where the corresponding instruction is related to a corresponding instruction loaded into the execution pipeline 210.

[0102] As shown in step 106, if a hit occurs, i.e., the corresponding data is cached, i.e., loaded and stored in the cache 222, then process 100 may branch to step 108. However, if a miss occurs, i.e., the corresponding data is not cached in the cache 222, then process 100 may branch to step 110.

[0103] As shown in step 108, in response to a hit prefetch access sent for the corresponding data, the joint scheduler 200 may associate the corresponding instruction with a valid indication and a pointer to a corresponding cache entry in the cache 222, where the corresponding cache entry stores the corresponding data.

[0104] As shown in steps 110 - 118, in response to a miss prefetch access sent for the corresponding data, the combined scheduler 200 may initiate a read cycle to load the corresponding data from the next level memory and cache it in the cache 222.

[0105] As shown in step 110, before loading the corresponding data from the next level memory and caching it in the cache 222, the combined scheduler 200 may block further accesses, whether prefetch accesses and / or demand accesses related to the corresponding instruction, until the corresponding data is cached and available in the cache 222.

[0106] To this end, the combined scheduler 200 may associate the corresponding instruction with a block tag indicating that the corresponding data related to the corresponding instruction is currently being cached, and each prefetch access and / or demand access related to the corresponding instruction is blocked from being sent until the corresponding data is fetched from the next level memory.

[0107] As shown in step 112, the combined scheduler 200 may initiate a read cycle to the next level memory to obtain the corresponding data and store (cache) it in the cache 222.

[0108] As shown in step 114, in response to successfully completing the initiated read cycle to load the corresponding data from the next level memory and cache it in the cache 222, the combined scheduler 200 may associate the corresponding instruction with a valid indication and a pointer to the corresponding cache entry in the cache 222, where the corresponding cache entry stores the cached corresponding data. This operation is similar to the operation of step 108 that the combined scheduler performs in response to a hit prefetch access.

[0109] As shown in step 118, in response to the successful completion of the prefetch cycle, the combined scheduler 220 may remove the block tag associated with the corresponding instruction, thereby allowing one or more other accesses to be sent, typically demand accesses related to the corresponding instruction, for which the corresponding data is loaded and cached in the cache 222.

[0110] As shown in step 120, the combined scheduler 220 may send a demand access related to the corresponding instruction.

[0111] Specifically, the combined scheduler 220 may check the valid indication associated with the corresponding instruction and thereby determine whether the corresponding data associated with the corresponding instruction is cached in the cache 222. In response to determining that the corresponding data has been cached, the combined scheduler 220 may use the pointer associated with the corresponding instruction to send the demand access to access the corresponding data in the cache entry pointed to in the cache 222, such as reading, writing, updating, etc.

[0112] This means that the demand access can be performed without initiating an additional lookup to check whether the corresponding data is cached, e.g., by traversing an address cache, i.e., the tag / index field of the cache 222, as can be done by existing methods.

[0113] Directly accessing the cache entry mapped by the pointer during the demand access and eliminating the need for the additional lookup can significantly reduce the latency of accessing data associated with the currently executing instruction, thereby significantly improving the performance of the processing unit 202. Additionally, using the pointer to access the corresponding data in the demand access can also significantly increase the bandwidth of data access, as the saved time is used to send and execute more prefetch accesses and / or demand-by accesses, thereby further improving the performance of the processing unit 202.

[0114] Optionally, in a case where one or more prefetch accesses and one or more demand accesses related to a common cache entry are simultaneously scheduled through corresponding prefetch ports and demand ports, in response to the completion of the prefetch access, a combined scheduler 200 may directly update, in the demand port, a pointer pointing to the common cache entry for one or more demand accesses.

[0115] This means that, while a corresponding prefetch access is in progress, the demand access is not blocked until the corresponding data is cached, and the combined scheduler 200 may send the demand access and update the pointers associated with them after the prefetch access is completed, at which time the demand access is already in the pipeline and waiting for its turn to be sent.

[0116] Specifically, the combined scheduler 200 may identify the cache entry that caches data by the simultaneous prefetch access and directly update the pointer associated with the corresponding demand access to map the cache entry that caches the data obtained by the simultaneous prefetch access, thereby bypassing the standard mechanism for non-simultaneous access.

[0117] Optionally, the combined scheduler 200 may be adapted to track the invalidation and / or eviction of data stored in each of the plurality of cache entries in cache 222 and update the valid indication associated with the corresponding instruction accordingly. This means that the combined scheduler 200 can monitor each cache entry pointed to (mapped) by a pointer associated with each instruction in pipeline 210, where the pointer is also associated with the valid indication, to detect the invalidation and / or eviction of data cached in the cache entry.

[0118] In response to the eviction of a corresponding cache entry storing a corresponding data, where the corresponding data is related to a corresponding one of the plurality of instructions loaded into execution pipeline 210, the combined scheduler 220 may associate the corresponding instruction with an invalid indication, which indicates that the corresponding data related to the corresponding instruction is no longer cached in cache 222.

[0119] The combined scheduler 200 may also initiate another prefetch cycle to load the corresponding data from the next level of memory into cache 222. After another prefetch access fails, steps 110 to 118 may be repeated, and the valid indication and the pointer associated with the instruction will be updated again to indicate that the corresponding data is cached in a cache entry pointed to by the relevant pointer.

[0120] Optionally, the combined scheduler 220 may be adapted to mark a corresponding cache entry with an active indication after each prefetch access hits. The combined scheduler circuit 220 may use the active indication to track (listen for) the invalidation and / or eviction of data stored in the cache entry marked with the active indication, where the active indication indicates that the corresponding cache entry is mapped by a pointer associated with one or more instructions loaded into execution pipeline 210.

[0121] By tracking the invalidation and / or eviction of data stored in cache 222, the combined scheduler 220 can significantly enhance the out-of-order execution of one or more instructions loaded into pipeline 210.

[0122] The combined scheduler 220 achieves this by performing steps 110 - step 118 of process 100, and in response to detecting invalidation and / or eviction, sending another prefetch cycle to load the corresponding invalidated and / or evicted data from the next level of memory 230, where the corresponding invalidated and / or evicted data is related to one or more instructions loaded into pipeline 210, so that the corresponding data is available in cache 222 during a demand access sent for the corresponding instruction.

[0123] Now refer to Figure 3A and Figure 3B ,Figure 3A and Figure 3B is a schematic diagram of an exemplary scheduler of a processing unit configured to schedule memory accesses. Specifically, scheduling memory accesses is known in the art.

[0124] As Figure 3A shown, a single-issue scheduler 310 known in the art may be adapted to issue multiple memory accesses to access data related to one or more instructions loaded in an execution pipeline (e.g., execution pipeline 210).

[0125] Specifically, the single-issue scheduler 310 may issue multiple memory accesses, whether they are prefetch accesses or demand accesses, as long as they are not blocked, i.e., they are ready to be issued.

[0126] Each new instruction assigned to the single-issue scheduler 310 may have its own blocked / ready status (some cases may be common). Instructions that are "ready to be issued" to one or more ports may be scheduled for issue.

[0127] Thus, the single-issue scheduler 310 may include a ready logic 312 implemented using one or more circuits, components, and / or elements, which is adapted to determine whether each memory access is ready or blocked.

[0128] The single-issue scheduler 31 may also include an access scheduler 314 that employs one or more scheduling algorithms, such as age-based scheduling, location-based scheduling, etc., to schedule one or more memory accesses that are not affected by any blocking conditions and are thus ready to be issued through one or more issue ports.

[0129] As Figure 3B shown, the ready logic 312 may be implemented, as is well known in the art, by separately monitoring the blocking status of each memory access (whether it is a prefetch access or a demand access) and generating a ready indication (e.g., a flag) accordingly.

[0130] Now refer to Figure 4A and Figure 4B , Figure 4A and Figure 4B is a schematic diagram of an exemplary combined scheduler of a processing unit configured to efficiently schedule memory accesses according to some embodiments of the present invention.

[0131] As Figure 4A shown, a combined (dual) issue scheduler 410 (e.g., combined scheduler 200) may be adapted to schedule one or more prefetch accesses and one or more demand accesses to access data related to one or more instructions loaded in an execution pipeline (e.g., execution pipeline 210).

[0132] Specifically, the dual-issue scheduler 410 can schedule one or more prefetch accesses and one or more demand accesses, and each access is scheduled through one or more corresponding issue ports, which means that prefetch accesses and demand accesses are processed differently from each other.

[0133] The dual-issue scheduler 410 can include ready logic 412 implemented using one or more circuits, components, and / or elements, which is adapted to determine the blocking / ready status of prefetch accesses and demand accesses respectively related to each instruction (entry) loaded in the execution pipeline 210 and assigned by the dual-issue scheduler 410 for scheduling.

[0134] This means that the ready logic 412 can associate prefetch accesses and demand accesses related to each instruction, but can determine the blocking / ready status separately for each access.

[0135] The dual-issue scheduler 410 can include a first scheduler, namely the prefetch access scheduler 414, which employs one or more scheduling algorithms, such as age-based scheduling, location-based scheduling, etc., for scheduling one or more prefetch accesses that are not affected by any blocking conditions and are thus ready to be issued via one or more issue ports.

[0136] The dual-issue scheduler 410 can also include a second scheduler, namely the demand access scheduler 416, which employs one or more scheduling algorithms, such as age-based scheduling, location-based scheduling, etc., for scheduling one or more demand accesses that are not affected by any blocking conditions and are thus ready to be issued through one or more issue ports.

[0137] As Figure 4B shown, the ready logic 412 can monitor the blocking status of prefetch accesses and demand accesses related to the same instruction respectively and generate corresponding ready indications (e.g., flags) accordingly. Thus, in response to determining that a specific prefetch access related to a specific instruction is ready to be issued, the ready logic 412 can issue a prefetch ready indication. Similarly, in response to determining that a specific demand access related to a specific instruction is ready to be issued, the ready logic 412 can issue a demand ready indication.

[0138] The description of the various embodiments of the present invention is for illustrative purposes only and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology found in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

[0139] It is expected that during the term of the patent for this application, many related systems, methods, and computer programs will be developed, and the scope of terms such as processor architecture, cache architecture, and scheduling algorithms is intended to include all such new technologies a priori.

[0140] As used herein, the term "about" means ±10%.

[0141] The terms "comprising," "includes," "containing," "having," and their variations mean "including but not limited to." This term encompasses the terms "consisting of" and "consisting essentially of."

[0142] The phrase "consisting essentially of" means that a composition or method may include additional ingredients and / or steps, provided that the additional ingredients and / or steps do not materially alter the basic and novel characteristics of the claimed composition or method.

[0143] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" may include multiple compounds, including mixtures thereof.

[0144] As used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any embodiment described as "exemplary" is not necessarily to be construed as superior to or better than other embodiments and / or to exclude combinations of features in other embodiments.

[0145] As used herein, the word "optional" means "provided in some embodiments and not provided in other embodiments." Any particular embodiment of the invention may include a plurality of "optional" features, unless such features are mutually conflicting.

[0146] Throughout this application, various embodiments of the invention may be presented in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the range description should be regarded as having specifically disclosed all the possible sub-ranges as well as the individual numerical values within that range. For example, a range description (e.g., from 1 to 6) should be considered to have specifically disclosed sub-ranges (e.g., from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc.) and the individual numbers within the range, e.g., 1, 2, 3, 4, 5, and 6. This applies regardless of how broad the range is.

[0147] Whenever a numerical range is indicated herein, it is intended to include any cited numeral (fractional or integer) within the indicated range. The phrases "range / range between" a first indicated numeral and a second indicated numeral and "range / range from" a first indicated numeral "to" a second indicated numeral are used interchangeably herein and are intended to include the first and second indicated numerals and all fractions and integers therebetween.

[0148] It should be appreciated that certain features of the invention described in separate embodiments for clarity may also be provided in combination in a single embodiment. Conversely, various features of the invention described in a single embodiment for brevity may also be provided separately or in any suitable sub-combination or in any suitable manner in any other described embodiment of the invention. Certain features described in various embodiments should not be considered essential features of those embodiments unless the embodiment cannot be implemented without those elements.

[0149] Although the present invention has been described in conjunction with its specific embodiments, it is obvious that many substitutions, modifications and changes will occur to those skilled in the art. Therefore, the present invention is intended to cover all substitutions, modifications and changes that fall within the spirit and broad scope of the appended claims.

[0150] It is the intention of the applicant that all publications, patents and patent applications mentioned in this specification should be incorporated by reference in their entirety into this specification, just as if each individual publication, patent or patent application was clearly and individually indicated to be incorporated by reference herein when cited. In addition, the citation or identification of any reference in this application should not be construed as an admission that such reference can serve as prior art for the present invention. To the extent that section headings are used, they should not be construed as necessarily limiting. In addition, any priority document of this application is hereby incorporated by reference in its entirety.

Claims

1. A joint scheduler, characterized in that: The joint scheduler comprises: a co-scheduler circuit adapted to dispatch pre-fetch accesses and demand accesses of data associated with a plurality of instructions loaded into an execution pipeline of at least one processing circuit, each pre-fetch access comprising checking whether a corresponding data is cached in one of a plurality of cache entries of at least one cache, each demand access comprising accessing the corresponding data; The joint scheduler circuit is adapted to: In response to each hit prefetch access issued for a corresponding data associated with a corresponding instruction of the plurality of instructions, associating the corresponding instruction with a valid indication and a pointer to a corresponding cache entry storing the corresponding data such that the demand access associated with the corresponding instruction uses the associated pointer to access the corresponding data in the at least one cache; and In response to each missed prefetch access sent for the corresponding data associated with the corresponding instruction of the plurality of instructions, a read cycle is initiated to load the corresponding data from a next level memory and cache the corresponding data in the at least one cache.

2. The joint scheduler according to claim 1, characterized in that: The joint scheduler circuit is suitable for: in response to the successful completion of a corresponding read cycle, the corresponding read cycle is initiated to load the corresponding data from the next level memory and cache the corresponding data in the at least one cache after a corresponding miss prefetch access, associating the corresponding instruction with the valid indication and the pointer to the corresponding cache entry, and the corresponding cache entry stores the cached corresponding data.

3. The joint scheduler according to claim 1, characterized in that: The joint scheduler circuit is suitable for: in response to the successful completion of a corresponding read cycle, the corresponding read cycle is initiated to load the corresponding data from the next level memory and cache the corresponding data in the at least one cache after a corresponding miss prefetch access, sending another prefetch access to update the valid indication and the pointer to the corresponding cache entry, and the corresponding cache entry stores the cached corresponding data.

4. The joint scheduler according to claim 1, characterized in that: The joint scheduler circuit is also adapted to track invalidation and / or eviction of data stored in each of the plurality of cache entries, and in response to eviction of the corresponding cache entry, the corresponding cache entry stores the corresponding data associated with the corresponding instruction of the plurality of instructions, associates the corresponding instruction with an invalidation indication, and initiates another prefetch cycle to load the corresponding data from the next level memory into the at least one cache.

5. The joint scheduler according to claim 1, characterized in that: The joint scheduler circuit is further adapted to: for each hitting prefetch access, mark the corresponding cache entry with an activity indication indicating that the corresponding cache entry is mapped by a pointer associated with at least one of the plurality of instructions, and the joint scheduler uses the activity indication to track invalidation and / or eviction of data stored in the corresponding cache entry.

6. The joint scheduler according to claim 1, characterized in that: The joint scheduler circuit is adapted to, for each prefetch access that missed, associate the corresponding instruction with a block tag indicating that each prefetch access and / or demand access associated with the corresponding instruction is blocked from being issued.

7. The joint scheduler according to claim 6, characterized in that: The co-scheduler circuit is adapted to delete the block tag associated with the corresponding instruction in response to successful completion of the prefetch cycle.

8. The joint scheduler according to claim 1, characterized in that: The joint scheduler circuit includes at least one prefetch port for sending the prefetch accesses and at least one demand port for sending the demand accesses, the at least one prefetch port being separate and independent from the at least one demand port.

9. The joint scheduler according to claim 8, characterized in that: At least one pre-fetch access and at least one demand access are sent simultaneously by the at least one pre-fetch port and the at least one demand port, respectively, independently.

10. The joint scheduler according to claim 9, characterized in that: For the at least one prefetch access and the at least one demand access related to a public cache entry sent simultaneously, corresponding to the completion of the at least one prefetch access, a pointer pointing to the public cache entry is directly updated in the at least one demand port for the at least one demand access.

11. A method for jointly scheduling pre-fetch access and demand access, characterized in that: The method for jointly scheduling pre-fetch access and demand access includes: Using a joint scheduler circuit, the joint scheduler circuit is adapted to send pre-fetch accesses and demand accesses of data associated with a plurality of instructions loaded into an execution pipeline of at least one processing circuit, each pre-fetch access including checking whether a corresponding data is cached in one of a plurality of cache entries of at least one cache, and each demand access including accessing the corresponding data; the joint scheduler circuit is adapted to: In response to each hit prefetch access issued for a corresponding data associated with a corresponding instruction of the plurality of instructions, associating the corresponding instruction with a valid indication and a pointer to a corresponding cache entry storing the corresponding data such that the demand access associated with the corresponding instruction uses the associated pointer to access the corresponding data in the at least one cache; and In response to each missed prefetch access sent for the corresponding data associated with the corresponding instruction of the plurality of instructions, a read cycle is initiated to load the corresponding data from a next level memory and cache the corresponding data in the at least one cache.

12. The method for jointly scheduling pre-fetch access and demand access according to claim 11, characterized in that: The joint scheduler circuit is suitable for: in response to the successful completion of a corresponding read cycle, the corresponding read cycle is initiated to load the corresponding data from the next level memory and cache the corresponding data in the at least one cache after a corresponding miss prefetch access, associating the corresponding instruction with the valid indication and the pointer to the corresponding cache entry, and the corresponding cache entry stores the cached corresponding data.

13. The method for jointly scheduling pre-fetch access and demand access according to claim 11, characterized in that: The joint scheduler circuit is suitable for: in response to the successful completion of a corresponding read cycle, the corresponding read cycle is initiated to load the corresponding data from the next level memory and cache the corresponding data in the at least one cache after a corresponding miss prefetch access, calling another prefetch access to update the valid indication and the pointer to the corresponding cache entry, and the corresponding cache entry stores the cached corresponding data.

14. The method for jointly scheduling pre-fetch access and demand access according to claim 11, characterized in that: The joint scheduler circuit is also adapted to track invalidation and / or eviction of data stored in each of the plurality of cache entries, and in response to eviction of the corresponding cache entry, the corresponding cache entry stores the corresponding data associated with the corresponding instruction of the plurality of instructions, associates the corresponding instruction with an invalidation indication, and initiates another prefetch cycle to load the corresponding data from the next level memory into the at least one cache.

15. The method for jointly scheduling pre-fetch access and demand access according to claim 11, characterized in that: The joint scheduler circuit is further adapted to: for each hitting prefetch access, mark the corresponding cache entry with an activity indication indicating that the corresponding cache entry is mapped by a pointer associated with at least one of the plurality of instructions, and the joint scheduler uses the activity indication to track invalidation and / or eviction of data stored in the corresponding cache entry.

16. The method for jointly scheduling pre-fetch access and demand access according to claim 11, characterized in that: The joint scheduler circuit is adapted to, for each prefetch access that missed, associate the corresponding instruction with a block tag indicating that each prefetch access and / or demand access associated with the corresponding instruction is blocked from being issued.

17. The method for jointly scheduling pre-fetch access and demand access according to claim 16, characterized in that: The co-scheduler circuit is adapted to delete the block tag associated with the corresponding instruction in response to successful completion of the prefetch cycle.

18. A joint scheduler, characterized in that: The joint scheduler comprises: A co-scheduler circuit is adapted to enhance out-of-order execution of a plurality of instructions loaded into an execution pipeline of at least one processing circuit by: tracking invalidation and / or eviction of data stored in each of a plurality of cache entries of at least one cache storing data associated with the plurality of instructions; and A prefetch cycle is issued to load invalid data and / or evicted data associated with at least one of the plurality of instructions, thereby making the corresponding data available in the at least one cache during a demand access issued for the at least one instruction.

19. The joint scheduler according to claim 18, characterized in that: The co-scheduler is adapted to: send the prefetch cycle in response to eviction and / or invalidation of at least one cache entry storing corresponding data, the corresponding data being associated with a corresponding instruction of the plurality of instructions, by: initiating the prefetch cycle to load the corresponding data from a next level memory into the at least one cache; and The corresponding instruction is associated with a valid indication and a pointer to a corresponding cache entry, the corresponding cache entry storing the corresponding data, so that the required access related to the corresponding instruction uses the associated pointer to access the corresponding data in the at least one cache.

20. The joint scheduler according to claim 19, characterized in that: The joint scheduler circuit is further adapted to: associate the corresponding instruction with the valid indication and the pointer by initiating another prefetch cycle; The joint scheduler circuit is adapted to: in response to a hit in the another prefetch cycle, associate the corresponding instruction with the valid indication and the pointer.