Zero Jitter Cache Queue Manager

By marking the cache behavior of the queue head and tail in the computer system's cache without evicting, the problem of cache bumps in queue data access is solved, and the zero bump effect is achieved and system performance is improved.

CN109840220BActive Publication Date: 2025-06-10INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201811269286.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-11-29
Filing Date
2018-10-29
Publication Date
2025-06-10
Estimated Expiration
2038-10-29

AI Technical Summary

Technical Problem

In computer systems, bumps are prone to occur when processing queue data in caches, resulting in performance degradation, especially when the order of storage and retrieval of queue data is opposite to the assumption of cache management.

Method used

Cache bumps are avoided by marking the head and tail cache behavior of the queue in the cache without eviction, ensuring that these critical data is not eviction frequently.

Benefits of technology

It realizes the zero bump effect when queue data access is accessed, improves the cache hit rate and system performance, and reduces the number of accesses to external memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109840220B_ABST
    Figure CN109840220B_ABST
Patent Text Reader

Abstract

This application provides a zero-thrashing cache queue manager. Various systems and methods for queue management in computer memory are described herein. A system for implementing a zero-thrashing cache queue manager performs operations including the steps of: receiving a memory access request for a queue; when the memory access request is to add data to the queue, writing the data to a queue tail cache line in the cache and protecting the queue tail cache line from eviction from the cache; and when the memory access request is to remove data from the queue, reading the data from the current queue head cache line in the cache, the current queue head cache line being protected from eviction from the cache.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments described herein generally relate to computer memory management, and more particularly to systems and methods for implementing a zero-thrash cache queue manager. Background Art

[0002] In a computer microarchitecture, a cache is memory used for temporarily storing information for faster access. Typically, a cache is implemented as a smaller, faster memory that is closer to the processor core. The cache is used to store copies of data from frequently accessed main memory locations. When a program requests data from a memory location, the data can be stored in the cache, anticipating that the data may be requested again in the near future. If the data is requested again, this results in a cache hit, and the data is provided from the faster cache rather than from a slower memory device such as dynamic random access memory (DRAM). Since cache memory is generally more expensive to manufacture and since cache memory is typically implemented on-die, efficient cache usage is crucial for ensuring maximum system performance and maintaining a small die size. Brief Description of the Drawings

[0003] In the drawings (which are not necessarily drawn to scale), the same numbers may describe similar components in different views. The same numbers with different letter suffixes may represent different instances of similar components. Some embodiments are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which:

[0004] Figure 1A and Figure 1B illustrates cache thrashing and a mechanism for removing queued cache thrashing according to an embodiment;

[0005] Figure 2 is a block diagram illustrating a queue cache mechanism according to an embodiment;

[0006] Figure 3A and Figure 3B is a block diagram illustrating cache management according to an embodiment;

[0007] Figure 4 is a flowchart illustrating a process for implementing a zero-thrash cache queue manager according to an embodiment;

[0008] Figure 5 is a flowchart illustrating a method for implementing a zero-thrash cache queue manager according to an embodiment; and

[0009] Figure 6FIG. 0 is a block diagram of an example machine on which any one or more of the techniques (e.g., methods) discussed herein may be performed. DETAILED DESCRIPTION

[0010] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of some example embodiments. It will be apparent, however, to one skilled in the art that the present disclosure may be practiced without these specific details.

[0011] When data is requested from memory, the data is copied into the cache in fixed-size blocks called "cache lines" or "cache blocks". When the processor needs to read or write a location in main memory, the processor checks the cache. If the memory location is in the cache, this is considered a cache hit, and the processor can read from or write to the cache. If the memory location is not in the cache, this is considered a cache miss, and a new cache entry is created for the data read from the location in main memory. If the cache was previously full before the cache miss, an eviction policy is used to determine which cache entry to delete or rewrite to make room for the new cache entry.

[0012] Cache thrashing involves frequently evicting the "wrong" cache lines such that they are read from and written to external memory more than once. Due to the excessive cache misses, thrashing results in a performance that is worse than the performance without any cache at all. As a result, the cache experiences many cache misses, such that the benefits of the cache are largely offset.

[0013] Cache thrashing may be experienced more in certain data structures such as queues, in which the order of storing and retrieving data is contrary to many of the assumptions in cache management. For example, in many cases, cache management assumes spatial and temporal locality of subsequent data accesses. This is generally a reasonable assumption because data that was most recently stored in the cache is typically more likely to be the data that will be accessed in the near future. However, in a queuing structure, the data that was stored first (the oldest data) is the first to be dequeued (the most recent data request). Thus, due to the first-in, first-out (FIFO) nature of the accesses associated with a queue, a standard cache may experience cache thrashing.

[0014] The managed cache described in this document takes into account the nature of the accesses associated with a queue in order to ensure that thrashing will not be caused regardless of the usage pattern. Thus, generally, each entry can be written to and read from main memory at most once, resulting in zero thrashing with respect to queue data.

[0015] The mechanisms described herein can be used for queues of any size, which provides an advantage over other implementations that attempt to solve cache thrashing by using increased memory size and limiting the queue size. Local on-chip resources are not scalable for large queues, so the queue size is limited to a relatively small number of entries compared to queues located in main memory.

[0016] Figure 1A and Figure 1B illustrates cache thrashing and a mechanism for removing cache thrashing for queuing according to an embodiment. The most recently used (MRU) cache list 100 is used to track cache lines and their recency of last use. Assume the cache has room for four cache lines. The MRU cache list 100 is organized such that the most recently used cache line is at the top and the earliest used cache line is at the bottom. The queue 102 includes data present in cache lines A, B, C, D, and E, with the data in cache line A being the head of the queue and the data in cache line E being at the tail. When the queue 102 is accessed or created, blocks A, B, C, and D are read into the cache. The data in the earliest used block A is at the bottom of the MRU list 100. Thus, when block E is accessed, block A is evicted. Now, when the queue 102 is accessed again in order to process the data stored in the queue 102, the queue 102 is accessed in a FIFO manner, and block A is accessed first. Since block A is not present in the cache, block A is fetched from external memory (e.g., main memory). Block B is the LRU block, so block B is evicted. Each successive block results in a cache miss and a read from external memory. The end result is complete cache thrashing and high overhead in memory operations.

[0017] Figure 1B illustrates a mechanism for removing cache thrashing in queuing operations. Certain cache lines are protected from eviction. Specifically, the cache lines containing the head and tail of the queue are marked such that they cannot be evicted. Since the queue is designed to read from the head and write to the tail, protecting these data portions ensures that thrashing does not occur. Thus, in Figure 1BIn the figure, when queue 102 is read into the cache, the head (block A) is marked as "not evictable". This is depicted by an asterisk ('*') in cache line A of the MRU list 100. When block E is loaded into the cache, cache line A is not evicted due to its protected state. Instead, cache line B is evicted (since it is the second oldest data). Additionally, cache line E is marked as "not evictable". As queue 100 grows, cache line E is accessed and additional data can be written to the address space in cache line E until the address space is full. Then a new cache line is fetched from the external memory, marked as "not evictable", and placed into the cache and MRU list 100. Similarly, when using queue 102 and dequeuing data, cache line A is used. After all of cache line A has been exhausted (e.g., dequeued and processed), then block B is read into the cache and placed into MRU list 100 according to the MRU policy. Since block B is now the head of the queue, block B is marked as "not evictable".

[0018] Generally, based on the recognition that a cache line is the head or tail of queue 102, these cache lines are marked as "protected from eviction". Additionally, when all data is exhausted (all contained entries have been popped from the head), the cache line with the head of queue 102 can be marked as "immediately invalid". This step can be performed so that data is not written back from the cache to the external memory. The only property required to guarantee this zero thrashing mechanism is to have a minimum number of cache lines that is twice the number of queues. In other words, there needs to be at least a cache line for the head and another cache line for the tail of the queue.

[0019] Note that in a queuing application, reading means a pop operation (or dequeue operation), while writing means a push operation (or enqueue operation). Therefore, when a cache line is filled with data, it no longer needs protection. The cache line filled with data can be evicted based on the cache policy, but a new tail is available to eliminate thrashing. Similarly, when the application reads a complete cache line from the cache, this means that the cache line does not contain any interesting information, and the cache line can be immediately invalidated without actually writing anything back to the external memory. A new queue head can be read into the cache and used. This feature has the additional benefit that for a (fully cached) shallow queue, even if queue 102 is continuously written to and read with new data, queue 102 will not access the external memory at all.

[0020] The extended protection of this mechanism protects three cache lines: one for the head of the queue, one for the tail of the queue, and one for the next data after the head of the queue. The third cache line can be used as a prefetch buffer, allowing the process using the queue to move away from the head and continue processing without waiting for the cache line to be fetched from external memory. To implement a dual-cache-line mechanism or a triple-cache-line mechanism, the cache maintains a track of each queue and the position of the cache line that was last read into or written to for each queue.

[0021] Figure 2 FIG. [FIGURE NUMBER] is a block diagram illustrating a queue cache mechanism according to an embodiment. Data can be pushed onto or popped from the queue (e.g., enqueued or dequeued respectively) (operation flow 200). The queue cache 202 includes a cache manager 204 that operates to protect the cache line containing the tail of the queue when data is pushed onto the queue. The cache line containing the tail of the queue is written to the cache store 206 and can be written out to external memory when the cache line is full (operation flow 208).

[0022] Similarly, when data is popped from the queue, the cache manager 204 operates to protect the cache line containing the head of the queue. When the cache line is completely read out due to a queue pop, the cache line can be immediately invalidated. The next cache line that holds the next data in the queue can be read into the cache store (when there is a cache miss) and protected so that the next cache line cannot be swapped out to external memory under the cache management policy.

[0023] In addition, although Figure 2 FIG. [FIGURE NUMBER] illustrates one cache manager 204, it should be understood that multiple cache managers can be used. For example, for each core in a multi-core processor, one cache memory can be used. This may be useful in, for example, an L1 data cache. Alternatively, a single cache manager can be used for a cache shared among several cores, such as an L3 cache. In another embodiment, the cache manager for the queue is separate from other data caches on the chip and can be shared across one or more cores for queue use. The cache lines for long queues can be direct mapped, M-way associative, or fully associative. Except for the direct mapping exception, the cache can use a FIFO replacement policy, random, or some LRU approximation.

[0024] Figure 3A and Figure 3B FIG. [FIGURE NUMBER] is a block diagram illustrating cache management according to an embodiment. In Please note that the "[FIGURE NUMBER]" placeholders in the translation of ID=4 and ID=12 need to be replaced with the actual figure numbers in the original document.Figure 3A In this case, the cache store 300 (e.g., cache store 206) can be configured to store two queues: "Queue 0" and "Queue 1", which are referred to as "Q0" and "Q1" respectively. The cache store 300 includes a plurality of cache lines (represented by blocks in the cache store 300). Queue Q0 includes a head cache line 302 and a tail cache line 304. Unprotected cache lines 306A, 306B, and 306C (collectively 306) represent the middle part of Queue Q0. These unprotected cache lines 306 can be written out to external memory according to a cache management policy to make room for other cache contents. Maintaining the head and tail of Queue Q0 eliminates thrashing of the process using Queue Q0.

[0025] Block 308 is the cache line that was previously the head cache line of Queue Q0, and this block 308 can be empty or invalidated after the processing of Queue Q0 has moved to the current head cache line 302.

[0026] Queue Q1 includes a head cache line 310 and a tail cache line 312, similar to the head cache line and tail cache line for Queue Q0. Unprotected cache lines 314A, 314B, 314C, 314D, and 314E (collectively 314) can be used as regular cache lines according to a cache management policy.

[0027] Figure 3B is another illustration of the cache store 300 that uses three protected cache lines for Queue Q0. The head cache line 302, the tail cache line 304, and the prefetch cache line 316 of Q0 are protected together, which ensures a cache hit for the process using Queue Q0 when the process exhausts the data in the head cache line 302. It should be understood that additional prefetch cache lines can be used, such that more than one prefetch cache line can be available for a given queue. This is a design consideration that sacrifices memory usage in exchange for faster queue processing. Additionally, when space and design permit, multiple queues can use two, three, or more protected cache lines for each queue. For example, Figure 3BExamples include queue Q0 that uses three protected cache lines and queue Q1 that uses two protected cache lines. As another example, queue Q0 can use three protected cache lines, and queue Q1 can also use three protected cache lines. As yet another example, queue Q0 can use three protected cache lines, queue Q1 can also use three protected cache lines, and queue Q2 can use two protected cache lines. As yet another example, queue Q0 can use two protected cache lines, queue Q1 can use three protected cache lines, and queue Q2 can use four protected cache lines. Any number of combinations and permutations of cache line protection can be used, and these combinations and permutations are considered to be within the scope of the present disclosure.

[0028] Figure 4 is a flowchart illustrating a process 400 for implementing a zero - thrashing cache queue manager according to an embodiment. In modern architectures, a memory access from a core through the memory subsystem is a combination of a specific request (e.g., read, write, etc.), the physical memory address required, and possibly data. Data moves around most memory subsystems in 64 - byte quantities called cache lines. When a cache line is copied into a cache entry, the cache entry is filled. A cache entry can include data from the cache line, a memory location, and flags. The memory location can be part of a physical address, a virtual address, an index, etc. In an example, the memory location can include a portion of the following information: a physical address (e.g., the n most significant bits), an index that describes which cache entry the data has been placed in, and a block offset that specifies the byte position within the cache line of a particular cache entry.

[0029] The flags include a valid flag, a dirty flag, and a protected flag. The valid flag is used to indicate whether the cache entry has loaded valid data. At power - on, the hardware sets all valid bits in all cache entries to "invalid". An invalid cache entry can be evicted without writing the data out to external memory (e.g., main memory).

[0030] The dirty bit is used to indicate whether the data in the cache entry has changed since the data was read into the cache from external memory. Dirty data is a cache line that has been changed by the processor and whose data has not yet propagated back to external memory.

[0031] The protected flag is used to indicate cache entries that should not be evicted. The protected flag may also be referred to as the "do not evict" (DNE) flag or the "do not evict" (DNE) bit. In the systems described in this document, the protected flag is used to protect cache entries that contain the head and tail of a queue. Additional cache lines may be protected, such as the prefetch portion that is the next one after the current head of the queue in the queue.

[0032] During a memory read operation, if the cache has the requested physical address in a cache entry, the cache returns the data. If not, the cache requests the data from deeper in the memory subsystem and evicts some cache entry to make room. If the evicted cache entry has been modified, that cache entry must be written to the deeper memory subsystem as part of the eviction. This means that the read stream may slow down because earlier write sets must be pushed deeper into the memory subsystem.

[0033] During a memory write operation, if the cache does not have the cache line in a cache entry, the cache reads the cache line from deeper in the memory subsystem. The cache evicts some other physical address from its cache entry to make room for the cache line. Although the write may only change some of the 64 bytes, the read is necessary to obtain all 64 of those bytes. When the cache entry is written to for the first time, the cache entries for that physical address in all other caches are invalidated. This action makes the first write to a cache entry more costly than later writes.

[0034] Using the implementation described here for a queue cache, cache misses are unlikely. Instead, when a queue is created by a process and the first element is added to the queue, the cache line containing the data for that element is marked as "do not evict". This cache entry represents both the head and the tail of the queue because it has only one element and exists in only one cache line. As the queue expands across several cache lines, there may subsequently be a separate cache line for the head and another for the tail. Since the head cache line and the tail cache line are never evicted from the cache, there are never cache misses, and thus, the zero thrashing property of a managed queue cache.

[0035] Now turn to Figure 4In the flowchart, at 402, a memory access request for the queue is received. At decision block 404, it is determined whether the memory access request is a write operation (e.g., a push operation or an enqueue operation) or a read operation (e.g., a pop operation or a dequeue operation). When the memory access request is a write operation, at operation 406, the data in the cache line is updated with the data being enqueued, and the cache entry is marked as "dirty" due to the modified state (operation 408). The cache entry including the tail of the queue is marked as "non-evictable" (operation 410). The cache entry can be marked as "non-evictable" after each write operation. Other conventional cache operations can occur to store memory address information, mark the cache entry as valid, invalidate other cache entries with the same physical address, and so on. Process 400 returns to operation 402 to receive the next memory access request for the queue.

[0036] When the memory access request is a read operation, the cache entry for the head of the queue is accessed (operation 412). If the read operation is accessing the last data entry in the cache line, the cache entry is depleted, and the cache entry can be immediately marked as "invalid" (operation 414).

[0037] In an alternative implementation, a pre-fetched queue head is maintained. As the queue grows over time, the next queue head cache entry can be identified and marked as "non-evictable" as part of the memory allocation and cache management process. However, when a queue head cache entry is depleted, there is no guarantee that the third cache line in order is in the cache. Thus, to maintain the pre-fetched queue head after the cache entry is depleted, a new pre-fetched queue head can be read from external memory into the cache and marked as "non-evictable". Depending on whether the pre-fetched head cache line is maintained, operation 416 can be optionally performed. Operation 416 can read another cache line into the cache and mark the cache line as "non-evictable" for pre-fetching the next queue head "on deck". If the next queue head cache line is the same as the cache line for the tail (e.g., the multiple elements have been dequeued such that the queue is contained in two cache lines, one for the current head and one for the tail), then operation 416 can terminate without reading data into the cache.

[0038] Figure 5FIG. 0 is a flowchart illustrating a method 500 for implementing a zero-latency cache queue manager according to an embodiment. At 502, a memory access request for a queue is received. The memory access can be received in hardware, e.g., at a cache controller on a die having one or more processor cores. Alternatively, the memory access can be received at a higher level in hardware or software, such as at a memory management unit, a device driver, an operating system library, etc.

[0039] At 504, when the memory access request is to add data to the queue, the data is written to a queue tail cache line in the cache. Protect the queue tail cache line from being evicted from the cache.

[0040] In an embodiment, a flag in the cache entry is used to protect the queue tail cache line from being evicted from the cache, where the cache entry contains the queue tail cache line. In a further embodiment, method 500 includes: setting the flag in the cache entry to protect the queue tail cache line from being evicted each time data is written to the queue tail cache line. In a related embodiment, method 500 includes: marking the cache entry containing the queue tail cache line as dirty as a result of writing data to the queue tail cache line.

[0041] In one embodiment, method 500 includes: when reading data from the current queue head cache line depletes the data in the current queue head cache line, marking the cache entry as invalid.

[0042] At 506, when the memory access request is to remove data from the queue, data is read from the current queue head cache line in the cache. Protect the current queue head cache line from being evicted from the cache.

[0043] In an embodiment, method 500 includes maintaining a prefetch queue head cache line, which is the next queue head cache line in the queue relative to the current queue head cache line. In a further embodiment, maintaining the prefetch queue head cache line includes: marking the prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, where the prefetch queue head cache entry contains the prefetch queue head cache line. In a related embodiment, maintaining the prefetch queue head cache line includes: obtaining the prefetch queue head cache line from external memory; and marking the prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, where the prefetch queue head cache entry contains the prefetch queue head cache line.

[0044] Each embodiment may be implemented in one or a combination of hardware, firmware, and software. Embodiments may also be implemented as instructions stored on a machine-readable storage device that can be read and executed by at least one processor to perform the operations described herein. The machine-readable storage device may include any non-transitory mechanism for storing information in a form readable by a machine, such as a computer. For example, the machine-readable storage device may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, and other storage devices and media.

[0045] The processor subsystem may be used to execute instructions on the machine-readable medium. The processor subsystem may include one or more processors, each having one or more cores. Additionally, the processor subsystem may be disposed on one or more physical devices. The processor subsystem may include one or more specialized processors, such as a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), or a fixed-function processor.

[0046] Examples as described herein may include logic or multiple components, modules, or mechanisms, or may operate on logic or multiple components, modules, or mechanisms. Each module may be hardware, software, or firmware that is communicatively coupled to one or more processors to implement the operations described herein. Each module may be a hardware module, and as such, each module may be considered a tangible entity capable of performing specified operations and configured or arranged in a particular manner. In an example, circuitry (e.g., internally or relative to external entities such as other circuitry) may be arranged as a module in a prescribed manner. In an example, all or part of one or more computer systems (e.g., stand-alone client or server computer systems) or one or more hardware processors may be configured by firmware or software (e.g., instructions, an application portion, or an application) to operate as a module for performing prescribed operations. In an example, the software may reside on a machine-readable medium. In an example, when executed by the underlying hardware of the module, the software causes the hardware to perform the prescribed operations. Thus, the term hardware module is understood to encompass a tangible entity that is physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., transiently) configured (e.g., programmed) to operate in the specified manner or to perform part or all of any of the operations described herein. Considering examples in which modules are configured temporarily, each of these modules need not be instantiated at any one time. For example, in the case where a module includes a general hardware processor configured using software, the general hardware processor may be configured as respective different modules at different times. The software may configure the hardware processor accordingly, e.g., to constitute a particular module at one instance of time and a different module at a different instance of time. Each module may also be a software or firmware module that operates to perform the methods described herein.

[0047] As used in this document, circuitry or a circuit may include, for example, alone or in any combination: hardwired circuitry; programmable circuitry such as a computer processor including one or more separate instruction processing cores; state machine circuitry; and / or firmware storing instructions executed by the programmable circuitry. The circuitry, circuit system, or module may be embodied, jointly or separately, as a circuit system forming part of a larger system, such as, for example, an integrated circuit (IC), a system on a chip (SoC), a desktop computer, a laptop computer, a tablet computer, a server, a smart phone, and so on.

[0048] As used in any embodiment herein, the term “logic” may refer to firmware and / or circuitry configured to perform any of the foregoing operations. The firmware may be embodied as code, instructions, or an instruction set and / or data hard-coded (e.g., non-volatile) in a memory device and / or circuitry.

[0049] As used in any embodiment herein, "circuitry" can include, for example, hardwired circuitry, programmable circuitry, state machine circuitry, logic, and / or firmware that stores instructions executed by the programmable circuitry, either alone or in any combination. The circuitry can be embodied as an integrated circuit, such as an integrated circuit chip. In some embodiments, the circuitry can be at least partially formed by a processor circuitry that executes code and / or instruction sets (e.g., software, firmware, etc.) corresponding to the functions described herein, thereby transforming a general-purpose processor into a special-purpose processing environment to perform one or more of the operations described herein. In some embodiments, the processor circuitry can be embodied as a stand-alone integrated circuit or can be incorporated as one of several components into an integrated circuit. In some embodiments, the various components and circuitry of a node or other system can be combined in a system-on-chip (SoC) architecture.

[0050] Figure 6 FIG. 600 is a block diagram of an example of a machine in the form of a computer system in which a set of executable instructions or sequence of instructions can cause the machine to perform any one of the methods discussed herein. In alternative embodiments, the machine operates as a stand-alone device or can be connected (e.g., networked) to other machines. In a networked deployment, the machine can operate as a server or a client in a server-client network environment, or it can act as a peer machine in a peer-to-peer (or distributed) network environment. The machine can be a head-mounted display, a wearable device, a personal computer (PC), a tablet PC, a hybrid tablet, a personal digital assistant (PDA), a mobile phone, or any machine capable of executing instructions (sequentially or otherwise) specifying actions to be taken by that machine. Further, while only a single machine is shown, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein. Similarly, the term "processor-based system" shall be taken to include any collection of one or more machines controlled or operated by a processor (e.g., a computer) to individually or jointly execute instructions to perform any one or more of the methods discussed herein.

[0051] The example computer system 600 includes at least one processor 602 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both, a processor core, a computing node, etc.), a main memory 604, and a static memory 606, and these components communicate with each other via a link 608 (e.g., a bus). The computer system 600 may further include a video display unit 610, an alphanumeric input device 612 (e.g., a keyboard), and a user interface (UI) navigation device 614 (e.g., a mouse). In one embodiment, the video display unit 610, the input device 612, and the UI navigation device 614 are incorporated into a touchscreen display. The computer system 600 may additionally include a storage device 616 (e.g., a drive unit), a signal generation device 618 (e.g., a speaker), a network interface device 620, and one or more sensors (not shown), such as a global positioning system (GPS) sensor, a compass, an accelerometer, a gyroscope, a magnetometer, or other sensors.

[0052] The storage device 616 includes a machine-readable medium 622, on which is stored a set or multiple sets of data structures and instructions 624 (e.g., software), and the set or multiple sets of data structures and instructions 624 embody one or more of the methods or functions described herein, or are utilized by one or more of the methods or functions described herein. During the execution of the instructions 624 by the computer system 600, the instructions 624 may also reside entirely or at least partially within the main memory 604, the static memory 606, and / or within the processor 602, and the main memory 604, the static memory 606, and the processor 602 also constitute machine-readable media.

[0053] Although the machine-readable medium 622 is illustrated as a single medium in the example embodiment, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database and / or associated cache and server) that store one or more instructions 624. The term "machine-readable medium" should also be considered to include any tangible medium that can store, encode, or carry instructions that are executed by a machine and cause the machine to execute any one or more of the methods of the present disclosure, or that can store, encode, or carry data structures utilized by or associated with such instructions. The term "machine-readable medium" should accordingly be considered to include, but not be limited to: solid-state memories as well as optical and magnetic media. Specific examples of machine-readable media include non-volatile memories, as examples, including but not limited to: semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0054] A transmission medium can be used to further transmit or receive instructions 624 over a communication network 626 via a network interface device 620 using any one of several well-known transmission protocols (e.g., HTTP). Examples of communication networks include: local area networks (LANs), wide area networks (WANs), the Internet, mobile phone networks, plain old telephone (POTS) networks, and wireless data networks (e.g., Bluetooth, Wi-Fi, 3G, and 4G LTE / LTE-A, 5G, DSRC, or WiMAX networks). The term "transmission medium" shall be considered to include any non-transitory medium that can store, encode, or carry instructions executable by a machine, and includes digital or analog communication signals or other non-transitory media for facilitating the communication of such software.

[0055] Additional Notes and Examples:

[0056] Example 1 is a system for implementing a zero-jitter cache queue manager, the system comprising: a memory including instructions and a processor subsystem that, when executing the instructions, causes the processor subsystem to perform operations including the steps of: receiving a memory access request for a queue; when the memory access request is to add data to the queue, writing the data to a queue tail cache line in the cache and protecting the queue tail cache line from being evicted from the cache; and when the memory access request is to remove data from the queue, reading the data from the current queue head cache line in the cache, the current queue head cache line being protected from being evicted from the cache.

[0057] In Example 2, the subject matter of Example 1 includes, wherein a flag in a cache entry is used to protect the queue tail cache line from being evicted from the cache, the cache entry containing the queue tail cache line.

[0058] In Example 3, the subject matter of Example 2 includes, setting the flag in the cache entry to protect the queue tail cache line from being evicted each time data is written to the queue tail cache line.

[0059] In Example 4, the subject matter of Examples 2-3 includes, marking the cache entry containing the queue tail cache line as dirty as a result of writing data to the queue tail cache line.

[0060] In Example 5, the subject matter of Examples 2-4 includes, when reading data from the current queue head cache line depletes the data in the current queue head cache line, marking the cache entry as invalid.

[0061] In Example 6, the subject matter of Examples 1 - 5 includes maintaining a prefetch queue head cache line that is the next queue head cache line in the queue relative to the current queue head cache line.

[0062] In Example 7, the subject matter of Example 6 includes maintaining the prefetch queue head cache line by: marking a prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, where the prefetch queue head cache entry contains the prefetch queue head cache line.

[0063] In Example 8, the subject matter of Examples 6 - 7 includes maintaining the prefetch queue head cache line by: obtaining the prefetch queue head cache line from external memory; and marking a prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, where the prefetch queue head cache entry contains the prefetch queue head cache line.

[0064] Example 9 is a method for implementing a zero - thrashing cache queue manager, the method including: receiving a memory access request to a queue; when the memory access request is to add data to the queue, writing the data to a queue tail cache line in the cache and protecting the queue tail cache line from being evicted from the cache; and when the memory access request is to remove data from the queue, reading data from the current queue head cache line in the cache, where the current queue head cache line is protected from being evicted from the cache.

[0065] In Example 10, the subject matter of Example 9 includes protecting the queue tail cache line from being evicted from the cache using a flag in a cache entry that contains the queue tail cache line.

[0066] In Example 11, the subject matter of Example 10 includes setting the flag in the cache entry to protect the queue tail cache line from being evicted each time data is written to the queue tail cache line.

[0067] In Example 12, the subject matter of Examples 10 - 11 includes marking the cache entry that contains the queue tail cache line as dirty as a result of writing data to the queue tail cache line.

[0068] In Example 13, the subject matter of Examples 10 - 12 includes marking the cache entry as invalid when reading data from the current queue head cache line depletes the data in the current queue head cache line.

[0069] In Example 14, the subject matter of Examples 9-13 includes maintaining a prefetch queue head cache line that is the next queue head cache line in the queue relative to the current queue head cache line.

[0070] In Example 15, the subject matter of Example 14 includes, where maintaining the prefetch queue head cache line includes marking a prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, the prefetch queue head cache entry containing the prefetch queue head cache line.

[0071] In Example 16, the subject matter of Examples 14-15 includes, where maintaining the prefetch queue head cache line includes: obtaining the prefetch queue head cache line from an external memory; and marking a prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, the prefetch queue head cache entry containing the prefetch queue head cache line.

[0072] Example 17 is at least one machine-readable medium including instructions that, when executed by a machine, cause the machine to perform the operations of any one of the methods of Examples 9-16.

[0073] Example 18 is an apparatus including a device for performing any one of the methods of Examples 9-16.

[0074] Example 19 is a device for implementing a zero-thrashing cache queue manager, the device including: means for receiving a memory access request to a queue; means for writing data to a queue tail cache line in the cache and protecting the queue tail cache line from being evicted from the cache when the memory access request is to add data to the queue; and means for reading data from a current queue head cache line in the cache when the memory access request is to remove data from the queue, the current queue head cache line being protected from being evicted from the cache.

[0075] In Example 20, the subject matter of Example 19 includes, where a flag in a cache entry is used to protect the queue tail cache line from being evicted from the cache, the cache entry containing the queue tail cache line.

[0076] In Example 21, the subject matter of Example 20 includes means for setting the flag in the cache entry to protect the queue tail cache line from being evicted each time data is written to the queue tail cache line.

[0077] In Example 22, the subject matter of Examples 20 - 21 includes apparatus for marking a cache entry that includes the queue tail cache line as dirty as a result of writing data to the queue tail cache line.

[0078] In Example 23, the subject matter of Examples 20 - 22 includes apparatus for invalidating a cache entry when reading data from the current queue head cache line depletes the data in the current queue head cache line.

[0079] In Example 24, the subject matter of Examples 19 - 23 includes apparatus for maintaining a prefetch queue head cache line that is the next queue head cache line in the queue relative to the current queue head cache line.

[0080] In Example 25, the subject matter of Example 24 includes wherein the apparatus for maintaining the prefetch queue head cache line includes apparatus for marking a prefetch queue head cache entry to protect the prefetch queue head cache line from eviction from the cache, the prefetch queue head cache entry including the prefetch queue head cache line.

[0081] In Example 26, the subject matter of Examples 24 - 25 includes wherein the apparatus for maintaining the prefetch queue head cache line includes: apparatus for obtaining the prefetch queue head cache line from an external memory; and apparatus for marking a prefetch queue head cache entry to protect the prefetch queue head cache line from eviction from the cache, the prefetch queue head cache entry including the prefetch queue head cache line.

[0082] Example 27 is at least one machine - readable medium including instructions for implementing a zero - thrashing cache queue manager, the instructions when executed by a machine causing the machine to perform operations including the steps of: receiving a memory access request for a queue; when the memory access request is to add data to the queue, writing the data to a queue tail cache line in the cache and protecting the queue tail cache line from eviction from the cache; and when the memory access request is to remove data from the queue, reading data from a current queue head cache line in the cache, the current queue head cache line being protected from eviction from the cache.

[0083] In Example 28, the subject matter of Example 27 includes wherein a flag in a cache entry is used to protect the queue tail cache line from eviction from the cache, the cache entry including the queue tail cache line.

[0084] In Example 29, the subject matter of Example 28 includes an instruction that causes the machine to perform the following operations: setting the flag in the cache entry to protect the queue tail cache line from being evicted each time data is written to the queue tail cache line.

[0085] In Example 30, the subject matter of Examples 28-29 includes an instruction that causes the machine to perform the following operations: marking the cache entry containing the queue tail cache line as dirty as a result of writing data to the queue tail cache line.

[0086] In Example 31, the subject matter of Examples 28-30 includes an instruction that causes the machine to perform the following operations: marking the cache entry as invalid when reading data from the current queue head cache line depletes the data in the current queue head cache line.

[0087] In Example 32, the subject matter of Examples 27-31 includes an instruction that causes the machine to perform the following operations: means for maintaining a prefetch queue head cache line that is the next queue head cache line in the queue relative to the current queue head cache line.

[0088] In Example 33, the subject matter of Example 32 includes where maintaining the prefetch queue head cache line includes: marking the prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, the prefetch queue head cache entry containing the prefetch queue head cache line.

[0089] In Example 34, the subject matter of Examples 32-33 includes where maintaining the prefetch queue head cache line includes: obtaining the prefetch queue head cache line from external memory; and marking the prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, the prefetch queue head cache entry containing the prefetch queue head cache line.

[0090] Example 35 is at least one machine-readable medium including instructions that, when executed by a processor subsystem, cause the processor subsystem to perform operations to implement any one of Examples 1-34.

[0091] Example 36 is a device including means for implementing any one of Examples 1-34.

[0092] Example 37 is a system for implementing any one of Examples 1-34.

[0093] Example 38 is a method for implementing any of Examples 1-34.

[0094] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The accompanying drawings illustrate specific embodiments that may be practiced. These embodiments are also referred to herein as "examples". Such examples may include elements other than those shown or described. However, examples including the elements shown or described are also contemplated. In addition, examples using any combination or arrangement of those elements (or one or more aspects thereof) shown or described herein, or referring to a particular example (or one or more aspects thereof) shown or described herein, or referring to other examples (or one or more aspects thereof) shown or described herein are also contemplated.

[0095] Publications, patents, and patent documents cited in this document are incorporated herein by reference in their entirety as if individually incorporated by reference. In the case of inconsistent usage between this document and those documents incorporated by reference, the usage in the incorporated (s) cited document is supplementary to the usage in this document; for irreconcilable inconsistencies, the usage in this document prevails.

[0096] In this document, as is common in patent documents, the term "a" is used to include one or more than one, independently of any other instance or usage of "at least one" or "one or more". In this document, the term "or" is used to mean a non-exclusive or, such that unless otherwise indicated, "A or B" includes "A but not B", "B but not A", and "A and B". In the appended claims, the terms "including" and "in which" are used as the ordinary English equivalents of the corresponding terms "comprising" and "wherein". In addition, in the appended claims, the terms "including" and "comprising" are open-ended, that is, a system, apparatus, article, or process that includes elements other than those recited after such terms in the claim is still considered to fall within the scope of that claim. In addition, in the appended claims, the terms "first", "second", and "third", etc. are used only as labels and are not intended to indicate a numerical order of their objects.

[0097] The above description is intended to be illustrative and not restrictive. For example, the above-described examples (or one or more aspects thereof) may be used in conjunction with other embodiments. That is, one of ordinary skill in the art may use other embodiments by reviewing the above description. The abstract allows the reader to quickly ascertain the nature of the technical disclosure. It should be understood that submitting the abstract is not intended to limit or interpret the scope or meaning of the claims. Further, in the above Detailed Description, the various features may be grouped together in order to streamline the disclosure. However, the claims may not recite every feature disclosed herein, as the features of an embodiment may be a subset of the features described. Additionally, an embodiment may include fewer features than those disclosed in a particular example. Accordingly, the appended claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate embodiment. The scope of the embodiments disclosed herein should be determined with reference to the appended claims along with the full scope of equivalents to which such claims are entitled.

Claims

1. A system for implementing a zero - thrashing cache queue manager, the system comprises: A processor subsystem, the processor subsystem is configured to: Receive a memory access request for a queue; When the memory access request is to add data to the queue, write the data to a queue - tail cache line in the cache, the queue - tail cache line is protected from being evicted from the cache by: marking the queue - tail cache line to indicate that the cache entry containing the queue - tail cache line should not be evicted; and When the memory access request is to remove data from the queue, read the data from the current queue - head cache line in the cache, the current queue - head cache line is protected from being evicted from the cache by: marking the current queue - head cache line to indicate that the cache entry containing the current queue - head cache line should not be evicted.

2. The system according to claim 1, wherein, A flag in the cache entry is used to protect the queue - tail cache line from being evicted from the cache, the cache entry contains the queue - tail cache line.

3. The system according to claim 2, wherein, The processor subsystem is configured to set the flag in the cache entry to protect the queue - tail cache line from being evicted each time data is written to the queue - tail cache line.

4. The system according to claim 2, wherein, As a result of writing data to the queue - tail cache line, the processor subsystem is configured to mark the cache entry containing the queue - tail cache line as dirty.

5. The system according to claim 2, wherein, When reading data from the current queue - head cache line depletes the data in the current queue - head cache line, the processor subsystem is configured to mark the cache entry as invalid.

6. The system according to claim 1, wherein, The processor subsystem is configured to maintain a pre - fetch queue - head cache line, the pre - fetch queue - head cache line is the next queue - head cache line in the queue relative to the current queue - head cache line.

7. The system according to claim 6, wherein, To maintain the pre - fetch queue - head cache line, the processor subsystem is configured to: mark the pre - fetch queue - head cache entry to protect the pre - fetch queue - head cache line from being evicted from the cache, the pre - fetch queue - head cache entry contains the pre - fetch queue - head cache line.

8. The system according to claim 6, wherein, To maintain the pre - fetch queue - head cache line, the processor subsystem is configured to: obtain the pre - fetch queue - head cache line from external memory; and mark the pre - fetch queue - head cache entry to protect the pre - fetch queue - head cache line from being evicted from the cache, the pre - fetch queue - head cache entry contains the pre - fetch queue - head cache line.

9. A method for implementing a zero-thrashing cache queue manager, the method comprises: Receiving a memory access request for a queue; When the memory access request is to add data to the queue, writing the data to a queue tail cache line in the cache, the queue tail cache line being protected from eviction from the cache by: marking the queue tail cache line to indicate that the cache entry containing the queue tail cache line should not be evicted; and When the memory access request is to remove data from the queue, reading the data from the current queue head cache line in the cache, the current queue head cache line being protected from eviction from the cache by: marking the current queue head cache line to indicate that the cache entry containing the current queue head cache line should not be evicted.

10. The method according to claim 9, wherein, A flag in the cache entry is used to protect the queue tail cache line from being evicted from the cache, the cache entry containing the queue tail cache line.

11. The method according to claim 10, further comprises: Setting the flag in the cache entry to protect the queue tail cache line from being evicted each time data is written to the queue tail cache line.

12. The method according to claim 10, further comprises: Marking the cache entry containing the queue tail cache line as dirty as a result of writing data to the queue tail cache line.

13. The method according to claim 10, further comprises: When reading data from the current queue head cache line depletes the data in the current queue head cache line, marking the cache entry as invalid.

14. The method according to claim 9, further comprises: Maintaining a prefetch queue head cache line, the prefetch queue head cache line being the next queue head cache line in the queue relative to the current queue head cache line.

15. The method according to claim 14, wherein, Maintaining the prefetch queue head cache line comprises: marking a prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, the prefetch queue head cache entry containing the prefetch queue head cache line.

16. The method according to claim 14, wherein, Maintaining the prefetch queue head cache line comprises: obtaining the prefetch queue head cache line from external memory; and marking a prefetch queue head cache entry to protect the prefetch queue head cache line from being evicted from the cache, the prefetch queue head cache entry containing the prefetch queue head cache line.

17. At least one machine-readable medium comprising instructions that, when executed by a machine, cause the machine to perform the operations of the method according to any one of claims 9-16.

18. A device for implementing a zero-jitter cache queue manager, the device comprising: means for receiving a memory access request for a queue; means for writing data to a queue tail cache line in the cache when the memory access request is to add data to the queue, the queue tail cache line being protected from eviction from the cache by: tagging the queue tail cache line to indicate that the cache entry containing the queue tail cache line should not be evicted; and means for reading data from a current queue head cache line in the cache when the memory access request is to remove data from the queue, the current queue head cache line being protected from eviction from the cache by: tagging the current queue head cache line to indicate that the cache entry containing the current queue head cache line should not be evicted.

Citation Information

Patent Citations

  • High-performance queue implementing of multiprocessor system

    CN101346692A

  • Method for dynamically dividing shared high-speed caches and circuit

    CN102609362A