Probabilistic tracking of memory access

A probabilistic counter increment method in memory access tracking reduces traffic and resource usage, enabling efficient memory page migration to optimize memory performance.

GB2636849APending Publication Date: 2025-07-02ARM LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2023019987
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-07-02

AI Technical Summary

Technical Problem

Existing memory access tracking methods are processing-intensive due to the need to count every access to every memory region, leading to inefficient traffic management.

Method used

A probabilistic approach is employed to increment counters associated with memory addresses, reducing the need for continuous updates by only incrementing counters based on a probabilistic mechanism, such as random number comparisons with thresholds, thereby minimizing traffic and allowing efficient tracking of memory access frequency.

Benefits of technology

This method effectively reduces traffic and resource utilization while accurately determining memory access frequency, enabling efficient migration of memory pages to faster or logically closer locations, thus optimizing memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

There is provided an apparatus that includes storage circuitry for storing one or more addresses and one or more counters each associated with one of the one or more addresses. Receive circuitry recei
Need to check novelty before this filing date? Find Prior Art

Description

The present technique relates to data processing. In particular, the present technique might have relevance to the field of memory systems. It is desirable to be able to determine how frequently accessed certain regions or blocks of memory are. This makes it possible to, for instance, promote more frequently accessed memory regions to faster memory and potentially to move less frequently accessed memory regions to slower memory. A difficulty that is encountered is that counting every access to every region of memory is processing / traffic intensive - every access would involve a counter being updated. Viewed from a first example configuration, there is provided an apparatus comprising: storage circuitry configured to store one or more addresses and one or more counters each associated with one of the one or more addresses; and receive circuitry configured to receive a request comprising an address, wherein in response to receiving the request, one of the one or more counters in the storage circuitry that is associated with the address that is in the request is incremented probabilistically. Viewed from a second example configuration, there is provided a data processing method comprising: storing one or more addresses and one or more counters each associated with one of the one or more addresses; receiving a request comprising an address; and in response to receiving the request, one of the one or more counters that is associated with the address that is in the request is incremented probabilistically. Viewed from a third example configuration, there is provided a computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising: data structures configured to store one or more addresses and one or more counters each associated with one of the one or more addresses; and receive program logic configured to receive a request comprising an address, wherein in response to receiving the request, one of the one or more counters in the data structures that is associated with the address that is in the request is incremented probabilistically. The present technique will be described further, by way of example only, with reference to embodiments thereof as illustrated in the accompanying drawings, in which: Figure 1 schematically illustrates an example of a system in accordance with some examples; Figure 2 shows an example of the claimed apparatus in accordance with some examples; Figure 3 shows a way of probabilistically increasing the counters based on random numbers; Figure 4 shows an alternative arrangement that allows for probabilistically increasing counters based on random sampling; Figure 5 shows two example ways in which sampling can occur; Figure 6 illustrates, in the form of a flow chart, a method for altering the sampling rate; and Figure 7 illustrates the process of determining migration in more detail; and Figure 8 illustrates a simulator implementation that may be used. Before discussing the embodiments with reference to the accompanying figures, the following description of embodiments and associated advantages is provided. In accordance with one example configuration there is provided an apparatus comprising: storage circuitry configured to store one or more addresses and one or more counters each associated with one of the one or more addresses; and receive circuitry configured to receive a request comprising an address, wherein in response to receiving the request, one of the one or more counters in the storage circuitry that is associated with the address that is in the request is incremented probabilistically. By determining how frequently accessed a particular section (e.g. a page) of memory is, it is possible to make decisions regarding how that memory should be used. In practice, keeping a perfect track of exactly how often each region of memory is accessed can be inefficient due to the increased traffic. The present technique therefore proposes a probabilistic increment in the counter that indicates how ‘hot’ or frequently accessed an area of memory is based on requests relating to that memory (e.g. requests that specify an address used to access the area of memory). By doing this probabilistically, not every request always results in an increment of the associated counter (e.g. for the memory block that contains or is addressed by the address given in the request). Consequently, counter updates are not continually performed and so an indication of access frequency can be made without significant increase in traffic. This is in contrast to a situation in which every request is logged. In some examples, the request is a translation request from the address in a first address domain to a translated address in a second address domain; and the storage circuitry is configured to store one or more translated addresses, each associated with one of the one or more addresses. In such examples, the storage circuitry (and / or the apparatus itself) acts as a translation looksaide buffer (TLB) that provides translations between address domains. Since the TLB provides entries relating to pages of memory, this is a practical way of determining access frequency for blocks of memory when the blocks of memory are at a page level of granularity. By providing the data in the TLB, the counters can be updated more quickly and efficiently than if the counters were stored in something like main memory. Of course, in other embodiments, the apparatus may take the form of a dedicated data structure that might sit, for instance, in the load / store unit. The memory blocks need not be page sized and count instead be the size of, for instance, a cache line. In some examples, the second address domain is either a physical address domain or an intermediate physical address domain; the first address domain is either an intermediate physical address domain or a virtual address domain; and the first address domain and the second address domain are different. Virtual address domains may be used within software in order to give a particular view of memory that may not correspond with how the memory is physically arranged. This makes it possible to expand the range of physical memory by paging the memory in and out of a backing store (e.g. a hard disk) when it is not being used. It also makes it possible to provide memory security by providing a virtual memory domain that is unique to an operating system or application. Since the view of memory does not include memory areas that are provided to other operating systems or applications, it is significantly harder for an application or operating system to access memory that has not been assigned to it. Intermediate physical addresses are provided in order to establish a multi-stage translation mechanism. This might be used in more complicated systems such as hypervisors where multiple levels of memory translation are performed in order to hide memory from both operating systems and applications (for instance). In some examples, the one or more counters are incremented probabilistically in response to a randomly generated number being compared to a threshold value. One way in which the probabilistic increment can take place is by generating a random number. If the random number meets some threshold (which may be dependent on the value of the counter) then the counter is incremented. For example, when a request for a translation of an address X is received, the probability with which the counter is incremented from 0 to 1 could be 0.25 (e.g. a randomly generated number between 0 and 1 is generated, and the threshold is 0.25). Statistically, one might expect four requests for translations (or accesses) to address X to be required for the counter to be incremented in this way and so a counter value of' 1 ’ could be equated to four accesses without a counter being incremented four times (even though the four accesses is not guaranteed). The skilled person will appreciate that the comparison could be that the randomly generated number is less than the threshold, but other comparisons may also be possible. The counters themselves could be stored in memory. In some examples, the counters are stored in the associated page table entries. In some examples, the counters can be cached in entries of the TLB (or indeed a dedicated cache) rather than in the page table entries in memory themselves. Such caching might take place until the counter is to be incremented, at which point the counter value can be ‘written back’ to the official page table entry (e.g. in memory). In some examples the one or more counters are N-bit counters, where N is an integer greater than one. Clearly with an N-bit counter, the maximum number that can be represented by the counter is equal to 2N - 1. The probability with which the counter is incremented at each stage, combined with the value of N indicate the maximum number of accesses that can be probabilistically represented. In some examples, the threshold value decreases as a value of the one or more counters increases. That is, as the counter increases it becomes harder (it is less likely) that a further counter increase will occur. Note that there does not necessarily need to be a fixed pattern with respect to the probabilities. For instance, the increase from 0 to 1 could be automatic (that is, as soon as the memory portion is accessed, the counter is incremented) whereas an increase from 1 to 2 could be 1 / 2A6 and an increase from 2 to 3 could be 1 / 2A8. In some examples, a rate at which the threshold value decreases as the value of the one or more counters increases is greater than linear. That is, the change in probability from N to N+l is not the same amount as the change in probability from N+l to N+2, i.e. P(N+1)-P(N) is not the same as P(N+2)-P(N+l). For example, the probability with which each counter increment occurs could decrease by 0.25 at each successive value. That is, if an increase from 0 to 1 has a probability of 0.25 (1 / 2A2) then the probability of an increase from 1 to 2 could be 0.0625 (1 / 2A4). Furthermore, the probability of an increase from 2 to 3 could be 0.015625 (1 / 2A6). In this case, the change in probability is exponential. In some examples, the apparatus comprises: promotion circuitry configured to migrate memory pages having a physical address whose corresponding value of the one or more counters is above a predetermined value. When the counter for a memory address reaches a predetermined value it can be assumed that the page encompassing that address is ‘hot’ and therefore frequently accessed. The page can therefore be ‘promoted’ resulting in the page being migrated to different circuitry. Note that the promotion circuitry could include general purpose processing circuitry that executes software to perform the above functionality and migrates memory pages using either a general purpose CPU or a dedicated data movement engine. In some examples, the promotion circuitry is configured to migrate the memory pages from regular-access memory to frequent-access memory which is faster or logically closer to a processor than the regular-access memory. The frequent-access memory could be faster in the sense that it has a lower latency with the processor so that demanded memory pages can be returned more quickly. Typically, memory that is logically closer to the processor has a lower latency and can therefore return demanded memory pages more quickly that memory further from the processor. In some examples, the one or more counters are incremented probabilistically as a consequence of the request falling within a sample of requests received by the apparatus. Rather than considering absolutely all translation requests, the apparatus may consider a ‘sample’ of all translation requests. Each of the requests within that sample are considered and used to update the counter as appropriate. This is probabilistic in the sense that there is only a probability that a given access will appear within the sample. However, it is expected that the sample will be representative of all accesses. In some examples, a size of the sample of the requests is one in every M of the requests, where M is an integer greater than one. There are a number of ways of taking a sample. In these examples, the sample is considered to comprise every M’th translation request. In this way, the requests are distributed over an extended period. In some examples, a size of the sample of the requests is defined by a window in the requests. In these examples, the samples are clustered together. This makes it possible to process the sample without waiting for an extended period of time to elapse. However, using such a window can sometimes be susceptible to bias if a particular task was taking place across the sampling period as the sample would be influenced by the task taking place. The window size could be defined by a number of accesses or by a time period. In some example, the size of the sample of the requests is dynamically changeable each epoch. The epoch defines the point at which a reassessment is to be made regarding where data is to be migrated to. Typically it indicates the start of a sampling period (a resting period follows the sampling period so that sampling is not always being performed, which would cause a significant increase in activity in the memory system) In some examples, promotion circuitry configured to migrate memory pages having a physical address whose corresponding value of the one or more counters is above a predetermined value. When the counter for a memory address reaches a predetermined value it can be assumed that the page encompassing that address is ‘hof and therefore frequently accessed. The page can therefore be ‘promoted’ resulting in the page being migrated to different circuitry. In some examples, the promotion circuitry is configured to migrate the memory pages from regular-access memory to frequent-access memory which is faster or logically closer to a processor than the regular-access memory. The frequent-access memory could be faster in the sense that it has a lower latency with the processor so that demanded memory pages can be returned more quickly. Typically, memory that is closer to the processor has a lower latency and can therefore return demanded memory pages more quickly that memory further from the processor. In some examples, the size of the sample of the requests is set such that a number of the memory pages having a physical address in the storage circuitry whose corresponding value of the counter is above the threshold is equal to or less than a number of memory pages that can be stored to the frequent-access memory. The sample size can be set (or refined over a number of epochs / sample periods) so that the number of ‘hot’ memory pages discovered roughly corresponds with the number of pages that can be moved to the frequent-access memory. This can be calculated mathematically -for instance, if in one epoch 50 ‘hot’ pages are discovered in 20 ms and the memory can support 150 memory pages then the sampling process might be run for 60 ms. In addition to or instead of this, the sampling process can be adjusted so that it keeps increasing each epoch until a number of pages slightly exceeds the capacity of the frequent-access memory and decreasing the length of the sampling process if it results in significantly too many ‘hot’ pages (e.g. using a trial and error mechanism). In some examples, in response to the counter becoming saturated, the storage circuitry is configured to store a record in a buffer. When the counter becomes saturated a log entry can be made in the buffer. This makes it possible to gather a more complete view of only the ‘hot’ pages. Furthermore, the list of such pages can be seen by software, provided that the buffer is architecturally addressable (for instance, the buffer might be the general purpose DRAM that is accessible by software). This allows more dynamic decisions to be made. For instance, it may be desirable to keep blocks of memory pages together in memory even if one of the pages is frequently accessed. In some examples, the counter of the one or more address translation is reset each epoch. That is, the sampling process may be repeated with frequently accessed or ‘hot’ pages potentially changing over time. Particular embodiments will now be described with reference to the figures. Figure 1 schematically illustrates an example of a system 2. It will be appreciated that this is simply a high level representation of a subset of components of the system and the system may include many other components not illustrated. The system 2 comprises processing circuitry 4 for performing data processing in response to instructions decoded by an instruction decoder 6. The instruction decoder 6 decodes instructions fetched from an instruction cache 8 to generate control signals 10 for controlling the processing circuitry 4 to perform corresponding processing operations represented by the instructions. The processing circuitry 4 may include one or more execution units for performing operations on values stored in registers 14 to generate result values to be written back to the registers. For example the execution units could include an arithmetic / logic unit (ALU) for executing arithmetic operations or logical operations, a floating-point unit for executing operations using floating-point operands and / or a vector processing unit for performing vector operations on operands including multiple independent data elements. The processing circuitry also includes a memory access unit (or load / store unit) 15 for controlling transfer of data between the registers 14 and the memory system. In this example, the memory system includes the instruction cache 8, a level 1 data cache 16, a level 2 cache 17 shared between data and instructions, and main memory 18. It will be appreciated that other cache hierarchies are also possible - this is just one example. A memory management unit (MMU) 20 is provided for providing address translation functionality to support memory accesses triggered by the load / store unit 15. The MMU has a translation lookaside buffer (TLB) 22 for caching a subset of entries from page table stored in the memory system 16, 17, 18, 19. Each page table entry may provide an address translation mapping for a corresponding page of addresses and may also specify access control parameters, such as access permissions specifying whether the page is a read only region or is both readable and writable, or access permissions specifying which privilege levels can access the page. Attempts to access pages where permission is not given may result in a fault 60. In this example, the memory system 16, 17 is made up from two memories 18, 19 backed by DRAM for instance. The memories include a local memory 18 and a main memory 19. The local memory 18 is nearer to the processing circuitry 4 than the main memory 19, which may be on a different chip and may require off-chip communication to take place. Consequently, data can be accessed more quickly from the local memory 18 than the main memory 19. It will be appreciated that it is desirable for frequently accessed data (data required by the processing circuitry 4) to be in the fastest memory that is practical so that the benefits of the faster memory system component can be realised. Clearly the location of the data will depend on the granularity of data being referred to. It may therefore be preferable to store the pages in the local memory 18 rather than the main memory 19 Movement between the memories 18, 19 can be performed by promotion circuitry 24. The extent of accesses can be tracked in a buffer 26, which may form part of the one of the memories 18 or could be a separate dedicated structure. In addition, parameter storage 28 is used to store parameters such as an indication of how ‘hot’ or frequently accessed a section of memory must be in order to get promoted to the faster storage 18. In this example, the claimed apparatus includes the TLB 22, but could include other components such as the MMU 20, the system 2, or even a component including the system 2. Furthermore, the present technique is not limited to a TLB 22 but could instead be a separate component that tracks the usage of portions of memory (e.g. via memory access requests). Figure 2 shows an example of the claimed apparatus 100 in accordance with some examples. The apparatus includes storage circuitry 102, which in this case stores a number of page table entries (PTEs). Each PTE is associated with a page of memory (e.g. of 64 kB), which can be addressed using a virtual address (VA). The virtual address is an example of the claimed first address domain and the physical address is an example of the claimed second address domain. In this case, the apparatus takes the form of the TLB 22 and therefore receives requests for translation from a virtual address to a physical address at the receive circuitry 106. A lookup is then performed on the storage circuitry 102 to search for the corresponding virtual address, which is returned using transmit circuitry 108. For completeness, when a lookup fails in the storage 102, this results in a page table walk taking place in which the translation is searched for using the page tables that may be stored in the memory system. The result is then returned for storage in the TLB 22 so that it can be accessed more quickly in the future. So in this example, a request for the translation of virtual address Ox 190da02e is received. This is found in the storage circuitry 102 and the resulting physical address 0xdl27 is returned. For each PTE, a counter 104 is provided, which gives an indication of the number of times that the translation has been accessed (and so, by extension, the number of times the page of memory has been accessed). In this example, the counter 104 for the returned entry (currently set to 2) may be incremented. For completeness, each PTE in this example also includes an application space ID (ASID), which indicates an application running on the processing circuitry 4 to which the page belongs, and a validity identifier (V), which indicates whether the PTE is valid (1) or not (0). Invalid entries are ignored. Other entries (such as those for indicating permissions) may also be present but are not shown. Note that in other examples, the claimed apparatus could take the form of a device that sits alongside or within the load / store unit 15 and monitors memory accesses (rather than translation requests for memory accesses). Such a device could be adaptable to the size of memory block that is considered. In particular, such a device could be used to migrate cache lines throughout the memory system. However, by adding the counters 104 to the page table entries of the TLB 22, it is possible to migrate pages of memory without significant additional hardware. It is desirable to increment the counters 104 probabilistically so that not every access to every block of memory is recorded - thereby reducing the amount of traffic / activity. By recording the accesses probabilistically, it is possible to instead infer the number of accesses without definitively recording them all. In this example, the TLB 22 acts as a cache for the counters 104, the official value for which are stored in the official page table entries (e.g. in memory). If and when the counters 104 are to be updated, the new value of the counter is written back to memory. Of course, as already explained, a dedicated cache could be used for the counters 104 and indeed in some cases the TLB 22 can be forgone in its entirety and the page table entries in memory can be updated directly. In some examples, multiple processors may seek to update the counters. This situation is described in more detail below. TLBs may provide translations between virtual and physical addresses but can also provide translations between virtual and intermediate physical addresses and between intermediate physical addresses and physical addresses - that is, the translation might be multi-stage. In these examples, the we are assuming that the page table entries are lead-entries. That is, they are either virtual to physical translations or intermediate address to physical translations. One way of achieving this is shown in Figure 3. Here, every access (e.g. translation request to the TLB 22 or memory access) has a probability of incrementing the counter. The probability P changes (decreases) as the counter increases. Here, the counter is a 3-bit value thereby allowing the representation of numbers between 0 and 7 (inclusive). Each time an access is made, a random number is generated (between 0 and 1) and the counter is incremented if the generated random number is less than P In this example, the probability of a counter change from 0 to 1 is 1. That is to say that when the counter is 0, the counter value will increment to 1 regardless of the random number generated. From 1 to 2, the probability is 1 / (2A6). From 2 to 3, the probability is 1 / (2A8). From 3 to 4, the probability is 1 / (2A10). From 4 to 5, the probability is 1 / (2A12). From 5 to 6, the probability is 1 / (2A14). From 6 to 7, the probability is 1 / (2A16). If the counter is 7, then the probability of further increase is 0 (i.e. no further increase happens) because this would saturate the counter. An entry may be written to a log in the buffer 26 to indicate that the counter has been saturated. It is therefore less likely that the counter will increase as the counter value increases and so very highly accessed pages of memory are limited in the number of counter increments that they cause (thereby limiting traffic and activity caused as a consequence of tracking the requests). In some scenarios, there may be multiple processors, each with their own TLB (or otherwise cached version of the counters). This can result in the locally cached version of the counters becoming out of date - e.g. if another processor causes its copy of the counter value to be incremented. When the counter is to be written back to memory (e.g. when it is to be incremented) , it may be the case that the official counter value (e.g. in the page table entry in memory) has already been increased. In some examples, the increment that has newly occurred is simply discarded and the updated version of the counter is copied from memory into the local cache (e.g. the TLB). In other examples, the newly updated version of the counter from the page table entry in memory is used to re-determine the probability of the counter change, and the counter update check is made again (i.e. another random number is generated and checked against the newly updated threshold). The update to the page table entry can be made using either a compare-and-swap (CAS) operation or a read-modify-write (RmW) operation. Such operations allow for the atomic updating of a value, which may be useful for providing consistency where multiple requesters or MMUs may be operating. Figure 4 shows an alternative arrangement. The apparatus 300 is generally similar to the apparatus 100 shown in Figure 2 except that the counter 304 is a 1-bit counter that is always incremented (i.e. without a random probability) when an access is made to the corresponding page of memory. In these examples, the probabilistic incrementing is achieved by not considering every access that occurs, but instead only considering a proportion of them. There is therefore a less than 100% probability that any given access will be cause the counter 304 to be incremented. Consequently, the more frequently a block of memory (e.g. a page of memory) is accessed, the more likely that the counter associated with that block of memory will be incremented. In this example, any page whose counter is increased (i.e. to 1) might be promoted to the faster / local memory 18 (with other pages being demoted back down to the main memory 19). This technique relies on more frequently accessed pages being more likely to fall within the sample being taken. Such pages are therefore more likely to have their counter incremented. By decreasing the sample size, it becomes less likely that infrequently accessed pages will have the counter incremented and so will be less likely to be promoted. Furthermore by redetermining the position of pages on a regular basis, it is expected that frequently accessed pages will regularly find themselves in the faster memory 18 with less frequently access pages being promoted less often (i.e. being placed in the slower main memory 19). This setup is also amenable to changing access patterns in that infrequently accessed pages that have a small burst of activity can be temporarily placed in to the faster local memory 18 in order to benefit their burst of activity with faster access. In some examples, each time the counter of an entry is incremented, then a further counter 306 is also incremented. This further counter 306 can be implemented in hardware and thereby using to quickly provide a count of the number of entries that are to be promoted to the faster memory / cache. Figure 5 shows two ways in which sampling can occur. In the first example (sampling 1), one in every M requests are sampled. That is, a request is taken, and any relevant counters are incremented, then the next M-l requests are ignored (i.e. have no effect). In the example shown in Figure 5, M is set to 5 so that l-in-5 requests are sampled. In the second example in Figure 5 (sampling 2), a window of n requests are sampled and remaining requests are then ignored. The former sampling technique therefore spreads out the requests. This is helpful, as it tends to produce a more representative sample. In particular, the second example is prone to selecting a set of requests that are all part of the same function and that therefore might all hit the same page (by virtue of spatial locality). For instance, a large number of requests might be adjacent in a situation where an array is being iterated. In this situation, the window might correspond with this access pattern and lead to a high number of accesses to the same page of memory. Regardless of the sampling process being used, at some point in time an epoch occurs. The epochs are usually equally spaced out, but this need not be so. At the end of the epoch, the counters 304 are reset and a new process begins. The end of the epoch might also mark the point where promotion between caches or memories takes place (particularly where 1-in-M sampling is performed) since it is at this point where it is not whether a page is ‘hot’ or not. In the case of the windowed sampling (sampling 2), the promotion (and demotion) can happen after the sampling window of n samples has been handled and this makes it possible to spread out the traffic cost of moving pages of data between the local memory 18 and the main memory 19. Note that in this example, n (the window size) is defined as a number of requests. In other examples, it can be defined as a period of time. In each of the sampling techniques it is possible for the extent of sampling to change (e.g. at each epoch). This will reduce the maximum number of requests that can be promoted to the local memory 18. In some examples, promotion is based on the hotness of a page over a series of epochs. For example, a page may be promoted if it is ‘hot’ (e.g. accessed in the sample of accesses) over either a series of S epochs or perhaps R out of S epochs with pages meeting these requirements being promoted and pages not meeting the requirements being demoted. Software could also track or rank the hotness of pages, e.g. over the most recent S epochs with more hot pages being promoted and less hot pages being demoted. Figure 6 illustrates, in the form of a flow chart 500, a method for altering the sampling rate. At a step 502, the sampling process begins and the counters are adjusted (incremented) as appropriate as requests come in. Either during the process or at the end of the epoch at step 504, pages are migrated based on their counter value. In particular, pages whose counters are above the threshold (0 in the case of the examples shown in Figure 5) are moved to the local memory 18 and other pages are moved to the main memory 19. At step 506, it is determined whether the number of pages to be migrated will cause the limit of use of the local memory 18 to be exceeded (e.g. whether its capacity will be nearly or actually exceeded). If so, then the sampling rate of the next epoch (M or n) is reduced by, e.g. 10% at step 508. Otherwise, the sampling rate is increased by 5% at step 510. In either case, at step 512, the process waits until the next epoch and the counters are reset. The process then returns to step 502. The altered sampling rate can be stored in a hardware register for efficiency purposes. In addition, the capacity and location of the buffer 26 can be configurable by software using hardware registers. It may also be stored in a dedicated hardware structure. In this way, the rate of sampling is changed so that the number of pages promoted to the local memory 18 is kept as high as possible such that the limit is not reached. This avoids a situation in which either the number of pages is too many for the local memory 18 to handle or is too few such that little benefit is gained from the promotion. Of course in this example, the control loop is simple and increases the sampling by 5% where more capacity exists and decreases the sampling by 10% where too much sampling is being performed. That is, if the capacity is close to being exceeded, it backs off quicker than it builds up. Other control loops are of course usable. Figure 7 illustrates step 504 (the process of determining migration) in more detail. In practice, this step might also be used in the system shown in Figure 3 where a larger counter is provided than a 1 -bit counter. At step 600, the counter of the next page table entry (or entry having an associated counter) is considered. Then, at step 602, it is determined whether the counter is greater than the predetermined value. If not, then the page (or block of memory such as a cache line) to which the entry / counter relates (e.g. is counting accesses) is migrated to slower memory 19 at step 604. Otherwise, it is migrated to faster memory 18 at step 606. The process then returns to step 600. Note that for efficiency purposes, in some embodiments, whenever the counter reaches the predetermined value, data regarding the entry is stored (e.g. in a log or in the buffer 26). This reduces the need to iterate through the page table entries to determine which pages should be promoted. Instead, software can read the log / buffer in order to determine the pages to be promoted. If the log / buffer reaches its capacity then processing of the log / buffer to promote entries is immediately scheduled and further updates are disabled until that processing is complete. Figure 8 illustrates a simulator implementation that may be used. Whilst the earlier described embodiments implement the present invention in terms of apparatus and methods for operating specific processing hardware supporting the techniques concerned, it is also possible to provide an instruction execution environment in accordance with the embodiments described herein which is implemented through the use of a computer program. Such computer programs are often referred to as simulators, insofar as they provide a software based implementation of a hardware architecture. Varieties of simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 730, optionally running a host operating system 720, supporting the simulator program 710. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and / or multiple distinct instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations which execute at a reasonable speed, but such an approach may be justified in certain circumstances, such as when there is a desire to run code native to another processor for compatibility or re-use reasons. For example, the simulator implementation may provide an instruction execution environment with additional functionality which is not supported by the host processor hardware, or provide an instruction execution environment typically associated with a different hardware architecture. An overview of simulation is given in “Some Efficient Architecture Simulation Techniques”, Robert Bedichek, Winter 1990 USENIX Conference, Pages 53 - 63. To the extent that embodiments have previously been described with reference to particular hardware constructs or features, in a simulated embodiment, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuitry may be implemented in a simulated embodiment as computer program logic. Similarly, memory hardware, such as a register or cache, may be implemented in a simulated embodiment as a software data structure. In arrangements where one or more of the hardware elements referenced in the previously described embodiments are present on the host hardware (for example, host processor 730), some simulated embodiments may make use of the host hardware, where suitable. The simulator program 710 may be stored on a computer-readable storage medium (which may be a non-transitory medium), and provides a program interface (instruction execution environment) to the target code 700 (which may include applications, operating systems and a hypervisor) which is the same as the interface of the hardware architecture being modelled by the simulator program 710. Thus, the program instructions of the target code 700 may be executed from within the instruction execution environment using the simulator program 710, so that a host computer 730 which does not actually have the hardware features of the apparatus 2 discussed above can emulate these features. In the present application, the words “configured to..are used to mean that an element of an apparatus has a configuration able to carry out the defined operation. In this context, a “configuration” means an arrangement or manner of interconnection of hardware or software. For example, the apparatus may have dedicated hardware which provides the defined operation, or a processor or other processing device may be programmed to perform the function. “Configured to” does not imply that the apparatus element needs to be changed in any way in order to provide the defined operation. Although illustrative embodiments of the invention have been described in detail herein with reference to the accompanying drawings, it is to be understood that the invention is not limited to those precise embodiments, and that various changes, additions and modifications can be effected therein by one skilled in the art without departing from the scope and spirit of the invention as defined by the appended claims. For example, various combinations of the features of the dependent claims could be made with the features of the independent claims without departing from the scope of the present invention.

Claims

1. An apparatus comprising:storage circuitry configured to store one or more addresses and one or more counters each associated with one of the one or more addresses; andreceive circuitry configured to receive a request comprising an address, whereinin response to receiving the request, one of the one or more counters in the storage circuitry that is associated with the address that is in the request is incremented probabilistically.

2. The apparatus according to claim 1, whereinthe request is a translation request from the address in a first address domain to a translated address in a second address domain; andthe storage circuitry is configured to store one or more translated addresses, each associated with one of the one or more addresses.

3. The apparatus according to claim 2, whereinthe second address domain is either a physical address domain or an intermediate physical address domain;the first address domain is either an intermediate physical address domain or a virtual address domain; andthe first address domain and the second address domain are different.

4. The apparatus according to any one of claims 1-3, whereinthe one or more counters are incremented probabilistically in response to a randomly generated number being compared to a threshold value.

5. The apparatus according to any preceding claim, whereinthe one or more counters are N-bit counters, where N is an integer greater than one.

6. The apparatus according to claim 5, whereinthe threshold value decreases as a value of the one or more counters increases.

7. The apparatus according to claim 6, whereina rate at which the threshold value decreases as the value of the one or more counters increases is greater than linear.

8. The apparatus according to any preceding claim, comprising:promotion circuitry configured to migrate memory pages having a physical address whose corresponding value of the one or more counters is above a predetermined value.

9. The apparatus according to claim 8, whereinthe promotion circuitry is configured to migrate the memory pages from regular-access memory to frequent-access memory which is faster or logically closer to a processor than the regular-access memory.

10. The apparatus according to any one of claims 1-3, whereinthe one or more counters are incremented probabilistically as a consequence of the request falling within a sample of requests received by the apparatus.

11. The apparatus according to claim 10, whereina size of the sample of the requests is one in every M of the requests, where M is an integer greater than one.

12. The apparatus according to claim 10, whereina size of the sample of the requests is defined by a window in the requests.

13. The apparatus according to any one of claims 10-12, whereinthe size of the sample of the requests is dynamically changeable each epoch.

14. The apparatus according to any one of claims 12-13, comprising:promotion circuitry configured to migrate memory pages having a physical address whose corresponding value of the one or more counters is above a predetermined value.

15. The apparatus according to claim 14, whereinthe promotion circuitry is configured to migrate the memory pages from regular-access memory to frequent-access memory which is faster or logically closer to a processor than the regular-access memory.

16. The apparatus according to claim 15, whereinthe size of the sample of the requests is set such that a number of the memory pages having a physical address in the storage circuitry whose corresponding value of the counter is above the threshold is equal to or less than a number of memory pages that can be stored to the frequent-access memory.

17. The apparatus according to any preceding claim, whereinin response to the counter becoming saturated, the storage circuitry is configured to store a record in a buffer.

18. The apparatus according to any preceding claim, whereinthe counter of the one or more address translation is reset each epoch.

19. A data processing method comprising:storing one or more addresses and one or more counters each associated with one of the one or more addresses;receiving a request comprising an address; andin response to receiving the request, one of the one or more counters that is associated with the address that is in the request is incremented probabilistically.5 20. A computer program for controlling a host data processing apparatus toprovide an instruction execution environment, the computer program comprising:data structures configured to store one or more addresses and one or more counters each associated with one of the one or more addresses; and10 receive program logic configured to receive a request comprising anaddress, whereinin response to receiving the request, one of the one or more counters in the data structures that is associated with the address that is in the request is incremented probabilistically.15

Citation Information

Patent Citations

  • Method and apparatus for protecting memory devices via a synergic approach

    US20230114414A1