Memory address access frequency
The described apparatus and method address inefficiencies in memory access tracking by using trace and processing circuitry to reduce data volume and storage needs, facilitating efficient memory access frequency analysis and placement strategies.
Patent Information
- Application Number
- GB2023018221
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-11
AI Technical Summary
Existing memory access tracking methods face challenges in balancing storage requirements and data granularity, leading to inefficiencies in determining memory access frequencies and placement strategies.
An apparatus and method utilizing trace circuitry to generate memory access traces, buffer circuitry for binning, and processing circuitry to update frequency bins, employing approximate counting algorithms to reduce data volume while maintaining access frequency information.
This approach efficiently stores memory access frequency data with reduced bandwidth consumption and storage requirements, enabling intelligent data placement and analysis.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present technique relates to data processing and particularly the field of memory accessing. It may be desirable to understand which areas of memory are being accessed over a period of time. For instance, this can be useful in working out where to place data. For instance, data that is accessed frequently might be placed in areas of memory that have lower latency and / or higher bandwidth. Meanwhile, data that is accessed together might be Toad balanced’ so that it can be simultaneously accessed in a short space of time. One way of doing this might be to gather trace data regarding data accesses and store them. However, such data is large, which requires a large amount of storage. This could be overcome by storing less data (e.g. only the most recent data). But this inevitably results in less information being available. There is therefore a tradeoff in terms of the amount of storage that must be supplied and the amount of information available to make intelligent decisions on data placement, as well as bandwidth for capturing the data. A further trade-off is flexibility versus efficiency in that we want to be flexible in what data we can collect and analyse while doing so in an efficient manner. Viewed from a first example configuration, there is provided an apparatus comprising: trace circuitry configured to generate a trace indicating a series of memory addresses of memory accesses made to a memory; buffer circuitry configured to perform binning on the memory addresses to produce buffer-circuitry-maintained-frequency-bins that indicate access frequencies of the memory addresses, and processing circuitry, separate from the buffer circuitry, and configured to execute a stream of instructions to update processing-circuitry-maintained-frequency-bins to indicate access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins. Viewed from a second example configuration, there is provided a method comprising: generating a trace indicating a series of memory addresses of memory accesses made to a memory; performing binning on the memory addresses to produce buffer-circuitry-maintained-frequency-bins that indicate access frequencies of the memory addresses; and executing a stream of instructions to update processing-circuitry-maintained-frequency-bins to indicate access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins. Viewed from a third example configuration, there is provided trace circuitry configured to generate a trace indicating a series of memory addresses of memory accesses made to a memory; buffer circuitry configured to perform binning on the memory addresses to produce buffer-circuitry-maintained-frequency-bins that indicate access frequencies of the memory addresses; and processing circuitry, separate from the buffer circuitry, and configured to execute a stream of instructions to update processing-circuitry-maintained-frequency-bins to indicate access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins. The present technique will be described further, by way of example only, with reference to embodiments thereof as illustrated in the accompanying drawings, in which: Figure 1 illustrates a data processing apparatus in accordance with some examples; Figure 2 shows an example of how the buffer-circuitry-maintained-frequency-bins in storage can be updated in response to notifications about memory accesses being received at the buffer circuitry, e.g. as part of a trace sample; Figure 3 shows an example of the update process; Figure 4 illustrates an example in which the rewriteable storage contains different sets of instructions to be executed by the processing circuitry; Figures 5 A and 5B illustrate flowcharts that show methods of handling the data process in accordance with some examples; and Figure 6 shows an example of implementing the technique on a chip. Before discussing the embodiments with reference to the accompanying figures, the following description of embodiments and associated advantages is provided. In accordance with one example configuration there is provided an apparatus comprising: trace circuitry configured to generate a trace indicating a series of memory addresses of memory accesses made to a memory; buffer circuitry configured to perform binning on the memory addresses to produce buffer-circuitry-maintained-frequency-bins that indicate access frequencies of the memory addresses; and processing circuitry, separate from the buffer circuitry, and configured to execute a stream of instructions to update processing-circuitry-maintained-frequency-bins to indicate access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins. In the above examples, the trace contains indications of memory accesses made to a memory (e.g. a memory hierarchy) and in particular identifies the addresses that are being accessed. The trace is used to update a series of ‘bins’ that are stored by buffer circuitry. In particular, these bins indicate how many times an address (or a range of addresses) have been accessed over a defined period. At a point in time, a stream of instructions is executed by processing circuitry. These instructions are used to update a further set of bins that are maintained by the processing circuitry. Either or both sets of bins may be implemented using approximate counting algorithms (see, for instance https: / / en,wikipedia.org / wiki / Lossy Count Algorithm). The update is based on the contents of the bins in the buffer circuitry. By performing binning using hardware it is possible to reduce the amount of data that is to be transmitted using memory bandwidth. In particular, data regarding the exact order of memory accesses is mostly disregarded since this is not relevant to determining the frequency with which the memory locations are hit. Furthermore, the binning enables some degree of aggregation to occur by tracking that, for instance, a particular address (or a range of addresses) was accessed several times, rather than listing every individual access, which would result in repeated data. The need for excessive hardware storage is resolved by enabling long term data to be stored in the processing-circuitry-maintained-frequency-bins, which can for instance be stored in main memory or even in backing storage as desired. The combination of the processing circuitry and the buffer circuitry therefore make it possible to store a large amount of memory address access frequency data without excessive bandwidth consumption. In some examples, the trace circuitry is configured to generate the trace for a subset of all memory accesses made by at least one accessing element to the memory. The trace need not contain a complete list of memory accesses. Instead, statistical approaches can be made such as the trace containing details of every N’th memory access or memory accesses being randomly assigned to the trace. In either case, since the trace will contain a statistical sample of all memory accesses, it can be assumed that the analysis of the sample will be reflective of the type and number of memory accesses actually occurring. In particular, if only 1 / N memory accesses are part of the trace (through either round-robin selection or random selection as discussed above), one might expect, statistically, for the frequency access numbers to be approximately 1 / N’th of their true value. Relatively speaking, the access frequency distribution should remain approximately the same. In some examples, at least one of the buffer-circuitry-maintained-frequency-bins and the processing-circuitry-maintained-frequency-bins have a size greater than one. Although it is possible to achieve an element of amalgamation at the buffer circuitry by making each bin a size of ‘ 1’ (e.g. since multiple memory accesses to the same single address will be compressed to one entry), a further degree of compression can be achieved if the bins are larger than one. That is to say that each bin refers to a range of addresses and the frequency associated with that bin is incremented whenever a memory access occurs to any memory address within that range. In some examples, at least one of the buffer-circuitry-maintained-frequency-bins and the processing-circuitry-maintained-frequency-bins have a size equal to a page size of the memory. In these examples, the range of each bin corresponds with a memory page. Consequently, any access made to an address within that page causes the corresponding entry to be incremented. The size of a memory page will be dependent on the memory architecture and typically defines the range of virtual addresses for which a single entry is provided in a page table that provides translations from virtual to physical addresses. In some examples, in response to receiving notification of one of the memory accesses to one of the memory addresses: the buffer circuitry is configured to increment a frequency of one of the buffer-circuitry-maintained-frequency-bins that corresponds with the one of the memory addresses if the one of the buffer-circuitry-maintained-frequency-bins that corresponds with the one of the memory addresses is present, and to add a new entry to the buffer-circuitry-maintained-frequency-bins that corresponds with the one of the memory addresses otherwise, in dependence on a further condition. The binning process therefore increases the frequency for an entry that corresponds with the incoming memory access. If the memory access is to a memory location that is not present (i.e. is not reflected by one of the bins) then provided another condition is met, a new entry having a bin into which that memory access can be inserted will be created. This condition could be that there is sufficient storage in the buffer circuitry for example (possibly after another entry has been dropped depending on eviction criteria). In some examples, the processing circuitry is configured to update the processing-circuitry-maintained-frequency-bins to indicate the access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins by amalgamating the processing-circuitry-maintained-frequency-bins and the buffer-circuitry-maintained-frequency-bins. The amalgamation can take a number of different forms. In some examples, the amalgamation is achieved by adding the totals of corresponding bins together. Bins that are not present are assumed to have a value of 0. In cases where the bins are aligned at the same point and are the same length (e.g. starting at address 0 and 4k each), this is a simple addition process. In other examples where bins are of different sizes, it might be necessary to work out how much overlap there is between the bins and interpolate accordingly. For instance, a buffer-circuitry bin that covers 6k-10k might provide 50% of its total to the processing-circuitry bin 4k-8k and 50% of its total to the processing-circuitry 8k-10k bin. Alternatively, a 6k-7k buffer circuitry bin and a 7k-8k buffer circuitry bin might have both of their totals added to a 4k-8k processing circuitry bin. In some examples, the apparatus comprises: buffer bin storage circuitry configured to store the buffer-circuitry-maintained-frequency-bins; and processing circuitry bin storage circuitry configured to store the processing-circuitry-maintained-frequency-bins, wherein a capacity of the processing circuitry bin storage circuitry is greater than a capacity of the buffer bin storage circuitry. Since the storage capacity (in terms of the number of bins and / or the frequency that can be assigned to the bin) used for the processing-circuitry-maintained-frequency-bins is greater than that used for the buffer-circuitry-maintained-frequency-bins, more information can be stored. Thus, the overall amount of trace data that is to be stored can be kept high. In addition, since the processing-circuitry-maintained-frequency-bins are programmable, various schemes could be used to further reduce the space required. The data could also be stored in a way that makes likely computations easier. One likely computation is finding the bins with the highest or lowest access frequency. In some examples, the buffer-circuitry-maintained-frequency-bins relate to those of the memory accesses occurring within a predetermined preceding window. The window could be a sliding window if, for instance, the time at which each entry was added to the buffer circuitry was kept. In these situations, there may be no need for the contents of the buffer circuitry to be reset periodically provided the sliding window coincides with the amalgamation frequency (e.g. so that a data access is not amalgamated twice). In some examples, the processing circuitry is configured, in response to executing the stream of instructions to update the processing-circuitry-maintained-frequency-bins, to reset the buffer-circuitry-maintained-frequency-bins. In these examples, when the amalgamation takes place, the buffer-circuitry-maintained-frequency-bins are completely reset (e.g. to 0). The buffer circuitry therefore ‘begins again’ and starts counting again from zero. This can be useful when considering the principle of spatial locality. In particular, it is expected that memory accesses that occur next to each other are likely to be close to each other spatially. That is, that they will affect memory addresses that are close together. By resetting the buffer bins with some frequency, it means that tracking of bins that were accessed long ago is generally unlikely to continue. This frees up the buffer circuitry to track other accesses that have occurred more frequently. In some examples, the apparatus comprises: instruction storage circuitry configured to store the stream of instructions, wherein the instruction storage circuitry is rewritable. In these examples, therefore, the instructions that are used to update the processing-circuitry-maintained-frequency-bins can be changed as required, e.g. to perform updates in different ways. In some examples, the stream of instructions is changeable at runtime. In some examples, as the processing circuitry is being used, the instructions stored in the instruction storage circuitry are changed. In some examples, the processing circuitry itself may execute instructions to change the instructions that are stored in the instruction storage circuitry to perform the update, e.g. at runtime. In some examples, the processing-circuitry-maintained-frequency-bins are groups into sets of processing-circuitry-maintained-frequency-bins; and for each of the set in the sets of processing-circuitry-maintained-frequency-bins, the stream of instructions is selected from a plurality of streams of instructions and applied to those of the processing-circuitry-maintained-frequency-bins in that set. In these examples, different forms of update can be performed for different bins. In particular, the update might be performed differently for different applications. For example, for one application that performs highly localised spatial data, it might not be relevant what data was being accessed several minutes ago and so the update being performed might only look at the most recent window of time as reported by the buffer circuitry. For another application (which might be being executed simultaneously), memory access might be more distributed and so it might be useful to consider the memory accesses and particularly which memory addresses are being accessed most often over a long period of time. This might therefore result in an update being performed that adds the values held by the buffer circuitry to values held by the processing circuitry. Particular embodiments will now be described with reference to the figures. Figure 1 illustrates a data processing apparatus 100 in accordance with some examples. An accessing element 110 (which may be one of several) issues memory access requests to a memory. Notification and details of these accesses is sent to trace circuitry 120, which generates a trace from a sample of the notifications. There are a number of ways in which the trace may be generated. For instance, every N’th memory access from a given requester may be included within the trace. Alternatively, every time the trace circuitry 120 receives information of a memory access, there could be a random chance that the data is included with the trace that is generated by the trace circuitry. In some examples, all memory accesses are filtered according to particular characteristics so that the sampling is achieved based on those characteristics. Such characteristics might include memory accesses within an address range, memory accesses that are demand accesses vs prefetches, memory accesses that miss a level of cache, or memory accesses that stall processor execution beyond a threshold number of cycles, or any of these in combination. It is of course possible, in some embodiments, that the trace that is generated will not be a sample and will instead contain a record of every memory access that occurs. The trace that is generated is provided to buffer circuitry 130 and is used to fill a number of buckets that are maintained by the buffer circuitry 130 in storage 160 (which may form part of the buffer circuitry 130 itself). In this example, the buckets of the buffer circuitry 130 have a size that is commensurate with a page size of the memory that is being accessed - 4k in this example. As shown in this example, the buffer circuitry has seen 11 accesses to addresses at or above 0 and below 4k and 42 accesses to addresses at or above 4k and below 8k. Being as the buffer circuitry 130 provides dedicated hardware storage, the buffer circuitry 130 is able to operate quickly. The buffer circuitry may even form part of the accessing element 110 thereby limiting the effect on any bus bandwidth. Processing circuitry 140 also maintains a separate set of bins in a different storage 170 (such as DRAM). In this example, the bins maintained by the processing circuitry 140 are more extensive than those maintained by the buffer circuitry 130. That is, a greater number of bins can be stored and the frequency that can be expressed for each bin is also higher. Indeed, the storage 170 used to store the bins of the processing circuitry 140 may be capable of storing frequency data for every page in the memory system. Periodically, an update is performed by the processing circuitry 140 to update its bins in its storage 170. The update that is to be performed is handled by instructions 150 that are stored in rewriteable storage and thus the instructions themselves may be changed, e g. at runtime. The update can take a number of forms, but is based on the bins that are maintained by the buffer circuitry 130. As a consequence of the above, it is possible to keep track of all of the sampled memory accesses, rather than merely a subset of them, by making use of the large storage 170 available to the processing circuitry. On the other hand, since a first level of amalgamation takes place in the storage 160 used by the buffer circuitry 130, the amount of bandwidth used in storing all of the access data is reduced as compared to a situation where the same samples are stored directly by the processing circuitry 140. Note that in this example, elements have been separated out in order to improve readability. However, several components could in practice be the same. For instance the processing circuitry 140 and the accessing elements (or at least one of them) 110 could be the same or they could be components of the same element. Furthermore, the memory being accessed and the DRAM storage 170 could be the same, and either of these could be the same as the rewriteable storage 150 used to store the instructions. Other possibilities also exist. Figure 2 shows an example of how the buffer-circuitry-maintained-frequency-bins in storage 160 can be updated in response to notifications about memory accesses being received at the buffer circuitry 130. In this example, the bins are of size 4k, which is also the page size. A first memory read request to an address 3845 is received. This causes the frequency count for the bin ‘>= 0 and <4k’ in the storage circuitry 160 to be incremented by 1. A memory write request is then received. This writes the value stored in register r5 to the memory address 16718. In this case, there is no bin that covers the address 16718. If storage is available in the storage circuitry 160 then a new bin is created (‘>= 16k and <20k’) and this is set to an initial value (e.g. 1). Then, another memory write request occurs, which writes the value in register r5 to address 20718. In this case, there is no remaining storage to store a further new bin. Consequently, this memory access notification / trace item is dropped. Other conditions might affect whether a particular memory access is kept or not when an existing bin does not cover the address of the memory access. The result of the modifications is a set of bins that covers addresses in the ranges >= 0 and <4k, >= 4k and <8k, or >= 16k and <20k. As compared to the initial states of the bins (e.g. the frequency counts that were initially present), the count for ‘>= 0 and <4k’ has incremented by one, and the count for ‘>= 4k and <8k’ has remained the same because no access has been made to an address within that range. It is of course possible for other update techniques to be used. In some examples, the lack of storage space to add a new bin may cause an update of the processing-circuitry-maintained-frequency-bins (as shown in Figure 3), for instance. Figure 3 shows an example of the update process. Here, the processing-circuitry-maintained-frequency-bins in storage 170 are updated with reference to the buffer-circuitry-maintained-frequency bins 160. Here, the update involves accumulation, which is achieved by adding bins together. For instance, the bin ‘>= 0 and <4k’ in the processing-circuitry-maintained-frequency-bins is initially 8773. The same bin in the buffer-circuitry-maintained-frequency-bins is 11. These two values are added together so that the bin obtained a value of 8784 in the updated processing-circuitry-maintained-frequency-bins. Similarly, for the ‘>= 4k and <8k’ bin, the frequency is updated from 3 by 42 to give an updated value of 45. A non-present bin is assumed to have a value of zero. Hence, the bin ‘>= 8k and <12k’, which has a value of 99 in the processing-circuitry-maintained-frequency-bins remains at that value since there is no corresponding bin in the buffer-circuitry-maintained-frequency-bins. Thus, the value remains at 99. There are other forms of accumulation that can take place. In particular. Figure 4 illustrates an example in which the rewriteable storage 150 contains different sets of instructions to be executed by the processing circuitry 140. Each set of instructions produces a different update mechanic and each set of instructions in the rewritable storage can be replaced (e.g. at runtime). In this way, the update process can be tailored for different situations without the hardware being modified. Here, each set of instructions is applied to a different set of bins - each set of bins corresponding to a different application. For example, A first set of instructions (inst set 1) provides an addition mechanism as previously illustrated. Here, there are two bins that are subject to the policy (‘>= 0 and <4k’ and ‘>= 4k and <8k’). These bins have values of 8773 and 3 respectively. The corresponding bins in the buffer circuitry at update time are 11 and 42. Hence, the amalgamation by addition process causes the bins maintained by the processing circuitry to become 8784 (8773 + 11) and 45 (42 + 3) respectively. A second set of instructions (inst set 2) provides a DROP mechanism. That is, each time the update is to occur, the new set of bins is merged with the old set of bins and the bin having the lowest frequency that is subject to these instructions is dropped. There are only two bins to which this policy is applied (‘>= 8k and <12k’ and ‘>= 12k and <16k’). The first bin starts with a value of 99. At the time of update, the corresponding bin maintained by the buffer circuitry has a value of 3. The merged value is therefore 102. The second bin has a value of 6 but has no value maintained by the buffer circuitry. The merged value is therefore 6. The smallest bin is therefore the ‘>= 12k and <16k’ bin and so this bin is dropped. A third set of instructions (inst set 3) provides a sort mechanism. That is, each time the update is to occur, the bins are merged according to their ranges and then sorted in descending order of frequency. The only new bin to which this policy is applied f>= 24k and <28k’) starts at a value of 99. At the time of update, the corresponding bin maintained by the buffer circuitry has a value of 7. Consequently, the value of the bin in the processing-circuitry-maintained-frequency bins is replaced with the value ‘106’. The bins maintained by the buffer circuitry include ‘>= 28k and <32k’ with a frequency of 3 and ‘>= 32k and <36k’ with a frequency of 55. These bins are therefore sorted in the order ‘>= 24k and <28k’, ‘>= 32k and <36k’, ‘>= 28k and <32k’. Another example of a set of instructions would be a balancing instruction to adapt the size of the bins in order to trade off accuracy for storage size. For instance, bins can be joined by adding the ranges and the counts within. Meanwhile bins can be split by assuming that the counts are evenly divided across a split. A next iteration can then use the newly joined / split bins. Figure 5A illustrates a first flowchart 500 that shows a first method of handling the update process. At a step 505, a trace entry is considered, e.g. at the buffer circuitry 130. The trace entry is added to the buffer circuitry 130 at step 510 as appropriate. That is, if a relevant bin already exists then it is incremented. If not, a new bin is created that encapsulates the trace entry. It is then determined whether the buffer is full. If not, then the process returns to step 505. Otherwise, at step 520, an update is to be performed on the frequency bins maintained by the processing circuitry and so an amalgamation takes place between those bins and the bins stored by the buffer circuitry. Then at step 525, the buffer circuitry 130 is reset (e.g. the contents are deleted). Note that multiple trace entries may be received from the trace circuitry 120 at a same time. In this case, the same flowchart is followed with each of the trace entries being considered one at a time. In this example, the buffer circuitry is erased each time it is filled in the amalgamation process. Figure 5B illustrates a second flowchart 530, which shows an alternative implementation. In this implementation, the buffer acts as a First In First Out (FIFO) buffer. Here, the trace entry is again considered at a step 535 and the entry is added to the buffer circuitry 130 in step 540 as appropriate (see Figure 2). That is, either an existing bin count is incremented or a new bin is created. At a step 545, it is determined whether any new entry to the buffer (i.e. a new bin) would cause the buffer to overfill. If not, then the process simply returns to step 535 to handle the next trace entry. If so, then at step 550, the oldest bin is popped from the buffer to allow the new entry (the new bin) to be added. The oldest bin in then amalgamated with the bins maintained by the processing circuitry at step 555. The process then returns to step 535. In this second example, the update process is quicker because only a single bin needs to be amalgamated at a time. Furthermore, since bins are kept for longer before being removed from the buffer circuitry, more buffering is performed before they are amalgamated. That is, in the example of Figure 5A, a widely distributed series of accesses may cause the creation of numerous bins, which are quickly removed and then amalgamated. In contrast, in the example of Figure 5B, bins are kept until they must be removed to make way for another bin. The counter values are therefore able to reach higher values. Of course, other variants are also possible. For instance, rather than popping and amalgamating a single entry in Figure 5B, a group of the oldest entries could be popped and amalgamated at steps 550 and 555. In other examples, the process may be similar to Figure 5A but rather than amalgamating and resetting the buffer when it is full, this could take place after a period of time elapses. This would help to solve a problem in which a large stream of accesses to the same address causes it to take a long time for the buffer to fill up and therefore be amalgamated with the processing-circuitry based bins. Of course these techniques can be combined as well. Another way that this can be achieved is to use a list of memory accesses rather than the bins. Each memory access can be stored together with its associated time, and any access that falls outside the sliding window can be deleted. When amalgamation is to take place at step 555, a count of the entries that would fall into each bin at the processing circuitry 140 is performed, and the bins maintained by the processing circuitry 140 are updated. In effect, in this example, bins can be replicated and each bin stores exactly one entry. Concepts described herein may be embodied in computer-readable code for fabrication of an apparatus that embodies the described concepts. For example, the computer-readable code can be used at one or more stages of a semiconductor design and fabrication process, including an electronic design automation (EDA) stage, to fabricate an integrated circuit comprising the apparatus embodying the concepts. The above computer-readable code may additionally or alternatively enable the definition, modelling, simulation, verification and / or testing of an apparatus embodying the concepts described herein. For example, the computer-readable code for fabrication of an apparatus embodying the concepts described herein can be embodied in code defining a hardware description language (HDL) representation of the concepts. For example, the code may define a register-transfer-level (RTL) abstraction of one or more logic circuits for defining an apparatus embodying the concepts. The code may define a HDL representation of the one or more logic circuits embodying the apparatus in Verilog, SystemVerilog, Chisel, or VHDL (Very High-Speed Integrated Circuit Hardware Description Language) as well as intermediate representations such as FIRRTL. Computer-readable code may provide definitions embodying the concept using systemlevel modelling languages such as SystemC and SystemVerilog or other behavioural representations of the concepts that can be interpreted by a computer to enable simulation, functional and / or formal verification, and testing of the concepts. Additionally or alternatively, the computer-readable code may define a low-level description of integrated circuit components that embody concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSII. The one or more netlists or other computer-readable representation of integrated circuit components may be generated by applying one or more logic synthesis processes to an RTL representation to generate definitions for use in fabrication of an apparatus embodying the invention. Alternatively or additionally, the one or more logic synthesis processes can generate from the computer-readable code a bitstream to be loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA may be deployed for the purposes of verification and test of the concepts prior to fabrication in an integrated circuit or the FPGA may be deployed in a product directly. The computer-readable code may comprise a mix of code representations for fabrication of an apparatus, for example including a mix of one or more of an RTL representation, a netlist representation, or another computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus embodying the invention. Alternatively or additionally, the concept may be defined in a combination of a computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus and computer-readable code defining instructions which are to be executed by the defined apparatus once fabricated. Such computer-readable code can be disposed in any known transitory computer-readable medium (such as wired or wireless transmission of code over a network) or non-transitory computer-readable medium such as semiconductor, magnetic disk, or optical disc. An integrated circuit fabricated using the computer-readable code may comprise components such as one or more of a central processing unit, graphics processing unit, neural processing unit, digital signal processor or other components that individually or collectively embody the concept. Concepts described herein may be embodied in a system comprising at least one packaged chip. The apparatus described earlier is implemented in the at least one packaged chip (either being implemented in one specific chip of the system, or distributed over more than one packaged chip). The at least one packaged chip is assembled on a board with at least one system component. A chip-containing product may comprise the system assembled on a further board with at least one other product component. The system or the chip-containing product may be assembled into a housing or onto a structural support (such as a frame or blade). As shown in Figure 6, one or more packaged chips 400, with the apparatus described above implemented on one chip or distributed over two or more of the chips, are manufactured by a semiconductor chip manufacturer. In some examples, the chip product 400 made by the semiconductor chip manufacturer may be provided as a semiconductor package which comprises a protective casing (e.g. made of metal, plastic, glass or ceramic) containing the semiconductor devices implementing the apparatus described above and connectors, such as lands, balls or pins, for connecting the semiconductor devices to an external environment. Where more than one chip 400 is provided, these could be provided as separate integrated circuits (provided as separate packages), or could be packaged by the semiconductor provider into a multi-chip semiconductor package (e.g. using an interposer, or by using three-dimensional integration to provide a multi-layer chip product comprising two or more vertically stacked integrated circuit layers). In some examples, a collection of chiplets (i.e. small modular chips with particular functionality) may itself be referred to as a chip. A chiplet may be packaged individually in a semiconductor package and / or together with other chiplets into a multi-chiplet semiconductor package (e.g. using an interposer, or by using three-dimensional integration to provide a multi-layer chiplet product comprising two or more vertically stacked integrated circuit layers). The one or more packaged chips 400 are assembled on a board 402 together with at least one system component 404 to provide a system 406. For example, the board may comprise a printed circuit board. The board substrate may be made of any of a variety of materials, e.g. plastic, glass, ceramic, or a flexible substrate material such as paper, plastic or textile material. The at least one system component 404 comprise one or more external components which are not part of the one or more packaged chip(s) 400. For example, the at least one system component 404 could include, for example, any one or more of the following: another packaged chip (e.g. provided by a different manufacturer or produced on a different process node), an interface module, a resistor, a capacitor, an inductor, a transformer, a diode, a transistor and / or a sensor. A chip-containing product 416 is manufactured comprising the system 406 (including the board 402, the one or more chips 400 and the at least one system component 404) and one or more product components 412. The product components 412 comprise one or more further components which are not part of the system 406. As a non-exhaustive list of examples, the one or more product components 412 could include a user input / output device such as a keypad, touch screen, microphone, loudspeaker, display screen, haptic device, etc.; a wireless communication transmitter / receiver; a sensor; an actuator for actuating mechanical motion; a thermal control device; a further packaged chip; an interface module; a resistor; a capacitor; an inductor; a transformer; a diode; and / or a transistor. The system 406 and one or more product components 412 may be assembled on to a further board 414. The board 402 or the further board 414 may be provided on or within a device housing or other structural support (e.g. a frame or blade) to provide a product which can be handled by a user and / or is intended for operational use by a person or company. The system 406 or the chip-containing product 416 may be at least one of: an end-user product, a machine, a medical device, a computing or telecommunications infrastructure product, or an automation control system. For example, as a non-exhaustive list of examples, the chip-containing product could be any of the following: a telecommunications device, a mobile phone, a tablet, a laptop, a computer, a server (e.g. a rack server or blade server), an infrastructure device, networking equipment, a vehicle or other automotive product, industrial machinery, consumer device, smart card, credit card, smart glasses, avionics device, robotics device, camera, television, smart television, DVD players, settop box, wearable device, domestic appliance, smart meter, medical device, heating / lighting control device, sensor, and / or a control system for controlling public infrastructure equipment such as smart motorway or traffic lights. In the present application, the words “configured to..are used to mean that an element of an apparatus has a configuration able to carry out the defined operation. In this context, a “configuration” means an arrangement or manner of interconnection of hardware or software. For example, the apparatus may have dedicated hardware which provides the defined operation, or a processor or other processing device may be programmed to perform the function. “Configured to” does not imply that the apparatus element needs to be changed in any way in order to provide the defined operation. Although illustrative embodiments of the invention have been described in detail herein with reference to the accompanying drawings, it is to be understood that the invention is not limited to those precise embodiments, and that various changes, 5 additions and modifications can be effected therein by one skilled in the art without departing from the scope and spirit of the invention as defined by the appended claims. For example, various combinations of the features of the dependent claims could be made with the features of the independent claims without departing from the scope of the present invention. 10
Claims
1. An apparatus comprising:trace circuitry configured to generate a trace indicating a series of memory addresses of memory accesses made to a memory;buffer circuitry configured to perform binning on the memory addresses to produce buffer-circuitry-maintained-frequency-bins that indicate access frequencies of the memory addresses; andprocessing circuitry, separate from the buffer circuitry, and configured to execute a stream of instructions to update processing-circuitry-maintained-frequency-bins to indicate access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins.
2. The apparatus according to claim 1, whereinthe trace circuitry is configured to generate the trace for a subset of all memory accesses made by at least one accessing element to the memory.
3. The apparatus according to any preceding claim, whereinat least one of the buffer-circuitry-maintained-frequency-bins and the processing-circuitry-maintained-frequency-bins have a size greater than one.
4. The apparatus according to any preceding claim, whereinat least one of the buffer-circuitry-maintained-frequency-bins and the processing-circuitry-maintained-frequency-bins have a size equal to a page size of the memory.
5. The apparatus according to any preceding claim, whereinin response to receiving notification of one of the memory accesses to one of the memory addresses:the buffer circuitry is configured to increment a frequency of one of the buffer-circuitry-maintained-frequency-bins that corresponds with the one of the memory addresses if the one of the buffer-circuitry-maintained-frequency-bins that corresponds with the one of the memory addresses is present, andto add a new entry to the buffer-circuitry-maintained-frequency-bins that corresponds with the one of the memory addresses otherwise, in dependence on a further condition.
6. The apparatus according to any preceding claim, whereinthe processing circuitry is configured to update the processing-circuitry-maintained-frequency-bins to indicate the access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins by amalgamating the processing-circuitry-maintained-frequency-bins and the buffer-circuitry-maintained-frequency-bins.
7. The apparatus according to any preceding claim, comprisingbuffer bin storage circuitry configured to store the buffer-circuitry-maintained-frequency-bins; andprocessing circuitry bin storage circuitry configured to store the processing-circuitry-maintained-frequency-bins, whereina capacity of the processing circuitry bin storage circuitry is greater than a capacity of the buffer bin storage circuitry.
8. The apparatus according to any preceding claim, whereinthe buffer-circuitry-maintained-frequency-bins relate to those of the memory accesses occurring within a predetermined preceding window.
9. The apparatus according to any preceding claim, whereinthe processing circuitry is configured, in response to executing the stream of instructions to update the processing-circuitry-maintained-frequency-bins, to reset the buffer-circuitry-maintained-frequency-bins.
10. The apparatus according to any preceding claim, whereinthe processing circuitry is configured to execute the stream of instructions to update the processing-circuitry-maintained-frequency-bins every predetermined period.
11. The apparatus according to any preceding claim, comprising:instruction storage circuitry configured to store the stream of instructions, whereinthe instruction storage circuitry is rewritable.
12. The apparatus according to any preceding claim, whereinthe stream of instructions is changeable at runtime.
13. The apparatus according to any preceding claim, whereinthe processing-circuitry-maintained-frequency-bins are groups into sets of processing-circuitry-maintained-frequency-bins; andfor each of the set in the sets of processing-circuitry-maintained-frequency-bins, the stream of instructions is selected from a plurality of streams of instructions and applied to those of the processing-circuitry-maintained-frequency-bins in that set.
14. A method comprising:generating a trace indicating a series of memory addresses of memory accesses made to a memory;performing binning on the memory addresses to produce buffer-circuitry-maintained-frequency-bins that indicate access frequencies of the memory addresses, andexecuting a stream of instructions to update processing-circuitry-maintained-frequency-bins to indicate access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins.
15. A non-transitory computer-readable medium to store computer-readablecode for fabrication of an apparatus comprising:trace circuitry configured to generate a trace indicating a series of memory addresses of memory accesses made to a memory;buffer circuitry configured to perform binning on the memory addresses to produce buffer-circuitry-maintained-frequency-bins that indicate access frequencies of the memory addresses; andprocessing circuitry, separate from the buffer circuitry, and configured to execute a stream of instructions to update processing-circuitry-maintained-frequency-bins to indicate access frequencies of the memory addresses based on the buffer-circuitry-maintained-frequency-bins.
16. A system comprising:the apparatus of any preceding claim, implemented in at least one packaged chip;at least one system component; anda board, whereinthe at least one packaged chip and the at least one system component are assembled on the board.
17. A chip-containing product comprising the system of claim 16 assembledon a further board with at least one other product component.
Citation Information
Patent Citations
Memory scanner to accelerate page classification
US11237981B1