Range Prefetch Instructions

JP2024523244A5Pending Publication Date: 2025-05-19ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023576042
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-23
Filing Date
2022-05-18
Publication Date
2025-05-19

AI Technical Summary

Technical Problem

Existing data processing systems face challenges with prefetching techniques, such as hardware prefetchers requiring multiple demand accesses to learn patterns and consuming significant resources, while software prefetching is difficult to time correctly and lacks microarchitecture independence.

Method used

Implementing a range prefetch instruction that specifies first and second address range parameters and a stride parameter to control prefetching, allowing for more accurate and efficient prefetching across various microarchitectures.

Benefits of technology

Improves prefetch coverage and reduces resource consumption by allowing prefetchers to learn patterns quickly and adapt to different microarchitectures, enhancing performance and reducing development complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In response to the instruction decoder decoding a range prefetch instruction that specifies the first and second address range specification parameters and the stride parameter, the prefetch circuitry controls prefetching of data from a plurality of specified address ranges into at least one cache according to the first and second address range specification parameters and the stride parameter. A starting address and a size of each specified range depend on the first and second address range specification parameters. The stride parameter specifies an offset between starting addresses of successive specified ranges. The use of range prefetch instructions improves programmability and helps improve the balance between prefetch coverage and circuit area of ​​the prefetch circuitry.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present technique relates to the field of data processing, and more particularly, to prefetching.

[0002] Technology background Prefetching is used in data processing systems to place data into a cache in advance of the time that the data is required by processing circuitry, thereby improving performance by removing the latency associated with handling a cache miss or translation lookaside buffer miss from critical timing paths during the actual demand access made by the processing circuitry. Summary of the Invention

[0003] At least some examples provide an apparatus comprising: an instruction decoder for decoding instructions; a processing circuit for performing data processing in response to decoding of the instructions by the instruction decoder; at least one cache for caching data for access by the processing circuit; and a prefetch circuit for prefetching data into the at least one cache, wherein in response to the instruction decoder decoding a range prefetch instruction specifying first and second address range specification parameters and a stride parameter, the prefetch circuit is configured to control prefetching of data from a plurality of specified address ranges to the at least one cache in response to the first and second address range specification parameters and the stride parameter, wherein a starting address and a size of each specified range are dependent on the first and second address range specification parameters, and the stride parameter specifies an offset between starting addresses of successive specified ranges of the plurality of specified ranges.

[0004] At least some examples provide a method including: decoding an instruction; performing data processing in response to decoding the instruction by the instruction decoder; caching data in at least one cache for access by the processing circuitry; and prefetching data into the at least one cache, wherein in response to decoding a range prefetch instruction specifying first and second address range specification parameters and a stride parameter, the prefetching of data from a plurality of specified address ranges to the at least one cache is controlled in response to the first and second address range specification parameters and the stride parameter, a starting address and a size of each specified range being dependent on the first and second address range specification parameters, and the stride parameter specifying an offset between starting addresses of consecutive specified ranges of the plurality of specified ranges.

[0005] Further aspects, features and advantages of the present technology will become apparent from the following description of examples which should be read in conjunction with the accompanying drawings. [Brief description of the drawings]

[0006] [Figure 1] 1 is a schematic diagram of an example of a data processing system including a prefetch circuit; [Diagram 2] FIG. 13 illustrates an example of a range prefetch instruction. [Diagram 3] 1 illustrates an example of using first and second address range specification parameters of a range prefetch instruction to control prefetching from a given specified range. [Figure 4] 13 shows an alternative code example using a single addressing prefetch instruction. [Diagram 5] Here is a code example that uses the range prefetch instruction: [Figure 6] 13 shows a specific example of the encoding of a range prefetch instruction. [Figure 7] 1 illustrates the use of stride and count parameters to control prefetching of multiple ranges of addresses at intervals of a specified stride. [Figure 8] 1 illustrates a convolution operation that can benefit from the use of a range prefetch instruction. [Figure 9] 1 illustrates different examples of micro-architectural implementations of prefetch circuits. [Figure 10] 1 illustrates different examples of micro-architectural implementations of prefetch circuits. [Figure 11] 1 illustrates different examples of micro-architectural implementations of prefetch circuits. [Figure 12] FIG. 13 is a flow diagram showing control of prefetching based on a range prefetch instruction. [Figure 13-1] 1 shows a flow diagram illustrating in more detail the control of prefetching based on first and second address range specification parameters. [Figure 13-2] 1 shows a flow diagram illustrating in more detail the control of prefetching based on first and second address range specification parameters. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0007] One approach to performing prefetching may be to use hardware prefetching, where a prefetch circuit implemented in hardware monitors the addresses of demand access requests issued by a processing circuit to request access to data, so that a hardware prefetcher can learn the pattern of accesses being made and then issue prefetch requests to bring data for a given address into the cache in advance of expected demand accesses to the same address. With a typical hardware prefetcher, the prefetcher can learn from addresses accessed on demand and from microarchitectural cues such as cache misses to determine how to issue future prefetch requests without explicit hints provided by software based on programmer-specified or compiler-specified parameters of the software prefetch instruction. Although typical hardware prefetching schemes can be effective, they may have some drawbacks. A hardware prefetcher may require a reasonable number of demand accesses to be monitored before the prefetcher can train on and latch onto the stream of addresses found in the demand accesses, and therefore it may be difficult to perform effective prefetching of data for the initial portion of the stream of addresses. If a particular software workload operates on a very small stream, the hardware prefetcher may not have enough samples to identify the stream and prefetch it in a timely manner, since the stream may already have finished by the time the prefetcher has gained sufficient confidence in its prediction. This leads to a loss of prefetcher coverage and therefore a loss of performance. The hardware prefetcher may also need to store a large amount of prefetch prediction state data to track demand access behavior and use it to evaluate the confidence of the predicted address access pattern. For example, a typical hardware prefetcher training unit may take up a hardware budget of tens of kilobytes, which consumes a lot of power and energy.Another problem with typical hardware prefetchers is that they usually do not know when the stream of addresses ends, and therefore, even if the prefetcher is successfully trained to predict future accesses in a particular stream of addresses, the hardware prefetcher tends to continue to prefetch lines a distance beyond the end of the stream after the software being executed has already moved out of that stream and started operating on data in other address regions. Such inaccurate prefetching can waste available memory system bandwidth and cause cache pollution (e.g., replacement of more useful cache entries with over-prefetched data), which can reduce the cache hit ratio for other accesses and therefore cost performance.

[0008] Another approach to prefetching may be to use software prefetching, where a programmer or compiler includes a prefetch instruction in the program code to be executed that specifies a single target address of data expected to be used in the future, and a prefetch circuit responds to the single-address-specifying prefetch instruction to control the prefetching of data corresponding to the single address. For example, a prefetch circuit may respond to a single-address-specifying software prefetch instruction to issue a request for a single cache line of data to be brought into the cache. By providing a software prefetch instruction, a programmer or compiler can provide an explicit indication of future intent to access a particular address to prepare the cache for a later demand access.

[0009] However, software prefetching using such single-addressing prefetch instructions still has some drawbacks. Programmers may find it very difficult to add these software prefetch instructions to their code. Since each instruction specifies a single address, in practice the code needs to include many such software prefetch instructions corresponding to different addresses that the code is expected to access later. The point in the program code where to insert the prefetch instructions is not always obvious to the software developer. Also, since the effectiveness of the prefetch depends on how far in advance of a particular demand access the corresponding single-addressing range prefetch instruction is inserted, it may be difficult to control the timeliness of the prefetch with this approach. If the single-addressing prefetch instruction is included too early, the prefetched data may already have been evicted from the cache by the time the subsequent demand access occurs, wasting resources associated with performing the prefetch. If a single addressing range prefetch instruction is included too late, the prefetched data may not yet be resident in the cache by the time a subsequent demand access is encountered, and thus the demand access may still encounter a cache miss. Thus, with a typical software prefetch instruction that specifies a single address, getting the timing right is a trial and error process, which means that software developers spend a lot of development time figuring out the best location to add the software prefetch instruction.

[0010] Another problem with conventional software prefetching is that program code using single-address software prefetch instructions becomes specific to the particular microarchitecture of the processor core and memory system. Code written according to a particular instruction set architecture can run on a wide variety of different processing systems with different microarchitecture designs, which may experience different latencies when data is fetched from memory to the cache and may have different cache hit / miss ratios, depending on characteristics such as the number of levels of cache provided, the size of cache storage provided at a given cache level, the cache line size (the size of the basic unit of memory transferred between memory and cache), the number of other sources of access requests competing for memory system bandwidth, and many other particular characteristics of the particular processor implementation in hardware. These factors mean that when code runs on one microarchitecture platform, the preferred timing for a prefetch may be earlier than when the same code runs on another microarchitecture platform. This means that it is very difficult to write code using single-address software prefetch instructions that can execute across a range of microarchitectures with efficient performance. The development costs of generating different versions of code for different microarchitectures become very high.

[0011] In an embodiment described below, an apparatus includes an instruction decoder for decoding instructions and a processing circuit for performing data processing in response to the decoding of the instructions by the instruction decoder. At least one cache is provided for caching data for access by the processing circuit. A prefetch circuit is provided for prefetching data into the at least one cache. The instruction decoder supports a range prefetch instruction specifying first and second address range specification parameters and a stride parameter. In response to the instruction decoder decoding the range prefetch instruction, the prefetch circuit controls prefetching of data from a plurality of specified address ranges into the cache in response to the first and second address range specification parameters and the stride parameter, a starting address and a size of each specified range being dependent on the first and second address range specification parameters, and the stride parameter specifying an offset between starting addresses of successive specified ranges of the plurality of specified ranges.

[0012] Thus, the range prefetch instruction may be used by software to signal that some specified range of memory addresses will be accessed in the future. The stride parameter allows a single instruction to provide a prefetch hint that there is expected to be a subsequent access pattern of accessing several blocks of data of a particular size that are discontinuous in address space and separated by a fixed stride interval. This pattern of access may be useful for many machine learning applications, such as machine learning models that use convolutions of kernels of weights with matrices of activation data to train the machine learning model. Such convolutions may rely on applying kernels to different positions in the matrix, and for each kernel position, this may require reading several relatively short chunks of data representing portions of different rows or columns in the matrix, with gaps between the address ranges accessed for the activation value used for a given kernel position. By specifying a stride pattern of discontinuous ranges, it becomes much easier to efficiently train the prefetch circuitry to learn to effectively prefetch such a pattern of accesses. It will be appreciated that such folding is just one example of an access pattern that can benefit from a range prefetch instruction that specifies the bounds and stride of a range of addresses, and that other operations can also benefit from this.

[0013] The prefetch circuitry can then respond to these hints to control how data from the specified address range is prefetched into the cache. Compared to the single address-specified prefetch instruction described above, the range prefetch instruction is much easier for a programmer or compiler to use, since a single range prefetch instruction can be included that specifies the address associated with the portion of a given data structure that is expected to be processed, rather than inserting multiple prefetch instructions corresponding to individual addresses. In addition to simplifying software development, this also significantly reduces the number of prefetch instructions that need to be included in the code, reducing the amount of fetch, decode, and issue bandwidth consumed by the prefetch instruction, allowing for greater throughput for other types of instructions. A programmer using a range prefetch instruction does not need to worry about the timeliness of the insertion of a given prefetch instruction relative to a demand access, since the timeliness can be controlled in hardware by the prefetch circuitry within the bounds of the specified address range provided as a hint by the range prefetch instruction. The use of first and second address range specification parameters and a stride parameter in the range prefetch instructions also means that applications become much less microarchitecture and system architecture dependent compared to the use of prefetch instructions that specify a single address, making it much easier to write program code that can run across a range of different microarchitecture processor implementations.

[0014] On the other hand, compared to a typical hardware prefetch that does not use any software-specified hints explicitly provided by the prefetch instruction, the use of a range prefetch instruction is beneficial because by explicitly specifying parameters that identify the boundaries of a range of addresses expected to be accessed in the future, this means that the prefetch circuitry can be quickly prepared to prefetch from those ranges, rather than requiring a long initial training period when accesses begin to be made from a stream of addresses with predictable characteristics. Thus, in practice, when accesses are made to data within a particular range corresponding to a given data structure, in a pure hardware prefetch implementation, it is unlikely that the first part of that range will be successfully prefetched, as the prefetcher may still be gaining confidence. In contrast, the use of a range prefetch instruction allows the start of the range to be identified using the first address range-specified parameters, thus significantly reducing warm-up time and improving prefetch success rates, and therefore performance. A range prefetch instruction may be particularly useful when the stream of demand addresses that follow a predictable pattern is very small and difficult to effectively prefetch using a pure hardware prefetcher.

[0015] Similarly, since the size of a given one of the address ranges may be identifiable from the first and second address range specification parameters, this reduces the problem of over-prefetching beyond the end of the data structure being accessed, avoiding unnecessary consumption of prefetch bandwidth and possible eviction of other data by over-prefetched data that may cause performance issues in a pure hardware prefetcher. In general, for a given amount of hardware circuit area budget, a prefetcher that uses hints provided by range prefetch instructions may provide greater prefetch coverage (a greater proportion of demand accesses that hit previously prefetched data when they would otherwise miss in the cache) and therefore better performance than a comparable hardware prefetcher that learns prefetching purely from implicit queues inferred from monitoring microarchitectural information such as demand accesses and access hit / miss information. On the other hand, for a given level of prefetch coverage, a prefetcher that may be controlled based on range prefetch instructions may require less circuit area than a conventional hardware prefetcher that does not support a range prefetch queue received from a range prefetch instruction. This is because to maintain a comparable level of performance without these range prefetch queues, a much larger set of training data would typically need to be stored and compared to demand access.

[0016] So, in summary, by supporting range prefetch instructions, this can provide a better balance between prefetch coverage (performance) and circuit area and power consumption than is possible using traditional software and hardware prefetching schemes, and also allows for better programmability such that it is easier for software developers to write code that includes prefetch instructions.

[0017] The first and second address range specification parameters may be any two software-specified parameters of the range prefetch instruction that may be used to identify the boundaries of a selected one of the address ranges. For example, the selected range may be a first range of a plurality of address ranges identified using the instruction. The boundaries of the other ranges may be determined by applying an offset indicated by the stride to the boundaries of the selected one of the ranges.

[0018] The first and second address range parameters may be specified in different ways. For example, the first and second address range parameters may each specify a memory address as an absolute address or by referencing a reference address other than the first and second address range parameters themselves. For example, the first address range parameter may specify a starting address of the selected range and the second address range parameter may specify an ending address of the selected range absolutely as an offset relative to a reference address such as a program counter indicating the current point reached during execution.

[0019] However, since it is relatively unlikely that the size of the specified range will be larger than a certain amount, fewer bits can be used when one of the start address and end address is specified relative to the other. In some examples, the start address may be specified relative to the end address. However, it is often more natural for the start address of the specified range to be encoded as a base address in the operand of the range prefetch instruction, and the end address to be expressed, for example, by a size parameter or offset value relative to the start address. Thus, one useful example may be for a first address range specification parameter to include a base address (start address) of the specified range, and for a second address range specification parameter to include a range size parameter that specifies the size of each of the specified ranges. This allows the end address of the specified range to be determined based on applying an offset corresponding to the specified size to the base address. The specified base and size may specify the boundaries of the first specified range, and then the stride may be used to identify the start / end of subsequent ranges in the strided range group. Each of the ranges in the strided group of ranges may have the same size.

[0020] The first and second address range specification parameters may be encoded in the range prefetch instruction in different ways, for example, each of the first and second address range specification parameters may be encoded using an immediate value directly encoded in the instruction encoding of the range prefetch instruction, or using a register identified by a corresponding register specifier in the encoding of the range prefetch instruction.

[0021] Where the first address range specification parameter encodes a base address and the second address range specification parameter encodes a range size, a useful example may be that the first address range specification parameter is encoded in a first register specified by the range prefetch instruction and the second address range specification parameter is encoded in a second register specified by the range prefetch instruction, or by an immediate value specified directly in the encoding of the range prefetch instruction.

[0022] Regardless of how the address range specification parameter is encoded, the encoding of the address range specification parameter may, in some instances, support indicating a given specified address range as spanning a block of addresses corresponding to a number of bytes other than an exact power of two. For example, the size of the specified range may be any selected number of bytes up to a certain maximum size. The maximum size supported for a given specified address range may be larger than the size of a single cache line, where a cache line is the size of a unit of data that may be transferred between caches or between a cache and memory in a single transfer. Thus, the range of addresses that may be specified by the range prefetch instruction is not limited to a single cache line, but may be a range of any size, and thus is not limited to blocks of power of two size, allowing software to select the specified range to match the boundaries of a portion of a particular data structure that the software processes, which is generally not possible with a single address specification prefetch instruction that brings a single cache line into a cache corresponding to a block of bytes of a particular power of two size, and thus requires further single address specification prefetch instructions to be executed to prefetch other portions of a data structure that spans multiple cache lines or has a size that is not a power of two. Similarly, the encoding of the stride parameter may support indicating the offset as a number of bytes other than an exact power of two.

[0023] Alternatively, other examples may limit the encoding of size and / or stride to units of a power-of-two number of bytes, and therefore may not support specification of size or stride at single-byte granularity, for example, a size or stride parameter may specify a multiple of a particular base unit size, e.g., 2, 4, or 8 bytes.

[0024] It will be appreciated that other techniques for encoding the range size or stride may also be used.

[0025] In some examples, the range prefetch instruction may also specify a count parameter that indicates how many ranges separated by a stride offset should be prefetched. This allows the prefetcher to determine when to stop prefetching from multiple ranges. The stride and count parameters may be encoded in different ways, such as using an immediate value or a value in a register. The stride and count values ​​may be encoded using different registers than the registers used for the first or second address range specification parameters.

[0026] However, in one example where the first address range parameter specifies a base address and the second address range parameter is specified as a size or offset relative to the base address indicated by the first address range parameter, this may require fewer bits for the second address range parameter than the first, and therefore there may be some spare bits in the register used to provide the second address range parameter. Thus, in some cases, at least one of the stride parameter and the count parameter may be encoded in the same register as the second address range parameter. This reduces the number of register specifiers required in the range prefetch instruction encoding and saves encoding space that may be reused for other purposes.

[0027] It will be appreciated that although the range prefetch instruction supports the stride and count parameters described above, the range prefetch instruction may also support encodings in which only a single range of addresses to be prefetched is indicated, with the boundaries of that range being determined based on the first and second address range specification parameters. For example, when the count or stride parameters have a particular default value, this may indicate that only a single range should be prefetched. Thus, while it is not required that every instance of a range prefetch instruction triggers prefetching from multiple ranges separated in an instance of stride, the range prefetch instruction nevertheless has an encoding that allows the stride parameter (and optionally the count parameter) to be determined such that a programmer has the option of using a single instruction to prepare a prefetcher to prefetch data from addresses in multiple ranges separated by a fixed stride interval.

[0028] The first and second address range specification parameters and the stride parameter provided by the range prefetch instruction may not be the only information used by the prefetch circuitry to control the prefetching of data from the specified address range in at least one cache. The prefetching may also depend on at least one of monitoring the addresses or loaded data associated with the demand access requests issued by the processing circuitry to request access to the data, and microarchitectural information associated with the prefetch request or the results of the demand access request. Thus, while an indication of the boundaries of the range of addresses expected to be required by the demand accesses may be useful to prime the prefetcher, reduce prefetch warm-up time, avoid excessive prefetching, and simplify software development for the reasons discussed above, within the specified range, the prefetch hardware may have the flexibility to make its own decisions regarding which addresses within the specified range should be prefetched and when to issue those prefetches. By monitoring microarchitectural information indicative of the address and / or data of the demand access request, as well as the prefetch request or the result of the demand access request, the prefetch circuitry can make microarchitecturally dependent decisions to enable improved prefetch coverage for a given microarchitectural implementation, but guided by the range specification parameters of the range prefetch instruction, which may require less circuit area budget than an alternative hardware prefetch where such range hints were not provided. Thus, prefetching may be dynamically adjusted by the prefetch circuitry in hardware, but guided by information about range boundaries provided by the range prefetch instruction as defined by software.For example, in scenarios where it is recognized that data loaded from memory can be used, for example, as a pointer used to determine the address of a subsequent access, the loaded data associated with a demand access request can be used to control prefetching, and information from the loaded data of one demand access can be used to predict which addresses may be accessed in the future.

[0029] When microarchitectural information is used to regulate prefetching, this microarchitectural information may include any one or more of prefetch linefill latency (indicative of the latency between issuing a prefetch request and corresponding prefetched data being returned to at least one cache), demand access frequency (how often demand access requests are issued by the processing circuitry), demand access hit / miss information (e.g., providing information indicative of the percentage of demand access requests that hit or miss at a given level of cache, and in some cases this information may be maintained for different levels of cache), demand access prefetch hit information (e.g., indicative of the percentage of demand access requests that hit previously prefetched data as opposed to hitting previously fetched data on demand in response to a demand access), and prefetch utility information (e.g., an indication of the percentage of prefetch requests whose prefetched data is subsequently used by a demand access before the prefetched data is evicted from the cache). Any given implementation need not consider all of these types of microarchitectural information. Any one or more of these types may be considered.

[0030] The prefetch circuitry may control when a prefetch request is issued for a given address within a number of specified ranges based on monitoring addresses or loaded data associated with demand access requests issued by the processing circuitry to request access to data, or microarchitectural information associated with the prefetch request or the results of the demand access request. This differs from conventional software prefetch instructions, which typically act as a direct trigger for a specified address to be prefetched, rather than leaving flexibility for the hardware to determine the exact timing. Thus, the prefetch range instruction can be viewed as a mechanism for software to prime the prefetcher with hints regarding future address access patterns, rather than a trigger for the prefetch to necessarily be performed correctly, so that the hardware can make the final decision regarding prefetch timing depending on the microarchitectural queue, and the software can be more platform independent.

[0031] In some examples, the prefetch circuitry may include a sparse access pattern detection circuit for detecting a sparse pattern of accessed addresses within a given specified range of a plurality of ranges of addresses indicated by the range prefetch instruction based on monitoring the addresses of the demand access requests, and selecting which particular addresses within the specified range should be prefetched based on the detected sparse pattern. Thus, a microarchitecture implementation having a sparse access pattern detection circuit may benefit from the use of a range prefetch instruction even when the expected usage pattern is sparse and all addresses within the given specified range are unlikely to be accessed. For example, the sparse access pattern detection circuit may include a training circuit that monitors a sequence of addresses to detect strided address patterns or alternating sequences of stride intervals within a stream of demand addresses, and may store predicted state information indicative of the detected patterns and associated confidence levels to learn which addresses are most effective to prefetch. However, guided by the address range specification parameters from the range prefetch instruction, it can more quickly prime the sparse access pattern detection circuitry to enable better prefetching of the initial portion of the specified range, and hints from the range prefetch instruction can indicate to the sparse access pattern detection circuitry when the detected pattern ends, thus avoiding excessive prefetching.

[0032] Note that if a range prefetch instruction specifies multiple specified ranges using a stride parameter, the sparse access pattern detection circuitry may still be applied to detect further strided patterns of address accesses within each of those specified ranges. In that case, the specified ranges identified by the range prefetch instruction may be interpreted as boundaries of regions where future demand accesses may be required, but the sparse access pattern detection circuitry may learn which particular addresses within those regions are actually accessed and train the prefetch circuitry accordingly.

[0033] Other implementations may not have a sparse access pattern detection circuit and may choose to use a simpler hardware implementation of a prefetch circuit, which may assume that the access pattern within a given specified range is a dense access pattern, such that every address within the specified range should be prefetched.

[0034] Another aspect of prefetch control that may be controlled in hardware may be an advance distance used to control how far ahead the stream of prefetch addresses of prefetch requests issued by the prefetch circuitry is compared to the stream of demand addresses of demand access requests issued by the processing circuitry to request access to data. Based on the advance distance, the prefetch circuitry may control when a prefetch request for a given address is issued.

[0035] It may be useful for the prefetch circuitry to define the advance distance independent of explicit specification of the advance distance by the range prefetch instruction (or implicit specification of the advance distance based on the position of the range prefetch instruction relative to other code). For example, the setting of the advance distance may be based on a particular microarchitectural implementation of the prefetch circuitry, which may be independent of the software code since the same software code may be executed on different microarchitectural implementations. By not requiring software to control the advance distance (unlike an approach using a single addressing prefetch instruction, where the programmer must consider the point at which the software prefetch instruction is inserted based on the expected advance distance that is estimated to be useful), this makes it much easier for software developers to generate platform-independent code that can execute efficiently across a range of microarchitectures. For example, the prefetch circuitry may adjust the advance distance based on microarchitectural information associated with the outcome of the prefetch request for the demand access request. This microarchitectural information may be any of the types of microarchitectural information previously mentioned.

[0036] Different microarchitecture implementations of the prefetch circuitry may respond to a range prefetch instruction in different ways. A simple approach may be to simply issue prefetch requests corresponding to addresses in the specified range (or at least addresses in the first of the specified range) directly in response to the range prefetch instruction. In practice, however, this may risk that prefetching later parts of the specified range wastes memory bandwidth, because by the time the demand access stream reaches later parts of the specified range, the previously prefetched data is likely already evicted from the cache and therefore no longer benefits from the prefetch. Another approach is that the prefetch request is not issued directly in response to the range prefetch instruction, but the range specification parameters and stride can be used to store information about the boundaries of the specified address range, and then, when a demand access request for an address in the specified range is detected, this can be used to quickly latch onto the newly detected address stream and trigger further prefetching of subsequent addresses in the range. However, while this approach may still provide better performance than a pure hardware prefetching approach that does not use any software prefetch instructions, waiting for a demand access to reach a given specified address range may run the risk of prefetch coverage being less effective for the first portion of a given specified address range, because by the time the first demand access reaches that range, it may be too late to prefetch data from that first portion in time for the demand access to take advantage of the prefetched data.

[0037] Thus, one useful approach may be that in response to the instruction decoder decoding a range prefetch instruction, the prefetch circuitry may trigger a prefetch of data from an initial portion of a first of the multiple specified ranges to at least one cache. This prefetch of the initial portion may be performed directly in response to the range prefetch instruction, regardless of whether any demand accesses associated with addresses in the specified range have been detected. By prefetching data from the initial portion in direct response to the range prefetch instruction, this may improve prefetch coverage and subsequent demand access cache hit rates, and reduce warm-up penalties associated with conventional hardware prefetchers.

[0038] The size of the initial portion prefetched in response to the range prefetch instruction may be defined in hardware by the prefetch circuitry, independent of the explicit specification of the size of the initial portion by the range prefetch instruction. Thus, software does not need to specify how much of the specified range should be directly prefetched as the initial portion in response to the range prefetch instruction. The size of the initial portion may be fixed for a particular microarchitecture by the hardware designer, or may be dynamically (at run-time) adjustable based on microarchitectural information regarding the results of the prefetch request or the demand access request, as described above. This allows different microarchitectures executing the same code to make their own decisions about how much initial data to prefetch in response to the range prefetch instruction. The size of the initial portion may depend on parameters such as the number of cache levels provided, the size of the cache capacity at different levels of the cache, the latency in accessing the memory, and the frequency or amount of other accesses made to the memory system.

[0039] After prefetching the first portion of the specified range, the prefetching of data from a given address within the remaining portion of the first specified range may not occur directly in response to the range prefetch instruction. Instead, the prefetch circuitry may trigger the prefetching of data from a given address within the remaining portion in response to detecting that the processing circuitry issues a demand access request that specifies a target address that is ahead of the given address by no more than the given distance. By postponing the prefetching of data from a given address until the demand access reaches an address within the given distance of the given address, this increases the likelihood that the prefetched data for the given address will still be present in the given level of cache when a subsequent demand access request for that address is issued, increasing the cache hit rate and improving performance. Again, the size of the given distance may not be software defined, but may be hardware determined by the prefetch circuitry independent of the explicit specification of the given distance by the range prefetch instruction. Again, the given distance may be a fixed distance set for a particular microarchitectural implementation (which may differ from the given distance set on other microarchitectural implementations supporting the same program code) or may be adjustable at run-time based on microarchitectural information as described above. It may be useful for one microarchitecture to have a given distance larger than another to account for different delays in bringing data into a given level of cache.

[0040] The size of the first portion and the given distance mentioned in the previous paragraph are both examples of parameters representing the aforementioned advance distance that govern how far ahead of the stream of prefetch addresses is issued ahead of the stream of demand addresses.

[0041] When a range prefetch instruction specifies multiple ranges using a stride parameter, prefetching of second, third, and further of the multiple specified ranges may be controlled in different ways. For example, the prefetch circuitry may monitor demand accesses to determine timing for these streams. For example, prefetching from a given specified range may begin when a demand access arrives within a given distance of the end of the previous specified range.

[0042] Also, some implementations of prefetch circuitry may not necessarily take the same approach to controlling prefetch timing. For example, the preferred prefetch timing may be different for ranges below a particular threshold compared to ranges above that threshold. It will be appreciated that the particular prefetch timing control example described herein is merely an example and other implementations may choose different approaches.

[0043] The prefetch circuit may be configured such that, following decoding of the range prefetch instruction by the instruction decoder, prefetching for the given specified range may be stopped when the prefetch circuit identifies, based on the second address range specification parameter, that the end of the given specified range has been reached by the stream of prefetched addresses prefetched by the prefetch circuit or the stream of demand addresses specified by the demand access requests issued by the processing circuit, which would not be possible with a typical hardware or software prefetching scheme using a single address specification prefetch instruction. Thus, by preventing excessive prefetching beyond the end of the specified range, it is possible to avoid polluting the cache with unnecessary data that is unlikely to be used, helping to preserve cache capacity for other data that may be more useful, improving performance.

[0044] In some examples, the range prefetch instruction may also specify one or more prefetch hint parameters to provide hints regarding data associated with the specified range, and the prefetch circuitry may control prefetching of data from the specified address range into at least one cache in response to the one or more prefetch hint parameters. For example, the prefetch hint parameters may be any information other than information identifying the addresses of the specified range that provides a cue regarding estimated usage of data in the specified range that may be used to control prefetching.

[0045] For example, the one or more prefetch hint parameters may include any one or more of the following: A type parameter indicating whether the prefetched data is data expected to be the target of a load operation, data expected to be the target of a store operation, or an instruction expected to be the target of an instruction fetch operation. If the data is expected to be the target of a load or store operation, this may trigger the data to be prefetched into a data cache, and if the prefetched data is expected to be an instruction, the prefetch may be to an instruction cache. Thus, the target cache (which is the recipient of the prefetched data) may be selected according to the type parameter. Also, an indication of whether the data is expected to be the target of a load or store operation may be useful in selecting the coherency state in which the prefetched data is required. For example, in a system with multiple caches that cache data from a shared memory system, a coherency scheme may be used to manage coherency between different cached copies for a given address. If the data is expected to undergo a store operation such that software using the cached data updates the data value associated with the corresponding address, it may be preferable to obtain the cached copy of the data in a unique coherency state such that other copies of the data from the same address in other caches are invalidated to avoid other caches holding stale data after the store operation. On the other hand, for data fetched into a cache that only receives load operations, since the load is not expected to overwrite the data, it may be possible to place the data in the cache in a shared coherency state where corresponding versions of the data from the same address may still exist in other caches. It may therefore be useful for the prefetch circuitry to control the coherency state in which data from a specified address range is prefetched based on a type parameter provided as a prefetch hint parameter of the range prefetch instruction.

[0046] A data use timing parameter indicating at least one of an indication of how soon the prefetched data associated with the specified range is estimated to be required by the processing circuitry following a range prefetch instruction, and a target level in the cache hierarchy to which the data is prefetched. The data use timing parameter may be used to influence the advance distance at which a prefetch request is issued for the estimated demand access, and / or which level of cache hierarchy data is prefetched. The data use timing parameter may be expressed in different ways. In some cases, it may simply be a vague relative indication of whether the data may be needed "soon" or "less than soon" (or may indicate three or more qualitative levels of expected use timing), but need not quantify any particular use timing. In some cases, how the prefetch circuitry responds to a particular use timing parameter value may be microarchitecture dependent, and thus software may use the data use timing parameter to express a vague intent of the relative timing of when access to the specified address is likely to be required, but different microarchitectural implementations may respond to that vague indication in different ways. Other implementations may specify data usage timing parameters more explicitly, for example, by specifying a particular level in the cache hierarchy at which the data is prefetched.

[0047] A data reuse timing parameter indicating at least one of an indication of an estimated interval between a first demand access to the prefetched data and a further demand access to the prefetched data, or a selected level of the cache hierarchy from which the prefetched data is preferably evicted if it is replaced from a given level of the cache hierarchy. The data reuse timing parameter may be used to control cache replacement of entries in the cache. While the data reuse timing parameter described above is a hint regarding the timing of a first demand access to the prefetched data after the prefetch, the data reuse timing parameter is a hint regarding the timing of a next access after the first demand access. When prefetching data, the data reuse timing parameters may be used to set metadata associated with the prefetched data in a cache, which may be used to determine whether to evict the prefetched data when another address is allocated to the cache (e.g., cache entries indicated as less likely to be reused soon may be favored for eviction compared to cache entries indicated as more likely to be reused soon), or, if the prefetched data is evicted, to which subsequent level of cache the evicted data should be assigned. Again, the data reuse timing parameters may be a relatively precise indication of the expected interval between a first demand access and a further demand access (e.g., an approximate indication of the number of instructions or number of memory accesses), or an indication of a particular level of cache from which to evict the data, or even a vague relative indication of whether the data is likely to be needed "soon" or "less than soon" following a first demand access to that data (or an indication of three or more qualitative levels of expected data reuse timing). Again, the particular manner in which a particular micro-architectural implementation of a prefetch circuit or cache uses the data reuse timing parameters may vary from implementation to implementation.

[0048] · An access pattern hint indicating the relative density / sparseness of the pattern of accesses to addresses in the specified range, or the sparse access pattern type. This type of prefetch hint may be useful to allow the prefetch circuitry to determine whether it should prefetch all addresses in the specified range, or only some addresses. For example, the access pattern hint may be used by the prefetch circuitry to determine whether it is worth expending resources in training the sparse access pattern detection circuitry described above to learn a particular sparse access pattern within a given specified range. In some cases, the access pattern hint may simply indicate some relative indication of density or sparseness, such as, for example, a single bit indicating whether the accesses are dense (i.e., all or nearly all addresses are expected to be accessed) or sparse (fewer addresses are expected to be accessed from the specified range). Other implementations may provide finer granularity of the specification of the relative density / sparseness. Some examples may specify a particular sparse access pattern type. For example, a sparse memory access pattern type may indicate whether the access pattern is likely to be strided, spatial, temporal, random, etc. Such hints about expected access types can be useful to simplify the training of a sparse pattern detector.

[0049] A reusability parameter indicating the likelihood that data prefetched from a specified address range is expected to be used once or multiple times. Such a reusability parameter can be used to control how long data prefetched from a specified address range is retained in the cache after it has been prefetched or accessed by a demand access. For example, the reusability parameter can indicate whether the stream of accesses to the specified address range is expected to be of a non-temporal (streaming) type of access, where each access to a given address within those ranges uses the data once and then the address is likely not needed again, or whether the accesses are a temporal pattern of accesses, where data prefetched from a given address within the specified range is likely to be accessed more than once or more than a certain number of times. Thus, the reusability parameter can be used by the prefetch circuitry to control storage in the cache of some metadata associated with the prefetched data, and the cache can be used to determine whether to prioritize the prefetched data for eviction after the prefetched data has been accessed by a demand access.

[0050] A no-prefetch hint parameter indicating that prefetching of data from the specified range should be inhibited. In some examples, the prefetch range instruction may also support an encoding in which a no-prefetch hint parameter may be provided indicating that the prefetch circuitry should not prefetch any data from the specified range. This may be useful because it allows the prefetch circuitry to avoid expending training resources in learning information about access patterns within the specified range. This may be useful to conserve prefetch training resources for other ranges of addresses where the benefit of prefetching may be greater.

[0051] Any given implementation of the range prefetch instruction need not support all of these types of prefetch hint parameters. Any one or more of these prefetch hint parameters may be supported (or none of these prefetch hint parameters may be supported). The encoding of the prefetch hint parameters may vary widely from implementation to implementation. In some cases, the encoding of two or more of these hint parameters may be combined into a single field, with different prefetch hint code values ​​specified in that field being assigned to different patterns of prefetched hint combinations. Thus, it is possible, but not required, to identify each prefetch hint parameter using an entirely separate parameter of the instruction.

[0052] In some examples, the prefetch circuitry may set cache replacement hint information or cache replacement policy information associated with data prefetched from the specified range specified by the range prefetch instruction in response to one or more prefetch hint parameters. The cache may then control replacement of cache entries based on the cache replacement hint information or cache replacement policy information set by the prefetch circuitry. Here, the cache replacement policy information may be information used directly by the cache to control replacement of cache entries, e.g., to control the selection of which of a candidate set of entries should be the entry to be replaced with the newly allocated data. For example, the cache replacement policy information may be least recently used (LRU) information to determine which of a set of candidate cache entries has been accessed least recently, or may specify a priority scheme indicating a relative level of priority of a given cache entry for replacement. In general, any known cache replacement policy may be used, but in the approach where the range prefetch instruction specifies prefetch hint parameters, these hints may be used to set this cache replacement policy information associated with a given cache entry when data from the specified range is prefetched into that entry. Alternatively, the use of a prefetch hint parameter may not directly affect the setting of replacement policy information associated with a particular cache entry, but may simply be used to set cache replacement hint information that may indirectly control cache replacement, such as a hint that entries associated with addresses in a given address range may be more or less favorable for eviction. Cache replacement hint information may be information indicating that a given entry of the cache may be evicted (or marked as prioritized for eviction) after a first demand access hits that entry.While there are many ways in which cache replacement can be implemented through prefetch hint parameters, it will be appreciated that in general, the prefetch circuitry can use prefetch hint parameters for a range of prefetch instructions to influence cache replacement, thereby making it more likely that the cache hit rate can be maintained higher based on the expected usage patterns exhibited by the software using the prefetch hint parameters.

[0053] Examples of data processing devices FIG. 1 shows a schematic example of a data processing apparatus 2. The apparatus includes an instruction decoder 4 for decoding program instructions fetched from an instruction cache or memory. Based on the decoded instructions, the instruction decoder 4 controls a processing circuit 6 to perform data processing operations corresponding to the instructions. The processing circuit 6 has access to registers 8 used to provide operands of the instructions and to store results of the processed instructions. The processing circuit 6 includes several execution units 10, 12 for executing different classes of program instructions. By way of example, FIG. 1 shows a processing circuit 6 including an arithmetic / logic unit (ALU) 10 for performing arithmetic or logical operations and a load / store unit 12 for performing load operations for loading data from a memory system into the registers 8 or store operations for storing data from the registers 8 into the memory system. It will be understood that these are only some examples of possible execution units and that other types of execution units may be provided or that multiple execution units of the same type may be provided.

[0054] The memory system in this example includes an instruction cache (not shown in FIG. 1 for simplicity), several data caches 14, 16, 18 implemented in a cache hierarchy, and a memory 20. In this example, the cache hierarchy includes three levels of caches, including a level 1 data cache 14, a level 2 cache 16, and a level 3 cache 18. Some of these caches may also be used for instructions (e.g., a level 2 or level 3 cache may be shared for data and instructions). Other examples may have a different number of caches in the hierarchy, or a different organization of data caching relative to instruction caching.

[0055] When a load / store operation is executed in response to a load / store instruction decoded by the instruction decoder 4, the virtual address specified using the operand of the load / store instruction is translated into a physical address by the memory management unit (MMU) 22 based on address mapping information specified in a page table entry of a page table structure stored in the memory system. Information from the page table structure may be cached in a translation lookaside buffer (TLB) within the MMU 22.

[0056] The address of a load / store operation is looked up in the cache, and if there is a hit at a given level of the caches 14, 16, 18, the data may be made available faster than if the data had to be fetched from memory 20. A hit in a cache higher in the cache hierarchy (e.g., level 1 data cache 14) allows the data to be made available faster than a hit at a lower level of the cache hierarchy closer to memory 20, but higher level caches, such as level 1, have smaller capacity, and therefore there is a trade-off between cache access latency and capacity at each level.

[0057] The data access requests issued by the load / store unit 12 are demand access requests issued on demand based on addresses specified by load / store instructions of the executing software program. If the data is not brought into the caches 14, 16, 18 until the program reaches a point of execution where the corresponding load / store instruction is executed, there may be additional latency caused by a cache miss since there is a delay while the data is brought from a lower level cache or memory. Also, if the virtual address for a demand access misses in the TLB 24, the MMU 22 may need to perform a page table walk to traverse a page table structure stored in memory to find address mapping information for translating the virtual address. Such a page table walk may cause significant delays in processing the demand access.

[0058] The prefetch circuitry 30 is provided to bring data into a given level of the caches 14, 16, 18 prior to the time that the processing circuitry 6 issues a demand access request corresponding to the prefetched address. This means that delays associated with a TLB or cache miss may be encountered when performing the prefetch rather than at the time of the demand access, and thus when the prefetch is successful, this reduces the critical timing path of processing a demand access that would otherwise miss in the cache.

[0059] The prefetch circuitry 30 may use a variety of sources to maintain a set of prefetch state information 32 that may be used to predict which addresses may be required by future demand accesses, and may issue prefetch requests 34 that request that data for a particular prefetch address be brought to a particular level of the cache 14, 16, 18 based on the prefetch state information 32. For example, the information used to predict the prefetch address may include a stream of demand addresses 36 specified by demand access requests issued by the load / store unit 12, microarchitectural information 38 that indicates information about the results of the demand access requests or prefetch requests 34 issued by the load / store unit 12, as well as software-defined hints 40 obtained from prefetch instructions decoded by the instruction decoder 4. The demand addresses 36 may be either virtual addresses or physical addresses, and the prefetch circuitry 30 may be trained on either.

[0060] Software vs. Hardware Prefetching Software prefetching is used to pre-fill cache lines before they are accessed in the future. This removes the latency of cache fills and TLB fills from the critical path during the actual demand access, thereby improving performance. Traditional software prefetching only brings up one cache line at a time (because the software prefetch instruction only specifies a single address). An alternative is to rely on a hardware prefetcher to learn the pattern of accesses and issue prefetches in advance of demand accesses that achieve the same benefit. However, both traditional software prefetching and hardware prefetching have several drawbacks.

[0061] Disadvantages of traditional software prefetching: 1. Programmers know that adding a traditional software prefetch instruction that specifies a single address is difficult. The pattern of accesses is not obvious in software. Plus, getting the timeline right is a trial and error process. Getting it to work perfectly is very difficult.

[0062] 2. A kernel with traditional software prefetching is micro-architecture and system architecture specific. Even the size of a cache line (the basic unit of memory that is prefetched) varies between implementations, which will break when run on different systems.

[0063] 3. If the system supports a scalable vector architecture, where the instruction set architecture allows the microarchitect designer to choose different sizes for the vector registers, and the instruction set architecture is designed so that the same code can run on different microarchitectures with different vector sizes, then developing code that uses software prefetch instructions becomes more difficult because the vector length is unknown, and therefore it is not possible to add prefetches at the correct rate in a general way, even if conservative assumptions about cache line lengths are made.

[0064] 4. If it turns out that software prefetching actually hurts performance instead of improving it (e.g., the prefetch timeliness was incorrect so there was no performance benefit to the prefetch and the prefetch evicted other data which in turn caused additional cache misses), then there is no way to fix this at run time and the code would need to be rewritten.

[0065] 5. Traditional software prefetching targets a single line. The tradeoff between code bloat, complexity, and accuracy is difficult. For example, software may end up over-prefetching to avoid replicating loops that omit prefetching.

[0066] Disadvantages of traditional hardware prefetching: 1. The hardware prefetcher makes several demand accesses to train and latch into the stream. If the stream is small, the hardware prefetcher does not have enough samples to identify the stream and prefetch in time. This leads to a loss of coverage.

[0067] 2. If there are too many parallel streams with interleaved accesses between them, the hardware prefetcher finds it more difficult to identify and latch onto the streams. A training unit of the hardware prefetcher can take up tens of KB in the hardware budget and can also consume a lot of power and energy.

[0068] 3. The hardware prefetcher does not know when a particular stream of accesses will end, and therefore tends to prefetch lines beyond the end of the stream. Such inaccurate prefetching can take away available bandwidth and cause cache pollution.

[0069] Range Prefetch Instructions 2 illustrates a range prefetch instruction that can help address these shortcomings. The instruction set architecture supported by the instruction decoder 4 and processing circuitry 6 may include a range prefetch instruction that specifies first and second address range specification parameters 50, 52, as well as a stride parameter 56. In this example, the instruction also specifies an additional prefetch hint 54 and a count parameter 58.

[0070] In this example, the first address range specification parameter 50 specifies a range start address (also called a base address) that identifies the start of the specified address range. The second address range specification parameter 52 specifies the size of the specified range, and thus implicitly identifies the end address of the range relative to the start address. The end address can be determined by adding a value corresponding to the specified value to the start address. An alternative way of specifying the first and second address range specification parameters could be to define the start address using a size or offset parameter defined for the end address. It is also possible that rather than specifying one of the start / end addresses relative to the other, both the start and end addresses could be specified as absolute addresses or relative to a separate reference address other than the range boundary address. The size of the address range may be encoded using parameters that support sizes that are any arbitrarily chosen non-power-of-two number of bytes, and thus the size is not limited to power-of-two block sizes.

[0071] For some encodings of range prefetch instructions, the instruction may specify a single range of addresses. For example, this may be the case when the stride parameter 56 is set to a default value such as 0, or the count parameter 58 is set to a default value such as 0 or 1, or another part of the instruction encoding has a value indicating that only one range is to be prefetched.

[0072] In other encodings, the range prefetch instruction can represent multiple ranges separated by an interval of the specified stride parameter 56, and the count parameter 58 can indicate how many ranges should be represented. These alternative options are described in more detail below with respect to Figures 6 and 7.

[0073] In the example of Figure 2, both the first and second address range parameters are encoded as data stored in a register specified by the range prefetch instruction, which includes register specifiers Xn, Xm to identify the registers that store the first and second address range parameters. However, other approaches can use immediate values ​​to specify the address range parameters, for example, the range size can be defined as an immediate value. The stride value 56 and count value 58 can also be specified in further registers or immediate values, but in this example they are specified in the same register as the range size 52, as described in Figure 6 discussed below.

[0074] Optionally, the range prefetch instruction may also specify one or more additional prefetch hints that may be useful in controlling the prefetching. For example, the range prefetch instruction shown in FIG. 2 includes a prefetch operation field 54, which may be specified in a register referenced by the instruction or as an immediate value, and may have an encoding that identifies one or more of several types of prefetch hint information, such as: Load / store / instruction type information identifying whether addresses within the range defined by parameters 50, 52 are expected to be used by a future load operation, store operation, or instruction fetch - this may be useful for controlling whether to prefetch into a data or instruction cache, or for controlling the coherency state set for prefetched data; Reusability parameters, such as an indication of whether addresses within the specified range(s) are likely to be used for a streaming access pattern of workload in which data from each address is likely to be used only once (or a relatively small number of times), or a temporal pattern of access in which data from a given address within the specified range is likely to be used more than once (or a relatively large number of times); Data usage timing parameters, such as an indication of a target level in the cache hierarchy where the data is to be prefetched, or a relative indication of the time that the data associated with the specified range is estimated to be required by a processing circuit following a range prefetch instruction, which may be useful for controlling when a prefetch operation is issued and / or into which level of the cache hierarchy the prefetched data is prefetched; Data reuse timing parameters, such as an indication of a selected level of cache (e.g., L2 or L3) that should be evicted when prefetched data is replaced in a given level of cache (e.g., L1), or a relative indication of when a further demand access to the address of the prefetched data is expected to be issued immediately following a first demand access to the same address; Access pattern hints that indicate the relative density / sparseness of the pattern of accesses to addresses within a specified range for a particular sparse access pattern type, such as whether accesses are likely to be strided, spatial, temporal, or multi-strided; No-prefetch hint information indicating that subsequent program code is unlikely to use data from addresses within the specified ranges defined by the first and second address range specification parameters 50, 52.

[0075] Any given implementation of a range prefetch instruction may choose to implement any one or more of these types of prefetches and need not encode all of them. Prefetch hint information may be useful in allowing prefetch circuitry 30 to modify how it controls the issuance of prefetch requests for addresses within a specified range.

[0076] Therefore, a range prefetch instruction is proposed that signals one or more ranges of memory addresses to be accessed in the future. When the range prefetch instruction is executed on a microarchitecture having a prefetch circuit 30, the range prefetch instruction can program a hardware backend of the prefetch circuit 30, which is responsible for issuing prefetches in a timely manner by using microarchitectural cues, such as current activity patterns, and tracking progress within the installed ranges. The range prefetch instruction allows the hardware backend to bypass the training stage and start issuing accurate prefetches directly.

[0077] The range prefetch instruction can handle dense or sparse accesses to the encoded address range. The range prefetch instruction also encodes stride and count parameters to handle more complex access patterns. The range prefetch instruction can also rely on hardware to detect complex access patterns within the encoded address range(s), such as multi-stride patterns, spatial patterns, temporal patterns, etc. The hardware backend can also provide hints to the cache replacement policy to make smarter replacement decisions (e.g., the prefetch circuitry 30 can control the setting of cache replacement information associated with cache entries that is used to determine which cache entries to replace when new data needs to be allocated to a given cache 14, 16, 18, e.g., setting cache entry metadata used in Least Recently Used (LRU), Re-Reference Interval Prediction (RRIP), or any other known cache replacement policy).

[0078] From an architectural standpoint, the range prefetch instruction is not required to perform any operation, and it is acceptable for a particular microarchitecture processor implementation to ignore the range prefetch instruction and treat it as a no-operation (NOP) instruction. For example, a processor implementation that does not have prefetch circuitry 30 or that does not support the use of hints 40 available from the range prefetch instruction may treat the range prefetch instruction as a NOP.

[0079] However, there are several advantages in a micro-architectural implementation that supports a prefetch circuit 30 that can use hints from range prefetch instructions.

[0080] Addressing the shortcomings of traditional software prefetching: One range prefetch instruction can be provided for the entire set of one or more ranges. This instruction can be placed once outside the inner loop. The hardware backend is responsible for issuing the prefetches in a timely manner. Thus, the application becomes more micro-architecture and system architecture independent. The associated backend also allows for detection of whether range prefetching is actually detrimental to performance. For example, the prefetch circuitry 30 has the flexibility to selectively turn off prefetching for certain ranges if prefetching is found to hurt performance.

[0081] Addressing the shortcomings of traditional hardware prefetching: Because the range prefetch instruction provides a precise hint about the starting address and range of a stream of addresses, the range prefetch instruction and its simple hardware backend can achieve much higher accuracy while using only a fraction of the hardware overhead of a traditional hardware prefetcher.

[0082] Advantages of programmability: As can be seen from Figure 5 described below, adding a range prefetch instruction to an existing application is very simple and elegant. Thus, in addition to the performance benefits, ease of programming is the primary benefit of this instruction. As an example, in one use case with a depthwise convolution kernel, engineers spent half a day trying to add a traditional software prefetch instruction to the depthwise convolution kernel, resulting in a slight speedup. In contrast, adding a range prefetch instruction took about 10 minutes, and the performance improvement was much greater than when a traditional software prefetch instruction was used.

[0083] Example of range-based prefetch control Figure 3 illustrates an example of how prefetch circuitry 30 may control prefetching based on parameters of a range prefetch instruction. Figure 3 illustrates control over a single range, e.g., the first of multiple ranges specified by a range prefetch instruction when a stride parameter is used to encode multiple ranges.

[0084] In response to the instruction decoder 4 decoding the range prefetch instruction, address parameters 50, 52 (and additional prefetch hints 54, if provided) may be provided to the prefetch circuitry 30. For example, the prefetch circuitry may record information identifying a prefetch state start address and an end address in the prefetch state storage device 32. In response to decoding the prefetch instruction, the prefetch circuitry may issue a prefetch request 34 corresponding to an address within an initial portion of a range, the initial portion being of a particular size J.

[0085] If the overall size of the range is very small (e.g., smaller than the size J defined for the first portion), e.g., only two or three cache lines, then the entire range may be prefetched in direct response to the range prefetch instruction and there may be no need to perform subsequent additional prefetches in response to demand access requests for that range.

[0086] However, for larger ranges specified by a range prefetch instruction, it may not be efficient to prefetch the entire range in response to decoding the range prefetch instruction, because there is a risk that prefetching for later addresses in the range will be useless (because by the time the demand access issued by load / store unit 12 actually reaches the later address, the prefetched data may already have been evicted from the cache in response to accesses to other addresses that may trigger replacements in the cache). Thus, the first portion prefetched in response to a range prefetch instruction may not cover the entire range between the range start address and the range end address.

[0087] For addresses within the remaining portion of the specified range, as shown in FIG. 3, data for prefetch address #A may be prefetched when the prefetch circuit detects that the demand access has reached or passed an address (#AD) from the address 36 of the demand access, where D is an advance distance parameter that controls the margin by which the prefetch stream leads the demand access stream.

[0088] The size J of the first portion and the distance D are both examples of advance distance parameters that affect the timeliness of the prefetch. These parameters can be defined by the prefetch hardware of the circuit 30, independent of any particular software-defined parameters specified by the range prefetch instruction. J and D can be fixed or dynamically adjustable based on microarchitecture information 38, which can include information on, for example, cache hit / miss rates, demand access frequency, prefetch latency, prefetch success rates, or prefetch usefulness. Because different microarchitecture platforms experience different latencies when accessing memory, it can be useful for the advance distances J, D to be controlled in hardware rather than software, and therefore it can be useful to determine how far ahead of the demand access stream the prefetch stream should be issued in response to a queue from the microarchitecture.

[0089] It will be appreciated that FIG. 3 is just one way of controlling prefetching based on range information defined in a range prefetch instruction and that other techniques may be used.

[0090] If a range prefetch instruction specifies multiple ranges, then for subsequent ranges in the stride pattern after the first range, the first portion of size J may be prefetched at various times, such as directly in response to a range prefetch instruction similar to the first range, or when the demand access stream arrives within a given distance of the end of the previous range. Addresses in the remaining portions of the subsequent ranges may be prefetched when the demand stream arrives within distance D of those addresses, similar to what was shown for the first range in FIG.

[0091] Code examples demonstrating ease of programming 4 and 5 show example code that illustrates the programmability advantages of using range prefetch instructions instead of software prefetch instructions that each specify a single address. In this example, for ease of explanation, it is assumed that there are no strided patterns, and therefore only a single range is encoded by each instance of the instruction. However, it will be appreciated that similar advantages apply to strided ranges.

[0092] 4 shows an example of a program loop that has been modified by a programmer or compiler to include a conventional software prefetch instruction PF to prefetch addresses in advance of expected access times. Without the prefetch instruction, the loop would look like this:

number

[0093] Thus, each iteration of the loop performs some arbitrary operation on the data from the first and second buffers (the processing operation is labeled "process" to indicate that no specific operation is intended and it can be any operation). One element of the first buffer is combined with two different elements of the second buffer. Without adding any prefetching operations, the loop is reasonably compact and easy to write.

[0094] If a single addressing software prefetch instruction were used to control the prefetch in this example, this would make the code much more complicated since the prefetch instruction would need to be included a certain distance before the instruction that actually uses the corresponding address specified by the prefetch instruction. Thus, the prefetch instruction "PF" embedded within the processing loop itself must specify an address a certain distance ADV ahead of the corresponding address accessed by the processing instruction itself. This means that it is not possible to include a software prefetch instruction within the main program loop to prefetch data for the first part of a loop iteration having a loop count value of i less than ADV. Thus, as shown in FIG. 4, an additional preliminary loop is added to include a prefetch instruction to prepare the cache with prefetched data to be processed in the first part of the ADV iteration of the main loop. Also, since the main loop now includes a prefetch instruction to prefetch a distance before the main demand access contained in the loop, if this loop is simply executed up to the required number of main data processing loop iterations N, the prefetch instruction may over-fetch beyond the end of the data structure being operated on, wasting cache resources. Thus, in the example of FIG. 4, the main loop may execute for a number of iterations (N-ADV) less than the required number of iterations N, and an additional tail loop is added which performs data processing for a certain number of remaining iterations without prefetch instructions included in the tail loop.

[0095] Thus, the use of single-addressing prefetch instructions adds a lot of additional loop control overhead because instead of processing only a single loop, there are now three loops, which introduces overhead, and when the high-level code shown in FIG. 4 is compiled into a hardware-supported instruction set, additional branch and compare instructions are required to compare the loop counter i to a threshold to determine whether to continue executing further iterations of each loop, and to control the program flow according to the comparison. Since branches may be mispredicted, including these extra loops may hurt performance. An alternative would be to evict the tail loop and simply accept the excess prefetching that would occur if the main loop were to iterate for a full N iterations, but this may waste memory system bandwidth in making unnecessary prefetches and may hurt performance because unnecessarily prefetched data may evict other data, which may then cause additional cache misses.

[0096] Another problem is that the preferred value of advance distance ADV may be different for different microarchitecture platforms, and may even vary on the same microarchitecture platform based on the amount of other patterns of accesses to memory that may occur when the code is executed. Thus, it is very difficult for a software developer or compiler to set the advance distance ADV in a manner that is appropriate for a range of microarchitectures. This can make code of the type shown in Figure 4 very platform dependent, significantly increasing software development costs.

[0097] In contrast, by using range prefetch instructions as shown in FIG. 5, the software developer does not need to consider the relative timing of the prefetcher for individual addresses. Instead, the software developer simply includes several range prefetch instructions before the original processing loop. Each range prefetch instruction specifies a first and a second address range specification parameter that defines the start address and the end address of a range of addresses corresponding to the data structure to be processed. For example, in the code example shown in FIG. 5, one range prefetch instruction specifies the start address and the size of a first buffer, and a second range prefetch instruction specifies the start address and the size of a second buffer. This provides useful information to the prefetcher 30, allowing the prefetcher to prefetch in the first part of two address ranges corresponding to the first and second buffers, and then the timeliness of subsequent prefetches of data within the address ranges can be controlled in hardware without requiring explicit software programming. This avoids the need for a set-up and tail loop as shown in FIG. 4, and also significantly reduces the number of prefetch instructions inserted, saving instruction fetch and decode bandwidth.

[0098] Strided Range Encoding Figure 6 shows a more detailed example of an instruction encoding for a range prefetch instruction. The instruction encoding includes: a first register identifier field Rn that specifies the register providing the range start address 50; a second register identifier field Rm that specifies the register that provides the range size parameter 52 (shown as "Length" in FIG. 6), the stride parameter 56, and the count parameter 58; and A prefetch operation field Rp that provides the additional prefetch hint 54 described above. The prefetch operation field Rp can be interpreted as either a register identifier field that identifies the register that provides the prefetch hint, or as an immediate value that identifies the prefetch hint.

[0099] The remaining bits of the instruction encoding other than fields Rn, Rm, Rp either provide opcode bits that identify this instruction as a range prefetch instruction, or may be used to signal different variants of the range prefetch instruction, or may be reused for other purposes.

[0100] The encoding of the register data stored in the register identified in the Rn field of the instruction is shown at the bottom of FIG. 6. The range size (length), stride and count parameters 52, 56, 58 are all encoded in the same register. In this example, the length and stride parameters are each specified using 22 bits, and the count is specified using 16 bits, meaning that all three values ​​can fit within a single 64-bit register that can be identified using the Rm field. Of course, other sizes of the length, stride and count parameters are possible. Also, in this example, the length and stride values ​​are signed values ​​that can be either positive or negative. It is also possible to define the length and stride as unsigned values ​​that are always positive. The range size and stride values ​​are encoded as any software-specified number of bytes, which is not limited to an exact power-of-two number of bytes.

[0101] 7 illustrates the use of first and second address range specification parameters 50, 52, a stride parameter 56, and a count parameter 58 to control the prefetching of data from multiple address ranges separated by intervals of a specified stride. A range start address 50 (denoted as B to indicate a base address) identifies the address at the start of a first of a group of ranges represented by the range prefetch instruction. A range size parameter 52 (or length L) indicates the size of each of the groups of ranges. Thus, the end address of the first range in the group is at address B+L. A stride parameter S indicates the offset between the start addresses of successive ranges in the group, such that a second range in the group spans from address B+S to B+S+L, a third range in the group spans from address B+2S to B+2S+L, and so on. The count value C indicates the total number of ranges for which prefetching should be controlled based on the range prefetch instruction, so that the final range in the group can span from address B+(C-1)S to B+(C-1)S+L.

[0102] Alternatively, if it is considered implicit to always instruct the range prefetch circuit to subject at least one range to prefetch control (where a count value of 0 may not be useful), then the count C can be defined to indicate the total number of ranges minus 1, and thus in this case the end of the stream is at address B+CS+L, and therefore the total number of ranges is C+1, and the first range does not require explicit encoding, thus accounting for the fact that this can allow a larger maximum number of ranges to be encoded in a particular number of bits.

[0103] However, it may be more intuitive for use by software developers if the count field 58 encodes the total number of ranges C, including the first range in the group (the C=0 encoding may be reused to encode other information, such as additional prefetch hints, for example).

[0104] 8 illustrates a use case where this group of strided address ranges may be useful to control prefetching. Machine learning algorithms (such as convolutional neural networks) may involve convolution operations where a relatively small kernel matrix of weights is convolved with a larger matrix of input activations, and the output matrix produced in the convolution operation depends on various combinations of multiplying the kernel weights with the input activations for different combinations of kernel / activation positions, and logically, the kernel is effectively swept across different positions of the input matrix, and the multiplications performed in the convolution operation represent the products of the kernel weights and activations correspondingly positioned when the kernel is at a given position relative to the matrix.

[0105] Such convolutions may require many load / store operations to memory to load in various sets of activations for different kernel locations and store back the processing results. The data loaded for a given kernel location may correspond to several relatively small portions of a row or column (e.g., the portions represented by the shaded portions of the matrix at the top of FIG. 8), and therefore they may follow a pattern of access such as that shown in FIG. 7, where there are several ranges of interest separated by intervals of regular strides. It will be appreciated that this is a simplification of the convolution operation for ease of explanation, and that in practice the layout of the activation data in memory may not match the logical arrangement of a matrix as shown in FIG. 8 (e.g., if there are multiple layers of activations, the elements associated with a given column or row of a given layer may not actually be contiguous in memory). Nevertheless, it is common for such convolution operations to have access to several ranges that are discontinuous in memory but of equal size, with a regular offset between the start of one range and the start of the next range.

[0106] Thus, by providing range prefetch instructions that support specification of a stride parameter as shown in Figures 6 and 7, a single instruction that prepares the prefetch circuitry 30 can control the precise prefetching of several ranges of addresses that are discontinuously located in the address space with much less training overhead than if pure hardware prefetching techniques were used, while being much easier to program than the option of using software prefetch instructions that each specify a single address.

[0107] Exemplary Microarchitectural Options for Prefetch Circuit 30 9-11 show various examples of possible micro-architectural implementations of the prefetch circuit 30. FIG.

[0108] 9 shows a first example in which the prefetch circuitry comprises a hardware prefetcher 80 that can prefetch data at a predicted address according to the address 36 of the demand access using prefetch training according to any known scheme. For example, this can be based on stride detection (with a single stride or multiple different strides) in the demand address stream.

[0109] Compared to a typical hardware prefetcher, the hardware prefetcher 80 is modified to supplement its training using additional prefetch hints from range prefetch instructions without requiring a dedicated prefetch state machine separate from the existing state machine implemented for training the prediction state 32 based on monitoring the address 36 of the demand access. A typical hardware prefetcher may already have control logic that can monitor the address 36 of the demand access, detect a pattern such as a stride sequence, and gradually adapt the confidence of the observed sequence, and when the confidence reaches a certain level, the hardware prefetcher 80 may start issuing prefetch addresses to a prefetch address queue 82 according to the pattern detected in the demand address stream 36, and issue a prefetch request 34 to the memory system to prefetch data from the indicated prefetch address in the queue 82.

[0110] When range prefetch instructions are supported, the range bounds and stride from the range prefetch instruction prepare the prediction state 32 to settle more quickly into a reliably predicted stream of prefetch addresses. For example, the starting address from the range prefetch instruction (or the starting address 50 plus a multiple of the stride parameter 56) can be used to initialize a new entry in the prefetch prediction state 32, immediately recognizing that a predictable stream of addresses follows that starting address, whereas with standard hardware prefetching techniques it may take some time for the hardware prefetcher to detect from the demand address stream 36 that a demand access has reached that region of the address space and settle into a reliable prediction.

[0111] Another approach is that when a range prefetch instruction is encountered that identifies that a prefetch should be performed for a given range of addresses, the confidence of the corresponding entry in the prefetch prediction state 32 can be increased, thus indicating a higher confidence than would normally be set if the same demand address stream 36 were encountered without executing the range prefetch instruction.

[0112] Similarly, the end address of the range identified by the range prefetch instruction can be programmed into an entry in the prefetch prediction state 32 so that the hardware prefetcher 80 can cease generating prefetch requests when the prefetch stream exceeds the end address or when the demand address stream 36 is observed to have moved beyond the end address of the specified range.

[0113] 10 illustrates a second example in which, in addition to a hardware prefetcher 80 that may operate according to conventional hardware prefetching techniques based on monitoring the demand address stream 36, the prefetch circuitry 30 also includes a range prefetcher 90 that executes a separate state machine to respond to prefetch hints 40 received from range prefetch instructions. In this approach, the range prefetcher 90 can operate in parallel with the hardware prefetcher 80, and both the hardware prefetcher 80 and the range prefetcher 90 can provide prefetch addresses to a prefetch address queue 82 to trigger corresponding prefetch requests 34 to be issued to the memory system.

[0114] The range prefetcher 90 can maintain a set of prediction states 92 that are separate from the prefetch prediction states 32 maintained by the hardware prefetcher, although the range prefetch prediction states 92 can be simpler. For example, the range prefetch prediction state 92 can simply record the parameters 50, 52, 54, 56, 58 of the range prefetch instruction, but may not need to include other fields included in the prefetch prediction state 32, such as a confidence field or a field that records multiple possible candidates for a detected access pattern. Other examples can encode the prediction state 92 in other ways, for example, assigning separate entries of the prediction state to multiple ranges encoded using the stride parameter of the range prefetch instruction.

[0115] In the approach shown in Figure 10, when a range prefetch instruction is decoded by the instruction decoder 4, the range prefetcher can control prefetching in the manner shown in Figure 3, where for a first specified range of a set of one or more ranges encoded by the range prefetch instruction, an initial portion of size J is first prefetched in response to the installation of a new entry in range prefetch state table 92, and then subsequent prefetches are controlled with timing that means the prefetch remains a certain distance D ahead of the observed pattern of demand addresses 36. Microarchitectural information 38 can be used to set the size of the initial portion J and the advance distance D, as previously described.

[0116] The range prefetcher 90 in this example may also provide a signal 94 to the hardware prefetcher 80 to inhibit the hardware prefetcher 80 from performing disciplined prefetching in the ranges in which the range prefetcher 90 is performing prefetching. If the no-prefetch hints described in FIG. 2 are supported, this signal may also be used to inhibit prefetching within ranges indicated by the reach range prefetch instruction as not needing to be prefetched at all.

[0117] Thus, by utilizing hints available from range prefetch instructions and suppressing training by the hardware prefetcher 80 within those ranges, providing a range prefetcher 90 that can perform more accurate predictions with a smaller amount of prediction state information than the hardware prefetcher 80, this conserves resources in the hardware prefetcher 80 to detect other patterns of addresses that are not explicitly indicated as useful to prefetch by the software prefetch instructions. This means that for a given size of circuit area for the prefetch circuitry 30, there are more training resources available for other addresses outside the ranges handled by the range prefetcher 90, so overall performance can be improved, or, for a given amount of prefetch coverage, the size of the hardware prefetcher 80 and prediction state storage 32 can be reduced while maintaining a comparable level of performance. This approach can therefore provide a better balance between performance and circuit area and power costs.

[0118] 11 shows another possible micro-architecture implementation, in which a range prefetcher 90 is provided in the prefetch circuitry 30, but there is no dedicated hardware prefetcher 80 to detect access patterns outside the ranges indicated by the range prefetch instruction hints 40. In this case, prefetching may not be performed at all outside these ranges. This approach may be much less costly in terms of circuit area. The range prefetcher 90 in this example may operate similarly to that described in FIG. 10.

[0119] Optionally, as shown in FIG. 11, the range prefetcher 90 can include a sparse access pattern detection circuit 96 to perform further hardware controlled training of predictions of sparse patterns of accesses within the range being processed by the range prefetcher 90. This can be useful to address cases where software indicates an entire range or ranges where address accesses are likely, but not all addresses are necessarily accessed, and therefore prefetching data from each address within the range can lead to excessive prefetching that impacts performance if it causes replacement of other data in the cache. The sparse access pattern detection circuit 96 can function in a similar manner to the hardware prefetcher 80 shown in the previous example in that it monitors the demand address sequence 36 to maintain a prediction state 32 that can represent information about the type of access pattern detected and the associated level of confidence, to enable a prefetch request to be generated for a predicted address that is predicted based on the detected sparse access pattern. However, compared to the hardware prefetcher 80 shown in FIG. 9 or FIG. 10, the sparse access pattern detection circuit 96 and corresponding prediction state 32 may be very low cost in terms of circuit area and power consumption because the training performed by the sparse access pattern detection circuit 96 is applied only to addresses within the range being processed by the range prefetcher 90, rather than to the entire address space, and much less prediction state data needs to be generated and compared.

[0120] If additional prefetch hints 54 are available from the range of the prefetch instruction, this may affect the sparse access pattern detection. For example, the dense / sparse indicator may be used to determine whether to invoke the sparse access pattern detection circuit 96 for a given range tracked in the range tracking data 92. If the corresponding range prefetch instruction indicates that the pattern is a dense access pattern, the power cost of training the sparse access pattern detection circuit 96 may be avoided by suppressing use of the sparse access pattern detection circuit for that range, which also saves limited prediction state storage 32 for training on other ranges. On the other hand, when the dense / sparse indicator specifies a sparse access pattern, training may be activated using the sparse access pattern detection 96, and if a sparse access pattern type is specified in the prefetch hints 54, this may help the sparse access pattern detection circuit 96 settle on a reliable prediction more quickly, or determine which particular form of pattern detection should be applied to a given range of addresses.

[0121] The same code can be executed across any of the microarchitectural implementations of the prefetcher shown in Figures 9 to 11, and it will be understood that there may be many other ways in which the range information 50, 52 and additional prefetch hints 54 can be used by the prefetchers to improve their predictions.

[0122] Although not shown in any of FIGS. 9-11, some examples of the prefetch circuit may also include some address translation circuitry similar to the MMU 22 of the processing circuit 6 shown in FIG. 1. Like the MMU 22, the address translation circuitry of the prefetch circuit may include a relatively small TLB 24. By providing some translation logic within the prefetch circuit 30, this allows addresses to be translated for prefetch purposes without consuming bandwidth for translation of demand addresses, which may be more important to performance. Which addresses are translated by the address translation circuitry of the prefetch circuit 30 may depend on whether the prefetch is trained on a virtual address or a physical address. If the training is based on a virtual address, the translation may be of a prefetch address specified for a prefetch request issued to the memory system. If the training is based on a physical address, the translation may be of a specified range of addresses based on an operand of a range prefetch instruction.

[0123] Alternatively, in other examples, any address translation required for prefetching may be handled by the main MMU 22 associated with processing circuitry 6, which also handles translation for demand accesses.

[0124] Exemplary Methods 12 is a flow diagram illustrating a method of performing prefetching. In step 200, the instruction decoder 4 decodes a range prefetch instruction that specifies first and second address range specification parameters 50, 52 and a stride parameter 56. In step 202, in response to decoding the range prefetch instruction, the prefetch circuitry 30 controls the prefetching of data from one or more specified address ranges into at least one cache 14, 16, 18 in response to the first and second address range specification parameters and the stride (and also in response to one or more additional prefetch hints, if specified by the range prefetch instruction). Some examples may also encode a count parameter to control the prefetching.

[0125] 13 is a second flow diagram illustrating in more detail how to control prefetching based on a range prefetch instruction. In step 210, the instruction decoder 4 decodes a range prefetch instruction that specifies first and second address range specification parameters 50, 52, a stride parameter 56, and a count parameter 58 (and optionally a prefetch hint 54). In step 211, the prefetch circuitry 30 initializes a running count of how many ranges specified by the range prefetch instruction have already been processed, and the first of the specified ranges is treated as the current specified range to be processed.

[0126] In response to the range prefetch instruction, in step 212, the prefetch circuit 30 prefetches data from an initial portion of the first specified address range. The size of the initial portion is defined by the hardware of the prefetch circuit 30 and may be independent of the information specified in the range prefetch instruction. Optionally, the size of the initial portion may be adjusted based on monitoring of demand accesses and / or microarchitectural information 38 regarding the prefetch request or the results of the demand accesses, but in some examples the size of the initial portion may also be determined using hints regarding the size of the initial portion provided by software or information specifying a link with a previous range to aid in prefetching the initial portion of the range (nevertheless, even if hints provided by software are available, the hardware still has the flexibility to deviate from the size of the initial portion indicated by software if microarchitectural cues indicate that an initial portion of a different size is estimated to be better in terms of performance). Microarchitectural cues such as cache miss rates, the percentage of prefetch requests that are considered successful because there are subsequent demand accesses that hit the same address, or information about the latency of processing a prefetch request or a demand access request may be useful in determining how far ahead in the demand access stream to issue a prefetch request to the corresponding address, and therefore, if a larger advance distance is desired, it may be useful to increase the size of the initial portion prefetched in step 212.

[0127] Once the initial portion has been prefetched, the prefetch circuitry 30 monitors the address 36 of the demand access. In step 214, it is detected that the demand access has reached a target address that is less than or equal to a given distance away from a given address within the remainder of the current specified range. Again, the given distance is defined by the prefetch circuitry hardware. Once the demand access has reached the target address, in step 215, a prefetch of data up to the given address is performed. Again, the given distance may be adjusted based on monitoring of the demand access and / or micro-architectural information regarding the results of the prefetch or demand access requests.

[0128] In step 216, the prefetch circuit 30 detects whether the demand access is within a given distance from the end of the current specified range, and a running count value maintained by the prefetch circuit 30 to track how many ranges are still to be processed indicates that there is at least one more specified range to be prefetched after the current range. The given distance in step 216 may be the same as the given distance in step 214, or it may be different. If the demand access address is within a given distance from the end of the current range and there is at least one more range to be prefetched, in step 217 the prefetch circuit 30 issues a prefetch request for an address within the first portion of the next specified range after the current specified range in the strided group of ranges. Step 217 is omitted if the demand access is not yet within a given distance of the end of the current specified range, or if the running count indicates that this is the last range to be prefetched.

[0129] In step 218, the prefetch circuit 30 also checks whether the address 36 of the demand access or the address of the issued prefetch request has reached or exceeded the end of the current specified range. If neither the demand address stream nor the prefetch address stream has reached the end of the specified range, the method continues to loop through steps 214, 216, and 218 until either the demand access or the prefetch access has exceeded the end of the specified range.

[0130] When the end of the specified range is reached, prefetching is stopped for the specified range in step 220 .

[0131] In step 222, the prefetch circuit 30 checks whether the running count value (maintained internally within the prefetch circuit 30 to track how many of the non-contiguous ranges in the strided set have already been processed) indicates that there is at least one more group of non-contiguous ranges that have not yet been processed. If there is at least one more range to prefetch, then in step 224 the next specified range becomes the current specified range. Thus, the start / end addresses of the next range are addresses offset by the stride parameter relative to the start / end addresses of the range that previously served as the current range. The running count value used to track how many ranges remain for processing is updated to indicate how many ranges remain to be prefetched, and the method then returns to step 214 to continue monitoring demand accesses to determine when to issue a prefetch request for the next range.

[0132] Thus, if stride and count parameters are used to define multiple ranges, each range is processed in a corresponding manner by a pass through steps 214 to 224, and finally, if it is determined in step 222 that there are no more ranges to prefetch, prefetching is discontinued in step 226 for the range represented by the range prefetch instruction.

[0133] 13 shows certain steps being performed sequentially in a particular order, it will be appreciated that other implementations may reorder the steps and perform some steps in parallel. For example, the checks in steps 214, 216, and 218 may be performed in parallel or in a different order.

[0134] It will also be appreciated that the approach illustrated in FIG. 13 is merely an example and that other implementations may use the prefetch instruction parameters provided by the range prefetch instruction in different ways.

[0135] In this application, the term "configured to..." is used to mean that an element of an apparatus has a configuration capable of performing a defined operation. In this context, "configuration" refers to a method of arrangement or interconnection of hardware or software. For example, an apparatus may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not imply that an apparatus element needs to be modified in any way to provide the defined operation.

[0136] Although exemplary embodiments of the present invention are described in detail herein with reference to the accompanying drawings, it will be understood that the invention is not limited to these precise embodiments, and that various changes and modifications can be made to the embodiments by those skilled in the art without departing from the scope of the present invention as defined by the appended claims.

Claims

1. An apparatus comprising: an instruction decoder for decoding an instruction; a processing circuit for performing data processing in response to decoding of the instruction by the instruction decoder; at least one cache for caching data for access by said processing circuitry; a prefetch circuit for prefetching data into the at least one cache; and in response to the instruction decoder decoding a range prefetch instruction specifying first and second address range specification parameters and a stride parameter, the prefetch circuitry is configured to control prefetching of data from a plurality of specified ranges of addresses into the at least one cache in accordance with the first and second address range specification parameters and the stride parameter, a starting address and a size of each specified range being dependent on the first and second address range specification parameters, and the stride parameter specifying an offset between starting addresses of successive specified ranges of the plurality of specified ranges.

2. 2. The apparatus of claim 1, wherein the first address range specification parameter includes a base address of a selected one of the plurality of specified ranges, and the second address range specification parameter includes a range size parameter that specifies a size of each of the plurality of specified ranges.

3. The apparatus of claim 1 or 2, wherein the range prefetch instruction also specifies a count parameter indicating a number of ranges within the plurality of specified ranges.

4. 3. The apparatus of claim 1, wherein the stride parameter is encoded in the same register as one of the first and second address range specification parameters.

5. In response to the instruction decoder decoding the range prefetch instruction, the prefetch circuitry controls prefetching of data from the plurality of specified ranges of addresses into the at least one cache. the first and second address range designation parameters and the stride parameter; at least, monitoring addresses or loaded data associated with demand access requests issued by the processing circuitry to request access to data; The apparatus of claim 1 or 2, configured to control based on both a prefetch request or micro-architectural information associated with the outcome of the demand access request.

6. 6. The apparatus of claim 5, wherein the micro-architectural information includes at least one of prefetch linefill latency, demand access frequency, demand access hit / miss information, demand access prefetch hit information, and prefetch usefulness information.

7. 3. The apparatus of claim 1, wherein the prefetch circuitry is configured to control when a prefetch request is issued for a given address within the multiple specified ranges based on monitoring an address or loaded data associated with a demand access request issued by the processing circuit to request access to data, or based on micro-architectural information associated with the prefetch request or a result of the demand access request.

8. 3. The apparatus of claim 1, wherein the prefetch circuitry comprises a sparse access pattern detection circuit that detects a sparse pattern of accessed addresses within a given designated range of the plurality of designated ranges based on monitoring the addresses of demand access requests, and selects which particular addresses within the given designated range should be prefetched based on the detected sparse pattern.

9. 3. The apparatus of claim 1, wherein the prefetch circuitry is configured to control the timing of issuing prefetch requests based on an advance distance defined by the prefetch circuitry, the advance distance indicating how far a stream of prefetch addresses of the prefetch requests has advanced compared to a stream of demand addresses of demand access requests issued by the processing circuitry to request access to data.

10. 10. The apparatus of claim 9, wherein the prefetch circuitry is configured to define the advance distance independent of an explicit specification of the advance distance by the range prefetch instruction.

11. 10. The apparatus of claim 9, wherein the prefetch circuitry is configured to adjust the advance distance based on micro-architectural information associated with the prefetch request or an outcome of the demand access request.

12. 3. The apparatus of claim 1, wherein in response to the instruction decoder decoding the range prefetch instruction, the prefetch circuitry is configured to trigger a prefetch of data from an initial portion of a first specified range of the plurality of specified ranges into the at least one cache.

13. 13. The apparatus of claim 12, wherein the prefetch circuitry is configured to define the size of the initial portion independent of an explicit specification of a size of the initial portion by the range prefetch instruction.

14. 13. The apparatus of claim 12, wherein the prefetch circuitry is configured to trigger a prefetch of data from the given address within a remainder of the first specified range in response to detecting that the processing circuitry issues a demand access request that specifies a target address that is no more than a given distance ahead of the given address.

15. 15. The apparatus of claim 14, wherein the prefetch circuitry is configured to define the given distance independent of an explicit specification of the given distance by the range prefetch instruction.

16. 3. The apparatus of claim 1, wherein following decoding of the range prefetch instruction, the prefetch circuitry is configured to stop prefetching for a given specified range when the prefetch circuitry identifies, via a stream of prefetch addresses prefetched by the prefetch circuitry or a stream of demand addresses specified by a demand access request issued by the processing circuitry, that the prefetch circuitry has reached an end of a given specified range of the plurality of specified ranges.

17. 3. The apparatus of claim 1, wherein the range prefetch instruction also specifies one or more prefetch hint parameters for providing hints for data associated with the plurality of specified ranges, and the prefetch circuitry is configured to control prefetching of data from the plurality of specified ranges into the at least one cache in response to the one or more prefetch hint parameters.

18. The one or more prefetch hint parameters include: a type parameter indicating whether the prefetched data is data expected to be the target of a load operation, data expected to be the target of a store operation, or an instruction expected to be the target of an instruction fetch operation; a data usage timing parameter indicating at least one of: an indication of how soon the prefetched data is estimated to be required by the processing circuitry after the range prefetch instruction; and a target level in a cache hierarchy to which the data is prefetched. a data reclamation timing parameter indicative of at least one of: an indication of an estimated interval between a first demand access to the prefetched data and a further demand access to the prefetched data; or a selected level of the cache hierarchy from which the prefetched data is preferably evicted if it is to be replaced from a given level of the cache hierarchy; an access pattern hint indicating a relative density / sparseness of a pattern of accesses to addresses within the plurality of specified ranges, or a sparse access pattern type; a reusability parameter indicating the likelihood that the prefetched data is expected to be used once or multiple times; and a no-prefetch hint parameter indicating that prefetching of data from the plurality of specified ranges should be inhibited.

19. the prefetch circuitry is configured to set cache replacement hint information or cache replacement policy information associated with data prefetched from the plurality of specified ranges specified by the range prefetch instruction in response to the one or more prefetch hint parameters; 20. The apparatus of claim 17, wherein the cache is configured to control replacement of cache entries based on the cache replacement hint information or cache replacement policy information set by the prefetch circuitry.

20. 1. A method comprising: Decoding the instruction; and performing data processing in response to decoding of the instruction by an instruction decoder; caching data in at least one cache for access by the processing circuitry; prefetching data into the at least one cache; 1. A method according to claim 1, wherein in response to decoding a range prefetch instruction specifying first and second address range specification parameters and a stride parameter, prefetching of data from a plurality of specified ranges of addresses into the at least one cache is controlled in accordance with the first and second address range specification parameters and the stride parameter, a starting address and a size of each specified range being dependent on the first and second address range specification parameters, and the stride parameter specifying an offset between starting addresses of successive specified ranges of the plurality of specified ranges.