Method and system for retrieving data for an accelerator
By detecting data page access attempts on the accelerator and dividing the array into a multi-level subarray, selecting data pages for prefetching based on page access conditions, solving the problem of inefficient data transmission between the accelerator and the main storage unit, and improving the performance of heterogeneous computer systems, especially when processing multi-dimensional array data.
Patent Information
- Application Number
- CN202080072816.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-12
- Filing Date
- 2020-10-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-10-30
AI Technical Summary
In heterogeneous computer systems, the data transmission efficiency between the accelerator and the main storage unit is inefficient, resulting in system performance degradation, especially when processing multidimensional array data, existing prefetching methods cannot effectively predict and optimize data access patterns.
By detecting unstored data page access attempts on the accelerator, dividing the array into a multi-level subarray, and selecting data pages for prefetching based on the page access conditions, the accelerator's memory management unit is used to achieve efficient data transmission.
It improves the data access efficiency of the accelerator, reduces page failures, and optimizes system performance, especially when processing multi-dimensional array data, improving the overall operating speed and stability of the system.
Smart Images

Figure CN114616553B_ABST
Abstract
Description
[0001] This disclosure claims priority to a U.S. application having application number 16 / 900,215, filed on June 12, 2020, which claims priority to a U.S. Provisional Application having application number 62 / 940,178, filed on November 25, 2019, the disclosures of both of which are incorporated herein by reference. Technical Field
[0002] This disclosure generally relates to accelerators, and more particularly to methods, systems, and non-transitory computer-readable media for retrieving data for us via an accelerator. Background Art
[0003] Heterogeneous computer systems employing accelerators have become an important part of many modern computer systems. Many such computer systems employ a unified virtual memory architecture that allows a central processing unit and an accelerator to share a virtual memory space. While the unified management of the virtual memory space is beneficial, it also presents challenges. Memory oversubscription, frequent page migrations between main storage units, and inefficient memory usage can potentially degrade system performance. An important part of addressing these challenges is prefetching, but if prefetching is performed inefficiently, it can in turn reduce system performance. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, a system, and a non-transitory computer-readable medium for retrieving data for an accelerator. The method includes: detecting an access attempt to a first data page not stored on a main storage unit of the accelerator, the first data page corresponding to a part of a multi-dimensional array; in response to detecting the access attempt to the first data page: partitioning the array into sub-arrays by: partitioning the array into a plurality of first-level sub-arrays, and partitioning the first-level sub-arrays into a plurality of second-level sub-arrays, wherein a first first-level sub-array contains the first data page; selecting pages for prefetching, wherein the selecting of pages for prefetching includes: if a first second-level sub-array meets a page access volume condition, selecting all pages in the first second-level sub-array for prefetching, wherein the first second-level sub-array contains the first data page; and transferring the first data page and any data pages selected for prefetching from a storage system connected to the accelerator to the main storage unit.
[0005] Part of the additional objects and advantages of the disclosed embodiments are set forth in the following description, and part of them will be apparent from the description, or learned through the practice of the embodiments. The objects and advantages of the disclosed embodiments are realized and attained by the elements and combinations particularly pointed out in the claims.
[0006] It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory and are not intended to limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The following detailed description and the drawings illustrate embodiments of the present disclosure in various aspects. The various features shown in the drawings are not drawn to scale.
[0008] Figure 1 is a simplified diagram of a general Von Neumann architecture.
[0009] Figure 2 is a simplified diagram of a basic Von Neumann architecture.
[0010] Figure 3 is a simplified diagram of the internal architecture of a central processing unit. Von Neumann architecture
[0011] Figure 4 is a simplified diagram of a variant of the Von Neumann architecture.
[0012] Figure 5 is a simplified diagram of an operating system structure.
[0013] Figure 6 is a schematic diagram of how a virtual paging system applying page tables maps virtual addresses to physical addresses.
[0014] Figure 7 is a simplified diagram of a modified Von Neumann architecture adopted by many modern computer systems.
[0015] Figure 8 is also a simplified diagram of an architecture adopted by many modern computer systems.
[0016] Figure 9 is a simplified diagram of an accelerator.
[0017] Figure 10 is a schematic diagram of a recursive, bottom-up process for determining which pages of a multi-dimensional data structure to prefetch according to some embodiments of the present disclosure.
[0018] Figure 11 is a schematic diagram of an exemplary computer system for applying the method for fetching data for an accelerator provided by some embodiments of the present disclosure.
[0019] Figure 12 is a simplified diagram of a memory management unit (MMU) of an accelerator provided according to some embodiments of the present disclosure.
[0020] Figure 13 is a flowchart of an exemplary method for retrieving data for an accelerator provided according to some embodiments of the present disclosure.
[0021] Figure 14 is a more detailed schematic diagram of an exemplary accelerator architecture provided according to some embodiments of the present disclosure.
[0022] Figure 15 is a schematic diagram of an exemplary accelerator core architecture provided according to some embodiments of the present disclosure.
[0023] Figure 16 is a schematic diagram of an alternative architecture of an exemplary accelerator provided according to some embodiments of the present disclosure.
[0024] Figure 17 is a schematic diagram of an exemplary cloud system installed with an accelerator provided according to some embodiments of the present disclosure.
[0025] Figure 18 is a simplified diagram for illustrating how a host system determines and executes the selection and prefetching of data pages in an nth-level subarray provided according to some embodiments of the present disclosure. Specific embodiments
[0026] Exemplary embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. The following description refers to the accompanying drawings, where the same numbers in different drawings represent the same or similar elements unless otherwise specified. The implementations described in the following description of the exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with aspects related to the present disclosure described in the appended claims. Specific aspects of the present disclosure are described in more detail below. The terms and definitions provided herein shall prevail if they conflict with the terms and / or definitions incorporated by reference.
[0027] For most modern computer systems, the design, structure, and performance of the computer system's memory are particularly important. Due to various factors stemming from the physical properties of the materials used in most modern computing hardware (e.g., the use of silicon), the performance of the computer system's memory is often a bottleneck in system performance. For this reason, in most modern systems, a large amount of system resources (e.g., a significant proportion of the circuitry) are devoted to complex memory subsystems to reduce and mitigate the impact of this memory bottleneck.
[0028] To a large extent, the importance of memory to computer system performance can be attributed to the basic underlying architecture used by most modern computer systems. Of course, the use of this architecture itself is driven by various underlying physical constraints, which makes it a better choice (currently) than alternative architectures, although it introduces potential performance bottlenecks.
[0029] To better understand why memory is often a performance bottleneck in many computer systems, it is necessary to provide an overview of their basic design. Starting from the most fundamental, in an abstract sense, a computer can be considered a device capable of automatically performing a series of operations. These operations are primitive, which means that the computer system can execute these operations without additional explanation. Informally speaking, the computer system understands these operations. Due to this requirement, primitive operations are usually (but not necessarily) relatively simple, resulting in most such operations being some kind of arithmetic - or logic - based process (e.g., "add these two numbers"). Of course, these operations can be combined to obtain more complex results, which is essentially a program, i.e., a series of operations dedicated to a specific task. Additionally, an operation almost always has a related label, namely an instruction, which can be used to identify (and instruct the computer to execute) a specific operation.
[0030] This basic definition of "computer" is a very broad one because it encompasses most systems capable of automatically performing a series of operations, regardless of their structure or architecture (i.e., regardless of their implementation details). Due to this openness, this definition allows a large number of possible systems and general system architectures to be used to build a "computer". This variation mainly involves which primitive operations are available, how the instructions that determine which operations to execute are processed, and how the inputs and outputs of these operations are handled. For computer science and related fields, almost any of these architectures can be used, although simpler and easier - to - use architectures (e.g., "Turing machines") are usually chosen. However, to implement a computer in physical reality, one must contend with various physically imposed limitations, especially those imposed by the properties of available materials. Ultimately, based on various factors, most modern computer systems find the most suitable architecture to be the von Neumann architecture or variants based on the von Neumann architecture.
[0031] Conceptually, a computer based on the von Neumann architecture can be considered as two interacting subsystems: a data processing subsystem and a data storage subsystem. As its name implies, the data processing subsystem is responsible for performing the various (atomic) operations of the computer system. The main task of the data processing subsystem is to perform the actual computations executed by the computer system. On the other hand, the data storage subsystem is responsible for storing the programs (which are composed of instructions) that direct the data processing subsystem. Additionally, the data storage subsystem is also responsible for storing the various data that are inputs to (or outputs of) the various operations performed by the data processing subsystem. These data can be included at the initialization / startup of the program (e.g., before or at the time when the data processing subsystem starts executing the operation indicated by the first instruction of the program), or these data can originate from the stored output of some operations previously performed by the data processing subsystem.
[0032] Modern computer systems generally employ a generalized variant of the von Neumann architecture. In these generalized variants, the data processing subsystem can contain a variety of data processing elements. Similarly, in the generalized von Neumann architecture, the data storage subsystem can contain a variety of data storage elements. Typically, these data processing subsystems include a core element called the central processing unit (CPU). Generally speaking, the CPU is capable of properly handling various tasks (e.g., the CPU has a variety of atomic operations available). The CPU is also typically responsible for managing and coordinating the activities of other data processing elements. These other data processing elements are usually heterogeneous. That is, these data processing elements have different performance characteristics and different performance for a particular workload. These additional data processing elements are called accelerators, and they are typically dedicated to improving the performance of specific operations and workloads and usually at the expense of reducing the performance of other workloads.
[0033] Similarly, the data storage subsystem typically contains a core element called the primary storage unit (PSU). Typically, the PSU can be directly accessed by the data processing subsystem (e.g., the CPU) and is responsible for storing the instructions and data used by the data processing subsystem. Generally speaking, the PSU is optimized for speed, which means that the PSU can quickly handle data transfers with the CPU (and typically the data processing elements). However, as a trade-off, the PSU typically has low storage density, high storage cost, and volatility, which means that the PSU does not retain data when power is lost. Therefore, data storage systems typically also employ multiple heterogeneous data storage elements. That is, these data storage elements have different performance characteristics and different performance for a particular workload. These additional data storage elements are called data storage devices, and they are typically dedicated to increasing data storage and non-volatility and usually at the expense of reducing speed and increasing latency.
[0034] Figure 1 is a simplified diagram of a general Von Neumann architecture. According to Figure 1 , the Von Neumann architecture 100 consists of a data processing subsystem 110 and a data storage subsystem 120. Also shown in the figure is I / O 130, which can provide inputs and outputs to the data processing subsystem 110 and the data storage subsystem 120. As Figure 1 shown, the data processing subsystem 110 consists of a variety of data processing elements, and data processing elements 111, 112, 113, and 114 are shown in the figure. Similarly, the data storage subsystem also consists of a variety of data storage elements, and data storage elements 121, 122, 123, and 124 are shown in the figure. As indicated by the arrows, the data processing subsystem 110 and the data storage subsystem 120 can communicate with each other and transfer data.
[0035] In the basic version of the Von Neumann architecture, the data processing subsystem consists only of a CPU, and the data storage subsystem consists only of a PSU. Generally, the PSU can be directly accessed by the data processing subsystem (e.g., the CPU) and is responsible for storing the instructions and data used by the data processing subsystem. During operation, the data processing subsystem and the data storage subsystem are closely intertwined, especially the CPU and the PSU. To operate, the CPU typically needs to retrieve the next instruction (represented as data) from the PSU that indicates the next operation the CPU is going to perform. This instruction (more precisely, the operation indicated by this instruction) can then trigger an interaction with more data on the PSU (e.g., retrieving inputs for arithmetic / logical operations, storing outputs of arithmetic / logical operations, etc.).
[0036] Figure 2 is a simplified diagram of a basic Von Neumann architecture. According to Figure 2 , the basic Von Neumann architecture consists of a data processing subsystem 210 and a data storage subsystem 220. As shown in the figure, the data processing subsystem 210 consists only of one data processing element: the central processing unit 211. Similarly, the data storage subsystem 220 also consists only of one data storage element: the main storage unit 221. Also shown in the figure is I / O 230, which can provide inputs and outputs to the central processing unit 211 and the main storage unit 221. Some internal components of the CPU 211 are also shown in the figure, including a control unit 212, an arithmetic logic unit 213, and processor registers 214.
[0037] Frequent communication requirements make the speed at which the data storage subsystem responds or reacts to interactions with the data processing subsystem a potential bottleneck in the performance of the entire computer system. This bottleneck is known as the "von Neumann bottleneck". Unfortunately, in most modern computer systems, the von Neumann bottleneck is a significant problem. For various reasons, the speed of operation / interaction of the processing subsystem (especially the CPU) with the storage subsystem is significantly faster than the speed of operation / interaction of the storage subsystem (especially the PSU) with the processing subsystem. This trend is sometimes referred to as the "memory wall" or "bandwidth wall". Due to this speed difference, most modern computer systems employ various strategies to reduce or mitigate the impact of the von Neumann bottleneck and the memory wall.
[0038] To understand some of the other strategies in use, it is necessary to understand in more detail how the data processing subsystem interacts with the data storage subsystem, paying particular attention to how the CPU interacts with the PSU. In turn, to understand how the CPU interacts with the PSU, it is necessary to understand some of the internal designs and architectures of most CPUs (which also applies to most other data processing elements) and most PSUs. First, as a general matter, most modern computers are based on binary logic. This means that almost all data in a computer is stored / represented as a sequence of bits (a unit of information that can only be one of two symbols, typically represented as "0" and "1"). Each set of discrete bits represents a binary number (e.g., a number in base 2, such as 0b0001101110001111, where "0b" indicates that the number is in base 2 and the number is 7055 in base 10). For better readability, binary numbers are usually written as hexadecimal numbers (i.e., in base 16) (e.g., the previous binary number 0b0001101110001111 is 0x1B8F, where "0x" indicates that the number is in base 16).
[0039] For a data storage subsystem, most PSUs are composed of multiple basic units called "memory locations". Each memory location points to a set of discrete data, which usually has the same fixed length / size. A "memory location" can be identified and addressed by a unique identifier associated with the memory location called a physical address. The "first" physical address usually starts numbering from 0x00000000 and continues to increment by 1 for each sequential unit (e.g., for 32-bit addresses, 0x00000001, 0x00000002, 0x00000003,..., 0xFC4A95E3, etc., and for 64-bit addresses, 0x0000000000001, 0x00000000002,..., 0x000009BCFC4A95E3, etc.). Since most modern computers are based on binary data, the size of the memory locations in most PSUs is some multiple of bits (note that a memory location can be composed of smaller units called memory cells, which usually store the value of a single bit; however, memory cells usually cannot be directly accessed). For most modern systems, the size of a memory location is 8 bits, also called a byte. This memory design is called "byte addressing". Similarly, since most modern computers use binary numbers to represent data, physical addresses are usually also binary numbers with a fixed number of bits, and for convenience, physical addresses are usually also a fixed number of bytes (e.g., 16-bit (2-byte) physical addresses, 32-bit (4-byte) physical addresses, 64-bit (8-byte) physical addresses). Discrete data blocks larger than one byte are stored in a sequential address sequence. Note that although all data stored on most PSUs is binary, the interpretation of what the number represents (e.g., its data type, such as a string, floating point number, etc.) may vary based on the program / instructions being executed or the operation being performed.
[0040] As for the data processing subsystem, the CPU (and most data processing elements) have a set of operations / actions that they can perform, which are typically basic. Each of these operations / actions is represented and identified by an instruction, and due to the binary nature of the computer, instructions are usually represented by binary numbers. The design of the entire instruction set used by the CPU (and the operations / actions represented by the instruction set) is called the instruction set architecture. At a high level, the instruction set architectures of most modern computer systems divide instructions into two main categories: memory operations and arithmetic / logical operations. This is commonly referred to as the load-store architecture. The reason for this separation is that most CPUs (and typically most data processing elements) cannot directly use the data stored in the PSU in arithmetic operations (e.g., as input), and most CPUs also cannot directly use the output of arithmetic operations to store data on the PSU. Instead, to use the data on the PSU, the data must first be transferred / stored in the processor registers of the CPU. Processor registers are a smaller memory holding area used to hold the data that the CPU is using (or generating). Processor registers are usually of a fixed size, called a word. Similar to the physical addresses of the PSU, a word is a certain number of bits, and for convenience, the size of a word is usually also multiple bytes (e.g., a word size of 16 bits (2 bytes), 32 bits (4 bytes), 64 bits (8 bytes)). Only the data located in these processor registers can be used as input for arithmetic / logical operations, and only the data located in these registers can be stored onto the PSU.
[0041] In a CPU that uses the load-store architecture, to transfer (e.g., copy) data into the processor registers of the CPU, the CPU must be instructed to read the data from the PSU and copy the data into the processor registers. The CPU reads the specific data and uses a memory operation to copy the data into a specific register (using the corresponding instruction to indicate this operation). Similarly, the CPU is instructed to write the specific data in the register into the PSU using a different memory operation (represented by a corresponding different instruction). Therefore, to perform an arithmetic / logical operation using two different numerical values, the CPU usually has to fetch a load instruction, fetch the first input, fetch a second load instruction, fetch the second input, fetch a third instruction, and perform the indicated operation using the two numerical values just fetched as input.
[0042] Figure 3 is a simplified diagram of the internal architecture of the central processing unit as described above. According to Figure 3, the basic architecture of the central processing unit includes a control unit 310, an arithmetic logic unit 320, and processor registers 330. The control unit 310 includes an instruction register 311 and a program counter 312. The control unit 310 also has a connection to the PSU (not shown), and the control unit 310 can retrieve data and instructions from the PSU. As described above, the control unit 310 can retrieve instructions from the PSU and store these instructions in the instruction register 311. The control unit 310 can also use the program counter 312 to store the physical address on the PSU where the next instruction to be executed is located. As Figure 3 shown, the control unit 310 is connected to the arithmetic logic unit 320. Based on the instructions stored in the instruction register 311, the control unit 310 can configure the arithmetic logic unit to perform various operations. The arithmetic logic unit 320 can perform these operations by changing the data stored in the processor registers 330, where the processor registers 331, 332, 333, 334, 335, and 336 are shown here.
[0043] Note that some systems use an instruction set architecture called register-memory architecture. In these instruction sets, arithmetic operations can specify data on the PSU, seemingly allowing the data on the PSU to be used by arithmetic / logic operations (e.g., used as input or set by output) without first being copied to or from the processor registers. However, in most such systems, these instructions are actually implemented by a deeper, more fundamental layer of instructions called microcode, which breaks a given instruction into several smaller, more fundamental microinstructions. Then, the set of microinstructions performs the above process, i.e., before the CPU uses the input data, it first retrieves the input data and places it into the processor registers. This means that the microinstructions retrieve the input data (located on the PSU and not currently in the processor registers) from the PSU and place the retrieved input data into the processor registers before performing any arithmetic / logic operations using the input data. Similarly, this also means that the microinstructions first place the output data generated by performing the arithmetic / logic operations into the processor registers before performing any memory operations using the output data.
[0044] Thus, in a high-level overview of the basic von Neumann architecture, the CPU of the data storage subsystem typically interacts with the PSU of the storage subsystem of the memory in the pattern of (1) fetch an instruction -> (2) execute the instruction -> (3) fetch an instruction -> (4) execute the instruction -> …. In almost every case, the CPU retrieving its next instruction involves fetching the instruction (represented as data) from the PSU. Additionally, if the fetched instruction is a memory operation, the CPU will also interact with the PSU (e.g., fetch data from the PSU or store / modify data on the PSU). Thus, the CPU constantly interacts with the PSU to determine its next operation (i.e., fetch its next instruction) and, in many cases, execute that operation.
[0045] Given the frequency at which the CPU and the PSU interact with each other (and the frequency at which the data processing subsystem and the data storage subsystem interact with each other), being able to reduce or improve the impact of the von Neumann architecture or the memory wall, e.g., the trend of the PSU being slower than the CPU is a core consideration for the performance of the entire computer system. To achieve this goal, modern computer systems adopt various strategies to reduce or mitigate the performance difference between the data processing subsystem and the data storage subsystem. One of the oldest and most basic strategies is to optimize the PSU to help reduce the speed difference. For this very reason, the PSU is typically optimized to have high speed and low latency.
[0046] One of the oldest and most basic strategies is to heavily optimize the PSU, the main component of most data storage subsystems, to have high speed and low latency. While this helps reduce the performance gap between the data processing subsystem (or the CPU of its main components) and the data storage subsystem, high-speed optimizing the PSU brings various trade-offs. For example, the PSU typically has a lower storage density (e.g., less data stored per unit volume) and usually has a higher storage cost (e.g., higher cost per unit of data storage). However, most importantly, the optimization for high speed and low latency results in most PSUs being volatile, meaning that the PSU will not retain data if power is lost.
[0047] In a technical sense, non-volatile memory is not absolutely necessary for a computer system. If there is some way to load the necessary data (e.g., the program to be executed) from the outside and some way to retrieve any desired output (e.g., the result of the program), the normal operation of the computer system only requires a way to store and retain data during execution. This function is exactly what the PSU provides. This is similar to the human “working memory”, i.e., the nervous system responsible for retaining / remembering information for the currently ongoing task. However, while not absolutely necessary, most modern computer systems find the ability to store data long-term to be necessary.
[0048] Thus, since having non-volatile storage is often necessary, most data storage subsystems also employ an additional component known as a secondary storage unit (SSU). Just as the PSU is analogous to a human's "working memory", the SSU is analogous to a human's "long-term memory", which is a nervous system responsible for retaining / remembering memories and information that are not currently in use, typically long-term and potentially permanent. Considering the role of the SSU in the data storage subsystem, unlike the PSU, the SSU is typically optimized for high storage capacity and long-term data retention, including non-volatility. As a consequence, therefore, the SSU is typically slower than the PSU both in terms of bandwidth (the amount of data that can be transferred within a time interval) and latency (the time taken to respond to an I / O request).
[0049] Thus, in most computer systems, the data storage subsystem has a PSU and an SSU, where the PSU serves as a high-speed repository for data related to the processing / tasks being executed by the CPU (or other data processing elements of the data processing subsystem), while the SSU serves as a slow repository for data and information unrelated to any currently executing processing / tasks. However, due to the low performance of the SSU (e.g., slower speed and higher latency) creating a greater speed differential between the SSU and the CPU (or other data processing elements), most computer systems do not directly transfer data from the SSU to the data processing subsystem.
[0050] Although a direct exchange between the CPU and the SSU is not theoretically impossible, most computer systems are structured such that for the CPU (or other elements of the data processing subsystem) to access data stored on the SSU, the data on the SSU must first be transferred to the PSU, and then that data is transferred to the data processing subsystem. Evidently, the reason for this data storage hierarchy is to prevent the data processing subsystem (e.g., the CPU) from spending and idling a large amount of computing time (e.g., multiple clock cycles) while the data is retrieved from the SSU. Instead, when the data processing subsystem is handling other tasks, the data is first read into the PSU. At some point after the transfer to the PSU is complete, the data processing subsystem can then retrieve the data (and process the related task) from the faster PSU, thus avoiding a large waste of computing time.
[0051] Figure 4 is a simplified diagram of a variant of the von Neumann architecture as described above. According to Figure 4, the architecture consists of a data processing subsystem 410 and a data storage subsystem 420. As shown in the figure above, the processing subsystem 410 consists of only one data processing element: the central processing unit 411, while the data storage subsystem 420 consists of two data storage elements: the main storage unit 421 and the secondary storage unit 422. The figure also shows an I / O 430, which can provide input and output to the central processing unit 411 and the main storage unit 421. Figure 4 Some internal components of the central processing unit 411 are also shown, including a control unit 412, an arithmetic logic unit 413, and processor registers 414. The connection between the central processing unit 411 and the main storage unit 421 is also shown in the figure. As shown by these connections, data can be transferred between the central processing unit 411 and the main storage unit 421. The connection between the main storage unit 421 and the secondary storage unit 422 is also shown in a similar manner in the figure.
[0052] Of course, modern computer systems employ a variety of additional strategies to improve the von Neumann bottleneck in addition to optimizing the speed of the PSU. In many cases, these strategies also simultaneously address other issues imposed on the physical constraints of the computer system. This is worth noting because when dealing with these other constraints, trade - offs are sometimes made with further improving the von Neumann bottleneck.
[0053] In real - life computer systems, one of these additional complexities involves the operating system (OS). The operating system can be thought of as the “central” or “master” program that manages the hardware of the computer system and various low - level operations. At a high level, one of the main benefits of the operating system is that it allows other programs / applications to safely ignore (i.e., abstract away) the various underlying details of the computer system and other applications that may be running simultaneously. As part of this abstraction, the operating system typically manages the operation of the data storage subsystem, a task known as “memory management”. One of the most useful and widespread aspects of memory management used by the operating system is the use of “paged virtual memory”
[0054] Figure 5 is a simplified diagram of the structure of the operating system described above. According to Figure 5, a computer system can be composed of various hardware 520, and these hardware 520 can have an operating system 510 executed thereon. The operating system 510 can have various system components that provide various functions, and system components 511, 512, and 513 are shown in the figure. An example of such a system component is memory management. Running on top of the operating system are user applications 500, and applications 501, 502, and 503 are shown in the figure. The user applications 500 can utilize various features and services provided by the operating system 510 through system components 511, 512, and 513 to interface with the hardware 520 in a simplified manner.
[0055] As background, most modern computer systems provide the ability to allow multiple programs to execute simultaneously. This feature is typically implemented by the operating system, in which this feature is not called concurrent execution but is more commonly referred to as multitasking. Multitasking can be achieved by interleaving the execution of various programs, giving each program a certain amount of run time. This process is called "time-sharing management". Multitasking can also utilize a CPU (or other data processing element) with multiple cores by allowing different programs to be executed on different cores. In either case, allowing multiple programs to be in execution simultaneously can lead to potential conflicts in their memory usage. Specifically, as described above, programs use various memory addresses (and the represented memory locations) to store their data. However, when two programs attempt to use the same physical address space, conflicts can occur. In fact, such conflicts are particularly likely to occur because the address spaces used by most programs start from 0x00000000 until reaching the total amount of memory that the application can use. This can cause programs to overwrite or otherwise modify each other's data. In addition to data corruption, this typically results in incorrect results (e.g., if the changed value is used as an input for other operations), and usually causes the program whose data has been changed to fail (e.g., crash).
[0056] Paged virtual memory, as a means of the operating system, transparently (to the program) coordinates the use of memory by any executing program to solve these problems. This essentially abstracts the details of coordinating memory usage (from the program's perspective), that is, allows a program to be written without having to consider which other programs it may be running concurrently with. In particular, paged virtual memory is a combination of two related memory management techniques: paged memory and virtual memory.
[0057] In a computer system that utilizes virtual memory, the operating system allocates a virtual address space for each program. From the perspective of the program, the "virtual address space" is the same as the "physical address space". However, since each program is allocated in its own virtual address space, each program can freely use any virtual memory address without having to consider the memory usage of any other concurrently executing programs. To actually implement the virtual address space of each program, the operating system can allocate a sufficiently large portion of the physical address space and create a data structure that records the mapping between a program's virtual memory addresses and the corresponding physical memory addresses from the allocated physical address space. Then this data structure can be used to map the program's virtual addresses to the corresponding physical addresses. By ensuring that this portion of the physical address space is only mapped to the program's virtual address space, the operating system can ensure that no two programs attempt to use the same physical address, thus avoiding the memory coordination problem of concurrently executing multiple programs. Virtual addressing is typically provided with hardware assistance by the CPU through its Memory Management Unit (MMU). The MMU mainly performs the conversion from virtual memory addresses to physical addresses, which speeds up the processing.
[0058] However, "virtual memory" has several inefficiencies. First, without some kind of memory segmentation scheme (such as paging), the physical address space allocated to a program's virtual address space must be contiguous. This leads to the problem of memory fragmentation, reducing the size of the available memory on the PSU that is effective. Second, in some versions of virtual memory, each executing program must be allocated a physical address space until the program finishes execution. Due to the limited capacity of the PSU, this can quickly exhaust the available memory on the PSU, leaving insufficient memory to execute additional programs. Paging is a segmentation scheme that can be applied to virtual memory to avoid the first problem. Paging also naturally applies to another variant called demand paging, which helps to solve the second problem.
[0059] Specifically, paged virtual memory involves dividing the virtual memory of each program into segments of equal size, called pages (sometimes called "virtual pages"), with a typical page size of 4KB. Similarly, the physical memory of the system is divided into blocks of equal size, called page frames (sometimes also called "physical pages"). To improve efficiency, the size of the page frame is usually the same as the size of the page (however, occasionally the page frame is increased or decreased by a power of 2). Then, instead of mapping the entire virtual address space of a program to some (contiguous) physical address space, each page of the program's virtual memory is mapped to a page frame of physical memory. The operating system maintains a data structure called a page table that records the physical frame to which each page is mapped. While this version of paging avoids the fragmentation problem that would occur if the virtual memory of a program had to be mapped to contiguous segments of physical memory, it does not solve the problem of having to map all of the virtual memory of a program to physical memory, regardless of whether all of that virtual memory is currently in use. However, a variant of paging called demand paging solves this problem.
[0060] Figure 6 is a schematic diagram of how a virtual paging system that applies a page table maps virtual addresses to physical addresses. Specifically, Figure 6 illustrates a two-level page table for mapping virtual addresses to physical addresses. First, in each page table system, a control register 601 contains the starting physical address of the highest-level page table. In a two-level page table scheme, the highest-level page table is the page directory. Each entry in the page directory is the starting physical address of a next-level (and also lowest-level) page table, and each entry in the page table is the address of the physical frame that actually stores the required data. In operation, the virtual address (a combination of page directory index 611, page table index 631, and physical page offset 651) in a two-level page table scheme is divided into three components: the first component (page directory index 611) is the index of the corresponding page table in the page directory 620, the second component (page table index 631) is the address index of the corresponding physical frame in the page table 640. The third component is the offset into the physical frame and is not mapped. Instead, the third component is appended to the physical frame address to determine the complete physical address to which the virtual address maps.
[0061] Demand paging is essentially the same as paged virtual memory, except that demand paging takes advantage of the existence of the SSU. In demand paging, only some pages of a program are assigned (supported / stored in) page frames, and typically these are the pages that are currently in use. For pages not supported by the page frames, these unsupported pages are stored (swapped out) to the SSU. This portion of the SSU used for this purpose is typically referred to as the swap space. This allows the operating system or computer system to allocate more virtual memory than physical memory at one time, because the operating system can freely swap the pages supported by the page frames at any given time. By analogy, one can think of the pages stored on the SSU as being somewhat similar to the concept of a human's "mid-term memory". Typically, the SSU is used to store data for long-term retention, and most of the data is not currently being used at any given time. This is similar to a human's "long-term memory". However, the portion of the SSU used for the swap space is not typically intended for long-term retention. Instead, the data stored in the swap space is only used in the intermediate future (e.g., on the order of minutes), rather than at the literal current moment or in the short-term future (e.g., on the order of seconds). This is similar to a human's short mid-term memory, which stores information that is relevant to the current task but is not used for operations that are to be performed at the exact foreseeable moment.
[0062] Since demand paging involves storing some pages on the PSU and other pages on the SSU, the operating system should have some way to determine which pages should be on the PSU and which pages should be on the SSU. At the start of an application, there is an initial question of which pages should be placed on the PSU during initialization. Thereafter, when the application is running, there are two separate but related questions, namely when to transfer additional pages from the SSU to the PSU (page-in) and when to transfer pages (if any) from the PSU back to the SSU (page-out).
[0063] Since a page generally must reside in the PSU for the CPU to execute its instructions (or access its data), a program must have at least one page in the PSU when it starts execution. Thereafter, at least when a new page needs to be referenced, the new page is brought in. More precisely, referencing / accessing a page (more precisely, a virtual address on the page) that is not currently brought in (e.g., currently swapped to the SSU) is called a page fault. Whenever a page fault occurs, the operating system intervenes to bring in the referenced page (i.e., bring the referenced page from the SSU into the PSU). Variants of demand paging may also choose to load other non-referenced pages (called prefetching), or load pages without a page fault occurring. Almost every operating system also performs the reverse operation, i.e., evict currently brought-in pages (e.g., currently loaded into the PSU) to the swap space in the PSU. The criteria used to decide which page to evict are called page replacement algorithms, and different algorithms use different criteria. Although theoretically page eviction is not required, the problem introduced by demand paging, i.e., ensuring available space is reserved in the PSU, may occur again, e.g., all available physical memory, i.e., each page frame, is eventually used. This may lead to memory exhaustion and cause problems if the program attempts to access a page that has been evicted.
[0064] The basic variant of demand paging is "pure demand paging", which performs paging only when a page fault occurs. While this does solve the problem of limited available physical memory on the PSU, pure demand paging introduces another problem by exacerbating the impact of the memory bottleneck. Specifically, since pages are loaded only when a page fault occurs, whenever a page is accessed for the first time, the operating system typically intervenes to load the referenced page onto the PSU. Since pages are relatively small, this is likely to occur frequently. However, the problem is that because of the large speed difference between the SSU and various components of the data processing subsystem (such as the CPU), when a page fault occurs, the CPU usually has to wait a long time (relative to the time taken for each clock cycle) to fetch the relevant page from the SSU back into the freshly mapped page frame in the PSU. To help reduce the occurrence of page faults, many operating systems use a variant of demand paging called anticipatory paging, which is the opposite of pure demand paging.
[0065] In pre-paging, as the name suggests, the operating system attempts to predict which pages a process will soon access or reference (i.e., the operating system attempts to predict which virtual memory addresses the process will attempt to interact with in the near future), and then pre-loads, e.g., pre-fetches, those pages that are predicted to be accessed soon. By the time the access occurs, the page has been mapped to a page frame and its data has been loaded into the corresponding memory location in the PSU. Unfortunately, however, it is not always possible to know exactly in advance which page (or more specifically, which virtual addresses) a process will attempt to reference, at least not without running a full simulation of the program, which is redundant. Therefore, every implementation of pre-paging uses heuristics to guess which pages (more specifically, which virtual addresses) will be referenced in the near future, and then pre-loads those pages (more specifically, the pages containing the virtual addresses). Page replacement algorithms also use heuristics to guess which pages will not be referenced in the near future, and then evict those pages.
[0066] The heuristics used by many variants of pre-paging are based on the locality of reference. The locality of reference refers to the tendency of a program (especially a specific instruction in the program) to repeatedly reference nearby memory locations in a short period of time. This in turn means that a reference to a page (virtual address) implies that the nearby pages (virtual addresses) may also be accessed. Pre-paging schemes can take advantage of this by, when a page fault occurs, bringing in the page that triggered the page fault as well as the adjacent pages of the page that triggered the page fault (if the adjacent pages have been evicted), since the adjacent pages may also be accessed in the near future. Typically, these functions are provided with hardware support by the MMU of the CPU.
[0067] While the paged virtual memory system described above helps improve some of the performance degradation caused by the von Neumann bottleneck, many performance degradations still exist, if not most. For this reason, most modern computer systems employ an additional strategy to further reduce performance loss: caching. More specifically, almost all modern components of the data processing subsystem, especially the CPU, use internal memory called hardware caches. The function of the hardware cache is very similar to that of a higher-level memory above the PSU but below the processor registers inside. More specifically, the hardware cache is usually faster than the PSU (just as the PSU is faster than the SSU), but it also has a smaller storage capacity than the PSU (just as the storage capacity of the PSU is smaller than that of the SSU). The hardware cache is similar to the PSU in many ways, except that, for convenience, the hardware cache only stores a copy of the data on the PSU. In other words, the cache is used to hold a copy of frequently accessed data so that the data can be accessed more quickly than retrieving it from the PSU. If the data is modified, the data can be kept in the cache (if the data is still in use and / or being modified) before being written back to the PSU. Thus, in effect, the cache holds the current copy of the data.
[0068] The interaction between the hardware cache and the PSU is also similar to the interaction between the PSU and the SSU, because the CPU (or other data processing element) containing the hardware cache decides which data to keep in the hardware cache and also decides which of the retained data to remove (or write back if the data has been modified but not yet written to the PSU). This is largely similar to the memory management function performed by the operating system when deciding which pages to bring into the PSU and which pages to evict to the SSU. The criteria for deciding when to fetch data back into the cache or when to evict data from the cache are called cache algorithms. This is somewhat complicated because most modern CPUs (and most other modern data processing elements) have several levels of hardware caches (called multi-level caches), which form a cache hierarchy. The hardware caches are usually labeled L1, L2, etc., where L1 is located closest to the processor registers (i.e., at the top of the cache hierarchy), L2 is located the next closest to the processor registers, and so on. The multi-caches follow the general trend of the memory hierarchy, i.e., the hardware cache closer to the processor registers (e.g., the L1 cache) is faster but has a smaller storage capacity, while the hardware cache closer to the PSU is slower but has a larger storage capacity. The way data moves between different caches is roughly the same as the way data moves between the outermost cache and the PSU. However, the difference between the hardware cache and the PSU is that instead of the operating system being responsible for bringing data in and out, it is the CPU (or other data storage element) that is responsible for moving data (and determining the data to be moved) between the PSU and the last-level cache or between different levels of caches.
[0069] As mentioned above, most modern computer systems use a modified version of the basic von Neumann architecture. Most modern computers consist not only of a data processing subsystem with a CPU, but also contain other data processing elements. These other data processing elements, called accelerators, are more specialized than the CPU. Different from the CPU, an accelerator uses a special architecture that is more efficient for the specific tasks it is designed for (e.g., greater serial speed, more parallelism, less energy consumption). This is where the term "accelerator" comes from, and an accelerator is dedicated to accelerating one or more workloads it is designed for. However, as the price of accelerating on a set of tasks, an accelerator tends to perform worse on other sets of tasks. A common example of such an accelerator is the GPU, which is faster than the CPU on tasks that require processing a large number of simple operations, such tasks often appearing in machine learning applications.
[0070] Figure 7 is a simplified diagram of the improved von Neumann architecture adopted by many modern computer systems. As Figure 7 shown, this architecture consists of a data processing subsystem 710 and a data storage subsystem 720. According to Figure 7 , the data processing subsystem 710 consists of various data processing elements, including a central processing unit 711 and accelerators 715. As shown in the figure, the data processing subsystem 710 can have multiple central processing units, with CPU 712 and CPU 713 shown in the figure, and multiple accelerators, with accelerator 716 and accelerator 717 shown in the figure. Similarly, the data storage subsystem 720 also consists of various data storage elements, including a main storage unit 721 and a secondary storage unit 725. As shown in the figure, the data storage subsystem 720 can have multiple main storage units, with main storage unit (PSU) 722 and 723 shown in the figure, and multiple secondary storage units, with secondary storage unit (SSU) 726 and 727 shown in the figure. Also shown in the figure is a bus 740 that connects the data processing subsystem 710 and the data storage subsystem 720. The bus 740 can also connect the internal components of the data processing subsystem 710 and the data storage subsystem 720 together. In addition, the bus 740 can connect the data processing subsystem 710 and the data storage subsystem 720 to I / O 730, which can provide input and output to these subsystems.
[0071] Figure 8 is also a simplified diagram of the architecture adopted by many modern computer systems, especially heterogeneous computer systems. According to Figure 8, the host system 801 may include a central processing unit 802, an interconnect unit 808, and a data storage subsystem 805. The central processing unit 802 may include cores 803 and a CPU MMU 804. Similarly, the data storage subsystem 805 may include a main storage unit 806 and a secondary storage unit 807. The CPU 802 may communicate with the data storage subsystem 805 (and ultimately with the main storage unit 806 and the secondary storage unit 807) through the interconnect unit 808. The host system 801 may also be connected to a plurality of accelerator units through the interconnect unit 807, and accelerator units 820 and 830 are shown in the figure. Still referring to Figure 8 the accelerator unit shown, such as accelerator 820, consists of a core (such as core 821), an MMU (such as MMU 822), and a main storage unit (such as main storage unit 813).
[0072] However, despite this difference, the accelerator does have a specific, associated similarity with the CPU; the accelerator still fetches the instructions to be executed and the data to be used from the PSU. In some computer systems, the accelerator may fetch the instructions and data from the PSU into an internal hardware cache, just as the CPU does. However, since complex systems usually have to be used to ensure data consistency, such systems bring complexity. In other words, it is necessary to use additional systems to ensure that the accelerator does not access the data on the PSU that is being modified by another data processing element (e.g., the CPU or another accelerator), since this may lead to contention or the use of outdated and invalid data. Additionally, having the accelerator share the PSU with the CPU results in contention for storage space between the accelerator and the CPU, which may reduce the performance of the system.
[0073] As an alternative approach, in some computer systems, the accelerator has its own internal PSU that is only used by that accelerator. In this approach, in order for the accelerator to fetch instructions or data, the instructions are first copied from the central PSU to the accelerator's PSU. In other words, an additional step is added, and the accelerator's internal PSU processes the central PSU in a manner similar to how the CPU processes the SSU. Note that at this level of complexity, there may be scenarios in some systems where the accelerator can directly retrieve data from the central PSU or directly transfer data from the SSU to its internal PSU. While this solves some of the problems associated with sharing the central PSU described above, giving the accelerator its own internal PSU introduces another inefficiency. Namely, in a basic implementation, the accelerator's internal PSU is opaque; instead, the accelerator's internal PSU is exposed to the programs using the accelerator, and these programs must be responsible for its use. Coupled with the lack of coordination between the virtual memory addresses of the programs running on the GPU and the virtual addresses of the same programs running on the CPU, this means that zero-copy operations are not possible.
[0074] To address this limitation, some computer systems use systems to implement unified memory, such as heterogeneous system architectures. This allows for the coordination of virtual memory used on the internal PSUs of the central PSU and the accelerator, so that pointers can be freely shared between programs running on the accelerator and programs running on the CPU (or other accelerators using the central PSU). In other words, unified memory allows the CPU and the accelerator (such as a GPU) to use the same virtual memory space (for a given program).
[0075] However, when using a system with unified memory, such as a heterogeneous system architecture, for reasons of consistency, the pages used by the accelerator may still have to be transferred into the PSU of the accelerator. However, like the PSU of the system, the space of the PSU of the accelerator is limited. Therefore, the way pages move between the PSU of the accelerator and the PSU of the system is roughly the same as the way pages move between the system PSU and the SSU. However, careful management is required to avoid problems such as page thrashing. In many cases, the data on the PSU of the accelerator is managed by the MMU of the accelerator, the driver of the accelerator running on the host system / OS, or the cooperation between these two components. For example, the memory management function of the PSU of the accelerator is managed by the MMU of the accelerator, the driver of the accelerator running on the host system / OS, or the cooperation between these two components.
[0076] Figure 9 is a simplified diagram of the accelerator described above. As Figure 9 shown, the computer system 900 may include a host 920, and the host 920 is connected to an accelerator 901. The host 920 may include a CPU 921 and a data storage subsystem 922. The data storage subsystem 922 may include multiple connected memory systems, and the connected memory systems 923, 924, and 925 are shown in the figure. Generally, the connected memory systems 923, 924, and 925 may be various memory devices, such as DRAM (e.g., main storage units), SSDs (e.g., secondary storage units), HDDs, magnetic tape drives, etc. Additionally, the CPU 921 is connected to the data storage subsystem 922 and can exchange data with the data storage subsystem 922. Regarding the accelerator 901, according to Figure 9, the accelerator 901 includes an accelerator core 902, a main storage unit 906, a memory management unit 910, and an I / O hub 912. As shown in the figure, the accelerator core 902 may include multiple accelerator cores, and accelerator cores 903, 904, and 905 are shown in the figure. Similarly, the main storage unit 906 may include multiple main storage units, and main storage units 907, 908, and 909 are shown in the figure. In addition, the memory management unit 910 is connected to the accelerator core 902, the main storage unit 906, and the I / O hub 912. The I / O hub 912 itself is connected to the host 920.
[0077] A key part of the memory management of the accelerator is to prefetch pages to avoid unnecessary page faults, because retrieving pages from the system's CPU usually takes a lot of time. However, previous prefetching methods were less efficient. For example, sequential prefetching involves prefetching a fixed number of data pages, which usually results in too little prefetched data and causes many unnecessary page faults. As another example, many prefetching schemes assume that spatial locations are all linear, just like the layout of physical memory. In other words, a particularly important aspect of virtual and physical address spaces is that they are linear and sequential. For example, page 0xFC4A95E3 follows immediately after page 0xFC4A95E2. Many prefetching schemes assume that the data used by a process is similar sequential data. In other words, if a process is using data from page 0xFC4A95E2, then the process is likely to use data from pages adjacent to page 0xFC4A95E2 (not necessarily immediately adjacent, such as 0xFC4A95E1 and 0xFC4A95E3). Although this assumption applies to some processes and some data types, it completely fails in multi-dimensional arrays (arrays with two or more dimensions) because the way a multi-dimensional array is stored in virtual memory may be different from the way the process accesses the individual values of the array.
[0078] For example, take a two-dimensional array as an example.
[0079] This array is stored by rows in memory. So, x0 is stored at (0xA92EF3150), followed by x1 stored at (0xA92EF3151), x2 stored at (0xA92EF3152), x3 stored at (0xA92EF3153), x4 stored at (0xA92EF3154), x5 stored at (0xA92EF3155), and finally, x 15 is stored at (0xA92EF315F). In other words, this array is stored in memory in the format, where i represents the i-th row and j represents the j-th item in the i-th row. In contrast, when using the array The process accesses column by column, that is, the process accesses x0 at memory location (0xA92EF3150), followed by x4 at (0xA92EF3154), x8 at (0xA92EF3152), and x9 at (0xA92EF3153). 12 , accessing x1 at (0xA92EF3154) and x5 at (0xA92EF3155), until finally, accessing x at storage location (0xA92EF315F). 15 In other words, the array is Format. This operation often occurs in matrix multiplication, as the rows of the first matrix are dot-multiplied with the columns of the second. If the prefetch strategy does not take this into account, the prefetch system will hurt system performance because it will transfer data unnecessarily (and may cull other data to make room).
[0080] In order to effectively and efficiently pre-fetch data to the main memory of the accelerator using a unified virtual memory in a heterogeneous computer system, some embodiments of the present disclosure detect an access attempt to the first data page that is not currently stored in the main memory unit of the accelerator (for example, by Figure 12 The memory management unit 1110 of the accelerator or through Figure 18 1811 of the memory management component 1812 of the accelerator). In some embodiments, detecting an attempt to access a data page that is not currently stored on the main storage unit of the accelerator includes detecting a page fault. The page fault may be the result of a core running on the accelerator attempting to access a virtual memory address that is not currently stored on the main storage unit of the accelerator. For example, the page corresponding to the virtual memory address is stored on the main storage unit of the host system, or is exported to the SSU of the host system. In some embodiments, detecting an attempt to access a first data page that is not currently stored on the main storage unit of the accelerator is a prediction that the core running on the accelerator may access the page in the near future. For example, if the core has been accessing consecutive pages in a substantially consecutive order, it can be predicted that the core will attempt to access a page that is slightly further in order and has not yet been loaded into the PSU of the accelerator.
[0081] After detecting an access attempt to a first data page that is not currently stored on a main storage unit of the accelerator, some embodiments may determine (e.g., by Figure 12 The memory management unit 1110 of the accelerator or Figure 18The memory management component 1811) whether the first data page corresponds to a part of an array with multiple dimensions (e.g., one-dimensional array, two-dimensional array, three-dimensional array, four-dimensional array, etc.). For example, if it corresponds to a three-dimensional array, the first data page can store multiple array entries from array entry
[127]
[63] [0] to data entry
[127]
[63]
[1023] . Generally, whether a particular page and the data stored at each memory address on that page correspond to an array can be determined in various ways. For example, in some embodiments, metadata is created to indicate which pages / virtual memory addresses correspond to the array and what the dimensions of the array are. These metadata can be created by the core using the array, or by dedicated instructions / source code for creating metadata, or by analyzing the memory and the access patterns of the array stored in the memory by the core.
[0082] If it is determined that the first data page is part of an array, then in some embodiments, the array is divided into multiple sub-arrays (e.g., by Figure 12 the array partitioning unit 1206 or by Figure 18 the array partitioning processing unit 1817). For clarity, the sub-arrays created by dividing the complete array can be referred to as level-1 sub-arrays. Any sub-sub-arrays created by dividing the level-1 sub-arrays can be referred to as level-2 sub-arrays. Generally, the sub-arrays obtained by dividing the level-n sub-arrays are called level-(n + 1) sub-arrays. For simplicity, the complete undivided array can be referred to as the level-0 sub-array. In some embodiments, the level-1 sub-array containing the first data page can be further divided into level-2 sub-arrays. Generally speaking, in some embodiments, the level-n sub-array containing the first data page can be recursively divided into level-(n + 1) sub-arrays until a division stop condition is met. Note that in many embodiments, the array is only logically divided into sub-arrays, but the array or the memory storing the array does not have to physically change. In other words, the sub-arrays may just be references to specific bounded parts of the array, rather than having physically independent copies of those parts of the array. The information of each sub-array can be recorded as metadata and used in the process of responding to an access attempt to the first data page.
[0083] Figure 10 is a schematic diagram illustrating how an array is recursively divided into sub-arrays. As Figure 10 shown, the array (composed of sub-arrays SA1, SA2, SA3, SA4, SA5, SA6, SA7, and SA8) is divided into 8 sub-arrays. Then these sub-arrays are divided into sub-sub-arrays. Figure 10It also shows how to recursively determine whether a sub-array (or array) should be prefetched. Taking the sub-array SA1 as an example, the sub-array SA1 is divided into 8 sub-sub-arrays. Then these sub-sub-arrays are evaluated to determine whether a threshold number of pages have been retrieved previously. If so, such sub-sub-arrays are marked as meeting the threshold condition (the condition here is that more than half of the pages have been retrieved previously), and these sub-sub-arrays are shown in gray. After determining this, it is determined whether there are enough sub-sub-arrays among the sub-sub-arrays that make up SA1 that meet the threshold condition (here, more than half of the sub-sub-arrays meet). Since 6 out of the 8 sub-sub-arrays of the sub-array SA1 meet the threshold condition, SA1 as a whole is marked as meeting the threshold condition, which means that each page in SA1 (for example, each page in each sub-sub-array that makes up SA1) is also prefetched. Then the process continues, and SA1 is now considered to have met the threshold condition for the next higher-level sub-array.
[0084] In some embodiments, the division stop condition may be based on whether a specific level of sub-array has been reached (e.g., whether the 4th-level sub-array has been reached). In some embodiments, the division stop condition may be based on whether the sub-array is below a specific size (e.g., whether the number of entries in the sub-array is less than a specific number, whether the number of data pages in the sub-array is less than a specific number, or whether the sub-array is a 1D array or its dimension is below a specific size). In some embodiments, the division stop condition may also be based on the activity of the accelerator. In some embodiments, the division stop condition may further be based on the total number of pages being used by the core to which the page belongs. In other words, the division stop condition may be based on the core and the total size of all its data (e.g., when all the pages of the core are retrieved into the main storage unit of the accelerator, the size of the main storage unit occupied by the core). In some embodiments, the division stop condition may further be based on the number of pages or the total size of the array (or an nth-level sub-array). Additionally, in some embodiments, the division stop condition may be based on an evaluation strategy. The evaluation strategy may include criteria / algorithms for determining the optimal granularity level in the array block being evaluated (e.g., the optimal size of various sub-arrays).
[0085] In some embodiments, after the division stop condition is met, it can be determined whether the lowest-level sub-array (i.e., the sub-array with the highest dimension) containing the first data page meets the page access volume condition (e.g., through Figure 12 the PAV determination unit 1207 or by Figure 18PAV processing unit 1818). If the nth-level sub-array meets the page access volume condition, all pages in the nth-level sub-array are selected for prefetching. Then, this process is recursively repeated to evaluate the (n - 1)th-level sub-array containing the nth-level sub-array to determine whether it meets the page access volume condition. If so, all pages in the (n - 1)th-level sub-array are also selected for prefetching. Generally, this process can be recursively repeated until the nth-level sub-array being evaluated does not meet the page access volume condition.
[0086] In some embodiments, the page access volume condition can be based on a threshold number of data pages being fetched onto the PSU of the accelerator for the highest-level sub-array, or for other nth-level sub-arrays, a threshold number of the (n - 1)th-level sub-arrays they form also meet the page access volume condition of the (n - 1)th-level sub-array. In some embodiments, the page access volume condition can be based on the activity of the accelerator. In some embodiments, the page access volume condition can also be based on the total number of pages being used by the core to which the first data page belongs. In other words, the page access volume condition can be based on the core and the total size of all its data (e.g., when all data pages of the core are fetched into the main storage unit of the accelerator, the size of the main storage unit occupied by the core). Similarly, in some embodiments, the page access volume condition can also be based on the number of pages adjacent to the first data page that have already been retrieved (or, equivalently, although not yet retrieved). In other words, the page access volume condition can be based on how much additional data needs to be retrieved for a given page access volume condition. In some embodiments, the page access volume condition can also be based on the number of pages or the total size of the array (or an nth-level sub-array). Additionally, in some embodiments, the page access volume condition can be based on the fetch policy. The fetch policy can include criteria / algorithms for determining the optimal number of pages to fetch / prefetch, and the optimal threshold for setting the page access volume condition to achieve this.
[0087] In some embodiments, the pages that have been selected for prefetching are included in determining whether a sub-array meets its page access volume condition. For example, in some embodiments, the lowest-level (e.g., smallest) level-n sub-array containing data pages is considered first. If a threshold number of pages within the sub-array at this level (e.g., stored in the main storage unit of the accelerator) have been retrieved, then the remaining pages can also be selected for prefetching. Then, some embodiments can consider the next highest-level (e.g., the next smallest level) level-n sub-array (e.g., the level-(n + 1) sub-array). Similarly, if a threshold number of pages have been retrieved from the level-n sub-arrays that make up the level-(n + 1) sub-array, then each page of the level-(n + 1) sub-array that has not been retrieved (e.g., each page in each level-n sub-array that makes up the level-(n + 1) sub-array) can be selected for prefetching. In other words, if a level-(n + 1) sub-array has a threshold number of level-n sub-arrays that meet their page access volume conditions, then the level-(n + 1) sub-array meets the threshold criteria for having all its pages selected for prefetching. Or, if a threshold number of pages have been previously retrieved from a level-(n + 1) sub-array, then each page in that level-(n + 1) sub-array is selected for prefetching. In both cases, the decision to prefetch pages at a lower level is considered (e.g., if pages of a level-n sub-array are being prefetched, then when deciding whether to prefetch pages for the level-(n + 1) sub-array in a recursive process, these pages can be considered as having been retrieved). This process can continue to repeat until the page access volume condition is not met.
[0088] After one of the level-n sub-arrays fails to meet the page access volume condition, the process may end. After the process ends, in some embodiments, the first data page and the data of all the pages selected for prefetching are transferred from the system memory coupled to the accelerator to the main storage unit of the accelerator (e.g., through Figure 12 the acquisition unit 1202 or through Figure 18 the data migration control module 1814). In some embodiments, this involves retrieving data from the main storage unit being used by the CPU of the host system coupled to the accelerator (e.g., Figure 11 the data storage subsystem 1122 in Figure 18 or the storage subsystem 1823 in
[0089] In addition, after detecting an attempt to access the first data page and determining that the first data page is part of an array, some embodiments evaluate the activity of the accelerator in response to detecting the attempt to access the first data page (e.g., through Figure 12the activity evaluation unit 1204 or through Figure 18 the activity processing unit 1815 of the accelerator). In some embodiments, the activities of the accelerator include the activities of the cores running on the accelerator. The core activities include more general kernel activities such as the number of currently running cores, the memory size used by the cores, the memory size referenced / accessed by the cores, and so on. The activities of the accelerator also include the activities of specific cores, for example, the activities of a specific core or a group of cores, such as the duration of core operation, the memory size used by the cores, the memory size actually referenced / accessed by the cores, the memory access pattern of the cores (e.g., the pages referenced / accessed and the relationships between them), etc. The activities of the accelerator can also include recent or historical core activities (including general and specific). For example, some embodiments can consider the activities in the most recent x seconds to evaluate or determine the average value or trend of the activities of the running cores, specific cores, or a group of cores. The activities of the accelerator can also include the predicted future behavior of the cores, whether it is a prediction of the behavior of the cores in general or a prediction of the behavior dedicated to a specific core or a group of cores.
[0090] In addition, in some embodiments, the activities of the accelerator can include the usage of the main storage unit of the accelerator. In some embodiments, the usage of the main storage unit can include more general usage, such as the total amount of pages of stored data, the total amount of remaining free / unused space, the overall storage access pattern of the pages stored in the main storage unit, the relative arrangement of these memory accesses, etc. The activities of the accelerator can also include more specific usages, such as the total amount of pages stored for a specific core, the overall storage access pattern of a specific core to these pages, the residence time of a page or a group of pages in the main storage unit, the number of times a page or a group of pages is referenced / accessed, the time since a page or a group of pages was last referenced / accessed, whether a page or a group of pages is frequently used by different processors (e.g., the CPUs of the host system), whether a specific page shows signs of thrashing, etc. The activities of the accelerator can also include the recent or historical usage of the main storage unit (including general and specific). For example, some embodiments consider the usage of the main storage unit in the most recent x seconds to evaluate or determine the average value or trend of the usage of the main storage unit, whether it is regarding the usage of specific pages or a group of pages belonging to a specific core, or regarding the usage of a single page in a group of pages. The activities of the accelerator can also include the predicted usage of the main storage unit (including general usage or usage based on specific pages).
[0091] In addition, in some embodiments, the activities of the accelerator include additional information related to the accelerator, such as the number of pending kernels (e.g., kernels queued up for execution but not yet running), the characteristics of the pending kernels, the current temperature of the accelerator, the remaining headroom between the current temperature and the maximum operating temperature, the current power usage of the accelerator, etc. The activities of the accelerator can also include recent or historical information of these additional information related to the accelerator. For example, some embodiments can consider the data of these additional information for the most recent x seconds to evaluate the average usage of the main storage unit or determine the trend of its usage. The activities of the accelerator can also include predicted future data based on this additional information.
[0092] Figure 11 FIG. 4 is a schematic diagram of an exemplary computer system 1100 for managing the main storage unit of an accelerator provided by some embodiments of the present disclosure. Although the system 1100 is generally described, the system 1100 can be implemented on any computer configured to execute, for example Figure 12 the methods described, such as servers, desktop computers, laptops, tablets, and so on. As Figure 11 shown, the computer system 1100 includes a host 1120, and the host 1120 is connected to an accelerator 1101. The host 1120 includes a CPU 1121 and a data storage subsystem 1122. The data storage subsystem 1122 includes a plurality of connected memory systems, and the connected memory systems 1123, 1124, and 1125 are shown in the figure. Generally, the connected memory systems 1123, 1124, and 1125 are various memory devices, such as DRAM (e.g., main storage unit), SSD (e.g., secondary storage unit), HDD, magnetic tape drives, etc. Additionally, the CPU 1121 is connected to the data storage subsystem 1122 and exchanges data with the data storage subsystem 1122.
[0093] Regarding the accelerator 1101, according to Figure 11 FIG. 5, the accelerator 1101 includes accelerator cores 1102, a main storage unit 1106, a memory management unit 1110, and an I / O hub 1112. As shown in the figure, the accelerator cores 1102 can include multiple accelerator cores, and the accelerator cores 1103, 1104, and 1105 are shown in the figure. Similarly, the main storage unit 1106 can include multiple main storage units, and the main storage units 1107, 1108, and 1109 are shown in the figure. Additionally, the memory management unit 1110 also has a prefetch engine 1111, and the prefetch engine 1111 can implement, for example Figure 13The method described in. In addition, the memory management unit 1110 is connected to the accelerator core 1102, the main storage unit 1106, and the I / O hub 1112. The I / O hub 1112 itself is connected to the host 1120.
[0094] Figure 12 is a simplified diagram of a memory management unit (MMU) of an accelerator according to some embodiments of the present disclosure. According to Figure 12 , the memory management unit 1110 (from Figure 11 ) has a page fault handling unit 1201, a fetch unit 1202, a page table 1203, and a prefetch engine 1111 (from Figure 11 ). The prefetch engine 1101 itself has an activity evaluation unit 1204, a merging unit 1205, an array partitioning unit 1206, and a PAV determination unit 1207. As shown by the input arrow, the memory management unit 1110 of the accelerator can receive a page fault. When a page fault is received, the page fault handler 1201 determines whether the page is part of an array. If the page is part of an array, the page fault handling unit 1201 can respond to the page fault by preparing to retrieve the page referenced in the page fault. The page fault handling unit 1201 can do this immediately or wait for a period of time before execution. In any case, after the page fault handling unit 1201 starts to respond to the page fault, it can prompt the activity evaluation unit 1204 to evaluate the activity of the accelerator based on the memory state of the accelerator and information about the core (e.g., information about the processes running on the accelerator).
[0095] Based on the activity evaluation of the activity evaluation unit 1204 and the page requested by the page fault, the array partitioning unit 1206 can partition the data array into sub-arrays until a partitioning stop condition is met. The PAV determination unit 1207 can then use the sub-arrays created by the array partitioning unit 1206 to recursively determine which sub-arrays meet the page access volume condition. If a sub-array meets its page access volume condition, the PAV determination unit 1207 can select all the pages in the sub-array for prefetching. The PAV determination unit 1207 can recursively perform this operation on the sub-array containing the first data page until the page access volume condition is not met. The fetch unit 1202 can then retrieve the first data page and any data pages selected for prefetching. After retrieving the data pages, the fetch unit 1202 can create entries for them in the page table 1203. The fetch unit 1202 can then store the retrieved pages in the main storage unit of the accelerator (e.g., Figure 11 one of the main storage units 1106 in, such as PSU 1107).
[0096] In addition, after the page fault handling unit 1201 receives a first page fault, but before the PAV determination unit 1207 starts to determine whether the sub-arrays from the array partitioning unit 1206 meet their page access volume conditions, the page fault handling unit 1201 may receive a second page fault caused by referencing a second data page, which is also part of the array. In this case, the page fault handling unit may prompt the merging unit 1205 to determine whether the page faults are merged into one fetch operation. Based on the activity assessment of the activity assessment unit 1204 and the pages requested in the page faults, the merging unit 1205 may determine whether the first data page referenced in the first page fault and the second data page referenced in the second page fault should be fetched as part of one sub-array (and select other pages for prefetching in the sub-array).
[0097] To this end, the merging unit 1205 may cooperate with the PAV determination unit 1207 to decide whether to set the page access volume condition for the smallest sub-array containing the first data page and the second data page, so as to fetch the two data pages together. If the PAV determination unit 1207 uses the input from the merging unit 1205 to set the page access volume size for the smallest sub-array containing the two pages, such that the sub-array meets the page access volume condition, then the merging unit 1205 merges the fetching of the first data page and the second data page into one fetch operation. Otherwise, the PAV determination unit 1207 may consider fetching the second data page, but will not set the page access volume condition to fetch the two pages together. The merging unit 1205 may also report to the page fault handling unit 1201 that the second page error will not be merged with the first page fault, which may prompt the page fault handling unit 1201 to respond to the second page fault by itself (e.g., in the same way as handling the first page fault).
[0098] Typically, the prefetch engine 1111 may include hardware, software, or a combination of hardware and software. For example, the fetch unit 1202 may be implemented by circuitry on a GPU that receives a memory address and then uses the memory address to retrieve pages from the connected memory system and store them on the PSU of the accelerator. Similarly, in some embodiments, the memory management unit 1110 of the accelerator may have circuitry (the page fault handling unit 1201 is shown in the figure) for initially receiving page faults from the cores running on the accelerator. This circuitry may in turn trigger a page fault handling routine. The page fault handling routine may be firmware on the GPU, and the page fault handling routine triggers this firmware to start execution on the GPU (in a manner similar to how a page fault exception of an application on a CPU triggers a context switch to kernel mode and executes the page fault handling unit of the kernel). Then, the page fault handling routine may implement the activity evaluation unit 1204 by checking the memory state (e.g., the content of the PSU of the accelerator) to evaluate the activity of the accelerator. The page fault handling unit may also implement the array partitioning unit 1206 by using the activity evaluation and the metadata of the array to determine which subarrays the array will be partitioned into. The page fault handling unit may then implement the PAV determination unit 1207 by using the subarrays generated by the array partitioning unit 1206 to set page access volume conditions, determine whether the subarrays meet the page access volume conditions, and select the pages of the subarrays that meet the page access volume conditions for prefetching.
[0099] However, in some embodiments, the activity evaluation unit 1204 (or a part of the activity evaluation unit 1204) may be replaced by a hardware-implemented circuit. For example, the memory management unit 1110 of the accelerator may have circuitry dedicated to monitoring the pages accessed or referenced by the cores running on the accelerator. This information may be stored in a buffer inside the memory management unit 1110 of the accelerator. Then, when a page fault occurs, the page fault handling unit may trigger a firmware routine that uses the information stored in the buffer and may combine it with an analysis of the state and usage of the PSU of the accelerator to arrive at a more complete activity evaluation. In a similar manner, the PAV determination unit 1207 may utilize firmware to select the page access volume conditions and determine which subarrays meet the page access volume conditions, but then use dedicated circuitry to convert the subarrays that meet the page access volume conditions into addresses of data pages for retrieval and then retrieve these addresses from the connected memory system. Additionally, a part of the fetch unit 1202 may be implemented as firmware. For example, the dedicated circuitry within the fetch unit 120 performs the retrieval operation but calls firmware to update the page table 1203. Figure 13 is a flowchart of an exemplary method for retrieving data for an accelerator according to some embodiments of the present disclosure. As Figure 13As shown, in step 1302, an access attempt to a first data page corresponding to a portion of the array is detected. For example, the memory management unit 1110 of the accelerator detects an instruction in which a core running on the accelerator references a specific virtual memory address. Next, in step 1303, it is determined whether the first data page is stored in the main storage unit of the accelerator. This can be performed, for example, by the memory management unit 1110, which can refer to the page table triggering the core to determine whether a page containing the referenced virtual memory address has been assigned a corresponding physical address in the main storage unit of the accelerator. If the first data page is stored in the main storage unit of the accelerator, the method returns to step 1302. On the other hand, if the first data page is not stored in the main storage unit of the accelerator, the method proceeds to step 1304.
[0100] In step 1304, the array is divided into level-1 sub-arrays. For example, the array partitioning unit 1206 of the prefetch engine 1111 can divide the array into 2 阵列的维度 sub-arrays of the same size. This may involve creating metadata to indicate the pages or memory addresses used to mark the boundaries of each sub-array. Then, in step 1305, it is determined whether the current level-n sub-array (the first iteration is the level-1 sub-array) meets the partitioning stop condition. If, in step 1305, the current level-n sub-array does meet the partitioning stop condition, the method proceeds to step 1307. On the other hand, if the current level-n sub-array does not meet the partitioning stop condition, the method proceeds to step 1306. In step 1306, the current level-n sub-array containing the first data page is divided into level-(n + 1) sub-arrays. Then, the method returns to step 1305, and the level-(n + 1) sub-arrays created in step 1306 become the current level-n sub-arrays.
[0101] As described above, if, in step 1305, the current level-n sub-array does meet the partitioning stop condition, then the method proceeds to step 1307. In step 1307, it is determined whether the current level-n sub-array containing the first data page meets the page access volume condition. For example, the PAV determination unit 1207 of the prefetch engine 1111 determines whether, for the lowest-level level-n sub-arrays, a threshold number of data pages (e.g., 50%, including the level-n sub-arrays) have currently been transferred to the PSU of the accelerator. Similarly, the PAV determination unit 1207 of the prefetch engine 1111 determines whether, for higher-level level-n sub-arrays, a threshold number of level-(n + 1) sub-arrays (e.g., 50%, including the level-n sub-arrays) have met their respective access volume conditions. If it is determined that the current level-n sub-array containing the first data page does not meet the page access volume condition, the method proceeds to step 1309.
[0102] On the other hand, if it is determined that the n - th level sub - array currently containing the first data page does satisfy the page access volume condition, then the method proceeds to step 1308. In step 1308, all pages of the n - th level sub - array currently containing the first data page (which have not been retrieved or selected for retrieval) are selected for pre - fetching. For example, the PAV determination unit 1207 of the pre - fetch engine 1111 uses the metadata generated by the array partitioning unit 1206 to determine all pages of the current n - th level sub - array. Then, the method returns to step 1307, where the (n + 1) - th level sub - array containing the n - th level sub - array becomes the new n - th level sub - array.
[0103] As described above, if it is determined in step 1307 that the n - th level sub - array currently containing the first data page does not satisfy the page access volume condition, then the method proceeds to step 1309. In step 1309, the first data page and any data pages selected for pre - fetching are fetched from the memory system connected to the accelerator into the main storage unit of the accelerator. This may involve, for example, the acquisition unit 1202 accessing the connected memory system (e.g., Figure 11 the connected memory system 1122 in
[0104] Figure 14 is a more detailed schematic diagram of the architecture of an exemplary accelerator according to some embodiments of the present disclosure. Specifically, Figure 14 a neural network processing architecture 1400 is shown, which can be an accelerator unit for artificial neural networks, an FPGA, an ASIC, or various other types of accelerators. As Figure 14 shown, the processing architecture 1400 may include an accelerator processing system 1402, a memory controller 1404, a direct memory access (DMA) unit 1406, a global memory 1408, a joint test action group (JTAG) / test access port (TAP) controller 1410, a peripheral interface 1412, a bus 1414, etc. It should be understood that the accelerator processing system 1402 can perform algorithmic operations (e.g., operations of machine learning) based on the communicated data.
[0105] The accelerator processing system 1402 may include a command processor 1420 and multiple accelerator cores 1430. The command processor 1420 can be used to control and coordinate one or more accelerator cores, and accelerator cores 1431, 1432, 1433, 1434, 1435, 1436, 1437, 1438, and 1439 are shown in the figure. Each of the accelerator cores 1430 can provide a set of synapse / neuron circuits (e.g., for an artificial neural network) for parallel computing. For example, Figure 14The first layer of the accelerator core 1430 therein may provide circuitry representing the input layer of an artificial neural network, while the second layer of the accelerator core 1430 may provide circuitry representing the hidden layer of the artificial neural network. In some embodiments, the accelerator processing system 1402 may be implemented as one or more GPUs, NPUs, TPUs, FPGAs, ASICs, or other heterogeneous accelerator units.
[0106] For example, the accelerator core 1430 may include one or more processing elements, each processing element including a single instruction multiple data (SIMD) architecture that includes one or more processing units configured to perform one or more operations (e.g., multiplication, addition, multiply-accumulate, etc.) based on instructions received from the command processor 1420. To perform operations on the communicated data packets, the accelerator core 1430 may include one or more processing elements for processing the information in the data packets. Each processing element may include any number of processing units. In some embodiments, the accelerator core 1430 may be considered a tile or the like. In some embodiments, multiple accelerator cores 1430 may be communicatively coupled to each other. For example, multiple accelerator cores 1430 may be coupled to a single directional ring bus to support efficient pipelining for large neural network models. The structure of the accelerator core 1430 will be described in detail below with reference to Figure 15 the structure of the accelerator core 1430 will be described in detail below with reference to
[0107] The processing architecture 1400 of the accelerator may also communicate with the host unit 1440. The host unit 1440 may be one or more processing units (e.g., X86 central processing units). As Figure 14 shown, the host unit 1440 may be associated with a host memory 1442. In some embodiments, the host memory 1442 may be an internal memory or an external memory associated with the host unit 1440. In some embodiments, the host memory 1442 may include a host disk, which is an external memory configured to provide additional memory to the host unit 1440. The host memory 1442 may be a double data rate synchronous dynamic random access memory (e.g., DDR SDRAM), etc. Compared with on-chip memory (serving as a higher-level cache) integrated within one or more processors, the host memory 1442 may be configured to store a large amount of data at a slower access speed. The data stored in the host memory 1442 may be transferred to the processing architecture 1400 of the accelerator for executing the neural network model.
[0108] In some embodiments, a host system having a host unit 1440 and a host memory 1442 may include a compiler (not shown). A compiler is a program or computer software that converts computer code written in one programming language into NPU instructions to create an executable program. In the application of machine learning, the compiler may perform various operations, such as preprocessing, lexical analysis, parsing, semantic analysis, converting the input program into an intermediate representation, code optimization, and code generation, or a combination thereof. For example, the compiler may compile a neural network to generate static parameters, such as the connections between neurons and the weights of neurons.
[0109] In some embodiments, the host system 1440 may push one or more commands to the accelerator processing system 1402. As described above, these commands may be further processed by the command processor 1420 of the accelerator processing system 1402, temporarily stored in the instruction buffer of the processing architecture 1400 of the accelerator, and then assigned to one or more corresponding accelerator cores (e.g., accelerator cores 1431 and 1432) or processing elements. Some commands may instruct the DMA unit 1406 to load instructions (generated by the compiler) and data from the host memory 1442 into the global memory 1408. The loaded instructions may then be assigned to each accelerator core assigned with a corresponding task, and these accelerator cores may process these instructions.
[0110] It can be understood that the first few instructions received by the accelerator core 1430 may instruct the accelerator core 1430 to load / store data from the host memory 1442 into one or more local memories of the accelerator core (e.g., Figure 14 local memory 1512). Each of the accelerator cores 1430 may then start an instruction pipeline, which includes fetching instructions from the instruction buffer (e.g., via a sequencer), decoding the instructions (e.g., via the DMA unit 1406), generating local memory addresses (e.g., corresponding to operands), reading source data, performing or loading / storing operations, and then writing back the results.
[0111] The command processor 1420 may interact with the host unit 1440 and transfer relevant commands and data to the accelerator processing system 1402. In some embodiments, the command processor 1420 may interact with the host unit 1440 under the supervision of a kernel mode driver (KMD). In some embodiments, the command processor 1420 may modify the relevant commands for each accelerator core such that the accelerator cores 1430 may work in parallel as much as possible. The modified commands may be stored in the instruction buffer. In some embodiments, the command processor 1420 may be configured to coordinate one or more accelerator cores to execute in parallel.
[0112] The memory controller 1404, also known as the memory management unit, can manage reading or writing data from a specific memory block in the global memory 1408 having on-chip memory blocks (e.g., second-generation high-bandwidth memory blocks (HBM2)), where the on-chip memory blocks serve as the main storage unit / main memory of the accelerator. For example, the memory controller 1404 can manage reading / writing data from outside the accelerator processing system 1402 (e.g., from the DMA unit 1406 or a DMA unit corresponding to another NPU) or from inside the accelerator processing system 1402 (e.g., from the local memory of the accelerator core, e.g., the accelerator core 1431, and via the command processor 1420 controlled by a 2D grid). Additionally, although Figure 14 only one memory controller is shown, it is understood that more than one memory controller may be provided in the accelerator unit 1400. For example, each memory block (e.g., HBM2) within the global memory 1408 has a memory controller. In some embodiments, the global memory 1408 can store instructions and data from the host memory 1442 via the DMA unit 1406. Then, these instructions can be distributed to the instruction buffers of each accelerator core assigned with a corresponding task, and these accelerator cores process these instructions accordingly.
[0113] The memory controller 1404 can generate memory addresses and initiate memory read or write cycles. The memory controller 1404 may include several hardware registers that can be written to and read by one or more processors. These registers can include a memory address register, a byte count register, one or more control registers, and other types of registers. These registers can specify a combination of a source, a destination, a transfer direction (reading from / writing to an input / output (I / O) device), the size of the transfer unit, the number of bytes transferred in one burst, or other typical characteristics of the memory controller. The memory controller 1404 may also have a prefetch engine (not shown), such as Figure 11 and Figure 12 the prefetch engine 1111 in
[0114] The DMA unit 1406 can assist in transferring data between the main memory 1442 and the global memory 1408. For example, the DMA unit 1406 can assist in loading data or instructions from the host memory 1442 into the local memory of the accelerator core 1430. The DMA unit 1406 can also assist in transferring data between multiple accelerators. In addition, the DMA unit 1406 can assist in transferring data between multiple NPUs (e.g., the accelerator processing system 1402 implemented on the NPU). For example, the DMA unit 1406 can assist in transferring data between multiple accelerator cores 1430 or within each accelerator core. The DMA unit 1406 can allow off-chip devices to access on-chip and off-chip memories without causing a CPU interruption. Therefore, the DMA unit 1406 can also generate memory addresses and initiate memory read or write cycles. The DMA unit 1406 can also include several hardware registers that can be written to and read by one or more processors, and these hardware registers include memory address registers, byte count registers, one or more control registers, and other types of registers. These registers can specify a combination of source, destination, transfer direction (read from or write to an I / O device), size of the transfer unit, or the number of bytes transferred in one burst. It should be understood that the accelerator unit 1400 may include a second DMA unit that can be used to transfer data between other multiple neural network processing architectures to allow multiple neural network processing architectures to communicate directly without involving the host CPU.
[0115] The JTAG / TAP controller 1410 can specify a dedicated debug port to implement a serial communication interface (e.g., the JTAG interface) to access the NPU with low overhead without requiring direct external access to the system address and data buses. The JTAG / TAP controller 1410 can also have an on-chip test access interface (e.g., the TAP interface) that implements a protocol for accessing a set of test registers that are used to present the chip logic levels and device capabilities of individual parts.
[0116] The peripheral interface 1412 (e.g., a Peripheral Component Interconnect Express (PCIe) interface), if present, is used as (and typically is) an inter-chip bus to provide communication between the accelerator unit 1400 and other devices. The bus 1414 (e.g., an I2C bus) includes an intra-chip bus and an inter-chip bus. The intra-chip bus interconnects all the internal components required by the system architecture. Although not all components are connected to each other, all components have some connections to other components with which they need to communicate. The inter-chip bus connects the NPU to other devices, such as off-chip memory or peripherals. For example, the bus 1414 can provide high-speed communication across the accelerator cores and can also (via the accelerator processing system 1402) connect the accelerator cores 1430 to other units such as off-chip memory or peripherals. Typically, if the peripheral interface 1412 (e.g., an inter-chip bus) is present, the bus 1414 is only related to the intra-chip bus, although in some implementations, the bus 1414 can still be related to dedicated inter-bus communication.
[0117] The accelerator processing system 1402 can be configured to perform operations based on an artificial neural network. Although in some embodiments of the present disclosure, the accelerator processing architecture 1400 can be used for convolutional neural networks, it should be understood that the accelerator processing architecture 1400 can be used for various neural networks, such as deep neural networks (DNNs), recurrent neural networks (RNNs), etc. In addition, some embodiments can be configured for various processing architectures, such as CPUs, GPGPUs, GPUs, NPUs, TPUs, FPGAs, ASICs, any other type of heterogeneous accelerator processing units (HAPUs), etc.
[0118] In operation, according to some embodiments of the present disclosure, the DMA unit 1406 can be used to transfer an artificial neural network from the host memory 1442 to the accelerator unit 1400. The host unit 1440 can be connected to the accelerator processing architecture 1400 via the peripheral interface 1412. In some embodiments, the artificial neural network and intermediate values of the artificial neural network can be stored in the global memory 1408 controlled by the memory controller 1404. Finally, the artificial neural network is run on the AI processor 1402, and the command processor 1420 manages the processing of the artificial neural network inputs.
[0119] Figure 15 is a schematic diagram of the architecture of a core of an exemplary accelerator according to some embodiments of the present disclosure. As Figure 15 shown, the accelerator core 1501 (e.g., Figure 14 the accelerator core 1430) can include one or more operation units, such as a first operation unit 1502 and a second operation unit 1504, a memory engine 1506, an sequencer 1508, an instruction buffer 1510, a constant buffer 1515, a local memory 1512, etc.
[0120] One or more operation units may include a first operation unit 1502 and a second operation unit 1504. The first operation unit 1502 may be configured to perform operations on received data (e.g., a matrix). In some embodiments, the first operation unit 1502 may include one or more processing units configured to perform one or more operations (e.g., multiplication, addition, multiply-accumulate, element-wise operations, etc.). In some embodiments, the first operation unit 1502 is configured to accelerate the execution of convolution operations or matrix multiplication operations. The second operation unit 1504 is configured to perform pooling operations, interpolation operations, region of interest (ROI) operations, etc. In some embodiments, the second operation unit 1504 includes an interpolation unit, a pooling data path, etc.
[0121] The memory engine 1506 may be configured to perform data copying inside the corresponding accelerator core 1501 or between two accelerator cores. The DMA unit 1406 (from Figure 14 ) may assist in copying data inside the corresponding accelerator core 1501 or between two accelerator cores. For example, the DMA unit 1406 may support the memory engine 1506 in performing data copying from local memory (e.g., Figure 15 local memory 1512) to the corresponding operation unit. The memory engine 1506 may also be configured to perform matrix transpose to make the matrix suitable for use in the operation unit.
[0122] The sequencer 1508 may be coupled to the instruction buffer 1510 and is configured to fetch commands and distribute the commands to the components of the accelerator core 1501. For example, the sequencer 1508 may distribute a convolution command or a multiplication command to the first operation unit 1502, a pooling command to the second operation unit 1504, or a data copy command to the memory engine 1506. The sequencer 1508 may also be configured to monitor the execution of neural network tasks and parallelize subtasks of the neural network tasks to improve execution efficiency. In some embodiments, the first operation unit 1502, the second operation unit 1504, and the memory engine 1506 may run in parallel under the control of the sequencer 1508 according to the instructions stored in the instruction buffer 1510.
[0123] The instruction buffer 1510 may be configured to store instructions belonging to the corresponding accelerator core 1501. In some embodiments, the instruction buffer 1510 is coupled to the sequencer 1508 and provides instructions to the sequencer 1508. In some embodiments, the instructions stored in the instruction buffer 1510 may be processed by the command processor 1420 (from Figure 14)Transmitted or modified. The constant buffer 1514 can be configured to store constant values. The constant values stored in the constant buffer 1514 can be used by operation units such as the first operation unit 1502 or the second operation unit 1504 for operations such as batch normalization, quantization, de - quantization, etc.
[0124] The local memory 1512 can provide a storage space with fast read / write speeds. To reduce possible interactions with the global memory, the storage space of the local memory 1512 can be implemented with a large capacity. With a large - capacity storage space, most data accesses can be performed within the accelerator core 1501, thereby reducing the latency caused by data access. In some embodiments, to minimize the latency and energy consumption caused by data loading, a static random - access memory (SRAM) integrated on the chip can be used as the local memory 1512. In some embodiments, the local memory 1512 can have a capacity of 192MB or more. According to some embodiments of the present disclosure, the local memory 1512 can be evenly distributed on the chip to alleviate dense wiring and heat - generation problems.
[0125] Figure 16 is a schematic diagram of an alternative architecture of an exemplary accelerator according to some embodiments of the present disclosure. Similar to Figure 14 the exemplary accelerator architecture, Figure 16 shows an accelerator architecture 1600 on which some embodiments of the present disclosure can be implemented. In various embodiments, the accelerator architecture 1600 can be an accelerator for artificial neural networks, an FPGA, an ASIC, or various other types of accelerators. As Figure 16 shown, the architecture 1600 can include a heterogeneous computing unit (HCU) 1601 and corresponding host unit 1610 and host memory 1611, etc. It should be understood that the HCU 1601 can be a dedicated computing device for facilitating neural network computing tasks. For example, the HCU 1601 can perform algorithm operations (e.g., machine - learning operations) based on the communicated data. The HCU 1601 can be an accelerator, such as a GPU, NPU, TPU, FPGA, ASIC, etc.
[0126] The HCU 1601 can include one or more computing units 1602, a memory hierarchy 1605, a controller 1606, and an interconnect unit 1607. Each computing unit 1602 can read data from the memory hierarchy 1605, write data to the memory hierarchy 1605, and perform algorithm operations (e.g., multiplication, addition, multiply - accumulate, etc.) on the data. In some embodiments, the computing unit 1602 can include multiple engines for performing different operations. For example, as Figure 16As shown, the computing unit 1602 may include a dot product engine 1603, a vector engine 1604, etc. The dot product engine 1603 may perform dot product operations such as multiplication and convolution. The vector engine 1604 may perform vector operations such as addition.
[0127] The memory hierarchy 1605 may have on-chip memory blocks (e.g., 4 blocks of HBM2) as the main storage unit / main memory. The memory hierarchy 1605 may also have a memory controller or a memory management unit not shown. The memory hierarchy 1605 may store data and instructions and provide high-speed access to the data and instructions it stores to other components such as the computing unit 1602 and the interconnect unit 1607. The interconnect unit 1607 may transfer data between the HCU 1602 and other external components such as a host unit or another HCU. The interconnect unit 1607 may include a PCIe interface 1608 and an inter-chip connection 1609. The PCIe interface 1608 provides communication between the HCU and the host unit 1610 or Ethernet. The inter-chip connection 1609 serves as an inter-chip bus and connects the HCU to other devices such as other HCUs, off-chip memories, or peripherals. The memory management unit of the memory hierarchy 1605 may also have a prefetch engine (not shown), e.g., Figure 11 the prefetch engine 1111 in []. The prefetch engine in the memory management unit may control the number of pages prefetched from the host memory 1611.
[0128] The controller 1606 may control and coordinate the operations of other components such as the computing unit 1602, the interconnect unit 1607, and the memory hierarchy 1605. For example, the controller 1606 may control the dot product engine 1603 or the vector engine 1604 in the computing unit 1602 and the interconnect unit 1607 to facilitate parallelization among these components.
[0129] The host memory 1611 may be an off-chip memory, such as the memory of the host CPU. For example, the host memory 1611 may be a DDR memory (such as DDR SDRAM), etc. Compared with the on-chip memory integrated within one or more processors (as a higher-level cache), the host memory 1611 is configured to store a large amount of data at a slower access speed. The host unit 1610 may be one or more processing units (e.g., an X86 CPU). In some embodiments, the host system having the host unit 1610 and the host memory 1611 may include a compiler (not shown). A compiler is a program or computer software that converts computer code written in a programming language into instructions for the HCU 1601 to create an executable program. In the application of machine learning, the compiler may perform various operations, such as preprocessing, lexical analysis, parsing, semantic analysis, converting the input program into an intermediate representation, code optimization, and code generation, or a combination thereof.
[0130] Figure 17 FIG. shows a schematic diagram of an exemplary cloud system 1706 equipped with a neural network processing architecture 1701 according to an embodiment of the present disclosure. As Figure 17 shown, the cloud system 1706 may provide cloud services with artificial intelligence (AI) capabilities and may include multiple computing servers (e.g., computing servers 1707 and 1708). In some embodiments, the computing server 1707 may be installed with an accelerator architecture 1400 ( Figure 14 ) or an accelerator architecture 1600 ( Figure 16 ). For simplicity and clarity, Figure 17 the neural network processing architecture 1701 shown is a simplified version of the accelerator architecture 1700.
[0131] With the assistance of the neural network processing architecture 1600, the cloud system 1706 may provide extended AI capabilities such as image recognition, face recognition, translation, 3D modeling, etc. It can be understood that the neural network processing architecture 1600 may be deployed to computing devices in other forms. For example, the neural network processing architecture 1600 may also be integrated in computing devices such as smartphones, tablets, and wearable devices. Additionally, although Figures 14 to 17 a specific architecture is shown, it can be understood that any HCU or accelerator providing the ability to perform parallel computing may be used.
[0132] Figure 18 is a simplified diagram for schematically showing how a host system, such as through a device driver, decides and implements the selection and prefetching of data pages from an nth-level sub-array. More precisely, Figure 18illustrates how the CPU 1822 of the host system 1821 executes the driver 1810 of the accelerator 1801 to enable the host 1820 to interface with the accelerator 1801. Specifically, the host system 1821 can utilize the memory management component 1811 of the driver 1810 to determine which pages are to be prefetched along with the first data page referenced in a page fault. According to Figure 18 , the GMMU 1803 of the graphics processing unit (GPU) 1802 sends information about the page fault, as well as core information (e.g., information about the core that triggered the page fault) and memory status information (e.g., information about the memory usage of the main storage unit of the accelerator) to the memory management component 1811 of the driver 1810.
[0133] The memory management component 1811 can respond to the page fault through the page fault handling unit 1813, which can determine whether the data page referenced in the page fault corresponds to a part of an array. If the page fault handling unit 1813 determines that the data page does correspond to a part of an array, the page fault handling unit 1813 can trigger the memory management component 1811 to evaluate the activity of the accelerator using the active processing unit 1815 of the accelerator based on the memory status of the accelerator and information about the core (e.g., information about the processes running on the accelerator). In some embodiments, the active processing unit 1815 of the accelerator can be software executed on the CPU 1822 that analyzes the current state and usage of the PSU of the accelerator (or provides this information as part of the page fault) to generate an activity assessment of the accelerator.
[0134] Based on the activity assessment of the active processing unit 1815 of the accelerator and the pages requested in the page fault, the array partitioning processing unit 1817 can partition the array into sub-arrays until a partitioning stop condition is met. In some embodiments, the array partitioning processing unit 1817 can be software executed on the CPU 1822 that uses the activity assessment generated by the active processing unit 1815 to recursively partition the array into sub-arrays until the partitioning stop condition is met and creates metadata recording the boundaries of each sub-array.
[0135] Next, the PAV handler 1818 then uses the sub-arrays created by the array partitioning processing unit 1817 to recursively determine which sub-arrays meet the page access volume condition. If a sub-array meets its page access volume condition, the PAV processor 1818 can select all the pages in the sub-array for prefetching. The PAV processor 1818 can recursively perform this operation on the sub-array containing the first data page until the page access volume condition is not met. Then, the data migration control module 1814 can retrieve the first data page and any data pages selected for prefetching. These pages can be retrieved from the memory subsystem 1823 of the host system 1821 and may pass through the interface with the CPU 1822. Finally, the data migration control module 1814 can then send these pages (e.g., perform page migration) to the GPU 1802.
[0136] In some embodiments, partitioning stop conditions are determined for each level of sub-arrays, and these conditions may be different. For example, in some embodiments, the partitioning stop condition depends on the number of pages contained in the sub-array. In contrast, in some embodiments, for each level of sub-arrays, the partitioning stop condition can be the same (or the same for certain levels of sub-arrays). Similarly, in some embodiments, for each instance of responding to an access attempt to a data page not stored on the main storage unit of the accelerator, partitioning stop conditions are determined, and these conditions may be different. Conversely, in some embodiments, partitioning stop conditions are not determined for each page fault. For example, in some embodiments, the partitioning stop condition is statically set for each core at the time the core starts execution. In other embodiments, the partitioning stop condition may instead change only at set intervals. For example, in some embodiments, the partitioning stop condition can only change after a certain period of time has elapsed or a certain number of page faults have occurred. Similarly, in some embodiments, the partitioning stop condition can be changed when certain events occur.
[0137] In addition, in some embodiments, the partitioning stop condition is set based on the array (e.g., no sub-array partitioning) so that no pages are selected for prefetching. In some embodiments, this is done when the main storage unit of the accelerator approaches a certain capacity threshold (e.g., when the unused rate of the PSU of the accelerator is less than 5%). In some embodiments, the partitioning stop condition is set based on the array when the core using the first data page approaches a certain number of pages (in terms of data) stored on the PSU of the accelerator. This can be relative (e.g., when the core uses more than 40% of the capacity of the PSU of the accelerator) or absolute (when the core uses more than 10 Gb of the storage capacity of the accelerator PSU).
[0138] In some embodiments, the partitioning stop condition can be based on an evaluation of the frequency and degree to which a sub-array meets the page access volume condition. For example, this can be useful in instances where it is detected whether a sub-array is finely partitioned (e.g., if certain levels of sub-arrays almost always meet the page access volume condition) or in instances where it is detected that a sub-array is coarsely partitioned (e.g., if there are often no sub-arrays that meet the page access volume condition). In some embodiments, this may involve, for example, recording what sub-arrays were created and whether they meet the page access volume condition. These logs can also include timestamps of when the processes occurred, keeping only the x most recent entries, keeping entries only for a certain period of time, or keeping entries only for x instances of page selections for prefetching.
[0139] In some embodiments, the partitioning stop condition is based on the partitioning stop condition used to respond to the previous page fault caused by the triggering core. Additionally, in some embodiments, the partitioning stop condition can also be based on an evaluation strategy. The evaluation strategy can include a criterion / algorithm by which the optimal granularity level (e.g., the optimal size of various sub-arrays) is determined when evaluating an array block. In some embodiments, the evaluation strategy is based on an evaluation of multiple data pages that make up the array, which have been retrieved and are currently stored in the main storage unit of the accelerator.
[0140] In some embodiments, page access volume conditions are determined for each level of sub-array, and these conditions may be different. For example, in some embodiments, the page access volume condition depends on the number of pages contained in the sub-array. In contrast, in some embodiments, the page access volume condition can be the same for each level of sub-array (or the same for a group of sub-arrays at certain levels). Similarly, in some embodiments, page access volume conditions are determined for each instance of an access attempt to a data page that is not stored in the main storage unit of the accelerator, and these conditions may not be the same. In contrast, in some embodiments, page access volume conditions are not determined for each page fault. For example, in some embodiments, the page access volume condition is statically set for each core at the time the core starts execution. In other embodiments, the page access volume condition instead changes only at set intervals. For example, in some embodiments, the page access volume condition changes only after a certain amount of time has elapsed or after a certain number of page faults have occurred. Similarly, in some embodiments, the page access volume condition changes when certain events occur.
[0141] In addition, in some embodiments, a page access volume condition is set such that no sub-array meets the page access volume condition. In some embodiments, this is done when the main storage unit of the accelerator approaches a certain capacity threshold (e.g., when the unused rate of the PSU of the accelerator is less than 5%). In some embodiments, when the core using the first data page approaches a certain number of pages (in terms of data) stored on the PSU of the accelerator, the page access volume condition can be set based on the array. This can be relative (e.g., when the core uses more than 40% of the storage capacity of the PSU of the accelerator) or absolute (when the core uses more than 10 Gb of storage capacity of the PSU of the accelerator).
[0142] In some embodiments, the page access volume condition is based on an evaluation of the frequency and degree to which the pages selected for prefetching are retrieved and then evicted within a certain time span. For example, this is useful for detecting instances where the page access volume condition is set too low (e.g., if most of the pages selected for prefetching are not used before being evicted). In some embodiments, this may involve, for example, recording which pages are selected for prefetching and the page access volume conditions used. It may also involve recording which pages are evicted. These logs can also include timestamps of when the eviction occurred, and these logs may only retain x latest entries, only retain entries for a certain period of time, or only retain entries for the pages of x retrievals. In some embodiments, the logs of the evicted pages can be evaluated periodically to determine whether the pages prefetch are often evicted before being accessed.
[0143] In some embodiments, the page access volume condition can be based on the page access volume condition for responding to the previous page fault caused by the trigger core. In addition, in some embodiments, the page access volume condition can also be based on the acquisition policy. The acquisition policy can include criteria / algorithms for determining the optimal number of pages to acquire / prefetch, and the optimal threshold for setting the page access volume condition to achieve this. In some embodiments, the acquisition policy can be based on an evaluation of the memory access pattern of the trigger core.
[0144] In some embodiments, an access attempt (e.g., page fault) to a data page not stored on the PSU of the accelerator is responded to immediately. In some other embodiments, the page fault is not responded to immediately. For example, in some embodiments, the page fault can only be responded to after a certain duration has elapsed since the page fault was received. In some embodiments, the page fault can only be responded to at a set interval. For example, in some embodiments, the page fault can only be responded to once every x microseconds.
[0145] In some embodiments, multiple access attempts to data pages that are not stored on the PSU of the accelerator are merged (e.g., merged together). For example, in some embodiments, after receiving a first page fault, but before responding to it, a second page fault is received from the same triggering core. In some embodiments, when multiple page faults from the same core have been received but not yet responded to, the two page faults can be merged together (e.g., by the merge unit 1205 of Figure 12 or by the merge processing unit 1816 of Figure 18 ). For example, in some embodiments, the first page (referenced by the first page fault) and the second page (referenced by the second page fault) can be compared. If the data of the two pages is close enough to each other, a page access volume condition is set such that the smallest subarray containing the data of the first page and the second page satisfies the page access volume condition. This will ensure that the second page data is selected for prefetching together with the first page data. Whether the page access volume condition can be set such that the smallest subarray containing the two data pages satisfies the page access volume condition depends on the activity assessment of the accelerator. In some embodiments, if the page access volume condition is not set such that the smallest subarray containing the two data pages satisfies its page access volume condition, the page faults may not be merged. Figure 12 or by the merge processing unit 1816 of Figure 18 Figure 18 For example, in some embodiments, the first page (referenced by the first page fault) and the second page (referenced by the second page fault) can be compared. If the data of the two pages is close enough to each other, a page access volume condition is set such that the smallest subarray containing the data of the first page and the second page satisfies the page access volume condition. This will ensure that the second page data is selected for prefetching together with the first page data. Whether the page access volume condition can be set such that the smallest subarray containing the two data pages satisfies the page access volume condition depends on the activity assessment of the accelerator. In some embodiments, if the page access volume condition is not set such that the smallest subarray containing the two data pages satisfies its page access volume condition, the page faults may not be merged.
[0146] Correspondingly, in some embodiments, after receiving a first page fault, a second page fault is received, and other data pages have been selected for prefetching, and the process of retrieving the data pages has started. In this case, in some embodiments, it is determined whether the data page referenced in the second page fault is included in the data pages selected for prefetching in response to the first page fault. If the page referenced in the second page fault is already included in the pages selected for prefetching, the second page fault can be merged with the first page fault because the second data page has already been retrieved as part of the response to the first page fault.
[0147] In some embodiments, retrieving one or more of the referenced pages is triggered when a core running on the accelerator attempts to access one or more of the referenced pages.
[0148] In some embodiments, the memory system from which the data pages are retrieved is the main storage unit of the host system (e.g., the system in which the accelerator is running). In some of these embodiments, the accelerator and the host system can implement a unified virtual memory system.
[0149] In some embodiments, the accelerator can be a graphics processing unit (GPU), a neural network processing unit (NPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC) or other heterogeneous accelerator unit.
[0150] In some embodiments, a non-transitory computer-readable storage medium storing instructions is also provided, and the instructions can be executed by a device to perform the above method. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAMs, PROMs, and EPROMs, FLASH-EPROMs, or any other flash memory, NVRAMs, caches, registers, any other storage chip or cartridge memory, and networked versions thereof. The device may include one or more processors, input / output interfaces, network interfaces, and / or memories.
[0151] It should be noted that relational terms such as "first" and "second" herein are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. Additionally, words such as "comprising", "having", "including", and other similar forms have equivalent meanings and are open-ended, because one or more items following any of these words do not mean an exhaustive listing of these items, nor are they limited to the one or more items listed.
[0152] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations, except where infeasible. For example, if it is stated that a component may include A or B, then unless otherwise specifically stated or infeasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then unless otherwise specifically stated or infeasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0153] It can be understood that the above embodiments can be implemented by hardware or software (program code) or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. When executed by a processor, the software can perform the disclosed method. The devices, modules, and other functional units described in this disclosure can be implemented by hardware or software or a combination of hardware and software. Those of ordinary skill in the art should also understand that the above devices, modules, and other functional units can be combined or further divided into multiple subunits.
[0154] The embodiments can be further described by the following sentences.
[0155] 1. A method for retrieving data for an accelerator, the method comprising:
[0156] detecting an access attempt to a first data page not stored on the main storage unit of the accelerator, the first data page corresponding to a portion of an array having multiple dimensions; and
[0157] In response to detecting an access attempt to the first data page:
[0158] Partition the array into sub - arrays by:
[0159] Partition the array into a plurality of first - level sub - arrays, and
[0160] Partition the first - level sub - arrays into a plurality of second - level sub - arrays, wherein the first - level sub - arrays contain the first data page;
[0161] Select pages for pre - fetching, wherein the selection of pages for pre - fetching includes:
[0162] If a first second - level sub - array meets the page access volume condition, select all pages in the first second - level sub - array for pre - fetching, wherein the first second - level sub - array contains the first data page; and
[0163] Transfer the first data page and any data pages selected for pre - fetching from the storage system connected to the accelerator to the main storage unit.
[0164] 2. The method according to claim 1, wherein after selecting all pages in the first second - level sub - array for pre - fetching, the selection of pages for pre - fetching further includes:
[0165] In response to the first second - level sub - array meeting the page access volume condition, if a first first - level sub - array meets the page access volume condition, select all pages in the first first - level sub - array for pre - fetching.
[0166] 3. The method according to claim 2, wherein after selecting all pages in the first first - level sub - array for pre - fetching, the selection of pages for pre - fetching further includes: In response to the first first - level sub - array meeting the page access volume condition, if the array meets the page access volume condition, select all pages in the array for pre - fetching.
[0167] 4. The method according to claim 3, further comprising: after partitioning the first - level sub - arrays into a plurality of second - level sub - arrays, partitioning the array into a plurality of sub - arrays by:
[0168] Partition the n - th level sub - array with the smallest level containing the first data page into (n + 1) - th level sub - arrays and continue to recursively partition the (n + 1) - th level sub - arrays including the first data page until a partitioning stop condition is met.
[0169] Among them, when the division stop condition is satisfied, k is the value of n, the value of n starts from 2 and ends at k, and before selecting all the pages in the first second-level subarray for prefetching, the selection of pages for prefetching is achieved through the following operations:
[0170] If the m-level subarray meets the page access volume condition, then select all the pages in the m-level subarray containing the first data page for prefetching, and if the (m - 1)-level subarray meets the page access volume condition, then continue to recursively select all the pages in the (m - 1)-level subarray containing the first data page for prefetching until the previous m-level subarray does not meet the page access volume condition or reaches the second-level subarray.
[0171] Among them, the value of m starts from k, and only when the value of m is 3, and the m-level subarray containing the first data page meets the page access volume condition, and the first second-level subarray meets the page access volume condition, will the pages of the first second-level subarray be selected for prefetching.
[0172] 5. The method according to claim 1, wherein when determining the page access volume condition, it includes the data pages that have been selected for prefetching.
[0173] 6. The method according to claim 1, wherein the size of the subarrays at each level is the same as the size of all the subarrays at the same level.
[0174] 7. The method according to claim 1, wherein each level of subarray has 2 x subarrays, and x is the dimension of the array.
[0175] 8. The method according to claim 4, wherein the division stop condition includes: the level of the subarray reaches a specific level, the size of the subarray is lower than a specific size, or the number of pages contained in the subarray is less than a specific number.
[0176] 9. The method according to claim 8, wherein the division stop condition includes:
[0177] The number of pages that have been retrieved or selected for prefetching reaches a threshold number, or,
[0178] The next lower-level subarrays that meet their respective page access volume conditions reach a threshold number.
[0179] 10. The method according to claim 9, wherein the page access volume condition requires:
[0180] When the space on the main storage unit of the accelerator that is not used by any core running on the accelerator is less than 5% of the capacity of the storage unit, all the next lower-level sub-arrays meet their respective page access volume conditions.
[0181] 11. The method according to claim 1, wherein the division stop condition, the page access volume condition, and the number of level-i sub-arrays divided into level-(i + 1) sub-arrays are based on the number of pages that make up the array.
[0182] 12. The method according to claim 1, wherein the storage system is the main storage unit of the host system, and the accelerator and the host system use a unified virtual storage system.
[0183] 13. The method according to claim 13, wherein the detecting an access attempt to the first data page includes: receiving a page fault triggered by a triggering core that attempts to access the first data page.
[0184] 14. The method according to claim 13, wherein the division stop condition, the page access volume condition, and the number of level-i sub-arrays divided into level-(i + 1) sub-arrays are based on the number of pages being used by the triggering core.
[0185] 15. The method according to claim 1, further comprising: evaluating the activity of the accelerator, wherein the division stop condition, the page access volume condition, and the number of level-i sub-arrays divided into level-(i + 1) sub-arrays are based on the activity evaluation of the accelerator.
[0186] 16. The method according to claim 15, wherein the activity of the accelerator includes:
[0187] The size of the main storage unit on the main storage unit of the accelerator that is used by the triggering core running on the accelerator, where the triggering core triggers an access to the first data page;
[0188] The size of the main storage unit on the main storage unit of the accelerator that is not used by any core other than the triggering core running on the accelerator;
[0189] The size of the main storage unit on the main storage unit of the accelerator that is not used by any core other than the triggering core running on the accelerator;
[0190] The size of the main storage unit on the main storage unit of the accelerator that is not used by any core other than the triggering core running on the accelerator;
[0191] The number of storage access patterns of the triggering core; or
[0192] The memory access pattern of any non-triggered core running on the accelerator.
[0193] 17. A system for retrieving data for an accelerator, the system comprising:
[0194] A main storage unit for storing a plurality of data pages;
[0195] A detection unit for detecting an access attempt to a first data page not stored on the main storage unit of the accelerator, the first data page corresponding to a part of an array having a plurality of dimensions;
[0196] An array partitioning unit for:
[0197] Partitioning the array into a plurality of first-level sub-arrays, and
[0198] Partitioning the first-level sub-arrays into a plurality of second-level sub-arrays, wherein the first-level sub-arrays contain the first data page;
[0199] A selection unit for selecting pages for prefetching, wherein the selection of pages for prefetching includes:
[0200] If a first second-level sub-array meets the page access volume condition, selecting all pages in the first second-level sub-array for prefetching, wherein the first second-level sub-array contains the first data page;
[0201] An acquisition unit for transferring the first data page and any data page selected for prefetching from a storage system connected to the accelerator to the main storage unit.
[0202] 18. The system according to claim 17, wherein the selection of pages for prefetching includes:
[0203] After selecting all pages in the first second-level sub-array for prefetching,
[0204] In response to the first second-level sub-array meeting the page access volume condition, if a first first-level sub-array meets the page access volume condition, selecting all pages in the first first-level sub-array for prefetching.
[0205] 19. The system according to claim 18, wherein the selection of pages for prefetching further includes: after selecting all pages in the first first-level sub-array for prefetching,
[0206] In response to the first first-level sub-array meeting the page access volume condition, if the array meets the page access volume condition, selecting all pages in the array for prefetching.
[0207] 20. The system according to claim 19, wherein the array partitioning unit is further configured to:
[0208] After partitioning the first-level sub-array into a plurality of second-level sub-arrays, partition the array into a plurality of sub-arrays by the following operations:
[0209] Partition the n-level sub-array of the smallest level containing the first data page into (n + 1)-level sub-arrays and continue to recursively partition the (n + 1)-level sub-arrays including the first data page until the partitioning stop condition is satisfied,
[0210] wherein, when the partitioning stop condition is satisfied, k is the value of n, the value of n starts from 2 and ends at k, and, before selecting all the pages in the first second-level sub-array for prefetching, the selection unit is further configured to:
[0211] If the m-level sub-array satisfies the page access volume condition, select all the pages in the m-level sub-array containing the first data page for prefetching, and, if the (m - 1)-level sub-array satisfies the page access volume condition, continue to recursively select all the pages in the (m - 1)-level sub-array containing the first data page for prefetching until the previous m-level sub-array does not satisfy the page access volume condition or reaches the second-level sub-array,
[0212] wherein, the value of m starts from k, and only when the value of m is 3, the m-level sub-array containing the first data page satisfies the page access volume condition, and the first second-level sub-array satisfies the page access volume condition, will the pages in the first second-level sub-array be selected for prefetching.
[0213] 21. The system according to claim 17, wherein each level of sub-array has 2 x sub-arrays, and x is the dimension of the array.
[0214] 22. The system according to claim 17, wherein the detection of the access attempt to the first data page includes: receiving a page fault triggered by a triggering core attempting to access the first data page.
[0215] 23. The system according to claim 17, further comprising: an evaluation unit configured to evaluate the activity of the accelerator, wherein the partitioning stop condition, the page access volume condition, and the number of the i-level sub-arrays partitioned into (i + 1)-level sub-arrays are evaluated based on the activity of the accelerator.
[0216] 24. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for retrieving data for an accelerator, the method comprising:
[0217] Detecting an attempt to access a first data page not stored on the main storage unit of the accelerator, the first data page corresponding to a portion of an array having multiple dimensions; and
[0218] In response to detecting the attempt to access the first data page:
[0219] Partitioning the array into sub-arrays by:
[0220] Partitioning the array into a plurality of first-level sub-arrays, and
[0221] Partitioning the first-level sub-arrays into a plurality of second-level sub-arrays, wherein the first-level sub-arrays contain the first data page;
[0222] Selecting pages for prefetching, wherein the selection of pages for prefetching includes:
[0223] If a first second-level sub-array meets a page access volume condition, selecting all pages in the first second-level sub-array for prefetching, wherein the first second-level sub-array contains the first data page; and
[0224] Transferring the first data page and any data pages selected for prefetching from a storage system connected to the accelerator to the main storage unit.
[0225] 25. The non-transitory computer-readable medium according to claim 24, wherein the selection of pages for prefetching further includes:
[0226] After selecting all pages in the first second-level sub-array for prefetching, in response to the first second-level sub-array meeting the page access volume condition, if a first first-level sub-array meets the page access volume condition, selecting all pages in the first first-level sub-array for prefetching.
[0227] 26. The non-transitory computer-readable medium according to claim 25, wherein the selection of pages for prefetching further includes:
[0228] After selecting all pages in the first first-level sub-array for prefetching, in response to the first first-level sub-array meeting the page access volume condition, if the array meets the page access volume condition, selecting all pages in the array for prefetching.
[0229] 27. The non-transitory computer-readable medium according to claim 26, wherein the instruction set is executable by at least one processor of a computer system to cause the computer system to further perform:
[0230] After dividing the first-level subarray into a plurality of second-level subarrays, divide the array into a plurality of subarrays by the following operations:
[0231] Divide the n-level subarray of the smallest level containing the first data page into (n + 1)-level subarrays and continue to recursively divide the (n + 1)-level subarrays including the first data page until a division stop condition is met,
[0232] wherein, when the division stop condition is met, k is the value of n, the value of n starts from 2 and ends at k, and, before selecting all pages in the first second-level subarray for prefetching, the selection of pages for prefetching is achieved by the following operations:
[0233] If the m-level subarray meets the page access volume condition, select all pages in the m-level subarray containing the first data page for prefetching, and if the (m - 1)-level subarray meets the page access volume condition, continue to recursively select all pages in the (m - 1)-level subarray containing the first data page for prefetching until the previous m-level subarray does not meet the page access volume condition or reaches the second-level subarray,
[0234] wherein, the value of m starts from k, and only when the value of m is 3, and the m-level subarray containing the first data page meets the page access volume condition, and the first second-level subarray meets the page access volume condition, will the pages in the first second-level subarray be selected for prefetching.
[0235] 28. The non-transitory computer-readable medium according to claim 24, wherein each level of subarray has 2 x subarrays, and x is the dimension of the array.
[0236] 29. The non-transitory computer-readable medium according to claim 24, wherein the detection of an access attempt to the first data page includes: receiving a page fault triggered by a triggering core attempting to access the first data page.
[0237] 30. The non-transitory computer-readable medium according to claim 24, wherein the instruction set is executable by at least one processor of a computer system to cause the computer system to further perform:
[0238] Evaluating the activities of the accelerator, wherein the division stop condition, the page access volume condition, and the number of the i-th level sub-arrays divided into the (i + 1)-th level sub-arrays are evaluated based on the activities of the accelerator.
[0239] In the foregoing specification, embodiments have been described with reference to numerous specific details, which may vary with the implementation. Certain adaptations and modifications can be made to the described embodiments. By considering the present disclosure and practice, other embodiments will be apparent to those skilled in the art. The present specification and embodiments are considered to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims. The step sequences shown in the figures are for illustrative purposes only and are not limited to any particular step sequence. Thus, those skilled in the art can understand that these steps can be executed in a different order while implementing the same method.
[0240] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications can be made to these embodiments. Thus, although specific terms are used, they are for general and descriptive purposes only and not for the purpose of limitation.
[0241] Considering the specification and practice of the embodiments disclosed herein, other embodiments will be apparent, such that the specification and examples are considered to be examples only, and the true scope and spirit of the disclosed embodiments are indicated by the following claims.
Claims
1. A method for retrieving data for an accelerator, the method comprising: Detecting an access attempt to a first data page not stored on a main storage unit of the accelerator, the first data page corresponding to a portion of an array having multiple dimensions; And In response to detecting the access attempt to the first data page: Dividing the array into sub-arrays until a division stop condition is met by: Dividing the array into a plurality of first-level sub-arrays, and Dividing the first-level sub-arrays into a plurality of second-level sub-arrays, wherein the first-level sub-arrays contain the first data page; Selecting pages for prefetching, wherein the selecting of pages for prefetching includes: If a first second-level sub-array meets a page access volume condition, selecting all pages in the first second-level sub-array for prefetching, wherein the first second-level sub-array contains the first data page; and Transferring the first data page and any data pages selected for prefetching from a storage system connected to the accelerator to the main storage unit; Among them, each level of sub-arrays has sub-arrays, and x is the dimension of the said array.
2. The method according to claim 1, wherein, After selecting all pages in the first second-level sub-array for prefetching, the selecting of pages for prefetching further includes: In response to the first second-level sub-array meeting the page access volume condition, if a first first-level sub-array meets the page access volume condition, selecting all pages in the first first-level sub-array for prefetching.
3. The method according to claim 2, wherein After selecting all pages in the first first-level sub-array for prefetching, the selecting of pages for prefetching further includes: in response to the first first-level sub-array meeting the page access volume condition, if the array meets the page access volume condition, selecting all pages in the array for prefetching.
4. The method according to claim 3, further comprising: After dividing the first-level sub-arrays into a plurality of second-level sub-arrays, dividing the array into a plurality of sub-arrays by: Dividing the smallest-level nth-level sub-array containing the first data page into (n + 1)th-level sub-arrays and continuing to recursively divide the (n + 1)th-level sub-arrays including the first data page until the division stop condition is met, wherein when the division stop condition is met, k is the value of n, the value of n starts from 2 and ends at k, and before selecting all pages in the first second-level sub-array for prefetching, the selection of pages for prefetching is achieved by: If an mth-level sub-array meets the page access volume condition, selecting all pages in the mth-level sub-array containing the first data page for prefetching, and if the (m - 1)th-level sub-array meets the page access volume condition, continuing to recursively select all pages in the (m - 1)th-level sub-array containing the first data page for prefetching until the previous mth-level sub-array does not meet the page access volume condition or reaches the second-level sub-array, wherein the value of m starts from k, and only when the value of m is 3, the mth-level sub-array containing the first data page meets the page access volume condition, and the first second-level sub-array meets the page access volume condition, will the pages of the first second-level sub-array be selected for prefetching.
5. The method according to claim 1, wherein, When determining the page access volume condition, it includes data pages that have been selected for prefetching.
6. The method according to claim 1, wherein The size of the sub-arrays at each level is the same as the size of all sub-arrays at the same level.
7. The method according to claim 4, wherein The division stop condition includes: the level of the sub-array reaches a specific level, the size of the sub-array is lower than a specific size, or the number of pages contained in the sub-array is less than a specific quantity.
8. The method according to claim 1, wherein The division stop condition includes: The number of pages that have been retrieved or selected for prefetching reaches a threshold number, or, The number of sub-arrays at the next lower level that meet their respective page access volume conditions reaches a threshold number.
9. The method according to claim 8, wherein The page access volume condition requires that: When the space on the main storage unit of the accelerator that is not used by any core running on the accelerator is less than 5% of the capacity of the storage unit, all sub-arrays at the next lower level meet their respective page access volume conditions.
10. The method according to claim 1, wherein, The division stop condition, the page access volume condition, and the number of sub-arrays at the i-th level divided into sub-arrays at the (i + 1)-th level are based on the number of pages that make up the array.
11. The method according to claim 1, wherein, The storage system is the main storage unit of the host system, and the accelerator and the host system use a unified virtual storage system.
12. The method according to claim 1, wherein The detection of the access attempt to the first data page includes: receiving a page fault triggered by a triggering core that attempts to access the first data page.
13. The method according to claim 12, wherein, The division stop condition, the page access volume condition, and the number of sub-arrays at the i-th level divided into sub-arrays at the (i + 1)-th level are based on the number of pages being used by the triggering core.
14. The method according to claim 1 further comprises: Evaluate the activity of the accelerator, where the division stop condition, the page access volume condition, and the number of sub-arrays at the i-th level divided into sub-arrays at the (i + 1)-th level are based on the activity evaluation of the accelerator.
15. The method according to claim 14, wherein, The activity of the accelerator includes: The size of the main storage unit on the main storage unit of the accelerator used by the triggering core running on the accelerator, and the triggering core triggers the access to the first data page; The size of the main storage unit on the main storage unit of the accelerator that is not used by any non-triggering core running on the accelerator; The size of the main storage unit on the main storage unit of the accelerator that is not used by any non-triggering core running on the accelerator; The size of the main storage unit on the main storage unit of the accelerator that is not used by any non-triggering core running on the accelerator; The number of storage access patterns of the triggering core; or The memory access pattern of any non-triggering core running on the accelerator.
16. A system for retrieving data for an accelerator, the system includes: A main storage unit for storing a plurality of data pages; A detection unit for detecting an access attempt to a first data page not stored on the main storage unit of the accelerator, the first data page corresponding to a part of an array having multiple dimensions; An array division unit for dividing the array into sub-arrays until a division stop condition is met: Dividing the array into a plurality of first-level sub-arrays, and Dividing the first-level sub-arrays into a plurality of second-level sub-arrays, where the first-level sub-arrays contain the first data page; A selection unit for selecting pages for prefetching, where the selection of pages for prefetching includes: If the first and second level sub-arrays meet the page access volume condition, then all pages in the first and second level sub-arrays are selected for prefetching, where the first and second level sub-arrays contain the first data page; An acquisition unit for transferring the first data page and any data page selected for prefetching from a storage system connected to the accelerator to the main storage unit; Among them, each level of sub-arrays has sub-arrays, and x is the dimension of the said array.
17. The system according to claim 16, wherein, The selection of pages for prefetching includes: After selecting all pages in the first and second level sub-arrays for prefetching, In response to the first and second level sub-arrays meeting the page access volume condition, if the first and first level sub-arrays meet the page access volume condition, then all pages in the first and first level sub-arrays are selected for prefetching.
18. The system according to claim 17, wherein The selection of pages for prefetching further includes: after selecting all pages in the first and first level sub-arrays for prefetching, In response to the first and first level sub-arrays meeting the page access volume condition, if the array meets the page access volume condition, then all pages in the array are selected for prefetching.
19. The system according to claim 18, wherein The array partitioning unit is further configured to: After partitioning the first level sub-array into multiple second level sub-arrays, partition the array into multiple sub-arrays by the following operations: Partition the nth level sub-array with the smallest level containing the first data page into the (n + 1)th level sub-array and continue to recursively partition the (n + 1)th level sub-array containing the first data page until the partitioning stop condition is met, where, when the partitioning stop condition is met, k is the value of n, the value of n starts from 2 and ends at k, and, before selecting all pages in the first and second level sub-arrays for prefetching, the selection unit is further configured to: If the mth level sub-array meets the page access volume condition, then all pages in the mth level sub-array containing the first data page are selected for prefetching, and, if the (m - 1)th level sub-array meets the page access volume condition, then continue to recursively select all pages in the (m - 1)th level sub-array containing the first data page for prefetching until the previous mth level sub-array does not meet the page access volume condition or reaches the second level sub-array, where, the value of m starts from k, and only when the value of m is 3, the mth level sub-array containing the first data page meets the page access volume condition, and the first and second level sub-arrays meet the page access volume condition, will the pages of the first and second level sub-arrays be selected for prefetching.
20. The system according to claim 16, wherein The detection of an access attempt to the first data page includes: receiving a page fault triggered by a triggering core attempting to access the first data page.
21. The system according to claim 16, further comprising: An evaluation unit for evaluating the activity of the accelerator, where the partitioning stop condition, the page access volume condition, and the number of the ith level sub-arrays partitioned into the (i + 1)th level sub-arrays are evaluated based on the activity of the accelerator.
22. A non-transitory computer-readable medium that stores a set of instruction sets, the instruction sets being executable by at least one processor of a computer system to cause the computer system to execute a method for retrieving data for an accelerator, the method including: Detect an attempt to access a first data page that is not stored on the main storage unit of the accelerator, where the first data page corresponds to a part of an array having multiple dimensions; and In response to detecting an attempt to access the first data page: Divide the array into sub-arrays by the following operations until a division stop condition is met: Divide the array into a plurality of first-level sub-arrays, and Divide the first-level sub-arrays into a plurality of second-level sub-arrays, where the first-level sub-arrays contain the first data page; Select pages for prefetching, where the selection of pages for prefetching includes: If a first second-level sub-array meets the page access volume condition, select all pages in the first second-level sub-array for prefetching, where the first second-level sub-array contains the first data page; and Transfer the first data page and any data pages selected for prefetching from the storage system connected to the accelerator to the main storage unit; wherein, each level of sub-arrays has sub-arrays, and x is the dimension of the array.
23. The non-transitory computer-readable medium according to claim 22, wherein, The selection of pages for prefetching further includes: After selecting all pages in the first second-level sub-array for prefetching, in response to the first second-level sub-array meeting the page access volume condition, if a first first-level sub-array meets the page access volume condition, select all pages in the first first-level sub-array for prefetching.
24. The non-transitory computer-readable medium according to claim 23, wherein, The selection of pages for prefetching further includes: After selecting all pages in the first first-level sub-array for prefetching, in response to the first first-level sub-array meeting the page access volume condition, if the array meets the page access volume condition, select all pages in the array for prefetching.
25. The non-transitory computer-readable medium according to claim 24, wherein, The instruction set can be executed by at least one processor of a computer system to cause the computer system to further perform: After dividing the first-level sub-arrays into a plurality of second-level sub-arrays, divide the array into a plurality of sub-arrays by the following operations: Divide the smallest-level nth-level sub-array containing the first data page into (n + 1)th-level sub-arrays and continue to recursively divide the (n + 1)th-level sub-arrays including the first data page until the division stop condition is met, where, when the division stop condition is met, k is the value of n, the value of n starts from 2 and ends at k, and, before selecting all pages in the first second-level sub-array for prefetching, the selection of pages for prefetching is achieved by the following operations: If an mth-level sub-array meets the page access volume condition, select all pages in the mth-level sub-array containing the first data page for prefetching, and if the (m - 1)th-level sub-array meets the page access volume condition, continue to recursively select all pages in the (m - 1)th-level sub-array containing the first data page for prefetching until the previous mth-level sub-array does not meet the page access volume condition or reaches the second-level sub-array, where, the value of m starts from k, and only when the value of m is 3, and the mth-level sub-array containing the first data page meets the page access volume condition, and the first second-level sub-array meets the page access volume condition, will the pages in the first second-level sub-array be selected for prefetching.
26. The non-transitory computer-readable medium according to claim 22, wherein, The detection of an access attempt to the first data page includes: receiving a page fault triggered by a triggering core attempting to access the first data page.
27. The non-transitory computer-readable medium according to claim 22, wherein, The instruction set can be executed by at least one processor of a computer system to cause the computer system to further perform: Evaluating the activity of the accelerator, wherein the partitioning stop condition, the page access volume condition, and the number of level-i sub-arrays partitioned into level-(i + 1) sub-arrays are evaluated based on the activity of the accelerator.
Citation Information
Patent Citations
Method for accelerating convolution neutral network hardware and AXI bus IP core thereof
CN104915322A
Conditional prefetching
US20140229682A1
Supporting large pages in hardware prefetchers
US20150278099A1