A Method and Device for Accelerating Hardware Page Table Traversal
By adding a computing unit to the processor cache, directly calculating the next-level base address, the problem of low efficiency of hardware page table traversal is solved and faster address conversion is achieved.
Patent Information
- Application Number
- CN201911195523.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2039-11-28
AI Technical Summary
In the prior art, the hardware page table traversal efficiency is low, resulting in slow conversion speed of virtual address to physical address, and the existing acceleration scheme increases processor cost or excessive overhead.
Add a calculation unit to each level of the processor cache, and directly calculate the next level base address through the base address and offset address, avoiding returning to the arithmetic logic unit in the processor core for calculations, and improving the base address access efficiency.
The hardware page table traversal process is accelerated, the conversion efficiency of virtual address to physical address is improved, and the number of data transmissions in the processor core is reduced.
Smart Images

Figure CN112860600B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer storage, and particularly to a method and apparatus for accelerating hardware page table traversal. Background Art
[0002] In a processor architecture, the Memory Management Unit (MMU) in a processor core is responsible for translating the virtual address (VA) used by an application program into a physical address (PA). In the case of applying a paging management mechanism, the conversion from VA to PA requires querying a page table. Since part of the page table is stored in the Translation Lookaside Buffer (TLB) in the MMU, generally, the TLB can assist the MMU in completing the virtual address conversion (for example, if the required page table base addresses are exactly stored in the TLB (i.e., TLB hit), then the conversion from VA to PA can be completed); otherwise, a hardware page table walk (HPTW) needs to be performed on the stored page table to obtain the final PA.
[0003] Currently, in most architectures of all commercial processors, after a TLB miss, a hardware-automatic page table traversal process will be triggered, and then the TLB will be refilled. However, currently, HPTW is completed serially, and the number of memory accesses is directly related to the number of page table levels (for example, if the page table is divided into four levels, at least four data accesses are required to complete the conversion from VA to PA). Specifically, after a certain-level page table base address hits a certain-level cache, the obtained page table base address is returned to the Load Store Unit (LSU) in the core, and the next-level page table base address is calculated through the Arithmetic and Logic Unit (ALU). During the above page table query process, the data obtained each time needs to be returned to the LSU in the processor core for further processing, resulting in low page table traversal efficiency.
[0004] In the prior art, generally the following two acceleration schemes are provided to accelerate HPTW.
[0005] Solution 1: Add an extra cache in the MMU to cache the intermediate-level page table base addresses required during the HPTW process, such as the third-level to fourth-level page table base addresses. Query this cache after a TLB miss; if a hit occurs in this cache, directly jump to the corresponding level of the page table (for example, query this cache after a TLB miss; if a hit on the third-level page table is found in this cache, retrieve the base address of this page table; then continue with the HPTW to query the remaining levels of the page table). However, adding an extra cache in the processor core will increase the production cost; in addition, due to the small capacity of the extra cache, it cannot guarantee hits for all queries, and it is difficult to effectively improve the memory access efficiency.
[0006] Solution 2: Add a prefetch engine in each level of the cache to prefetch the data in the linked list structure in each level of the cache, so there is no need to retrieve the memory access address back to the LSU in the core for processing. However, a TLB structure (the TLB structure is used to translate the VA to the PA) needs to be added in each level of the cache, otherwise data prefetching cannot be performed. Therefore, adopting Solution 2 will result in too much overhead.
[0007] Therefore, how to improve the process of hardware page table traversal, accelerate the page table traversal and improve the address translation efficiency is an urgent problem to be solved. Summary of the Invention
[0008] The embodiments of the present invention provide a method and device for accelerating hardware page table traversal, which can improve the process of hardware page table traversal, accelerate the page table traversal and improve the address translation efficiency.
[0009] In a first aspect, the embodiments of the present invention provide a device for accelerating page table traversal, including a control unit of the i-th level of the first cache and a computing unit of the i-th level of the first cache; i = 1, 2,..., N, and N is a positive integer; wherein,
[0010] The control unit of the i-th level of the first cache is configured to:
[0011] Receive a memory access request, where the memory access request includes a K-th level base address and a K-th level high-order address, and K is a positive integer;
[0012] Judge whether the i-th level of the first cache stores the K-th level base address according to the K-th level base address;
[0013] In the case where it is judged that the i-th level of the first cache stores the K-th level base address, send the K-th level base address to the computing unit of the i-th level of the first cache.
[0014] The computing unit of the i-th level of the first cache is configured to:
[0015] Determine the base address of the (K + 1)th level according to the base address of the Kth level and the offset address of the Kth level, where the offset address of the Kth level is determined according to the high-order address of the Kth level.
[0016] In the embodiment of the present invention, by adding a computing unit to each level of the first cache, after each level of the first cache obtains the base address, it can directly calculate the base address of the next level according to the base address and the corresponding offset. Specifically, the control unit of the ith level of the first cache receives a memory access request. On the premise that it is determined that the ith level of the first cache (i.e., the current level of cache) stores the base address of the Kth level, the control unit sends the base address of the Kth level in the memory access request to the computing unit of the current level of cache. The current level of cache calculates the base address of the (K + 1)th level (i.e., the base address of the next level) according to the base address of the Kth level and the offset address determined from the high-order address included in the memory access request. Different from the prior art in which any level of cache, after determining the base address of the Kth level, needs to return the base address of this level to the memory access unit in the processor core, and then calculate the base address of the next level through the arithmetic logic unit in the processor core and continue to query the base address of the next level; in the embodiment of the present invention, no matter at which level of the first cache the base address of the Kth level is determined, the base address of the next level can be calculated by shifting through the computing unit added to the current level of cache, improving the access efficiency of the base address, accelerating the process of traversing the hardware page table, and ultimately improving the conversion efficiency from the virtual address to the physical address.
[0017] In a possible implementation manner, the control unit of the ith level of the first cache is further configured to: in the case where it is determined that the ith level of the first cache does not store the base address of the Kth level and i ≠ N, send the memory access request to the control unit of the (i + 1)th level of the first cache. In the embodiment of the present invention, in the case where the ith level of the first cache determines that it does not store the base address of the Kth level and there is a next level of cache, a memory access request is sent to the next level of cache (i.e., the (i + 1)th level of the first cache), supplementing another result of determining whether the ith level of the first cache stores the base address of the Kth level, making the embodiment of the present invention more complete to avoid execution errors or system failures, etc.
[0018] In a possible implementation, the device further includes a home node; the home node is coupled to the i-th level first cache; the control unit of the i-th level first cache is further configured to: when it is determined that the i-th level first cache does not store the K-th level base address, send the memory access request to the home node. The home node is used to receive the memory access request. In the embodiments of the present invention, after adding the description of the home node, when it is determined that the i-th level first cache does not store the K-th level base address, the memory access request can be sent not only to the (i + 1)-th level first cache, but also to the home node at the same time. The processing situation of the N-th level first cache is further improved. Sending the memory access request to the home node is beneficial to judging the storage situation of the base address in all caches.
[0019] In a possible implementation, the home node is further configured to: after receiving the memory access request, determine whether each level of the N-level first cache stores the K-th level base address according to the K-th level base address; when it is determined that the target first cache stores the K-th level base address, send the memory access request to the control unit of the target first cache. In the embodiments of the present invention, after the home node receives the memory access request, it judges whether each level of the first cache has a cache storing the required K-th level base address. After it is determined that the target first cache (a certain level or several levels of caches) stores the K-th level base address, the memory access request is sent to this level of cache. It is possible to avoid the situation where the memory is accessed even though a certain cache stores the base address. For example, the third-level cache determines that it does not have the base address of the third-level page table itself, while the second-level cache stores the base address of the third-level page table. Since the third-level cache will not return to the second-level cache for access, the home node can be responsible for returning the request to the second-level cache.
[0020] In a possible implementation, the device further includes: a memory controller coupled to the home node, and a memory coupled to the memory controller; the home node includes a buffer; the home node is further configured to: when it is determined that each level of the first cache does not store the K-th level base address, send the memory access request to the memory controller. The memory controller is configured to: determine the K-th level base address in the memory according to the K-th level base address; send the K-th level base address to the home node. The memory is used to store the K-th level base address. The buffer is used to determine the (K + 1)-th level base address according to the K-th level base address and the K-th level offset address. In the embodiments of the present invention, when the home node finds that all caches do not store the required base address, it will send the memory access request to the memory controller, instructing the memory controller in the processor chip to obtain the required base address from the memory. The possible situations in the foregoing embodiments are supplemented, avoiding execution errors, and ensuring that the page table traversal process can continue even when all caches do not store the required base address.
[0021] In a possible implementation, the device further includes a second cache coupled to the home node, where the second cache includes a control unit of the second cache and a computing unit of the second cache; the home node is further configured to: after receiving the memory access request, determine whether the second cache stores the Kth-level base address; and in the case where it is determined that the second cache stores the Kth-level base address, send the memory access request to the control unit of the second cache. In the embodiments of the present invention, the case where the processor chip has multiple processor cores (including the second cache) is supplemented. When the cache of one processor core does not store the required base address, the home node can check whether the required base address is in the caches of other cores to continue the page table query.
[0022] In a possible implementation, the device further includes a memory controller coupled to the home node and a memory coupled to the memory controller; the home node includes a buffer buffer; the home node is further configured to: in the case where it is determined that none of the first-level caches and the second cache stores the Kth-level base address, send the memory access request to the memory controller. The memory controller is configured to: determine the Kth-level base address in the memory according to the Kth-level base address; and send the Kth-level base address to the home node. The memory is used to store the Kth-level base address. The buffer is configured to determine the (K + 1)th-level base address according to the Kth-level base address and the Kth-level offset address. In the embodiments of the present invention, in the case where the processor chip has multiple processor cores, when it is detected by the home node that none of the caches store the required base address, a corresponding instruction is sent to the memory controller, so as to obtain the required base address from the memory in a timely manner to ensure that the page table base address can still be queried when the cache cannot be hit.
[0023] In a possible implementation, the device further includes a memory management unit coupled to the first cache of the first level. The memory management unit includes a third cache; the third cache is used to store the base address of the Kth level, where K = 1, 2, …, M, and M is an integer greater than 1; the memory management unit is configured to: before sending a memory access request to the first cache of the first level, determine whether the third cache stores the base address of the first level; in the case where it is determined that the third cache stores the base address of the first level, obtain the base address of the first level; in the case where it is determined that the third cache does not store the base address of the first level, send the memory access request to the first cache of the first level. In the embodiment of the present invention, a third cache (i.e., a newly added cache) is newly added to the memory management unit, and the newly added cache can store partial page table base addresses. When a TLB miss occurs, it can be checked in the newly added cache whether there are all the page table base addresses required for the VA-to-PA conversion. If the newly added cache is hit, all the page table base addresses can be directly obtained, so that there is no need to perform a page table traversal and the required physical address can be quickly obtained.
[0024] In a possible implementation, the memory access request further includes a base address identifier of the Kth level base address, and the base address identifier of the Kth level base address is used to indicate the level of the Kth level base address. In the embodiment of the present invention, the form of the memory access request is supplemented. A specific domain or data segment can be added to the memory access request for the cache control unit or the home node to identify that the memory access request is a query and access to the base address of which level.
[0025] In a possible implementation, the control unit of the first cache of the ith level is further configured to: in the case of determining the base address of the (M + 1)th level, send the base address of the (M + 1)th level to the memory management unit. The memory management unit is further configured to receive the base address of the (M + 1)th level. In the embodiment of the present invention, the possible processing situation of the cache after determining the last level base address (i.e., the base address of the (M + 1)th level) is supplemented. For example, if a certain level of cache determines the last level base address, the base address can be sent to the memory management unit. Further, after the memory management unit receives the last page base address (i.e., the last level base address), the physical address is obtained by adding the page offset.
[0026] Second aspect, an embodiment of the present invention provides a method for accelerating page table traversal, including: receiving a memory access request through a control unit of a first cache at the i-th level, where the memory access request includes a base address at the K-th level and a high-order address at the K-th level, K is a positive integer, and i = 1, 2,..., N; through the control unit of the first cache at the i-th level, determining whether the first cache at the i-th level stores the base address at the K-th level according to the base address at the K-th level; in the case where it is determined that the first cache at the i-th level stores the base address at the K-th level, sending the base address at the K-th level to a computing unit of the first cache at the i-th level through the control unit of the first cache at the i-th level; determining a base address at the (K + 1)-th level according to the base address at the K-th level and an offset address at the K-th level through a computing unit of the first cache at the i-th level, where the offset address at the K-th level is determined according to the high-order address at the K-th level.
[0027] In a possible implementation manner, the method further includes: in the case where it is determined that the first cache at the i-th level does not store the base address at the K-th level and i ≠ N, sending the memory access request to a control unit of a first cache at the (i + 1)-th level through the control unit of the first cache at the i-th level.
[0028] In a possible implementation manner, the method further includes: in the case where it is determined that the first cache at the i-th level does not store the base address at the K-th level, sending the memory access request to a home node through the control unit of the first cache at the i-th level.
[0029] In a possible implementation manner, the method further includes: after receiving the memory access request, determining, through the home node, whether each level of the first cache at the N levels stores the base address at the K-th level according to the base address at the K-th level; in the case where it is determined that a target first cache stores the base address at the K-th level, sending the memory access request to a control unit of the target first cache.
[0030] In a possible implementation manner, the method further includes: in the case where it is determined that none of the first caches at each level stores the base address at the K-th level, sending the memory access request to a memory controller through the home node; determining the base address at the K-th level in the memory according to the base address at the K-th level through the memory controller; sending the base address at the K-th level to the home node through the memory controller; determining the base address at the (K + 1)-th level through a buffer (buffer) of the home node according to the base address at the K-th level and the offset address at the K-th level.
[0031] In a possible implementation manner, the method further includes: after receiving the memory access request, determining, through the home node, whether a second cache stores the base address at the K-th level; in the case where it is determined that the second cache stores the base address at the K-th level, sending the memory access request to a control unit of the second cache through the home node.
[0032] In a possible implementation, the method further includes: when it is determined that neither the first cache of each level nor the second cache stores the Kth-level base address, sending the memory access request to the memory controller through the home node;
[0033] Determining the Kth-level base address in the memory through the memory controller according to the Kth-level base address; sending the Kth-level base address to the home node through the memory controller; determining the (K + 1)th-level base address through the buffer (buffer) of the home node according to the Kth-level base address and the Kth-level offset address.
[0034] In a possible implementation, the method further includes: before sending a memory access request to the first cache of the first level, determining whether the third cache stores the first-level base address through the memory management unit; when it is determined that the third cache stores the first-level base address, obtaining the first-level base address through the memory management unit; when it is determined that the third cache does not store the first-level base address, sending the memory access request to the first cache of the first level through the memory management unit.
[0035] In a possible implementation, the memory access request further includes a base address identifier of the Kth-level base address, and the base address identifier of the Kth-level base address is used to indicate the level of the Kth-level base address.
[0036] In a possible implementation, after the home node receives the memory access request, it may first determine whether the first cache stores the Kth-level base address, and then determine whether the second cache stores the Kth-level base address. Optionally, after the home node receives the memory access request, it simultaneously determines whether the first cache and the second cache store the Kth-level base address. Further optionally, if any cache stores the required Kth-level base address, the home node sends a corresponding memory access request to it; the home node can accept the results fed back by multiple caches, and use the highest-level base address as the final traversal result among the numerous results, and then continue the traversal process. Or, in the process where the traversal has been completed in multiple caches, the home node can directly feedback the final base address to the memory management unit through the cache that has completed the last step of the traversal (i.e., the cache that obtained the last base address).
[0037] In a third aspect, an embodiment of the present invention provides an apparatus for accelerating page table traversal, including: N first caches of the first level; each of the first caches of the N first levels includes a calculation unit and a control unit; the N first caches store one or more of the M-level base addresses; N is an integer greater than 0, and M is an integer greater than 1; wherein,
[0038] The control unit of the first cache at the i-th level is configured to: receive a memory access request at the K-th level, where the memory access request at the K-th level includes a K-th level base address and a K-th level high-order address; 0 < K ≤ M, K is an integer; i = 1, 2,..., N; determine whether the first cache at the i-th level stores the K-th level base address according to the K-th level base address; and in the case where it is determined that the first cache at the i-th level stores the K-th level base address, send the K-th level base address to the computing unit of the first cache at the i-th level.
[0039] The computing unit of the first cache at the i-th level is configured to: determine a (K + 1)-th level base address according to the K-th level base address and the K-th level offset address; the K-th level offset address is determined according to the K-th level high-order address.
[0040] In a possible implementation, the device further includes a home node; the home node is coupled to the first cache at the N-th level; the control unit of the first cache at the N-th level is further configured to: in the case where it is determined that the first cache at the N-th level does not store the K-th level base address, send the memory access request at the K-th level to the home node. The home node is configured to receive the memory access request at the K-th level.
[0041] In a possible implementation, the home node is further configured to: after receiving the memory access request at the K-th level, determine whether each level of the first cache stores the K-th level base address according to the K-th level base address; and in the case where it is determined that the first cache at the p-th level stores the K-th level base address, send the memory access request at the K-th level to the control unit of the first cache at the p-th level, where p = 1, 2,..., N and p ≠ i.
[0042] In a possible implementation, the method further includes: before sending a memory access request at the first level to the first cache at the first level, determining whether a third cache stores the first level base address through a memory management unit; in the case where it is determined that the third cache stores the first level base address, obtaining the first level base address through the memory management unit; and in the case where it is determined that the third cache does not store the first level base address, sending the memory access request at the first level to the first cache at the first level through the memory management unit.
[0043] In a fourth aspect, an embodiment of the present invention provides a terminal, which includes a processor configured to support the terminal in executing corresponding functions in a method for accelerating hardware page table traversal provided in the second aspect. The terminal may further include a memory for coupling with the processor and storing necessary program instructions and data of the terminal. The terminal may further include a communication interface for the terminal to communicate with other devices or communication networks.
[0044] Fifth aspect, an embodiment of the present invention provides a chip system, which may include: an accelerated page table traversal device as described in the first aspect above, and an auxiliary circuit coupled to the accelerated page table traversal device.
[0045] Sixth aspect, an embodiment of the present invention provides an electronic device, which may include: an accelerated page table traversal device as described in the first aspect above, and discrete devices coupled to the outside of the accelerated page table traversal device.
[0046] Seventh aspect, an embodiment of the present invention provides a chip system, which can execute any method involved in the second aspect above, so that related functions can be realized. In a possible design, the chip system further includes a memory for storing necessary program instructions and data. The chip system can be composed of chips or can include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments.
[0048] Figure 1 It is a schematic diagram of a hardware page table traversal process provided by an embodiment of the present invention;
[0049] Figure 2 It is a schematic diagram of a system architecture provided by an embodiment of the present invention;
[0050] Figure 3 It is a schematic diagram of an application architecture provided by an embodiment of the present invention;
[0051] Figure 4 It is another schematic diagram of an application architecture provided by an embodiment of the present invention;
[0052] Figure 5 It is a schematic diagram of an accelerated hardware page table traversal device provided by an embodiment of the present invention;
[0053] Figure 6 It is a schematic diagram of a specific accelerated hardware page table traversal device provided by an embodiment of the present invention;
[0054] Figure 7 It is a schematic diagram of the internal structure of a cache provided by an embodiment of the present invention;
[0055] Figure 8 It is a schematic diagram of a page table query process provided by an embodiment of the present invention;
[0056] Figure 9 It is a corresponding Figure 8 schematic diagram of the interaction of some hardware provided by an embodiment of the present invention;
[0057] Figure 10 It is a schematic diagram of the command format of a memory access request provided by an embodiment of the present invention;
[0058] Figure 11 It is another schematic diagram of the page table query process provided by an embodiment of the present invention;
[0059] Figure 12 It is a kind of corresponding Figure 11 Schematic diagram of the interaction of part of the hardware;
[0060] Figure 13 It is yet another schematic diagram of the page table query process provided by an embodiment of the present invention;
[0061] Figure 14 It is a kind of corresponding Figure 13 Schematic diagram of the interaction of part of the hardware;
[0062] Figure 15 It is still another schematic diagram of the page table query process provided by an embodiment of the present invention;
[0063] Figure 16 It is a kind of corresponding Figure 15 Schematic diagram of the interaction of part of the hardware;
[0064] Figure 17 It is a schematic diagram of an accelerated hardware page table traversal device provided by an embodiment of the present invention;
[0065] Figure 18 It is a schematic diagram of another specific accelerated hardware page table traversal device provided by an embodiment of the present invention;
[0066] Figure 19 It is a schematic diagram of the page table query process in a multi-core scenario provided by an embodiment of the present invention;
[0067] Figure 20 It is a kind of corresponding Figure 19 Schematic diagram of the interaction of part of the hardware;
[0068] Figure 21 It is a schematic diagram of an accelerated hardware page table traversal device in a multi-core scenario provided by an embodiment of the present invention;
[0069] Figure 22 It is a schematic diagram of an accelerated hardware page table traversal method provided by an embodiment of the present invention;
[0070] Figure 23 It is a schematic diagram of another accelerated hardware page table traversal method provided by an embodiment of the present invention;
[0071] Figure 24It is a schematic structural diagram of a chip provided by an embodiment of the present invention. Detailed implementation manners
[0072] Next, the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0073] The terms "first", "second", "third", and "fourth" in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order; and the objects described by the terms "first", "second", "third", and "fourth" etc. may also be the same object, or there may be an inclusion or other relationship between each other. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0074] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0075] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a computing device and the computing device can both be components. One or more components can reside in a process and / or an execution thread, and the components can be located on one computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media storing various data structures. Components can communicate, for example, through local and / or remote processes according to signals having one or more data packets (such as data from two components interacting with another component between a local system, a distributed system, and / or a network, such as interacting with other systems through signals on the Internet).
[0076] First, some terms in this application are explained to facilitate understanding by those skilled in the art.
[0077] (1) The physical address, also known as the real address or binary address, is the memory address that exists in electronic form on the address bus and enables the data bus to access a specific storage unit in the main memory. The addresses are numbered starting from 0 and incremented by 1 sequentially. Therefore, the physical address space of the memory grows linearly.
[0078] (2) In the paging management mechanism (i.e., paging mechanism), the virtual address is transformed to generate the physical address.
[0079] (3) Memory is used to temporarily store the operation data in the processor and the data exchanged with external memories such as hard disks. During the operation of the computer, the processor will transfer the data to be operated to the memory for operation and then transfer the result out after the operation is completed. Dynamic Random Access Memory (DRAM) has high cost performance and good scalability and is the main part of general memory.
[0080] (4) The Central Processing Unit (CPU) is used to interpret computer instructions and process the data in the computer. It is responsible for reading instructions, decoding and executing instructions in the computer. The central processing unit mainly includes two parts, namely the controller and the arithmetic unit, and also includes a cache memory and the data and control buses that realize the connection between them.
[0081] (5) The Arithmetic & Logical Unit (ALU) is the core component of the processor and mainly performs various arithmetic and logical operation operations, such as the four arithmetic operations of addition, subtraction, multiplication and division, logical operations such as AND, OR, NOT, XOR, and operations such as shift, comparison and transfer.
[0082] (6) The Memory Management Unit (MMU), sometimes called the paged memory management unit (PMMU), or simply the memory management unit, is a computer hardware that is responsible for handling the memory access requests of the Central Processing Unit (CPU). Its functions include the conversion of virtual addresses to physical addresses (i.e., virtual memory management), memory protection, and control of the central processor cache; in a relatively simple computer architecture, it is responsible for bus arbitration and memory bank switching.
[0083] (7) The cache, short for the high-speed buffer memory, is located between the CPU and the main memory DRAM. It is a memory with a small capacity but high speed, usually composed of static random access memory (SRAM). As long as the SRAM is powered on, the data stored in it can be constantly maintained. SRAM can generally be divided into the following five major parts: memory cell array, row / column address decoder, sense amplifier, control circuit, and drive circuit. Specifically, the cache can save a part of the data that the CPU has just used or is cycling through. If the CPU needs to use this part of the data again, it can directly call it from the cache, which speeds up the data access speed and reduces the CPU's waiting time. The cache is generally divided into level 1 cache (L1 cache), level 2 cache (L2 cache), level 3 cache (L3 cache), and so on. Among them, the L1 cache is mainly integrated inside the CPU, the L2 cache is integrated on the motherboard or inside the CPU, and the L3 cache is integrated on the motherboard or inside the CPU. When it is inside the CPU, the L3 cache is shared by multiple processor cores.
[0084] (8) The buffer is a reserved storage space with a certain capacity used to buffer input or output data. According to whether it corresponds to an input device or an output device, the buffer is divided into an input buffer and an output buffer.
[0085] (9) The core is the heart of the processor, used to complete all calculations, accept / store commands, process data, etc. The cores of various processors have a fixed logical structure, involving the layout of logical units such as the first-level cache, second-level cache, execution unit, instruction-level unit, and bus interface.
[0086] (10) The Translation Lookaside Buffer (TLB), also known as page table cache, translation bypass cache, or translation lookaside buffer, is a cache of the CPU used to improve the translation speed from virtual address to physical address. All current desktop and server processors (such as x86) use TLB. TLB has a fixed number of slots for storing tag page table entries that map virtual addresses to physical addresses. Its search key is the virtual memory address, and its search result is the physical address. If the requested virtual address exists in the TLB, a very fast matching result will be given, and then the obtained physical address can be used to access the memory. If the requested virtual address is not in the TLB, the tag page table will be used for virtual-to-physical address translation, and the access speed of the tag page table is much slower than that of the TLB. In some systems, the tag page table is allowed to be swapped to secondary storage, so the virtual-to-physical address translation may take a very long time.
[0087] (11) The page table is a special data structure placed in the page table area of the system space, storing the correspondence between logical pages and physical page frames. The logical address space is described by pages of a fixed size, and the physical memory space is described by page frames of the same size. The operating system implements the page mapping from logical pages to physical page frames and is also responsible for the management of all pages and the control of process execution.
[0088] (12) The Memory Controller is a bus circuit controller used to manage and plan the transfer speed from memory to the CPU; it can be a separate chip or integrated into a related large chip. It performs necessary control over memory access according to certain timing rules, including the control of address signals, data signals, and various command signals, enabling the CPU to use the storage resources of the memory according to its needs.
[0089] (13) The Point of Serialization (PoS), also known as the Home Node / Ordering Point, is a key point for maintaining coherence among multiple cores in a multi-core processor. At the PoS, it is possible to monitor whether the data in the caches of all processor cores has been modified and its status to ensure data coherence among the caches. Therefore, in this application, processing HPTW through the PoS can ensure that old page table data will not be fetched and the data correctness is guaranteed.
[0090] (14) The CPU pipeline technology is a technology that decomposes instructions into multiple steps and overlaps the operations of each step of different instructions, so as to realize the parallel processing of several instructions and accelerate the program running process. Each step of the instruction has its own independent circuit for processing. After each step is completed, it proceeds to the next step, while the previous step processes subsequent instructions. The pipeline structure of the processor is the most basic element of the processor microarchitecture, which bears and determines the details of other microarchitectures of the processor.
[0091] (15) The bus is the key to the multi-core connection of the processor and is used to transfer information between the core and the cache, memory, etc.
[0092] (16) The Extended Page Table (EPT) consists of four-level page tables, namely the page map level 4 table (PML4), the page-directory-pointer table (PDPT), the page-directory (PD), and the page table (PT).
[0093] (17) Hardware Page Table Walk (HPTW) is the process by which the hardware module queries the page table. The virtual address used when the program accesses memory needs to be converted into a physical address before the memory can be accessed. After a TLB miss occurs, the hardware can complete the traversal of the page table to find the missing page table. For example, a 48-bit virtual address is finally determined to obtain the corresponding physical address after a complete traversal. Please refer to Figure 1 , Figure 1 which is a schematic diagram of a hardware page table traversal process provided by an embodiment of the present invention. As Figure 1 shown, the physical base address of the page table is stored in the page base address register. After a TLB miss occurs, the MMU part in the processor core provides the HPTW function, and uses the physical address in CR3 as the base address of the first-level page table. After obtaining this data, the PML4 in VA is used as the address offset [47:39] to query and obtain the base address of the second-level page table; after splicing [38:30] of VA and continuing to query until the final physical address is obtained. The HPTW process needs to query the L2 cache. If the page table hits in the L2 cache, the above process can be accelerated to a certain extent (without having to fetch data from memory every time), otherwise the above HPTW process needs to perform four memory access operations to complete.
[0094] (18) The on-chip bus protocol provides a special mechanism that can integrate the processor into other Intellectual Property (IP) cores and peripherals.
[0095] (19) The base address, or physical base address, simply referred to as the base address, can be understood as the basic address where data (such as a page table or a page) is stored in the memory, and it is the calculation reference for the relative offset (offset address). For example, in the case of a multi-level page table, after determining the base address of a certain level of the page table, according to the corresponding offset address, a page table entry in this level of the page table can be determined as the base address of the next level of the page table.
[0096] First, a system architecture on which the embodiments of the present invention are based will be described. Please refer to Figure 2 , Figure 2 is a schematic diagram of a system architecture provided by an embodiment of the present invention; as Figure 2 shown, the system architecture includes a processor chip 10, a processor core 11, a bus 12, a memory controller 14, and a memory 20. Among them, inside the processor chip 10, there are a processor core 11, a bus 12, and a memory controller 14. The processor core 11 may include processor core 1, processor core 2,..., and processor core Q (Q is an integer greater than 0). Multiple processor cores and the memory controller 14 are all connected to the bus 12. The processor chip 10 processes data through the processor core 11 (for example, in the embodiments of the present invention, the processor core is responsible for converting the virtual address into a physical address. Specifically, in the case of only one processor core, this core can complete the process of hardware page table traversal through other devices on the chip; in the case of multiple processor cores and when the cache of a certain core does not store all the base addresses, data interaction occurs between multiple cores, and cooperate with other on-chip devices to complete the process of hardware page table traversal); the processor chip 10 interacts with the memory 20 through the memory controller 14, such as reading data or writing data, etc. The bus is a channel for data interaction between multiple cores and other components (such as the memory controller, etc.).
[0097] Based on the foregoing system architecture, taking a certain processor core as an example, next, one of the application architectures on which the embodiments of the present invention are based will be described. Please refer to Figure 3 , Figure 3 is a schematic diagram of an application architecture provided by an embodiment of the present invention; as Figure 3 shown, the embodiments of the present invention can be applied to the processor chip 10; the processor chip 10 is connected to the memory 20. The processor chip 10 includes a processor core 11, a bus 12, a home node 13, and a memory controller 14. The embodiments of the present invention do not limit the specific connection manner between modules, units, or devices. Among them,
[0098] The memory 20 is used to store each level of the base address in the M-level physical base address (i.e., the M-level base address), where M is an integer greater than 1. For example, in the case of a four-level page table, the memory stores the base addresses of the first-level page table to the fourth-level page table, a total of 4 levels of base addresses.
[0099] The memory controller 14 is used to receive the K-level memory access request sent by the home node; according to the relevant information of the K-level base address in the K-level memory access request, obtain the corresponding K-level base address from the memory, and return the K-level base address to the home node; K is an integer greater than 0.
[0100] The home node 13 is used for:
[0101] Receive a request, and the type of the request may include HPTW requests and other types of requests.
[0102] Determine whether the received request is a memory access request (i.e., an HPTW request).
[0103] On the premise that the home node determines that the request is a memory access request, identify the level of the target page table base address queried by the memory access request (for example, for a K-level memory access request, the queried is the K-level page table base address); optionally, through the home node buffer (including the identification module), identify the level of the target page table base address.
[0104] Determine which level of the N-level first cache stores the base address of the target page table.
[0105] After the home node determines which level of the N-level first cache stores the base address of the target page table, send the memory access request to the first cache that stores the base address of the target page table.
[0106] In the case where it is determined that none of the N-level first caches stores the K-level base address, send the K-level memory access request to the memory controller. Optionally, the home node calculates the (K + 1)-level base address based on the obtained K-level base address, and further queries according to the (K + 1)-level base address. For example, if the home node obtains the 3-level base address from the memory and calculates the 4-level base address, it can immediately continue to search for the 4-level base address in the home node PoS.
[0107] The processor core 11 may include a processor core pipeline 110, an arithmetic unit 111, a memory access unit 112, a memory management unit 113, and a first cache 114. Among them, the first cache 114 may include N-level first caches, such as the 1st-level first cache, the 2nd-level first cache, the 3rd-level first cache,..., the i-th level first cache,..., and the Nth-level first cache, etc.; i = 1, 2,..., N, that is, the value of i can be 1, 2, 3,... or N, and N is an integer greater than 0; the embodiment of the present invention does not limit the number of multi-level caches included in the first cache.
[0108] The processor core pipeline 110 is used to process instructions between units or devices such as the arithmetic unit 111 and the memory access unit 112 in parallel. It can be understood that the arithmetic unit or execution unit may include an arithmetic logic unit (ALU).
[0109] The memory access unit 112 is used to send the virtual address provided by the application program to the memory management unit (MMU), and receive the physical address corresponding to the virtual address fed back by the MMU.
[0110] Optionally, when the MMU is not enabled, the memory access unit 112 can interact with the first cache. The memory access unit 112 can obtain the page base address returned by the first cache 114, and then send the page base address to the ALU for calculation processing. For example, the page base address is fed back to the ALU, and the ALU adds the page base address and the page offset to obtain the final required physical address.
[0111] The memory management unit 113 is used to send a first-level memory access request to the first-level first cache. For example, after a TLB miss occurs, after determining the TLB miss page table base address (in the case of a four-level page table, the TLB stores the base addresses of the 1-4 level page tables or does not store the base addresses of the 1-4 page tables, and there is no situation where only part of the level page table base addresses are stored), then send the first-level memory access request (including the first-level page table base address and the high-order address in the virtual address); receive the (M + 1)-level base address, that is, the base address of a certain page in the last page table. Optionally, the memory management unit 113 may include a translation lookaside buffer 1130. It can be understood that the memory management unit 113 is part of the storage unit.
[0112] The first cache 114 is used to store the base addresses of one or more levels of page tables in the M-level page table base addresses.
[0113] Optionally, the processor core 11 may be Figure 3 the kernel shown, or may include multiple kernels. For example, please refer to Figure 4 , Figure 4 which is another schematic diagram of the application architecture provided by the embodiment of the present invention. As Figure 4As shown, the processor core 11 includes multiple cores, such as processor core 1, processor core 2, …, processor core Q (Q is an integer greater than 1). Among them, the home node 13 can include multiple home nodes. For example, at the connection of each core and the bus, a corresponding home node can be set. The PoS and the Nth-level first cache (such as the L3 cache is a cache shared by multiple cores outside the core) are logically and functionally independent; the PoS can be logically located anywhere inside the processor. In the processor design, the two can be designed separately (that is, located at different physical positions in the processor); or, in order to accelerate the interaction between the two, the two are placed together or closely adjacent in the physical structure. The embodiments of the present invention do not limit the organizational structures of the PoS and the L3 cache.
[0114] Specifically, in Figure 4 the architecture shown, the home node is used for:
[0115] When it is determined that any one of the second cache and the Nth-level first cache stores the Kth-level base address, the home node sends a Kth-level memory access request to the corresponding cache;
[0116] When it is determined that neither the second cache nor the Nth-level first cache stores the Kth-level base address, a Kth-level memory access request is sent to the memory controller.
[0117] It should be noted that for other contents of the home node (such as receiving memory access requests, identifying base address levels, etc.), please refer to the relevant descriptions in the foregoing Figure 1 and will not be elaborated here.
[0118] The processor core 11 may further include a second cache 115; the second cache includes (N - 1) levels of first caches, such as the first-level first cache, the second-level first cache, the third-level first cache, …, the ith-level first cache, …, and the (N - 1)th-level first cache, etc. It can be understood that in the case of multiple cores, the first cache includes the Nth-level first cache and the second caches 115 of all cores; in the case of a single core, the first cache includes all N levels of first caches. It can be understood that in the case of multiple cores, for a certain core, the first cache can be the cache within the core and the cache shared with other cores (for example, Figure 3 the Nth-level first cache shown; optionally, the shared cache can be one or more) in general; the second cache is the general term for the caches within other cores.
[0119] It should be noted that Figure 4 the processor core pipeline 110, arithmetic unit 111, memory access unit 112, memory management unit 113, etc. shown in Figure 3 the same units or devices as those in Figure 3The related descriptions and the explanations of some of the foregoing terms will not be elaborated herein.
[0120] It should be noted that this application can be specifically applied to all of the above caches and home nodes to accelerate the HPTW process.
[0121] Combined with Figure 3 the application architecture described above, a hardware page table traversal acceleration device related to an embodiment of the present invention will be described below. Please refer to Figure 5 , Figure 5 which is a schematic diagram of a hardware page table traversal acceleration device provided by an embodiment of the present invention; as Figure 5 shown, it mainly describes the interaction between the memory management unit 113, the first cache 114, the home node 13, the memory controller 14, and the memory 20 involved in Figure 1 . The memory management unit 113 is connected to the secondary cache 1141; the tertiary cache 1142 is connected to the secondary cache 1141 and is also connected to the memory controller 14 through the bus 12. Among them, a home node 13 can be set at the node where the tertiary cache 1142 is connected to the bus. The memory controller 14 in the processor chip 10 is connected to the memory 20. For the related content of the components or units involved in the embodiments of the present invention, please refer to the descriptions of the foregoing embodiments, which will not be elaborated herein.
[0122] In the case of N = 2 and M = 4, one or more levels of the four-level base addresses are stored in these two levels of the first cache, that is, only some of the base addresses may be stored in the first cache or all of the base addresses may be stored. Please refer to Figure 6 , Figure 6 which is a schematic diagram of a specific hardware page table traversal acceleration device provided by an embodiment of the present invention; as Figure 6 shown, the processor chip includes a processor core 11 and a tertiary cache (L3 cache, that is, the second-level first cache), and the processor core 11 includes a secondary cache (L2 cache, that is, the first-level first cache). In the figure, the L3 cache is an off-core cache, but the embodiments of the present invention are not limited thereto, that is, the L3 cache can also be an on-core cache. It should be noted that Figure 6 this is only an exemplary situation. For the case of a multi-level cache including a four-level cache, a five-level cache, a six-level cache, etc. that can store page table base addresses, reference can be made to Figure 6 and the descriptions of the corresponding embodiments.
[0123] Among them, for the internal structure of the cache, please refer to Figure 7 , Figure 7 which is a schematic diagram of the internal structure of a cache provided by an embodiment of the present invention; as Figure 7 shown, the i-th level of the first cache includes the control unit of the i-th level of the first cache and the computing unit of the i-th level of the first cache.
[0124] The computing unit of the first cache at the i-th level is used to determine the base address at the (K + 1)-th level according to the base address at the K-th level and the offset address at the K-th level; the offset address at the K-th level is determined according to the high-order address at the K-th level in the virtual address; K is an integer greater than 0 and less than or equal to M, and i = 1, 2,..., N. For example, according to the base address at the first level and the offset address at the first level (address offset or offset), the base address at the second level is determined; assuming there are a total of 4 levels of page tables, the base address at the fifth level is determined according to the base address at the fourth level and the offset address at the fourth level (the base address at the fifth level is the physical base address of a certain page in the fourth-level page table); therefore, the base address can be the page table or the physical base address of the page. Specifically, for example, when the base address of the first-level page table hits the L2 cache, after the L2 cache determines the base address of the first-level page table, it determines which item in the first-level page table is the base address of the second-level page table according to the offset address; assuming that [47:39] (i.e., the high-order address at the first level) in the 48-bit virtual address is 000100000, which corresponds to 32, and on the premise that 32 is within the serial number range of the first-level page table, the cache checks the 32nd item in the first-level page table and obtains the physical base address of the second-level page table stored therein, that is, determines the base address of the second-level page table. For the specific description of base address determination and address concatenation, please refer to the explanation of some terms (19), which will not be elaborated here.
[0125] Based on Figure 3 the architecture and related devices shown, the four page table query processes that may be involved in the embodiments of the present invention are described in the case of N = 2 and M = 4.
[0126] Please refer to Figure 8 and Figure 9 , Figure 8 is a schematic diagram of a page table query process provided by an embodiment of the present invention; as Figure 8 shown, the value of K in the K-th level memory access request, the K-th level base address, and the (K + 1)-th level base address includes 1, 2, 3,..., M. Figure 9 is an interactive schematic diagram of some hardware corresponding to Figure 8 provided by an embodiment of the present invention; assuming that the L2 cache hits all four levels of page table base addresses (i.e., each level of the 4-level base address is stored in the L2 cache), as Figure 9 shown, in the device for accelerating hardware page table traversal shown in the foregoing Figure 6 , its various functional modules can perform corresponding operations according to the following timing, and the specific steps are as follows:
[0127] Step 1 (S1): The control unit 71 of the secondary cache receives the first-level memory access request (including the first-level base address and the first-level high-order address).
[0128] Specifically, receive the first-level memory access request from the MMU113; for the command format of the memory access request sent from the MMU to the secondary cache, please refer to Figure 10 , Figure 10 which is a schematic diagram of the command format of a memory access request provided by an embodiment of the present invention; as Figure 10 shown, the command format may include a high-order address, a request type, the base address of the i-th level page table, and a bit field. Combining the foregoing embodiments, in the example of a 48-bit virtual address and a 4-level page table, the corresponding command format is {high-order address [47:12], memory access request (i.e., HPTW request), base address of the i-th level page table, and bit field [1:0]}. Taking the command format of the first-level memory access request as an example, as Figure 9 shown, [47:12] is the high-order address in the virtual address except for the low-order address of this segment other than the page offset offset [11:0], including [47:39] (i.e., the first-level high-order address), [38:30] (i.e., the second-level high-order address), [29:21] (i.e., the third-level high-order address), and [20:12] (i.e., the fourth-level high-order address).
[0129] Optionally, determine whether the request is a memory access request according to the data in the area of the "request type".
[0130] Optionally, after determining that a certain request is an HPTW request, determine which level of page table the HPTW request queries according to the bit field in the command format. Taking {[47:12], memory access request, base address of the first-level page table, 00} as an example, the bit field 00 is used to indicate that the memory access request queries the base address of the second-level page table.
[0131] Optionally, the control unit of the secondary cache receives the first-level memory access request from the MMU; wherein, the high-order address included in the first-level memory access request may be [47:12].
[0132] Optionally, after the MMU triggers a page table traversal, the process of the hardware page table traversal starts from obtaining the base address of the starting page table in the CR3 and queries until the base address of the target page table at this level is queried. For example, on the premise of a four-level page table, if there is only a page table entry missing in the TLB that contains the base address of the second-level page table, then the traversal process ends when the base address of the second-level page table is queried.
[0133] Step 2 (S2): The control unit 71 of the secondary cache determines that the secondary cache has the first-level base address based on the first-level base address; then it locates the stored first-level base address according to the first-level high-order address. Specifically, it determines the high-order part of the first-level base address, that is, the tag of this address; it determines whether this tag is stored in the secondary cache, and if this tag is stored in the secondary cache, it can be determined that the first-level base address is stored in the secondary cache. It can be understood that the tag is a part of the base address, and it can be considered that a part of the base address is used as the identifier for determining whether there is storage; optionally, the received first-level base address is compared with all the stored base addresses to determine whether the secondary cache stores the same base address.
[0134] Step 3 (S3): The control unit 71 of the secondary cache sends the first-level base address and the first-level high-order address to the computing unit of the secondary cache. Specifically, the computing unit sends the first-level base address and the high-order address (such as [47:12]); it can be understood that the high-order address includes the first-level high-order address (such as [47:39]). Optionally, after determining that the memory access request queries the page table level, the high-order address corresponding to the level can be determined from the high-order address. For example, if the current memory access request queries the first-level page table base address, then it is determined whether the secondary cache hits the first-level page table base address according to the tag of the first-level page table base address, and the first-level page table base address is queried after a hit.
[0135] Step 4 (S4): The computing unit 72 of the secondary cache calculates the second-level base address by adding the offset address to the first-level base address.
[0136] Specifically, after determining the first-level base address, the corresponding first-level offset address is concatenated to the first-level base address to obtain the second-level base address. Based on the first-level base address, the first-level offset address is further added to determine the second-level base address (for the description of this calculation process, please refer to the relevant description of the foregoing embodiments and will not be elaborated here), that is, in the first-level page table, it is determined which page table entry in the first-level page table is the base address of the second-level page table through the first-level offset address.
[0137] Step 5 (S5): After the computing unit 72 of the secondary cache obtains the second-level base address, it sends the second-level base address to the control unit 71 of the secondary cache.
[0138] Step 6 (S6): The control unit 71 of the secondary cache determines that the secondary cache has the second-level base address according to the second-level high-order address (the second-level high-order address is included in the high-order address); then it locates the stored second-level base address according to the second-level base address.
[0139] Specifically, please refer to the description in Step 2 above and will not be elaborated here.
[0140] Step 7 (S7): The control unit 71 of the secondary cache sends the second-level base address and the second-level high-order address to the computing unit of the secondary cache.
[0141] Specifically, please refer to the description of the foregoing step 3, which will not be elaborated herein.
[0142] ……
[0143] Step (L-1): The calculation unit 72 of the secondary cache calculates the fifth-level base address (i.e., the page base address in the fourth-level page table) by adding the offset address to the fourth-level base address.
[0144] Specifically, after obtaining the fourth-level base address, the page table entry in the fourth-level page table is determined according to the fourth-level offset address, and the page base address corresponding to the required physical address is thus determined.
[0145] Step L (SL, hereinafter all represented by step L / SL for the last step): The calculation unit 72 of the secondary cache feeds back the fifth-level base address to the control unit 71 of the secondary cache.
[0146] Optionally, the control unit 71 of the secondary cache may feed back this page base address to the MMU113. Further optionally, after obtaining the fifth-level base address, the MMU may complete the conversion from VA to PA according to the page base address and the page offset.
[0147] In the embodiment of the present invention, by adding a calculation unit to each level of the first cache, after each level of the first cache obtains the base address, it can directly calculate the base address of the next level according to the base address and the corresponding offset. Specifically, the control unit of the i-th level of the first cache receives the K-th level memory access request. On the premise of determining that the i-th level of the first cache (i.e., the current level of cache) stores the K-th level base address, it sends the K-th level base address in the K-th level memory access request to the calculation unit of the current level of cache. The current level of cache calculates the (K + 1)-th level base address (i.e., the base address of the next level) according to the K-th level base address and the offset address determined from the high-order address included in the K-th level memory access request. Different from the prior art that after any level of cache determines the K-th level base address, it needs to return the base address of this level to the memory access unit in the processor core, and then calculate the base address of the next level through the arithmetic logic unit in the processor core, and then continue to query the base address of the next level; in the embodiment of the present invention, no matter which level of the first cache determines the K-th level base address, the base address of the next level can be calculated by shifting through the calculation unit added to the current level of cache, improving the access efficiency of the base address, accelerating the traversal process of the hardware page table, and ultimately improving the conversion efficiency from virtual address to physical address.
[0148] It should be noted that the identifiers S1, S2, S3, etc. shown in the figure correspond to steps such as step 1, step 2, step 3, etc. The following embodiments also use the same identifiers and will not be described again later. It can be understood that the corresponding parts of the identifiers shown in this application are exemplary descriptions.
[0149] Please refer toFigure 11 and Figure 12 , Figure 11 is another schematic diagram of the page table query process provided by an embodiment of the present invention; as Figure 11 shown, the memory access request can be a K-level memory access request, and the base address can be a K-level base address; wherein, the value of K includes 1, 2, 3, …, M. Figure 12 is an interaction schematic diagram of a part of the hardware corresponding to Figure 11 of the present invention; assuming that the L2 cache hits the first-level base address (i.e., the first-level page table base address), and the L3 cache hits the second-fourth level base addresses; as Figure 12 shown, in the device for accelerating the hardware page table traversal shown in the foregoing Figure 6 , each of its functional modules can perform corresponding operations according to the following timing sequence, and the specific steps are as follows:
[0150] Step 1 (S1): The control unit 71 of the secondary cache receives the first-level memory access request.
[0151] Step 2 (S2): The control unit 71 of the secondary cache determines whether the secondary cache has the first-level base address according to the first-level base address; then finds the stored first-level base address according to the first-level base address.
[0152] Step 3 (S3): The control unit 71 of the secondary cache sends the first-level base address and the first-level high-order address to the computing unit of the secondary cache.
[0153] Step 4 (S4): The computing unit 72 of the secondary cache calculates the second-level base address by adding the offset address to the first-level base address.
[0154] Step 5 (S5): After the computing unit 72 of the secondary cache obtains the second-level base address, it sends the second-level base address to the control unit 71 of the secondary cache.
[0155] Step 6 (S6): The control unit 71 of the secondary cache determines that the secondary cache does not have the second-level base address according to the second-level base address; then sends the second-level memory access request to the control unit 81 of the tertiary cache.
[0156] Specifically, in the case where it is determined that the secondary cache (the second-level first cache) does not store the second-level base address and i≠N, the second-level memory access request is sent to the tertiary cache (i.e., the third-level first cache). In the case of i = N, for example, after it is determined that the tertiary cache does not store the second-level base address, the tertiary cache can send the second-level memory access request to the home node.
[0157] Optionally, after determining that the secondary cache does not have the second-level base address, the control unit 71 of the secondary cache may also send a second-level memory access request to the home node at the same time. It can be understood that in the embodiments of the present invention, the operations performed by the home node and the control unit 81 of the tertiary cache after receiving the second-level memory access request are basically the same. Therefore, only the interaction relationship between the tertiary cache and the secondary cache is described in the figure.
[0158] Step 7 (S7): The control unit 81 of the tertiary cache receives the second-level memory access request, and the second-level memory access request includes the second-level base address and the second-level high-order address (which is also included in the high-order address).
[0159] Step 8 (S8): The control unit 81 of the tertiary cache determines whether the tertiary cache has the second-level base address according to the second-level base address; then it locates the stored second-level base address according to the second-level base address.
[0160] Step 9 (S9): The control unit 81 of the tertiary cache sends the second-level base address and the second-level high-order address to the computing unit of the tertiary cache.
[0161] Step 10 (S10): The computing unit 82 of the tertiary cache adds the corresponding offset address to the second-level base address to calculate the third-level base address.
[0162] ……
[0163] Step L (SL): The computing unit 82 of the tertiary cache feeds back the fifth-level base address to the control unit 81 of the tertiary cache.
[0164] It should be noted that the content similar to the foregoing embodiments in the embodiments of the present invention will not be described in detail. Please refer to the description of the corresponding steps in the foregoing embodiments.
[0165] Please refer to Figure 13 and Figure 14 , Figure 13 is another schematic diagram of the page table query process provided by the embodiments of the present invention; as Figure 13 shown, the memory access request may be a K-level memory access request, and the base address may be a K-level base address; wherein, the value of K includes 1, 2, 3,..., M. Figure 14 is an interaction schematic diagram of some hardware corresponding to Figure 13 provided by the embodiments of the present invention; assuming that the L2 cache hits the first-level base address, the third-level base address, and the fourth-level base address, and the L3 cache hits the second-level base address; as Figure 14 shown, in the device for accelerating the hardware page table traversal shown in the foregoing Figure 6 , each of its functional modules may perform corresponding operations according to the following timing sequence. The specific steps are as follows:
[0166] Step 1 (S1): The control unit 71 of the secondary cache receives the first-level memory access request.
[0167] Step 2 (S2): The control unit 71 of the secondary cache determines that the secondary cache has the first-level base address according to the first-level base address; then it locates the stored first-level base address according to the first-level base address.
[0168] Step 3 (S3): The control unit 71 of the secondary cache sends the first-level base address and the first-level high address to the computing unit of the secondary cache.
[0169] Step 4 (S4): The computing unit 72 of the secondary cache adds the offset address to the first-level base address to calculate the second-level base address.
[0170] Step 5 (S5): After the computing unit 72 of the secondary cache obtains the second-level base address, it sends the second-level base address to the control unit 71 of the secondary cache.
[0171] Step 6 (S6): The control unit 71 of the secondary cache determines that the secondary cache does not have the second-level base address according to the second-level base address; then it sends the second-level memory access request to the control unit 81 of the tertiary cache.
[0172] Optionally, after the control unit 71 of the secondary cache determines that the secondary cache does not have the second-level base address, it can also send the second-level memory access request to the home node at the same time. It can be understood that in the embodiments of the present invention, the operations performed by the home node and the control unit 81 of the tertiary cache after receiving the second-level memory access request are basically the same, so only the interaction relationship between the tertiary cache and the secondary cache is described in the figure.
[0173] Step 7 (S7): The control unit 81 of the tertiary cache receives the second-level memory access request, and the second-level memory access request includes the second-level base address and the second-level high address (also included in the high address).
[0174] Step 8 (S8): The control unit 81 of the tertiary cache determines that the tertiary cache has the second-level base address according to the second-level base address; then it locates the stored second-level base address according to the second-level base address.
[0175] Step 9 (S9): The control unit 81 of the tertiary cache sends the second-level base address and the second-level high address to the computing unit of the tertiary cache.
[0176] Step 10 (S10): The computing unit 82 of the tertiary cache adds the corresponding offset address to the second-level base address to calculate the third-level base address.
[0177] Step 11 (S11): After the tertiary cache unit 82 obtains the third-level base address, it sends the third-level base address to the control unit 81 of the tertiary cache.
[0178] Step 12 (S12): The control unit 81 of the tertiary cache determines that the tertiary cache does not have the third-level base address based on the third-level base address; then it sends a third-level memory access request to the home node 13.
[0179] Specifically, when sending a memory access request from the control unit 81 of the tertiary cache to the home node PoS13, a new command is added compared to the requests in the prior art. The requests coming out of the cache need to be sent to the PoS for unified processing. To implement the embodiments of the present invention, a new command needs to be added to the request command. In the AMBA CHI protocol command encoding, the Opcode field is the command format, and 0x3B - 0x3F can be reserved commands (i.e., reserve commands). The embodiments of the present invention can implement the HPTW request by using the 0x3B command. Optionally, an additional Addr[47:12] field (used to identify the high-order address in the original virtual address) and a Level[1:0] field (indicating which level of page table the PoS is currently processing) are added to the Request flit. It can be understood that after adding a new command in the prior art, the length of the original command increases. For example, the original command length is 137, and after the increase, the command length is 137 (original command length) + 36 (i.e., the length of the high-order address in the original virtual address) + 2 (i.e., the length of the Level[1:0] field) = 175.
[0180] Step 13 (S13): The home node 13 receives the third-level memory access request.
[0181] Step 14 (S14): After the home node 13 receives the third-level memory access request, it determines whether each first-level cache stores the third-level base address based on the third-level base address. In the case where it is determined that the secondary cache stores the third-level base address, it sends a third-level memory access request to the secondary cache.
[0182] Specifically, it is determined which one or which caches among all the caches store the third-level base address based on the third-level base address. Among them, the structure of the PoS has two columns of data. The first column stores the Tag (i.e., the physical address) of the cache line, and the second column stores in which cache (such as the secondary cache) of which level the cache line is located. For example, if the second column stores that the cache line is stored in which secondary cache, then 0001 can represent being in the first L2 cache, and 1111 can represent that the data exists in all four L2 caches.
[0183] Optionally, the home node also includes a buffer. For example, four new entries are added, and the four entries have the same structure as the home node, but can be specifically used for HPTW (that is, only HPTW requests enter the above four entries for processing, for example, the first-level page table enters the first entry, or the two columns of data in the first entry only store relevant information of the first-level page table). After the home node 13 receives the third-level memory access request, the buffer in the home node determines which level of the first cache stores the third-level base address according to the tag of the third-level base address. For example, the home node first determines that the memory access request is a query for the third-level base address based on the bit field (such as 10), and then determines which cache stores the third-level base address (for example, it determines that there is a third-level base address in the second-level cache).
[0184] Step 15 (S15): The control unit 71 of the secondary cache receives the level 3 memory access request. Optionally, after the secondary cache receives the level 3 memory access request, the control unit 71 of the secondary cache can determine again whether the level 3 base address is stored in the cache of this level according to the memory access request.
[0185] …
[0186] Step L(SL): the calculation unit 72 of the L2 cache feeds back the L5 base address to the control unit 71 of the L2 cache.
[0187] Optionally, the control unit 71 of the secondary cache feeds back the fifth-level base address to the MMU.
[0188] It should be noted that the contents of each step in the embodiment of the present invention that are similar to those in the aforementioned embodiment will not be repeated here, and please refer to the description of the corresponding steps in the aforementioned embodiment.
[0189] See also Figure 15 and Figure 16 , Figure 15 FIG. 2 is another schematic diagram of a page table query process provided by an embodiment of the present invention; Figure 15 As shown, the memory access request can be a K-th level memory access request, and the base address can be a K-th level base address; wherein the value of K includes 1, 2, 3, ..., M. Figure 16 This is a corresponding embodiment provided by the present invention. Figure 15 Schematic diagram of the interaction of some hardware; assuming that L2 cache hits the first-level base address and the third-level base address, L3 cache hits the second-level base address, and the fourth-level base address is stored in memory; Figure 16 As shown in the aforementioned Figure 6 In the device for accelerating hardware page table traversal shown in the figure, each functional module thereof can perform corresponding operations according to the following timing sequence, and the specific steps are as follows:
[0190] Step 1 (S1): The control unit 71 of the secondary cache receives a first-level memory access request.
[0191] Step 2 (S2): The control unit 71 of the secondary cache determines that the secondary cache has the first-level base address according to the first-level base address; then it locates the stored first-level base address according to the first-level base address.
[0192] Step 3 (S3): The control unit 71 of the secondary cache sends the first-level base address and the first-level high address to the computing unit of the secondary cache.
[0193] Step 4 (S4): The computing unit 72 of the secondary cache calculates the second-level base address by adding the offset address to the first-level base address.
[0194] Step 5 (S5): After the computing unit 72 of the secondary cache obtains the second-level base address, it sends the second-level base address to the control unit 71 of the secondary cache.
[0195] Step 6 (S6): The control unit 71 of the secondary cache determines that the secondary cache does not have the second-level base address according to the second-level base address; then it sends a second-level memory access request to the control unit 81 of the tertiary cache.
[0196] Optionally, after determining that the secondary cache does not have the second-level base address, the control unit 71 of the secondary cache can also send a second-level memory access request to the home node at the same time. Or, in step 6, only send a second-level memory access request to the home node. It can be understood that in the embodiments of the present invention, the operations performed by the home node and the control unit 81 of the tertiary cache after receiving the second-level memory access request are basically the same, so only the interaction relationship between the tertiary cache and the secondary cache is described in the figure.
[0197] Step 7 (S7): The control unit 81 of the tertiary cache receives the second-level memory access request, and the second-level memory access request includes the second-level base address and the second-level high address (also included in the high address).
[0198] Step 8 (S8): The control unit 81 of the tertiary cache determines that the tertiary cache has the second-level base address according to the second-level base address; then it locates the stored second-level base address according to the second-level base address.
[0199] Step 9 (S9): The control unit 81 of the tertiary cache sends the second-level base address and the second-level high address to the computing unit of the tertiary cache.
[0200] Step 10 (S10): The computing unit 82 of the tertiary cache adds the corresponding offset address to the second-level base address to calculate the third-level base address.
[0201] Step 11 (S11): After the tertiary cache unit 82 obtains the third-level base address, it sends the third-level base address to the control unit 81 of the tertiary cache.
[0202] Step 12 ( S12 ): the control unit 81 of the third-level cache determines that the third-level cache does not have a third-level base address according to the third-level base address; and then sends a third-level memory access request to the home node 13 .
[0203] Step 13 (S13): The home node 13 receives the level 3 memory access request.
[0204] Step 14 (S14): After receiving the level 3 memory access request, the home node 13 determines whether each level 1 cache stores the level 3 base address according to the level 3 base address. If it is determined that the level 2 cache stores the level 3 base address, the level 3 memory access request is sent to the level 2 cache.
[0205] Optionally, the home node further includes a buffer. After the home node 13 receives the level 3 memory access request, the buffer determines which level of first cache stores the level 3 base address according to the tag of the level 3 base address. For example, the home node determines that the level 3 base address is stored in the level 2 cache.
[0206] Step 15 (S15): The control unit 71 of the second-level cache receives a third-level memory access request.
[0207] Optionally, after the secondary cache receives the third-level memory access request, the control unit 71 of the secondary cache may determine again whether the third-level base address is stored in the cache at this level according to the memory access request.
[0208] Step 16 (S16): After the control unit 71 of the secondary cache determines that the secondary cache has a third-level base address according to the third-level base address, it searches for the stored third-level base address according to the third-level base address.
[0209] Step 17 ( S17 ): The control unit 71 of the secondary cache sends the third-level base address and the third-level high-order address to the calculation unit 72 of the secondary cache.
[0210] Step 18 (S18): The calculation unit 72 of the secondary cache adds the offset address to the third-level base address to calculate the fourth-level base address.
[0211] Step 19 ( S19 ): After the calculation unit 72 of the secondary cache obtains the 4th level base address, it sends the 4th level base address to the control unit 71 of the secondary cache.
[0212] Step 20 ( S20 ): The control unit 71 of the secondary cache determines that the secondary cache does not have a fourth-level base address according to the fourth-level base address; and then sends a fourth-level memory access request to the control unit 81 of the third-level cache and the home node 13 .
[0213] Step 21 (S21): The home node 13 and the control unit 81 of the third-level cache both receive the fourth-level memory access request.
[0214] Specifically, in the embodiments of the present invention, optionally, after the tertiary cache receives a fourth-level memory access and does not store the fourth-level base address, it is determined that the current-level cache does not have the fourth-level base address.
[0215] Step 22 (S22): After the home node 13 receives a fourth-level memory access request, it is determined that neither the secondary cache nor the tertiary cache stores the fourth-level base address, and a fourth-level memory access request is sent to the memory controller 14.
[0216] Specifically, after the home node receives a fourth-level memory access request, it is determined that the request is a fourth-level memory access request. A request is sent to the memory controller, instructing the memory controller 14 to obtain the fourth-level base address from the memory 20. Optionally, the home node includes a buffer, and the buffer is used to determine which level of page table the request is for querying, and after the determination is completed, it processes from the current-level query onwards.
[0217] Step 23 (S23): The memory controller 14 receives a fourth-level memory access request.
[0218] Step 24 (S24): After the memory controller obtains the fourth-level base address from the memory according to the fourth-level memory access request, it sends the fourth-level base address to the home node 13. Specifically, the data interaction between the memory controller and the memory is not described in detail here.
[0219] Step 25 (S25): The home node 13 calculates the fifth-level base address based on the fourth-level base address and the fourth-level offset address.
[0220] Optionally, the buffer of the home node may further include a buffer calculation unit; the buffer is also used to perform a shift calculation on the base address. Further optionally, after the home node obtains the fifth-level base address, it sends the fifth-level base address to the MMU.
[0221] It should be noted that the content similar to the foregoing embodiments in each step of the embodiments of the present invention will not be repeated here, and please refer to the description of the corresponding steps in the foregoing embodiments.
[0222] It can be understood that the page table query process involved in the embodiments of the present invention may include but is not limited to the four processes provided above. For example, in the case of storing the fourth-level cache, the fifth-level cache, and other-level caches, etc., the description in the foregoing diagrams can be referred to.
[0223] Based on Figure 3 the architecture and related devices shown below, a hardware-accelerated page table traversal device involved in the embodiments of the present invention will be described. Please refer to Figure 17 , Figure 17 is a schematic diagram of a hardware-accelerated page table traversal device provided by the embodiments of the present invention; as Figure 17As shown in the figure, a third cache 1131 is added to the memory management unit 113. After a TLB miss occurs when the memory management unit 113 queries the TLB, it first queries whether the page table base address is stored in the third cache. If the third cache is hit and the entire process of page table traversal can be completed relying on the third cache, the solution provided in the foregoing embodiment does not need to be executed. Otherwise, for the missing page table base address in the TLB 1130 and the third cache 1131, the page table traversal solution provided in the foregoing embodiment of the present invention is continued to be executed.
[0224] Based on Figure 4 the architecture shown in the figure, a page table query process that may be involved in an embodiment of the present invention is described in the case of N = 2 and M = 4.
[0225] First, a specific accelerated hardware page table traversal device in the case of N = 2 and M = 4 is described below.
[0226] Please refer to Figure 18 , Figure 18 which is a schematic diagram of another specific accelerated hardware page table traversal device provided by an embodiment of the present invention; as Figure 18 shown in the figure, the processor chip 10 includes a processor core 11, and the processor core 11 may include a processor core 1 and a processor core 2. The second cache 115 of the processor core 2 is a secondary cache (i.e., the first-level second cache), and in the embodiment of the present invention, there is only one level of secondary cache, and the secondary cache may be the second cache 115. The tertiary cache is an off-core cache shared by the processor core 1 and the processor core 2, that is, the second-level first cache. The first-level first cache and the first-level second cache can both be secondary caches. A home node 13 can be set at the connection between the processor core 1 and the bus, and a home node 13 can be set at the connection between the processor core 2 and the bus. The home node 13 may include multiple home nodes, but the functions of the multiple home nodes are the same.
[0227] Please refer to Figure 19 , Figure 19 which is a schematic diagram of a page table query process in the case of multiple cores provided by an embodiment of the present invention; as Figure 19 shown in the figure, as Figure 19 shown in the figure, the memory access request may be a K-level memory access request, and the base address may be a K-level base address; wherein, the value of K includes 1, 2, 3,..., M. Figure 20 is an interactive schematic diagram of some hardware corresponding to Figure 19 ; assuming that the L2 cache (i.e., the secondary cache in the first cache) hits the first-level base address and the third-level base address, the L3 cache (i.e., the tertiary cache in the first cache) hits the second-level base address, and the fourth-level base address is stored in the secondary cache of the second cache; as Figure 20 shown in the figure, in the foregoing Figure 18In the device for accelerating the hardware page table traversal shown, each functional module can perform corresponding operations according to the following timing sequence. The specific steps are as follows:
[0228] Step 1 (S1): The control unit 71 of the secondary cache receives the first-level memory access request.
[0229] Step 2 (S2): The control unit 71 of the secondary cache determines whether there is a first-level base address in the secondary cache according to the first-level base address; then it locates the stored first-level base address according to the first-level base address.
[0230] Step 3 (S3): The control unit 71 of the secondary cache sends the first-level base address and the first-level high-order address to the computing unit of the secondary cache.
[0231] Step 4 (S4): The computing unit 72 of the secondary cache calculates the second-level base address by adding the offset address to the first-level base address.
[0232] Step 5 (S5): After the computing unit 72 of the secondary cache obtains the second-level base address, it sends the second-level base address to the control unit 71 of the secondary cache.
[0233] Step 6 (S6): The control unit 71 of the secondary cache determines that there is no second-level base address in the secondary cache according to the second-level base address; then it sends the second-level memory access request to the control unit 81 of the tertiary cache.
[0234] Step 7 (S7): The control unit 81 of the tertiary cache receives the second-level memory access request, and the second-level memory access request includes the second-level base address and the second-level high-order address (also included in the high-order address).
[0235] Step 8 (S8): The control unit 81 of the tertiary cache determines that there is a second-level base address in the tertiary cache according to the second-level base address; then it locates the stored second-level base address according to the second-level base address.
[0236] Step 9 (S9): The control unit 81 of the tertiary cache sends the second-level base address and the second-level high-order address to the computing unit of the tertiary cache.
[0237] Step 10 (S10): The computing unit 82 of the tertiary cache adds the corresponding offset address to the second-level base address to calculate the third-level base address.
[0238] Step 11 (S11): After the computing unit 82 of the tertiary cache unit obtains the third-level base address, it sends the third-level base address to the control unit 81 of the tertiary cache.
[0239] Step 12 (S12): The control unit 81 of the tertiary cache determines that there is no third-level base address in the tertiary cache according to the third-level base address; then it sends the third-level memory access request to the home node 13.
[0240] Step 13 (S13): The home node 13 receives a third-level memory access request.
[0241] Step 14 (S14): After the home node 13 receives the third-level memory access request, it determines whether each level of the first cache stores the third-level base address according to the third-level base address. In the case where it is determined that the second-level cache stores the third-level base address, a third-level memory access request is sent to the second-level cache.
[0242] Step 15 (S15): The control unit 71 of the second-level cache receives the third-level memory access request.
[0243] Step 16 (S16): After the control unit 71 of the second-level cache determines that the second-level cache has the third-level base address according to the third-level base address, it locates the stored third-level base address according to the third-level base address.
[0244] Step 17 (S17): The control unit 71 of the second-level cache sends the third-level base address and the third-level high-order address to the computing unit 72 of the second-level cache.
[0245] Step 18 (S18): The computing unit 72 of the second-level cache calculates the fourth-level base address by adding the offset address to the third-level base address.
[0246] Step 19 (S19): After the computing unit 72 of the second-level cache obtains the fourth-level base address, it sends the fourth-level base address to the control unit 71 of the second-level cache.
[0247] Step 20 (S20): The control unit 71 of the second-level cache determines that the second-level cache does not have the fourth-level base address according to the fourth-level base address; then it sends a fourth-level memory access request to the control unit 81 of the third-level cache and the home node 13.
[0248] Step 21 (S21): The home node 13 receives the fourth-level memory access request.
[0249] Optionally, the control unit 81 of the third-level cache receives the fourth-level memory access request.
[0250] Step 22 (S22): After the home node 13 receives the fourth-level memory access request, if it determines that neither the second-level cache nor the third-level cache stores the fourth-level base address and the second-level cache in the second cache stores the fourth-level base address, it sends a fourth-level memory access request to the second-level cache 115 in the second cache.
[0251] Step 23 (S23): The second-level cache 115 of the second cache receives the fourth-level memory access request sent by the home node 13.
[0252] Step 24 (S24): After the second-level cache 115 of the second cache determines the fourth-level base address stored in the second-level cache according to the received fourth-level memory access request, the second-level cache 115 of the second cache sends the fourth-level base address to the home node 13.
[0253] Step 25 (S25): The home node 13 determines the fifth-level base address according to the fourth-level base address and the fourth-level offset address.
[0254] In the case of multiple cores (that is, when there is a first cache, a second cache, and the processor chip is connected to the memory in the processor chip), assuming that the L2 cache (i.e., the secondary cache in the first cache) hits the first-level base address and the third-level base address, the L3 cache (i.e., the tertiary cache in the first cache) hits the second-level base address, and the fourth-level base address is stored in the memory. The foregoing Figure 19 and Figure 20 In the corresponding embodiment, the specific steps are as follows:
[0255] Step 1 (S1): The control unit 71 of the secondary cache receives the first-level memory access request.
[0256] ……
[0257] Step 21 (S21): The home node 13 receives the fourth-level memory access request.
[0258] (Steps 1 - Step 21 can refer to Steps 1 - Step 21 in the foregoing Figure 19 and Figure 20 corresponding embodiment, which will not be elaborated here.)
[0259] Step 22 (S22): After the home node 13 receives the fourth-level memory access request, it determines that neither the secondary cache nor the primary cache stores the fourth-level base address, and sends the fourth-level memory access request to the memory controller 14.
[0260] Step 23 (S23): The memory controller 14 receives the fourth-level memory access request.
[0261] Step 24 (S24): After obtaining the fourth-level base address from the memory according to the fourth-level memory access request, the memory controller sends the fourth-level base address to the home node 13. Specifically, the data interaction between the memory controller and the memory will not be elaborated here.
[0262] Step 25 (S25): The home node 13 calculates the fifth-level base address according to the fourth-level base address and the fourth-level offset address.
[0263] It can be understood that the secondary cache in the second cache may also include a control unit and a calculation unit, which can refer to the description in the first cache and will not be elaborated here. Other possible implementation manners and specific descriptions of the above steps can refer to the description of the foregoing Figure 15 and Figure 16 corresponding embodiment and will not be elaborated here.
[0264] It can be understood that in the case of multiple processor cores, the base address calculation node can occur within a multi-core or at the home node. For example, after the third-level base address is calculated in the second-level cache of processor core 1, the third-level base address is hit in the third-level cache, and the fourth-level base address is obtained through the calculation unit of the third-level cache. However, the fourth-level base address is not hit in the third-level cache. It is found through the home node that the fourth-level base address is hit in the second-level cache of processor core 2, and a fourth-level memory access request is sent to the second-level cache of processor core 2 through the home node. For specific descriptions, reference can be made to the descriptions of the foregoing Figure 15 and Figure 16 corresponding embodiments, which will not be elaborated here.
[0265] It should be noted that the steps and hardware structures in the embodiments of the present invention that are similar to those in the foregoing embodiments will not be elaborated. Please refer to the descriptions of the corresponding steps in the foregoing embodiments. Among them, the i-th level of the first cache (for example, the second-level cache, the third-level cache) in the foregoing embodiments all refers to a cache with the ability to store the page table base address, and does not involve the first-level cache in the prior art (since the current first-level cache does not store the page table base address). However, the situation where the first-level cache may store the page table base address in the future is not excluded. In the case where the first-level cache can store the page table base address, the first-level first cache can be the first-level cache. Otherwise, the first-level first cache in the embodiments of the present invention is generally the second-level cache. Moreover, the embodiments of the present invention do not limit the number and level of caches.
[0266] Based on Figure 4 the architecture and related devices shown below, a hardware-accelerated page table traversal device involved in the embodiments of the present invention will be described. Please refer to Figure 21 , Figure 21 which is a schematic diagram of a hardware-accelerated page table traversal device in the case of multiple cores provided by the embodiments of the present invention; as Figure 21 shown, a third cache 1131 is added to the memory management unit 113 of processor core 1. After the memory management unit 113 queries the TLB and a TLB miss occurs, it first queries whether the third cache stores all the required page table base addresses. If the third cache is hit and the entire process of page table traversal can be completed relying on the third cache, the solution provided in the foregoing embodiments does not need to be executed. Otherwise, for the page table base addresses missing in the TLB 1130 and the third cache 1131, the page table traversal solution provided in the foregoing embodiments of the present invention is continued to be executed.
[0267] It should be noted that in the case where the processor chip 10 has multiple processor cores 11 (such as processor core 1, processor core 2,..., processor core Q, etc.), each processor core can include a third cache. For example, the illustrated processor core 1 has a third cache 1131; then processor core 2 or processor core Q can also have a third cache.
[0268] Combine Figure 3 With the above application architecture, the following describes an accelerated hardware page table traversal method according to an embodiment of the present invention. Please refer to Figure 22 , Figure 22 FIG. is a schematic diagram of an accelerated hardware page table traversal method provided by an embodiment of the present invention; as Figure 22 shown, it may include step S2201 - step S2212; among them, the optional steps may include step S2205 - step S2212.
[0269] Step S2201: Receive a memory access request through the control unit of the i-th level first cache.
[0270] Specifically, the memory access request includes a K-th level base address and a K-th level high-order address. 0 < K ≤ M, K is an integer; i = 1, 2,..., N. In a possible implementation, the memory access request further includes a base address identifier of the K-th level base address, and the base address identifier is used to indicate the level of the K-th level base address.
[0271] Step S2202: Through the control unit of the i-th level first cache, determine whether the i-th level first cache stores the K-th level base address according to the K-th level base address.
[0272] Step S2203: When it is determined that the i-th level first cache stores the K-th level base address, send the K-th level base address to the computing unit of the i-th level first cache through the control unit of the i-th level first cache.
[0273] Step S2204: Through the computing unit of the i-th level first cache, determine the (K + 1)-th level base address according to the K-th level base address and the K-th level offset address.
[0274] Specifically, the K-th level offset address is determined according to the K-th level high-order address.
[0275] Step S2205: When it is determined that the i-th level first cache does not store the K-th level base address and i ≠ N, send the memory access request to the (i + 1)-th level first cache through the control unit of the i-th level first cache.
[0276] Step S2206: When it is determined that the i-th level first cache does not store the K-th level base address, send the memory access request to the home node through the control unit of the i-th level first cache.
[0277] Step S2207: After receiving the memory access request, through the home node, determine whether each level of the N-level first cache stores the K-th level base address according to the K-th level base address.
[0278] Step S2208: When it is determined that the target first cache stores the Kth level base address, the memory access request is sent to the control unit of the target first cache.
[0279] Step S2209: When it is determined that the Kth-level base address is not stored in each level of the first cache, the memory access request is sent to the memory controller through the home node.
[0280] Step S2210: determining the Kth level base address in the memory according to the Kth level base address through the memory controller.
[0281] Step S2211: Send the Kth level base address to the home node through the memory controller.
[0282] Step S2212: Determine the (K+1)th level base address through the buffer of the home node according to the Kth level base address and the Kth level offset address.
[0283] It should be noted that the method for accelerating hardware page table traversal described in the embodiment of the present invention can be found in the above Figures 8 - 17 The relevant description of the accelerated hardware page table traversal device in the device embodiment described in will not be repeated here.
[0284] Combination Figure 4 The application architecture is described below. Another method for accelerating hardware page table traversal involved in an embodiment of the present invention is described below. Figure 23 , Figure 23 is a schematic diagram of another method for accelerating hardware page table traversal provided by an embodiment of the present invention; Figure 23 As shown, it may include steps S2301 to S2317.
[0285] Step S2301: receiving a memory access request through the control unit of the i-th level first cache.
[0286] Specifically, the memory access request includes a K-th level base address and a K-th level high address; <K≤M,K为整数;i=1、2、...、N。在一种可能的实现方式中,所述访存请求还包括所述第K级基址的基址标识,所述基址标识用于指示所述第K级基址的级别。
[0287] Step S2302: determining, by the control unit of the i-th level first cache and according to the K-th level base address, whether the i-th level first cache stores the K-th level base address.
[0288] Step S2303: When it is determined that the i-th level of the first cache stores the K-th level base address, the control unit of the i-th level of the first cache sends the K-th level base address to the computing unit of the i-th level of the first cache.
[0289] Step S2304: The computing unit of the i-th level of the first cache determines the (K + 1)-th level base address according to the K-th level base address and the K-th level offset address.
[0290] Specifically, the K-th level offset address is determined according to the K-th level high address.
[0291] Step S2305: When it is determined that the i-th level of the first cache does not store the K-th level base address and i ≠ N, the control unit of the i-th level of the first cache sends the memory access request to the (i + 1)-th level of the first cache.
[0292] Step S2306: When it is determined that the i-th level of the first cache does not store the K-th level base address, the control unit of the i-th level of the first cache sends the memory access request to the home node.
[0293] Step S2307: After receiving the memory access request, the home node determines whether each level of the first cache in the N-level first cache stores the K-th level base address according to the K-th level base address.
[0294] Step S2308: When it is determined that the target first cache stores the K-th level base address, the memory access request is sent to the control unit of the target first cache.
[0295] Step S2309: After receiving the memory access request, the home node determines whether the second cache stores the K-th level base address.
[0296] Step S2310: When it is determined that the second cache stores the K-th level base address, the home node sends the memory access request to the second cache.
[0297] Step S2311: When it is determined that none of the first caches at each level and the second cache store the K-th level base address, the home node sends the memory access request to the memory controller.
[0298] Step S2312: According to the K-th level base address, the memory controller determines the K-th level base address in the memory.
[0299] Step S2313: The memory controller sends the K-th level base address to the home node.
[0300] Step S2314: Determine the (K + 1)-th level base address through the buffer of the home node according to the K-th level base address and the K-th level offset address.
[0301] Step S2315: Before sending a memory access request to the first cache of the first level, determine whether the third cache stores the first level base address through the memory management unit.
[0302] Step S2316: When it is determined that the third cache stores the first level base address, obtain the first level base address through the memory management unit.
[0303] It should be noted that on the premise that the third cache stores the first level base address, the third cache will also store the base addresses of the remaining levels for converting the target virtual address into a physical address.
[0304] Step S2317: When it is determined that the third cache does not store the first level base address, send the memory access request to the first cache of the first level through the memory management unit.
[0305] It should be noted that for the accelerated hardware page table traversal method described in the embodiments of the present invention, reference can be made to the relevant descriptions of the accelerated hardware page table traversal device in the device embodiment described above, which will not be elaborated here. Figures 19 - 21 in the relevant description of the accelerated hardware page table traversal device in the device embodiment described above, which will not be elaborated here.
[0306] As Figure 24 shown, Figure 24 is a schematic structural diagram of a chip provided by an embodiment of the present invention. The accelerated hardware page table traversal device in the foregoing embodiment can be implemented in the Figure 24 structure shown. The device includes at least one processor 241 and at least one memory 242. In addition, the device may further include general components such as an antenna, which will not be elaborated here.
[0307] The processor 241 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the above programs.
[0308] The memory 242 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.
[0309] Among them, the memory 242 is used to store the application program code for executing the above solution and is controlled by the processor 241 for execution. The processor 241 is used to execute the application program code stored in the memory 242. Specifically as follows:
[0310] Receive a memory access request through the control unit of the i-th level of the first cache. The memory access request includes the K-th level base address and the K-th level high-order address, where K is a positive integer and i = 1, 2,..., N; through the control unit of the i-th level of the first cache, determine whether the i-th level of the first cache stores the K-th level base address according to the K-th level base address; in the case where it is determined that the i-th level of the first cache stores the K-th level base address, send the K-th level base address to the calculation unit of the i-th level of the first cache through the control unit of the i-th level of the first cache; through the calculation unit of the i-th level of the first cache, determine the (K + 1)-th level base address according to the K-th level base address and the K-th level offset address, and the K-th level offset address is determined according to the K-th level high-order address.
[0311] Figure 24 When the shown chip is an acceleration hardware page table traversal device, the code stored in the memory 242 can execute the above Figure 22 Or Figure 23 The provided acceleration hardware page table traversal device method. For example, in the case where it is determined that the i-th level of the first cache does not store the K-th level base address and i ≠ N, send the memory access request to the control unit of the (i + 1)-th level of the first cache through the control unit of the i-th level of the first cache.
[0312] Alternatively, after receiving the memory access request, through the home node, determine whether each level of the first cache in the N-level first cache stores the K-level base address according to the K-level base address; in the case where it is determined that the target first cache stores the K-level base address, send the memory access request to the control unit of the target first cache.
[0313] Alternatively, in the case where it is determined that none of the levels of the first cache stores the K-level base address, send the memory access request to the memory controller through the home node; through the memory controller, determine the K-level base address in the memory according to the K-level base address; send the K-level base address to the home node through the memory controller; according to the K-level base address and the K-level offset address, determine the (K + 1)-level base address through the buffer buffer of the home node.
[0314] It should be noted that for the functions of the chip 24 described in the embodiments of the present invention, reference can be made to the relevant descriptions in the method embodiments described above Figures 22 - 23 and will not be elaborated here.
[0315] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0316] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0317] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical or other form.
[0318] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0319] In addition, each functional unit in the embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0320] Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server or a network device, etc., specifically, the processor in the computer device) to execute all or part of the steps of the methods in the various embodiments of the present application. Among them, the aforementioned storage medium may include: USB flash drives, mobile hard disks, magnetic disks, optical disks, read-only memory (ROM) or random access memory (RAM), etc., various media that can store program codes.
[0321] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. An apparatus for accelerating hardware page table traversal, characterized in that, It includes an N-level first cache; each level of the N-level first cache includes a computing unit and a control unit; the N-level first cache stores one or more levels of base addresses among the M-level base addresses, where M is an integer greater than 1; the i-th level of the N-level first cache includes the control unit of the i-th level of the first cache and the computing unit of the i-th level of the first cache; i = 1, 2,..., N, and N is a positive integer; where, The control unit of the i-th level of the first cache is used for: Receiving a memory access request, where the memory access request includes the K-th level base address and the K-th level high-order address, and K is a positive integer less than or equal to M; Judging whether the i-th level of the first cache stores the K-th level base address according to the K-th level base address; When it is judged that the i-th level of the first cache stores the K-th level base address, sending the K-th level base address to the computing unit of the i-th level of the first cache; The computing unit of the i-th level of the first cache is used for: Determining the (K + 1)-th level base address according to the K-th level base address and the K-th level offset address, where the K-th level offset address is determined according to the K-th level high-order address.
2. The device according to claim 1, characterized in that, The control unit of the i-th level of the first cache is further used for: When it is judged that the i-th level of the first cache does not store the K-th level base address and i ≠ N, sending the memory access request to the control unit of the (i + 1)-th level of the first cache.
3. The device according to claim 1 or 2, characterized in that, The device further includes a home node; the home node is coupled to the i-th level of the first cache; The control unit of the i-th level of the first cache is further used for: When it is judged that the i-th level of the first cache does not store the K-th level base address, sending the memory access request to the home node; The home node is used for receiving the memory access request to identify the level of the K-th level base address based on the memory access request.
4. The device according to claim 3, characterized in that, The home node is further used for: After receiving the memory access request, judging whether each level of the N-level first cache stores the K-th level base address according to the K-th level base address; When it is judged that the target first cache stores the K-th level base address, sending the memory access request to the control unit of the target first cache.
5. The device according to claim 4, characterized in that The device further includes: a memory controller coupled to the home node, and a memory coupled to the memory controller; the home node includes a buffer; The home node is further used for: When it is judged that each level of the first cache does not store the K-th level base address, sending the memory access request to the memory controller; The memory controller is used for: Determining the K-th level base address in the memory according to the K-th level base address; Sending the K-th level base address to the home node; The memory is used for storing the K-th level base address; The buffer is used for determining the (K + 1)-th level base address according to the K-th level base address and the K-th level offset address.
6. The device according to claim 4, characterized in that The device further includes a second cache coupled to the home node, and the second cache includes the control unit of the second cache and the computing unit of the second cache; The home node is further used for: After receiving the memory access request, judging whether the second cache stores the K-th level base address; When it is determined that the second cache stores the Kth - level base address, send the memory access request to the control unit of the second cache.
7. The device according to claim 6, characterized in that, The device further includes a memory controller coupled to the home node and a memory coupled to the memory controller; the home node includes a buffer buffer. The home node is further configured to: When it is determined that neither each level of the first cache nor the second cache stores the Kth - level base address, send the memory access request to the memory controller. The memory controller is configured to: Determine the Kth - level base address in the memory according to the Kth - level base address. Send the Kth - level base address to the home node. The memory is used to store the Kth - level base address. The buffer is configured to determine the (K + 1)th - level base address according to the Kth - level base address and the Kth - level offset address.
8. The device according to any one of claims 1-7, characterized in that, The device further includes a memory management unit coupled to the first - level first cache, and the memory management unit includes a third cache; the third cache is used to store the Kth - level base address, where K takes one or more values from 1, 2, …, M. The memory management unit is configured to: Before sending the memory access request to the first - level first cache, determine whether the third cache stores the first - level base address. When it is determined that the third cache stores the first - level base address, obtain the first - level base address. When it is determined that the third cache does not store the first - level base address, send the memory access request to the first - level first cache.
9. The device according to any one of claims 1-8, characterized in that, The memory access request further includes a base - address identifier of the Kth - level base address, and the base - address identifier of the Kth - level base address is used to indicate the level of the Kth - level base address.
10. A method for accelerating hardware page table traversal, characterized in that, A device applied to accelerate the traversal of the hardware page table. The device includes N - level first caches; each level of the N - level first caches includes a computing unit and a control unit; the N - level first caches store one or more levels of the M - level base addresses, where M is an integer greater than 1; among them, the ith - level first cache in the N - level first caches includes the control unit of the ith - level first cache and the computing unit of the ith - level first cache; i = 1, 2, …, N, and N is a positive integer; the method includes: Receive a memory access request through the control unit of the ith - level first cache, where the memory access request includes the Kth - level base address and the Kth - level high - order address, and K is a positive integer less than or equal to M. According to the Kth - level base address, determine whether the ith - level first cache stores the Kth - level base address through the control unit of the ith - level first cache. When it is determined that the ith - level first cache stores the Kth - level base address, send the Kth - level base address to the computing unit of the ith - level first cache through the control unit of the ith - level first cache. Determine the (K + 1)th - level base address according to the Kth - level base address and the Kth - level offset address through the computing unit of the ith - level first cache, where the Kth - level offset address is determined according to the Kth - level high - order address.
11. The method according to claim 10, wherein The method further includes: In the case where it is determined that the i-th level of the first cache does not store the K-th level base address and i≠N, the control unit of the i-th level of the first cache sends the memory access request to the control unit of the (i + 1)-th level of the first cache.
12. The method according to claim 10 or 11, characterized in that, The method further includes: In the case where it is determined that the i-th level of the first cache does not store the K-th level base address, the control unit of the i-th level of the first cache sends the memory access request to the home node.
13. The method according to claim 12, characterized in that, The method further includes: After receiving the memory access request, the home node determines whether each level of the N-level first cache stores the K-th level base address according to the K-th level base address; In the case where it is determined that the target first cache stores the K-th level base address, the memory access request is sent to the control unit of the target first cache.
14. The method according to claim 13, wherein The device further includes: a memory controller coupled to the home node, and a memory coupled to the memory controller; the home node includes a buffer buffer; the method further includes: In the case where it is determined that each level of the first cache does not store the K-th level base address, the home node sends the memory access request to the memory controller; The memory controller determines the K-th level base address in the memory according to the K-th level base address; The memory controller sends the K-th level base address to the home node; According to the K-th level base address and the K-th level offset address, the (K + 1)-th level base address is determined through the buffer buffer of the home node.
15. The method according to claim 13, wherein The device further includes a second cache coupled to the home node, and the second cache includes a control unit of the second cache and a computing unit of the second cache; the method further includes: After receiving the memory access request, the home node determines whether the second cache stores the K-th level base address; In the case where it is determined that the second cache stores the K-th level base address, the home node sends the memory access request to the control unit of the second cache.
16. The method according to claim 15, wherein The device further includes: a memory controller coupled to the home node, and a memory coupled to the memory controller; the home node includes a buffer buffer; the method further includes: In the case where it is determined that each level of the first cache and the second cache do not store the K-th level base address, the home node sends the memory access request to the memory controller; According to the K-th level base address, the memory controller determines the K-th level base address in the memory; The memory controller sends the K-th level base address to the home node; According to the K-th level base address and the K-th level offset address, the (K + 1)-th level base address is determined through the buffer buffer of the home node.
17. The method according to any one of claims 10 to 16, characterized in that, The device further includes a memory management unit coupled to the first cache of the first level, and the memory management unit includes a third cache; the third cache is used to store the K-th level base address, where K takes one or more values from 1, 2,..., M; the method further includes: Before sending the memory access request to the first cache of the first level, the memory management unit determines whether the third cache stores the first level base address; When it is determined that the third cache stores the first-level base address, obtain the first-level base address through the memory management unit; When it is determined that the third cache does not store the first-level base address, send the memory access request to the first-level first cache through the memory management unit.
18. The method according to any one of claims 10-17, characterized in that, The memory access request further includes a base address identifier of the K-level base address, and the base address identifier of the K-level base address is used to indicate the level of the K-level base address.
19. A chip system, characterized in that, The chip system is implemented by executing the method according to any one of claims 10-18.
Citation Information
Patent Citations
Apparatus and Method for Accelerated Hardware Page Table Walk
US20120331265A1
Selective prefetching of physically sequential cache line to cache line that includes loaded page table entry
US20150309936A1