A performance evaluation method, device, system, storage medium and equipment
Patent Information
- Application Number
- CN202210307791.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-03-25
AI Technical Summary
然而,在通过fio等工具进行测试时,会经过操作系统复杂的文件系统栈,而文件系统栈各个层均会对测试结果产生影响,从而引入许多不确定因素,以致测试结果不准确
[0010] In the above technical solution, since the embodiments of this specification initiate memory read and write operations by driving the FTE device, and since the programmable test engine device does not need to go through the complex file system stack of the operating system to access memory, the embodiments of this specification evaluate the collected performance data that can directly reflect the translation performance of the SMMU, rather than indirectly, thereby reducing the impact of many uncertain factors. Therefore, compared with methods using operating system user-space tools such as fio, the performance evaluation method of the system memory management unit proposed in the embodiments of this specification is more accurate.
Smart Images

Figure CN114691461B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of system testing technology, and in particular to a performance evaluation method, apparatus, system, storage medium, and device. Background Technology
[0002] In a CPU (Central Processing Unit) chip, the System Memory Management Unit (SMMU) is a crucial hardware module. Its function is to translate the virtual addresses of external devices into physical addresses, enabling these devices to access physical memory. Therefore, in CPU chip design, the performance of the SMMU directly impacts the performance of external devices, and consequently, the overall performance of the CPU chip.
[0003] However, in the current CPU chip design and verification process, there is a lack of effective and accurate methods for evaluating the performance of the SMMU. Currently, the commonly used method for evaluating SMMU performance by those skilled in the art is to use operating system user-space tools, such as fio, to test the read and write performance of the SMMU's backend peripheral disks, thereby determining the overall performance of the SMMU. However, when testing with tools like fio, the test involves navigating the complex file system stack of the operating system, and each layer of the file system stack affects the test results, introducing many uncertainties and leading to inaccurate results. Summary of the Invention
[0004] To overcome the problems existing in the related technologies, this specification provides a performance evaluation method, apparatus, system, storage medium, and device to address the deficiencies in the related technologies.
[0005] According to a first aspect of the embodiments of this specification, a performance evaluation method is provided, the method comprising: The programmable test engine device sends a memory access request to the system memory management unit, wherein the programmable test engine device is an external device connected to a computer device that loads the system memory management unit; The system memory management unit collects performance data generated when translating the memory address pointed to by the memory access request; The performance of the system memory management unit is evaluated based on the performance data.
[0006] According to a second aspect of the embodiments of this specification, a performance evaluation apparatus is provided, the apparatus comprising: A driver module is used to drive a programmable test engine device to send a memory access request to the system memory management unit, wherein the programmable test engine device is an external device connected to a computer device that loads the system memory management unit; The monitoring module is used to collect performance data generated when the system memory management unit translates the memory address pointed to by the memory access request; An evaluation module is used to evaluate the performance of the system memory management unit based on the performance data.
[0007] According to a third aspect of the embodiments of this specification, a performance evaluation system is provided, the system comprising: A physically connected computer device and a programmable test engine device; the computer device is equipped with a system memory management unit; The computer device is used for: The programmable test engine device sends a memory access request to the system memory management unit. The system memory management unit collects performance data generated when translating the memory address pointed to by the memory access request; The performance of the system memory management unit is evaluated based on the performance data.
[0008] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the above embodiments.
[0009] According to a fifth aspect of the embodiments of this specification, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in any of the above embodiments.
[0010] In the above technical solution, since the embodiments of this specification initiate memory read and write operations by driving the FTE device, and since the programmable test engine device does not need to go through the complex file system stack of the operating system to access memory, the embodiments of this specification evaluate the collected performance data that can directly reflect the translation performance of the SMMU, rather than indirectly, thereby reducing the impact of many uncertain factors. Therefore, compared with methods using operating system user-space tools such as fio, the performance evaluation method of the system memory management unit proposed in the embodiments of this specification is more accurate.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit the embodiments of this specification. Attached Figure Description
[0012] Figure 1This is a schematic diagram of a test procedure for testing SMMU performance using fio, as shown in one embodiment of this specification.
[0013] Figure 2 This is a flowchart illustrating a performance evaluation method for a system memory management unit according to an embodiment of this specification.
[0014] Figure 3 This is a block diagram of an evaluation apparatus for a system memory management unit according to an embodiment of this specification.
[0015] Figure 4 This is a schematic diagram of the architecture of a performance evaluation device for a system memory management unit, as shown in one embodiment of this specification.
[0016] Figure 5 This is a schematic diagram of the architecture of another system memory management unit performance evaluation device shown in one embodiment of this specification.
[0017] Figure 6 This is a block diagram of an evaluation apparatus for another system memory management unit according to an embodiment of this specification.
[0018] Figure 7 This is a schematic diagram of the architecture of another system memory management unit performance evaluation device shown in one embodiment of this specification.
[0019] Figure 8 This is a block diagram of an evaluation apparatus for another system memory management unit according to an embodiment of this specification.
[0020] Figure 9 This is a schematic diagram of the architecture of another system memory management unit performance evaluation device shown in one embodiment of this specification.
[0021] Figure 10 This is a schematic diagram of the architecture of a performance evaluation system for a system memory management unit, as shown in one embodiment of this specification.
[0022] Figure 11 This is a schematic diagram of the structure of a computing device hardware according to an embodiment of this specification. Detailed Implementation
[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those described in this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the embodiments described in this specification as detailed in the appended claims.
[0024] The terminology used in the embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the embodiments of this specification. The singular forms “a,” “described,” and “the” as used in the embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0025] It should be understood that although the terms first, second, third, etc., may be used to describe various information in the embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of the embodiments of this specification, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0026] Fio is a user-space tool in an operating system that allows for various read and write operations on the disk, thus testing the read and write performance of the disk and file system. In CPU chip design, the disk is a back-end peripheral of the SMMU (System Memory Management Unit). When the SMMU is enabled, the disk initiates read and write requests through the SMMU. Therefore, using fio can reflect SMMU performance to some extent. However, this method of testing SMMU performance using fio has significant drawbacks.
[0027] like Figure 1 As shown, Figure 1This specification provides an embodiment illustrating a test process for testing SMMU performance using fio. Taking the ext4 (Fourth extended filesystem) filesystem as an example, it describes the entire process of making read or write requests to an NVMe disk via fio. The NVMe disk refers to a disk with an interface specification and transport protocol of NVM Express (Non-Volatile Memory Express), and is a backend peripheral of the SMMU, initiating read and write requests through the SMMU. First, fio calls user-mode read / write system calls from the operating system's user space, then sequentially passes through the vfs (Virtual File Systems) layer, the filesystem layer, and the block layer, ultimately writing data to the NVMe disk. If the filesystem currently managing the disk is ext4, the vfs layer's read / write system call interface calls the ext4 filesystem layer's interface. Upon entering the block layer, the system begins submitting data and writing methods to the disk driver. The NVMe disk driver then invokes DMA (Direct Memory Access) to allocate memory and establish I / O page tables. At this point, the NVMe disk hardware uses the returned I / O virtual address to request address translation from the SMMU. Only after the SMMU returns the physical address is the data handed over to the NVMe disk for reading and writing, thus concluding the entire test process. It's worth noting that during this process, the SMMU's backend peripherals are not limited to just an NVMe disk; for example... Figure 1 In the process of NVMe disk requesting address translation, the network interface card (NIC) also makes address translation requests. Therefore, the computing resources of SMMU are also shared by the NIC.
[0028] The above method of using fio to test SMMU performance has the following drawbacks: First, this method results in a very deep call stack with too many influencing factors, which severely impacts the accuracy of SMMU performance evaluation. For example... Figure 1 As shown, fio read and write operations need to pass through the VFS layer, file system layer, and block layer in sequence. These intermediate layers are affected by other software modules of the operating system, such as operating system scheduling, memory allocation, and file system I / O, which further affects the performance accuracy of fio tests. Therefore, the test method of using fio for user-space read and write and then checking the file system I / O performance is affected by multiple layers, has a large test granularity, and ultimately leads to an inaccurate evaluation of SMMU performance.
[0029] Second, this approach lacks fine-grained isolation strategies at the SMMU level, thus failing to accurately reflect SMMU performance. For example... Figure 1 As shown, since the SMMU may have multiple backend peripherals at the same time, such as network cards and disks, if the network card also sends a translation request to the SMMU when the fio is reading and writing to the NVMe disk, the read and write operations on the disk initiated by the fio will be affected by the SMMU's response to the network card's translation request. Therefore, the disk read and write speed obtained by using the fio test cannot accurately reflect the speed of the SMMU's response to the disk address translation request, and thus cannot accurately evaluate the performance of the SMMU.
[0030] Third, this method only displays file read and write speeds, not SMMU performance. Since the evaluation method using the fio tool only shows the file system I / O rate and doesn't reveal any information about SMMU performance, this method is essentially insufficient for directly evaluating SMMU performance from a hardware perspective.
[0031] Because the performance evaluation methods of SMMU currently used in this field have the problems mentioned above, there is a lack of an effective and accurate performance evaluation method for SMMU in the current CPU chip design verification process.
[0032] In response, this specification proposes a performance evaluation method for a system memory management unit (SMMU), which can effectively and accurately evaluate the performance of the SMMU.
[0033] like Figure 2 As shown, Figure 2 This is a flowchart illustrating a performance evaluation method for a system memory management unit according to an embodiment of this specification, including the following steps: Step S201: Drive the programmable test engine device to send a memory access request to the system memory management unit; Step S202: Collect performance data generated when the system memory management unit translates the memory address pointed to by the memory access request; Step S203: Evaluate the performance of the system memory management unit based on the performance data.
[0034] In step S201, the Fabric Test Engine (FTE) is a back-end peripheral of the SMMU, connected to the computer device to which the SMMU belongs, and can initiate memory read / write operations. When the FTE initiates a memory read / write operation, the memory address pointed to by its memory read / write request is actually a virtual address. Therefore, the SMMU will respond to the memory read / write request to translate the virtual memory address into a physical memory address, so that the memory read / write operation can read the correct data from memory or write data to the correct location in memory.
[0035] Since the memory read and write requests initiated by FTE directly point to memory after the address is translated by SMMU, without needing to go through multiple layers of calls such as the VFS layer, file system layer and block layer, the performance of SMMU tested using FTE is less affected by other modules of the operating system, and its test results have higher accuracy.
[0036] Furthermore, in some embodiments, a specific device isolation method can be used to enable the SMMU to collect only the read and write traffic of the FTE, thereby reducing the impact of address translation requests initiated by other backend peripherals, such as network cards and disks, on the performance evaluation of the SMMU, and further improving the accuracy of the performance evaluation of the SMMU.
[0037] In a CPU chip, the SMMU can also be set to passthrough mode. When the SMMU is set to passthrough mode, it will not perform any operation on the addresses sent by external devices; that is, it will not perform the operation of translating virtual addresses into physical addresses, but will directly use the received addresses as physical addresses. In this case, the SMMU will not look up any page tables.
[0038] In some embodiments, the specific device isolation method described above may involve configuring the SMMU's response mode for memory access requests from all external devices other than the FTE, such as NVMe disks and network cards, connected to the computer device to which the SMMU belongs, to pass-through mode when it detects that other external devices are connected to the SMMU. After using this specific device isolation method, external devices such as NVMe disks and network cards will be in pass-through mode for the SMMU. That is, for the memory addresses of memory access requests from these external devices, the SMMU will not perform address translation operations, but will directly use them as physical addresses for further processing. Meanwhile, the FTE remains in its normal operating state, meaning its memory access operations can initiate address translation requests to the SMMU, and the SMMU will translate the memory address pointed to by the FTE's memory access request from a virtual address to a physical address. Therefore, by using the device isolation method described above, only the FTE can request address translation from the SMMU among all external devices connected to the computer. Thus, when the SMMU is enabled, it will only receive address translation requests from the FTE. Therefore, in this state, the translation information inside the SMMU will only record read and write requests from the FTE, thereby eliminating the influence of other external devices and further improving the accuracy of the SMMU's performance evaluation.
[0039] However, due to framework limitations, current Linux operating systems, when setting the SMMU to pass-through mode, do not differentiate between different external devices, preventing them from operating in different modes. Instead, all external devices behind the SMMU are set to pass-through mode. To address this, and to achieve more precise evaluation, a logic module can be added to the operating system kernel to allow the SMMU to be set to device-level pass-through mode. This would allow different external devices to operate in different modes: the FTE operates in normal mode, while other external devices operate in pass-through mode, achieving device isolation.
[0040] In step S202, when the SMMU translates the memory address pointed to by the memory access request of the FTE, it generates and records the corresponding execution data. This data can reflect the performance of the SMMU to a certain extent. Therefore, this data can be used as the performance data of the SMMU for the performance evaluation of the SMMU.
[0041] Specifically, SMMU performance data can include SMMU response data for FTE read / write requests. Specifically, this data can include the number of address translation requests (REQ) performed by the SMMU during the response to the FTE's memory access request, and the total time (TIME) for the SMMU to respond to the FTE's memory access request. By using the number of translation requests (REQ) completed by the SMMU and the time (TIME) consumed in completing these requests, the QPS (Queries Per Second) of translation requests completed by the SMMU can be calculated, serving as a relatively detailed performance metric for evaluating the SMMU's performance.
[0042] The formula for calculating QPS is as follows: Specifically, the higher the QPS of the SMMU, the more address translation requests the SMMU processes per unit of time, and the better the performance of the SMMU; conversely, the lower the QPS of the SMMU, the fewer address translation requests the SMMU processes per unit of time, and the worse the performance of the SMMU.
[0043] The SMMU contains two components: the TBU (Translation Buffer Unit) and the TCU (Translation Control Unit). The TBU caches page tables and the TLB (Translation Lookaside Buffer), while the TCU controls and manages address translation. During an address translation operation, the SMMU sequentially searches the STE table, CD table, and multiple page tables to translate the virtual address into the corresponding physical address. After completing an address translation operation, the SMMU caches the STE, CD, and page tables retrieved during that translation as historical translation data in the TBU. This allows the SMMU to first search the historical translation data for the corresponding entries in the TBU when an external device initiates another address translation request. If the SMMU can find the corresponding entry in the historical translation data, it avoids searching for the entry in memory, thus speeding up the lookup process. When the SMMU cannot find the corresponding entry in the historical translation data, the SMMU then controls the TCU to search for the corresponding entries in the ste table, cd table, and page table in memory in sequence.
[0044] In particular, the TLB in TBU can be used to cache the mapping relationship between virtual addresses and physical addresses. Therefore, when the memory address pointed to by the memory access request of an external device is recorded in the TLB, the physical address translated from that memory address can be quickly obtained, further speeding up the lookup. Therefore, TLB is also known as a fast table.
[0045] When SMMU performs a complete address translation operation, if cached historical translation data exists in TBU, it will directly search TBU for historical translation data corresponding to the current address translation request. If the corresponding historical translation data is found, it will directly use the historical translation data to perform the current address translation operation. If the corresponding historical translation data is not found, the TCU will then control the execution of the address translation operation, that is, search for the corresponding entries in the ste table, cd table, and page table in memory, thereby translating the address pointed to by the address translation request from a virtual address to a physical address.
[0046] Clearly, the data generated by the TBU and TCU during the address translation operation can also reflect the performance of the SMMU to some extent.
[0047] In some embodiments, the SMMU performance data may also include TBU performance data and TCU performance data.
[0048] Specifically, TBU performance data may include one or more of the following: Execution time cycles; The number of fast table misses is tlb_miss; Data transmission and reception count (transaction).
[0049] Specifically, the TCU's performance data may include one or more of the following: The number of cache misses is config_cache_miss; Read / write request access count config_struct_access; Execution time cycles; The number of times the fast table misses (tlb_miss); The number of page table lookups during translation: trans_table_walk_access; Data transmission and reception count (transaction).
[0050] In step S203, the performance of the SMMU can be evaluated based on the performance data collected by the SMMU when the corresponding FTE makes a memory access request.
[0051] For example, in some embodiments, the number of translation requests completed by the SMMU (Recorded Number of Translation Requests) REQ recorded in step S202 and the time TIME consumed by the SMMU to complete the translation operation can be used to calculate the number of translation requests completed by the SMMU (Recorded Number of Translation Requests) QPS. The SMMU's performance can then be evaluated based on the QPS. Specifically, the higher the SMMU's QPS, the more address translation requests the SMMU processes per unit time, and the better the SMMU's performance; conversely, the lower the SMMU's QPS, the fewer address translation requests the SMMU processes per unit time, and the worse the SMMU's performance. The calculation method for QPS can be found above and will not be repeated here.
[0052] For example, in some embodiments, the detailed data of the TBU and TCU recorded in step S202 can be displayed to reflect the performance of the SMMU through the real-time data of the TBU and TCU.
[0053] Furthermore, in some embodiments, historical performance data of TBU and TCU can be collected and recorded in step S202, and then the historical performance data of TBU and TCU can be compared and analyzed in step S203. For example, the highest value, lowest value or average value in the historical performance data of TBU and TCU can be calculated, and the performance of SMMU can be reflected from the long-term feedback performance data of TBU and TCU.
[0054] Furthermore, since historical translation data is cached in the SMMU, the SMMU's response speed to the FTE's memory access request varies significantly depending on whether historical translation data corresponding to the memory address pointed to by the FTE's memory access request exists in the cache.
[0055] Specifically, when the SMMU cache contains historical translation data corresponding to the memory address pointed to by the FTE's memory access request, the SMMU can directly complete the address translation operation based on the historical translation data. Therefore, the SMMU takes the shortest time to complete the address translation operation at this time, and its performance is the best.
[0056] When the SMMU's cache does not contain historical translation data corresponding to the memory address pointed to by the FTE's memory access request, the SMMU needs to search for the corresponding entries in the ste table, cd table, and page table in memory and load these entries into the cache. Therefore, the SMMU takes the longest time to complete the address translation operation at this time, and its performance is the worst.
[0057] In real-world business scenarios, SMMU may encounter both the best-case and worst-case scenarios when performing address translation operations.
[0058] In some embodiments, in order to more comprehensively demonstrate the performance of the SMMU, the best and worst performance of the SMMU can be simulated separately, thereby more accurately reflecting the performance of the SMMU.
[0059] In some embodiments, the SMMU cache can be cleared each time the FTE initiates a memory access request. This means deleting all cached entries in the SMMU, such as the ste table, cd table, and page table, as well as the TLB. This ensures that regardless of the memory address pointed to by the FTE's memory access request, the corresponding entries in the ste table, cd table, and page table must be searched sequentially in memory and loaded into the cache before the FTE's address translation request can be completed. This simulates the worst-case scenario for SMMU performance and tests the SMMU's performance under the worst-case conditions.
[0060] Specifically, before collecting performance data generated when the SMMU translates the memory address pointed to by the FTE's memory access request, the SMMU can be notified to check whether there is cached historical translation data. If historical translation data exists, the SMMU can be notified to clear all historical translation data before performing the address translation operation; if it does not exist, the SMMU can be notified to perform the address translation operation directly.
[0061] In some embodiments, after the FTE initiates a memory access request, the cache refresh may not be performed by default, allowing the SMMU to directly perform address translation operations using historical translation data in the cache, i.e., to execute a fast translation path, to simulate the best performance of the SMMU, thereby testing the performance of the SMMU under the best conditions.
[0062] Specifically, after collecting the memory access request initiated by the FTE and translating it, the SMMU can be notified to check whether there is historical translation data associated with the memory address pointed to by the memory access request of the FTE. If there is historical translation data associated with the memory address, the SMMU can be notified to obtain the historical translation data associated with the memory address and complete the address translation operation based on the historical translation data. If there is no historical translation data, the SMMU can be notified to perform the address translation operation.
[0063] In some embodiments, the best and worst performance of the SMMU can be simulated separately, and performance data under the best and worst conditions can be collected separately to evaluate the performance of the SMMU under two different conditions.
[0064] For example, in some embodiments, the best-performing SMMU scenario can be simulated, and the number of address translation requests completed by the SMMU and the time consumed in completing these requests can be collected to calculate the QPS value of the SMMU under the best performance condition, i.e., max_qps. Conversely, the worst-performing SMMU scenario can be simulated, and the number of address translation requests completed by the SMMU and the time consumed in completing these requests can be collected again to calculate the QPS value of the SMMU under the worst performance condition, i.e., min_qps. Based on the QPS obtained in both scenarios, a detailed QPS range, min_qps ~ max_qps, can be derived. This QPS range can also reflect the SMMU's performance to some extent.
[0065] It is worth noting that when calculating the QPS interval of the SMMU, you can calculate max_qps first and then min_qps, or you can calculate min_qps first and then max_qps. The embodiments in this specification do not restrict the calculation order of max_qps and min_qps.
[0066] Similarly, in some embodiments, the performance data of TBU and TCU under the best SMMU performance and the worst SMMU performance can be obtained separately, thereby reflecting more detailed SMMU performance data.
[0067] Corresponding to the performance evaluation method embodiment of the system memory management unit described above, this specification embodiment also provides a performance evaluation device for the system memory management unit.
[0068] like Figure 3 As shown, Figure 3 This is a block diagram of a performance evaluation device for a system memory management unit according to an embodiment of this specification, comprising the following modules: Driver module 310 is used to drive the programmable test engine device to send memory access requests to the system memory management unit; Monitoring module 320 is used to collect performance data generated when the system memory management unit translates the memory address pointed to by the memory access request; Evaluation module 330 is used to evaluate the performance of the system memory management unit based on performance data.
[0069] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0070] Specifically, such as Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the architecture of a system memory management unit performance evaluation device according to an embodiment of this specification. Here, SPT (SMMU Performance Test) refers to an embodiment of the system memory management unit performance evaluation device according to this specification. It should be understood that... Figure 4 The drive module 310 and monitoring module 320 shown, as well as the evaluation module 330 (not shown), can be either Figure 3 The corresponding module in the illustrated device embodiment may also be the corresponding module in other device embodiments. Figure 3 The illustrated embodiments and Figure 4 The embodiments shown may be the same embodiment or different embodiments.
[0071] Specifically, such as Figure 4 As shown, SPT includes a user space portion and a kernel space portion.
[0072] The user space part of SPT is mainly responsible for initiating memory read and write requests directly by calling the interface of the driver module 310 to drive the underlying FTE peripheral. When the read and write request is completed, it calculates and prints or displays the QPS of SMMU based on the feedback from the monitoring module 320. In addition, SPT will print or display various performance data that can reflect the performance of SMMU, such as the number of translation requests completed by SMMU, the time consumed by SMMU to complete the translation request, the performance data of TBU, and the performance data of TCU.
[0073] The kernel space of SPT is mainly responsible for executing the functions of each module.
[0074] Specifically, the driver module 310 can be set in the kernel space of SPT to control the external device FTE to initiate memory access requests. When the FTE initiates a memory access request, the computer device will send the memory address pointed to by the memory access request of the FTE to the SMMU, and the SMMU will translate the memory address from a virtual address to a physical address.
[0075] Specifically, the monitoring module 320 can also be set in the kernel space of SPT to monitor and collect real-time performance data of SMMU. The monitoring module 320 can send the collected SMMU performance data to the evaluation module 330 for further analysis, calculation and evaluation. Finally, the evaluation results of SMMU and various performance data are displayed in the user space screen or printed out.
[0076] Specifically, the performance data of the SMMU collected by the monitoring module 320 may include the SMMU's response to the read and write requests of the FTE. Specifically, the SMMU's response to the read and write requests of the FTE may include the number of address translation requests (REQ) performed by the SMMU in the process of responding to the memory access requests of the FTE, and the total duration (TIME) of the SMMU's response to the memory access requests of the FTE.
[0077] In some embodiments, such as Figure 5 As shown, Figure 5 This is a schematic diagram of the architecture of another system memory management unit performance evaluation device according to an embodiment of this specification. The SMMU includes a TBU and a TCU. When the FTE initiates an address translation request operation, the SMMU first checks the TBU for matching cached historical translation data. If it exists, the address translation operation is performed directly based on that historical translation data; otherwise, the address translation operation is performed through the TCU. Specifically, the performance data of the SMMU collected by the monitoring module 320 may also include the performance data of the TBU and TCU, to make the performance evaluation of the SMMU more comprehensive and accurate.
[0078] Specifically, the evaluation module 330 was not in Figure 4 As shown in the figure. In some embodiments, the evaluation module 330 can be set in the kernel space of the SPT for calculating and analyzing the performance data of the SMMU; in some embodiments, the evaluation module 330 can also be set in the user space of the SPT for displaying the evaluation results and various performance data of the SMMU; in some embodiments, the evaluation module 330 can also be set in both the kernel space and the user space. It should be understood that the specific setting method of the evaluation module 330 should be determined according to the actual evaluation method of the SMMU, and the embodiments in this specification do not limit it in this way.
[0079] In some embodiments, SPT can also use specific device isolation methods to enable SMMU to collect only the read and write traffic of FTE, thereby reducing the impact of address translation requests initiated by other backend peripherals, such as network cards and disks, on the performance evaluation of SMMU and further improving the accuracy of SMMU performance evaluation.
[0080] like Figure 6 As shown, Figure 6 This is a block diagram of a performance evaluation apparatus for another system memory management unit according to an embodiment of this specification, comprising the following modules: Driver module 310 is used to drive the programmable test engine device to send memory access requests to the system memory management unit; Monitoring module 320 is used to collect performance data generated when the system memory management unit translates the memory address pointed to by the memory access request; Evaluation module 330 is used to evaluate the performance of the system memory management unit based on performance data; The isolation module 340 is configured to set the response mode of the system memory management unit to the memory access request of the other external device to a pass-through mode when it is detected that the computer device is connected to other external devices.
[0081] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0082] like Figure 7 As shown, Figure 7 This is a schematic diagram of the architecture of another system memory management unit performance evaluation device according to an embodiment of this specification. In this device, the SMMU is configured to pass-through mode for memory access requests from all external devices except the FTE, such as NVMe disks and network cards. Therefore, among all external devices connected to the computer, only the FTE can request address translation from the SMMU. Thus, when the SMMU is enabled, it will only receive address translation requests from the FTE. Therefore, in this state, the translation information inside the SMMU will only record read and write requests from the FTE, thereby eliminating the influence of other external devices and further improving the accuracy of the SMMU performance evaluation.
[0083] In particular, Figure 7 In the illustrated embodiment, an isolation module 340 may be configured in the kernel space of the SPT. Figure 7 (Not shown in the diagram) is used to control the SMMU, enabling the SMMU to be configured in pass-through mode for all external devices except the FTE, thereby isolating memory access requests from other external devices. In some embodiments, the SPT may not have an isolation module 340, but instead use other methods outside the device to configure the SMMU in pass-through mode for all external devices except the FTE, thereby isolating memory access requests from other external devices.
[0084] In some embodiments, SPT can also control the SMMU to clear its cache or not clear its cache before the SMMU responds to the address translation request of the FTE, thereby simulating the worst-case performance of the SMMU and the best-case performance of the SMMU, respectively, to collect performance data of the SMMU under the worst-case performance and the best-case performance of the SMMU.
[0085] like Figure 8As shown, Figure 8 This is a block diagram of a performance evaluation apparatus for another system memory management unit according to an embodiment of this specification, comprising the following modules: Driver module 310 is used to drive the programmable test engine device to send memory access requests to the system memory management unit; Monitoring module 320 is used to collect performance data generated when the system memory management unit translates the memory address pointed to by the memory access request; Evaluation module 330 is used to evaluate the performance of the system memory management unit based on performance data; The control module 350 is used to control the system memory management system to clear the cache before each address translation operation, or not to clear the cache.
[0086] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0087] In some embodiments, the control module 350 can control the SMMU to clear the cache each time the FTE initiates a memory access request. That is, all entries in the cached ste table, cd table, page table, and TLB in the SMMU are deleted. This ensures that no matter what memory address the FTE's memory access request points to, the corresponding entries in the ste table, cd table, and page table need to be searched in memory in turn, and the found entries in the ste table, cd table, and page table are loaded into the cache before the FTE's address translation request can be completed. This simulates the worst-case scenario of SMMU performance and tests the performance of SMMU under the worst-case condition.
[0088] Specifically, before collecting performance data generated when the SMMU translates the memory address pointed to by the memory access request of the FTE, the control module 350 can notify the SMMU to check whether there is cached historical translation data. If historical translation data exists, the control module 350 can notify the SMMU to clear all historical translation data before performing the address translation operation. If historical translation data does not exist, the control module 350 can notify the SMMU to perform the address translation operation directly.
[0089] In some embodiments, the control module 350 may also not perform a cache refresh operation by default after the FTE initiates a memory access request, so that the SMMU can directly perform address translation operations through historical translation data in the cache, that is, execute a fast translation path to simulate the best performance of the SMMU, thereby testing the performance of the SMMU under the best conditions.
[0090] Specifically, after the control module 350 collects the memory access request initiated by the FTE and the SMMU translates it, it can notify the SMMU to check whether there is historical translation data associated with the memory address pointed to by the memory access request of the FTE. If there is historical translation data associated with the memory address, the control module 350 can notify the SMMU to obtain the historical translation data associated with the memory address and complete the address translation operation based on the historical translation data. If there is no historical translation data, the control module 350 can notify the SMMU to perform the address translation operation.
[0091] like Figure 9 As shown, Figure 9 This is a schematic diagram of the architecture of another system memory management unit performance evaluation device according to an embodiment of this specification. The SPT kernel space includes a control module 350, which can be configured to either not refresh the cache or refresh the cache, respectively, to simulate the best and worst performance scenarios of the SMMU, and to collect performance data under both scenarios, thereby evaluating the SMMU performance under these two different conditions.
[0092] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the embodiments in this specification, depending on actual needs. Those skilled in the art can understand and implement this without creative effort.
[0093] This specification also provides an embodiment of a system memory management unit performance evaluation system, such as... Figure 10 As shown, Figure 10 This is a schematic diagram illustrating the architecture of a performance evaluation system for a system memory management unit according to an embodiment of this specification. The performance evaluation system for the system memory management unit includes: The computer device 1001 and the programmable test engine device (FTE) 1002 are physically connected, and the computer device 1001 is equipped with a system memory management unit (SMMU) 1003.
[0094] The FTE 1002 is used to issue memory access requests to the SMMU 1003 under the drive of the computer device 1001.
[0095] The computer device 1001 can perform the following steps: The programmable test engine device sends a memory access request to the system memory management unit. The system's memory management unit generates performance data when translating the memory address pointed to by a memory access request; The performance of the system memory management unit is evaluated based on performance data.
[0096] In the embodiments of this specification, the computer device 1001 can also be used to execute the steps of any of the above-described method embodiments of the performance evaluation method for the system memory management unit of the embodiments of this specification.
[0097] For example, in some embodiments, computer device 1001 may also be used to configure the response mode of SMMU 1003 to memory access requests from all external devices other than FTE 1002, such as NVMe disks and network cards, to pass-through mode when it detects that other external devices, such as NVMe disks and network cards, are connected to the computer device to which SMMU 1003 belongs.
[0098] For example, in some embodiments, the performance data of SMMU 1003 may include data on SMMU 1003's response to read / write requests from FTE 1002. In some embodiments, the data on SMMU 1003's response to read / write requests from FTE 1002 may include the number of address translation requests (REQ) performed by SMMU 1003 in responding to memory access requests from FTE 1002, and the total duration (TIME) of SMMU 1003's response to memory access requests from FTE 1002.
[0099] For example, in some embodiments, the computer device 1001 can also be used to clear the cache of the SMMU 1003 each time the FTE 1002 initiates a memory access request. That is, all entries such as the ste table, cd table, and page table, as well as the tlb, cached in the SMMU 1003 are deleted. This ensures that no matter what memory address the memory access request of the FTE 1002 points to, the corresponding entries of the ste table, cd table, and page table need to be searched in memory in turn, and the found entries of the ste table, cd table, and page table are loaded into the cache in order to complete the address translation request of the FTE 1002. This simulates the worst-case performance of the SMMU 1003 and tests the performance of the SMMU 1003 under the worst-case scenario.
[0100] For example, in some embodiments, the computer device 1001 may, before collecting performance data generated when the SMMU 1003 translates the memory address pointed to by the memory access request of the FTE 1002, notify the SMMU 1003 to check whether there is cached historical translation data. If historical translation data exists, the computer device 1001 may notify the SMMU 1003 to clear all historical translation data before performing the address translation operation; if historical translation data does not exist, the computer device 1001 may notify the SMMU 1003 to perform the address translation operation directly.
[0101] For example, in some embodiments, the computer device 1001 may also be used to not perform cache refresh by default after the FTE 1002 initiates a memory access request, so that the SMMU 1003 can directly perform address translation operations through historical translation data in the cache, that is, execute a fast translation path, to simulate the best performance of the SMMU 1003, thereby testing the performance of the SMMU 1003 under the best conditions.
[0102] For example, in some embodiments, after the computer device 1001 collects the memory access request initiated by the FTE 1002 and the SMMU 1003 translates it, it can notify the SMMU 1003 to query whether there is historical translation data associated with the memory address pointed to by the memory access request of the FTE 1002. If there is historical translation data associated with the memory address, the computer device 1001 can notify the SMMU 1003 to obtain the historical translation data associated with the memory address and complete the address translation operation based on the historical translation data. If there is no historical translation data, the computer device 1001 can notify the SMMU 1003 to perform the address translation operation.
[0103] In the embodiments of this specification, the implementation process of the functions and roles of the computer device 1001 is detailed in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0104] This specification also provides a computer device, which includes at least a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in any of the foregoing embodiments.
[0105] Figure 11 This diagram illustrates a more specific hardware structure of a computing device provided in an embodiment of this specification. The device may include: a processor 1101, a memory 1102, an input / output interface 1103, a communication interface 1104, and a bus 1105. The processor 1101, memory 1102, input / output interface 1103, and communication interface 1104 are interconnected internally via the bus 1105.
[0106] The processor 1101 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification. The processor 1101 may also include a graphics card, such as an Nvidia Titan X graphics card or a 1080Ti graphics card.
[0107] The memory 1102 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1102 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1102 and is called and executed by the processor 1101.
[0108] Input / output interface 1103 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.
[0109] The communication interface 1104 is used to connect the communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0110] Bus 1105 includes a pathway for transmitting information between various components of the device, such as processor 1101, memory 1102, input / output interface 1103, and communication interface 1104.
[0111] It should be noted that although the above-described device only shows the processor 1101, memory 1102, input / output interface 1103, communication interface 1104, and bus 1105, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0112] This specification also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the foregoing embodiments.
[0113] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0114] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the embodiments of this specification can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the embodiments of this specification, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this specification.
[0115] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.
[0116] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. When implementing the embodiments of this specification, the functions of each module can be implemented in one or more software and / or hardware. Alternatively, some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0117] The above description is merely a specific implementation of the embodiments of this specification. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of the embodiments of this specification, and these improvements and modifications should also be considered within the protection scope of the embodiments of this specification.
Claims
1. A performance evaluation method, the method comprising: The programmable test engine device sends a memory access request to the system memory management unit. The programmable test engine device is an external device connected to the computer device that loads the system memory management unit. The memory access request sent by the programmable test engine device is directly directed to memory after the address is translated by the system memory management unit. If other external devices are detected to be connected to the computer device, the system memory management unit is configured to pass through mode for responding to memory access requests from the other external devices; wherein, when the system memory management unit is configured to pass through mode for responding to memory access requests from a specified device, the system memory management unit does not perform a translation operation on the memory address pointed to by the memory access request from the specified device. The system memory management unit collects performance data generated when translating the memory address pointed to by the memory access request; The performance of the system memory management unit is evaluated based on the performance data.
2. The method according to claim 1, wherein the performance data includes the number of times the system memory management unit translates the memory address when responding to the memory access request, and the total duration of the system memory management unit responding to the memory access request.
3. The method according to claim 1, before collecting the performance data generated when the system memory management unit translates the memory address pointed to by the memory access request, further includes the step of: The system memory management unit is notified to check if there is cached historical translation data. If the historical translation data exists, the system memory management unit is notified to clear the historical translation data before performing the translation operation. If the historical translation data does not exist, the system memory management unit is notified to perform the translation operation.
4. The method according to claim 1, before collecting the performance data generated when the system memory management unit translates the memory address pointed to by the memory access request, further includes the step of: The system memory management unit is notified to query whether historical translation data associated with the memory address exists. If the historical translation data exists, the system memory management unit is notified to retrieve the historical translation data associated with the memory address. If the historical translation data does not exist, the system memory management unit is notified to perform a translation operation.
5. A performance evaluation apparatus, the apparatus comprising: A driver module is used to drive a programmable test engine device to send a memory access request to a system memory management unit. The programmable test engine device is an external device connected to a computer device that loads the system memory management unit. The memory access request sent by the programmable test engine device is directly directed to memory after the address is translated by the system memory management unit. An isolation module is configured to, when the computer device is detected to be connected to other external devices, configure the response mode of the system memory management unit to memory access requests from the other external devices to a pass-through mode; wherein, when the response mode of the system memory management unit to memory access requests from a specified device is configured to a pass-through mode, the system memory management unit does not perform a translation operation on the memory address pointed to by the memory access request from the specified device; The monitoring module is used to collect performance data generated when the system memory management unit translates the memory address pointed to by the memory access request; An evaluation module is used to evaluate the performance of the system memory management unit based on the performance data.
6. A performance evaluation system, the system comprising: A physically connected computer device and a programmable test engine device; the computer device is equipped with a system memory management unit; The programmable test engine device is used to send a memory access request to the system memory management unit under the drive of the computer device; The computer device is used for: The programmable test engine device sends a memory access request to the system memory management unit. The memory access request sent by the programmable test engine device is directly translated by the system memory management unit and points to memory. If other external devices are detected to be connected to the computer device, the system memory management unit is configured to pass through mode for responding to memory access requests from the other external devices; wherein, when the system memory management unit is configured to pass through mode for responding to memory access requests from a specified device, the system memory management unit does not perform a translation operation on the memory address pointed to by the memory access request from the specified device. The system memory management unit collects performance data generated when translating the memory address pointed to by the memory access request; The performance of the system memory management unit is evaluated based on the performance data.
7. The system according to claim 6, wherein the performance data includes the number of times the system memory management unit translates the memory address when responding to the memory access request, and the total duration of the system memory management unit responding to the memory access request.
8. The system according to claim 6, wherein the computer device is further configured to: Before collecting performance data generated when the system memory management unit translates the memory address pointed to by the memory access request, the system memory management unit is notified to check whether there is cached historical translation data. If the historical translation data exists, the system memory management unit is notified to clear the historical translation data before performing the translation operation. If the historical translation data does not exist, the system memory management unit is notified to perform the translation operation.
9. The system according to claim 6, wherein the computer device is further configured to: Before collecting performance data generated when the system memory management unit translates the memory address pointed to by the memory access request, the system memory management unit is notified to query whether there is historical translation data associated with the memory address. If the historical translation data exists, the system memory management unit is notified to obtain the historical translation data associated with the memory address. If it does not exist, the system memory management unit is notified to perform the translation operation.
10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1-4.
11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method of any one of claims 1-4.
Citation Information
Patent Citations
Memory management unit, processing unit, system and memory access method
CN114185817A
Method and apparatus for handling missing memory page abnomality, and device and storage medium
WO2022057749A1