Memory performance detection method, device and equipment and memory medium

By performing address mapping and delay testing between memory boxes and host nodes based on the peripheral component interconnect topology in a large server cluster, the problem of inconsistent memory performance of host nodes is solved, efficient scheduling and consistency detection of memory resources is achieved, and operation and maintenance costs are reduced.

CN120540926APending Publication Date: 2025-08-26INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510526074.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In large server clusters, the memory performance detection of each host node is inconsistent, resulting in difficulty in scheduling memory resources, and the existing technology has problems such as waste of resources and high operation and maintenance costs.

Method used

The address mapping between the memory box and the host node is converted into extended memory that the host node can recognize, and delay testing is performed to determine the consistency of memory performance, and memory pooling is used to schedule memory resources.

Benefits of technology

It realizes fast and effective detection of the memory performance of the host node, eliminates memory space inconsistency, improves the efficiency of memory resource scheduling and the effectiveness of consistency testing, and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540926A_ABST
    Figure CN120540926A_ABST
Patent Text Reader

Abstract

The invention discloses a memory performance detection method, device and equipment and a memory medium, and relates to the technical field of memory performance detection. Memory with the same capacity of a data path port of a memory box is mapped to an endpoint port, the mapped memory is converted into an extended memory which can be identified by a host node, and then the memory performance of the extended memory is subjected to a consistency test, so that a memory mapping relationship can be quickly and effectively established; according to the embodiment of the invention, the host nodes can schedule the memory capacity required by the host nodes in the memory box, and the inconsistency caused by different mapped memory spaces is eliminated by configuring extended memories with the same capacity for the host nodes, so that the memory performance of the plurality of host nodes tends to be consistent on the premise, and the effectiveness of the consistency test is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of server memory management technology, and in particular to a memory performance detection method, apparatus, device, and memory medium. Background Art

[0002] With the widespread deployment of large-scale server clusters and cloud computing infrastructure, hyperscale cloud service providers are facing memory capacity challenges brought about by exponential data growth. Related technologies can utilize memory pooling to schedule memory resources across multiple host nodes. However, this can lead to inconsistent memory latency across host nodes.

[0003] Therefore, how to effectively perform memory performance testing on each host node is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] This application provides a memory performance detection method, apparatus, device and memory medium to effectively perform memory performance detection on each host node.

[0005] This application provides a memory performance detection method, including:

[0006] For each of the plurality of host nodes, based on a preset peripheral component interconnect topology, performing address mapping between a target memory corresponding to a data path port of the memory cartridge and an endpoint port of the host node to obtain an address mapping result, and converting the target memory into an extended memory recognizable by the host node based on the address mapping result; the extended memory of each host node having the same capacity;

[0007] Perform latency tests on the extended memory of multiple host nodes to obtain performance data for multiple host nodes;

[0008] Determine memory performance consistency of multiple host nodes based on performance data of multiple host nodes.

[0009] This application also provides a memory performance detection device, including:

[0010] a mapping module configured to perform, for each of the plurality of host nodes, address mapping between a target memory corresponding to a data path port of the memory cartridge and an endpoint port of the host node based on a preset peripheral component interconnect topology, obtain an address mapping result, and convert the target memory into an extended memory recognizable by the host node based on the address mapping result; the extended memory of each host node having the same capacity;

[0011] A test module is used to perform latency testing on the extended memory of multiple host nodes to obtain performance data of the multiple host nodes;

[0012] The determination module is used to determine the memory performance consistency of the multiple host nodes based on the performance data of the multiple host nodes.

[0013] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any one of the above-mentioned memory performance detection methods when executing the computer program.

[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned memory performance detection methods are implemented.

[0015] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned memory performance detection methods when executed by a processor.

[0016] Through the present application, for each host node in turn, based on the peripheral component interconnect topology, the memory of the same capacity of the data path port of the memory cartridge is mapped to the endpoint port and the mapped memory is converted into extended memory recognizable by the host node, and then the memory performance of the extended memory is tested for consistency. This can quickly and effectively establish a memory mapping relationship, enable the host node to schedule the memory capacity required by the host node in the memory cartridge, and eliminate inconsistencies caused by different mapped memory spaces by configuring the same capacity of extended memory for each host node. Under this premise, the memory performance of multiple host nodes is made consistent, thereby improving the effectiveness of the consistency test. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 A schematic diagram of an application scenario of the memory performance detection method provided in an embodiment of the present application;

[0019] Figure 2 A flowchart of a memory performance detection method provided in an embodiment of the present application;

[0020] Figure 3 A schematic diagram of a peripheral component interconnect topology structure preset in the memory performance detection method provided in an embodiment of the present application;

[0021] Figure 4 A schematic diagram of the structure of a memory performance detection device provided in an embodiment of the present application;

[0022] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION

[0023] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0025] With the widespread deployment of large-scale server cluster architectures and cloud computing infrastructure, hyperscale cloud service providers are facing memory capacity challenges brought about by the exponential growth of data volume.

[0026] In related technologies, a design concept of capacity redundancy can be adopted. However, this approach will inevitably cause waste of memory resources. In addition, in order to adapt to the load requirements of different scenarios, dozens of instance specifications can be preset. However, the increase in instance types leads to increased operation and maintenance costs.

[0027] In order to solve the above technical problems, the inventors of this application have discovered that a memory pooling approach can be used to build a cluster-level memory box (memory resource pool). When the physical server is short of memory, the memory pool resources can be called in real time to avoid the cost increase and resource rigidity caused by physical expansion. When the physical memory capacity does not need to be large, the memory resources can be automatically released back to the memory box to reduce resource consumption. The inventors also considered that when each host node uses the memory pool resources, there will be a delay in memory performance. How to detect this memory performance delay and the consistency of the scheduling and allocation performance of the memory in the memory box by each host node? Therefore, after creative work, the inventors found that the consistency of the memory allocation mapping between each host node and the memory box has a great impact on the consistency of memory performance. Based on this, the embodiment of this application provides a memory performance detection method.

[0028] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0029] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the memory performance detection method depends, the specific application environment architecture or specific hardware architecture is described here. Figure 1 , Figure 1 Schematic diagram of the application scenario of the memory performance detection method provided in the embodiment of this application. Figure 1 As shown, a memory performance testing device 100 is connected to a general-purpose computing unit and a switch device. The general-purpose computing unit includes multiple host nodes 101. The multiple host nodes 101 in the general-purpose computing unit are connected to a switch device 102 via a set of Compute Express Link (CXL) interconnect buses. The switch device 102 is connected to a memory cartridge 103 via another set of CXL interconnect buses. The CXL interconnect buses can utilize Clustered Direct Attach Copper Cable with Force Plug (CDFP). Optionally, the general-purpose computing module is adaptable to various platform servers. The switch device 102 can support the CXL 2.0 protocol and provide 32 high-speed interfaces. Each high-speed interface can support 32GT / s, x16 bandwidth, and supports any upstream and downstream configuration. The memory cartridge 103 can also support the CXL 2.0 protocol and can accommodate up to 64 fifth-generation Double Data Rate 5 Registered Dual In-Line Memory Modules (DDR5 RDIMMs).

[0030] In a specific implementation, the memory performance testing device 100 can sequentially perform address mapping between each host node 101 and the memory in the memory cartridge 103. During the address mapping process, based on a preset peripheral component interconnect topology, the target memory corresponding to the data path (DP) port of the memory cartridge 103 can be mapped to the corresponding endpoint (EP) port of the host node 101 to obtain an address mapping result. Based on the address mapping result, the target memory is converted into extended memory recognizable by the host node. Latency testing is performed on the extended memory of multiple host nodes 101 to obtain performance data of the multiple host nodes 101. Based on the performance data of the multiple host nodes 101, the memory performance consistency of the multiple host nodes 101 is then determined. The memory performance testing method provided in this embodiment maps the memory of the same capacity of the data path port of the memory cartridge to the endpoint port for each host node in turn based on the peripheral component interconnect topology, and converts the mapped memory into extended memory recognizable by the host node, and then performs a consistency test on the memory performance of the extended memory. This method can quickly and effectively establish a memory mapping relationship, so that the host node can schedule the memory capacity required by the host node in the memory cartridge. By configuring the extended memory of the same capacity for each host node, the inconsistency caused by different mapped memory spaces is eliminated. Under this premise, the memory performance of multiple host nodes is made consistent, thereby improving the effectiveness of the consistency test.

[0031] Figure 2 A flow chart of the memory performance detection method provided in the embodiment of the present application is shown as follows: Figure 2 As shown, an embodiment of the present application provides a memory performance detection method, which is described in detail as follows:

[0032] 201. For each of the multiple host nodes, based on a preset peripheral component interconnect topology, address mapping is performed between a target memory corresponding to a data path port of a memory cartridge and an endpoint port of the host node to obtain an address mapping result. Based on the address mapping result, the target memory is converted into extended memory recognizable by the host node; the extended memory of each host node has the same capacity.

[0033] The execution subject of this embodiment may be a terminal device or a server, for example Figure 1 The memory performance detection device shown.

[0034] In this embodiment, the peripheral component interconnect topology may be a peripheral component interconnect express (PCIe) topology. To improve consistency, the PCIe topology may be designed symmetrically, and the communication links connecting multiple host nodes to the memory cartridge via a switch device may be designed symmetrically so that the length and bandwidth of each link are designed to be consistent. The symmetrical PCIe topology significantly improves the memory performance consistency of multiple host nodes by balancing paths and bandwidth allocation. Figure 3 As shown in the figure, there are eight host nodes (host nodes 0, 1, 2, ..., 6, 7), each of which includes two central processing unit (CPU) cores. The switch device consists of two layers: layer 0 and layer 1. For symmetry, one of the two CPU cores of each host node is connected to layer 0 of the switch device, and the other is connected to layer 1 of the switch device. Furthermore, the other eight ports of layer 0 of the switch device are connected one-to-one with the eight memory modules of the memory cartridge, and the other eight ports of layer 1 of the switch device are connected one-to-one with the other eight memory modules of the memory cartridge.

[0035] 202. Perform a delay test on the extended memory of multiple host nodes to obtain performance data of the multiple host nodes.

[0036] Specifically, you can install a testing tool on your system, such as the Memory Latency Checker (MLC) tool. You can run the MLC tool through a script and use the numactl command to specify host node 0 as the test node to obtain performance data for host node 0. Then, perform latency tests on the remaining host nodes (for example, host nodes 1, 2, ..., 6, 7) in sequence to obtain performance data for each of the remaining host nodes.

[0037] 203. Determine memory performance consistency of the multiple host nodes based on performance data of the multiple host nodes.

[0038] Specifically, after obtaining the performance data of each host node, the performance data of each host node may be compared. The delay difference of each host node should be less than a preset value, such as 10%.

[0039] In some embodiments, if the delay difference is greater than a preset value, the links of adjacent ports in the switch device can be swapped and tested again to confirm whether the cause of the delay difference is related to the topology. If the delay difference is still greater than the preset value after the link is swapped, it indicates that the delay difference is not related to the topology. The method provided in this embodiment can quickly locate whether the performance bottleneck is caused by the physical topology (such as cable quality, port congestion) or the device itself (such as memory module failure) by dynamically swapping the links of adjacent ports of the switch and retesting the delay difference, thereby enabling targeted optimization. This method can accurately identify link imbalance problems, optimize the CXL switch port allocation strategy, and reduce the latency fluctuations of multiple nodes accessing the memory pool. At the same time, by quantifying the consistency standard through a preset threshold (such as 10%), it can ensure the efficient collaboration of heterogeneous computing clusters, avoid the performance degradation of a single node affecting the global task, and significantly improve the stability and reliability of scenarios such as AI training and real-time analysis.

[0040] As can be seen from the above description, the memory performance detection method provided in the embodiment of the present application maps the memory of the same capacity of the data path port of the memory box to the endpoint port based on the peripheral component interconnection topology for each host node in turn and converts the mapped memory into extended memory recognizable by the host node, and then performs consistency testing on the memory performance of the extended memory. It can quickly and effectively establish a memory mapping relationship, so that the host node can schedule the memory capacity required by the host node in the memory box, and by configuring the extended memory of the same capacity for each host node, eliminates the inconsistency caused by different mapped memory spaces. Under this premise, the memory performance of multiple host nodes is made consistent, thereby improving the effectiveness of the consistency test.

[0041] In some embodiments, step 201 performs address mapping of the target memory corresponding to the data path port of the memory cartridge with the endpoint port of the host node based on a preset peripheral component interconnect topology to obtain an address mapping result, which may specifically include: obtaining a first connection state of a first port of a switch device corresponding to the host node; if the first connection state is a normal working state, obtaining a peripheral component interconnect topology of the switch device; if the peripheral component interconnect topology is consistent with the preset peripheral component interconnect topology, obtaining the number of multiple downstream devices of the switch device in the peripheral component interconnect topology; if the number of devices is consistent with the preset number, obtaining device information of the multiple downstream devices; configuring a corresponding endpoint port for the host node based on the device information; setting the mode of the endpoint port to a memory expansion mode and configuring a required memory capacity for the endpoint port; based on the memory expansion mode, mapping the memory of the data path port of the memory cartridge to the endpoint port according to the required memory capacity; and associating the addresses of the memory mapping input and output of the host node with the host random access memory address and the memory of the data path port of the memory cartridge through a secondary global register to obtain an address mapping result. The memory performance detection method provided in this embodiment can accurately identify link anomalies and ensure the reliability of memory expansion configuration by automatically verifying the consistency of switch port status and preset topology; by dynamically setting the endpoint port to memory expansion mode and flexibly allocating capacity, it can adapt to the resource requirements of different business scenarios on demand and avoid wasting hardware resources; by uniformly associating the host memory and CXL memory pool addresses through the secondary global register, it can eliminate the risk of address conflicts, improve cross-node data access efficiency, reduce latency, and improve the effectiveness and reliability of memory latency consistency detection.

[0042] Specifically, you can first check the connection status of the first port on the switch device connected to the host node. For example, assuming the first port is port 1, you can enter the command xconn-control 0 portstatus 1. If the output contains a preset string, such as "Link: 0x44201000 [L0], [G3 x4]", it is determined that the first port is stable and in normal working state, such as L0 state, and subsequent mapping and testing can be performed.

[0043] After confirming that the switch and host nodes are properly connected, check the PCIe topology of the switch. This topology can help you understand the connection between the switch and the memory modules in the downstream memory cartridges. For example, compare the number of devices on the reserved PCIe bus (for example, 3-41) with the preset PCIe topology. If the comparison results are consistent, the connection between the switch and the memory cartridges is normal. For example, you can enter the command lspci –tv to view the topology.

[0044] After checking the PCIe topology of the switch, you can view the device information (such as device type and physical address) of the downstream devices to further understand the device information and facilitate subsequent memory mapping.

[0045] After obtaining the device information of the switch's downstream devices, the memory module information of the memory cartridge connected to the switch is also known. You can then configure the host node's EP port and map the memory between the host node's EP port and the memory cartridge's DP port. For example, you can enter a command (e.g., xconn-control 0 epporton 13) to configure port 13 on the host node as an EP port. You can also enter a command (e.g., xconn-control 0 portstatus 13) to view the port configuration. If the output contains a preset string (e.g., Ena: 0x00100000 [EP, LTSSM EN]), the connection is normal.

[0046] After configuring the EP port, you can configure the CXL mode of the EP port. To implement memory mapping, you can set the CXL mode of the EP port to memory expansion mode, that is, CXL Type 3 mode. After configuring the mode, you can verify the configuration by executing the command (for example, xconn-control 0 cxltype 13).

[0047] After the mode is successfully configured, you can configure the required memory capacity for the host node, that is, the size of the host-managed device memory (HDM). For example, you can set the HDM size to 512GB. You can verify the configuration by executing the command (for example, xconn-control 0 showhdmsize 13).

[0048] After configuring the mode and required memory capacity, you can perform memory mapping. For example, you can enter a command (e.g., xconn-control 0 poolmap 6 0 13 256) to map 64GB of memory from DP port 6 to EP port 13 (the CXL base address is 256GB). To prevent address overlap, segmented mapping can be performed. For example, you can repeatedly execute the command (e.g., xconn-control poolmap 6 [md(d): n] 13 [256 + 64*n]) to map the memory from DP port 6 to EP port 13 in multiple steps. Note that the addresses must not overlap.

[0049] After address mapping is complete, global address mapping can be performed using the Level 2 Global Register (L2G). For example, a command (e.g., xconn-control 0 poolmap 6 0 8 768 8) can be entered to ensure that all host random access memory (RAM) addresses, including memory-mapped input / output (MMIO), are mapped to corresponding L2G registers. 768 = 256GB (host RAM) + 512GB (HDM capacity). This results in the address mapping result. Using the Level 2 Global Register to uniformly associate host local memory with memory cartridge addresses eliminates address space fragmentation and enables transparent integration of physical memory resources. Hardware-level address translation and consistency management avoid redundant data copies and access conflicts, significantly reducing memory access latency in heterogeneous computing scenarios and improving the effectiveness of memory performance testing.

[0050] In some embodiments, before obtaining the first connection status of the first port of the switch device corresponding to the host node, the method may also include: scanning the downstream devices of the switch device to determine the available ports and unavailable ports of the switch device. The memory performance detection method provided in this embodiment, by pre-scanning the downstream devices of the switch and identifying available ports, can automatically eliminate interference from unconnected or faulty ports, accurately lock valid physical links, and avoid invalid configuration operations. This mechanism can expose hardware connection problems (such as loose cables and damaged ports) in advance, reduce the risk of configuration failures caused by physical layer anomalies, significantly improve the efficiency of CXL memory pool deployment and the success rate of system initialization, provide reliable underlying connection guarantees for multi-node heterogeneous computing environments, and enhance the effectiveness of memory performance consistency detection.

[0051] For example, a command may be input to scan downstream devices to display the corresponding skipped ports, for example, port 1 is a skipped port and is unavailable, while ports 4-7, 10, and 11 are DP ports and are available ports.

[0052] In some embodiments, the addresses of the memory mapping input and output of the host node and the host random access memory address and the memory of the data path port of the memory box are associated through a secondary global register. After obtaining the address mapping result, it can also include: obtaining the address information in the secondary global register of the endpoint port to confirm whether the address mapping result is accurate; if the address mapping result is accurate, obtaining the third connection status between the endpoint port and the switch device; if the third connection status is a normal working state, and the configuration and power interface of the switch device contains the range information of the required memory of the endpoint port, it is confirmed that the switch device recognizes the memory of the data path port normally.

[0053] Specifically, after completing the mapping and associating the L2G registers, you can view the corresponding L2G registers to confirm the address mapping results. For example, you can enter a command (such as xconn-control 0 showl2g 13) to view the L2G register of EP port 13, which displays the address mapping relationship.

[0054] In some embodiments, based on the address mapping result, converting the target memory into extended memory recognizable by the host node may include: loading a driver module for a direct access device in the host node system to obtain a direct access device, and loading the memory corresponding to the address mapping result into the direct access device; and converting the memory under the direct access device to a non-uniform memory access node to obtain extended memory recognizable by the host node. The memory performance detection method provided in this embodiment achieves seamless integration of heterogeneous memory resources by dynamically loading a direct access device driver and automatically converting CXL memory to a non-uniform memory access (NUMA) node. This allows the host system to identify the remote CXL memory pool as a localized NUMA node, eliminating application-layer adaptation costs. It also ensures low-latency access and resource isolation for critical services, improving the effectiveness of memory latency consistency detection.

[0055] Specifically, you can load the driver in the host node's system, for example, by running commands like modprobe device_dax and modprobe dax_hmem, to load the CXL memory in the memory cartridge into the Direct Access (DAX) device. You can also enter commands like ls / dev to check whether device dax0.0 exists. To check the capacity of device dax0.0, you can enter commands like daxctl list –u to view the device and properties in the Devices - Direct Access Devices (Device - Dax) area to confirm whether the expanded memory capacity is the preset capacity (for example, 512GB).

[0056] In some embodiments, converting the memory under the direct access device to the non-uniform memory access node may include: generating random data and writing the random data to the direct access device; reading target data corresponding to the random data from the direct access device; comparing the target data with the random data; if the comparison results are the same, confirming that the read and write verification of the direct access device is successful, and converting the memory under the direct access device to the non-uniform memory access node. The memory performance detection method provided in this embodiment can ensure the reliability of CXL memory hardware and address mapping by writing random data and verifying read and write consistency before converting to the NUMA node, avoiding silent data corruption caused by physical link anomalies or register configuration errors; through the automated data comparison mechanism, it can expose hardware defects (such as signal attenuation, storage unit failure) or protocol layer compatibility issues in advance, reduce runtime risks in the production environment, and ensure the reliability and effectiveness of memory performance detection.

[0057] For example, you can use / dev / random as the input device to generate 64 bytes of random content and perform a read / write comparison on the DAX device. Write the randomly generated data to the DAX device, then read the written data and compare it with the randomly generated data. If they match, verification succeeds. After the read / write verification is complete, run the command (for example, sudo daxctl reconfigure-device --mode=system-ram --force dax0.0) to convert the CXL memory on the DAX device to a NUMA node.

[0058] In some embodiments, based on a preset peripheral component interconnect (PCI) topology, address mapping the target memory corresponding to the data path port of the memory cartridge to the endpoint port of the host node to obtain an address mapping result may include: after the memory cartridge is powered on, powering on and initializing the switch device; obtaining a second connection status between the endpoint port of the memory cartridge and the switch device; and if the second connection status is normal, address mapping the target memory corresponding to the data path port of the memory cartridge to the endpoint port of the host node based on the preset PCI topology to obtain an address mapping result. The memory performance detection method provided in this embodiment ensures hardware link layer readiness by automatically initializing the switch and verifying the connection status of the memory cartridge endpoint port after powering on, thereby avoiding mapping failures due to physical layer anomalies (such as loose cables or signal attenuation). Through an address mapping mechanism under preset topology constraints, the memory views of multiple host nodes can be forced to align, eliminating address space fragmentation and ensuring data access consistency across nodes. Dynamically converting the CXL memory pool into recognizable extended memory through a standardized process can achieve plug-and-play of heterogeneous memory resources, achieve low latency, and improve the effectiveness and reliability of memory latency consistency detection.

[0059] For example, first plug in the AC power supply, turn on the memory cartridge, and then plug in the switch. The switch automatically starts up. After pressing the power button, the management controller, such as the management central processing unit (mCPU), starts up, and the operating system is accessed. Then, you enter your account information (username / password) to complete the system login. After completing the system login, you can enter the corresponding command to check the connection status of the switch. If the connection status is normal, execute the power-on command to perform subsequent configuration and address mapping steps.

[0060] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0061] Figure 4 This is a schematic diagram of the structure of the memory performance detection device provided in the embodiment of the present application. Figure 4 As shown, an embodiment of the present application further provides a memory performance detection device, including a mapping module 401, a testing module 402 and a determining module 403.

[0062] A mapping module 401 is configured to perform, for each of the plurality of host nodes, address mapping between a target memory corresponding to a data path port of the memory cartridge and an endpoint port of the host node based on a preset peripheral component interconnect topology, obtain an address mapping result, and convert the target memory into extended memory recognizable by the host node based on the address mapping result; the extended memory of each host node having the same capacity;

[0063] Testing module 402, configured to perform a delay test on the extended memory of multiple host nodes to obtain performance data of the multiple host nodes;

[0064] The determination module 403 is configured to determine the memory performance consistency of the plurality of host nodes based on the performance data of the plurality of host nodes.

[0065] The memory performance detection device provided in the embodiment of the present application maps the capacity of the DP port of the memory cartridge to the EP port based on the peripheral component interconnect topology, so that the host random access memory Host RAM addresses, including the memory mapped input / output (MMIO), are mapped to corresponding registers. This can quickly and effectively establish a memory mapping relationship, allowing the host node to schedule the memory capacity required by the host node in the memory cartridge. By configuring the same capacity of extended memory for each host node, consistency testing can be facilitated.

[0066] In some embodiments, the mapping module 401 is specifically used to: obtain a first connection state of a first port of a switch device corresponding to a host node; if the first connection state is a normal working state, obtain a peripheral component interconnect topology structure of the switch device; if the peripheral component interconnect topology structure is consistent with a preset peripheral component interconnect topology structure, obtain the number of multiple downstream devices of the switch device in the peripheral component interconnect topology structure; if the number of devices is consistent with the preset number, obtain device information of the multiple downstream devices; based on the device information, configure a corresponding endpoint port for the host node; set the mode of the endpoint port to a memory extension mode, and configure the required memory capacity for the endpoint port; based on the memory extension mode, map the memory of the data path port of the memory box to the endpoint port according to the required memory capacity; associate the address of the memory mapping input and output of the host node with the host random access memory address and the memory of the data path port of the memory box through a secondary global register to obtain an address mapping result.

[0067] In some embodiments, the mapping module 401 is further configured to scan downstream devices of the switch device to determine available ports and unavailable ports of the switch device.

[0068] In some embodiments, the mapping module 401 is also used to: obtain the address information in the secondary global register of the endpoint port, and confirm whether the address mapping result is accurate; if the address mapping result is accurate, obtain the third connection state between the endpoint port and the switch device; if the third connection state is a normal working state, and the configuration and power interface of the switch device contains the range information of the required memory of the endpoint port, then confirm that the switch device's memory recognition of the data path port is normal.

[0069] In some embodiments, the mapping module 401 is specifically used to: load the driver module of the direct access device in the system of the host node, obtain the direct access device, and load the memory corresponding to the address mapping result into the direct access device; convert the memory under the direct access device to the non-uniform memory access node to obtain the extended memory that can be recognized by the host node.

[0070] In some embodiments, the mapping module 401 is specifically used to generate random data and write the random data to a direct access device; read target data corresponding to the random data from the direct access device; compare the target data with the random data; if the comparison result is the same, it is confirmed that the read and write verification of the direct access device is successful, and the memory under the direct access device is converted to a non-uniform memory access node.

[0071] In some embodiments, the mapping module 401 is specifically used to: after the memory cartridge is powered on, power on and initialize the switch device; obtain a second connection state between the endpoint port of the memory cartridge and the switch device; if the second connection state is normal, then based on a preset peripheral component interconnect topology, perform address mapping between the target memory corresponding to the data path port of the memory cartridge and the endpoint port of the host node to obtain an address mapping result.

[0072] For the description of the features in the embodiment corresponding to the memory performance detection device, please refer to the relevant description of the embodiment corresponding to the memory performance detection method, and will not be repeated here.

[0073] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 5 As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the electronic device 50 further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus.

[0074] During the specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that the at least one processor 501 executes the above-mentioned memory performance detection method embodiment.

[0075] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0076] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0077] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0078] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0079] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned memory performance detection method embodiments when running.

[0080] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0081] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned memory performance detection method embodiments are implemented.

[0082] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned memory performance detection method embodiments.

[0083] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0084] The above is a detailed introduction to a memory performance detection method, device, equipment and memory medium provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A memory performance detection method, characterized in that: include: For each of the plurality of host nodes, based on a preset peripheral component interconnect topology, performing address mapping between a target memory corresponding to a data path port of a memory cartridge and an endpoint port of the host node to obtain an address mapping result, and converting the target memory into an extended memory recognizable by the host node based on the address mapping result; The extended memory of each host node has the same capacity; Performing a delay test on the extended memory of the plurality of host nodes to obtain performance data of the plurality of host nodes; Memory performance consistency of the multiple host nodes is determined based on the performance data of the multiple host nodes.

2. The memory performance detection method according to claim 1, characterized in that: The method of performing address mapping between the target memory corresponding to the data path port of the memory cartridge and the endpoint port of the host node based on a preset peripheral component interconnect topology to obtain an address mapping result includes: Obtaining a first connection status of a first port of a switch device corresponding to the host node; If the first connection state is a normal working state, obtaining a peripheral component interconnect topology structure of the switch device; If the peripheral component interconnect topology is consistent with the preset peripheral component interconnect topology, obtaining the number of multiple downstream devices of the switch device in the peripheral component interconnect topology; If the number of devices is consistent with the preset number, obtaining device information of the plurality of downlink devices; Based on the device information, configure a corresponding endpoint port for the host node; Setting the mode of the endpoint port to a memory expansion mode and configuring the required memory capacity for the endpoint port; Based on the memory expansion mode, memory mapping the data path port of the memory cartridge to the endpoint port according to the required memory capacity; The memory mapping input and output addresses of the host node, the host random access memory address and the memory of the data path port of the memory cartridge are associated through a secondary global register to obtain an address mapping result.

3. The memory performance detection method according to claim 2, characterized in that: Before obtaining the first connection status of the first port of the switch device corresponding to the host node, the method further includes: Scan downstream devices of the switch device to determine available ports and unavailable ports of the switch device.

4. The memory performance detection method according to claim 2, wherein: After associating the memory mapping input and output addresses of the host node with the host random access memory address and the memory of the data path port of the memory cartridge through the secondary global register and obtaining the address mapping result, the method further includes: Obtaining address information in the secondary global register of the endpoint port to confirm whether the address mapping result is accurate; If the address mapping result is accurate, obtaining a third connection state between the endpoint port and the switch device; If the third connection state is a normal working state, and the configuration and power interface of the switch device includes the range information of the required memory of the endpoint port, it is confirmed that the memory recognition of the data path port by the switch device is normal.

5. The memory performance detection method according to any one of claims 1 to 4, characterized in that: The converting the target memory into extended memory recognizable by the host node based on the address mapping result includes: Loading a driver module of a direct access device in a system of a host node to obtain a direct access device, and loading a memory corresponding to the address mapping result into the direct access device; The memory under the direct access device is converted to a non-uniform memory access node to obtain an extended memory that can be recognized by the host node.

6. The memory performance detection method according to claim 5, characterized in that: The converting the memory under the direct access device to the non-uniform memory access node includes: generating random data, and writing the random data into the direct access device; Reading target data corresponding to the random data from the direct access device; Comparing the target data with the random data; If the comparison result is the same, it is confirmed that the read and write verification of the direct access device is successful, and the memory under the direct access device is converted to a non-uniform memory access node.

7. The memory performance detection method according to any one of claims 1 to 4, characterized in that: The method of performing address mapping between the target memory corresponding to the data path port of the memory cartridge and the endpoint port of the host node based on a preset peripheral component interconnect topology to obtain an address mapping result includes: After the memory cartridge is powered on, power on and initialize the switch. Acquire a second connection state between the endpoint port of the memory cartridge and the switch device; If the second connection state is a normal state, address mapping is performed between the target memory corresponding to the data path port of the memory cartridge and the endpoint port of the host node based on a preset peripheral component interconnect topology to obtain an address mapping result.

8. A memory performance detection device, characterized in that: include: a mapping module configured to perform, for each of the plurality of host nodes, address mapping between a target memory corresponding to a data path port of a memory cartridge and an endpoint port of the host node based on a preset peripheral component interconnect topology, obtain an address mapping result, and convert the target memory into an extended memory recognizable by the host node based on the address mapping result; The extended memory of each host node has the same capacity; A testing module, configured to perform a delay test on the extended memory of the plurality of host nodes to obtain performance data of the plurality of host nodes; A determination module is used to determine the memory performance consistency of the multiple host nodes based on the performance data of the multiple host nodes.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the memory performance detection method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the memory performance detection method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Computing system and computing method

    CN120980077A

  • CXL memory test method and device, electronic equipment and test network

    CN121641154A