Storage device hot plug method and device
By deploying CXL driver and memory hot-swap management modules on computing devices to manage the physical storage space of CXL devices, the problem that computing devices in the prior art cannot identify and manage CXL devices with constant power is solved, and the hot-swap management of the equipment is realized, reducing maintenance costs.
Patent Information
- Application Number
- CN202311695013.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art cannot identify and manage CXL-based physical storage devices without shutting down the computing device, resulting in power failure when replacing the storage device, which increases maintenance and testing costs.
By deploying CXL driver and memory hot-swap management modules on the computing device, listening to insert events of the CXL device, managing the physical storage space of the target device, generating management data and passing it back to the CXL device, enabling it to perform read and write operations.
It realizes hot-swap management of CXL devices and memory without interrupting services, reducing maintenance and testing costs.
Smart Images

Figure CN120144494A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technologies, and in particular, to a method and device for hot plugging a storage device. Background Art
[0002] Compute Express Link (CXL) is a development standard for high-speed and high-capacity central processing unit (CPU) to device and CPU to memory connections, designed specifically for high-performance data center computers. An extended storage device based on CXL, simply referred to as a CXL device, can effectively increase the storage capacity of a computing device. However, a computing device usually only supports identifying the memory of a CXL device in the case of cold plugging and then reporting it to the operating system of the computing device. This results in the need to power off the computing device and interrupt the operation of the service when replacing a CXL device or the memory on a CXL device, thereby increasing the maintenance and debugging costs. Summary of the Invention
[0003] Embodiments of this application provide a method, device, computing device, computer storage medium, and computer program product for hot plugging a storage device, which can implement hot plugging of a physical storage device based on CXL.
[0004] In a first aspect, an embodiment of this application provides a method for hot plugging a storage device. The method runs on a computing device and includes: receiving an insertion event, where the insertion event is an event that a target device is hot inserted into the computing device, and the insertion event includes first information, where the first information includes a unique identification number and a capacity of the target device, and the target device is a Compute Express Link (CXL) device or a memory connected to a CXL device; performing physical storage space management on the target device according to the capacity to obtain management data of the target device, where the management data includes an offset and a physical space address; and sending the management data to the CXL device, so that the CXL device can perform read and write operations according to the management data.
[0005] In this embodiment, hot insertion means that the computing device accesses a CXL device or inserts a memory (such as a memory module) into the connected CXL device without shutting down. The computing device listens for insertion events reported by the CXL device and performs hot insertion processing on the inserted target device, including managing the physical storage space of the target device, obtaining corresponding management data, and returning it to the CXL device. In this way, based on the stored management data, the CXL device can decode a read / write request received from the computing device and thus locate the specific physical space address pointed to by the request to perform read and write operations. In this embodiment, the computing device can implement insertion management of a physical storage device based on CXL without interrupting the service, which helps to reduce the maintenance and debugging costs.
[0006] In some possible examples, receiving an insertion event includes: receiving an insertion event reported by a CXL device through a CXL driver running on a computing device.
[0007] In this example, a CXL driver is deployed on the computing device to communicate with the CXL device, so as to be able to detect and identify the event information reported by the CXL device, and be able to transfer the event information to the operating system kernel of the computing device for processing.
[0008] In some possible examples, physical storage space management of a target device is performed according to capacity to obtain management data of the target device, including: the CXL driver running on the computing device transfers the first information to the memory hotplug management module running on the computing device through a target communication method, and the memory hotplug management module is a program for supporting memory hotplug; through the memory hotplug management module, unified addressing is performed on the physical storage space of the target device according to capacity to obtain management data.
[0009] In this example, the target communication method may include function calls, process communication, system calls, etc., but is not limited thereto. The computing device reuses the memory hotplug management module of the kernel to perform hot insertion management on the CXL-based physical storage device, and uses the CXL driver to interact with the memory hotplug management module to transfer messages. Such a communication mechanism realizes the perception ability and processing ability of the operating system for the hot insertion of the CXL-based physical storage device, so that a CXL-based physical storage device can be added in the state where the computing device is not shut down, realizing memory hot expansion.
[0010] In some possible examples, after obtaining the management data of the target device, the method includes: transferring the management data to the memory driver running on the computing device through the memory hotplug management module; sending the management data to the CXL device through the memory driver via the CXL driver running on the computing device.
[0011] In this way, the memory driver in the kernel obtains the management data of the target device and sends it back to the corresponding CXL device, and the kernel can normally use the storage space of the target device based on the management data and request the CXL device to perform corresponding read and write operations.
[0012] In some possible examples, after sending the management data to the CXL device, the method includes: when the Fabric Manager (FM) running on the computing device senses that the target device is hot-inserted into the computing device, obtaining second information of the target device from the CXL device, and FM is a program for managing the CXL device, and the second information includes the unique identification number and capacity of the target device; through FM, the physical storage space of the target device is divided into multiple physical regions according to the second information for use by multiple computing devices in the cluster where the computing device is located.
[0013] In this example, if the CXL device acts as a DAX device and is managed by the FM deployed on the computing device, when the FM senses that a target device is hot-plugged, it can request the CXL device to obtain relevant information to partition the physical storage space of the target device, obtaining multiple physical regions for allocation to each computing device in the cluster where the computing device is located, so as to achieve resource sharing of the memory provided by the CXL physical storage device in the cluster.
[0014] In some possible examples, after sending the management data to the CXL device, the method includes: mapping the physical storage space to the logical storage space through the memory driver according to the management data; creating, through the memory driver, a memory management page and an attribute file for the target device according to the logical storage space, where the memory management page is used to manage the mapping relationship between the physical storage space and the logical storage space, and the attribute file is used to provide the information of the logical storage space for the application layer to use; triggering the online operation of the logical storage space by writing an online command to the attribute file.
[0015] In this example, if the CXL device serves as the extended memory of the computing device, the memory driver running in the operating system kernel of the computing device can further process the management data passed by the memory hot-plug management module, map the physical storage space of the target device to the corresponding logical storage space, and perform corresponding page table management and online operations for use by the processes in the application layer.
[0016] In some possible examples, the method further includes: receiving an uninstall instruction for instructing to hot-remove the target device from the computing device, where the uninstall instruction includes the unique identification number of the target device; releasing the physical storage space of the target device and clearing the management data according to the unique identification number when the target device is in an idle state; sending an uninstall message to the CXL device to instruct the CXL device to clear the management data stored in itself.
[0017] In this example, the computing device can also receive the uninstall instruction through the CXL driver, and this uninstall instruction can be manually input into the computing device through the command line. Thus, the computing device can release the physical storage space of the target device to be uninstalled from the sparse memory model according to the uninstall instruction, clear the relevant management data, and instruct the CXL device to synchronize and update to complete the hot-uninstall operation without interrupting the operation of the computing device, which is beneficial for maintenance.
[0018] In some possible examples, before releasing the physical storage space of the target device according to the unique identification number, the method includes: determining the usage status of the target device by the Fabric Manager (FM) of the computing device according to the unique identification number, where the usage status includes at least an idle status or an occupied status; when the target device is in the occupied status, reclaiming the physical storage space of the target device from the target computing device that occupies the target device through the FM, so that the target device is restored to the idle status, and the target computing device is a computing device or in the same cluster as the computing device.
[0019] In this example, when the FM deployed on the computing device is responsible for managing CXL devices, before hot-unloading the target device, the FM needs to determine the usage status of the target device and perform operations such as space reclaiming, so that the hot-removal operation process is performed on the target device in the idle status, which helps to ensure the stable operation of the cluster service.
[0020] In some possible examples, before releasing the physical storage space of the target device according to the unique identification number, the method includes: determining the usage status of the target device by the memory driver running on the computing device according to the unique identification number, where the usage status includes at least an idle status or an occupied status; when the target device is in the occupied status, reclaiming all the storage space of the target device through the memory driver, so that the target device is restored to the idle status.
[0021] In this example, if the CXL device is managed by the computing device kernel, the memory driver is responsible for determining the usage status of the target device and performing operations such as space reclaiming before hot-unloading the target device, so as to ensure the stable operation of the processes on the computing device.
[0022] In some possible examples, before releasing the physical storage space of the target device according to the unique identification number, the method includes: when the target device is in the idle status, writing a take-offline command to the property file of the target device to trigger the take-offline operation of the corresponding logical storage space, where the property file is used to provide the information of the logical storage space for the application layer to use; releasing the mapping relationship between the physical storage space and the corresponding logical storage space through the memory driver, and clearing the memory management pages created according to the logical storage space; notifying the memory hot-plug management module running on the computing device through the memory driver to release the physical storage space of the target device.
[0023] Second aspect, an embodiment of the present application provides a hot plugging device for a storage device, the device includes: a receiving module and a processing module, wherein, the receiving module is configured to: receive an insertion event, the insertion event is an event that a target device is hot plugged into a computing device, the insertion event includes first information, the first information includes the unique identification number and capacity of the target device, and the target device is a Compute Express Link (CXL) device or a memory connected to the CXL device; the processing module is configured to: manage the physical storage space of the target device according to the capacity to obtain management data of the target device, wherein the management data includes an offset and a physical space address; send the management data to the CXL device, so that the CXL device can perform read and write operations according to the management data.
[0024] In some possible examples, the receiving module is specifically configured to: receive the insertion event reported by the CXL device through the CXL driver running on the computing device.
[0025] In some possible examples, the processing module is specifically configured to: enable the CXL driver running on the computing device to transfer the first information to the memory hot plugging management module running on the computing device through a target communication method, and the memory hot plugging management module is a program for supporting memory hot plugging; through the memory hot plugging management module, uniformly address the physical storage space of the target device according to the capacity to obtain management data.
[0026] In some possible examples, the processing module is further configured to: transfer the management data to the memory driver running on the computing device through the memory hot plugging management module; through the memory driver, send the management data to the CXL device through the CXL driver running on the computing device.
[0027] In some possible examples, the processing module is further configured to: when the Fabric Manager (FM) running on the computing device senses that the target device is hot plugged into the computing device, obtain second information of the target device from the CXL device, and the FM is a program for managing the CXL device, and the second information includes the unique identification number and capacity of the target device; through the FM, divide the physical storage space of the target device into multiple physical regions according to the second information for use by multiple computing devices in the cluster where the computing device is located.
[0028] In some possible examples, the processing module is further configured to: map the physical storage space to a logical storage space according to the management data through the memory driver; through the memory driver, create a memory management page and an attribute file of the target device according to the logical storage space, the memory management page is used to manage the mapping relationship between the physical storage space and the logical storage space, and the attribute file is used to provide the information of the logical storage space for the application layer to use; by writing an online command to the attribute file, trigger the online operation of the logical storage space.
[0029] In some possible examples, the receiving module is further used to receive an uninstall instruction, the uninstall instruction is used to instruct to hot-unplug the target device from the computing device, and the uninstall instruction includes a unique identification number of the target device; the processing module is further used to release the physical storage space of the target device according to the unique identification number and clear the management data when the target device is in an idle state; and send an uninstall message to the CXL device to instruct the CXL device to clear the management data stored therein.
[0030] In some possible examples, the processing module is also used to: determine the usage status of the target device according to the unique identification number through the structure manager FM of the computing device, and the usage status includes at least an idle state or an occupied state; when the target device is in the occupied state, reclaim the physical storage space of the target device from the target computing device occupying the target device through FM so that the target device is restored to the idle state, and the target computing device is a computing device or is in the same cluster as the computing device.
[0031] In some possible examples, before releasing the physical storage space of the target device according to the unique identification number, the method includes: determining the usage status of the target device according to the unique identification number through a memory driver running on the computing device, the usage status including at least an idle state or an occupied state; when the target device is in the occupied state, reclaiming all the storage space of the target device through the memory driver to restore the target device to the idle state.
[0032] In some possible examples, the processing module is also used to: when the target device is in an idle state, trigger an offline operation on the corresponding logical storage space by writing an offline command to the target device's attribute file, and the attribute file is used to provide information about the logical storage space to the application layer; release the mapping relationship between the physical storage space and the corresponding logical storage space through the memory driver, and clear the memory management page created according to the logical storage space; notify the memory hot-swap management module running on the computing device through the memory driver to release the physical storage space of the target device.
[0033] In a third aspect, an embodiment of the present application provides an electronic device, comprising: at least one memory for storing programs; and at least one processor for executing programs stored in the memory; wherein, when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.
[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0035] Fifth aspect, an embodiment of the present application provides a computer program product, characterized in that when the computer program product runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0036] Sixth aspect, an embodiment of the present application provides a chip, characterized in that it includes at least one processor and an interface; the at least one processor obtains program instructions or data through the interface; the at least one processor is used to execute the program line instructions to implement the method described in the first aspect or any possible implementation manner of the first aspect.
[0037] It can be understood that the beneficial effects of the above second aspect to the sixth aspect can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here. Description of the Drawings
[0038] Figure 1 is a schematic diagram of an application scenario of an embodiment of the present application;
[0039] Figure 2 is a schematic diagram of the structure of a computing device provided by an embodiment of the present application;
[0040] Figure 3A is a schematic diagram of the process of the kernel in the computing device of an embodiment of the present application for controlling the hot insertion of the memory on the CXL device;
[0041] Figure 3B is a schematic diagram of the process of the kernel in the embodiment of the present application for controlling the hot removal of the memory based on the CXL device;
[0042] Figure 3C is a schematic diagram of the process of the FM in the embodiment of the present application for controlling the hot insertion of the memory on the CXL device;
[0043] Figure 3D is a schematic diagram of the process of the FM in the embodiment of the present application for controlling the hot removal of the memory on the CXL device;
[0044] Figure 4 is a schematic diagram of the process of a method for hot plugging a storage device provided by an embodiment of the present application;
[0045] Figure 5 is a schematic diagram of the process of hot insertion of a target device in a specific embodiment of the present application;
[0046] Figure 6 is a schematic diagram of the process of hot removal of a target device in a specific embodiment of the present application;
[0047] Figure 7 is a schematic diagram of the process of hot insertion of a target device in a specific embodiment of the present application;
[0048] Figure 8 It is a schematic flowchart of hot insertion of a target device in a specific embodiment of the present application;
[0049] Figure 9 It is a schematic flowchart of hot removal of a target device in a specific embodiment of the present application;
[0050] Figure 10 It is a schematic structural diagram of a hot pluggable device for a storage device provided by an embodiment of the present application;
[0051] Figure 11 It is a schematic structural diagram of a chip provided by an embodiment of the present application. Detailed implementation manners
[0052] The term "and / or" in this article is an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this article represents an "or" relationship between associated objects. For example, A / B represents A or B.
[0053] The terms "first" and "second" in the description and claims of this article are used to distinguish different objects, rather than to describe a specific order of objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of response messages.
[0054] In the embodiments of the present application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0055] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.
[0056] To facilitate the understanding of the technical solutions of the embodiments of the present application, the following explains the technical terms and abbreviations involved in this article.
[0057] CXL protocol: also known as the Computing Express Link protocol, is a high-speed serial protocol that enables fast and reliable data transmission between different components within a computing device, such as the central processing unit (CPU) and accelerator, memory buffer, and smart network interface card (Smart NIC). CXL protocols include CXL.io protocol, CXL.cache protocol, CXL.memory protocol, etc.
[0058] Fabric manager (FM): manages the CXL storage space of the memory pool composed of one or more CXL devices, is responsible for sharding these CXL storage spaces, and then allocating them to various computing devices, and recording mapping information. FM can be deployed on any computing device, or on a CXL switch (switch) and other devices, and only one computing device, CXL switch or other device can run FM at the same time.
[0059] Basic Input Output System (BIOS): is a set of programs fixed to a read-only memory on the motherboard of the computer. It is the first program to run when the computer is turned on. It includes the most important basic input and output programs of the computer, the self-test program after powering on, and the system self-starting program. BIOS can also read and write specific information of system settings. Its main function is to provide the lowest-level and most direct hardware settings and control for the computer.
[0060] Advanced Configuration and Power Interface (ACPI): It is an operating system power management and hardware configuration interface and is an open standard. ACPI defines the hardware abstraction interface between the system firmware (BIOS or unified extensible firmware interface (UEFI)) and the operating system.
[0061] Hot plugging: plugging and unplugging physical devices (such as memory, solid-state drives, etc.) without turning off the power to the computing device (such as a server).
[0062] Memory hot-plug: The Linux (an operating system) kernel supports memory hot-plug technology. It supports increasing or decreasing physical memory while the computing device is running and the operating system is running normally.
[0063] Firmware (FW): A program written into an erasable programmable read-only memory (EPROM) or an electrically erasable programmable read-only memory (EEPROM). Firmware refers to the "driver" of a device stored inside the device. Through the firmware, the operating system can achieve the running actions of a specific machine according to the standard device driver.
[0064] Command Line Interface (CLI): A user interface that supports users to input instructions through the keyboard. After the computer receives the instructions, it executes them.
[0065] Baseboard Management Controller (BMC): It can implement a series of monitoring and control functions, and the objects of operation are system hardware. For example, it monitors the temperature, voltage, fans, power supply, etc. of the system and makes corresponding adjustment work to ensure that the system is in a healthy state. It can be responsible for recording various hardware information and log records for prompting users and subsequent problem location. The BMC is an independent system that can communicate with other hardware on the server (such as CPU, memory, etc.) through a physical channel and can also interact with the BIOS and the operating system (OS).
[0066] Page Frame Number (PFN): Physical memory is divided into fixed-size partitions, called page frames (or page frames). The number of a page frame is called the page frame number (or page frame number, memory block number), and the page frame number starts from 0. A certain number of page frames form a memory segment.
[0067] Designated Vendor-Specific Extended Capability (DVSEC): This function defines a configuration register structure that can be implemented by the vendor and provides a consistent hardware / software interface. The DVSEC can provide prompts in the DVSEC function definition to allow the system software to determine whether the hot-pluggable ports are extensible and indicate the number of buses that need to be reserved.
[0068] Buddy (memory allocation system): It is a memory allocation mechanism in the Linux system kernel. By organizing physical memory page frames, it reasonably allocates and reclaims physical memory pages, enabling the quick combination of memory allocation and adjacent memory to alleviate memory fragmentation.
[0069] List interface: A common data structure interface that provides a set of methods for accessing, inserting, deleting, and traversing list elements.
[0070] Host-Managed Device Memory (HDM): The CXL.memory protocol makes the memory of the device be uniformly addressed in the system, so that it seems to be the main memory of the server (Host) to the CPU and can be directly accessed. The device memory after unified addressing is called HDM.
[0071] Generally, a computing device can only support identifying CXL devices in the cold start mode. The CXL device (including the CXL controller and one or more memories connected to the CXL controller) is installed on the computing device and powered on and initialized in the cold start mode. At this time, the BIOS program can detect the CXL device, obtain the status information, memory configuration information, etc. of the CXL device and record them, and then transfer the recorded information to the kernel at the end of the BIOS life cycle. In this way, the kernel can know the CXL device after startup and complete the corresponding memory space mapping. However, this cold plug-and-play technology does not support the plugging and unplugging of some memories or the whole board of the CXL device on the computing device without power-off. This results in that the computing device must be powered off when replacing the CXL device or its memory, interrupting the business operation. And the power-on restart time of large computing devices (such as servers) is 5 to 10 minutes. Therefore, this cold plug-and-play technology greatly increases the maintenance and debugging cost, especially in the cluster computing scenario, which will greatly affect the business operation.
[0072] Since the CXL device is an external device, the way it communicates with the kernel is different from that of the memory on the motherboard (such as a memory module), that is, the CXL device cannot communicate with the kernel like the motherboard memory. Therefore, how to notify the operating system to operate on the CXL device after physically identifying the hot plugging and unplugging of the CXL device on the computing device is still an unsolved problem.
[0073] To achieve the hot plugging and unplugging of the CXL device and reduce the maintenance cost while expanding the memory based on the CXL device, the embodiment of the present application provides a method for hot plugging and unplugging a storage device. This method mainly enables information transfer at the software level between the CXL device and the operating system of the computing device, so that the CXL device inserted when the computing device is not powered off can be recognized and processed by the operating system. In this way, the hot insertion and hot removal of the CXL device can be performed without interrupting the business running on the computing device.
[0074] To facilitate the understanding of the technical solution of the embodiment of the present application, an application scenario of the embodiment of the present application will be introduced first below.
[0075] Exemplarily, Figure 1 shows a schematic diagram of an application scenario of the embodiment of the present application. As Figure 1As shown, a computing cluster 10 may include multiple computing devices 100, and the computing devices 100 may communicate with each other based on protocols such as the hypertext transfer protocol (HTTP). In this example, the CXL device 20 may be connected to multiple computing devices 100, and one of the computing devices 100 may allocate the memory space provided by the CXL device 20 to each of the computing devices 100 in the cluster 10 through the deployment of FM30.
[0076] In this example, the CXL device 20 may include a circuit board 23, a CXL controller 21 and a memory 22 disposed on the circuit board, and the memory 22 is connected to the CXL controller 21. Among them, the memory 22 may be a storage chip, which is attached to the circuit board 23 of the CXL device together with the CXL controller, so that the memory 22 can be hot-plugged together with the entire CXL device 20. Alternatively, the memory 22 may be a memory module, and there is a memory slot on the circuit board 23 of the CXL device 20, and the memory 22 can be inserted into the memory slot. Then, the memory 22 can be hot-plugged independently or together with the entire CXL device 20. Exemplarily, the memory 22 may be a volatile memory or a non-volatile memory, and the CXL controller 21 can manage each memory 22.
[0077] The CXL controller 21 may be a multi-headed CXL expansion control chip, so that the CXL controller 21 can be connected to multiple computing devices 100.
[0078] Taking the hot-plugging of the memory 22 on the CXL device 20 without powering off the computing device 100 as an example. In this example, in the non-power-off scenario, the CXL controller 21 can identify the hardware information of the newly inserted memory 22 on the circuit board 23 (including but not limited to the unique identification number, capacity, bandwidth, latency, etc.), and report the hardware information to each computing device 100. Then, the kernel of the operating system 101 of the computing device 100 detects the capacity of the newly inserted memory 22 according to the hardware information, divides the memory 22 into corresponding physical pages, obtains the offset and physical space address of the physical pages, maps these physical pages to the virtual memory space, and the physical space address and offset of the memory 22 are also passed back to the CXL device 20 for storage.
[0079] When the FM30 senses that a memory 22 is inserted into the CXL device 20 in the computing cluster 10, it can obtain information such as the unique identification number and capacity of the memory 22 for storage space allocation, and provide it for use by each computing device 100 in the computing cluster 10. It can be understood that when the computing device 100 uses the storage space of the allocated memory 22, it is based on management data such as the physical space address and offset of the memory 22 stored in its own internal storage.
[0080] In this example, if a certain memory 22 on the CXL device 20 needs to be unloaded, the operation and maintenance personnel input the hardware information of the memory 22 to the FM30 through the command line to trigger the unloading process. The FM30 determines whether the storage space provided by the memory 22 is occupied by any computing device 100. If it is occupied, the space is first reclaimed from these computing devices 100. If the memory 22 is idle (or has been reclaimed), the FM30 notifies the operating system 101 of the computing device 100 to gradually unload the virtual memory space and physical space of the memory 22, clear the relevant records of the memory 22, and synchronize and update the CXL device 20. In this way, the user can pull out the memory 23 from the circuit board 23.
[0081] It can be understood that the computing cluster 10 can perform hot-plugging processing on the entire CXL device 20 based on a similar principle through the FM30, without interrupting the cluster service, which is beneficial to reducing the maintenance and debugging costs during the memory expansion process.
[0082] In other application scenarios, the CXL device 20 can also be directly connected to a certain computing device 100 through a cable or other means, and the operating system 101 of the computing device 100 can implement the entire hot-plugging process for the CXL device 20 or the memory 22 based on a similar principle.
[0083] Next, a computing device provided by an embodiment of the present application will be introduced with reference to the accompanying drawings.
[0084] Exemplarily, Figure 2 A schematic structural diagram of a computing device is shown. The computing device 100 can be a server, a computer, or a virtualized device (such as a virtual machine or a container, etc.). Among them, the server can be a physical server, a cloud server, or a GPU server. As Figure 2 shown, the computing device 100 can include a processor 110, a memory 120, and a communication interface 130. The processor 110, the memory 120, and the communication interface 130 can be deployed on the motherboard 11 and connected through a bus or other means.
[0085] In this embodiment, the processor 110 is the computing core and control core of the computing device 100. In some embodiments, the processor 110 may execute some or all of the steps of the method provided in this embodiment. As an example, the processor 110 may be a central processing unit (CPU), a system on chip (SOC), a processor integrated on the SOC, a separate processor chip, or a controller, etc.: The processor 110 may further include dedicated processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), etc. The processor 110 may be a processor group composed of multiple processor chips, and the multiple processor chips are coupled to each other through one or more buses.
[0086] The memory 120 provides storage space and can be used as the memory device of the server to store computer program instructions and data, such as operating systems (OS) 101 like Windows system, Linux system, etc., as well as programs such as the CXL driver 121, the memory hot plug management module 122, the memory driver 123, the CXL command line interface module 126, etc., but not limited to these. In addition, the memory 210 may also be used to store data such as the memory management table 125, the property file (such as the sysfs file) 125, etc.
[0087] Exemplarily, the memory 120 may be a non - volatile memory, such as an embedded multi - media card (EMMC), a universal flash storage (UFS), or a read - only memory (ROM), or other types of static storage devices that can store static information and instructions. It may also be a volatile memory, such as a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read - only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other computer - readable storage medium that can be used to carry or store program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0088] The communication interface 130 is mainly used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present application.
[0089] In addition, the computing device 100 may further include a baseboard management controller BMC150, and the baseboard management controller BMC150 communicates with other hardware (such as the CPU, etc.) on the computing device 100 through a physical channel.
[0090] In this embodiment, the computing device 100 can also be connected to the CXL device 20 through a cable or other means. The FM30 can be deployed in the processor 110 to perform hot - plug management and storage space configuration and usage for the connected CLX device 20. Alternatively, the FM30 can also be deployed in the BMC150 to manage and configure the CXL device 20.
[0091] It should be noted that when the FM30 is deployed on multiple devices (including the computing device 100) in a computing cluster 10, different FM30s back up each other, and only one FM30 is in the running state in the cluster at the same time. The storage space provided by the CXL device 20 accessed by a certain computing device 100 can be allocated by the FM30 for use by other computing devices 100 in the computing cluster 10.
[0092] Next, the functions of the relevant software programs in the computing device 100 in this embodiment will be introduced.
[0093] In this embodiment, referring again to Figure 1As shown, the computing device 100 can be connected to the CXL device 20. When the CXL driver 121 of the computing device 100 is run by the processor 110, it can be used to implement communication between the kernel and the CXL device 20, identify the events reported by the CXL device 20, and parse the device information of the CXL device 20 from the events (including but not limited to the unique identification number of the CXL device 20, the number of memories 22 inserted on each CXL controller 21, etc.), parse the hardware information of the memory 22 (including the unique identification number of the memory, capacity, bandwidth, latency, etc.), and so on.
[0094] The memory hot-plug management module 122 is a software module that runs in the operating system 101 of the computing device 100 and supports memory hot-plug based on the ACPI technology or the memory hot-plug technology. In this embodiment, the memory hot-plug management module 122 may specifically include an ACPI module I221 and a memory hot-plug module 1222. The ACP module I221 and the memory hot-plug module 1222 can respectively define different hardware configuration interfaces of the operating system 101. When the CXL driver 121 is run by the processor 110, it can communicate with the ACPI module 1221, and transfer the device information of the newly inserted CXL device 20 and the hardware information of the corresponding memory 22 to the ACPI module 1221 through the interface defined by the ACPI module 1221, and then the ACPI module 1221 transfers this information to the memory hot-plug module 1222. Alternatively, when the CXL driver 121 is run by the processor 110, it can also directly communicate with the memory hot-plug module 1222, and transfer the device information of the newly inserted CXL device 20 and the hardware information of the memory 22 through the interface defined by the memory hot-plug module 1222.
[0095] Exemplarily, the memory hot-plug module 1222 can adopt mature memory hot-plug technology in the art, support allowing dynamic addition or removal of physical memory devices at runtime, and realize dynamic management of memory resources by updating the memory management data structure and memory segment mapping relationship of the kernel, and provide a corresponding event notification mechanism so that the user space and application programs can respond to memory hot-plug events. Specifically, in this example, the memory hot-plug module 1222 can detect the insertion of physical memories, thereby initiating detection of the attributes and capacities of these memories, and initializing to obtain the management data of the memory 22 (such as the starting page frame number, ending page frame number, total number of pages, physical space address, etc.). The memory hot-plug module 1222 can also update the corresponding memory segment mapping relationship according to the insertion event or removal event during the memory hot-plug process, and support the kernel to dynamically allocate and release the capacities provided by these expanded memories 22.
[0096] The memory driver 123 is a memory management subsystem in the operating system 101 of the computing device 100. When the memory driver 123 is run by the processor 110, it can obtain the management data of the newly inserted CXL device 20 or memory 22 from the memory hot-plug management module 122, and perform management operations on the extended memory, such as creating a memory management table 125, creating a memory attribute file (such as a sysfs file) 124, etc. In this way, by writing to the attribute file 124, the corresponding logical storage space (memory space) of the CXL device 20 and the memory 22 can be triggered to go online and be visible to the user. Similarly, when unloading the corresponding logical storage space of the CXL device 20 and the memory 22, the memory driver 123 can also make the logical storage space invisible to the user through operations on the attribute file 124 to achieve hot unloading of the CXL device 20 or the memory 22.
[0097] The CXL command-line interface module 126 can communicate directly with the CXL driver 121 or indirectly communicate with the CXL driver 121 through the corresponding software development kit (SDK). The user can input instructions through the CXL command-line interface module 126 to trigger the underlying CXL driver 121 to obtain relevant information about the CXL device 20 and its memory 22 and transfer it to the operating system 101 or FM30.
[0098] FM30 can uniformly manage and allocate the storage space provided by the CXL device 20 accessed by the computing device 100. In some examples, FM30 can also be responsible for controlling the hot plug and unplug of the CXL device 20 and its memory 22.
[0099] In this example, when the computing device 100 is deployed as a server in a cluster, when other servers (not running FM30) apply to the FM30 of the computing device 100 for binding storage space, the FM30 divides the managed storage space into multiple physical regions, so as to allocate and bind physical regions to the servers according to their requests, and records the allocation situation in the mapping table 132. The mapping table 132 mainly records the relationship between physical regions and the mapped memories, and the relationship between physical regions and the bound computing devices.
[0100] The BMC 150 can also be deployed to implement the FM30, so as to manage the relevant information of the CXL device 20 and its memory 22, such as the above-mentioned device information, hardware information, and location information (such as the location where the CXL device 20 is connected to the motherboard 11, and the location information of the memory 22 in the slot of the CXL device 20), etc. In this way, the location information can be viewed with the unique identification number of the memory 22 and the unique identification number of the CXL device 20, so as to perform operations such as allocation management or unloading of target devices according to these identification numbers.
[0101] In one scenario, a CXL device is connected to a computing device to expand the memory of the computing device. Next, in combination with Figure 3A the hot plug-in process shown in Figure 3B and
[0102] the uninstallation (removal) process shown in Figure 3A the specific hot plugging process of the memory 22 on the CXL device 20 by the memory of the computing device 100 will be introduced.
[0103] In this implementation, when the computing device 100 is powered on, when a memory 22 is inserted into the CXL device 20, as
[0104] The CXL driver 121 identifies this event as an insertion event, parses out the device information of the CXL device 20 and the hardware information of the memory 22 from it, and transfers them to the memory hotplug management module 122 of the operating system 101 kernel through step 2. Then, the memory hotplug management module 122 can perform physical page partitioning and unified addressing on the physical storage space of the memory 22 according to this hardware information, obtain the PFN of each physical page, the physical space address and offset of the physical page, etc., and thus generate the memory 22 management data (including the starting PFN, total number of pages, starting physical space address, offset, etc.) and transfer it to the memory driver 123 through step 4 (it can also be together with the device information and hardware information).
[0105] After receiving the management data, the memory driver 123 will return the management data to the CXL device 20 through steps 6 and 7. At the same time, the memory driver 123 also processes as follows according to whether the storage space provided by the memory 22 is allocated for the kernel space or the user space: If the storage space is allocated for the kernel space, the memory management module 1231 in the memory driver 123 maps each physical page of the memory 22 to the corresponding virtual page according to the management data through step 5, and creates a memory management table 125 to manage the mapping from the physical storage space of the memory 22 to the logical storage space. And the memory driver 123 also creates an attribute file 124 for the logical storage space of the memory 22. In this example, the attribute file 124 is a sysfs file. The sysfs file is an attribute file created by the linux system kernel for the new storage device. Under the sysfs file, a new memory node can be created corresponding to the newly inserted memory 22. In this way, writing "online" (online) to the state file (used to describe the state of the memory node) of the new memory node under the sysfs file can bring the logical memory space of the memory 22 online and make the logical memory space visible to the user. It can be understood that the online logical memory space is managed by the memory allocation system in the kernel ( Figure 3A not marked in), so that it can be provided for the kernel-mode processes of the computing device 100 to use.
[0106] If the storage space of the memory 22 is allocated for the user space, the memory driver 123 will hand over this management data to the memory allocator 127 for mapping from the physical storage space to the logical storage space for use by user-mode processes, which will not be elaborated here.
[0107] So far, the hot insertion operation of the computing device 100 for the memory 22 is completed. It can be understood that a part of the storage space of the memory 22 can be the kernel space, and the other part can be used as the user space, or the entire storage space of the memory 22 can be used as the kernel space. And the order of execution of step 5 and step 6 by the memory driver 123 is not limited.
[0108] In addition, after receiving the management data transmitted by the CXL driver 121, the CXL device 20 stores it. In this way, when the CXL device 20 receives a read / write request, it can determine the specific physical space address pointed to by the address in the request by decoding the address in the request, so as to perform the read / write operation.
[0109] In this implementation, when the computing device 100 performs a hot-unloading operation on the memory 22 on the CXL device 20, refer to Figure 3B as shown, which specifically includes:
[0110] First, the kernel of the operating system 101 can obtain the unloading instruction input manually in the CXL command-line interface module 126 through step 1. The unloading instruction includes at least the unique identification number (such as the SN number) of the memory 22 (target device) to be unloaded, and may also include the unique identification number of the CXL device 20, etc. The memory driver 123 in the kernel hands over the memory 22 as a part of the kernel space to the memory management module 1231 for processing according to the unique identification number of the memory 22, and hands over the memory 22 as a part of the user space to the memory allocator 127 for processing.
[0111] Among them, the processing of the kernel space includes: the memory management module 1231 determines the usage status (idle or occupied) of the kernel space according to the unique identification number of the memory 22. If it is idle, a logical storage space unloading operation is performed. If it is occupied, the space is first reclaimed from the process occupying this part of the space, and after the space is reclaimed, the logical storage space unloading is performed. Specifically, continue to refer to Figure 3B , when the memory management module 1231 performs a logical storage space unloading on the kernel space provided by the memory 22 through step 2, it first performs a write operation on the attribute file 124 of the memory 22, so that the logical storage space of the memory 22 is invisible to the user. Then, the mapping relationship between the corresponding physical storage space of the memory 22 and this logical storage space is released, and the record in the memory management table 125 is cleared, that is, the unloading of the logical storage space of the memory 22 is completed.
[0112] Similarly, for the part of the memory 22 as the user space, if it is occupied, the memory allocator 127 also first reclaims the space, and then releases the mapping relationship between this part of the logical storage space and the corresponding physical storage space of the memory 22, and completes the unloading of the logical storage space.
[0113] Next, after completing the logical space unloading, the memory driver 123 notifies the memory hot-swap management module 122 to perform physical storage space unloading on the memory 22. The memory hot-swap management module 122 executes step 3, releases the memory capacity provided by the memory 22 from the sparse memory model according to the unique identification number of the memory 22, and clears the memory information (including management data) related to the memory 22. Then, the memory hot-swap management module 122 notifies the CXL driver 121 to transmit the unloading information to the CXL device 20 where the memory 22 is located through steps 4, 5, and 6, instructing the CXL device 20 to unload the memory 22. In this way, the CXL device 20 verifies the unique identification number of the memory 22 contained in the unloading information according to the unloading information. If the verification is passed, the records related to the memory 22 (including management data such as the physical space address) are cleared, and step 7 is executed to return the message of unloading completion to the CXL command line interface module 126 to inform the operator. At this point, the operator can unplug the memory 22 from the CXL device 20.
[0114] In addition, in some possible implementations, the Figure 3A Hot plugging and Figure 3B The hot uninstall process shown in the figure has a similar operating principle, and performs hot insertion and hot uninstall operations on the entire CXL device 20 connected with the memory 22.
[0115] In another scenario, multiple computing devices 100 in computing cluster 10 are connected to CXL device 20 (eg Figure 1 As shown), and the storage space of the CXL device 20 is divided and managed by FM40, then the next step is to combine Figure 3C The hot-plug process shown and Figure 3D The hot uninstallation flow shown introduces the specific process of hot plugging the memory 22 on the CXL device 20 by the FM 30 .
[0116] In this implementation, during the operation of computing cluster 10, when CXL device 20 detects that a new memory 22 is inserted, such as Figure 3C As shown, the insertion process includes: the CXL device 20 reports the insertion event of the memory 22 to each connected computing device 100 through step 1, and then each computing device 100 performs steps 2 to 4, performs unified addressing of the physical storage space division of the memory 22, and transmits the generated management data to the memory driver 123. It can be understood that each computing device 100 and the CXL device 20 interactively perform steps 1 to 4, and Figure 3A The execution principles of steps 1-4 in the example shown in are similar and will not be repeated here.
[0117] Next, the memory driver 123 returns the obtained management data together with the identity identifier of the affiliated computing device 100 to the CXL device 20 through steps 5 and 6. In this way, the CXL device 20 will receive the management data associated with the identity identifiers of each computing device 100 and store them. Among them, the management data of the memory 22 returned by each computing device 100 may be different.
[0118] In this implementation, after the FM30 senses that a new device is inserted into the CXL device 20, it will obtain the unique identification number and capacity of the newly inserted memory 22 (and can also obtain the device information of the CXL device 20) from the CXL device 20 through step 7, and then perform space partitioning according to the capacity to obtain multiple physical regions for allocation to each computing device 100 in the computing cluster 10 for use. Then, each computing device 100 can use the corresponding physical region allocated by the FM30 based on the management data of the memory 22 stored by itself. For example, map the physical address space of this region to the logical storage space for the process to use. In this way, when the CXL device 20 receives a read / write request from the computing device 100, it can determine which physical space address allocated to which computing device the address in the request points to by decoding the identity identifier of the computing device and the carried address in the request, so as to locate to the corresponding position to perform the read / write operation.
[0119] In this implementation, when the FM30 performs the hot-unloading operation on the memory 22 on the CXL device 20, refer to Figure 3D as shown, specifically including:
[0120] The FM30 can obtain the unloading instruction input by the operator in the CXL command line interface module 126 through step 1. The unloading instruction includes at least the unique identification number of the memory 22 to be unloaded. The FM30 determines the computing devices 100 bound to each physical region of the memory 22 and their usage status (idle or occupied) according to the unique identification number. If the memory 22 is idle, the FM30 executes step 2 to notify each computing device 100 to perform the unloading operation. If a certain physical region is occupied, the FM30 will first reclaim the space from the corresponding computing device 100 and then execute step 2 to instruct the computing device 100 to perform the unloading operation.
[0121] Specifically, when the memory driver 123 of the computing device 100 receives the offloading operation notification transmitted by the FM30, it causes the memory hot-plug management module 122 to execute step 3 to perform physical storage space offloading on the memory 22, and the memory hot-plug management module 122 notifies the CXL device to offload the memory 22 through steps 4, 5, and 6. The offloading result is returned to the CXL command line interface module 126 through step 7 so that the operator can be aware, facilitating the operator to unplug the memory 22 from the CXL device 20. In this example, the execution principles of steps 3 to 7 are similar to those of steps 3 to 7 in the example shown above Figure 3B and will not be elaborated here.
[0122] In addition, in some possible implementation manners, operations similar to the Figure 3C hot insertion shown and Figure 3D the hot offloading process shown can be adopted to perform hot insertion and hot offloading operations on the entire CXL device 20 connected with the memory 22.
[0123] In this way, the computing device 100 can also achieve hot plugging of the CXL device or the memory of the CXL device without powering off, without interrupting business operations, which helps to reduce the maintenance and debugging costs.
[0124] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the computing device 100. In other embodiments of the present application, the computing device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0125] Next, based on the content described above, a method for hot plugging a storage device provided in the embodiments of the present application will be introduced. It can be understood that this method is proposed based on the content described above, and some or all of the content in this method can refer to the description above.
[0126] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a method for hot plugging a storage device provided in the embodiments of the present application. It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Taking the execution on the computing device 100 shown in Figure 1 as an example, as shown in Figure 4 , this method may include:
[0127] S410, the computing device receives an insertion event.
[0128] In this embodiment, the insertion event may be an event reported by the CXL device 20 that the target device is hot-inserted into the computing device 100. The target device may be the CXL device 20 or the memory 22 connected to the CXL device 20 (which may be a volatile memory, a non-volatile memory, etc., such as a memory module, pmem). The hot insertion means that the computing device 100 accesses the target device in the powered-on state.
[0129] Exemplarily, during the operation of the computing device 100, if the computing device 100 newly accesses a CXL device 20, or one or more memories 22 are newly inserted on the CXL device 20 connected to the computing device 100, the CXL device 20 may generate a corresponding insertion event and send it to the computing device 100. It can be understood that the computing device 100 can monitor the insertion event through the running CXL driver 121.
[0130] Exemplarily, the first information is encapsulated in the insertion event. The first information may specifically include the device information of the CXL device 20 and the hardware information of the memory 22 on the CXL device 20. It can be understood that when the target device is the CXL device 20, the device information of the CXL device 20 (including the unique identification number of the device 20) and the hardware information of all the memories 22 inserted thereon (including the unique identification number and capacity of the memory 22) need to be encapsulated in the insertion event. When the target device is one or more memories 22, the first message may include the hardware information of these memories 22 and the device information of the CXL device 20 where they are located.
[0131] S420, perform physical storage space management on the target device according to the capacity to obtain management data of the target device.
[0132] In this embodiment, the CXL driver 121 running on the computing device 100 parses the first information from the insertion event, and then passes the first information to the kernel of the operating system 101 through a target communication method such as function call, process communication, or system call, so as to perform physical storage space management on the target device according to the target capacity, such as dividing physical pages one by one and uniformly addressing them to obtain corresponding management data. The management data may include the offsets and physical space addresses of each physical page. In addition, it may also include the page frame numbers of the physical pages, the total number of pages, the starting physical space address, the length, etc.
[0133] As an example, the first information may be passed to the memory hotplug management module 122 in the kernel to perform the above physical storage space management through the memory hotplug management module 122 in the kernel.
[0134] S430, send the management data to the CXL device, so that the CXL device can perform read and write operations according to the management data.
[0135] In this embodiment, the kernel of the computing device 100 can send management data back to the CXL device 20 through the CXL driver 121 for storage. In this way, when the CXL device 20 receives a read / write request subsequently, it can decode the address in the request based on the management data to determine the physical space address indicated in the request and perform read / write operations.
[0136] In some examples, if the storage space of the target device is managed by the kernel of the computing device 100, in addition to sending the management data back to the CXL device 20, it is also necessary to pass the management data to the kernel driver 123 of the kernel to map the physical storage space of the target device to the corresponding logical storage space and store the mapping relationship, so as to be used by corresponding kernel-mode or user-mode processes.
[0137] In other examples, if the storage space of the target device is managed by the FM30 of the computing device 100, after the FM30 senses the insertion of the target device, it can also obtain information such as the unique identification number and capacity of the target device from the CXL device 20, so as to divide the storage space of the target device into regions for use by computing devices in need in the computing cluster 10.
[0138] In this embodiment, the computing device 100 can, through a set of notification mechanisms with the CXL device 20 and with the support of relevant program modules running in relevant drivers and reused kernels, enable the kernel to identify and process the hot insertion of the CXL device 20 and its memory 22, realizing the hot memory expansion of the computing device 100. In this way, the computing device 100 connected to the CXL device 20 can debug and replace the CXL device 20 and the memory 22 without power-off and without interrupting business operations, greatly reducing the cluster operation and maintenance costs.
[0139] The following describes the process of hot plugging and unplugging the memory 22 on the CXL device 20 with reference to the accompanying drawings.
[0140] In one implementation, one computing device 100 is connected to one CXL device 20, and the memory 22 can be inserted into the CXL device 20 through a memory slot. Then, the operating system 101 of the computing device 100 is responsible for the hot plugging and unplugging management of the memory 22. Specifically, in the hot insertion stage of the memory 22, as Figure 5 shown, the method for hot plugging and unplugging the storage device can specifically include:
[0141] S510, receiving an insertion event through the CXL driver running on the computing device.
[0142] In this example, the insertion event is used to represent the hot insertion of the memory 22 (i.e., the target device) on the CXL device 20 connected to the computing device 100.
[0143] When the memory 22 is inserted into the CXL device 20 while the computing device 100 is not powered off, the CXL device 20 will, through the CXL firmware X23, pull up a link in the CXL mode at the downstream port following the CXL protocol, generating an interrupt signal. After the CXL controller 21 recognizes the interrupt signal, it will execute step S500 to detect the newly inserted memory 22, initialize the memory according to the CXL protocol (configure the DVSEC function), and read the hardware information of the memory 22. The hardware information includes the unique identification number and capacity of the memory 22, and may also include, but is not limited to, the bandwidth and latency of the memory 22.
[0144] Next, the CXL controller 21 will encapsulate the DVSEC initialization information, the hardware information, and the device information of the CXL device 20 (including but not limited to the unique identification number of the CXL device 20, the number of memories inserted on the CXL device 20, etc.) into an insertion event and transmit it to the computing device 100 side. The CXL driver 121 running on the computing device 100 can receive this insertion event.
[0145] S520, perform physical storage space management on the target device according to the capacity to obtain management data of the target device.
[0146] In this example, step S520 may specifically include the following S521 to S522:
[0147] In S521, the CXL driver transmits the first information to the memory hot-plug management module running on the computing device through the target communication method.
[0148] In this example, if the CXL driver 121 determines that the received event is an insertion event, it will continue to parse the first information from the insertion event, and then, through the above target communication method, the CXL driver 121 transmits the first information to the memory hot-plug management module 122 running in the kernel.
[0149] In some specific examples, the CXL driver 121 can directly communicate with the memory hot-plug module 1222 in the memory hot-plug management module 122 to transmit the first information to the memory hot-plug module 1222 for parsing and processing.
[0150] In some other examples, the CXL driver 121 can transfer the first piece of information to the ACPI module 1221 in the memory hot-plug management module 122 by interacting with the ACPI module 1221. Subsequently, the ACPI module 1221 generates an ACPI table based on the first piece of information to record the first piece of information in the ACPI table, and then transfers the ACPI table to the memory hot-plug module 1222. Since ACPI technology supports communication with the operating system 101 and the ACPI table is data directly accessible by the kernel, the memory hot-plug module 1222 can directly read it. Therefore, transferring information through the ACPI module 1221 helps reduce the implementation difficulty of the execution logic.
[0151] S522, through the memory hot-plug management module, uniformly address the physical storage space of the target device according to the capacity to obtain management data.
[0152] In this example, the memory hot-plug module 1222 learns that a memory 22 is inserted based on the obtained first piece of information. It will detect the capacity of the memory 22 according to the first piece of information, search for a suitable address space in the sparse memory model of the computing device 100 to adapt to the space of the memory 22, divide the space of the memory 22 into physical pages (page frames), and uniformly address them to obtain the page frame numbers (PFNs) of each physical page, as well as the physical space addresses and offsets of each physical page. Among them, the PFN is a globally unique identifier associated with the page frame. In this way, the memory hot-plug module 1222 initializes to obtain the management data of the memory 22. The management data includes the PFNs of each physical page of the memory 22, the starting PFN, the total number of pages, the offsets and physical space addresses of each physical page, the starting physical space address, and the length, etc., and may also include information such as the device status (such as the organizational structure of the physical page, usage information, etc.), but is not limited thereto.
[0153] Subsequently, the memory hot-plug module 1222 can transfer the management data of the newly inserted memory 22 to the memory driver 123 of the kernel of the operating system 101 through a callback function.
[0154] S530, through the memory driver, send the management data to the CXL device via the CXL driver running on the computing device.
[0155] In this step, the memory driver 123 transmits the management data back to the CXL device 20 via the CXL driver 121. After receiving the management data such as the physical space address and offset sent by the computing device 100, the CXL device 20 saves them through step S531, so that when receiving read / write operation requests from the computing device 100 later, it can decode according to the address included in the request and address the corresponding physical space address to perform the corresponding read / write operation.
[0156] S540. According to the management data, map the physical storage space of the target device to the logical storage space through the memory driver.
[0157] In this step, if the physical storage space (which can be partial or all) of the target device is used as the kernel space, the memory management module 1231 in the memory driver 123 can map the physical pages of the memory 22 that serve as the kernel space to the corresponding virtual pages according to the management data, generating the corresponding mapping relationship.
[0158] S550. Through the memory driver, create the memory management page and attribute file of the target device according to the logical storage space.
[0159] In this step, the memory management module 1231 of the memory driver 123 creates a new memory management table 125 according to this mapping relationship to complete the mapping management of the physical storage space of the memory 22 and the corresponding logical storage space. By way of example and not limitation, the memory management table 125 may include a kernel direct mapping table, virtual memory mapping pages, virtual memory mapping tables, etc., which facilitate the kernel to manage and use the storage space provided by the memory 22 based on these memory management tables 125. It can be understood that the virtual mapping page is used to record the virtual pages of the memory mapping table, the virtual memory mapping table is used to describe the mapping relationship between the physical pages and the corresponding virtual pages of the memory 22, and the kernel direct mapping table can be used to describe the memory area that the kernel can directly map.
[0160] Next, according to the logical memory space of the memory 22, the memory driver 123 creates a corresponding sysfs file for the memory 22 and creates a memory node (i.e., memory logical unit) corresponding to the memory 22 under the sysfs file. It can be understood that creating the sysfs file and memory node can be automatically completed through a script or can be completed using the probe tool in the kernel, which will not be elaborated here.
[0161] S560. Trigger the online operation of the logical storage space by writing the online command to the attribute file.
[0162] In this step, by writing "online" (online command) to the state file of the memory node created by the memory driver 123 under the sysfs file, the logical memory space of the memory 22 can be brought online in the operating system 101, making this logical memory space visible to users.
[0163] It can be understood that in this example, all virtual pages of the online logical memory space can be set to ZONE_NORMAL (direct mapping area, which is also the memory area directly mappable by the above kernel) or ZONE_MOVABLE (movable area), and handed over to the memory allocation system (buddy system) in the kernel for management. It can be understood that these memory pages can be set manually or through an automated script.
[0164] In this example, after setting the logical storage space of the memory 22 to be visible to users, this part of the space can be used by kernel-mode processes.
[0165] In other examples, if the physical storage space of the memory 22 is used as the user space, then during the execution of the S540 stage, the memory allocator 127 of the memory driver 123 will complete the mapping of this part of the physical storage space to the corresponding logical storage space according to the management data, record the mapping relationship, and then this part of the logical storage space can be used by user-mode processes.
[0166] So far, the hot plug-in process of the memory 22 on the CXL device 20 on the computing device 100 ends. In this way, the computing device 100 in the computing cluster 10 can realize the hot plug-in recognition and processing of physical memories without interrupting the business.
[0167] In this implementation, when it is necessary to remove the memory 22 on the CXL device 20 from the computing device 100, refer to Figure 6 , the method may specifically further include:
[0168] S610, receiving an uninstallation instruction through the memory driver.
[0169] In this example, when the computing cluster is running the business, the user can input an uninstallation instruction through the CXL command line interface 126 to indicate a hot uninstall operation on one or more specified memories 22 on the specified CXL device 20, and describe the unique identification number of this memory 22 in the instruction to be transmitted to the memory driver 123 running on the computing device 100.
[0170] S620, determining the usage status of the target device according to the unique identification number through the memory driver running on the computing device.
[0171] In this step, the memory driver 123 may first determine the usage status (free or occupied) of the memory 22 according to the unique identification number of the memory 22 to be unloaded.
[0172] Exemplarily, if the storage space of the memory 22 is in an idle state, the following step S630 is continued to be executed.
[0173] Exemplarily, if the storage space of the memory 22 is in an occupied state, the memory driver 123 reclaims the space through the following S621 and then executes S630:
[0174] At S621, when the target device is in an occupied state, all storage space of the target device is reclaimed through a memory driver to restore the target device to an idle state.
[0175] Exemplarily, if data or running processes are stored in the storage space of the memory 22, the data or processes are migrated to other memory spaces, so that the memory 22 enters an idle state. It can be understood that the memory 22 as a kernel space can be reclaimed by the memory management module 1231, and the memory 22 as a user space can be reclaimed by the memory allocator 127.
[0176] S630, when the target device is in an idle state, write an offline command to the attribute file of the target device to trigger the offline operation of the corresponding logical storage space. In this example, the memory driver 123 in the kernel searches for the corresponding sysfs file according to the unique identification number of the memory 22, and writes an "offline" offline command (such as "echo offline> / sys / devices / system / memory / memoryXXX / state" command) to the state file of the corresponding memory node under the sysfs file through the shell command "echo", so that the logical memory space provided by the memory 22 is invisible to the user.
[0177] S640: Release the mapping relationship between the physical storage space and the corresponding logical storage space through the memory driver, and clear the memory management page created according to the logical storage space.
[0178] In this step, the memory driver 123 releases the mapping relationship between the physical page and the virtual page of the memory 22, releases the kernel direct mapping table, releases the virtual memory mapping page, releases the virtual memory mapping table, and completes the unloading of the logical storage space of the memory 22.
[0179] Next, the memory hot-plug management module 122 running on the computing device is notified through the memory driver 123 to release the physical storage space of the target device.
[0180] It can be understood that S63O and S640 are executed when the physical storage space of the memory 22 is used as the kernel space.
[0181] When the physical storage space of the memory 22 is used as the user space, when the target device is in the idle state, the mapping relationship between the physical storage space and the corresponding logical storage space is released, and then the memory hot-plug management module 122 running on the computing device is notified through the memory driver 123 to release the physical storage space of the target device.
[0182] S650, through the memory hot-plug management module, releases the physical storage space of the target device according to the unique identification number and clears the management data.
[0183] In this step, the kernel memory driver 123 can pass the unique identification number of the memory 22 to the memory hot-plug module 1222 in the memory hot-plug management module 122 through a callback function or a system call. Then, the memory hot-plug module 1222 finds all the blocks (blocks) corresponding to the memory 22 in the sparse memory model according to the management data, resets the allocation status of these blocks to make them unavailable, and clears the relevant records (such as the physical page frame number and physical space address of the memory 22), completing the unloading of the physical storage space of the memory 22.
[0184] S660, the computing device passes the unloading message to the CXL device through the CXL driver to instruct the CXL device to clear the management data stored by itself.
[0185] In this example, the memory hot-plug management module 122 interacts with the CXL device 20 through the CXL driver 121, and sends an unloading message including the unique identification number of the memory 22 to be unloaded and the device information (unique identification number, communication port, etc.) of the CXL device 20 where it is located to the CXL device 20.
[0186] S670, the CXL device performs the unloading operation of the target device according to the unloading message.
[0187] In this example, the CXL device 20 detects whether it is responsible for managing the memory 22 according to the received unloading message. If not, the CXL device 20 returns a message indicating unloading failure to the CXL command line interface 126 to indicate that the target device is not found. If so, the CLX device 20 resets the corresponding HDM decoder and the root ports (RootPorts) to the computing device 100, enters the standard PCIe hot remove process, releases the corresponding CXL.io resources, thereby clearing the relevant records about the memory 22. After completion, a message indicating that the unloading is completed is returned to the CXL command line interface 126 module to notify the operation and maintenance personnel.
[0188] In this way, after seeing the returned message, the operation and maintenance personnel can pull out the memory 22 on the CXL device 20, complete the hot plugging out of the memory on the CXL device 20, which is convenient for debugging and replacing the memory on the CXL device, flexibly manage the memory capacity, and does not need to interrupt the operation of the computing device 100, reducing the operation and maintenance cost and the energy consumption cost.
[0189] In some possible implementation manners, the principle of hot plugging and hot unplugging the memory 22 as described above can also be used to perform hot plugging and hot unplugging on the entire CXL device. Among them, as Figure 7 shown, the process of hot plugging the entire CXL device 20 includes:
[0190] S710, receive the insertion event reported by the CXL device through the CXL driver running on the computing device.
[0191] In this example, when the computing cluster is running services, when the CXL device 20 (i.e., the target device) connected with the memory 22 accesses the computing device 100, the CXL device 20 is powered on. Then, step S700 is executed on the CXL device 20, and the CXL firmware X23 is powered on and runs. As described in the CXL protocol, it enters the power-on initialization process: configure the bus resources of the corresponding stack, trigger link training (perform adjustment and configuration of link signal quality, rate, link width, etc.), after the link training is successful, detect whether the stack space is in the CXL mode; if so, configure the CXL ports, enter the CXL accelerator (which can connect the CXL controller 21 and each memory 22, not marked in the figure) configuration space, detect the CXL controller 21 and initialize it, detect the CXL.memory and initialize it (including but not limited to configuring the DVSEC function), execute the CXL.IO enumeration process with reference to the PCIe enumeration process, and so on. In addition, a driver is also bound to the memory 22 on the CXL device 20, and the power-on initialization process ends.
[0192] After the initialization process, the CXL device 20 encapsulates the hardware information (including but not limited to capacity, bandwidth, latency, etc.) of each memory 22 connected thereto, together with the device information (including but not limited to the unique identification number of the CXL device 20, the number of memories inserted on the CXL device 20, etc.) into an insertion event, and sends it to the CXL driver 121 at the computing device 100 end.
[0193] S720, perform physical storage space management on the target device according to the capacity to obtain the management data of the target device.
[0194] In this step, if the CXL driver 121 determines that the received event is an insertion event of a newly added CXL device, it continues to parse the first information from the insertion event. Then, based on the first information, a CXL node of the CXL device 20 is created at the computing device 100. It can be understood that this CXL node is the file record of the CXL device 20.
[0195] Next, the CXL driver 121 passes the first information to the memory hotplug management module 122 for processing similar to S520 in the above example, completes the allocation and unified addressing of the physical storage spaces of all the memories 22 on the CXL device 20, and thus generates corresponding management data and returns it to be stored on the CXL device 20, so as to perform corresponding read and write operations according to the management data. It can be understood that the process of returning the management data to the CXL device 20 can also refer to the description of S530 in the above example.
[0196] Moreover, operations similar to S540 to S560 in the above example are also performed on the computing device 100 to map the physical storage spaces of the respective memories on the CXL device 20 to corresponding logical storage spaces for use by processes in the user space or kernel space, which will not be elaborated here either.
[0197] So far, the operation process of hot-inserting the CXL device 20 into the computing device 100 ends.
[0198] In this implementation, if it is necessary to unload the CXL device 20 from the computing device 100, the unloading process is generally similar to the unloading process shown in the above Figure 6 example. The main difference is that in this implementation, when the CXL device 20 is to be unplugged from the computing device 100, the user can input an unloading instruction (including the unique identification number of the CXL device 20) through the CXL command-line interface 126 to indicate the unloading of the specified CXL device 20. And, in the stage where the CXL device 20 performs the unloading operation of the target device according to the unloading message, the CXL device 20 will detect whether it is the device described in the unloading message. If not, the CXL device 20 notifies the CXL driver 121 of the computing device 100 that the unloading fails and the target device is not found. If so, the CXL device 20 enters the power-off process and returns a message indicating that the unloading is completed to the CXL command-line interface module 126 to notify the operation and maintenance personnel, so as to unplug the CXL device 20.
[0199] In some implementations, multiple computing devices 100 in the computing cluster 10 are connected to a CXL device 20, and memories 22 can be inserted through memory slots on the CXL device 20. Then, the FM30 (deployed on one computing device 100) running in the computing cluster 10 is responsible for the hotplug management of the memories 22. Specifically, in the hot-insertion stage of the memories 22, as Figure 8As shown, the method for hot plugging the storage device may specifically include:
[0200] S810, receiving an insertion event reported by a CXL device through a CXL driver running on a computing device;
[0201] S820, performing physical storage space management on a target device according to its capacity to obtain management data of the target device;
[0202] S830, sending the management data to the CXL device via the memory driver through the CXL driver running on the computing device.
[0203] In this example, the execution principles of S810 to S830 are similar to the relevant descriptions of steps S510 to S530 in the above example. The main difference is that in this example, multiple computing devices 100 are connected to one CXL device 20. For each computing device 100, physical storage space management is performed on the memory 22 inserted on the CXL device 20 to obtain a set of management data. Moreover, the management data returned by each computing device 100 to the CXL device 20 is associated with the identity identifier of the computing device 100 itself. On the CXL device 20 side, the management data returned by each computing device 100 is associated and stored together with the corresponding computing device identity identifier through step S831. In this way, when any computing device 100 sends a read / write request, the CXL device 20 uses the corresponding management data to decode the address in the request based on the computing device identity identifier carried in the request, so as to find the specific physical storage address to perform the read / write operation.
[0204] S840, when a structure manager FM running on the computing device senses that the target device is hot - inserted into the computing device, obtaining second information of the target device from the CXL device.
[0205] In this step, when FM30 senses that the memory 22 is inserted into the computing device 100, a DAX file (a file that supports direct access by user - state software) is created for the memory 22 through the DAX driver, and the list interface in the CXL command - line interface module 126 is called to instruct the CXL device 20 where the newly inserted memory 22 is located to report device information, as well as information such as the unique identification number and capacity of the memory 22 (i.e., second information). It can be understood that when obtaining this information and data, the CXL command - line interface module 126 can obtain them from the CXL device 20 by running a script to call the CXL driver 121.
[0206] S850, through the FM, dividing the physical storage space of the target device into multiple physical regions according to the second information for use by multiple computing devices in the cluster where the computing device is located.
[0207] In this step, FM30 divides the physical storage space of the memory 22 according to the capacity of the memory 22, obtains multiple physical regions, and updates the mapping table 132 to record the mapping information about the memory 22. As an example, the mapping table 132 may at least include two parts. One part is used to manage the physical information of the target device (the memory 22 in this example), including but not limited to the unique identification number of the target device, the unique identification number of the CXL device where it is located, the corresponding relationship with the CXL device (such as in which memory slot), the capacity, and so on. The other part of the mapping table 132 is used to manage the physical region management information of the target device, including but not limited to the region information (region number) of the divided physical regions, the capacity corresponding to each physical region, the server IP allocated to the computing cluster, and so on.
[0208] In this example, when each computing device 100 uses the physical region provided by the memory 22 allocated by FM30, it can determine the physical space address range corresponding to this physical region based on the management data of the memory 22 stored locally in the computing device 100.
[0209] In some possible examples, the first information may further include the offset and starting physical space address of the memory 22. Since the management data after addressing the memory 22 by each computing device 100 is different, these offset and starting physical space addresses are also associated with the identity identifier of the corresponding computing device 100. In this way, when FM allocates a physical region to a certain computing device 100, it can also inform the computing device 100 of the corresponding offset and starting physical space address of this physical region on the memory 22 according to the identity identifier of the computing device 100.
[0210] For example. Suppose the offset in the management data of the memory 22 by the computing device A is A1 and the starting physical space address is A2, the offset in the management data of the memory 22 by the computing device B is B1 and the starting physical space address is B2, and FM30 divides the memory 22 into region C and region D. If region C is allocated to the computing device A, then FM30 will inform the computing device A of the offset A1, the starting physical space address A2, the capacity and number of region C on the memory 22, and so on. In this way, the computing device A can determine which part of the space of the memory 22 the allocated region C corresponds to based on the management data of the memory 22 stored by itself, and then use it. Similarly, if region C is allocated to the computing device B, then FM30 will inform the computing device B of the offset B1, the starting physical space address B2, the capacity and number of region C on the memory 22, and so on.
[0211] In this implementation, when it is necessary to hot-remove the memory 22 from the computing cluster 10, the FM30 performs hot-unloading management on the memory 22, then as Figure 9 shown, the method may further include:
[0212] S910, receiving an unloading instruction through the FM.
[0213] In this example, when the computing cluster is running services, the user can input an unloading instruction through the CXL command-line interface module 126, indicating to perform a hot-unloading operation on one or more specified memories 22 on the specified CXL device 20, and describe the unique identification number of the memory 22 in the instruction to be transmitted to the FM30 running on the computing device 100.
[0214] S920, through the FM, determining the usage status of the target device according to the unique identification number.
[0215] In this example, the FM30 may first determine the usage status (idle or occupied) of the memory 22 according to the unique identification number of the memory 22 to be unloaded.
[0216] Exemplarily, if the storage space of the memory 22 is in an idle state, the following step S930 is continued.
[0217] Exemplarily, if the storage space of the memory 22 is in an occupied state, after the FM30 performs space recovery through the following S921, S930 is then executed:
[0218] In S921, when the target device is in an occupied state, the FM recovers the physical storage space of the target device from the target computing device that occupies the target device, so that the target device returns to an idle state.
[0219] In this step, the FM30 can determine the target computing device that occupies the storage space of the memory 22 by checking the mapping table 132 and obtain the identity information (such as IP address information) of these target computing devices. It can be understood that the target computing device can be the computing device 100 where the FM30 is located, or a device located in the same cluster as the computing device 100.
[0220] Then, the FM30 sends space recovery information according to the IP address of the target computing device to recover the storage space provided by the memory 22 from the target computing device. The target computing device will transfer the data stored in the memory 22 to other idle storage spaces, so that the memory 22 returns to an idle state.
[0221] S930, in the case where the target device is in an idle state, the structure manager clears the record about the target device.
[0222] In this example, after FM30 determines that the storage space of memory 22 is idle, FM30 calls the DAX-driven interface to clear the DAX file corresponding to memory 22, and then updates mapping table 132 to clear the relevant records regarding memory 22, such as records of the space partitioning of this memory 22, records of being allocated for use by the target server, etc., thereby preventing the storage space of this memory 22 from being reallocated. For memory 22 whose relevant records have been cleared, FM30 no longer performs any operations on it.
[0223] S940, when the target device is in an idle state, release the physical storage space of the target device according to the unique identification number and clear the management data.
[0224] In this example, FM30 can communicate with the kernel of operating system 101 of computing device 100 to notify memory driver 123 to perform a hot-unloading operation on offload memory 22. Then, memory driver 123 can pass the unique identification number of memory 22 to memory hot-plug module 1222 in memory hot-plug management module 122 through a callback function or system call, etc. Subsequently, memory hot-plug module 1222 locates all the blocks (blocks) corresponding to memory 22 in the sparse memory model according to the management data, resets the allocation status of these blocks to make them unavailable, and clears the relevant records (such as the physical page frame number and physical space address of memory 22), thus completing the unloading of the physical storage space of memory 22.
[0225] S950, the computing device passes an unloading message to the CXL device through the CXL driver to instruct the CXL device to clear the management data stored by itself;
[0226] S960, the CXL device performs the unloading operation of the target device according to the unloading message.
[0227] In this example, the execution principle of S950 to S960 is similar to the principle of S660 to S670 in the above example, and will not be elaborated here.
[0228] So far, the management of the hot-insertion process of memory 22 on CXL device 20 in computing cluster 10 ends.
[0229] It can be understood that similar to the principle of FM30's hot-plug control of memory 22, the hot-plug of the entire CXL device 20 can also be controlled by FM30, which will not be elaborated here.
[0230] Based on the method in the above embodiments, an embodiment of the present application provides a storage device hot-plug device. Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a device provided by an embodiment of the present application.
[0231] As Figure 10 shown, the hot pluggable storage device 900 may include: a receiving module 901 and a processing module 902. Among them, the receiving module 901 is configured to: receive an insertion event, where the insertion event is an event that a target device is hot inserted into a computing device, and the insertion event includes first information, and the first information includes a unique identification number and a capacity of the target device. The target device is a Compute Express Link (CXL) device or a memory connected to a CXL device; the processing module 902 is configured to: perform physical storage space management on the target device according to the capacity to obtain management data of the target device, where the management data includes an offset and a physical space address; the processing module 902 is further configured to send the management data to the CXL device, so that the CXL device can perform read and write operations according to the management data..
[0232] It should be understood that the above device is used to execute the method in the above embodiment. For the corresponding program modules in the device, their implementation principles and technical effects are similar to those described in the above method. The working process of the device can refer to the corresponding process in the above method, which will not be elaborated here.
[0233] Based on the method in the above embodiment, an embodiment of the present application provides an electronic device. The electronic device may include: at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, when the program stored in the memory is executed, the processor is configured to execute the method in the above embodiment.
[0234] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program runs on a processor, it causes the processor to execute the method in the above embodiment.
[0235] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product, characterized in that when the computer program product runs on a processor, it causes the processor to execute the method in the above embodiment.
[0236] Based on the method in the above embodiment, an embodiment of the present application further provides a chip. Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a chip provided by an embodiment of the present application. As Figure 11 shown, the chip 1100 includes one or more processors 1101 and an interface circuit 1102. Optionally, the chip 1100 may further include a bus 1103. Among them:
[0237] The processor 1101 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1101 or the instructions in the form of software. The above-mentioned processor 1101 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0238] The interface circuit 1102 can be used for sending or receiving data, instructions, or information. The processor 1101 can utilize the data, instructions, or other information received by the interface circuit 1102 for processing, and can send the processed information through the interface circuit 1102.
[0239] Optionally, the chip 1100 further includes a memory. The memory may include a read-only memory and a random access memory, and provide operation instructions and data to the processor. A part of the memory may also include a non-volatile random access memory (NVRAM).
[0240] Optionally, the memory stores executable software modules or data structures. The processor can execute corresponding operations by calling the operation instructions stored in the memory (the operation instructions can be stored in the operating system).
[0241] Optionally, the interface circuit 1102 can be used to output the execution result of the processor 1101.
[0242] It should be noted that the respective functions corresponding to the processor 1101 and the interface circuit 1102 can be implemented through hardware design, can also be implemented through software design, or can be implemented in a combination of software and hardware. There is no limitation here.
[0243] It should be understood that each step of the above method embodiment can be completed by the logic circuit in the form of hardware in the processor or the instructions in the form of software.
[0244] It can be understood that the magnitudes of the sequence numbers of the respective steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the respective steps in the above embodiments can be selectively executed according to the actual situation, can be partially executed, or can be fully executed. There is no limitation here.
[0245] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0246] The method steps in the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), register, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.
[0247] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0248] It can be understood that the various digital numbers involved in the embodiments of the present application are only for convenience of description and are not used to limit the scope of the embodiments of the present application.
Claims
1. A method for hot plugging a storage device, characterized in that, the method runs on a computing device, and the method includes: receiving an insertion event, where the insertion event is an event that a target device is hot plugged into the computing device, and the insertion event includes first information, and the first information includes a unique identification number and a capacity of the target device, and the target device is a Compute Express Link (CXL) device or a memory connected to the CXL device; performing physical storage space management on the target device according to the capacity to obtain management data of the target device, where the management data includes an offset and a physical space address; sending the management data to the CXL device, so that the CXL device can perform read and write operations according to the management data.
2. The method according to claim 1, characterized in that, the performing physical storage space management on the target device according to the capacity to obtain management data of the target device includes: the CXL driver running on the computing device transfers the first information to a memory hot plug management module running on the computing device through a target communication method, and the memory hot plug management module is a program for supporting memory hot plug; through the memory hot plug management module, uniformly addressing the physical storage space of the target device according to the capacity to obtain the management data.
3. The method according to claim 1 or 2, characterized in that, after obtaining the management data of the target device, the method includes: transferring the management data to a memory driver running on the computing device through the memory hot plug management module; through the memory driver, sending the management data to the CXL device via the CXL driver running on the computing device.
4. The method according to claim 3, characterized in that, after sending the management data to the CXL device, the method includes: when a Fabric Manager (FM) running on the computing device senses that the target device is hot plugged into the computing device, obtaining second information of the target device from the CXL device, and the FM is a program for managing the CXL device, and the second information includes a unique identification number and a capacity of the target device; through the FM, partitioning the physical storage space of the target device according to the second information to obtain a plurality of physical regions for use by a plurality of computing devices in the cluster where the computing device is located.
5. The method according to claim 3, characterized in that, after sending the management data to the CXL device, the method includes: mapping the physical storage space to a logical storage space according to the management data through the memory driver; through the memory driver, creating a memory management page and an attribute file of the target device according to the logical storage space, where the memory management page is used to manage the mapping relationship between the physical storage space and the logical storage space, and the attribute file is used to provide information of the logical storage space for use by the application layer. An online operation on the logical storage space is triggered by writing an online command to the attribute file.
6. The method according to any one of claims 1 to 3, It is characterized in that The method further comprises: receiving an uninstall instruction, the uninstall instruction being used to instruct hot unplugging the target device from the computing device, the uninstall instruction including a unique identification number of the target device; When the target device is in an idle state, releasing the physical storage space of the target device according to the unique identification number, and clearing the management data; An uninstall message is sent to the CXL device to instruct the CXL device to clear the management data stored in the CXL device.
7. The method according to claim 6, It is characterized in that Before releasing the physical storage space of the target device according to the unique identification number, the method includes: Determining, by the configuration manager FM of the computing device, a usage state of the target device according to the unique identification number, wherein the usage state at least includes an idle state or an occupied state; When the target device is in an occupied state, the physical storage space of the target device is reclaimed from the target computing device occupying the target device through the FM, so that the target device is restored to an idle state. The target computing device is the computing device or is in the same cluster as the computing device.
8. The method according to claim 7, It is characterized in that Before releasing the physical storage space of the target device according to the unique identification number, the method includes: Determining, by a memory driver running on the computing device, a usage state of the target device according to the unique identification number, the usage state at least including an idle state or an occupied state; When the target device is in an occupied state, all storage space of the target device is reclaimed through the memory driver to restore the target device to an idle state.
9. The method according to claim 8, It is characterized in that Before releasing the physical storage space of the target device according to the unique identification number, the method includes: When the target device is in an idle state, a logoff command is written into a property file of the target device to trigger a logoff operation on the corresponding logical storage space, wherein the property file is used to provide information of the logical storage space to an application layer for use; Removing the mapping relationship between the physical storage space and the corresponding logical storage space through the memory driver, and clearing the memory management page created according to the logical storage space; The memory driver notifies a memory hot-plug management module running on the computing device to release the physical storage space of the target device.
10. A computing device, It is characterized in that include: at least one memory for storing a program; at least one processor, configured to execute the program stored in the memory; Wherein, when the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-9.
Citation Information
Cited By
Hot plug control system and method, electronic equipment and storage medium
CN120508521A
Hot-swap control system and method, electronic device, and storage medium
CN120508521B
Hot plug control device and method
CN120670350A
Hot plug control apparatus and method
CN120670350B
Storage device, electronic device, data storage method and storage medium
CN120723169A