Hot-swap method for storage device and device
By implementing the hot-swap management method on the computing device, the hot-swap problem of the CXL equipment cannot be identified and managed in the prior art under constant power, and efficient hot-swap and uninstall of the equipment is realized, reducing maintenance costs.
Patent Information
- Application Number
- PCT/CN2024/099552
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-11
- Filing Date
- 2024-06-17
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art cannot identify and manage hot-swap of CXL-based physical storage devices without shutting down the computing device, resulting in power failure when replacing the storage device, increasing maintenance and testing costs.
By implementing a hot-swap management method on a computing device, receiving insertion events, managing the physical storage space of the target device, generating management data, and sending it to a CXL device, enabling it to perform read and write operations.
It realizes hot plugging and hot unloading of CXL devices or memory while the computing device is not shut down, reducing maintenance and testing costs and improving the reliability of business operations.
Smart Images

Figure CN2024099552_19062025_PF_FP_ABST
Abstract
Description
Storage device hot-swap method and device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 11, 2023, with application number 202311695013.0 and application name “A method and device for hot-swapping storage devices”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of storage technology, and in particular to a method and device for hot-plugging a storage device. Background Art
[0003] Compute Express Link (CXL) is a high-speed, high-capacity standard for connecting central processing units (CPUs) to devices and CPUs to memory, specifically designed for high-performance data center computers. CXL-based expansion storage devices, referred to as CXL devices, can effectively increase the storage capacity of computing devices. However, computing devices typically only recognize CXL device memory during cold-swap operations and report this to the computing device's operating system. This means that replacing a CXL device or its memory requires powering off the computing device, interrupting service operations and increasing maintenance and commissioning costs.
[0004] Summary of the Invention
[0005] Embodiments of the present application provide a storage device hot-swap method, apparatus, computing device, computer storage medium, and computer program product, which can implement hot-swap of physical storage devices based on CXL.
[0006] In a first aspect, an embodiment of the present application provides a method for hot-plugging a storage device, which runs on a computing device and includes: receiving an insertion event, where the insertion event is an event of hot-plugging a target device into the computing device, the insertion event including first information, where the first information includes a unique identification number and capacity of the target device, where the target device is a computing quick connect (CXL) device or a storage device connected to the CXL device; managing the physical storage space of the target device based on the capacity to obtain management data of the target device, where the management data includes an offset and a physical space address; and sending the management data to the CXL device, so that the CXL device can perform read and write operations based on the management data.
[0007] In this embodiment, hot-swapping refers to connecting a computing device to a CXL device or inserting storage (such as a memory stick) into a connected CXL device without shutting down. The computing device listens for insertion events reported by the CXL device and performs hot-swapping processing on the target device, including managing the target device's physical storage space, obtaining corresponding management data, and returning it to the CXL device. This allows the CXL device to decode read and write requests from the computing device based on the stored management data, locating the specific physical space address targeted by the request and performing the read and write operations. In this embodiment, the computing device can manage the insertion of physical storage devices via CXL without interrupting service, reducing maintenance and commissioning costs.
[0008] In some possible examples, receiving the insertion event includes: receiving, by a CXL driver running on the computing device, the insertion event reported by the CXL device.
[0009] In this example, a CXL driver is deployed on the computing device to communicate with the CXL device, thereby being able to detect and identify event information reported by the CXL device and pass the event information to the computing device operating system kernel for processing.
[0010] In some possible examples, physical storage space of a target device is managed according to capacity to obtain management data of the target device, including: a CXL driver running on the computing device transmits first information to a memory hot-swap management module running on the computing device through a target communication method, where the memory hot-swap management module is a program for supporting memory hot-swap; and through the memory hot-swap management module, the physical storage space of the target device is uniformly addressed according to capacity to obtain management data.
[0011] In this example, the target communication method may include, but is not limited to, function calls, process communication, and system calls. The computing device reuses the kernel's memory hot-swap management module to manage the hot insertion of CXL-based physical storage devices. The CXL driver interacts with the memory hot-swap management module to exchange messages. This communication mechanism enables the operating system to perceive and process the hot insertion of CXL-based physical storage devices, enabling the addition of CXL-based physical storage devices without shutting down the computing device, thus achieving hot memory expansion.
[0012] In some possible examples, after obtaining management data of the target device, the method includes: passing the management data to a memory driver running on the computing device through a memory hot-swap management module; and sending the management data to the CXL device via a CXL driver running on the computing device through the memory driver.
[0013] In this way, the memory driver in the kernel obtains the management data of the target device and returns it to the corresponding CXL device. The kernel can then use the storage space of the target device normally based on the management data and request the CXL device to perform corresponding read and write operations.
[0014] In some possible examples, after sending the management data to the CXL device, the method includes: when a fabric manager FM running on the computing device senses that a target device has been hot-plugged into the computing device, obtaining second information of the target device from the CXL device, where the FM is a program for managing CXL devices, and the second information includes a unique identification number and capacity of the target device; and using the FM, dividing the physical storage space of the target device into regions based on the second information to obtain multiple physical regions, to provide for use by multiple computing devices in the cluster where the computing device is located.
[0015] In this example, if the CXL device acts as a DAX device and is managed by the FM deployed on the computing device, when the FM senses that a target device has been hot-plugged, it can request relevant information from the CXL device to divide the physical storage space of the target device into blocks, and obtain multiple physical areas to allocate to each computing device in the cluster where the computing device is located, thereby realizing resource sharing of the memory provided by the CXL physical storage device in the cluster.
[0016] In some possible examples, after sending the management data to the CXL device, the method includes: mapping the physical storage space to the logical storage space through the memory driver according to the management data; creating a memory management page and a property file of the target device according to the logical storage space through the memory driver, the memory management page is used to manage the mapping relationship between the physical storage space and the logical storage space, and the property file is used to provide the information of the logical storage space to the application layer for use; and triggering the online operation of the logical storage space by writing an online command to the property file.
[0017] In this example, if the CXL device serves as the extended memory of the computing device, the memory driver running in the kernel of the computing device's operating system can further process the management data passed by the memory hot-plug management module, map the physical storage space of the target device to the corresponding logical storage space, and perform corresponding page table management and online operations to provide it to the application layer process for use.
[0018] In some possible examples, the method further includes: receiving an uninstall instruction for instructing hot unplugging the target device from the computing device, the uninstall instruction including a unique identification number of the target device; when the target device is in an idle state, releasing physical storage space of the target device according to the unique identification number and clearing management data; and sending an uninstall message to the CXL device to instruct the CXL device to clear the management data stored therein.
[0019] In this example, the computing device can also receive an uninstall instruction through the CXL driver. The uninstall instruction can be manually input into the computing device through a command line. Thus, the computing device can release the physical storage space of the target device to be uninstalled from the sparse memory model according to the uninstall instruction, clear the relevant management data, and instruct the CXL device to synchronize and update to complete the hot uninstall operation without interrupting the operation of the computing device, which is conducive to maintenance.
[0020] In some possible examples, before releasing the physical storage space of the target device according to the unique identification number, the method includes: determining the usage status of the target device according to the unique identification number through the structure manager FM of the computing device, the usage status including at least an idle state or an occupied state; when the target device is in the occupied state, reclaiming the physical storage space of the target device from the target computing device occupying the target device through FM so that the target device is restored to the idle state, and the target computing device is a computing device or is in the same cluster as the computing device.
[0021] In this example, when the FM deployed on the computing device is responsible for managing the CXL device, before hot-uninstalling the target device, the FM needs to determine the usage status of the target device and perform operations such as space reclamation, so that the hot removal operation process is performed on the target device in the idle state, which is conducive to ensuring the stable operation of the cluster business.
[0022] In some possible examples, before releasing the physical storage space of the target device according to the unique identification number, the method includes: determining the usage status of the target device according to the unique identification number through a memory driver running on the computing device, where the usage status includes at least an idle state or an occupied state; when the target device is in the occupied state, reclaiming all the storage space of the target device through the memory driver to restore the target device to the idle state.
[0023] In this example, the CXL device is managed by the computing device kernel, and the memory driver is responsible for determining the usage status of the target device before hot unloading the target device and performing operations such as space reclamation to ensure stable operation of processes on the computing device.
[0024] In some possible examples, before releasing the physical storage space of the target device based on the unique identification number, the method includes: when the target device is in an idle state, triggering an offline operation on the corresponding logical storage space by writing an offline command to the target device's attribute file, and the attribute file is used to provide information about the logical storage space to the application layer for use; releasing the mapping relationship between the physical storage space and the corresponding logical storage space through the memory driver, and clearing the memory management page created based on the logical storage space; and notifying the memory hot-swap management module running on the computing device through the memory driver to release the physical storage space of the target device.
[0025] In a second aspect, an embodiment of the present application provides a storage device hot-plugging device, which includes: a receiving module and a processing module, wherein the receiving module is used to: receive an insertion event, where the insertion event is an event in which a target device is hot-plugged into a computing device, and the insertion event includes first information, where the first information includes a unique identification number and capacity of the target device, and the target device is a computing quick connect CXL device or a memory connected to the CXL device; the processing module is used to: manage the physical storage space of the target device according to the capacity, and obtain management data of the target device, where the management data includes an offset and a physical space address; and send the management data to the CXL device, so that the CXL device can perform read and write operations based on the management data.
[0026] In some possible examples, the receiving module is specifically configured to receive, through a CXL driver running on the computing device, an insertion event reported by the CXL device.
[0027] In some possible examples, the processing module is specifically used to: enable the CXL driver running on the computing device to transmit the first information to the memory hot-swap management module running on the computing device through the target communication method, the memory hot-swap management module being a program used to support memory hot-swap; through the memory hot-swap management module, the physical storage space of the target device is uniformly addressed according to the capacity to obtain management data.
[0028] In some possible examples, the processing module is further configured to: transmit the management data to a memory driver running on the computing device through the memory hot-swap management module; and send the management data to the CXL device through the CXL driver running on the computing device through the memory driver.
[0029] In some possible examples, the processing module is further used to: enable a fabric manager FM running on the computing device to sense that a target device is hot-plugged into the computing device, and obtain second information of the target device from the CXL device, where the FM is a program for managing CXL devices, and the second information includes a unique identification number and capacity of the target device; and divide the physical storage space of the target device into regions based on the second information through the FM to obtain multiple physical regions for use by multiple computing devices in the cluster where the computing device is located.
[0030] In some possible examples, the processing module is also used to: map the physical storage space to the logical storage space through the memory driver according to the management data; create the memory management page and attribute file of the target device according to the logical storage space through the memory driver, the memory management page is used to manage the mapping relationship between the physical storage space and the logical storage space, and the attribute file is used to provide the information of the logical storage space to the application layer for use; and trigger the online operation of the logical storage space by writing the online command to the attribute file.
[0031] In some possible examples, the receiving module is further configured to receive an uninstall instruction, the uninstall instruction being used to instruct hot removal of the target device from the computing device, the uninstall instruction including a unique identification number of the target device; the processing module is further configured to release physical storage space of the target device and clear management data based on the unique identification number when the target device is in an idle state; and send an uninstall message to the CXL device to instruct the CXL device to clear the management data stored therein.
[0032] In some possible examples, the processing module is also used to: determine the usage status of the target device according to the unique identification number through the structure manager FM of the computing device, and the usage status includes at least an idle state or an occupied state; when the target device is in the occupied state, reclaim the physical storage space of the target device from the target computing device occupying the target device through FM so that the target device is restored to the idle state, and the target computing device is a computing device or is in the same cluster as the computing device.
[0033] In some possible examples, before releasing the physical storage space of the target device according to the unique identification number, the method includes: determining the usage status of the target device according to the unique identification number through a memory driver running on the computing device, where the usage status includes at least an idle state or an occupied state; when the target device is in the occupied state, reclaiming all the storage space of the target device through the memory driver to restore the target device to the idle state.
[0034] In some possible examples, the processing module is also used to: when the target device is in an idle state, trigger an offline operation on the corresponding logical storage space by writing an offline command to the target device's attribute file, and the attribute file is used to provide information about the logical storage space to the application layer for use; release the mapping relationship between the physical storage space and the corresponding logical storage space through the memory driver, and clear the memory management page created based on the logical storage space; notify the memory hot-swap management module running on the computing device through the memory driver to release the physical storage space of the target device.
[0035] In a third aspect, an embodiment of the present application provides an electronic device, comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory; wherein, when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.
[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0037] In a fifth aspect, an embodiment of the present application provides a computer program product, characterized in that when the computer program product runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0038] In the sixth aspect, an embodiment of the present application provides a chip, characterized in that it includes at least one processor and an interface; at least one processor obtains program instructions or data through the interface; and at least one processor is used to execute program line instructions to implement the method described in the first aspect or any possible implementation of the first aspect.
[0039] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] FIG1 is a schematic diagram of an application scenario of an embodiment of the present application;
[0041] FIG2 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0042] FIG3A is a schematic diagram of a flow chart of hot-plug control of a memory on a CXL device by a computing device core according to an embodiment of the present application;
[0043] FIG3B is a schematic diagram of a process for controlling hot unloading of memory on a CXL device by a kernel in an embodiment of the present application;
[0044] FIG3C is a flow chart of FM controlling hot insertion of memory on a CXL device according to an embodiment of the present application;
[0045] FIG3D is a schematic diagram of a process flow of FM controlling hot unloading of memory on a CXL device in an embodiment of the present application;
[0046] FIG4 is a flow chart of a method for hot-swapping a storage device according to an embodiment of the present application;
[0047] FIG5 is a schematic diagram of a process for hot-plugging a target device in a specific embodiment of the present application;
[0048] FIG6 is a schematic diagram of a process of hot unloading a target device in a specific embodiment of the present application;
[0049] FIG7 is a schematic diagram of a process of hot-plugging a target device in a specific embodiment of the present application;
[0050] FIG8 is a schematic diagram of a process of hot-plugging a target device in a specific embodiment of the present application;
[0051] FIG9 is a schematic diagram of a process of hot unloading a target device in a specific embodiment of the present application;
[0052] FIG10 is a schematic structural diagram of a storage device hot-swap device provided in an embodiment of the present application;
[0053] FIG11 is a schematic structural diagram of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.
[0055] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.
[0056] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0057] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0058] To facilitate understanding of the technical solutions of the embodiments of the present application, the technical terms and abbreviations involved in this document are explained below.
[0059] The CXL protocol, also known as the Compute Express Link protocol, is a high-speed serial protocol that enables fast and reliable data transmission between different components within a computing device, such as the central processing unit (CPU) and accelerators, memory buffers, and smart network interface cards (Smart NICs). CXL protocols include CXL.io, CXL.cache, and CXL.memory.
[0060] The fabric manager (FM) manages the CXL memory space of the memory pool composed of one or more CXL devices. It is responsible for sharding this CXL memory space, allocating it to various compute devices, and recording mapping information. The FM can be deployed on any compute device or on a device such as a CXL switch. Only one compute device, CXL switch, or other device can run the FM at a time.
[0061] The Basic Input / Output System (BIOS) is a set of programs stored in a read-only memory (ROM) on the computer's motherboard. It is the first program to run upon startup. It includes the computer's most important basic input / output (BIO) programs, the post-boot self-test program, and the system startup program. The BIOS can also read and write detailed system configuration information. Its primary function is to provide the lowest-level, most direct hardware configuration and control for the computer.
[0062] Advanced Configuration and Power Interface (ACPI): An open standard for operating system power management and hardware configuration. ACPI defines a hardware abstraction interface between system firmware (BIOS or Unified Extensible Firmware Interface (UEFI)) and the operating system.
[0063] Hot swapping: plugging and unplugging physical devices (such as memory, solid-state drives, etc.) without shutting down the computing device (such as a server).
[0064] Memory hot-plug: The Linux kernel supports memory hot-plugging. This allows adding or removing physical memory while a computing device is running and the operating system is operating normally.
[0065] Firmware (FW): A program written to an erasable programmable read-only memory (EPROM) or electrically erasable programmable read-only memory (EEPROM). Firmware is the device's internal "driver program." It enables the operating system to operate a specific device according to standard device drivers.
[0066] Command line interface (CLI): A user interface that allows users to enter commands through the keyboard, and the computer executes the commands after receiving them.
[0067] Baseboard management controller (BMC): This controller performs a range of monitoring and control functions for system hardware. For example, it monitors system temperature, voltage, fans, power supplies, and other functions, making adjustments to ensure a healthy system. It also records hardware information and logs for user notifications and troubleshooting. The BMC is an independent system that communicates with other server hardware (such as the CPU and memory) through physical channels and interacts with the BIOS and operating system (OS).
[0068] Page frame number (PFN): Physical memory is divided into fixed-size partitions called page frames (or page frames). Page frames are numbered as page frame numbers (or page frame numbers, memory block numbers), starting at 0. A certain number of page frames form a memory segment.
[0069] Designated vendor-specific extended capability (DVSEC): This capability defines a configuration register structure that can be implemented by vendors while providing a consistent hardware / software interface. DVSEC can provide hints in the DVSEC capability definition to allow system software to determine whether a hot-pluggable port is extendable and indicate the number of busses that need to be reserved.
[0070] Memory allocation system (buddy): It is the memory allocation mechanism in the Linux system kernel. It organizes physical memory page frames, reasonably allocates and recycles physical memory pages, and allows memory allocation and adjacent memory merging to be carried out quickly, which is used to alleviate memory fragmentation.
[0071] List interface: is a common data structure interface that provides a set of methods for accessing, inserting, deleting, and traversing list elements.
[0072] Host-managed device memory (HDM): The CXL.memory protocol addresses device memory uniformly across the system, making it appear to the server's (host) CPU as if it were its own main memory and directly accessible. This uniformly addressed device memory is called HDM.
[0073] Typically, computing devices can only identify CXL-based devices during a cold boot. A CXL device (including a CXL controller and one or more memory devices connected to it) is installed on the computing device and initialized during a cold boot. At this point, the BIOS detects the CXL device, obtains and records its status information, including its memory configuration, and passes this information to the kernel at the end of the BIOS lifecycle. This allows the kernel to detect the CXL device and perform the corresponding memory space mapping upon boot. However, this cold-swap technology does not support swapping out or replacing partial memory or entire boards on a CXL device without powering down the computing device. This means that replacing a CXL device or its memory requires a power outage, disrupting service operations. Large computing devices, such as servers, typically require 5 to 10 minutes to power down and restart. Therefore, this cold-swap technology significantly increases maintenance and commissioning costs, significantly impacting service operations, particularly in cluster computing scenarios.
[0074] Since CXL devices are external devices, the way they communicate with the kernel is different from the way memory on the motherboard (such as memory sticks). That is, CXL devices cannot communicate with the kernel like motherboard memory. Therefore, how to notify the operating system to operate the CXL device after the computing device physically recognizes the hot plug of the CXL device remains an unresolved problem.
[0075] To enable hot-swapping of CXL devices, while simultaneously reducing maintenance costs while expanding memory capacity based on CXL devices, embodiments of the present application provide a method for hot-swapping storage devices. This method primarily involves software-level information exchange between the CXL device and the computing device's operating system, enabling the operating system to recognize and process CXL devices inserted while the computing device is powered on. This allows hot-swapping and hot-unswapping of CXL devices without interrupting services on the computing device.
[0076] To facilitate understanding of the technical solution of the embodiment of the present application, an application scenario of the embodiment of the present application is first introduced below.
[0077] For example, Figure 1 illustrates an application scenario of an embodiment of the present application. As shown in Figure 1 , a computing cluster 10 may include multiple computing devices 100, each of which can communicate with each other based on protocols such as the Hypertext Transfer Protocol (HTTP). In this example, a CXL device 20 can connect to multiple computing devices 100, and one computing device 100 can allocate memory space provided by the CXL device 20 to each computing device 100 in the cluster 10 by deploying a FM 30.
[0078] In this example, the CXL device 20 may include a circuit board 23, a CXL controller 21, and memory 22 disposed on the circuit board. The memory 22 is connected to the CXL controller 21. The memory 22 may be a memory chip attached to the circuit board 23 of the CXL device along with the CXL controller. This allows the memory 22 to be hot-swapped along with the entire CXL device 20. Alternatively, the memory 22 may be a memory stick, inserted into a memory slot provided on the circuit board 23 of the CXL device 20. This allows the memory 22 to be hot-swapped individually or along with the entire CXL device 20. For example, the memory 22 may be either volatile or non-volatile memory, and the CXL controller 21 may manage each memory 22.
[0079] The CXL controller 21 may be a multi-head CXL expansion control chip, so that the CXL controller 21 can be connected to multiple computing devices 100 .
[0080] Take, for example, the hot-swapping of memory 22 on a CXL device 20 without powering off the computing device 100. In this example, without powering off, the CXL controller 21 can identify the hardware information (including but not limited to the unique identification number, capacity, bandwidth, latency, etc.) of the newly inserted memory 22 on the circuit board 23 and report this hardware information to each computing device 100. The kernel of the operating system 101 of the computing device 100 then performs a capacity check on the newly inserted memory 22 based on the hardware information, divides the memory 22 into corresponding physical pages, obtains the offsets and physical space addresses of the physical pages, and maps these physical pages to the virtual memory space. The physical space address and offset of the memory 22 are also transmitted back to the CXL device 20 for storage.
[0081] When FM 30 detects that a memory 22 is inserted into a CXL device 20 in computing cluster 10, it can obtain information such as the unique identification number and capacity of the memory 22 to allocate storage space for use by each computing device 100 in computing cluster 10. It will be understood that when a computing device 100 uses the allocated storage space of the memory 22, it does so based on management data such as the physical address and offset of the memory 22 stored within the computing device 100.
[0082] In this example, to uninstall a specific memory 22 on a CXL device 20, the operator inputs the hardware information of the memory 22 to FM 30 via a command line, triggering the uninstall process. FM 30 then determines whether the storage space provided by the memory 22 is occupied by any computing devices 100. If so, it first reclaims space from these computing devices 100. If the memory 22 is free (or has been reclaimed), FM 30 notifies the operating system 101 of the computing device 100 to gradually uninstall the virtual and physical memory space of the memory 22, clearing the relevant records of the memory 22 and synchronizing the CXL device 20. This allows the user to remove the memory 23 from the circuit board 23.
[0083] It is understandable that the computing cluster 10 can, based on a similar principle, perform hot-swap processing on the entire CXL device 20 through the FM 30 without interrupting cluster services, thereby reducing maintenance and commissioning costs during memory expansion.
[0084] In other application scenarios, the CXL device 20 may also be directly connected to a computing device 100 via a cable or other means. The operating system 101 of the computing device 100 may implement the entire hot-swap process of the CXL device 20 or the memory 22 based on similar principles.
[0085] Next, a computing device provided in an embodiment of the present application is introduced with reference to the accompanying drawings.
[0086] For example, FIG2 shows a schematic diagram of the structure of a computing device. The computing device 100 can be a server, a computer, or a virtualized device (such as a virtual machine or a container, etc.). Among them, the server can be a physical server, a cloud server, or a GPU server. As shown in FIG2 , the computing device 100 may include a processor 110, a memory 120, and a communication interface 130. The processor 110, the memory 120, and the communication interface 130 can be deployed on a motherboard 11 and connected via a bus or other means.
[0087] In this embodiment, the processor 110 is the computing core and control core of the computing device 100. In some embodiments, the processor 110 can perform some or all of the steps of the method provided in this embodiment. As an example, the processor 110 can be a central processing unit (CPU), a system on chip (SOC), a processor integrated on an SOC, a separate processor chip or controller, etc.: the processor 110 can also include a dedicated processing device, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), etc. The processor 110 can be a processor group composed of multiple processor chips, and the multiple processor chips are coupled to each other through one or more buses.
[0088] Memory 120 provides storage space and can serve as a server's memory device for storing computer program instructions and data, such as, but not limited to, operating systems (OS) 101 (e.g., Windows, Linux, etc.), CXL drivers 121, memory hot-swap management module 122, memory drivers 123, and CXL command-line interface module 126. Memory 210 can also store data such as memory management tables 125 and attribute files (e.g., sysfs files) 125.
[0089] Exemplarily, the memory 120 may be a non-power-off volatile memory, such as an embedded multi media card (EMMC), universal flash storage (UFS) or read-only memory (ROM), or other types of static storage devices that can store static information and instructions. It may also be a power-off volatile memory (volatile memory), such as a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions. It may also be an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other computer-readable storage medium that can be used to carry or store program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0090] The communication interface 130 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0091] In addition, the computing device 100 may further include a baseboard management controller BMC 150 , and the baseboard management controller BMC 150 communicates with other hardware (such as a CPU, etc.) on the computing device 100 through a physical channel.
[0092] In this embodiment, the computing device 100 can also be connected to the CXL device 20 via a cable or other means. The processor 110 can deploy the FM 30 to perform hot-swap management and storage space configuration for the connected CXL device 20. Alternatively, the FM 30 can be deployed in the BMC 150 to manage and configure the CXL device 20.
[0093] It should be noted that when FM30 is deployed on multiple devices (including computing device 100) in a computing cluster 10, the different FM30s serve as backups for each other, and only one FM30 is operational in the cluster at any one time. The storage space provided by the CXL device 20 connected to a computing device 100 can be allocated by the FM30 to other computing devices 100 in the computing cluster 10.
[0094] The functions of the relevant software programs in the computing device 100 in this embodiment are introduced below.
[0095] In this embodiment, referring again to FIG. 1 , the computing device 100 can be connected to the CXL device 20 . The CXL driver 121 of the computing device 100 , when executed by the processor 110 , can be used to implement communication between the kernel and the CXL device 20 , identify events reported by the CXL device 20 , parse device information of the CXL device 20 from the events (including but not limited to the unique identification number of the CXL device 20 , the number of memories 22 inserted in each CXL controller 21 , etc.), and parse hardware information of the memories 22 (including the unique identification number, capacity, bandwidth, and latency, etc.).
[0096] The memory hot-plug management module 122 is a software module running in the operating system 101 of the computing device 100 and implemented based on ACPI technology or memory hot-plug technology to support memory hot-plugging. In this embodiment, the memory hot-plug management module 122 specifically includes an ACPI module 1221 and a memory hot-plug module 1222. The ACPI module 1221 and the memory hot-plug module 1222 can each define different hardware configuration interfaces for the operating system 101. When running on the processor 110, the CXL driver 121 can communicate with the ACPI module 1221. The ACPI module 1221 transmits device information of a newly inserted CXL device 20 and hardware information of the corresponding memory 22 to the ACPI module 1221 via the interface defined by the ACPI module 1221. The ACPI module 1221 then transmits this information to the memory hot-plug module 1222. Alternatively, when executed by the processor 110 , the CXL driver 121 may directly communicate with the memory hot-plug module 1222 , transmitting device information of the newly inserted CXL device 20 and hardware information of the memory 22 via an interface defined by the memory hot-plug module 1222 .
[0097] Exemplarily, the memory hot-plug module 1222 can adopt mature memory hot-plug technology in the field to support dynamic addition or removal of physical memory devices at runtime, realize dynamic management of memory resources by updating the kernel's memory management data structure and memory segment mapping relationship, and provide a corresponding event notification mechanism so that user space and application programs can respond to memory hot-plug events. Specifically, in this example, the memory hot-plug module 1222 can detect the insertion of physical memories, thereby starting the detection of the properties and capacity of these memories, and initializing the management data of the memory 22 (such as the starting page frame number, the ending page frame number, the total number of pages, the physical space address, etc.). The memory hot-plug module 1222 can also update the corresponding memory segment mapping relationship according to the insertion event or the removal event during the memory hot-plug process, supporting the kernel to dynamically allocate and release the capacity provided by these expanded memories 22.
[0098] Memory driver 123 is a memory management subsystem within operating system 101 of computing device 100. When executed by processor 110, memory driver 123 can retrieve management data for a newly inserted CXL device 20 or memory 22 from memory hot-swap management module 122 and perform extended memory management operations, such as creating a memory management table 125 and a memory attribute file (e.g., sysfs file) 124. Thus, by writing to attribute file 124, the logical storage space (memory space) corresponding to CXL device 20 and memory 22 can be brought online, making the logical storage space visible to the user. Similarly, when unmounting the logical storage space corresponding to CXL device 20 and memory 22, memory driver 123 can also operate on attribute file 124 to make the logical storage space invisible to the user, thereby enabling hot unmounting of CXL device 20 or memory 22.
[0099] The CXL command line interface module 126 can communicate directly with the CXL driver 121 or indirectly with the CXL driver 121 through a corresponding software development kit (SDK). Users can enter commands through the CXL command line interface module 126 to trigger the call of the underlying CXL driver 121 to obtain relevant information about the CXL device 20 and its memory 22 and pass it to the operating system 101 or FM 30.
[0100] FM 30 can centrally manage and allocate storage space provided by CXL device 20 connected to computing device 100. In some examples, FM 30 can also be responsible for controlling hot-plugging of CXL device 20 and its memory 22.
[0101] In this example, when computing device 100 is deployed as a server in a cluster, when other servers (not running FM30) request to bind storage space from FM30 of computing device 100, FM30 divides the managed storage space into multiple physical areas, allocates and binds physical areas to the servers upon request, and records the allocation details in mapping table 132. Mapping table 132 primarily records the relationships between physical areas and mapped storage, and between physical areas and bound computing devices.
[0102] BMC 150 can also implement the deployment of FM 30 to manage information related to CXL device 20 and its memory 22, such as the aforementioned device information, hardware information, and location information (such as the connection location of CXL device 20 to motherboard 11 and the location of memory 22 in the slot of CXL device 20). In this way, the location information can be viewed using the unique identification number of memory 22 and the unique identification number of CXL device 20, so as to perform operations such as allocation and removal of target devices based on these identification numbers.
[0103] In one scenario, a CXL device is connected to a computing device to expand the memory of the computing device. Next, the hot-swapping process of the computing device 100 core to the memory 22 on the CXL device 20 is described in conjunction with the hot-swapping process shown in FIG3A and the unloading (unplugging) process shown in FIG3B .
[0104] In this implementation, when the computing device 100 is powered on, and a memory 22 is inserted into the CXL device 20, as shown in FIG3A , the specific insertion process includes:
[0105] When the CXL device 20 detects a newly inserted memory card 22, it obtains the memory card 22's hardware information, generates a corresponding event, and transmits it to the CXL driver 121 of the computing device 100 via step 1. It will be appreciated that the CXL device 20 can run CXL firmware X23 (stored in a ROM) to identify the memory card 22 inserted into the memory slot, encode and decode the data in the memory card 22, and communicate with the CXL driver 121. Furthermore, the CXL firmware X23 can also perform read and write operations on the memory card 22, log read and write activity, and perform error injection, among other functions, to name a few.
[0106] CXL driver 121 identifies this event as an insertion event and parses it to obtain the device information of CXL device 20 and the hardware information of memory 22. The information is then passed to the memory hot-swap management module 122 of the operating system 101 kernel in step 2. Next, memory hot-swap management module 122 divides the physical storage space of memory 22 into physical pages and uniformly addresses them based on the hardware information. It obtains the PFN of each physical page, as well as the physical space address and offset of each physical page. This generates memory 22 management data (including the starting PFN, total number of pages, starting physical space address, offset, etc.), which is then passed to memory driver 123 in step 4 (optionally along with the device information and hardware information).
[0107] After receiving the management data, the memory driver 123 returns the management data to the CXL device 20 through steps 6 and 7. Simultaneously, the memory driver 123 performs the following processing based on whether the storage space provided by the memory 22 is allocated as kernel space or user space: If the storage space is allocated as kernel space, the memory management module 1231 in the memory driver 123 maps each physical page of the memory 22 to the corresponding virtual page according to the management data through step 5, creating a memory management table 125 to manage the mapping of the physical storage space of the memory 22 to the logical storage space. Furthermore, the memory driver 123 creates an attribute file 124 for the logical storage space of the memory 22. In this example, the attribute file 124 is a sysfs file, which is a property file created by the Linux kernel for new storage devices. Within the sysfs file, a new memory node can be created for the newly inserted memory 22. In this way, the state file (used to describe the state of the memory node) of the newly created memory node under the sysfs file is written with "online", which can make the logical memory space of the memory 22 online and visible to the user. It can be understood that the online logical memory space is managed by the memory allocation system in the kernel (not shown in Figure 3A), so that it can be provided to the kernel state process of the computing device 100.
[0108] If the storage space of the memory 22 is allocated as user space, the memory driver 123 will hand over the management data to the memory allocator 127 for mapping the physical storage space to the logical storage space to provide it to the user state process for use, which will not be repeated here.
[0109] At this point, the computing device 100 completes the hot-plug operation on the memory 22. It is understood that a portion of the storage space of the memory 22 can be used as kernel space, with the remainder as user space, or the entire storage space of the memory 22 can be used as kernel space. Furthermore, the order in which the memory driver 123 executes steps 5 and 6 is not limited.
[0110] Furthermore, the CXL device 20 stores the management data received from the CXL driver 121 . Thus, when the CXL device 20 receives a read or write request, it can decode the address in the request to determine the specific physical space address pointed to by the address, and thus perform the read or write operation.
[0111] In this implementation, when the computing device 100 performs a hot unloading operation on the memory 22 on the CXL device 20 , as shown in FIG. 3B , the process specifically includes:
[0112] First, the operating system 101 kernel can obtain an uninstall instruction manually inputted into the CXL command line interface module 126 in step 1. The uninstall instruction includes at least the unique identification number (e.g., SN) of the memory 22 (target device) to be uninstalled, and may also include the unique identification number of the CXL device 20. Based on the unique identification number of the memory 22, the kernel's memory driver 123 transfers the memory 22 as part of the kernel space to the memory management module 1231 for processing, and transfers the memory 22 as part of the user space to the memory allocator 127 for processing.
[0113] Among them, the processing of the kernel space includes: the memory management module 1231 determines the usage status of the kernel space (idle or occupied) according to the unique identification number of the memory 22. If it is idle, the logical storage space unloading operation is performed; if it is occupied, the space is first reclaimed from the process occupying this part of the space, and then the logical storage space is unloaded after the space is reclaimed. Specifically, referring to Figure 3B, when the memory management module 1231 performs the logical storage space unloading on the kernel space provided by the memory 22 through step 2, it first performs a write operation on the attribute file 124 of the memory 22, so that the logical storage space of the memory 22 is invisible to the user. Then, the mapping relationship between the corresponding physical storage space of the memory 22 and the logical storage space is released, and the records of the memory management table 125 are cleared, and the unloading of the logical storage space of the memory 22 is completed.
[0114] Similarly, for the part of memory 22 that serves as user space, if it is occupied, the memory allocator 127 also reclaims the space first, and then releases the mapping relationship between this part of the logical storage space and the corresponding physical storage space of memory 22, completing the unloading of the logical storage space.
[0115] After completing the logical space unloading, the memory driver 123 notifies the memory hot-swap management module 122 to perform a physical storage space unloading of the memory 22. The memory hot-swap management module 122 then executes step 3, releasing the memory capacity provided by the memory 22 from the sparse memory model based on the unique identification number of the memory 22 and clearing the memory information related to the memory 22 (including management data). The memory hot-swap management module 122 then notifies the CXL driver 121 through steps 4, 5, and 6 to transmit an unloading message to the CXL device 20 where the memory 22 resides, instructing the CXL device 20 to unload the memory 22. The CXL device 20 then verifies the unique identification number of the memory 22 contained in the unloading message. If the verification is successful, the CXL device 20 clears the records related to the memory 22 (including management data such as the physical space address) and executes step 7 to return a message indicating the unloading is complete to the CXL command line interface module 126 to notify the operator. At this point, the operator can remove the memory 22 from the CXL device 20.
[0116] Furthermore, in some possible implementations, the entire CXL device 20 connected to the memory 22 may be hot-inserted and hot-uninstalled using an operating principle similar to the hot-insertion process shown in FIG. 3A and the hot-uninstallation process shown in FIG. 3B .
[0117] In another scenario, multiple computing devices 100 in a computing cluster 10 are connected to a CXL device 20 (e.g., as shown in FIG1 ), and the storage space of the CXL device 20 is divided and managed by FM 40. Next, the specific process of hot-swapping and unswapping the memory 22 on the CXL device 20 by FM 30 is described in conjunction with the hot-swapping process shown in FIG3C and the hot-unswapping process shown in FIG3D .
[0118] In this implementation, when a CXL device 20 detects the insertion of a new memory device 22 during business operations in a computing cluster 10, as shown in FIG3C , the insertion process includes: The CXL device 20 reports the memory device 22 insertion event to each connected computing device 100 in step 1. Subsequently, each computing device 100 performs steps 2 through 4, uniformly addresses the physical storage space of the memory device 22, and transmits the generated management data to the memory driver 123. It will be appreciated that the interactive execution of steps 1 through 4 between each computing device 100 and the CXL device 20 is similar to the execution principle of steps 1 through 4 in the example shown in FIG3A , and will not be further described.
[0119] Next, the memory driver 123 returns the acquired management data, along with the identity of the computing device 100, to the CXL device 20 through steps 5 and 6. Thus, the CXL device 20 receives and stores the management data associated with the identity of each computing device 100. The management data returned by each computing device 100 for the memory 22 may differ.
[0120] In this implementation, after FM 30 detects the insertion of a new device into a CXL device 20, it obtains the unique identification number and capacity of the newly inserted memory 22 from the CXL device 20 (it may also obtain device information about the CXL device 20) in step 7. It then partitions the memory space based on the capacity, generating multiple physical regions for allocation to each computing device 100 in the computing cluster 10. Each computing device 100 can then use the corresponding physical region allocated by FM 30 based on its own stored management data regarding the memory 22, for example, by mapping the physical address space of the region to a logical storage space for use by its processes. Thus, when a CXL device 20 receives a read or write request from a computing device 100, it can decode the computing device identity and the address carried in the request to determine which physical space address the address in the request refers to, thereby locating the corresponding location to perform the read or write operation.
[0121] In this implementation, when the FM 30 performs a hot unloading operation on the memory 22 on the CXL device 20, as shown in FIG3D , the process specifically includes:
[0122] FM30 can obtain an uninstall instruction manually entered into CXL command line interface module 126 in step 1. The uninstall instruction includes at least the unique identification number of the memory 22 to be uninstalled. FM30 uses the unique identification number to determine the computing device 100 associated with each physical area of memory 22 and its usage status (idle or occupied). If the memory 22 is idle, FM30 executes step 2 to notify each computing device 100 to perform the uninstall operation. If a physical area is occupied, FM30 first reclaims space from the corresponding computing device 100 and then executes step 2 to instruct the computing device 100 to perform the uninstall operation.
[0123] Specifically, upon receiving the unmount operation notification from FM 30, memory driver 123 of computing device 100 causes memory hot-swap management module 122 to execute step 3 to physically unmount memory 22. Memory hot-swap management module 122 then notifies the CXL device 20 of the unmount operation through steps 4, 5, and 6. The unmount result is then returned to the CXL command-line interface module 126 in step 7 to inform the operator, facilitating the unmounting of memory 22 from CXL device 20. The execution principles of steps 3 through 7 in this example are similar to those of steps 3 through 7 in the example shown in FIG. 3B , and are not further detailed here.
[0124] Furthermore, in some possible implementations, the entire CXL device 20 connected to the memory 22 may be hot-plugged and hot-unloaded using an operating principle similar to the hot-plugging process shown in FIG. 3C and the hot-unloading process shown in FIG. 3D .
[0125] In this way, the computing device 100 can implement hot swapping of CXL devices or CXL device-based memories without powering off, without interrupting business operations, thereby reducing maintenance and commissioning costs.
[0126] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the computing device 100. In other embodiments of the present application, the computing device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0127] Next, based on the above description, a storage device hot-swap method provided by an embodiment of the present application is introduced. It is understandable that this method is proposed based on the above description, and part or all of the contents of this method can be referred to the above description.
[0128] Please refer to FIG4 , which is a flow chart illustrating a method for hot-swapping a storage device provided in an embodiment of the present application. It is understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. For example, the method may include the following steps, as shown in FIG4 , when executed on the computing device 100 shown in FIG1 :
[0129] S410: The computing device receives an insert event.
[0130] In this embodiment, the insertion event may be an event reported by the CXL device 20 indicating hot insertion of a target device into the computing device 100. The target device may be the CXL device 20 or the memory 22 (which may be volatile memory, non-volatile memory, such as a memory stick or PMEM) connected to the CXL device 20. The hot insertion refers to the connection of the computing device 100 to the target device while the computing device 100 is powered on.
[0131] For example, while computing device 100 is running a service, if a new CXL device 20 is connected to computing device 100, or if one or more storage devices 22 are newly inserted into a CXL device 20 connected to computing device 100, CXL device 20 may generate a corresponding insertion event and send it to computing device 100. It is understood that computing device 100 may monitor the insertion event via the running CXL driver 121.
[0132] For example, the insertion event encapsulates first information, which may specifically include device information of the CXL device 20 and hardware information of the storage 22 on the CXL device 20. It will be appreciated that when the target device is a CXL device 20, the insertion event must encapsulate the device information of the CXL device 20 (including the unique identification number of the device 20) and the hardware information of all storage devices 22 inserted therein (including the unique identification numbers and capacities of the storage devices 22). When the target device is one or more storage devices 22, the first message may include the hardware information of these storage devices 22 and the device information of the CXL device 20 in which they reside.
[0133] S420: Manage the physical storage space of the target device according to the capacity to obtain management data of the target device.
[0134] In this embodiment, the CXL driver 121 running on the computing device 100 parses the first information from the insertion event and then transmits the first information to the operating system 101 kernel via a target communication method such as a function call, process communication, or system call. This information then manages the physical storage space of the target device according to the target capacity, such as by dividing the physical pages and uniformly addressing them to obtain corresponding management data. This management data may include the offset and physical space address of each physical page, as well as the page frame number, total number of pages, starting physical space address, length, and so on.
[0135] As an example, the first information may be transferred to the memory hot-plug management module 122 in the kernel, so that the memory hot-plug management module 122 in the kernel manages the above-mentioned physical storage space.
[0136] S430: Send the management data to the CXL device, so that the CXL device can perform read and write operations based on the management data.
[0137] In this embodiment, the core of computing device 100 can transmit the management data back to CXL device 20 for storage via CXL driver 121. In this way, when CXL device 20 subsequently receives a read or write request, it can decode the address in the request based on the management data and determine the physical space address indicated in the request to perform the read or write operation.
[0138] In some exemplary embodiments, if the storage space of the target device is managed by the kernel of the computing device 100, in addition to returning the management data to the CXL device 20, the management data must also be passed to the kernel driver 123 of the kernel to map the physical storage space of the target device to the corresponding logical storage space, and store the mapping relationship, so as to provide it to the corresponding kernel mode or user mode process for use.
[0139] In other examples, if the storage space of the target device is managed by FM30 of computing device 100, then after FM30 senses the insertion of the target device, it can also obtain information such as the unique identification number and capacity of the target device from CXL device 20, thereby dividing the storage space of the target device into areas to provide it to computing devices in need in computing cluster 10.
[0140] In this embodiment, computing device 100 utilizes a notification mechanism with CXL device 20, supported by relevant drivers and program modules running in a reused kernel, to enable the kernel to recognize and process hot insertion of CXL device 20 and its memory 22, thereby enabling hot memory expansion of computing device 100. This allows computing devices 100 connected to CXL device 20 to commission and replace CXL device 20 and its memory 22 without powering down or interrupting service operations, significantly reducing cluster operation and maintenance costs.
[0141] The following describes the process of hot-swapping the memory 22 on the CXL device 20 with reference to the accompanying drawings.
[0142] In one implementation, a computing device 100 is connected to a CXL device 20. A memory 22 can be inserted into the CXL device 20 through a memory slot. The operating system 101 of the computing device 100 is responsible for hot-swapping the memory 22. Specifically, during the hot-swapping phase of the memory 22, as shown in FIG5 , the storage device hot-swapping method may include:
[0143] S510 : Receive an insert event through a CXL driver running on a computing device.
[0144] In this example, the insertion event is used to indicate the hot insertion of the memory 22 (ie, the target device) on the CXL device 20 connected to the computing device 100 .
[0145] When a memory device 22 is inserted into a CXL device 20 while the computing device 100 is powered on, the CXL device 20, through its CXL firmware X23, establishes a CXL mode link on the downstream port in accordance with the CXL protocol, generating an interrupt signal. After recognizing the interrupt signal, the CXL controller 21 executes step S500 to detect the newly inserted memory device 22, initialize it according to the CXL protocol (configure the DVSEC function), and read the hardware information of the memory device 22. This hardware information includes, but is not limited to, the unique identification number and capacity of the memory device 22, as well as information such as the bandwidth and latency of the memory device 22.
[0146] Next, the CXL controller 21 encapsulates the DVSEC initialization information, the hardware information, and the device information of the CXL device 20 (including but not limited to the unique identification number of the CXL device 20 and the number of memory devices inserted in the CXL device 20) into an insertion event and transmits it to the computing device 100. The CXL driver 121 running on the computing device 100 can receive the insertion event.
[0147] S520: Manage the physical storage space of the target device according to the capacity to obtain management data of the target device.
[0148] In this example, step S520 may specifically include the following steps S521 to S522:
[0149] At S521 , the CXL driver transmits the first information to a memory hot-swap management module running on the computing device through a target communication method.
[0150] In this example, CXL driver 121 determines that the received event is an insertion event, and then continues to parse the insertion event to obtain first information. Then, CXL driver 121 transmits the first information to memory hot-plug management module 122 running in the kernel through the target communication method.
[0151] In some specific examples, the CXL driver 121 may directly communicate with the memory hot-plug module 1222 in the memory hot-plug management module 122 to transmit the first information to the memory hot-plug module 1222 for parsing and processing.
[0152] In other examples, CXL driver 121 can exchange data with ACPI module 1221 in memory hot-plug management module 122 to pass the first information to ACPI module 1221. ACPI module 1221 then generates an ACPI table based on the first information, records the first information in the ACPI table, and then passes the ACPI table to memory hot-plug module 1222. Because ACPI technology supports communication with operating system 101, the ACPI table is data that the kernel can directly access and can be directly read by memory hot-plug module 1222. Therefore, passing information through ACPI module 1221 helps reduce the difficulty of implementing execution logic.
[0153] S522 , using the memory hot-swap management module, uniformly address the physical storage space of the target device according to capacity to obtain management data.
[0154] In this example, the memory hot-plug module 1222 learns that a memory 22 has been inserted based on the first information obtained. Based on the first information, the memory 22 is detected for capacity, and a suitable address space is found in the sparse memory model of the computing device 100 to accommodate the space of the memory 22. The space of the memory 22 is divided into physical pages (page frames), and the addresses are uniformly assigned to obtain the page frame number (PFN) of each physical page, as well as the physical space address and offset of each physical page. The PFN is a globally unique identifier associated with the page frame. In this way, the memory hot-plug module 1222 initializes and obtains management data for the memory 22. The management data includes the PFNs of the memory 22, the starting PFN, the total number of pages, the offset and physical space address of each physical page, the starting physical space address and length, etc., and may also include information such as device status (such as the organizational structure of the physical pages, usage information, etc.), but is not limited to this.
[0155] Next, the memory hot-plug module 1222 may transmit the management data of the newly inserted memory 22 to the memory driver 123 of the kernel of the operating system 101 through a callback function.
[0156] S530 : Send the management data to the CXL device via the memory driver and the CXL driver running on the computing device.
[0157] In this step, the memory driver 123 transmits the management data back to the CXL device 20 via the CXL driver 121. After receiving the management data such as the physical space address and offset sent by the computing device 100, the CXL device 20 saves it in step S531 so that when it subsequently receives a read or write operation request from the computing device 100, it can decode the address contained in the request and address the corresponding physical space address to perform the corresponding read or write operation.
[0158] S540 : Map the physical storage space of the target device to the logical storage space through the memory driver according to the management data.
[0159] In this step, if the physical storage space of the target device (which can be part or all) is used as kernel space, the memory management module 1231 in the memory driver 123 can map the physical page of the memory 22 used as kernel space to the corresponding virtual page according to the management data, and generate a corresponding mapping relationship.
[0160] S550: Create a memory management page and a property file of the target device according to the logical storage space through a memory driver.
[0161] In this step, the memory management module 1231 of the memory driver 123 creates a new memory management table 125 based on the mapping relationship to complete the mapping management between the physical storage space of the memory 22 and the corresponding logical storage space. By way of example and not limitation, the memory management table 125 may include a kernel direct mapping table, a virtual memory mapping page, a virtual memory mapping table, etc., so that the kernel can manage and use the storage space provided by the memory 22 based on these memory management tables 125. It can be understood that the virtual mapping page is a virtual page used to record the memory mapping table. The virtual memory mapping table is used to describe the mapping relationship between the physical page of the memory 22 and the corresponding virtual page. The kernel direct mapping table can be used to describe the memory area that the kernel can directly map.
[0162] Then, according to the logical memory space of the memory 22, the memory driver 123 creates a corresponding sysfs file for the memory 22, and creates a memory node (i.e., a memory logical unit) corresponding to the memory 22 under the sysfs file. It is understood that the creation of the sysfs file and the memory node can be automatically completed by a script, or can be completed using the probe tool in the kernel, which will not be repeated here.
[0163] S560 , triggering an online operation on the logical storage space by writing an online command to the property file.
[0164] In this step, by writing "online command" into the state file of the memory node created under the sysfs file through the memory driver 123, the logical memory space of the memory 22 can be put online in the operating system 101, making the logical memory space visible to the user.
[0165] It is understood that in this example, all virtual pages of the online logical memory space can be set to ZONE_NORMAL (direct mapping area, which is also the memory area that the kernel can directly map) or ZONE_MOVABLE (movable area), and handed over to the memory allocation system (buddy system) in the kernel for management. It is understood that these memory pages can be set manually or through automated scripts.
[0166] In this example, after the logical storage space of the memory 22 is set to be visible to the user, this space can be used by kernel-mode processes.
[0167] In other examples, if the physical storage space of the memory 22 is used as user space, then when executing stage S540, the memory allocator 127 of the memory driver 123 will complete the mapping of this part of the physical storage space to the corresponding logical storage space according to the management data and record the mapping relationship, so that this part of the logical storage space can be used by the user-mode process.
[0168] At this point, the hot insertion process of the memory 22 on the CXL device 20 on the computing device 100 is completed. In this way, the computing devices 100 in the computing cluster 10 can implement hot insertion recognition and processing of physical memory without interrupting services.
[0169] In this implementation, when the memory 22 on the CXL device 20 needs to be removed from the computing device 100, referring to FIG6 , the method may further include:
[0170] S610: Receive an uninstall instruction through a memory driver.
[0171] In this example, when the computing cluster is running services, a user can enter an unload instruction through the CXL command line interface 126 to instruct a hot unload operation on one or more specified memories 22 on a specified CXL device 20 . The instruction also describes the unique identification number of the memory 22 to be passed to the memory driver 123 running on the computing device 100 .
[0172] S620 , determining the usage status of the target device according to the unique identification number through a memory driver running on the computing device.
[0173] In this step, the memory driver 123 may first determine the usage status (free or occupied) of the memory 22 according to the unique identification number of the memory 22 to be unloaded.
[0174] Exemplarily, if the storage space of the memory 22 is in an idle state, the following step S630 is continued to be executed.
[0175] For example, if the storage space of the memory 22 is occupied, the memory driver 123 reclaims the space through the following S621 and then executes S630:
[0176] At S621 , when the target device is in an occupied state, all storage spaces of the target device are reclaimed through a memory driver, so that the target device is restored to an idle state.
[0177] For example, if data or running processes are stored in the storage space of the memory 22, the data or processes are migrated to other memory spaces, so that the memory 22 enters an idle state. It is understood that the portion of the memory 22 that is the kernel space can be reclaimed by the memory management module 1231, and the portion of the memory 22 that is the user space can be reclaimed by the memory allocator 127.
[0178] S630: When the target device is in an idle state, a logoff command is written to the target device's attribute file to trigger a logoff operation on the corresponding logical storage space. In this example, the memory driver 123 in the kernel searches for the corresponding sysfs file based on the unique identification number of the memory 22, and writes an "offline" logoff command to the state file of the corresponding memory node under the sysfs file through the shell command "echo" (e.g., "echo offline> / sys / devices / system / memory / memoryXXX / state" command), making the logical memory space provided by the memory 22 invisible to the user.
[0179] S640: Release the mapping relationship between the physical storage space and the corresponding logical storage space through the memory driver, and clear the memory management page created according to the logical storage space.
[0180] In this step, the memory driver 123 releases the mapping relationship between the physical page and the virtual page of the memory 22, releases the kernel direct mapping table, releases the virtual memory mapping page, releases the virtual memory mapping table, and completes the unloading of the logical storage space of the memory 22.
[0181] Next, the memory hot-plug management module 122 running on the computing device is notified via the memory driver 123 to release the physical storage space of the target device.
[0182] It can be understood that S630 and S640 are executed based on the physical storage space of memory 22 being used as kernel space.
[0183] When the physical storage space of the memory 22 is used as user space, when the target device is in an idle state, the mapping relationship between the physical storage space and the corresponding logical storage space is released, and then the memory hot-plug management module 122 running on the computing device is notified through the memory driver 123 to release the physical storage space of the target device.
[0184] S650, through the memory hot-swap management module, releases the physical storage space of the target device according to the unique identification number and clears the management data.
[0185] In this step, the kernel memory driver 123 can pass the unique identification number of the memory 22 to the memory hot-plug module 1222 in the memory hot-plug management module 122 through a callback function or system call. Then, the memory hot-plug module 1222 finds all blocks corresponding to the memory 22 in the sparse memory model based on the management data, resets the allocation status of these blocks to make them unavailable, and clears related records (such as the physical page frame number and physical space address of the memory 22), completing the unloading of the physical storage space of the memory 22.
[0186] S660: The computing device transmits an uninstall message to the CXL device via the CXL driver to instruct the CXL device to clear the management data stored therein.
[0187] In this example, the memory hot-swap management module 122 interacts with the CXL device 20 through the CXL driver 121 and sends an uninstall message including the unique identification number of the memory 22 to be uninstalled and the device information (unique identification number, communication port, etc.) of the CXL device 20 where it is located to the CXL device 20 .
[0188] S670: The CXL device performs an uninstall operation on the target device according to the uninstall message.
[0189] In this example, based on the received uninstall message, CXL device 20 checks whether it is responsible for managing memory 22. If not, CXL device 20 returns an uninstall failure message to CXL command line interface 126, indicating that the target device was not found. If so, CXL device 20 resets the corresponding HDM decoder and the root port to computing device 100, initiates a standard PCIe hot remove process, releases the corresponding CXL.io resources, and clears the relevant records related to memory 22. Upon completion, a complete uninstall message is returned to CXL command line interface 126 to notify operations and maintenance personnel.
[0190] In this way, after seeing the return message, the operation and maintenance personnel can unplug the memory 22 from the CXL device 20, completing the hot unplugging of the memory on the CXL device 20. This facilitates debugging and replacing the memory on the CXL device, flexibly manages the memory capacity, and does not require interrupting the operation of the computing device 100, thereby reducing operation and maintenance costs and energy consumption costs.
[0191] In some possible implementations, the entire CXL device can also be hot-inserted and hot-uninstalled using principles similar to those for hot-inserting and hot-uninstalling the memory 22. As shown in FIG7 , the process for hot-inserting the entire CXL device 20 includes:
[0192] S710 : Receive an insertion event reported by a CXL device through a CXL driver running on the computing device.
[0193] In this example, when a computing cluster is running services, when a CXL device 20 (i.e., a target device) connected to memory 22 is connected to computing device 100, the CXL device 20 is powered on. Then, step S700 is executed on the CXL device 20, and the CXL firmware X23 is powered on and executed. As described in the CXL protocol, the power-on initialization process begins: bus resources for the corresponding stack are configured, link training is triggered (adjusting link signal quality, rate, link width, etc.). After link training succeeds, the stack space is checked to see if it is in CXL mode. If so, CXL ports are configured, the CXL accelerator (which connects the CXL controller 21 and each memory 22, not shown) configuration space is entered, the CXL controller 21 is checked and initialized, CXL.memory is checked and initialized (including but not limited to configuring the DVSEC function), and the CXL.IO enumeration process is executed in accordance with the PCIe enumeration process. Furthermore, a driver is bound to the memory 22 on the CXL device 20, and the power-on initialization process ends.
[0194] After the initialization process, the CXL device 20 encapsulates the hardware information of each connected storage device 22 (including but not limited to capacity, bandwidth, latency, etc.) and device information (including but not limited to the unique identification number of the CXL device 20 and the number of storage devices inserted in the CXL device 20) as the first information into an insertion event and sends it to the CXL driver 121 on the computing device 100.
[0195] S720: Manage the physical storage space of the target device according to the capacity to obtain management data of the target device.
[0196] In this step, if the CXL driver 121 determines that the received event is an insertion event for a newly added CXL device, it proceeds to parse the insertion event to extract the first information. Next, based on the first information, a CXL node for the CXL device 20 is created on the computing device 100. It will be understood that this CXL node is a file record for the CXL device 20.
[0197] Next, the CXL driver 121 transmits the first information to the memory hot-swap management module 122 , which performs processing similar to S520 in the above example. This completes the allocation and unified addressing of physical storage space for all memories 22 on the CXL device 20 . Accordingly, corresponding management data is generated and returned to the CXL device 20 for storage, facilitating read and write operations based on the management data. It is understood that the process of returning management data to the CXL device 20 can also be referred to in the description of S530 in the above example.
[0198] Furthermore, computing device 100 also performs operations similar to those in S540 to S560 in the above example to map the physical storage space of each memory on CXL device 20 to the corresponding logical storage space for use by user mode or kernel mode processes, which will not be repeated here.
[0199] At this point, the operation flow of hot-plugging the CXL device 20 into the computing device 100 ends.
[0200] In this implementation, if a CXL device 20 needs to be uninstalled from computing device 100, the uninstall process is generally similar to the example uninstall process shown in FIG6 . The primary difference is that in this implementation, when a CXL device 20 is to be removed from computing device 100, the user can enter an uninstall command (including the unique identification number of the CXL device 20) through CXL command line interface 126 to instruct the uninstallation of the specified CXL device 20. Furthermore, when the CXL device 20 executes the uninstall operation for the target device according to the uninstall message, it checks whether it is the device described in the uninstall message. If not, the CXL device 20 notifies the CXL driver 121 of computing device 100 that the uninstallation failed and that the target device could not be found. If so, the CXL device 20 enters a power-down process and returns a message to the CXL command line interface module 126 indicating that the uninstallation is complete, allowing the operator to remove the CXL device 20.
[0201] In some implementations, multiple computing devices 100 in a computing cluster 10 are connected to a CXL device 20. A memory 22 can be inserted into the CXL device 20 via a memory slot. The FM 30 (deployed on a computing device 100) running in the computing cluster 10 is responsible for hot-swapping management of the memory 22. Specifically, during the hot-swapping phase of the memory 22, as shown in FIG8 , the storage device hot-swapping method may include:
[0202] S810 , receiving an insertion event reported by a CXL device through a CXL driver running on the computing device;
[0203] S820, managing the physical storage space of the target device according to the capacity, and obtaining management data of the target device;
[0204] S830 : Send the management data to the CXL device via the memory driver and the CXL driver running on the computing device.
[0205] In this example, the execution principles of steps S810 through S830 are similar to those described in the above example, with the primary difference being that, in this example, multiple computing devices 100 are connected to a single CXL device 20. Each computing device 100 manages the physical storage space of the memory 22 plugged into the CXL device 20, generating a set of management data. Furthermore, the management data returned by each computing device 100 to the CXL device 20 is associated with its own identity. The CXL device 20 associates and stores the management data returned by each computing device 100 with its corresponding computing device identity in step S831. Thus, when any computing device 100 sends a read or write request, the CXL device 20 uses the corresponding management data to decode the address in the request based on the computing device identity carried in the request, thereby locating the specific physical storage address to perform the read or write operation.
[0206] S840 : When the fabric manager FM running on the computing device senses that the target device is hot-plugged into the computing device, it obtains second information of the target device from the CXL device.
[0207] In this step, when FM 30 detects that memory 22 has been inserted into computing device 100, it creates a DAX file (a file that supports direct access by user-mode software) for the memory 22 through the DAX driver. It then calls the list interface in CXL command-line interface module 126 to instruct the CXL device 20 where the newly inserted memory 22 resides to report device information, including the unique identification number and capacity of the memory 22 (i.e., the second information). It will be appreciated that when obtaining this information and data, CXL command-line interface module 126 can invoke CXL driver 121 by running a script to retrieve it from CXL device 20.
[0208] S850 , dividing the physical storage space of the target device into regions according to the second information through FM to obtain multiple physical regions, so as to provide them to multiple computing devices in the cluster where the computing device is located for use.
[0209] In this step, FM 30 divides the physical storage space of memory 22 into regions based on its capacity, generating multiple physical regions. It then updates mapping table 132 to record mapping information for memory 22. By way of example, mapping table 132 may include at least two components. One component manages the physical information of the target device (memory 22 in this example), including but not limited to the target device's unique identification number, the unique identification number of the CXL device it resides on, its correspondence with the CXL device (e.g., memory slot), and capacity. The other component of mapping table 132 manages the target device's physical region management information, including but not limited to region information (region number) for the divided physical regions, the capacity corresponding to each physical region, and the IP addresses assigned to servers in the computing cluster.
[0210] In this example, when each computing device 100 uses the physical area provided by the memory 22 allocated by the FM 30 , it can determine the physical space address range corresponding to the physical area based on the management data of the memory 22 stored locally in the computing device 100 .
[0211] In some possible examples, the first information may also include the offset and starting physical space address of the memory 22. Because each computing device 100 uses different management data after addressing the memory 22, these offsets and starting physical space addresses are also associated with the identity of the corresponding computing device 100. In this way, when the FM allocates a physical area to a computing device 100, it can also inform the computing device 100 of the corresponding offset and starting physical space address of the physical area in the memory 22 based on the identity of the computing device 100.
[0212] Let's take an example. For example, computing device A gives the memory 22 management data with an offset of A1 and a starting physical space address of A2, computing device B gives the memory 22 management data with an offset of B1 and a starting physical space address of B2, and FM30 divides the memory 22 into region C and region D. If region C is assigned to computing device A, FM30 will inform computing device A of the offset A1, starting physical space address A2, and the capacity and number of region C on the memory 22. In this way, computing device A can determine which part of the memory 22 the allocated region C corresponds to based on the management data of the memory 22 stored by itself, and then use it. Similarly, if region C is assigned to computing device B, FM30 will inform computing device B of the offset B1, starting physical space address B2, and the capacity and number of region C on the memory 22.
[0213] In this implementation, when the memory 22 needs to be hot-unplugged from the computing cluster 10, the FM 30 performs hot-unloading management on the memory 22. As shown in FIG9 , the method may further include:
[0214] S910: Receive an uninstallation instruction via FM.
[0215] In this example, when the computing cluster is running a service, a user can input an unloading instruction through the CXL command line interface module 126 to instruct a hot unloading operation on one or more specified memories 22 on a specified CXL device 20 . The instruction also describes the unique identification number of the memory 22 for transmission to the FM 30 running on the computing device 100 .
[0216] S920: Determine the usage status of the target device according to the unique identification number through the FM.
[0217] In this example, the FM 30 may first determine the usage status (free or occupied) of the memory 22 according to the unique identification number of the memory 22 to be unloaded.
[0218] Exemplarily, if the storage space of the memory 22 is in an idle state, the following step S930 is continued to be executed.
[0219] For example, if the storage space of the memory 22 is occupied, the FM 30 reclaims the space through the following S921 and then executes S930:
[0220] At S921 , when the target device is in an occupied state, the physical storage space of the target device is reclaimed from the target computing device occupying the target device through FM, so that the target device is restored to an idle state.
[0221] In this step, FM30 can determine the target computing devices occupying the storage space of memory 22 by looking up the mapping table 132 and obtain the identity information (such as IP address information) of these target computing devices. It is understood that the target computing device can be the computing device 100 where FM30 is located, or it can be a device in the same cluster as the computing device 100.
[0222] Then, FM30 sends space recovery information according to the IP address of the target computing device to recover the storage space provided by the memory 22 to the target computing device. The target computing device will transfer the data stored in the memory 22 to other free storage spaces, so that the memory 22 returns to an idle state.
[0223] S930: When the target device is in an idle state, the structure manager clears the record about the target device.
[0224] In this example, after FM30 determines that the storage space of memory 22 is idle, FM30 calls the DAX driver interface to clear the DAX file corresponding to memory 22. FM30 then updates mapping table 132 to clear relevant records related to memory 22, such as records of space allocation for memory 22 and records of allocation to target servers, thereby preventing the storage space of memory 22 from being reallocated. FM30 no longer performs any operations on memory 22 with cleared related records.
[0225] S940, when the target device is in an idle state, release the physical storage space of the target device according to the unique identification number and clear the management data.
[0226] In this example, FM30 can communicate with the kernel of the operating system 101 of the computing device 100 to notify the memory driver 123 to perform a hot unloading operation on the unloaded memory 22. The memory driver 123 can then pass the unique identification number of the memory 22 to the memory hot-plug module 1222 in the memory hot-plug management module 122 through a callback function or system call. Then, the memory hot-plug module 1222 finds all blocks corresponding to the memory 22 in the sparse memory model based on the management data, resets the allocation status of these blocks to make them unavailable, and clears related records (such as the physical page frame number and physical space address of the memory 22), thereby completing the unloading of the physical storage space of the memory 22.
[0227] S950 , the computing device transmits an uninstall message to the CXL device via the CXL driver to instruct the CXL device to clear the management data stored therein;
[0228] S960: The CXL device performs the uninstall operation on the target device according to the uninstall message.
[0229] In this example, the execution principles of S950 to S960 are similar to the principles of S660 to S670 in the above example, and are not repeated here.
[0230] At this point, the hot-plug process management of the memory 22 on the CXL device 20 in the computing cluster 10 is completed.
[0231] It is understood that similar to the principle of hot-swap control of the memory 22 by the FM 30, the hot-swap control of the entire CXL device 20 can also be performed by the FM 30, which will not be described in detail here.
[0232] Based on the method in the above embodiment, the present invention provides a hot-swappable device for a storage device. Please refer to Figure 10, which is a schematic diagram of the structure of the device provided in the present invention.
[0233] As shown in Figure 10 , the storage device hot-swap apparatus 900 may include a receiving module 901 and a processing module 902. The receiving module 901 is configured to receive an insertion event, which is an event in which a target device is hot-plugged into a computing device. The insertion event includes first information, including the unique identification number and capacity of the target device. The target device is a Compute Express Connect (CXL) device or a storage device connected to the CXL device. The processing module 902 is configured to manage the physical storage space of the target device based on the capacity and obtain management data for the target device, including an offset and a physical space address. The processing module 902 is further configured to send the management data to the CXL device, enabling the CXL device to perform read and write operations based on the management data.
[0234] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method and will not be repeated here.
[0235] Based on the methods in the above embodiments, embodiments of the present application provide an electronic device. The electronic device may include: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is configured to execute the methods in the above embodiments.
[0236] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.
[0237] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product, characterized in that when the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0238] Based on the methods in the above embodiments, the present application also provides a chip. Please refer to Figure 11, which is a schematic diagram of the structure of a chip provided in the present application. As shown in Figure 11, the chip 1100 includes one or more processors 1101 and an interface circuit 1102. Optionally, the chip 1100 may also include a bus 1103.
[0239] The processor 1101 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 1101 or instructions in the form of software. The above-mentioned processor 1101 can be a general-purpose processor, a digital communicator (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods and steps disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0240] The interface circuit 1102 can be used to send or receive data, instructions or information. The processor 1101 can use the data, instructions or other information received by the interface circuit 1102 to process it, and can send the processing completion information through the interface circuit 1102.
[0241] Optionally, the chip 1100 further includes a memory, which may include a read-only memory and a random access memory, and provides operation instructions and data to the processor. Part of the memory may also include a non-volatile random access memory (NVRAM).
[0242] Optionally, the memory stores an executable software module or a data structure, and the processor can perform corresponding operations by calling an operation instruction stored in the memory (the operation instruction may be stored in an operating system).
[0243] Optionally, the interface circuit 1102 may be configured to output the execution result of the processor 1101 .
[0244] It should be noted that the corresponding functions of the processor 1101 and the interface circuit 1102 can be implemented through hardware design, software design, or a combination of hardware and software, and there is no limitation here.
[0245] It should be understood that each step of the above method embodiment can be completed by a hardware-based logic circuit or a software-based instruction in a processor.
[0246] It is understood that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in some possible implementations, the steps in the above embodiments can be selectively executed according to actual circumstances, and can be executed partially or completely, which is not limited here.
[0247] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0248] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0249] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0250] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
Claims
1. A storage device hot-swap method, characterized in that: The method is executed on a computing device and includes: receiving an insertion event, wherein the insertion event is an event of hot-plugging a target device into the computing device, the insertion event comprising first information, wherein the first information comprises a unique identification number and a capacity of the target device, and the target device is a computing fast connect CXL device or a memory connected to the CXL device; Performing physical storage space management on the target device according to the capacity to obtain management data of the target device, wherein the management data includes an offset and a physical space address; The management data is sent to the CXL device, so that the CXL device can perform read and write operations according to the management data.
2. The method according to claim 1, characterized in that The performing physical storage space management on the target device according to the capacity to obtain management data of the target device includes: The CXL driver running on the computing device transmits the first information to a memory hot-plug management module running on the computing device through a target communication method, wherein the memory hot-plug management module is a program for supporting memory hot-plug; The memory hot-plug management module uniformly addresses the physical storage space of the target device according to the capacity to obtain the management data.
3. The method according to claim 1 or 2, characterized in that: After obtaining the management data of the target device, the method includes: The management data is transmitted to a memory driver running on the computing device through the memory hot-swap management module; The management data is sent to the CXL device via the memory driver via a CXL driver running on the computing device.
4. The method according to claim 3, characterized in that After sending the management data to the CXL device, the method includes: When a fabric manager FM running on the computing device senses that the target device is hot-plugged into the computing device, acquiring second information of the target device from the CXL device, the FM being a program for managing the CXL device, the second information including a unique identification number and a capacity of the target device; Through the FM, the physical storage space of the target device is divided into regions according to the second information to obtain a plurality of physical regions, which are provided to a plurality of computing devices in the cluster where the computing device is located.
5. The method according to claim 3, characterized in that: After sending the management data to the CXL device, the method includes: According to the management data, mapping the physical storage space to the logical storage space through the memory driver; Creating a memory management page and a property file of the target device according to the logical storage space through the memory driver, wherein the memory management page is used to manage the mapping relationship between the physical storage space and the logical storage space, and the property file is used to provide information of the logical storage space to the application layer for use; An online operation on the logical storage space is triggered by writing an online command to the attribute file.
6. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: receiving an uninstall instruction, the uninstall instruction being used to instruct hot unplugging the target device from the computing device, the uninstall instruction including a unique identification number of the target device; When the target device is in an idle state, releasing the physical storage space of the target device according to the unique identification number, and clearing the management data; An uninstall message is sent to the CXL device to instruct the CXL device to clear the management data stored in the CXL device.
7. The method according to claim 6, characterized in that Before releasing the physical storage space of the target device according to the unique identification number, the method includes: Determining, by the configuration manager FM of the computing device, a usage state of the target device according to the unique identification number, wherein the usage state at least includes an idle state or an occupied state; When the target device is in an occupied state, the physical storage space of the target device is reclaimed from the target computing device occupying the target device through the FM, so that the target device is restored to an idle state, and the target computing device is used for the computing device. The device may be in the same cluster as the computing device.
8. The method according to claim 7, characterized in that Before releasing the physical storage space of the target device according to the unique identification number, the method includes: Determining, by a memory driver running on the computing device, a usage state of the target device according to the unique identification number, the usage state at least including an idle state or an occupied state; When the target device is in an occupied state, all storage space of the target device is reclaimed through the memory driver to restore the target device to an idle state.
9. The method according to claim 8, characterized in that Before releasing the physical storage space of the target device according to the unique identification number, the method includes: When the target device is in an idle state, a logoff command is written into a property file of the target device to trigger a logoff operation on the corresponding logical storage space, wherein the property file is used to provide information of the logical storage space to an application layer for use; Removing the mapping relationship between the physical storage space and the corresponding logical storage space through the memory driver, and clearing the memory management page created according to the logical storage space; The memory driver notifies a memory hot-plug management module running on the computing device to release the physical storage space of the target device.
10. A computing device, characterized in that include: at least one memory for storing a program; at least one processor, configured to execute the program stored in the memory; Wherein, when the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-9.
Citation Information
Patent Citations
Method for realizing hot plug on PCI EXPRESS (peripheral component interconnect express) in Linux
CN102508659A
Method and device for unloading mobile storage equipment
CN102662882A
Hot plugging method and device for persistent memory capable of being addressed through bytes
CN105260336A
Storage device and an element management method of the storage device
CN109840232A
Key management method, data protection method, system, chip and computer equipment
CN116011041A