A data reconstruction method, a solid-state storage unit and a storage system

By introducing a peer-to-peer interconnect network into the solid-state storage unit, the control plane and data plane are decoupled, which solves the problem of limited data flow efficiency between different smart solid-state drives, improves data reconstruction efficiency, and reduces host CPU resource consumption.

CN122152208APending Publication Date: 2026-06-05HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-12-05
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

The efficiency of data transfer between different smart solid-state drives is limited by the network bandwidth of the central processing unit, resulting in low efficiency during the data reconstruction process.

Method used

By introducing a peer-to-peer interconnect network into solid-state storage units, the control plane and data plane are decoupled, allowing solid-state storage units to read data directly from other storage devices. This avoids the limitations of the PCIe bus connected to the CPU or the bandwidth of the Internet, and enables direct data access through the peer-to-peer interconnect network.

Benefits of technology

It improves the efficiency of data transfer between different storage devices, reduces the resource consumption of the host CPU, and enhances the efficiency of the data reconstruction process and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122152208A_ABST
    Figure CN122152208A_ABST
Patent Text Reader

Abstract

The application discloses a data reconstruction method, a solid-state storage unit and a storage system, and relates to the technical field of storage. In the data reconstruction process, the solid-state storage unit only needs to interact with the host through the control plane command, the solid-state storage unit can directly read data from other storage devices, and the solid-state storage unit and the host do not need to interact through the data plane, so that the control plane and the data plane are decoupled. Moreover, the solid-state storage unit and other storage devices are connected through a peer-to-peer interconnection network to realize data access, and are no longer limited by the bandwidth of the PCIe bus connected by the CPU in the host or the bandwidth of the Internet accessed by the CPU, so that the data transmission time between the solid-state storage unit and other storage devices is reduced, the data flow efficiency between the solid-state storage unit and other storage devices is improved, the data processing efficiency of the solid-state storage unit is improved, and the resource consumption of the CPU in the host is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technology, and in particular to a data reconstruction method, a solid-state storage unit, and a storage system. Background Technology

[0002] Data reconstruction refers to integrating data from different sources into a unified data model to facilitate analysis and processing. As large-scale models and big data applications mature, data flow efficiency has become a crucial research area in computer architecture. To improve data flow efficiency, computing units are positioned close to storage components, offloading some data processing tasks from the host's central processing unit (CPU). Taking a smart solid-state drive (sSD) as an example, a smart SSD includes NAND flash memory and a field-programmable gate array (FPGA). The NAND flash memory provides large-capacity data storage, while the FPGA handles direct data processing within the smart SSD. During data reconstruction, the FPGAs within the smart SSDs interact with other smart SSDs via the CPU. This creates a data wall problem between different smart SSDs, limiting data flow efficiency between them. The network bandwidth between each smart SSD and the CPU is thus constrained during the data reconstruction process. Summary of the Invention

[0003] This application provides a data reconstruction method, a solid-state storage unit, and a storage system, which solves the problem that the data transfer efficiency between different storage devices is limited by the Internet bandwidth between the storage device and the CPU. This helps to reduce the transmission time required for data transfer and improve the data processing efficiency between different storage devices.

[0004] The technical solution adopted in this application is as follows.

[0005] Firstly, this application provides a data reconstruction method. This data reconstruction method is applied to a solid-state storage unit (SSU), which is connected to at least one storage device via a first bus, with each endpoint of the first bus interconnected peer-to-peer. The data reconstruction method includes: the SSU obtaining the address of a first storage device; the first storage device includes one or more of at least one storage device; the SSU obtaining first data associated with the first storage device based on the address of the first storage device; and the SSU reconstructing the first data to obtain multiple reconstructed data sets. Finally, the SSU writes the multiple reconstructed data sets into a target storage device within the at least one storage device.

[0006] In the first aspect of this application, the solid-state storage unit can directly read data from other storage devices during data reconstruction, and there is no need for data plane interaction between the solid-state storage unit and the host, thus achieving decoupling between the control plane and the data plane. Furthermore, the solid-state storage unit and other storage devices achieve data access through a peer-to-peer interconnect network, no longer limited by the bandwidth of the peripheral component interconnect express (PCIe) bus connected to the CPU in the host or the internet bandwidth accessed by the network card. This helps reduce data transfer time between the solid-state storage unit and other storage devices, improves data flow efficiency between the solid-state storage unit and other storage devices, and thus improves the data reconstruction efficiency of the solid-state storage unit.

[0007] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, the solid-state storage unit connects to the host via a first bus to obtain the address of the first storage device. This includes the solid-state storage unit receiving a data reconstruction command sent by the host, the data reconstruction command including the address of the first storage device. During the data reconstruction process, the solid-state storage unit only needs to interact with the host via control plane commands (i.e., data reconstruction commands). The solid-state storage unit can directly read data from other storage devices, and there is no need for data plane interaction between the solid-state storage unit and the host. This achieves decoupling between the control plane and the data plane, and also helps to reduce the CPU resource consumption in the host.

[0008] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, before the solid-state storage unit obtains the address of the first storage device, the data reconstruction method provided in this application further includes: in the event of a failure or alarm at an endpoint device connected to the first bus, the solid-state storage unit determines the first storage device to be read based on the reconstructed data range; or, in the event of an increase or decrease in the number of endpoint devices connected to the first bus, the solid-state storage unit determines the first storage device to be read based on the reconstructed data range. In the first aspect of this application, the solid-state storage unit offloads the computing power required by the CPU in the host, improves the data flow efficiency between the solid-state storage unit and other storage devices, and can also significantly reduce the resource consumption of the CPU in the host, allowing the CPU in the host to execute other computing devices or run applications during the data reconstruction process.

[0009] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, the solid-state storage unit stores access address ranges for each storage device in at least one storage device. The solid-state storage unit retrieves first data stored in the first storage device according to a data reconstruction command, including: the solid-state storage unit determining the access address associated with the first storage device from each access address range stored in the solid-state storage unit according to the data reconstruction type indicated by the data reconstruction command; and the solid-state storage unit reading the first data according to the access address associated with the first storage device.

[0010] Furthermore, before accessing other storage devices, the solid-state storage unit can obtain direct access to those other storage devices from the host, so that data access between the solid-state storage unit and other storage devices does not need to go through the host. This is beneficial to improving the data flow efficiency between the solid-state storage unit and other storage devices, and also helps to reduce the CPU resource consumption in the host.

[0011] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, in the event of a failure of the second storage device in at least one storage device, the first storage device stores data necessary for recovering the data in the second storage device. The aforementioned solid-state storage unit reconstructs the first data to obtain multiple reconstructed data, including: the solid-state storage unit performs erasure coding (EC) calculation on the first data to obtain multiple coded blocks; the coded blocks include at least one of data blocks or check blocks. Furthermore, the solid-state storage unit uses the coded block corresponding to the second storage device among the multiple coded blocks as multiple reconstructed data. In the EC redundancy protection scenario, the solid-state storage unit supports data recovery from the failed disk (second storage device) based on the read data, thereby improving the data security and robustness of the storage system including the solid-state storage unit and the aforementioned at least one storage device.

[0012] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, the EC algorithm used by the second storage device is the first algorithm. The solid-state storage unit performs EC calculation on the first data to obtain multiple coded blocks, including: the solid-state storage unit performs erasure coding calculation on the first data according to the first algorithm to obtain multiple coded blocks. In the first aspect of this application, the EC algorithm used in the data reconstruction process is consistent with the EC algorithm used by the faulty disk (second storage device), ensuring that the reconstructed data (coded blocks) obtained by the solid-state storage unit is consistent with the data stored in the second storage device, thus avoiding the problem of redundancy protection failure caused by data loss.

[0013] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, the second storage device uses the first algorithm as the EC algorithm. The solid-state storage unit performs EC calculation on the first data to obtain multiple coded blocks, including: the solid-state storage unit performs erasure coding calculation on the first data according to the second algorithm to obtain multiple coded blocks. Wherein, the first algorithm is EC(k+M), and the second algorithm is EC([ki]+[Mj]), where i and j are both natural numbers, and at least one of i and j is not 0. Here, k is the number of data blocks in an erasure coding stripe in the EC redundancy protection scenario, and M is the number of check blocks in an erasure coding stripe in the EC redundancy protection scenario. A data block refers to the stored original data, and a check block is the check data calculated based on the original data. In the first aspect of this application, compared with the EC algorithm used in the failed disk (second storage device), the EC algorithm used in the data reconstruction process has performed EC downgrading, that is, reducing the number of storage nodes required for EC redundancy protection. This ensures that in the event of a failure of the second storage device, the solid-state storage unit and at least one other storage device besides the second storage device can still support data redundancy protection, which is beneficial to improving data security and robustness.

[0014] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, when the data reconstruction type involves modifying the data format of a redundant array of independent disks (RAID), the solid-state storage unit (SSD) reconstructs the first data to obtain multiple reconstructed data sets. This includes the SSD modifying the RAID data format of the first data from a first RAID data format to a second RAID data format, resulting in multiple reconstructed data sets. For data reconstruction scenarios involving modifying the RAID data format, the SSD does not need to interact with the CPU / network card in the host during the RAID data format modification process. This reduces the high latency issues caused by interaction between the SSD and the host, frees up CPU / network card resources in the host, and improves the data transfer efficiency between different storage devices in the data reconstruction scenario.

[0015] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, the EC algorithm used by the first storage device is the first algorithm. When at least one storage device includes a newly added storage device connected to a solid-state storage unit via a first bus, the solid-state storage unit reconstructs the first data to obtain multiple reconstructed data, including: calculating erasure codes on the first data according to a third algorithm to obtain multiple coded blocks. Wherein, the first algorithm is EC(k+M), and the third algorithm is EC([k+i]+[M+j]), where i and j are both natural numbers, and at least one of i and j is not 0. Furthermore, the solid-state storage unit uses the multiple coded blocks as multiple reconstructed data. The EC algorithm used in the data reconstruction process has been upgraded, namely: increasing the number of storage nodes required for EC redundancy protection. In the event of a storage device failure, the data required to recover a set of original data is stored on a larger number of storage devices or storage nodes, improving the reliability of EC redundancy protection. Therefore, when adding disks or storage devices to a storage system, the EC algorithm used to protect the data is upgraded and then reconstructed. This allows the reconstructed data to be stored on a larger number of storage devices or storage nodes, avoiding data loss due to the failure of individual storage devices or solid-state storage units, and improving data security and robustness.

[0016] In conjunction with the data reconstruction method provided in the first aspect, in an optional implementation, after the solid-state storage unit writes multiple reconstructed data into the target storage device in at least one storage device, the data reconstruction method provided in this application further includes: the solid-state storage unit sending a data reconstruction response to the host, the data reconstruction response indicating that the data reconstruction process corresponding to the first storage device has been completed. During the data reconstruction process, the host does not need to be aware of the data interaction between the underlying solid-state storage unit and the storage device, the host will not experience a "busy" state leading to increased latency for other services, and the host's CPU resources will no longer be used for service-irrelevant scenarios, thus improving the performance of the solid-state storage unit and the computing node where the host resides.

[0017] In conjunction with the data reconstruction method provided in the first aspect, in one optional implementation, the aforementioned at least one storage device includes a combination of one or more of the following types: solid-state storage unit (SSU), hard disk drive (HDD), solid-state drive (SSD), magnetic tape, optical disk, dynamic random access memory (DRAM), dual in-line memory module (DIMM), storage class memory (SCM), non-volatile magnetic random access memory (MRAM), resistive random access memory (RRAM), ferroelectric random access memory (FeRAM), high bandwidth memory (HBM), or phase change memory (PCM), etc.

[0018] Secondly, this application provides another data reconstruction method. This data reconstruction method is applied to a storage system, which includes: a first solid-state storage unit and at least one storage device. The first solid-state storage unit is connected to the at least one storage device via a first bus, and the endpoint devices connected to the first bus are interconnected peer-to-peer. The data reconstruction method provided in this second aspect includes: the first solid-state storage unit obtaining the address of a first storage device; the first storage device includes one or more of at least one storage device. Furthermore, the first solid-state storage unit obtains first data associated with the first storage device based on the address of the first storage device. The first solid-state storage unit reconstructs the first data to obtain multiple reconstructed data. The first solid-state storage unit writes the multiple reconstructed data into a target storage device within the at least one storage device.

[0019] In the second aspect of this application, the solid-state storage unit only needs to interact with the host via control plane commands during data processing. The solid-state storage unit can directly read data from other storage devices, and no data plane interaction is required between the solid-state storage unit and the host, thus decoupling the control plane and data plane. Furthermore, the solid-state storage unit and other storage devices achieve data access through a peer-to-peer network, no longer limited by the bandwidth of the PCIe bus connected to the CPU in the host or the internet bandwidth accessed by the network card. This helps reduce data transfer time between the solid-state storage unit and other storage devices, improves data flow efficiency, and consequently increases the data processing efficiency of the solid-state storage unit. It also helps reduce the resource consumption of the CPU in the host.

[0020] In conjunction with the data reconstruction method provided in the second aspect, in one optional implementation, the first solid-state storage unit is connected to the host via a first bus, and the first solid-state storage unit obtains the address of the first storage device, including: the first solid-state storage unit receives a data reconstruction command sent by the host, the data reconstruction command including the address of the first storage device.

[0021] In conjunction with the data reconstruction method provided in the second aspect, in one optional implementation, before the first solid-state storage unit obtains the address of the first storage device, the data reconstruction method provided in the second aspect of this application further includes: in the event of a failure or alarm in the endpoint device connected to the first bus, the first solid-state storage unit determines the first storage device to be read based on the reconstructed data range; or, in the event of an increase or decrease in the number of endpoint devices connected to the first bus, the first solid-state storage unit determines the first storage device to be read based on the reconstructed data range.

[0022] In conjunction with the data reconstruction method provided in the second aspect, in one optional implementation, in the event of a failure of the second storage device in at least one storage device, the first storage device stores data necessary for recovering the data in the second storage device. The aforementioned first solid-state storage unit reconstructs the first data to obtain multiple reconstructed data, including: the first solid-state storage unit performing EC (Error Correction) calculations on the first data to obtain multiple coded blocks. Each coded block includes at least one of a data block or a check block; and the first solid-state storage unit uses the coded block corresponding to the second storage device among the multiple coded blocks as multiple reconstructed data.

[0023] Thirdly, this application provides a solid-state storage unit. The solid-state storage unit includes: a storage medium, a media controller, a communication interface, and a reconfiguration unit. The storage medium is used to store data. The media controller is connected to the storage medium and is used to manage the storage medium. The communication interface is connected to at least one storage device via a first bus, and the endpoint devices connected to the first bus are interconnected peer-to-peer. The reconfiguration unit is used to obtain a data reconfiguration command and execute the data reconfiguration method provided in the first aspect or any optional implementation thereof according to the data reconfiguration command.

[0024] Fourthly, this application provides a storage system. The storage system includes at least one storage device and a solid-state storage unit provided in the third aspect. The solid-state storage unit is connected to the at least one storage device via a first bus, and the devices at each endpoint of the first bus are interconnected peer-to-peer. The solid-state storage unit is used to acquire a data reconstruction command and execute a data reconstruction method provided in the first aspect or any optional implementation thereof according to the data reconstruction command. Alternatively, the first solid-state storage unit and the at least one storage device collaboratively execute a data reconstruction method provided in the second aspect or any optional implementation thereof.

[0025] Fifthly, this application provides a computer-readable storage medium. The computer-readable storage medium includes computer instructions. When the computer instructions are executed in a solid-state storage cell, the solid-state storage cell performs the operational steps of the method provided in the first aspect or any optional implementation thereof. Alternatively, when the computer instructions are executed in a storage system, the storage system performs the operational steps of the method provided in the second aspect or any optional implementation thereof.

[0026] Sixthly, this application provides a computer program product. This computer program product includes a computer program or instructions, which, when executed by an electronic device, cause the electronic device to perform operational steps of the method provided in the first aspect or any optional implementation thereof. Alternatively, the electronic device may perform operational steps of the method provided in the second aspect or any optional implementation thereof. Exemplarily, the electronic device may include, but is not limited to, a solid-state storage unit, a storage system, or a data access system including a solid-state storage unit or a storage system.

[0027] In a seventh aspect, this application provides a data access system. The data access system includes: one or more hosts, and a plurality of solid-state storage units provided in the third aspect. The one or more hosts are connected to the plurality of solid-state storage units via a first bus, and the endpoint devices of the first bus are interconnected peer-to-peer. A first host among the one or more hosts sends a data reconstruction command to a first solid-state storage unit among the plurality of solid-state storage units, the data reconstruction command including the address of the first storage device. The first solid-state storage unit is configured to execute, according to the data reconstruction command, the operation steps of the method provided in the first aspect or any optional implementation thereof. Alternatively, the first solid-state storage unit and the first host collaboratively execute the operation steps of the method provided in the second aspect or any optional implementation thereof.

[0028] The beneficial effects of aspects three through seven can be found in the description of any optional implementation in aspect one or two, and will not be repeated here. Based on the implementations provided in the above aspects, this application can be further combined to provide even more implementations. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the structure of a storage system provided in this application.

[0030] Figure 2 A software schematic diagram of a storage system provided in this application.

[0031] Figure 3 A flowchart illustrating a data reconstruction method provided in this application. Figure 1 .

[0032] Figure 4 A flowchart illustrating a data reconstruction method provided in this application. Figure 2 .

[0033] Figure 5 A flowchart illustrating a data reconstruction method provided in this application. Figure 3 .

[0034] Figure 6 A flowchart illustrating a data reconstruction method provided in this application. Figure 4 .

[0035] Figure 7 A flowchart illustrating a data reconstruction method provided in this application. Figure 5 .

[0036] Figure 8 A flowchart illustrating a data reconstruction method provided in this application. Figure 6 .

[0037] Figure 9A flowchart illustrating a data reconstruction method provided in this application. Figure 7 .

[0038] Figure 10 This is a schematic diagram of the structure of a reconfigurable unit provided in this application. Detailed Implementation

[0039] This application provides a data reconstruction method. During the data reconstruction process, the solid-state storage unit (SSD) only needs to interact with the host via control plane commands (i.e., data reconstruction commands). The SSD can directly read data from other storage devices, and there is no need for data plane interaction between the SSD and the host, thus achieving decoupling between the control plane and the data plane. Moreover, the SSD and other storage devices achieve data access through a peer-to-peer interconnection network, no longer limited by the bandwidth of the PCIe bus connected to the CPU in the host or the Internet bandwidth accessed by the network card. This helps reduce the data transmission time between the SSD and other storage devices, improves the data flow efficiency between the SSD and other storage devices, and thus improves the data reconstruction efficiency of the SSD. It also helps reduce the resource consumption of the CPU in the host.

[0040] The technical solutions involved in this application may be applied not only to current storage technologies or storage devices (such as solid-state storage cells), but also to future storage technologies or storage devices, or storage systems that include solid-state storage cells or storage devices. The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. A brief introduction to some concepts that may be involved in this application is provided below.

[0041] Storage medium: A storage material used to record sound, images, digital signals, or other signals. This storage material may include, but is not limited to, magnetic tape, magnetic disks (platters, discs), or optical disks. For example, magnetic tape refers to a strip-shaped material with a magnetic layer used to record sound, images, digital signals, or other signals. Similarly, magnetic disks may include, but are not limited to, hard disk drives (HDDs) and solid-state drives (SSDs).

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the application will now be described in further detail with reference to the accompanying drawings.

[0043] In the following description, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0044] Furthermore, in this application, directional terms such as "upper" and "lower" are defined relative to the orientation of the components shown in the accompanying drawings. It should be understood that these directional terms are relative concepts, used for relative description and clarification, and can change accordingly depending on the orientation of the components in the accompanying drawings.

[0045] During the data reconstruction process, the FPGAs in the smart SSDs interact with other smart SSDs through the CPU. There is a data wall problem between different smart SSDs, that is, the data flow efficiency between different smart SSDs is limited by the network bandwidth or bus bandwidth between each smart SSD and the CPU.

[0046] To address the aforementioned issues, the application scenarios of the embodiments of this application will be described below with reference to the accompanying drawings.

[0047] Figure 1 This application provides a schematic diagram of a storage system, which includes multiple computing nodes, such as computing nodes 1 to 3. Each computing node includes a host and a solid-state storage unit. The following description uses computing node 1 as an example to illustrate the host 10 and solid-state storage unit 11 included in computing node 1.

[0048] The host 10 can refer to a client, such as a desktop computer, server, laptop, or mobile device. The host 10 includes a CPU 101 and a network interface card (NIC) 102. The CPU 101 is used to process data access requests from outside the host 10 (server, client, or other storage system) and also to process requests generated internally within the host 10. For example, when the CPU 101 receives data access requests from a client via a communication interface, it temporarily stores the data in these requests in the host 10's memory. When the total amount of data in memory reaches a certain threshold, the CPU 101 sends the data stored in memory to the solid-state storage unit 11 for persistent storage via the communication interface. In some optional examples, the host 10 may include multiple CPUs 101; this application does not limit this.

[0049] Network interface card 102 can be a standard network interface card (NIC) or a smart NIC. A smart NIC, also known as a smart network adapter, not only performs the network transmission functions of a standard NIC but also provides a built-in programmable and configurable hardware acceleration engine. This improves application performance and significantly reduces CPU consumption in the host connected to the smart NIC during communication, providing more CPU resources for the application. For example, in a highly virtualized environment, CPU 101 in host 10 needs to run tasks related to an open virtual switch (OVS). Simultaneously, CPU 101 in host 10 also handles operations such as storage, online or offline encryption / decryption of data packets, deep packet inspection, firewall management, and complex routing. These operations not only consume significant CPU resources but also, due to CPU resource contention between different services, prevent the services from achieving optimal performance. NIC 102, as a hub connecting various services, accelerates these services. For example, when the network interface card 102 is a smart network interface card, the network interface card 102 includes a processor for performing acceleration functions. The processor may include, but is not limited to, processing chips such as data processing unit (DPU), graphics processing unit (GPU), and neural-network processing units (NPU), or it may refer to FPGA, tensor processing unit (TPU), microprocessor chip, digital signal processing (DSP), application-specific integrated circuit (ASIC), or one or more integrated circuit chips.

[0050] Optionally, the host 10 may also include memory, which refers to internal storage that directly exchanges data with the CPU 101. It can read and write data at any time and is very fast, serving as temporary data storage for the operating system or other running programs. Memory includes at least two types of storage; for example, it can be either random access memory (RAM) or read-only memory (ROM). For example, random access memory is DRAM or SCM. DRAM is a semiconductor memory and, like most random access memory (RAM), is a type of volatile memory device. However, DRAM and SCM are merely illustrative examples in this embodiment; memory may also include other random access memories, such as static random access memory (SRAM). For read-only memory, for example, it can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). In addition, the memory can also be a dual in-line memory module (DIMM), i.e., a module composed of dynamic random access memory (DRAM), or an SSD. In practical applications, the host 10 can be configured with multiple memory modules of different types. This embodiment does not limit the number or type of memory. Furthermore, the memory can be configured to have a power-saving function. The power-saving function means that when the operating system in the host 10 loses power and then is powered on again, the data stored in the memory will not be lost. Memory with a power-saving function is called non-volatile memory. Software programs are stored in the memory, and the CPU 101 can manage the hard disk by running the software programs in the memory. For example, the hard disk (such as solid-state storage unit 11) can be abstracted into a storage resource pool, and the storage resource pool can be provided to the server or users in the form of logical unit number (LUN). Here, the LUN is actually the hard disk seen on the server. Of course, some storage systems are also file servers themselves, and can provide shared file services to the server.

[0051] Please continue reading Figure 1 Solid-state storage unit 11, also known as the first solid-state storage unit, includes: multiple interfaces, a reconfiguration unit 112, a media controller 113, and a storage medium 114. The multiple interfaces include: a southbound interface, a northbound interface, and an eastbound interface.

[0052] In this embodiment, the southbound interface refers to the interface used for internal device connections, such as connecting the reconfiguration unit 112 and the media controller 113. The northbound interface refers to the interface used for external connections to other different devices, such as connecting the reconfiguration unit 112 to the bus 40. The eastbound interface refers to the interface used for external connections to other devices of the same type, such as connecting the reconfiguration unit 112 to the bus 40.

[0053] Bus 40 can also be referred to as the first bus or other names. In this embodiment, the endpoint devices connected to bus 40 are interconnected peer-to-peer. Peer-to-peer interconnection, also known as peer-to-peer connection (PC), is a high-bandwidth, high-quality cloud resource interconnection service that enables routing interoperability between devices. By configuring routing policies at both ends of the devices, interconnection can be achieved between devices in the same or different regions, or between the same or different users. Peer-to-peer interconnection does not depend on any independent hardware and does not have single-point-of-failure or bandwidth bottleneck issues. When multiple endpoint devices are connected through the peer-to-peer bus, the endpoint devices can communicate directly without data packets needing to be relayed through the public internet or CPU.

[0054] A communication network based on a peer-to-peer bus connection is called a peer-to-peer network. In a peer-to-peer network, each endpoint device supports hot-swapping and can automatically recover and retry network connections, ensuring high availability and fault tolerance of network connectivity. Furthermore, peer-to-peer networks can employ various network topologies, such as mesh, star, or other types, to meet the needs of more complex scenarios.

[0055] In this embodiment, the peer-to-peer bus 40 used to implement the peer-to-peer network may include, but is not limited to: unified bus (Ubus or UB), NVLink, etc. TM Computer Express Link (CXL) or other possible peer-to-peer interconnect buses are not limited to this application.

[0056] In the first optional scenario, the Unified Bus (UB) is also known as Lingqu. TM Bus, in Lingqu TM In a network composed of buses, various devices such as GPUs, DPUs, and CPUs can communicate directly with each other without the need for data to be relayed through the CPU in the host, which greatly improves the communication efficiency between different devices.

[0057] In the second optional scenario, NVLink TMIt is a vertically scalable interconnect bus that enables faster input of large datasets into models and rapid data exchange between CPUs and GPUs. Specifically, NVLink... TM It adopts a point-to-point structure and serial transmission, and is used for connection between CPU and GPU, and can also be used for interconnection between different GPUs.

[0058] In the third alternative scenario, CXL is a high-speed interconnect technology that provides higher data throughput and lower latency, helping to address the problem of high latency between CPUs and devices, and between devices.

[0059] This document uses a bus 40 with a UB as an example for illustrative purposes, but it should not be construed as meaning that the data reconstruction method provided in this application embodiment can only be applied to networks connected to the UB (UB network). In other words, the data reconstruction method provided in this application embodiment not only supports application in networks connected to the UB, but can also be applied to networks using NVLink. TM The connected network can also be used in CXL-connected networks, or other buses or networks that support peer-to-peer interconnection of endpoint devices.

[0060] In this article, an endpoint device refers to the receiving or sending end of data communication, which includes control command communication and service data communication. Control command communication is also called control plane communication, and service data communication is also called data plane communication.

[0061] Please see Figure 1 Bus 40 supports control plane communication between multiple hosts and multiple solid-state storage units. For example, CPUs 101, 201, and 301, network interface cards (NICs) 102, 202, and 302, and reconfiguration units 112, 212, and 312 communicate via bus 40. The control plane, also known as the management plane, refers to the transmission of control signaling or commands, rather than actual business data (such as voice data, image data, or others). Control commands are used to control call process establishment, process maintenance, or release of process resources. For example, control commands can be used to initiate data processing, including but not limited to: data reading, data writing, garbage collection (GC), data reconfiguration, EC verification, data compression / decompression, data migration, or other data reconfiguration processes.

[0062] Data reading refers to the reconstruction unit 112 reading data from the storage medium.

[0063] Data writing refers to the reconstruction unit 112 writing data into the storage medium.

[0064] The garbage collection (GC) process refers to the process where, if new data cannot be written to the storage area corresponding to garbage data in the storage medium, the reconstruction unit 112 needs to migrate the valid data in the storage block or page containing the storage area corresponding to the garbage data to increase the remaining available storage capacity in the storage medium. Garbage data refers to data that will not be read after GC, while valid data refers to data that will be read or used by processes or access devices (such as compute nodes, hosts, or solid-state storage units) after GC.

[0065] Data reconstruction refers to the process by which reconstruction unit 112 integrates data from different data sources into a unified data model for data analysis and computation.

[0066] EC (Error Correction) refers to a forward error correction technique primarily used in network transmission to prevent data packet loss. Storage systems utilize EC to improve storage reliability. Compared to multi-replica replication, EC achieves higher data reliability with less data redundancy.

[0067] Data compression refers to the process by which the reconstruction unit 112 calls a compression algorithm to compress the data in order to reduce the storage space occupied by the data.

[0068] Data decompression refers to the process by which the reconstruction unit 112 decompresses the compressed data, enabling the CPU 101 or other devices to perform business operations based on the decompressed data, such as playing audio or video, or performing other business types corresponding to the decompressed data.

[0069] Data migration refers to the refactoring unit 112 moving data from one storage area to another. These two storage areas can be located on the same storage medium or on different storage media.

[0070] The above data reconstruction process is merely an optional example provided in the embodiments of this application and should not be construed as a limitation of this application.

[0071] In this article, the data plane is also called the user plane. Data plane communication refers to the transmission of actual business data (also called real business data) between different solid-state storage units through the bus 40, such as voice data, image data or other types of business data.

[0072] Please see Figure 1Bus 40 also supports data plane communication between multiple solid-state storage units. Specifically, any two solid-state storage units among solid-state storage units 11, 21, and 31 can communicate via bus 40. For example, the data plane is also called the user plane, and data plane communication refers to the transmission of actual service data (also called real service data) by different solid-state storage units through this second UB network, such as voice data, image data, or other types of service data.

[0073] It is worth noting that the southbound interface, northbound interface, and eastbound interface mentioned above are merely optional names provided for embodiments of this application. In some feasible situations, the southbound interface is also called the internal interface or the first interface, the northbound interface is also called the first external interface or the second interface, and the eastbound interface is also called the second external interface or the third interface. This application does not limit the names of the multiple interfaces in the solid-state storage unit 11. In addition, depending on the number of devices that the solid-state storage unit 11 needs to connect, the solid-state storage unit 11 may also have other interfaces, such as power supply interfaces or other communication interfaces, which will not be elaborated here.

[0074] In the solid-state storage unit 11, the storage medium 114 is used to store data. This storage medium 114 can be one or more of the following: magnetic tape, optical disc, HDD media, SSD media, DRAM media, SRAM media, DIMM media, SCM media, non-volatile magnetic random access memory (MRAM) media, resistive random access memory (RRAM) media, ferroelectric RAM (FeRAM) media, high bandwidth memory (HBM) media, phase change memory (PCM) media, or other types of media, such as cache chips, flash memory chips (flash memory particles), or others. SCM is a composite storage technology that combines the characteristics of traditional storage devices and memory. Storage-class memory can provide faster read and write speeds than hard drives, but its access speed is slower than DRAM, and it is also cheaper than DRAM. For example, flash memory chips may include: XL-LAND, single-level cell (SLC), multi-level cell (MLC), trinary-level cell (TLC), quad-level cell (QLC), enterprise multi-level cell (eMLC), or others.

[0075] Media controller 113 is used to manage storage medium 114. Exemplarily, media controller 113 includes one or more processors and a cache. The processor is a CPU or other type of processing chip, used to process data access requests from outside the solid-state storage unit 11, and also to process requests generated internally by the solid-state storage unit 11. Exemplarily, when the processor receives a write data request sent by reconstruction unit 112, it temporarily stores the data in these write data requests in the cache. When the total amount of data in the cache reaches a certain threshold, the processor writes the data stored in the memory cache to the storage medium.

[0076] The reconstruction unit 112 is used to reconstruct data in the storage medium 114 or other storage media to obtain reconstructed data. The reconstruction unit 112 includes a calculation module and a command processing module. In one optional scenario, the calculation module and command processing module refer to two sets of integrated circuits with different functions in the reconstruction unit 112; in another optional scenario, the calculation module and command processing module refer to software modules or units configured on the integrated circuits included in the reconstruction unit 112. This application embodiment does not limit the specific implementation of the calculation module and command processing module.

[0077] The command processing module is used to receive control commands and decapsulate them; alternatively, the command processing module is used to encapsulate control information to obtain control commands and send the encapsulated control commands to other devices. For example, the control command includes one or more access addresses, the access addresses indicating storage areas located on storage medium 114 or other storage media.

[0078] This computing module is used to process data (such as data reconstruction, data compression / decompression, etc.) to obtain processed data. This computing module can offload the computing power of CPU 101 in host 10, reduce the resource consumption of CPU 101, and help improve the overall performance of computing node 1.

[0079] It is worth noting that the structure of computing node 1 described above is merely an example provided in the embodiments of this application and should not be construed as limiting this application. In some optional cases, a computing node may include multiple solid-state storage units, or a computing node may also include other memory, such as including but not limited to: HDD, SSD, DRAM, SRAM, DIMM, SCM, MRAM, RRAM, FeRAM, HBM, PCM, cache particles, flash memory chips (flash memory particles), or others, such as optical discs or magneto-electric disks (MED).

[0080] Please continue reading Figure 1The computing node 2 includes a host 20 and a solid-state storage unit 21. The host 20 includes a CPU 201 and a network interface card 202. The solid-state storage unit 21 includes multiple communication interfaces (northbound interface, southbound interface, and eastbound / westbound interface), a reconfiguration unit 212, a media controller 213, and a storage medium 214. The reconfiguration unit 212 includes a computing module and a command processing module.

[0081] Please continue reading Figure 1 The computing node 3 includes a host 30 and a solid-state storage unit 31. The host 30 includes a CPU 301 and a network interface card 302. The solid-state storage unit 31 includes multiple communication interfaces (northbound interface, southbound interface, and eastbound / westbound interface), a reconfiguration unit 312, a media controller 313, and a storage medium 314. The reconfiguration unit 312 includes a computing module and a command processing module.

[0082] The specific implementation of compute node 2 and compute node 3 can be found in the description of compute node 1 above, and will not be repeated here.

[0083] The above combination Figure 1 The structure of the storage system and the bus 40 used by each device have been described by way of example. The following is a description in conjunction with... Figure 2 The storage software of the storage system provided in this application embodiment and the peer-to-peer interconnection network corresponding to bus 40 are described below. Figure 2 A software schematic diagram of a storage system provided for this application. About Figure 2 The hardware structure of each computing node shown can be referenced. Figure 1 The description of that will not be repeated here.

[0084] Figure 1 The peer-to-peer interconnection network provided by the bus 40 shown includes: Figure 2 The first and second UB networks are shown. These are two abstract networks implemented based on bus 40. Both are peer-to-peer interconnection networks. For information on the characteristics and supported functions of peer-to-peer interconnection networks, please refer to [link to relevant documentation]. Figure 1 The description of that will not be repeated here.

[0085] exist Figure 2 In the first UB network shown, the first UB network is used to realize control plane communication between different endpoint devices. For example, the host 10 and the solid-state storage unit 11 transmit control commands, such as data reconstruction commands, through the first UB network.

[0086] exist Figure 2 In the illustrated second UB network, this second UB network is used to enable data plane communication between different endpoint devices. For example, service data is transmitted between solid-state storage units via the second UB network.

[0087] In addition, each host can also access the second UB network. For example, CPU 101, CPU 201, CPU 301, network interface card (NIC) 102, NIC 202, and NIC 302 support data plane communication with each solid-state storage unit. For instance, when host 10 receives a read / write request from another device, host 10 can communicate with one or more solid-state storage units via CPU 101 or NIC 102 to implement the access service corresponding to the read / write request.

[0088] Please continue reading. Figure 2 Regarding the software structure of each host, the following example uses host 10 as an illustration: host 10 has an operating system (OS) and storage software deployed in it.

[0089] The operating system (OS) is a program built into the host 10. The OS works in conjunction with various hardware components within the host 10 to interact with the user. Depending on the operating environment, operating systems (OS) can be categorized as: desktop operating systems, mobile operating systems, server operating systems, embedded operating systems, or others. The interface type provided by the operating system (OS) for user interaction can include, but is not limited to: command-line interface, graphical user interface, touch interface, natural user interface (NUI), or other interfaces.

[0090] Based on their operating environment, operating systems (OS) can be categorized as follows: Desktop Operating Systems (OS), Mobile Operating Systems (OS), Server Operating Systems (OS), Embedded Operating Systems (OS), and others. A desktop OS provides a "black box" for developers, allowing them to access its functionality through a series of standard system call functions. Server operating systems generally refer to those installed on mainframe computers, such as web servers, application servers, and database servers. Within a specific network, a server OS undertakes additional management, configuration, stability, and security functions. An embedded operating system is used in embedded systems. It is a versatile type of system software, typically including hardware-related low-level driver software, a system kernel, device driver interfaces, communication protocols, a graphical interface, and a standardized browser. An embedded operating system is responsible for allocating all software and hardware resources, scheduling tasks, and controlling and coordinating concurrent activities within the embedded system. It reflects the characteristics of the system in which it resides and can achieve the required functions by installing or removing certain modules. Mobile operating systems are mainly used in devices such as smartphones or smart tablets. They are developed from embedded operating systems and are specifically designed for mobile phones. In addition to the functions of embedded operating systems (such as process management, file system, network protocol stack, etc.), mobile operating systems also need to have power management for battery-powered systems, input / output for user interaction, embedded graphical user interface services that provide calling interfaces for upper-layer applications, low-level encoding and decoding services for multimedia applications, Java runtime environment, core wireless communication functions for mobile communication services, and upper-layer applications for smartphones.

[0091] Please see Figure 2 Each of the host units 10 to 30 is equipped with storage software, which is used to manage the various solid-state storage units (such as solid-state storage units 11 to 31) or other storage devices in the storage system. Figure 2 (not shown in the image), it can also be used to implement different clients ( Figure 2 (Not shown in the image) Data interaction between the computing nodes.

[0092] For example, the storage software is an application, such as a software module or unit that provides access functionality to the outside world. This software module can provide logical storage areas with management granularity that meet the access requirements of data access devices (such as hosts or clients). Such management granularity includes, but is not limited to, LUNs, data areas, data segments, data blocks, data pages, or others. In some optional cases, the logical storage area can refer to a memory pool. For example, the logical storage area can be supported by the physical storage areas included in each solid-state storage unit (such as storage media 114, storage media 214, or storage media 314 mentioned above). This application does not limit the type of storage media included in the physical storage area. Furthermore, after obtaining a data reconstruction request from a client, the storage software can also send a data reconstruction response corresponding to the data reconstruction request to the client.

[0093] Optionally, the storage software code is located on the host, or it can provide one or more calling interfaces to enable each host to use the relevant management functions of the storage software. The management functions supported by the storage software include, but are not limited to, one or more of the following: data read, data write, data migration, garbage collection, EC checksum, data reconstruction, data compression, data decompression, address mapping (or address translation), data snapshot, data replication (or data copy), data redundancy backup, or other data management functions. Therefore, this storage software is also called storage management software.

[0094] above Figure 1 and Figure 2 The content shown is merely an example of the storage system provided in the embodiments of this application and should not be construed as limiting this application. Depending on changes in user needs or adjustments to data reconstruction requirements, the storage system provided in the embodiments of this application may include more or fewer host or solid-state storage units. Other storage devices or memories may also be provided in the storage system, which is not limited in this application.

[0095] Below Figure 1 and Figure 2 Based on the accompanying drawings, the data reconstruction method provided in the embodiments of this application will be described by way of example. Figure 3 As shown, Figure 3 A flowchart illustrating a data reconstruction method provided in this application. Figure 1 .about Figure 3 The hardware and software implementations of the solid-state storage unit 11 and the host 10, the specific implementations of the first UB network and the second UB network, and the specific implementation of the storage software can be referred to the foregoing. Figure 1 and Figure 2 The content will not be repeated here.

[0096] Figure 3 The storage system shown also includes at least one storage device, such as a first storage device and a second storage device.

[0097] In a first alternative example, a storage device includes one of the aforementioned at least one storage device. For example, the first storage device is the aforementioned solid-state storage cell 21, and the second storage device is the aforementioned solid-state storage cell 31. Alternatively, the first storage device and the second storage device may be other types of memory or storage devices, which are not limited in this application.

[0098] In a second alternative example, a storage device includes multiple of the aforementioned at least one storage device. For example, the first storage device includes the aforementioned solid-state storage unit 21 and solid-state storage unit 31, and the second storage device includes one or more hard disks, which may refer to the aforementioned... Figure 1 or Figure 2 Other storage devices not shown, such as SSDs, HDDs, SSUs, or other types of storage media. The two optional examples above are merely examples of storage devices provided in the embodiments of this application and should not be construed as limiting this application.

[0099] Please see Figure 3 , Figure 3 The provided data reconstruction method can be used by Figure 1 or Figure 2 The storage system shown in this application executes a data reconstruction method, which includes the following steps S310 to S370.

[0100] S310, the storage software sends a data processing request to the host 10.

[0101] Correspondingly, host 10 receives data processing requests from storage software.

[0102] The data processing request may include, but is not limited to: read data request, write data request, data migration request, data reconstruction request, garbage collection request, compression request, decompression request, or other types of requests.

[0103] For example, the data processing request is generated by the storage software during operation.

[0104] For example, the data processing request is sent by the client to the storage software. The client can be a physical machine or a virtual machine, or a user or tenant on a cloud management platform, etc., which is not limited in this application.

[0105] As another example, the data processing request is sent by a host other than host 10 in the storage system to the storage software in host 10.

[0106] The above three methods are merely optional sources of data processing requests provided in the embodiments of this application and should not be construed as limiting this application. Depending on the data processing process, the source of the data processing request may also be other methods.

[0107] This application embodiment uses a data reconstruction request as an example to illustrate the data reconstruction method provided in this application embodiment. For example, the data reconstruction request may include, but is not limited to, information such as the address of the storage device and a reconstruction identifier. The reconstruction identifier is used to indicate the reconstruction method, reconstruction approach, reconstruction algorithm information, or data reconstruction type indicated by the data reconstruction request.

[0108] In response to the data processing request in S310, host 10 sends a data reconstruction command to solid-state storage unit 11.

[0109] Correspondingly, the solid-state storage unit 11 receives the data reconstruction command from the host.

[0110] In an optional scenario, the data reconstruction command includes the address of the first storage device. This address can be the disk identifier, disk address, UB endpoint address, or Internet Protocol (IP) address of the first storage device.

[0111] For example, the disk is identified as a drive letter. A drive letter is a relative identifier for a disk storage device by the operating system. The drive letter can be represented using 26 English characters plus a colon ":", but other representations are not limited in this application. For example, the address of the first storage device includes "SSU21:".

[0112] The difference between a data reconstruction command and the aforementioned data processing request is that a data reconstruction command is a control plane command and only includes control plane information, while a data reconstruction request includes not only control plane information but may also include data plane information (such as business data or the address of the first storage device).

[0113] In this embodiment, the data reconstruction command is used to indicate the data reconstruction type. For example, the data reconstruction type includes one or a combination of the following: disk recovery, storage expansion, or data format modification.

[0114] In the first optional example, the data reconstruction type includes: disk failure recovery. Disk failure recovery refers to the recovery of data from a disk (such as a solid-state storage unit or storage device) in a storage system after a failure or alarm occurs.

[0115] In the second optional example, the data reconstruction type includes: storage expansion. Storage expansion refers to migrating some data from other hard drives in the storage system to the new hard drive when a new hard drive (such as a solid-state storage unit or storage device) is added to the storage system.

[0116] In the third optional example, data reconstruction types include: data format modification. Data format modification refers to modifying the RAID data format in the storage system to adapt to the data organization of each hard drive (such as a solid-state storage unit or storage device) in the storage system.

[0117] In the fourth optional example, data reconstruction types include: disk failure recovery and storage expansion.

[0118] In the fifth optional example, data reconstruction types include: failed disk recovery and data format modification.

[0119] In the sixth optional example, data reconstruction types include: storage expansion and data format modification.

[0120] In the seventh optional example, data reconstruction types include: disk failure recovery, storage expansion, and data format modification.

[0121] The seven optional examples above are merely optional methods for data reconstruction types provided in the embodiments of this application and should not be construed as limiting this application. The data reconstruction type may change or be adjusted depending on different information such as changes in the state of the storage system and user needs, and this application does not limit it in this regard.

[0122] S330 and solid-state storage unit 11 respond to the data reconstruction command by sending a permission request to host 10.

[0123] Correspondingly, host 10 receives the permission request from solid-state storage unit 11.

[0124] This permission request instructs the host 10 to release direct access to other storage devices during the data reconstruction process, enabling the solid-state storage unit 11 to directly access those other storage devices. In other words, the data in those other storage devices can be directly read by the solid-state storage unit 11, or the solid-state storage unit 11 can directly write data to those other storage devices.

[0125] In one feasible example, the at least one storage device has only data storage functionality. Such at least one storage device may include, but is not limited to: magnetic tape, optical disc, HDD, SSD, DRAM, SRAM, DIMM, SCM, non-volatile magnetic random access memory (MRAM), resistive random access memory (RRAM), ferroelectric memory (FeRAM), high bandwidth memory (HBM), phase change memory (PCM), or other types.

[0126] In another feasible example, the at least one storage device has data storage functionality and data reconstruction functionality (or data computation functionality). This may include a memory and a reconstructor within the storage device. For example, the first storage device may refer to one or both of the aforementioned solid-state storage unit 21 or solid-state storage unit 31.

[0127] The two feasible examples above are merely optional methods for other storage devices provided in the embodiments of this application and should not be construed as limiting this application. In other optional methods, other storage devices may also include cache particles or flash memory chips, etc.

[0128] In this application embodiment, the storage device can use the same or different types of storage media as the solid-state storage unit 11, so that the data reconstruction method provided in this application embodiment can be applied to more types of storage systems, which is beneficial to improving the adaptability of the solid-state storage unit near the storage end to different types of storage systems.

[0129] S340, host 10 responds to the authorization request by sending an authorization response to solid-state storage unit 11.

[0130] Correspondingly, the solid-state storage unit 11 receives the delegation response sent by the host 10.

[0131] The decentralization response indicates that the solid-state storage unit 11 supports direct access to at least one of the aforementioned storage devices.

[0132] It is worth noting that steps S330 and S340 are not mandatory. In some optional implementations, the data reconstruction command in S320 may carry an authorization identifier, which indicates that the solid-state storage unit 11 supports direct access to at least one of the aforementioned storage devices. Alternatively, the host 10 may send an authorization response to the solid-state storage unit 11 before sending the data reconstruction command to it. Or, the host 10 may send an authorization response to the solid-state storage unit 11 after sending the data reconstruction command to it.

[0133] As an optional implementation, the solid-state storage unit 11 can send the aforementioned grant request via asynchronous message passing. Asynchronous message passing means that the solid-state storage unit 11 does not wait for messages from the host 10; even if the host 10 is shut down, the message passing will still succeed. During asynchronous message passing, the solid-state storage unit 11 does not need to wait for a response after sending the grant request and can continue processing other tasks.

[0134] In this embodiment, before accessing other storage devices, the solid-state storage unit 11 can obtain direct access to the other storage devices from the host 10, so that data access between the solid-state storage unit 11 and other storage devices does not need to go through the host, which is beneficial to improving the data flow efficiency between the solid-state storage unit 11 and other storage devices, and also beneficial to reducing the CPU resource consumption in the host 10.

[0135] S350, the solid-state storage unit 11 obtains the first data associated with the first storage device based on the address of the first storage device in the aforementioned data reconstruction command.

[0136] In one alternative implementation, the solid-state storage unit 11 stores the access address range of each of the aforementioned at least one storage device.

[0137] For example, solid-state storage unit 11 stores the access address ranges of solid-state storage units 11, 21, and 31. For instance, the access address range of solid-state storage unit 11 includes [0, 255], the access address range of solid-state storage unit 21 includes [256, 511], and the access address range of solid-state storage unit 31 includes [512, 767]. The access address ranges of each solid-state storage unit provided in this example are merely optional methods provided in the embodiments of this application and should not be construed as limiting this application.

[0138] Regarding the process of solid-state storage unit 11 acquiring the first data, the following is combined with Figure 4 An optional implementation method is provided. Figure 4 A flowchart illustrating a data reconstruction method provided in this application. Figure 2 Please see. Figure 4 The above-mentioned S350 includes the following steps S351 and S352.

[0139] S351, the solid-state storage unit 11 determines the access address associated with the first storage device from each access address range stored in the solid-state storage unit 11 according to the data reconstruction type indicated by the data reconstruction command.

[0140] The contents of the data reconstruction types can be found in the description of S320 above, and will not be repeated here.

[0141] In one alternative scenario, the access address is a physical storage address, such as a physical block address (PBA). The PBA indicates the actual location of the storage space represented by the access address within a solid-state storage cell or memory.

[0142] In another alternative scenario, the access address is a logical storage address, such as a logical block address (LBA). The LBA refers to the address used by the storage software in the host to manage different storage spaces within the storage pool.

[0143] The two feasible examples above are merely examples of access addresses provided in the embodiments of this application and should not be construed as limiting this application.

[0144] For example, the address of the first storage device includes "SSU21:", and the range of each access address stored in the solid-state storage unit 11 includes: SSU11: [0,255], SSU21: [256,511], SSU31: [512,767]. Then, the access address associated with the first storage device determined by the solid-state storage unit 11 includes: [256,511].

[0145] S352, Solid-state storage unit 11 reads first data according to the access address associated with the first storage device.

[0146] For example, the access address associated with the first storage device includes [256, 511], and the solid-state storage unit 11 reads the first data stored in the first storage device according to the access address.

[0147] In one alternative implementation, the first storage device includes a storage medium 114 in the solid-state storage unit 11, and the first data includes persistently stored data in the storage medium 114 and data stored in other storage devices in the first storage device.

[0148] In another alternative implementation, the first storage device does not include the solid-state storage unit 11, and the first data includes the data stored in the first storage device.

[0149] The optional implementations of the first storage device and the first data are merely examples provided in the embodiments of this application and should not be construed as limiting this application. In other optional implementations, the storage space indicated by the access address may also be located in a database or data center that is accessible by the solid-state storage unit 11.

[0150] Combination Figure 4As can be seen from the provided embodiments, after the solid-state storage unit 11 obtains the address of the first storage device, it can determine the access address associated with the first storage device from the access address range of local storage, and read the first data corresponding to the determined access address. During the processes of S351 and S352, the solid-state storage unit 11 does not need to interact with the CPU or network card in the host. The solid-state storage unit 11 offloads the processing resources of the host, releases the computing power of the CPU and network card in the host, and also helps to improve the data flow efficiency between the solid-state storage unit 11 and other storage devices (such as other SSUs).

[0151] Please continue reading. Figure 3 The data reconstruction method provided in this application embodiment also includes the following S360 and S370.

[0152] S360 and solid-state storage unit 11 reconstruct the first data to obtain multiple reconstructed data.

[0153] In S360 provided in this application embodiment, the data reconstruction process for reconstructing multiple data items may include, but is not limited to, one or a combination of the following: disk failure recovery, storage expansion, data format modification, data reading, data writing, garbage collection (GC) process, data reconstruction, EC verification, data compression / decompression, data migration, or other data reconstruction processes. For details on the various types of data reconstruction processes, please refer to the foregoing. Figure 1 The description of S320 will not be repeated here. For details on the implementation methods of fault disk recovery, storage expansion, and data format modification, please refer to the following... Figures 7 to 9 The relevant descriptions will not be repeated here.

[0154] Based on the contents of S320, S350 and S360, it can be seen that in addition to having traditional solid-state storage functions, solid-state storage unit 11 can also provide computing power to data reconstruction or data computing services through the computing power on the reconstruction unit, so as to offload the CPU computing power in host 10, thereby improving the performance of computing node 1 and storage system.

[0155] S370, solid-state storage unit 11 writes multiple reconstructed data into a target storage device in at least one storage device.

[0156] In the first optional scenario, the target storage device is the first storage device.

[0157] In the second alternative scenario, the target storage device includes newly added storage devices within the storage system.

[0158] In the third alternative scenario, the target storage device includes solid-state storage unit 11 and other storage devices.

[0159] The above three optional scenarios are merely optional methods for the target storage device provided in the embodiments of this application, and should not be construed as limiting this application.

[0160] Based on the contents of S320, S350, S360 and S370, it can be seen that during the data reconstruction process, the solid-state storage unit 11 only needs to interact with the host 10 to exchange control plane commands (i.e. data reconstruction commands). The solid-state storage unit 11 can directly read data from other storage devices, and there is no need for data plane interaction between the solid-state storage unit 11 and the host 10, thus achieving decoupling between the control plane and the data plane.

[0161] Furthermore, the solid-state storage unit 11 communicates with other storage devices via a peer-to-peer interconnect network (such as... Figure 2 The second UB network shown enables data access, no longer limited by the bandwidth of the PCIe bus connected to the CPU in the host 10 or the Internet bandwidth accessed by the CPU / network card. This helps to reduce the data transmission time between the solid-state storage unit 11 and other storage devices, improve the data flow efficiency between the solid-state storage unit 11 and other storage devices, thereby improving the data reconstruction efficiency of the solid-state storage unit 11, and also helps to reduce the resource consumption of the CPU in the host 10.

[0162] As an optional implementation, to improve the robustness of data reconstruction, in Figure 3 and Figure 4 On this basis, Figure 5 A flowchart illustrating a data reconstruction method provided in this application. Figure 3 For details on the implementation of S310 to S370, please refer to [link / reference]. Figure 3 and Figure 4 The provided embodiments are not described in detail here. Please refer to [link to relevant documentation]. Figure 5 Following the above-described S370, the data reconstruction method provided in this application embodiment further includes the following S380.

[0163] S380, solid-state storage unit 11 sends a data reconstruction response to host 10.

[0164] Corresponding to the process in S380, the host 10 receives the data reconstruction response sent by the solid-state storage unit 11.

[0165] The data reconstruction response indicates that the data reconstruction process for the first storage device has been completed.

[0166] Figure 5 Different from Figure 3 The point is: in Figure 5In the first storage device, there are solid-state storage units 11, 21, and 31, and the target storage device also includes solid-state storage units 11, 21, and 31. Therefore, during the process of reading the first data, solid-state storage unit 11 needs to read from storage media 114, 214, and 314 respectively; during the process of writing reconstructed data, solid-state storage unit 11 needs to write different reconstructed data to storage media 114, 214, and 314 respectively.

[0167] In combination with the above Figures 3 to 5 As can be seen from the embodiments, during the period from when the host 10 sends the data reconstruction command until it receives the reconstruction response from the solid-state storage unit 11, the CPU 101 and the network card 102 in the host 10 do not need to participate in the data reconstruction process. During this period, the computing power of the CPU 101 in the host 10 can be used to perform other services, such as computing tasks or storage tasks.

[0168] In other words, the solid-state storage unit 11 offloads the computing power required by the CPU 101 in the host 10. The CPU 101 only needs a minimal amount of computing power to support command forwarding on the control plane. This not only improves the data flow efficiency between the solid-state storage unit 11 and other storage devices, but also significantly reduces the resource consumption of the CPU 101 in the host 10. This allows the CPU 101 in the host 10 to execute other computing tasks or run applications during data processing, which is beneficial to improving the overall performance of computing node 1. In other words, the host 10 is unaware of the data interaction between the underlying solid-state storage units during data reconstruction. The host 10 will not experience a "busy" state that would increase latency for other services. CPU resources are no longer used in scenarios unrelated to business operations, thus improving the performance of computing node 1.

[0169] In addition, Figures 3 to 5 In the illustrated embodiment, the solid-state storage unit 11 obtains the address of the first storage device from the host 10. However, in some alternative implementations, the solid-state storage unit 11 also supports locally determining the address of the first storage device. The following section discusses this further. Figure 6 Provided as an example, Figure 6 A flowchart illustrating a data reconstruction method provided in this application. Figure 4 .about Figure 6 The hardware and software implementations of each component can be referred to the foregoing. Figure 1 or Figure 2 The description of that will not be repeated here.

[0170] Please see Figure 6 The data reconstruction method provided in this application includes the following steps S610 to S640.

[0171] S610, Solid State Storage Unit 11 obtains the address of the first storage device.

[0172] The first storage device includes one or more of at least one storage device.

[0173] The following describes several optional implementation methods to illustrate the triggering method of S610.

[0174] The first optional implementation: In the event of a failure or alarm at the endpoint device connected to the bus 40, the solid-state storage unit 11 determines the first storage device to be read based on the reconstructed data range.

[0175] In optional example 1, in the event of a failure of an endpoint device connected to bus 40, solid-state storage unit 11 determines the first storage device to be read based on the reconstructed data range.

[0176] For example, the solid-state storage unit 11 periodically checks the status of each endpoint device and determines whether the endpoint device connected to the bus 40 has failed. Alternatively, the endpoint device connected to the bus 40 may proactively report a fault status to each solid-state storage unit or the host.

[0177] The reconstructed data range is determined based on the access address range corresponding to each endpoint device of the connection bus 40. For example, if the faulty endpoint device is solid-state storage unit 31, the address range corresponding to solid-state storage unit 31 is [512, 767], and the data stored in solid-state storage unit 31 is determined based on the data in solid-state storage unit 11 and solid-state storage unit 21, then the first storage device includes solid-state storage unit 11 and solid-state storage unit 21.

[0178] In optional example 2, in the event of an alarm at an endpoint device connected to bus 40, solid-state storage unit 11 determines the first storage device to be read based on the reconstructed data range.

[0179] The second optional implementation: When the number of endpoint devices connected to the bus 40 increases or decreases, the solid-state storage unit 11 determines the first storage device to be read based on the reconstructed data range.

[0180] In optional example 3, when the number of endpoint devices connected to bus 40 increases, solid-state storage unit 11 determines the first storage device to be read based on the reconstructed data range.

[0181] For example, if the storage system 120 is expanded, such as when a user inserts a new storage device into the rack or server corresponding to the storage system 120, the solid-state storage unit 11 will detect the new storage device and determine the first storage device to be read based on the reconstructed data range.

[0182] In optional example 4, when the number of endpoint devices connected to bus 40 is reduced, solid-state storage unit 11 determines the first storage device to be read based on the reconstructed data range.

[0183] For example, if the storage system 120 is scaled down, and a user removes one or more storage devices from the rack or server corresponding to the storage system 120, the solid-state storage unit 11 will detect a storage device failure and determine the first storage device to be read based on the reconstructed data range.

[0184] The optional examples 1 to 4 above are merely optional triggering methods provided by the embodiments of this application and should not be construed as limiting this application. In some other optional implementations, the solid-state storage unit 11 periodically reconstructs the data in the storage system, including the solid-state storage unit 11, to improve the data security of the storage system.

[0185] S620, the solid-state storage unit 11 obtains the first data associated with the first storage device according to the address of the first storage device.

[0186] For details on the implementation of S620, please refer to the aforementioned S350 or... Figure 4 The description of that will not be repeated here.

[0187] S630 and solid-state storage unit 11 reconstruct the first data to obtain multiple reconstructed data.

[0188] For details on the implementation of S630, please refer to the aforementioned S360. Figure 7 , Figure 8 or Figure 9 The relevant descriptions will not be repeated here.

[0189] S640, solid-state storage unit 11 writes multiple reconstructed data into a target storage device in at least one storage device.

[0190] Combination Figure 6 As can be seen from the content, the solid-state storage unit 11 offloads the computing power required by the CPU 101 in the host 10, improves the data flow efficiency between the solid-state storage unit 11 and other storage devices, and can also significantly reduce the resource consumption of the CPU 101 in the host 10, so that the CPU 101 in the host 10 can execute other computing devices or run applications during the data reconstruction process.

[0191] The data reconstruction methods provided in the embodiments of this application will be further explained below in conjunction with different data reconstruction types.

[0192] (I) First data reconstruction scenario: recovery of faulty disk.

[0193] In the event of a failure of the second storage device in the aforementioned at least one storage device, the aforementioned first storage device stores data necessary for recovering the data in the second storage device. The process of reconstructing the first data using the solid-state storage unit 11... Figure 7 An optional implementation method is provided, such as Figure 7 As shown, Figure 7 A flowchart illustrating a data reconstruction method provided in this application. Figure 5 The aforementioned S360 or S630 includes the following S710 and S720.

[0194] S710 and solid-state storage unit 11 perform EC calculation on the first data to obtain multiple encoded blocks.

[0195] A coded block includes at least one of a data block or a parity block. A data block refers to the stored original data; it can also be called a raw data block. A parity block is the check data calculated based on the original data. In EC redundancy protection, the combination of k data blocks and m parity blocks calculated from those k data blocks is called an erasure code stripe (EC stripe).

[0196] The following explanation uses the EC algorithm employed by the second storage device as an example, specifically EC(k+M). In this paper, EC(k+M) represents the process of performing EC calculations on k data blocks to obtain M parity blocks. These k data blocks and M parity blocks are then stored in k+M storage nodes, with each storage node storing only one data block or one parity block from an erasure code stripe. This storage node can be the aforementioned solid-state storage unit, storage device, or other memory; this application does not limit its use.

[0197] In the first optional scenario, the solid-state storage unit 11 performs erasure coding calculations on the first data according to the first algorithm to obtain multiple coded blocks. That is, the EC algorithm used in the data reconstruction process is the same as the EC algorithm used in the failed disk (second storage device), so that the coded blocks obtained by the solid-state storage unit 11 are consistent with the data stored in the second storage device, avoiding the problem of redundancy protection failure caused by data loss.

[0198] In the second optional scenario, the solid-state storage unit 11 performs erasure coding calculations on the first data according to the second algorithm to obtain multiple coded blocks. The second algorithm is EC([ki]+[Mj]), where i and j are both natural numbers, and at least one of i and j is not 0.

[0199] For example, the first algorithm is EC(3+2) and the second algorithm is EC(2+2).

[0200] For example, the first algorithm is EC(3+2) and the second algorithm is EC(3+1).

[0201] For example, the first algorithm is EC(3+2) and the second algorithm is EC(2+1).

[0202] The first and second algorithms described above are merely optional methods provided in the embodiments of this application and should not be construed as limiting this application.

[0203] In this embodiment, compared to the EC algorithm used by the failed disk (second storage device), the EC algorithm used in the data reconstruction process has been downgraded, that is, the number of storage nodes required for EC redundancy protection has been reduced. This means that even if the second storage device fails, the solid-state storage unit 11 and at least one other storage device besides the second storage device can still support data redundancy protection, which is beneficial to improving data security and robustness.

[0204] S720, the solid-state storage unit 11 uses the coding block corresponding to the second storage device among multiple coding blocks as multiple reconstruction data.

[0205] Combination Figure 7 As can be seen from the provided embodiments, in the scenario of disk failure recovery, whether the original EC algorithm is used to recover data from the failed disk or the EC degradation method is used, it is beneficial to improve the data security and robustness of the storage system including the solid-state storage unit 11 and the aforementioned at least one storage device. Moreover, during the data reconstruction process, the solid-state storage unit 11 does not need to interact with the CPU or network card in the host. The solid-state storage unit 11 offloads the computing resources of the host, allowing the CPU or network card in the host to perform other types of services during the data reconstruction process, which is beneficial to improving the overall performance of the computing node.

[0206] (ii) Second data reconstruction scenario: storage expansion.

[0207] In the case where at least one of the aforementioned storage devices includes a newly added storage device connected to a solid-state storage unit via bus 40, the aforementioned first storage device stores data required to recover data from the second storage device. The process of reconstructing the first data using the solid-state storage unit 11... Figure 8 An optional implementation method is provided, such as Figure 8 As shown, Figure 8 A flowchart illustrating a data reconstruction method provided in this application. Figure 6 The aforementioned S360 or S630 includes the following S810 and S820.

[0208] S810 and solid-state storage unit 11 perform erasure coding calculation on the first data according to the third algorithm to obtain multiple coding blocks.

[0209] The first algorithm is EC(k+M), and the third algorithm is EC([k+i]+[M+j]), where i and j are both natural numbers, and at least one of i and j is not 0.

[0210] In scenario 1, the first algorithm is EC(3+2) and the third algorithm is EC(3+3).

[0211] In scenario 2, the first algorithm is EC(3+2) and the third algorithm is EC(4+2).

[0212] In scenario 3, the first algorithm is EC(3+2) and the third algorithm is EC(4+3).

[0213] The first and third algorithms described above are merely optional methods provided in the embodiments of this application and should not be construed as limiting this application.

[0214] S820 and solid-state storage unit 11 use multiple coded blocks as multiple reconstruction data.

[0215] In this embodiment, the EC algorithm used in the data reconstruction process has been upgraded, specifically by increasing the number of storage nodes required for EC redundancy protection. In the event of a storage device failure, the data needed to recover a set of original data is stored on a larger number of storage devices or nodes, improving the reliability of EC redundancy protection. Specifically, when a new disk / storage device is added to the storage system, the EC algorithm used to reconstruct the data to be protected is upgraded, ensuring that the reconstructed data is stored on a larger number of storage devices / nodes. This avoids data loss due to the failure of individual storage devices or solid-state storage units, improving data security and robustness.

[0216] (III) The third data reconstruction scenario: data format modification.

[0217] When the data reconstruction type is modifying the RAID data format, the process of reconstructing the first data for solid-state storage unit 11 is as follows: Figure 9 An optional implementation method is provided, such as Figure 9 As shown, Figure 9 A flowchart illustrating a data reconstruction method provided in this application. Figure 7 The aforementioned S360 or S630 includes the following S910.

[0218] S910 and solid-state storage unit 11 modify the RAID data format of the first data from the first RAID data format to the second RAID data format to obtain multiple reconstructed data.

[0219] RAID data format refers to the data format determined by the RAID level used in the storage system. RAID levels include: RAID 0, RAID 1, RAID 5, RAID 6, RAID 10, RAID 50, RAID 60, or others.

[0220] RAID 0: Striping (data is divided into blocks) but without redundancy, providing high read and write performance.

[0221] RAID 1: Mirroring, data is completely copied to another hard drive, providing fault tolerance.

[0222] RAID 5: Striping plus distributed parity provides data redundancy and read performance.

[0223] RAID 6: Similar to RAID 5, but offers a higher level of fault tolerance.

[0224] RAID 10: RAID 1+0 combines RAID 1 mirroring into RAID 0 striping, providing higher fault tolerance and read / write performance.

[0225] RAID 50: RAID 5 combined into RAID 0 provides higher performance and fault tolerance.

[0226] RAID 60: RAID 6 is combined into RAID 0, providing a higher level of performance and fault tolerance.

[0227] For a more detailed description of each RAID level, please refer to the content of the general technical documentation; this application will not elaborate on it further.

[0228] For example, the first RAID data format is RAID 5 level, and the second data format is RAID 50 level. For instance, RAID data formats can be marked according to RAID level qualifiers.

[0229] Combination Figure 9 As can be seen from the provided embodiments, for data reconstruction scenarios involving modification of RAID data format, the solid-state storage unit 11 does not need to interact with the CPU 101 / network card 102 in the host 10 during the modification of RAID data format. This reduces the problem of high service latency caused by interaction between the solid-state storage unit 11 and the host 10, frees up the computing resources of the CPU 101 / network card 102 in the host, and helps to improve the data flow efficiency between different storage devices in the data reconstruction scenario.

[0230] The above embodiments are illustrated using data reconstruction as an example, but the methods provided in this application can also be applied to scenarios such as data migration or garbage collection. The data reconstruction methods for data migration and garbage collection scenarios are described below by way of example.

[0231] I. Data Migration.

[0232] Data migration refers to moving data from one storage area to another.

[0233] Combination Figure 2 For example, if the source storage area for storing the target data is storage medium 214 and the target storage area for storing the target data is storage medium 314, then the data reconstruction method provided in this application embodiment can be applied to a data migration scenario. The reconstruction unit 112 can first read the target data from storage medium 214 and then write the target data directly to storage medium 314.

[0234] In data migration scenarios, the data flow between different solid-state storage units does not pass through the host. That is, there is no need for data plane interaction between the solid-state storage units and the host. Instead, the solid-state storage units communicate directly with each other. The solid-state storage units access each other through the second UB network. This is no longer limited by the bandwidth of the PCIe bus connected to the CPU in the host or the Internet bandwidth accessed by the CPU / network card. This helps to reduce the data transmission time between the solid-state storage units, improve the data flow efficiency between the solid-state storage units, and thus improve the data processing efficiency of the solid-state storage units. It also helps to reduce the resource consumption of the CPU and network card in the host.

[0235] II. Waste recycling.

[0236] Garbage collection refers to the process of migrating valid data from the storage block or page containing garbage data in a storage medium, where new data cannot be written to the garbage data storage area. This increases the remaining usable storage capacity in the storage medium. Garbage data refers to data that will not be read after garbage collection, while valid data refers to data that will be read or used by processes or access devices (such as compute nodes, hosts, or solid-state storage units) after garbage collection.

[0237] The following is combined Figure 2 By way of example, the method provided in this application embodiment is applied to a data migration scenario. The reconstruction unit 112 can first read all the valid data in multiple storage areas to be GC from the storage medium 214, merge and splice all the valid data, and then write the merged or spliced ​​data directly to the storage medium 214.

[0238] Before garbage collection, several storage areas in storage medium 214 awaiting garbage collection contained some garbage data, resulting in low resource utilization in these storage areas. After garbage collection, storage medium 214 only needs to provide a small number of storage areas to completely store the merged or spliced ​​data. This allows storage areas in multiple storage areas that do not contain merged or spliced ​​data to be used to store other data or metadata, thus improving the resource utilization of storage medium 214.

[0239] In summary, in the technical solution provided by this application, the CPU, network card, and each solid-state storage unit in the host can communicate via the first UB network for control plane communication, and the solid-state storage units can communicate via the second UB network for data plane communication. Control commands and service data during data processing are decoupled. During data plane communication, the reconfiguration unit in the solid-state storage unit interacts with other storage devices via the second UB network, offloading the CPU's computing resources from the host and reducing the CPU load. The CPU only needs minimal computing power to support information forwarding on the control plane and interaction with the storage software control plane, freeing up its computing power for other services.

[0240] It is understood that, in order to achieve the functions in the above embodiments, the solid-state storage unit and storage system include hardware structures and / or software modules corresponding to perform each function. Those skilled in the art should readily recognize that, based on the units and method steps of the various examples described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0241] The above text combines Figures 1 to 9 This document describes in detail the solid-state storage unit, storage system, and data reconstruction method provided according to embodiments of this application. The following section, in conjunction with... Figure 10 The controller provided in the embodiments of this application will be described by way of example.

[0242] Figure 10This application provides a schematic diagram of a reconstruction unit 1000, which includes a memory 1010 and at least one processing circuit 1020. The processing circuit 1020 can implement the data reconstruction method provided in the above embodiments, and the memory 1010 is used to store software instructions corresponding to the data reconstruction method. As an optional implementation, in hardware implementation, the reconstruction unit 1000 can refer to an integrated chip or chip system encapsulated with one or more processing circuits 1020. For example, when the reconstruction unit 1000 is used to implement the method steps in the above embodiments, the processing circuit 1020 included in the reconstruction unit 1000 executes the steps of the processor or controller in the above method and its possible sub-steps. In an optional case, the reconstruction unit 1000 may further include an interface circuit 1030, which can be used to send and receive data. For example, the interface circuit 1030 is used to receive control commands such as data processing commands / reconstruction commands, or to send data processing responses / reconstruction responses; the interface circuit 1030 can be used to implement the communication function of the reconstruction unit 1000. Therefore, in some examples, the interface circuit 1030 can also be referred to as the transceiver of the reconfiguration unit 1000. In the embodiments of this application, the interface circuit 1030, processing circuit 1020, and memory 1010 can be connected via a bus 1040, which can be divided into an address bus, data bus, control bus, etc. The bus 1040 can be a PCIe bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a computeexpress link (CXL), a cache coherent interconnect for accelerators (CCIX), or other types of buses, etc. For example, the communication interface provided by the interface circuit 1030 can be such as... Figure 1 The north-facing interface, south-facing interface, or east-west-facing interface are shown.

[0243] The reconstruction unit 1000 provided in this embodiment may be the reconstruction unit 112, reconstruction unit 212, or reconstruction unit 312, or other devices with data processing or data calculation functions, and this application does not limit it in this regard. For example, when other processing devices in the solid-state storage unit also have data processing or data calculation functions, the reconstruction unit 1000 may refer to the other processing devices in the aforementioned solid-state storage unit.

[0244] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as an SSD.

[0245] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Various equivalent modifications or substitutions can be conceived within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data reconstruction method, characterized in that, The method is applied to a solid-state storage unit, which is connected to at least one storage device via a first bus, and the devices at each endpoint of the first bus are interconnected peer-to-peer. The method includes: Obtain the address of a first storage device, wherein the first storage device includes one or more of the at least one storage device; Based on the address of the first storage device, obtain the first data associated with the first storage device; The first data is reconstructed to obtain multiple reconstructed data. The multiple reconstructed data are written to the target storage device in the at least one storage device.

2. The method according to claim 1, characterized in that, The solid-state storage unit is connected to the host via the first bus, and obtaining the address of the first storage device includes: The system receives a data reconstruction command sent by the host, the data reconstruction command including the address of the first storage device.

3. The method according to claim 1, characterized in that, Before obtaining the address of the first storage device, the method further includes: In the event of a failure or alarm at an endpoint device connected to the first bus, the first storage device to be read is determined based on the reconstructed data range; or, As the number of endpoint devices connected to the first bus increases or decreases, the first storage device to be read is determined based on the reconstructed data range.

4. The method according to any one of claims 1-3, characterized in that, The solid-state storage unit stores the access address range of each of the at least one storage device; The step of obtaining the first data stored in the first storage device according to the data reconstruction command includes: Based on the data reconstruction type indicated by the data reconstruction command, the access address associated with the first storage device is determined from each access address range stored in the solid-state storage unit; The first data is read based on the access address associated with the first storage device.

5. The method according to any one of claims 1-4, characterized in that, In the event of a failure of the second storage device in the at least one storage device, the first storage device stores data necessary for recovering data from the second storage device, wherein the data reconstruction of the first data yields multiple reconstructed data sets, including: The first data is subjected to erasure coding (EC) calculation to obtain multiple coding blocks; the coding blocks include at least one of data blocks or parity blocks. The coding block corresponding to the second storage device among the plurality of coding blocks is used as the plurality of reconstructed data.

6. The method according to claim 5, characterized in that, The second storage device uses the first algorithm as the EC algorithm, wherein the EC calculation is performed on the first data to obtain multiple coded blocks, including: The first data is subjected to erasure coding according to the first algorithm to obtain the plurality of coded blocks.

7. The method according to claim 5, characterized in that, The second storage device uses the first algorithm as the EC algorithm, wherein the EC calculation is performed on the first data to obtain multiple coded blocks, including: The first data is subjected to erasure coding calculation according to the second algorithm to obtain the plurality of coded blocks; Wherein, the first algorithm is EC(k+M) and the second algorithm is EC([ki]+[Mj]), where i and j are both natural numbers, and at least one of i and j is not 0.

8. The method according to any one of claims 1-7, characterized in that, When the data reconstruction type is modifying the RAID data format of a disk array, the first data is reconstructed to obtain multiple reconstructed data, including: The RAID data format of the first data is modified from the first RAID data format to the second RAID data format to obtain multiple reconstructed data.

9. The method according to any one of claims 1-8, characterized in that, The EC algorithm used by the first storage device is the first algorithm; when at least one storage device includes a newly added storage device connected to the solid-state storage unit via the first bus, the data reconstruction of the first data yields multiple reconstructed data, including: The first data is subjected to erasure coding calculation according to the third algorithm to obtain the plurality of coded blocks; Wherein, the first algorithm is EC(k+M) and the third algorithm is EC([k+i]+[M+j]), where i and j are both natural numbers, and at least one of i and j is not 0; The plurality of coded blocks are used as the plurality of reconstructed data.

10. The method according to claim 2, characterized in that, After writing the plurality of reconstructed data to the target storage device in the at least one storage device, the method further includes: A data reconstruction response is sent to the host, the data reconstruction response indicating that the data reconstruction process corresponding to the first storage device has been completed.

11. The method according to any one of claims 1-10, characterized in that, The at least one storage device includes one or a combination of the following types: solid-state storage unit (SSU), hard disk drive (HDD), solid-state drive (SSD), magnetic tape, optical disk, random access memory (DRAM), dual-line memory module (DIMM), storage-class memory (SCM), non-volatile magnetic random access memory (MRAM), resistive random access memory (RRAM), ferroelectric memory (FeRAM), high-bandwidth memory (HBM), or phase-change memory (PCM).

12. A data reconstruction method, characterized in that, The method is applied to a storage system, which includes: a first solid-state storage unit and at least one storage device, wherein the first solid-state storage unit is connected to the at least one storage device via a first bus, and the devices at each endpoint of the first bus are interconnected peer-to-peer. The method includes: The first solid-state storage unit obtains the address of the first storage device; the first storage device includes one or more of the at least one storage device. The first solid-state storage unit obtains the first data associated with the first storage device based on the address of the first storage device; The first solid-state storage unit reconstructs the first data to obtain multiple reconstructed data. The first solid-state storage unit writes the plurality of reconstructed data into the target storage device in the at least one storage device.

13. The method according to claim 12, characterized in that, The first solid-state storage unit is connected to the host via the first bus. The first solid-state storage unit obtains the address of the first storage device, including: The first solid-state storage unit receives a data reconstruction command sent by the host, the data reconstruction command including the address of the first storage device.

14. The method according to claim 12, characterized in that, Before obtaining the address of the first storage device, the method further includes: In the event of a failure or alarm at an endpoint device connected to the first bus, the first solid-state storage unit determines the first storage device to be read based on the reconstructed data range; or, When the number of endpoint devices connected to the first bus increases or decreases, the first solid-state storage unit determines the first storage device to be read based on the reconstructed data range.

15. The method according to any one of claims 12-14, characterized in that, In the event of a failure of the second storage device in the at least one storage device, the first storage device stores data necessary to recover the data in the second storage device; The first solid-state storage unit reconstructs the first data to obtain multiple reconstructed data, including: The first solid-state storage unit performs erasure coding (EC) calculation on the first data to obtain multiple coding blocks; the coding block includes at least one of a data block or a parity block. The first solid-state storage unit uses the coding block corresponding to the second storage device among the plurality of coding blocks as the plurality of reconstructed data.

16. A solid-state storage cell, characterized in that, include: Storage medium, used to store data; A media controller, connected to the storage medium, is used to manage the storage medium; The communication interface connects to at least one storage device via a first bus, and the endpoint devices connected to the first bus are interconnected peer-to-peer. The reconstruction unit is configured to: obtain a data reconstruction command and execute the method according to any one of claims 1-11 based on the data reconstruction command.

17. A storage system, characterized in that, include: At least one storage device and the solid-state storage unit as described in claim 16; The solid-state storage unit is connected to the at least one storage device via a first bus, and the devices at each endpoint of the first bus are interconnected peer-to-peer. The solid-state storage unit is used to obtain a data reconstruction command and execute the method of any one of claims 1-11 according to the data reconstruction command, or the first solid-state storage unit and the at least one storage device cooperate to execute the method of any one of claims 12-15.