A memory-computation separated storage system and a method for improving the network transmission efficiency of the system
By integrating lightweight TCP/IP protocol stack and RDMA technology in SSD and combining with the coroutine scheduling mechanism, the performance bottleneck of traditional storage systems in high concurrency and large throughput scenarios is solved, and low-latency and efficient data transmission and resource management is achieved, suitable for distributed storage and cloud computing.
Patent Information
- Application Number
- CN202510325051.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Traditional storage systems have performance bottlenecks in high concurrency and large throughput scenarios, the CPU burden is too heavy, and the network card bandwidth is limited, resulting in the system throughput cannot be linearly increased, and the multi-stage hardware structure increases power consumption and cost.
Sink network transmission and storage instruction processing to SSD, integrate lightweight TCP/IP protocol stack, RDMA technology and co-course scheduling mechanism, and use FPGA's PS and PL resources to realize UDP communication and RDMA data processing, reducing intermediate layer participation.
Reduce system delay and power consumption, achieve linear expansion of the number of SSDs and the total system bandwidth, improve transmission rate and response speed, optimize data paths, and improve throughput and concurrency performance.
Smart Images

Figure CN119835267B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network communication technologies. For example, it relates to a memory-computation separation storage system and a method for improving the network transmission efficiency of the system. Background Art
[0002] Traditional storage systems use a centralized architecture of "network card + CPU + PCIe + SSD" for data interaction. However, with the explosion of data volume and the increasing requirement for real-time performance, its limitations have gradually emerged. Data transmission needs to pass through multiple hardware levels, increasing latency. Moreover, the CPU bears an excessive burden during data parsing, protocol conversion, etc., becoming a performance bottleneck in high-concurrency scenarios. In addition, the network card bandwidth is limited and difficult to match the performance improvement of SSDs, resulting in the system throughput rate not increasing linearly or even decreasing in performance. At the same time, the multi-level hardware structure increases power consumption and cost, posing challenges to the operation of data centers and cloud services.
[0003] To break through the limitations of the traditional architecture, the concept of memory-computation separation has gradually been accepted and widely applied in the industry. In the memory-computation separation architecture, computing nodes and storage nodes are interconnected through a high-speed network. Computing nodes no longer rely on local storage but directly access remote storage devices through the network. This architecture makes resource allocation more flexible, and storage resources can also be centrally managed and dynamically expanded, thus enhancing the overall availability and efficiency of the system. However, although memory-computation separation has obvious advantages in theory, in practical applications, how to achieve low-latency and high-throughput cross-node data transmission is still a major challenge.
[0004] To solve this problem, the NVMe over Fabrics (NVMeoF) technology has emerged. The NVMeoF technology extends the NVMe storage protocol for local SSDs to remote devices, enabling computing nodes to directly access SSDs, thus significantly reducing data transmission latency and enhancing the system throughput. Although the NVMeoF technology alleviates the performance bottleneck of cross-node access to a certain extent, its implementation still relies on the CPU for packet processing and protocol conversion, which may still bring relatively high processing overhead in high-concurrency and large-traffic scenarios.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a general review nor is it intended to identify key / important elements or delineate the scope of protection of these embodiments. Instead, it serves as a preamble to the following detailed description.
[0007] The embodiments of the present disclosure provide a storage and computing separated storage system and a method for improving the network transmission efficiency of the system, aiming to optimize the data transmission efficiency of the storage system, increase the throughput, reduce the access latency, and enhance the processing ability for high-concurrency storage tasks.
[0008] The core idea of the present invention is to sink the network transmission and storage instruction processing to the SSD, so as to solve the data interaction problem between the storage and computing nodes in the storage and computing separated architecture. This method utilizes the lightweight lwIP protocol stack and the coroutine scheduling mechanism to give full play to the advantages of the PS and PL in the FPGA, and realizes the efficient processing of UDP communication, RDMA link establishment, and data reading and writing, thereby reducing power consumption, simplifying the hardware configuration, and improving the transmission rate and response speed.
[0009] The core of sinking the network communication and storage instruction processing functions directly into the SSD lies in integrating a network processing unit in the SSD, enabling each SSD to have independent network communication capabilities, thereby reducing or even eliminating the participation of intermediate layers such as the CPU, PCIe, and network card in the traditional storage node. In this way, not only can the system latency and power consumption be significantly reduced, but also the linear expansion between the number of SSDs and the total system bandwidth can be achieved, providing a new technical path for building a high-performance storage system.
[0010] In terms of network communication protocols, the traditional complete TCP / IP protocol stack is difficult to be directly applied in embedded systems or operating system-free environments due to its complex functions and large volume. For this reason, lightweight TCP / IP protocol stacks such as lwIP (LightWeight IP) have emerged. lwIP realizes complete network functions with relatively small resource occupancy, and its modular and streamlined design enables it to operate efficiently on resource-constrained embedded platforms. lwIP not only supports common network protocols such as TCP and UDP, but also has a high response speed and flexible interfaces, providing strong guarantees for systems with high real-time requirements.
[0011] In addition, to meet the requirements of low-latency and high-throughput data transmission, high-speed network communication technologies are also constantly evolving. Among them, the Remote Direct Memory Access (RDMA) technology has received extensive attention because it can directly access the memory of remote devices in the network without CPU intervention, thus realizing low-latency and high-rate data transmission. RDMA-based solutions, such as RoCE (RDMA over Converged Ethernet), achieve transmission effects similar to dedicated high-speed networks (such as Infiniband) in the Ethernet environment. Through RDMA, data can be directly transmitted from the sending end to the memory area of the receiving end, reducing multiple data copies and intermediate processing overheads in traditional data transmission, and significantly improving the overall performance of the system.
[0012] At the software design level, multi-task scheduling technology is also an important link in realizing high-concurrency data processing. In traditional operating systems, task scheduling often relies on thread or process mechanisms. However, in embedded systems or operating-system-free environments, these mechanisms are difficult to meet the requirements of real-time and lightweight due to large resource overheads. As a lightweight concurrent programming model, coroutines greatly reduce the overhead of context switching by implementing task switching in the user space. Coroutines can manage multiple tasks within a single thread, ensuring efficient and fast switching between tasks, thus meeting the requirements of task scheduling in high-concurrency scenarios. With the help of coroutines, multi-task scheduling not only realizes the efficient utilization of resources but also provides technical support for the stable operation of the system.
[0013] Based on the above analysis, the present invention proposes a memory-computation separated storage system and a method for improving the network transmission efficiency of the system. The memory-computation separated system includes an Ethernet IP core, a RoCE IP core, a DMA IP core, a MAC address encapsulation IP core, a packet shunting IP core, an address application unit, an ARM core, and a flash memory. The Ethernet IP core is used to establish a network connection with a computing node. After the network connection is established, communication starts. Packets are shunted by the packet shunting IP core. A part of the packets is transmitted to the ARM core for processing, and the other part flows to the RoCE IP core for RDMA packet processing. The RoCE IP core is responsible for processing and replying to Ack packets. At the same time, the MAC address encapsulation IP core adds a MAC header to the Ack packets replied by the ROCE IP core. The DMA IP core is used to process data interaction between the PL and the PS. The address application unit is responsible for the processing of the LWIP protocol, the distributed hash table DHT, the RDMA link establishment, and the address mapping module during Flash operations. The flash memory chip realizes non-volatile storage of data. The ARM core is responsible for the control and verification functions when reading and writing flash memory data.
[0014] As a further improvement, the storage system further includes a cache module for caching read and write data.
[0015] In some embodiments, the method is implemented based on the above storage system and includes:
[0016] Integrate the network protocol stack and storage instruction processing into Ethernet-SSD to enable SSD to have independent network communication capabilities, and adopt a lightweight TCP / IP protocol stack as the network protocol stack of Ethernet-SSD to optimize UDP or TCP communication;
[0017] Integrate a RoCE IP core and a DMA module for implementing remote direct memory access in Ethernet-SSD. The RoCE IP core is responsible for protocol parsing and data packing, and the processed data is sent to the memory or storage medium through the DMA module;
[0018] In the case of high-concurrency storage requests, Ethernet-SSD introduces a coroutine-based task scheduling method to manage and schedule multiple tasks. When multiple storage requests are received, Ethernet-SSD dynamically creates concurrent tasks through a coroutine pool. After the tasks are completed, the system automatically reclaims the relevant resources.
[0019] As a further improvement, after integrating the network protocol stack into Ethernet-SSD, protocol stack initialization is performed. Ethernet-SSD automatically configures the IP address, subnet mask, and default gateway of the device, completes the initialization of the network adapter, and establishes a network connection between Ethernet-SSD and the computing node. After the network connection is established, Ethernet-SSD starts to receive data packets from the computing node, parses the data packets, extracts the key information therein, and passes it to the storage management module inside Ethernet-SSD. The storage management module performs corresponding storage operations according to the extracted key information, and dynamically schedules tasks according to the priority of the storage request, the storage area, and the data size.
[0020] As a further improvement, the key information includes the storage operation type, target address, and data length.
[0021] As a further improvement, during the data transmission process of remote direct memory access, Ethernet-SSD and the computing node first exchange link establishment requests and confirmation information through the network to establish a transmission channel for storing data, and then perform data transmission. The process of establishing the transmission channel for storing data is as follows:
[0022] The computing node sends a link establishment request message containing its own QPN and PSN to the storage node;
[0023] After receiving the link establishment request message, the storage node Ethernet-SSD extracts the QPN and PSN of the computing node;
[0024] The storage node Ethernet-SSD extracts the transaction id that identifies this connection from the request link establishment message, fills the transaction id and the QPN and PSN of the storage node side into the reply message, and sends the reply message to the computing node;
[0025] The computing node extracts the QPN and PSN of the storage node from the reply message for subsequent composition of the RDMA message;
[0026] The computing node composes an RTU message and sends it to the storage node, indicating that the connection is established and RDMA communication can be performed;
[0027] After receiving the RTU message, the storage node establishes a connection and conducts RDMA communication.
[0028] As a further improvement, after the data transmission is completed, Ethernet-SSD generates an ACK confirmation packet and feeds back the transmission status to the computing node to ensure the integrity and reliability of the data transmission. When the computing node does not need RDMA communication, it actively disconnects the connection. The disconnection process is as follows:
[0029] The computing node initiates a disconnection request message to disconnect the specified QPN connection with the storage node;
[0030] After receiving the disconnection request message, the storage node assembles a reply message, keeps the transaction id consistent with the disconnection request message, and replies to the computing node;
[0031] After receiving the reply, the computing node destroys the QPN and disconnects the connection.
[0032] As a further improvement, the messages between the computing node and the storage node are sent through the CM message in the UDP callback function. After receiving the message sent by the computing node, the storage node parses the Attribute ID and Opcode fields in the message to determine the message type and perform corresponding operations.
[0033] As a further improvement, the coroutine-based task scheduling method realizes the switch between coroutines by saving and restoring the context information of the coroutines. The scheduling process includes coroutine creation, status management, context switching, coroutine switch condition check, and task scheduling.
[0034] The memory-computation separation storage system provided by the embodiments of the present disclosure and the method for improving the network transmission efficiency of the system can achieve the following technical effects:
[0035] (1) Integration of storage and network protocols to optimize the data path. By integrating the network protocol stack with the storage instruction processing into Ethernet-SSD, the SSD has independent network communication capabilities, reduces CPU and PCIe dependencies, reduces latency, and improves scalability.
[0036] (2) Optimization of the lightweight protocol stack for efficient data communication. The lwIP lightweight TCP / IP protocol stack is adopted to optimize UDP communication, reduce protocol conversion and redundant data interaction, and improve the storage request response speed.
[0037] (3) Efficient transmission based on RDMA to reduce CPU load. The RoCE protocol is adopted to achieve RDMA zero-copy data transmission without CPU intervention, greatly improving throughput and reducing storage access latency.
[0038] (4) Coroutine-based multi-task scheduling to improve concurrency performance. The task scheduling is optimized using the coroutine mechanism to reduce the overhead of context switching, enabling the Ethernet-SSD to operate efficiently under high-concurrency requests and improving the throughput rate.
[0039] (5) Hardware-software co-optimization to reduce system overhead. By combining the PS and PL resources of the FPGA and integrating hardware modules such as RoCE IP cores and DMA IP cores, end-to-end optimization is achieved, reducing intermediate links and improving the energy efficiency ratio.
[0040] The above general description and the following description are only exemplary and explanatory, and are not intended to limit the present application. Brief Description of the Drawings
[0041] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and among them:
[0042] Figure 1 is a schematic diagram of the RDMA link establishment process;
[0043] Figure 2 is a schematic diagram of the RDMA disconnection process;
[0044] Figure 3 is a schematic diagram of the modular Ethernet-SSD architecture. Detailed Description of the Embodiments
[0045] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below in conjunction with the drawings. The attached drawings are only for reference and explanation, and are not intended to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, a sufficient understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and devices can be shown in a simplified manner.
[0046] Terms such as "first" and "second" in the embodiments of the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion.
[0047] Unless otherwise specified, the term "plurality" means two or more.
[0048] In the embodiments of the present disclosure, the character " / " indicates an "or" relationship between the front and rear objects. For example, A / B means: A or B.
[0049] The term "and / or" is an associative relationship describing an object, indicating that there can be three relationships. For example, A and / or B means: A or B, or, the three relationships of A and B.
[0050] The term "correspond to" can refer to an associative relationship or a binding relationship. A corresponding to B means that there is an associative relationship or a binding relationship between A and B.
[0051] Embodiment 1
[0052] This embodiment discloses a storage-computation separation storage system. The specific implementation is a modular Ethernet-SSD, that is, an SSD integrated with a network protocol stack and a storage instruction processing function. As Figure 3 shown, it includes the following modules:
[0053] An Ethernet IP core, a RoCE IP core, a DMA IP core, a MAC address encapsulation IP core, a packet shunting IP core, an address application unit, an ARM core, a flash memory, and a cache. The Ethernet IP core provides a high-speed network connection through a 10G Ethernet network card IP and an sfp+ interface, so as to establish a network connection between a storage node (i.e., the storage system described in this embodiment) and a computing node. After the network connection is established, communication starts. The data packet is shunted by the packet shunting IP core. One part is transmitted to the ARM core for processing, and the other part flows to the RoCE IP core for RDMA packet processing. The RoCE IP core is responsible for processing and replying to the Ack packet. At the same time, the MAC address encapsulation IP core adds a MAC header to the Ack packet replied by the ROCE IP. The DMA IP core is used to process the data interaction between the PL (programmable logic) and the PS (processing system). The address application unit is responsible for the processing of the LWIP protocol, the distributed hash table DHT, the RDMA link establishment, and the address mapping module during Flash operations. To improve the data read / write efficiency, the system also includes a cache module, which is part of the DDR, reducing the read / write of the Flash memory and reducing the latency. The flash memory realizes the non-volatile storage of data. Finally, a real-time processing module is provided in the ARM core, which is responsible for the control and verification functions during the read / write of the flash memory data, including an error verification module to ensure the integrity and reliability of the data.
[0054] Embodiment 2
[0055] This embodiment discloses a method for improving the network transmission efficiency of a storage-computation separation storage system. This method is implemented based on the system described in claim 1 and includes:
[0056] 1. Integrate the network protocol stack and storage instruction processing into Ethernet-SSD, enabling SSD to have independent network communication capabilities. Adopt a lightweight TCP / IP protocol stack as the network protocol stack of Ethernet-SSD to optimize UDP or TCP communication. The lightweight TCP / IP protocol stack avoids the complexity and performance loss brought by traditional protocol stacks. This streamlined version of the TCP / IP protocol stack is specifically optimized for embedded environments, enabling Ethernet-SSD to operate independently without an operating system.
[0057] 2. Integrate the RoCE IP core and DMA module for implementing Remote Direct Memory Access (RDMA) in Ethernet-SSD. The RoCE IP core is responsible for protocol parsing and data packaging, and the processed data is sent to memory or storage media through the DMA module. Remote Direct Memory Access (RDMA) technology can further improve data transfer efficiency, avoiding the data transfer in traditional storage systems often relying on the CPU for protocol conversion and data transfer, reducing latency and resource consumption, and also significantly improving the system throughput.
[0058] 3. In the case of high-concurrency storage requests, Ethernet-SSD needs to efficiently manage and schedule multiple tasks. This embodiment introduces a coroutine-based task scheduling method to manage and schedule multiple tasks, optimizing the system's concurrent processing ability. Compared with traditional multi-threaded or multi-process scheduling, the coroutine mechanism can significantly reduce the overhead of context switching, thus improving the overall performance of the system. When receiving multiple storage requests, Ethernet-SSD dynamically creates concurrent tasks through a coroutine pool. After the tasks are completed, the system automatically reclaims relevant resources, avoiding the high overhead in traditional thread or process management. Coroutine scheduling can not only improve the concurrent execution efficiency of storage tasks, but also, by combining with the RDMA data flow scheduling strategy, achieve intelligent task allocation and execution, thereby enhancing the system throughput and load balancing ability.
[0059] After integrating the network protocol stack into the Ethernet-SSD, the protocol stack is initialized. The Ethernet-SSD automatically configures the device's IP address, subnet mask, and default gateway to complete the initialization of the network adapter, establishing a network connection between the Ethernet-SSD and the computing node. After the network connection is established, the Ethernet-SSD starts to receive data packets from the computing node, parse the data packets, and extract the key information. The key information includes the storage operation type, target address, and data length. Then, the key information is passed to the storage management module inside the Ethernet-SSD. The storage management module executes the corresponding storage operations based on the extracted key information and dynamically schedules tasks according to the priority of the storage requests, storage areas, and data sizes. High-priority requests are processed first, and low-priority requests are executed later, thus achieving efficient allocation of system resources. This scheduling mechanism can ensure that the system maintains low latency and high throughput in a high-concurrency environment.
[0060] During the data transfer process of remote direct memory access, the Ethernet-SSD and the computing node first exchange link establishment requests and confirmation messages through the network to establish a transmission channel for storing data, and then perform data transfer. As Figure 1 shown, the process of establishing a transmission channel for storing data is as follows:
[0061] The computing node sends a link establishment request message containing its own QPN and PSN to the storage node;
[0062] After receiving the link establishment request message, the storage node Ethernet-SSD extracts the QPN and PSN of the computing node;
[0063] The storage node Ethernet-SSD extracts the transaction id that identifies the connection from the request link establishment message, fills the transaction id and the QPN and PSN of the storage node into the reply message, and sends the reply message to the computing node;
[0064] The computing node extracts the QPN and PSN of the storage node from the reply message for subsequent composition of the RDMA message;
[0065] The computing node composes an RTU message and sends it to the storage node, indicating that the connection is established and RDMA communication can be performed;
[0066] After receiving the RTU message, the storage node establishes a connection and performs RDMA communication.
[0067] After the data transfer is completed, the Ethernet-SSD generates an ACK confirmation packet and feeds back the transfer status to the computing node to ensure the integrity and reliability of the data transfer; when the computing node does not require RDMA communication, it actively disconnects the connection, as Figure 2 shown, the disconnection process is as follows:
[0068] The computing node initiates a disconnection request message to disconnect the specified QPN connection with the storage node;
[0069] After receiving the disconnection request message, the storage node assembles a reply message, keeps the transaction id consistent with the disconnection request message, and replies to the computing node;
[0070] The computing node receives the reply, destroys the QPN, and disconnects the connection.
[0071] In this embodiment, the messages between the computing node and the storage node are sent through the CM (Communication Management) messages in the UDP callback function. After receiving the messages sent by the computing node, the storage node parses the Attribute ID and Opcode fields in the messages to determine the message type and perform corresponding operations. After the connection is successful, the key information such as the QPN and PSN of the Client side is extracted, and combined with the local QPN and PSN information, it is sent to the RoCE IP core to process subsequent RDMA Read / Write operations.
[0072] As shown below is the process of the UDP callback function handling CM messages:
[0073] UDP callback function handling CM message strategy:
[0074] Requirement: Extract and use key metadata when receiving CM messages, and handle successful connection establishment or disconnection;
[0075] Input: CM message;
[0076] Output: Reply Reply message or connection status;
[0077] Step 1: Extract the Attribute ID and Opcode fields from the CM message
[0078] attribute_id = CM message.AttributeID
[0079] opcode = CM message.Opcode
[0080] Step 2: Determine the message type according to the Attribute ID and Opcode fields
[0081] if attribute_id == 0x0010 && opcode == 0x64:
[0082] CM Req message processing
[0083] Extract information such as QPN and PSN of the Client side
[0084] Fill in the local QPN and PSN, and reply with Reply
[0085] Reply with Reply (CM message)
[0086] else if attribute_id == 0x0012 && opcode == 0x64:
[0087] CM Rej message processing
[0088] Reject connection establishment
[0089] else if attribute_id == 0x0014 && opcode == 0x64:
[0090] CM Rtu message processing
[0091] Connection successful
[0092] else if attribute_id == 0x0015 && opcode == 0x64:
[0093] CM Disconnect message processing
[0094] Disconnect the connection and destroy relevant resources
[0095] else:
[0096] Unknown message type, ignore or handle error
[0097] end if.
[0098] In this embodiment, the coroutine-based task scheduling method is used to efficiently manage the execution order and switching of multiple coroutines. This method realizes the switching between coroutines by saving and restoring the context information of coroutines, ensuring the efficient execution of concurrent tasks. The scheduling process includes coroutine creation, status management, context switching, coroutine switching condition checking, and task scheduling to ensure that tasks are executed as needed. The specific process is as follows:
[0099] Coroutine task scheduling strategy:
[0100] Requirement: Implement the switching between coroutines by saving and restoring the context information of coroutines, and perform task scheduling according to the states of coroutines (such as running, blocked, exited, etc.);
[0101] Input: Coroutine control block, including the state, ID, stack attributes, context information, etc. of each coroutine;
[0102] Output: The execution order and states of all coroutines;
[0103] while (there are runnable coroutines):
[0104] Step 1: Obtain the next coroutine to be executed
[0105] next_coroutine = obtain the next runnable coroutine
[0106] Step 2: If the current coroutine state is runnable, switch to this coroutine
[0107] if the current coroutine state == runnable:
[0108] Save the context of the current coroutine
[0109] Switch to the context of the next coroutine
[0110] else if the current coroutine state == exiting:
[0111] Destroy the coroutine
[0112] Return to the scheduler to continue scheduling other coroutines
[0113] else if the current coroutine state == blocked:
[0114] Wait for resources
[0115] Set the coroutine state to runnable
[0116] Add it to the runnable queue
[0117] else if the current coroutine state == dead:
[0118] continue
[0119] end while.
[0120] The technical solution of the embodiments of the present disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The foregoing storage medium may be a non-transitory storage medium, including: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes, or it may also be a transient storage medium.
[0121] By integrating a lightweight TCP / IP protocol stack, RDMA technology, coroutine scheduling mechanism, and hardware acceleration module in Ethernet-SSD, the invention effectively solves the performance bottleneck of traditional storage architectures in high-concurrency and high-throughput data access scenarios. By optimizing all aspects of the storage system, Ethernet-SSD not only provides efficient data transmission but also maximizes the utilization of task scheduling and resource management, ensuring the stability and efficiency of the system in high-load and high-concurrency environments. This innovative solution provides an efficient and reliable storage solution for fields such as distributed storage, cloud computing, and high-performance computing, and has broad application prospects.
[0122] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments merely represent possible variations. Unless explicitly required, the individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terms used in this application are only for describing the embodiments and do not limit the scope of protection. As used in the description herein, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms as well. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, or apparatus comprising the element. Herein, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the various embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method parts disclosed in the embodiments, the relevant parts may refer to the description of the method parts.
[0123] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. The skilled person may use different methods for each specific application to achieve the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. The skilled person can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0124] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the couplings, direct couplings, or communication connections shown or discussed between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components can be either physically separated or not. The components shown as units can be either physical units or not, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, the various functional units can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit.
Claims
1. A memory-computation separated storage system, characterized in that: It includes an Ethernet IP core, a RoCE IP core, a DMA IP core, a MAC address encapsulation IP core, a packet shunt IP core, an address application unit, an ARM core, and a flash memory. The Ethernet IP core is used to establish a network connection with a computing node. After the network connection is established, communication starts. Packets are shunted by the packet shunt IP core. One part is transmitted to the ARM core for processing, and the other part flows to the RoCE IP core for RDMA packet processing. The RoCE IP core is responsible for processing and replying with Ack packets. At the same time, the MAC address encapsulation IP core adds a MAC header to the Ack packets replied by the ROCE IP core. The DMA IP core is used to handle data interaction between the PL and the PS. The address application unit is responsible for the processing of the LWIP protocol, the distributed hash table DHT, the establishment of the RDMA link, and the address mapping during Flash operations. The flash memory chip realizes non-volatile storage of data. The ARM core is responsible for the control and verification functions when reading and writing flash memory data.
2. The memory and computing separated storage system according to claim 1, wherein: It also includes a cache module, which is used to cache read and write data.
3. A method for improving the network transmission efficiency of a memory-computation separated storage system, characterized in that, This method is implemented based on the storage system described in Claim 1 or 2, and includes: Integrate the network protocol stack and storage instruction processing into Ethernet-SSD, enabling SSD to have independent network communication capabilities, and adopt a lightweight TCP / IP protocol stack as the network protocol stack of Ethernet-SSD to optimize UDP or TCP communication; Integrate a RoCE IP core and a DMA module for implementing remote direct memory access in Ethernet-SSD. The RoCE IP core is responsible for protocol parsing and data packaging, and the processed data is sent to the memory or storage medium through the DMA module; In the case of high-concurrency storage requests, Ethernet-SSD introduces a coroutine-based task scheduling method to manage and schedule multiple tasks. When receiving multiple storage requests, Ethernet-SSD dynamically creates concurrent tasks through a coroutine pool. After the tasks are completed, the relevant resources are automatically recycled.
4. The method for improving the network transmission efficiency of the memory-computation separated storage system according to claim 3, wherein: After integrating the network protocol stack into Ethernet-SSD, perform protocol stack initialization. Ethernet-SSD automatically configures the IP address, subnet mask, and default gateway of the device, completes the initialization of the network adapter, and establishes a network connection with the computing node; after the network connection is established, Ethernet-SSD starts to receive packets from the computing node, parses the packets, extracts the key information therein, and passes it to the storage management module inside Ethernet-SSD. The storage management module performs corresponding storage operations according to the extracted key information.
5. The method for improving the network transmission efficiency of the memory-computation separated storage system according to claim 4, wherein: After receiving the key information, the storage management module dynamically schedules tasks according to the priority of the storage request, the storage area, and the data size; gives priority to processing high-priority requests and delays processing low-priority requests.
6. The method for improving the network transmission efficiency of the memory-computation separated storage system according to claim 4 or 5, characterized in that: The key information includes the storage operation type, the target address, and the data length.
7. The method for improving the network transmission efficiency of the memory-computation separated storage system according to claim 3, wherein: During the data transfer process of remote direct memory access, the Ethernet-SSD and the computing node first establish a link request and confirmation message through network switching to establish a transmission channel for storing data, and then perform data transfer. The process of establishing the transmission channel is as follows: The computing node sends a link establishment request message containing its own QPN and PSN to the storage node; After receiving the link establishment request message, the Ethernet-SSD of the storage node extracts the QPN and PSN of the computing node; The Ethernet-SSD of the storage node extracts the transaction id that identifies the connection from the link establishment request message, fills the transaction id and the QPN and PSN at the storage node end into the reply message, and sends the reply message to the computing node; The computing node extracts the QPN and PSN of the storage node from the reply message for subsequent composition of the RDMA message; The computing node composes an RTU message and sends it to the storage node, indicating that the connection is established and RDMA communication can be performed; After receiving the RTU message, the storage node establishes a connection and performs RDMA communication; The above QPN represents the queue number, and the PSN represents the packet sequence number.
8. The method for improving the network transmission efficiency of the memory-computation separated storage system according to claim 7, wherein: After the data transfer is completed, the Ethernet-SSD generates an ACK confirmation packet and feeds back the transmission status to the computing node to ensure the integrity and reliability of the data transfer; When the computing node does not need RDMA communication, it actively disconnects the connection. The disconnection process is as follows: The computing node initiates a disconnection request message to disconnect the specified QPN connection with the storage node; After receiving the disconnection request message, the storage node assembles a reply message, keeps the transaction id consistent with the disconnection request message, and replies to the computing node; The computing node receives the reply, destroys the QPN, and disconnects the connection.
9. The method for improving the network transmission efficiency of the memory-computation separated storage system according to claim 7 or 8, characterized in that: The messages between the computing node and the storage node are sent through the CM message in the UDP callback function. After the storage node receives the message sent by the computing node, it parses the Attribute ID and Opcode fields in the message to determine the message type and perform corresponding operations.
10. The method for improving the network transmission efficiency of the memory-computation separated storage system according to claim 3, wherein: The task scheduling method based on coroutines realizes the switching between coroutines by saving and restoring the context information of coroutines. The scheduling process includes coroutine creation, status management, context switching, coroutine switching condition checking, and task scheduling.
Citation Information
Patent Citations
Data reading method, data storage method, device and system
CN117311596A
Storage and calculation separation type hardware unloading serial acceleration system based on RoCE
CN118132262A