Data storage system, server, server cluster and data storage method
Patent Information
- Application Number
- PCT/CN2026/080710
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-02-28
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026080710_01102026_PF_FP_ABST
Abstract
Description
Data storage systems, servers, server clusters, and data storage methods Technical Field
[0001] This disclosure relates to the field of data storage, and more specifically, to a data storage system, a server, a server cluster, and a data storage method. Background Technology
[0002] Large model training and inference place dynamic and highly customized demands on the computing power, bandwidth, and capacity of graphics processing units (GPUs). While existing GPUs' high-bandwidth memory (HBM) is high-speed, it is not persistent. If data is lost due to power failure, it will force the training process to save checkpoints frequently, introducing additional latency, affecting the speed of AI training and inference, and making it difficult to fully adapt to high-performance computing scenarios.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This disclosure provides a data storage system, server, server cluster, and data storage method to at least address the technical problem in the related art where data storage systems affect the processing efficiency of large models.
[0005] According to one aspect of the present disclosure, a data storage system is provided, comprising: at least one video memory component and at least one block storage component; the video memory component includes a plurality of video memory units and communicates with a plurality of graphics processors and other video memory components respectively via a graphics card interconnect protocol, wherein the other video memory components are video memory components other than the video memory component in the at least one video memory component, and the video memory units in the at least one video memory component and the video memory units in the plurality of graphics processors constitute a video memory pool, wherein the video memory pool is configured to cache data accessed by the plurality of graphics processors; the block storage component communicates with the plurality of graphics processors via a direct connection protocol and is configured to persistently store the data accessed by the plurality of graphics processors.
[0006] According to one aspect of the present disclosure, a server is provided, including: a plurality of graphics processors and a data storage system of any one of the above embodiments.
[0007] According to one aspect of the present disclosure, a server cluster is provided, including: a plurality of servers described in the above embodiments, wherein the plurality of servers mutually back each other up.
[0008] According to one aspect of the present disclosure, a data storage method is provided, applied to a data storage system of any one of the above embodiments. The method includes: receiving a data access request sent by any one of a plurality of graphics processors; determining a target access object from a video memory pool or at least one block storage component based on the data access request, wherein the video memory pool is composed of video memory units in at least one video memory component and video memory units in a plurality of graphics processors; performing data access on the target access object based on the data access request to obtain a data access result; and sending the data access result to the graphics processor.
[0009] According to another aspect of the embodiments of this disclosure, a computer terminal is also provided, including: a memory storing an executable program; and a processor configured to run the program, wherein the program executes the methods in the various embodiments of this disclosure when it runs.
[0010] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is executed, it controls the device where the computer-readable storage medium is located to perform the methods of the various embodiments of the present disclosure.
[0011] According to another aspect of the embodiments of this disclosure, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this disclosure.
[0012] According to another aspect of the embodiments of this disclosure, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in various embodiments of this disclosure.
[0013] According to another aspect of the embodiments of this disclosure, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this disclosure.
[0014] In this embodiment, the system may include at least one video memory component and at least one block storage component. The video memory component comprises multiple video memory units and communicates with multiple graphics processors and other video memory components via a graphics interconnect protocol. The other video memory components are those other than the video memory component itself. The video memory units in the at least one video memory component and the video memory units in the multiple graphics processors constitute a video memory pool, which is configured to cache data accessed by the multiple graphics processors. The block storage component communicates with the multiple graphics processors via a direct connection protocol and is configured to persistently store the data accessed by the multiple graphics processors. It is noteworthy that the video memory component comprises multiple video memory units and communicates with multiple graphics processors and other video memory components via a graphics interconnect protocol. This design allows video memory units to be shared by multiple GPUs, thereby decoupling the fixed ratio of GPU computing power to video memory capacity and bandwidth, achieving the goal of flexible on-demand configuration of video memory resources. The communication links between the video memory units and multiple GPUs established via the graphics interconnect protocol constitute the video memory pool. The concept of a memory pool further expands the access range of GPUs, enabling multiple GPUs to access shared memory resources in parallel. This not only increases memory capacity but also improves memory bandwidth, thereby accelerating the training and inference process of large models. Through the design of memory and block storage components, the efficiency of large model training and inference can be significantly improved, solving the problem of data storage system bottlenecks limiting the processing efficiency of large intelligent computing models in related technologies. Furthermore, it improves data access speed and ensures persistent data storage. Thus, it solves the technical problem in related technologies where data storage systems affect the processing efficiency of large models.
[0015] It is worth noting that the above general description and the following detailed description are merely for illustrative and explanatory purposes and do not constitute a limitation thereof. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:
[0017] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data storage method according to an embodiment of the present disclosure;
[0018] Figure 2 is a structural block diagram of a computing environment according to an embodiment of the present disclosure;
[0019] Figure 3 is a structural block diagram of a service mesh according to an embodiment of the present disclosure;
[0020] Figure 4 is a schematic diagram of a data storage system according to an embodiment of the present disclosure;
[0021] Figure 5 is a schematic diagram of the internal architecture of a high-bandwidth video memory component according to an embodiment of the present disclosure.
[0022] Figure 6 is a schematic diagram of the storage of volatile tag data and persistent tag data in a video memory expansion component according to an embodiment of the present disclosure;
[0023] Figure 7 is a schematic diagram of the architecture of a high-bandwidth persistent memory controller according to an embodiment of the present disclosure;
[0024] Figure 8 is a schematic diagram of a low-latency non-volatile block storage design for a direct-connect graphics card according to an embodiment of the present disclosure;
[0025] Figure 9 is a schematic diagram of a low-latency non-volatile block storage used in direct connection with multiple graphics processors according to an embodiment of the present disclosure;
[0026] Figure 10 is a schematic diagram of a direct-connect graphics card memory expansion and storage device expansion according to an embodiment of the present disclosure;
[0027] Figure 11 is a schematic diagram of a distributed storage system constructed by a server cluster with direct-connected block storage pooling according to an embodiment of the present disclosure.
[0028] Figure 12 is a flowchart of a data storage method according to an embodiment of the present disclosure;
[0029] Figure 13 is a schematic diagram of a data storage device according to an embodiment of the present disclosure;
[0030] Figure 14 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] First, some nouns or terms that appear in the description of the embodiments of this disclosure shall be interpreted as follows:
[0034] High-bandwidth memory (HBM) is an advanced computer storage technology designed to provide higher bandwidth and lower power consumption than traditional DDR (Double Data Rate) memory. HBM achieves this by stacking memory chips and using a vertical channel (rather than a traditional planar layout), thereby reducing signal latency and increasing data transfer rates. In the GPU and high-performance computing fields, HBM is configured to accelerate data-intensive applications such as machine learning and graphics rendering because it provides faster data read and write speeds, meeting the real-time processing needs of these applications.
[0035] Persistent Memory (PMEM) is a novel memory technology that retains data even after power interruption, similar to disk storage, but with access speeds approaching those of Dynamic Random Access Memory (DRAM). This technology combines the advantages of storage and memory, providing an ideal choice for applications requiring data persistence and memory speed. In GPU servers and large-scale computing clusters, PMEM can serve as a persistent extension of GPU memory, providing fast access and data persistence guarantees for model checkpoints, configuration information, and more.
[0036] Unified Memory (UM) is a memory architecture that allows GPUs and CPUs to share the same memory space without explicitly copying data between CPU memory and GPU memory. This simplifies the programming model and improves data access efficiency because data can be directly shared between the GPU and CPU, reducing the overhead of data movement. In deep learning and high-performance computing, UM can significantly improve data processing speed and the overall performance of applications.
[0037] Currently, applications such as large-scale model training and inference have restructured the use of computing, interconnect, and storage infrastructure, significantly raising performance requirements for capacity, bandwidth, latency, and power consumption. GPU manufacturers have designed and launched cutting-edge products, providing GPU components with improved computing power, bandwidth, and capacity in fixed proportions. However, different models and scenarios have varying demands for computing power, bandwidth, and capacity; the fixed proportions of these capabilities are difficult to decouple, leading to situations where other capabilities are over-expanded to meet certain requirements. On the other hand, although GPUs are equipped with high-bandwidth video memory, this memory still lacks non-volatility. Without power, data in the video memory disappears, making it impossible to maintain a complete and timely state. To avoid forced recalculations due to failures during training, current training requires periodic pauses, with current data written to non-volatile storage in a checkpoint manner, thus affecting model training. Decoupling video memory capacity and bandwidth from computing power and configuring them on demand, while providing high-bandwidth, low-latency non-volatile storage, can help accelerate large-scale model training and inference.
[0038] The current solution pools multiple GPUs via a graphics interconnect protocol, thus forming a memory pool within each GPU. Furthermore, by running GPU-specific software on the CPU and GPU, the server's system memory—such as Dynamic Random Access Memory (DRAM) and a standard form of memory module (DIMM)—can be presented to the GPU in a unified memory mode for read and write access. Internally, the server uses a low-level link provided by the Peripheral Component Interconnect Express Switch (PCIe switch), with the GPU directly running the graphics interconnect protocol. The local solid-state drive (SSD) also connects to the CPU and GPU via PCIe, running the Non-Volatile Memory Host Controller Interface (NVMe) protocol. The SSD provides non-volatile storage within the server, ensuring data retention even without power.
[0039] Currently, GPUs are expensive, and to obtain larger capacity, high-bandwidth video memory, the number of GPUs needs to be increased, especially during large model training and inference, where there are significant demands on video memory capacity and bandwidth. Therefore, video memory pooling through GPU interconnect protocols can increase capacity and improve bandwidth, but at a significantly higher cost. While system memory mapped to the GPU through unified memory can increase the required video memory capacity, its bandwidth is lower than the GPU's video memory bandwidth performance, resulting in an imbalance in global video memory bandwidth performance, thus slowing down the progress of artificial intelligence (AI) computing. Compared to computation, storage implemented by Non-Volatile Memory Express Solid State Drives (NVMe SSDs) has significantly lower performance in terms of bandwidth and latency than video memory and even system memory. Therefore, loading and persistently writing data become bottlenecks for AI training and inference. The latency of SSD media and the heavy software stack of NVMe are both difficult to fully adapt to the needs of high-performance AI computing scenarios.
[0040] According to embodiments of this disclosure, a data storage method is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0041] The method embodiments provided in this disclosure can be executed in a mobile terminal, computer terminal, or similar computing device. FIG1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data storage method according to an embodiment of this disclosure. As shown in FIG1, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 configured to store data, and a transmission device 106 configured for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. It will be understood by those skilled in the art that the structure shown in FIG1 is only illustrative and does not limit the structure of the above-described electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in FIG1, or have a different configuration than shown in FIG1.
[0042] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in embodiments of this disclosure, the data processing circuitry serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0043] The memory 104 can be configured to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method in the embodiments of this disclosure. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the method in the above embodiments. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0044] The transmission device 106 is configured to receive or transmit data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module configured to communicate wirelessly with the Internet.
[0045] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0046] The hardware structure block diagram shown in Figure 1 can serve as an exemplary block diagram not only for the aforementioned computer terminal 10 (or mobile device) but also for the aforementioned server. In an optional embodiment, Figure 2 illustrates an example using the computer terminal 10 (or mobile device) shown in Figure 1 as a computing node in the computing environment 201. Figure 2 is a structural block diagram of a computing environment according to an embodiment of the present disclosure. As shown in Figure 2, the computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (shown as 210-1, 210-2, ... in the figure). Each computing node contains local processing and memory resources, and the end user 202 can remotely run applications or store data in the computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 in the computing environment 201, representing services "A", "D", "E", and "H", respectively.
[0047] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).
[0048] The services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.
[0049] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, as shown in Figure 2, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. The proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with similar Pods.
[0050] During operation, executing a user request from end user 202 requires calling one or more services in computing environment 201, and executing one or more functions of one service requires calling one or more functions of another service. As shown in Figure 2, service "A" 220-1 receives the user request from end user 202 from the ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to execute one or more functions.
[0051] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0052] In another alternative embodiment, FIG3 illustrates, in block diagram, an example of using the computer terminal 10 (or mobile device) shown in FIG1 above as a service mesh. FIG3 is a structural block diagram of a service mesh according to an embodiment of the present disclosure. As shown in FIG3, the service mesh 300 is mainly configured to facilitate secure and reliable communication between multiple microservices. Microservices refer to decomposing an application into multiple smaller services or instances and distributing them across different clusters / machines.
[0053] As shown in Figure 3, a microservice may include application service instance A and application service instance B, which together form the functional application layer of service mesh 300. In one implementation, application service instance A runs as a container / process 308 on machine / workload container group 314 (Pod), and application service instance B runs as a container / process 310 on machine / workload container group 316 (Pod).
[0054] In one implementation, application service instance A can be a product query service, and application service instance B can be a product order placement service.
[0055] As shown in Figure 3, application service instance A and grid proxy (sidecar) 303 coexist in machine / workload container group 314, and application service instance B and grid proxy 305 coexist in machine / workload container group 316. Grid proxy 303 and grid proxy 305 form the data plane layer of service mesh 300. Grid proxy 303 and grid proxy 305 run as containers / processes 304 and 306 respectively, and can receive requests 312 for product query services. Grid proxy 303 and application service instance A can communicate bidirectionally, and grid proxy 305 and application service instance B can also communicate bidirectionally. Furthermore, grid proxy 303 and grid proxy 305 can also communicate bidirectionally with each other.
[0056] In one implementation, traffic from application service instance A is routed to the appropriate destination via mesh proxy 303, and network traffic from application service instance B is routed to the appropriate destination via mesh proxy 305. It should be noted that the network traffic mentioned here includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), high-performance, general-purpose open-source frameworks (Google Remote Procedure Call, gRPC), and open-source in-memory data structure storage systems (Redis).
[0057] In one implementation, the functionality of the extended data plane layer can be achieved by writing custom filters for the proxy (Envoy) in service mesh 300. The service mesh proxy configuration can enable the service mesh to correctly proxy service traffic, achieving service interoperability and service governance. Mesh proxies 303 and 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0058] As shown in Figure 3, the service mesh 300 also includes a control plane layer. This control plane layer can consist of a set of services running in a dedicated namespace, managed by a managed control plane component 301 within machine / workload container groups (machine / Pods) 302. As shown in Figure 3, the managed control plane component 301 communicates bidirectionally with mesh agents 303 and 305. The managed control plane component 301 is configured to perform control and management functions. For example, it receives telemetry data from mesh agents 303 and 305 and can further aggregate this telemetry data. The managed control plane component 301 can also provide a user-facing Application Programming Interface (API) for these services, facilitating easier manipulation of network behavior and providing configuration data to mesh agents 303 and 305.
[0059] In the above-described operating environment, this disclosure provides a data storage system as shown in FIG4. FIG4 is a schematic diagram of a data storage system according to an embodiment of this disclosure. As shown in FIG4, the data storage system 400 includes: at least one video memory component 402 and at least one block storage component 404. The video memory component includes a plurality of video memory units 4021 and communicates with a plurality of graphics processors 406 and other video memory components 402 respectively through a graphics card interconnect protocol. The other video memory components are video memory components other than the video memory component in the at least one video memory component. The video memory units in the at least one video memory component and the video memory units in the plurality of graphics processors constitute a video memory pool 408. The video memory pool is configured to cache data accessed by the plurality of graphics processors. The block storage component 404 communicates with the plurality of graphics processors 406 through a direct connection protocol and is configured to persistently store the data accessed by the plurality of graphics processors.
[0060] The aforementioned data storage system is a system responsible for data storage, retrieval, and management, designed to provide high-performance, high-reliability storage solutions, especially for application scenarios with high computing demands.
[0061] The aforementioned memory module is a hardware module specifically configured to improve the memory access performance of a graphics processing unit (GPU). It contains multiple memory units, implemented using high-speed, high-bandwidth technologies designed to provide the GPU with the fast data read / write capabilities required by the GPU. The memory module communicates directly with multiple GPUs via a GPU interconnect protocol, allowing the GPUs to directly access the memory units within the memory module without the need for CPU intermediaries, thus reducing data transfer latency and increasing bandwidth. The memory module's role is to pool multiple physical memory units, forming a unified, larger memory pool for multiple GPUs to access. This allows for the flexible allocation of more memory resources to meet the high memory demands of GPUs during large model training or inference, while maintaining high-performance communication and data access.
[0062] The aforementioned block storage component is a storage device that provides persistent data storage capabilities in the form of blocks. Here, a block typically refers to a fixed-size data storage unit. The block storage component communicates with the GPU via a dedicated direct-connect protocol, allowing the GPU to directly read and write data stored on the medium without the need for complex file systems or operating systems as intermediaries. Non-volatile block storage components, such as the low-latency non-volatile block storage in the design, use persistent storage media that retain data even in the event of power failure. Through direct connection to the GPU, it provides low-latency data access, making data persistence operations during GPU training and inference more efficient.
[0063] The aforementioned graphics processing unit (GPU) is a processor specifically configured to accelerate graphics rendering, image processing, and parallel computing tasks. In the fields of AI and deep learning, GPUs have become critical computing resources due to their powerful parallel computing capabilities and high-efficiency data processing capabilities. GPUs contain a large number of computing units, such as stream processors, which can process a large number of matrix operations, convolution operations, etc. in parallel, making them ideal for training deep neural networks or executing complex image processing algorithms.
[0064] In the application, the GPU can not only access its local video memory, but also a video memory pool consisting of multiple video memory components and directly connected block storage components. This expands the GPU's data access capabilities and storage capacity, which is very helpful for processing large-scale datasets and accelerating the training and inference process of AI models.
[0065] The other memory components mentioned above refer to additional memory components in the system besides the memory component (402) described so far. These components also communicate with the GPU via the GPU interconnect protocol and can be incorporated into the memory pool, providing the GPU with unified, larger-capacity memory resources together with the current memory components. This design allows the system to flexibly expand memory capacity according to the GPU's computing needs, avoiding bottlenecks caused by insufficient memory capacity. In the system, the memory components and block storage components communicate directly with the GPU through their respective protocols, enabling the GPU to access data in the memory pool with high bandwidth and efficiently perform persistent data storage, thereby significantly improving the efficiency and reliability of AI computing tasks.
[0066] The aforementioned video memory units are configured as memory units for storing data processed by the graphics processing unit (GPU). They possess high bandwidth characteristics and are configured to accelerate data read and write operations. The aforementioned graphics processing unit is a processor adept at parallel processing and is commonly used in fields such as graphics rendering, scientific computing, and deep learning.
[0067] The aforementioned GPU interconnect protocols allow for high-speed communication between the GPU and memory units, configured to build high-bandwidth communication networks. The direct-connect protocols, on the other hand, are communication protocols designed to directly connect two or more hardware devices (such as processors and storage devices). Their purpose is to reduce intermediate steps in data transmission, thereby reducing latency and improving transmission efficiency. Direct-connect protocols are communication protocols between the GPU and block storage components, allowing the GPU to directly access persistent storage without going through the CPU or complex file systems, thus reducing the length of data transmission paths.
[0068] The aforementioned memory pool is a unified access resource pool composed of multiple memory units, configured to cache data accessed by the GPU, thereby improving data access efficiency and reducing latency.
[0069] As a low-level storage access method, the aforementioned block storage component provides low latency and non-volatile storage characteristics by communicating directly with the computing unit. It is an indispensable component of high-performance computing systems, especially in the field of AI, where it can improve data processing speed and system reliability.
[0070] In one optional embodiment, the data storage system 400 integrates a video memory component 402 and a block storage component 404 to construct a high-performance storage environment, aiming to accelerate the training and inference process of large-scale intelligent computing models. The video memory component 402 communicates with the GPU and other video memory components via a graphics card interconnect protocol to form a video memory pool, which is configured to efficiently cache data; the block storage component 404 communicates with the GPU via a direct connection protocol to provide a low-latency persistent storage solution, ensuring data persistence.
[0071] In the realm of modeling, the training and inference of large AI models place high demands on storage systems for both bandwidth and low latency. Graphics memory (GPU) components consist of multiple memory cells that form a high-speed, point-to-point communication link with the GPU via the GPU interconnect protocol. This design allows the GPU to directly access the memory cells within the GPU, forming a memory pool together with memory cells from other GPUs. This significantly improves data caching efficiency and reduces GPU latency when processing large amounts of data.
[0072] The block storage unit (corresponding to low-latency non-volatile block storage) communicates with the GPU via a specially designed direct-connect protocol, providing persistent storage functionality. Unlike the volatility of GPU memory, the block storage unit uses non-volatile storage media, maintaining data integrity even in the event of power failure. This direct-connect mechanism reduces intermediate steps in the data transmission path, effectively lowering the latency of persistent data storage. This allows the GPU to read and write data faster during training or inference, avoiding performance bottlenecks caused by data storage.
[0073] The data storage system 400 significantly improves the processing speed of large AI models by integrating a memory pool and low-latency non-volatile block storage. The memory pool allows GPUs to dynamically expand their memory capacity without affecting computational efficiency, eliminating the need to increase the number of GPUs, reducing costs, and improving resource utilization. The direct connection protocol of the block storage component addresses the need for data persistence while ensuring low-latency data access. This is crucial for avoiding redundant computations due to failures during training and for quickly loading model parameters during inference.
[0074] For example, taking a scenario involving the training of a large-scale image recognition model, in a traditional storage system, the GPU needs to frequently access storage devices to read training data and save model states, leading to a significant performance bottleneck. However, by implementing a data storage system 400, the GPU can directly communicate at high speed with the memory units in the memory pool, significantly reducing data transmission latency. Simultaneously, for data that needs to be persistently stored, such as model parameters and intermediate training results, the GPU can write directly through low-latency non-volatile block storage components, avoiding the additional latency associated with the CPU and file system, thereby greatly improving the efficiency and reliability of model training. This design not only solves the needs for data caching and persistent storage but also improves the overall data processing speed by reducing intermediate steps in the data access path, providing powerful support for the training and inference of large AI models.
[0075] Figure 5 is a schematic diagram of the internal architecture of a high-bandwidth video memory component according to an embodiment of the present disclosure. This architecture is designed to provide enhanced storage performance for the GPU while supporting persistent data storage. It can communicate directly with the GPU via the GPU interconnect interface, omitting the GPU stream processor (SP) and focusing on data storage and management. The controller plays a core role, managing the data in the cache and determining whether to write the data to high-bandwidth dynamic random access memory (DRAM) via high-bandwidth video memory (HBM) or to high-bandwidth persistent memory (PMEM) via high-bandwidth persistent memory (HBP) based on the data's volatile or persistent tag. Volatile data (such as temporary data during GPU computation) is written to HBM through the HBM controller, while persistent data (data that needs to be stored for a long time) is written to HBP through the HBP interface of the HBP controller. Before being written to persistent storage, the data in the cache undergoes bus scheduling to ensure the priority and order of data transmission. The peripheral circuitry is responsible for providing a stable operating environment, including power supply and various configuration management functions. This design allows the video memory components to provide a larger, high-bandwidth storage pool with persistent data capabilities. This is crucial for accelerating the training and inference processes of AI models, especially when dealing with large amounts of data and model parameters, effectively improving storage performance and data security.
[0076] Furthermore, Figure 5 illustrates the flexibility and scalability of the video memory components. Through direct communication with the GPU and other video memory components, a high-efficiency memory pool is formed, enabling data sharing among different GPUs and enhancing the collaborative capabilities and data processing speed of the GPU computing cluster. The HBP controller in Figure 5 manages persistent memory media and dynamic random access memory. Bus scheduling manages data flow and bandwidth overhead, responsible for prioritizing and scheduling data transmission between different components, ensuring efficient and orderly data flow within the system. The peripheral circuitry is a collective term for multiple circuits configured to implement testing, monitoring, debugging, power supply, configuration loading, and anti-interference auxiliary functions.
[0077] The aforementioned PMEM stores data with persistent storage tags, while DRAM stores data with volatile storage tags. In other words, PMEM provides persistent memory, and DRAM provides volatile memory; both offer the advantage of low write latency.
[0078] The internal architecture of the high-bandwidth video memory component described above is as follows: The front end runs the GPU interconnect protocol through the GPU interconnect interface. This interface and protocol reuse existing GPU designs, while the stream processor computing units in the GPU are not included in the video memory component. The controller can manage both DRAM and PMEM media. Memory data (including tags) from each GPU enters the cache of the video memory component through the GPU interconnect interface. Then, according to the tag (volatile or persistent) of the received write data, the video memory component controller writes the volatile tagged memory data into DRAM through the high-bandwidth memory controller, while writing the persistent tagged memory data into persistent memory through the high-bandwidth persistent memory interface disclosed herein. The HBM controller uses the existing DRAM design. Figure 5 shows that the system removes the GPU stream processing unit, reuses the existing GPU interconnect protocol and interface, expands the DRAM capacity, adds persistent memory, and realizes a video memory expansion component that can be incorporated into the video memory pool, providing a larger capacity high-bandwidth video memory pool.
[0079] Figure 6 is a schematic diagram illustrating the storage of volatile tag data and persistent tag data of a video memory expansion component according to an embodiment of the present disclosure. As shown in Figure 6, in this disclosure, access between high-bandwidth video memory components and between high-bandwidth video memory components and the graphics processing unit (GPU) is facilitated by different high-bandwidth video memory component caches being divided into D cache (volatile cache) and P cache (persistent cache). The GPU, according to its needs, labels its memory data as volatile or persistent type and stores the corresponding fields in volatile memory medium (DRAM) or persistent memory medium (PMEM). High-bandwidth video memory components are connected to the GPU and other high-bandwidth video memory components through the graphics interconnect structure (Fabric) to form DRAM memory pools and PMEM memory pools. Depending on application requirements, multiple high-bandwidth video memory components can be expanded into the memory pool, expanding it into a memory pool accessible to the GPU with high bandwidth (solid path in Figure 6). In addition to access from the GPU, data storage can also be scheduled between multiple high-bandwidth video memory components (dashed path in Figure 6), thereby achieving space reclamation, reducing space fragmentation, and improving video memory utilization. The triggering condition for this step can be a limited contiguous memory space, making it difficult to allocate large pages of memory. Please check information related to memory reclamation. Both volatile and persistent memory are accessed, allocated, and reclaimed via memory. When volatile memory is insufficient, computing applications can use persistent memory as video memory to accelerate GPU training or inference. Additionally, dynamic memory allocation (malloc) is possible in DRAM and PMEM, allowing memory space to be allocated dynamically during program execution.
[0080] The above steps can be used to create a system that includes at least one video memory (VRAM) component and at least one block storage component. The VRAM component comprises multiple memory cells and communicates with multiple graphics processors (GPUs) and other VRAM components via a GPU interconnect protocol. These other VRAM components are the VRAM components other than the VRAM component itself. The memory cells in the at least one VRAM component and the memory cells in the multiple GPUs constitute a VRAM pool, which is configured to cache data accessed by the multiple GPUs. The block storage component communicates with the multiple GPUs via a direct connection protocol and is configured to persistently store the data accessed by the multiple GPUs. It is noteworthy that the VRAM component comprises multiple memory cells and communicates with multiple GPUs and other VRAM components via a GPU interconnect protocol. This design allows multiple GPUs to share memory cells, thereby decoupling the fixed ratio of GPU computing power to memory capacity and bandwidth, achieving the goal of flexible on-demand configuration of VRAM resources. The communication links between the memory cells and multiple GPUs established via the GPU interconnect protocol constitute the VRAM pool. The concept of a memory pool further expands the access range of GPUs, enabling multiple GPUs to access shared memory resources in parallel. This not only increases memory capacity but also improves memory bandwidth, thereby accelerating the training and inference process of large models. Through the design of memory and block storage components, the efficiency of large model training and inference can be significantly improved, solving the problem of data storage system bottlenecks limiting the processing efficiency of large intelligent computing models in related technologies. Furthermore, it improves data access speed and ensures persistent data storage. Thus, it solves the technical problem in related technologies where data storage systems affect the processing efficiency of large models.
[0081] In the embodiments disclosed above, the video memory components include: a graphics interconnect interface, which is connected to multiple graphics processors and other video memory components respectively. The graphics interconnect interface runs a graphics interconnect protocol and is configured to transmit first data accessed by any graphics processor, wherein the first data includes a target tag; a volatile video memory unit, which is configured to store data containing volatile tags; a non-volatile video memory unit, which is configured to store data containing persistent tags; and a controller, which is connected to the graphics interconnect interface, the volatile video memory unit, and the non-volatile video memory unit, and is configured to write the first data to the volatile video memory unit or the non-volatile video memory unit, or read the first data from the volatile video memory unit or the non-volatile video memory unit based on the target tag.
[0082] The aforementioned video memory components communicate with multiple graphics processing units (GPUs) and other video memory components within the system via the GPU interconnect protocol, enabling efficient management and utilization of video memory resources. This GPU interconnect interface is a hardware interface that allows GPUs to exchange data directly at high speed without going through the CPU or main memory. This interface is particularly important in GPU clusters because it significantly reduces data transfer latency and increases bandwidth, making it one of the fundamental components for building high-performance computing systems.
[0083] The aforementioned GPU interconnect protocol is a communication protocol configured for data communication between GPUs. This type of protocol is designed to improve the communication efficiency of GPU clusters and support massively parallel computing. The aforementioned target tag is metadata that identifies data storage requirements and is configured to indicate whether data should be stored in volatile or non-volatile memory units, thereby allowing for appropriate allocation based on data characteristics and usage needs.
[0084] The aforementioned memory components establish communication links with the various GPUs and other memory components within the system via their GPU interconnect interface. The GPU interconnect protocol runs on this interface, enabling the GPU to transmit initial data to the memory components at high speed. The memory components internally contain volatile memory units (DRAM) and non-volatile memory units (PMEM). The controller schedules data to be written to or read from the appropriate type of memory unit based on the target tag in the initial data (indicating whether the data is volatile or persistent).
[0085] The aforementioned volatile memory units differ significantly from non-volatile memory units in function and characteristics, and their use cases and purposes are adapted to the different needs of GPU computing. Volatile memory units, such as DRAM (Dynamic Random Access Memory), are storage media that can store data while powered on, but lose the data when power is off. In GPU servers or high-performance computing systems, these memory units are mainly configured to store data that requires frequent reading and writing and fast access, such as temporary data, intermediate results, and cached data during computation. The characteristic of this type of data is that it is necessary when the GPU is performing tasks, but once the task ends or in the event of a power outage, its value is low or even unnecessary to retain.
[0086] Volatile memory units are configured to store data with volatile tags. These "volatile tags" are metadata added by the GPU to data to indicate its volatile properties. When the GPU generates data that needs to be accessed quickly, it marks this data as volatile and transmits it to the memory units via the GPU interconnect. The memory unit's controller, particularly the Volatile Memory Management Unit (HBM controller), stores this data in the volatile memory units to provide high-speed data access and processing capabilities.
[0087] The aforementioned non-volatile memory units, such as persistent memory (PMEM), are storage media that retain data even after a power outage. The application of such memory units in GPU servers is primarily to store data that needs to be retained after computational interruptions, such as model parameters and training states. This data is crucial during model training or inference and needs to be able to be quickly recovered and continue previous computational tasks even after an unexpected power outage.
[0088] Non-volatile memory units are configured to store data containing persistent tags. These "persistent tags" are also a type of metadata added by the GPU to the data, indicating the necessity of persistent storage. When the GPU generates data that requires persistent storage, it marks this data as persistent and transmits it to the memory unit via the GPU interconnect interface. The memory unit's controller, especially the Non-volatile Memory Management Unit (HBP controller), stores this data in the non-volatile memory units, ensuring that the data is retained even in the event of power failure or other abnormal conditions, thus guaranteeing the continuity of GPU tasks and the integrity of the data.
[0089] In scenarios involving the training of large models, GPUs generate a large number of temporary computation results. These results are frequently read and written during training, and therefore are stored in volatile memory units to ensure high-speed computation. Meanwhile, key model parameters and training states, which need to be retained for a long time during training and can be quickly recovered after a power outage, are marked as persistent data and stored in non-volatile memory units.
[0090] For example, suppose a deep learning network under training suddenly stops running due to a power failure. Temporary computation results in volatile memory units will be lost, but model parameters and training states stored in non-volatile memory units will be retained. Once the server restarts, the GPU can quickly restore to the training state before the power outage by reading the persistent data in the non-volatile memory units and continue training the model without having to start from scratch, greatly improving system reliability and efficiency. Through this design, the GPU server can effectively manage data storage and persistence while ensuring high-performance computing, thus providing an efficient and reliable data storage solution for large-scale AI model training and inference.
[0091] When the GPU generates or needs to access data, it transfers the data to the memory modules via the GPU interconnect interface. The memory module controller checks the data's target tag. If the data has a volatile tag, the controller writes the data to volatile memory units; if the data has a persistent tag, the controller writes the data to non-volatile memory units. When the GPU needs to read data, the controller similarly determines the data storage location based on the target tag and returns the data to the GPU.
[0092] For example, in scenarios involving large model training, the GPU generates a large number of intermediate results and model parameters during computation. Intermediate results, due to frequent access and the fact that they are not lost while powered on, are tagged as volatile and written to volatile memory by the controller. Model parameters, especially those data that need to be retained even after a power outage, are tagged as persistent and written to non-volatile memory by the controller. This way, even after a power outage and restart, model training can resume from the previously saved parameters, reducing the impact of training interruptions.
[0093] By decoupling memory capacity, bandwidth, and computing power, flexible allocation of memory resources can be achieved based on the specific needs of the GPU. This not only accelerates the computational efficiency of the GPU, especially in large model training and inference, but also improves data persistence and overall system reliability by providing high-bandwidth, low-latency non-volatile storage. For example, when processing large-scale neural network training, the GPU can efficiently cache intermediate results in volatile memory units, while model parameters can be stored in non-volatile memory units. This ensures that training can continue from the previous model parameter state after an unexpected power outage, avoiding data loss and the time-consuming process of training repetition.
[0094] For example, suppose an AI server cluster consists of servers equipped with multiple GPUs configured to train complex neural network models. During training, the GPUs generate a large number of intermediate computation results and model parameters that require persistent storage. This data is tagged with volatile or persistent labels by the GPUs during transmission and then transferred to the video memory (VRAM) unit via the GPU interconnect interface and protocol. The VRAM unit's controller, based on the labels, places the intermediate computation results into volatile memory units (such as DRAM), while the model parameters are placed into non-volatile memory units (such as PMEM). This mechanism ensures efficient model training, rapid processing and storage of intermediate results, and prevents the loss of critical model parameters due to power failures, thus improving the overall efficiency of AI training and the system's data security.
[0095] By designing a video memory component and a non-volatile block storage component, this disclosure solves the bottleneck problem of video memory and storage during GPU training and inference, while reducing costs, reducing space fragmentation, and improving video memory utilization, thereby accelerating the training and inference process of large models and providing a new approach for improving AI infrastructure.
[0096] In the above embodiments of this disclosure, the controller includes: a first cache connected to a graphics card interconnect interface and configured to cache first data; a volatile memory management unit connected to both the first cache and the volatile video memory unit and configured to write the first data to the volatile video memory unit or read the first data from the volatile video memory unit when the target tag is a volatile tag; and a non-volatile memory management unit connected to both the first cache and the non-volatile video memory unit and configured to write the first data to the non-volatile video memory unit or read the first data from the non-volatile video memory unit when the target tag is a persistent tag.
[0097] The aforementioned first cache is a high-speed storage unit in the controller configured to temporarily store data, providing initial storage before data is written to or read from the video memory unit. Caching significantly reduces access latency to the video memory unit and improves data transfer efficiency.
[0098] The aforementioned volatile memory management unit typically refers to the HBM controller, which manages data read and write operations on volatile memory units. It schedules data based on target tags to ensure data can be written or read quickly under high bandwidth conditions. Non-volatile memory management unit.
[0099] The aforementioned non-volatile memory management unit refers to the HBP controller, which is responsible for managing the data read and write operations of the non-volatile video memory unit, ensuring that the data can still be persistently stored even when power is off.
[0100] The controller's architecture described above is designed to efficiently process the initial data from the GPU. First, the data enters the controller's first cache via the GPU interconnect interface for initial storage and preparation. Next, based on the data's target label, the data is scheduled to either a volatile memory management unit or a non-volatile memory management unit for further processing.
[0101] In one optional embodiment, after receiving the first data, the first cache determines its storage requirements based on the target tag. If the target tag is a volatile tag, the HBM controller reads the data from the first cache and writes it into a volatile memory unit (such as DRAM). When the GPU needs to read the data, the HBM controller reads the data from the DRAM and returns it to the GPU through the first cache. If the target tag is a persistent tag, the HBP controller reads the data from the first cache and writes it into a non-volatile memory unit (such as PMEM) through the HBP interface. When the GPU needs to read the data, the HBP controller reads the data from the PMEM and returns it to the GPU through the first cache.
[0102] As the brain of the video memory component, the controller effectively manages data storage and retrieval through its internal first cache and two memory management units (HBM controller and HBP controller). The HBM controller and HBP controller focus on handling volatile and persistent data respectively, ensuring that data is accurately stored in the appropriate video memory units according to its characteristics, thereby meeting the needs of the GPU in different computing scenarios.
[0103] For example, in scenarios involving the training of large models, GPUs generate a large number of intermediate computation results and model parameters. Intermediate computation results, due to frequent access, are marked with volatile labels and written to volatile memory units (such as DRAM) via the HBM controller. Model parameters, especially those data that need to be persistently stored, are marked with persistent labels and written to non-volatile memory units (such as PMEM) via the HBP controller. Even in the event of a power outage, model parameters can be stored in PMEM, ensuring the continuity of training and the security of the data.
[0104] By designing the controller architecture to include a first cache, an HBM controller, and an HBP controller, this disclosure enables efficient management of GPU memory resources within a GPU server. This not only reduces data transmission latency and improves data processing speed, but also ensures the stability of model training and persistent data storage by providing non-volatile storage units, thereby accelerating the training and inference of large models, reducing overall costs, and improving resource utilization and computational efficiency.
[0105] For example, in scenarios where GPU server clusters are used for large-scale AI model training, the memory controller plays a crucial role. When the GPU generates a large amount of data during model training, this data is first cached in the first cache. For frequently used intermediate computation results, the HBM controller writes them to DRAM for high-bandwidth access. For model parameters that need to be persistently stored, the HBP controller writes them to PMEM through the HBP interface, ensuring non-volatile data storage. This mechanism greatly improves training efficiency, reduces performance bottlenecks caused by data persistence, and also lowers the overall expansion cost of the server cluster, providing strong technical support for efficient AI model training and inference.
[0106] In the above embodiments of this disclosure, the non-volatile memory management unit includes: a non-volatile memory management controller connected to a first cache, configured to encode first data using short error correction codes to obtain first encoded data, or to decode raw data obtained from the non-volatile video memory unit using short error correction codes to obtain first data; and a non-volatile memory management interface connected to both the non-volatile memory management controller and the non-volatile video memory unit, configured to write the first encoded data into the non-volatile video memory unit and read raw data from the non-volatile video memory unit.
[0107] The aforementioned non-volatile memory management controller is a core component of the video memory unit. It is responsible for processing the initial data containing persistent tags, including short error-correcting code encoding and decoding. Through encoding and decoding, the controller ensures the integrity and reliability of data during storage in the non-volatile video memory unit, allowing data to be correctly recovered and read even in the event of a power outage.
[0108] The aforementioned non-volatile memory management interface acts as a bridge between the non-volatile memory management controller and the non-volatile video memory unit. It is responsible for writing the encoded first data into the non-volatile video memory unit and reading the original data from the non-volatile video memory unit. This interface is the channel for data to enter and exit the non-volatile video memory unit and is also the key to realizing data storage and retrieval operations.
[0109] In one alternative embodiment, when the GPU generates first data that requires persistent storage, this data first reaches the first cache via the GPU interconnect interface and is then transferred to the non-volatile memory management controller. The non-volatile memory management controller encodes the first data using short error-correcting codes to generate first encoded data, enhancing the data's robustness against interference and error correction. The short error-correcting codes protect the integrity and accuracy of the data during storage and transmission. The encoded first data is then written to a non-volatile memory unit (such as persistent memory, PMEM) via the non-volatile memory management interface, achieving persistent data storage. When data needs to be read, the non-volatile memory management interface reads the original data (which may contain errors introduced by storage media aging or power supply issues) from the non-volatile memory unit and transmits the data to the non-volatile memory management controller. The non-volatile memory management controller decodes the original data using short error-correcting codes, detects and corrects any errors, and recovers the original first data. The first data is then returned to the GPU via the first cache, ensuring that the GPU can access accurate, complete, and persistent data.
[0110] For example, during the training and inference of large models, the GPU generates a large number of model parameters. These parameters need to remain constant throughout the training and inference process, and should be able to quickly resume training even after unexpected events such as power outages. At this time, the data is marked with persistent tags and transmitted to the non-volatile memory management unit (NMEM) of the video memory module via the GPU interconnect protocol. Suppose the GPU generates a 64-byte model parameter update during training, and this data is marked with a persistent tag. After receiving the data, the HBP controller encodes it with short error-correcting codes to ensure that even if a small number of errors are introduced during storage, they can be detected and corrected. The encoded data is written to the NMEM unit (such as PMEM) through the HBP interface. When the GPU needs to read these parameters, the HBP controller reads the raw data from the PMEM through the HBP interface, decodes it using short error-correcting codes, recovers the accurate model parameters, and returns them to the GPU through the first cache to continue the training task.
[0111] The introduction of non-volatile memory management units (HBP controller and HBP interface) significantly improves the storage quality and efficiency of persistent data in GPU servers. Through short error-correcting code encoding and decoding, not only is the data's resistance to interference enhanced, ensuring its integrity and accuracy, but the original bit error rate requirements of non-volatile storage media are also reduced, improving its reliability and lifespan. Furthermore, with the help of non-volatile memory units (such as PMEM), the GPU can directly access and update persistent data without going through the CPU or system memory, greatly reducing data transfer latency and improving the efficiency and speed of large model training and inference.
[0112] For example, in a scenario where a GPU server cluster is used for large-scale neural network model training, the GPU generates critical model parameters during training. These parameters not only need to be frequently updated during training but also need to be accurately recovered after accidents such as power outages. The GPU marks these parameters as persistent tags and transmits them to the HBP controller of the memory unit via the GPU interconnect interface. The HBP controller encodes this data with short error-correcting codes, improving the data's error resistance, and writes the data to non-volatile memory units (such as PMEM) for storage via the HBP interface. Even if the power supply is suddenly interrupted, the model parameters can be safely stored in the PMEM. When the server restarts, the GPU reads the data from the PMEM through the HBP controller, which automatically decodes the short error-correcting codes to recover the accurate model parameters. The GPU can then seamlessly continue the previous training task, greatly improving the continuity and efficiency of training and reducing the overall cost of model training. This structural design has significant beneficial effects on improving the storage hierarchy within the GPU server cluster and accelerating the large-scale model training and inference process.
[0113] Through the above design, the GPU server can effectively utilize non-volatile memory units to store critical persistent data, while improving the reliability and integrity of the data through short error correction code technology. Ultimately, it achieves the goal of improving the efficiency and performance of large-scale AI model training and inference while ensuring data persistence and reliability.
[0114] Figure 7 is a schematic diagram of the architecture of a high-bandwidth persistent memory controller according to an embodiment of the present disclosure. Figure 7 shows the architecture of the HBP high-bandwidth persistent memory controller designed in this disclosure. The controller retrieves data with persistent tags from the cache and matches it with a short error correction code based on the data size. For example, if the updated data is 64 bytes, and the user data protected by the short error correction code is 128 bytes, then the HBP controller needs to read the codeword containing the corresponding 64 bytes of short error correction code from the medium, and decode it to obtain 128 error-free bytes. The updated 64 bytes replace the corresponding 64 bytes, and the newly obtained 128 bytes are re-encoded with the short error correction code and written to persistent memory. It should be noted that one ECC code protects 128 bytes, of which 64 bytes are updated, while the remaining 64 bytes remain unchanged. Therefore, it is necessary to obtain these unchanged 64 bytes, combine them with the newly written 64 bytes to form a 128-byte array, and then perform error correction code (ECC encoding) on this new 128 bytes. The raw data read from the medium contains noise and needs to be decoded by ECC to obtain error-free data. Select 64 bytes that are not updated and combine them with the newly arrived 64 bytes (replacing the updated 64 bytes) to form 128 bytes, which are then protected by ECC encoding again.
[0115] Similarly, when reading 64 bytes, the codeword (including noise) of the short error-correcting code containing those 64 bytes is read. After decoding the short error-correcting code, 128 error-free bytes are obtained. The requested 64 bytes are then selected and sent to the cache. Therefore, as shown in Figure 7, the HBP memory controller is designed with a read-modify-write (RMW) unit. Afterwards, data is written to the persistent memory (PMEM) chip via the short error-correcting code encoder, scrambler, and media channels (channel 1, ..., channel N). The read path involves reading the original data via the media channel, restoring the data order by the scrambler, and then restoring the error-free data via the short error-correcting code decoder. A management path exists outside the data channels, primarily implementing logic (RMW logic) to physical (L2P) address mapping, media scheduling, and other functions. These management functions are implemented by multiple microprocessors running firmware, namely microprocessor 1, ..., microprocessor M, etc., communicating via a bus. Peripheral circuitry provides functions such as power supply, testing, diagnostics, configuration loading, and register settings.
[0116] In the above embodiments of this disclosure, the non-volatile memory management controller includes: an access unit configured to modify and write non-volatile data read from a non-volatile video memory unit based on first data to obtain data to be written, or to read first decoded data obtained from a non-volatile video memory unit to obtain first data; a short error correction code codec connected to the access unit, configured to encode the data to be written using short error correction codes to obtain first encoded data, or to decode the original data to obtain first decoded data; and a back-end controller including at least one first medium channel, respectively connected to at least one medium particle in the short error correction code codec and the non-volatile video memory unit, configured to transmit the first encoded data or the original data.
[0117] The access unit is a Read-Modify-Write (RMW) unit, dedicated to handling data reading, modification, and writing operations, and is particularly suitable for scenarios involving partial data updates. When it is necessary to modify part of the data in a non-volatile memory unit (such as PMEM) (for example, updating 64 bytes of data, the value here is only a limitation), the RMW logic reads a 128-byte codeword containing the 64 bytes of updated data from the PMEM, and obtains the error-free 124 bytes of data through ECC decoding. Then, the RMW logic replaces the corresponding 64 bytes in the old data with the new data (64 bytes), while keeping the remaining 64 bytes of data unchanged, thus forming a new data to be written (i.e., a 128-byte data block).
[0118] First, the GPU sends a read command to the non-volatile memory unit to read the data that needs to be modified or the current non-volatile data. Because this data uses non-volatile technology, it remains persistent even when the system is without power. Next, the data read from the non-volatile memory unit is loaded into the cache or the GPU's local storage. At this point, the GPU modifies the read data based on the new data (the first data), for example, updating model parameters, training state, or other computationally relevant data fields. After the data modification is complete, the GPU sends the modified data back to the memory controller, which encodes the data using error correction codes. Error correction codes are a coding technique used to detect and correct errors that occur during storage or transmission, effectively improving data reliability. Finally, the encoded modified data is written back to the non-volatile memory unit. This step ensures that even in the event of a sudden power outage or other abnormal situation, the latest state of the data is persistently saved and not lost.
[0119] By introducing non-volatile memory units and error-correcting code encoding, this approach aims to address the efficiency issues of data updates and persistent storage in traditional GPU computing architectures. In applications such as deep learning, which have high requirements for data persistence, frequent read and write operations accelerate storage device wear, while data loss or corruption severely impacts the continuity of model training and the accuracy of results. By combining data modification operations with non-volatile memory units, this approach not only reduces reliance on and pressure on main memory but also enhances data robustness and reliability through error-correcting code encoding, thereby improving the efficiency and security of the entire GPU server cluster in handling large-scale model training.
[0120] For example, in a large-scale language model training scenario, the GPU is updating the model's weight matrix. According to the process described in the claims, the GPU first reads the current weight matrix (non-volatile data) from non-volatile memory, and then modifies the weight matrix based on the latest training results (first data). Before writing the modified data back to the non-volatile memory, it is encoded with error correction codes by the memory controller to ensure the integrity and accuracy of the data during storage. Even if a system failure or power interruption occurs during training, the GPU server can recover the modified weight matrix from the non-volatile memory and continue training without starting from scratch, which greatly improves the speed and efficiency of model training. Simultaneously, the use of error correction codes ensures that even if the non-volatile storage medium is slightly damaged, the GPU can recover correct data from other redundant storage units, ensuring the continuity of the training process and data security.
[0121] The aforementioned read-modify-write combined error-correcting code encoding process provides crucial data persistence and efficiency support for GPU server clusters when handling large-scale model training and inference tasks, overcoming the limitations of traditional storage technologies in data persistence and storage efficiency. By utilizing non-volatile memory units and error-correcting codes, data reliability and storage speed are improved, providing a stable data foundation for GPU computing.
[0122] The aforementioned short error-correcting code encoder / decoder is closely connected to the access unit and is configured to perform ECC encoding and decoding of data. Before data is written to a non-volatile memory unit (such as PMEM), the short error-correcting code encoder / decoder encodes the data to generate ECC codewords containing redundant information to ensure data integrity and reliability. When reading data, the short error-correcting code encoder / decoder decodes the codewords read from the PMEM, corrects potential errors, and recovers accurate data.
[0123] The aforementioned back-end controller includes at least one first media channel connected to the media particles in the short error-correcting code codec and non-volatile memory units (such as PMEM). The media channel is responsible for transmitting encoded data (first encoded data) or raw data to the PMEM media particles, and for transmitting data from the PMEM media particles to the codec for decoding. The back-end controller provides a direct physical link between the data and the PMEM media particles, enabling fast data access.
[0124] The aforementioned first decoded data refers to the data processed by the short error-correcting code decoder. It is usually the raw data read from a non-volatile memory unit (such as PMEM), which may contain errors introduced by external factors such as storage medium characteristics or power interruptions. The decoding process can detect and correct these errors, thereby recovering the original accurate data, i.e., the "first decoded data".
[0125] Conversely, the first encoded data mentioned above refers to the data generated after adding redundant error correction information through a short error correction code encoder before the data is written to the non-volatile memory unit. The encoding process aims to improve the reliability and durability of data during storage, so that even if the storage medium suffers some degree of damage, the original data can be recovered through the redundant information.
[0126] The aforementioned media particles are the basic storage units that constitute non-volatile memory cells (such as PMEM), responsible for storing the actual data. Storage devices typically contain multiple media particles, each capable of independent read and write operations. The independence and addressability of the media particles are crucial foundations for achieving efficient data storage and management.
[0127] For example, during the training of a large model on a GPU server cluster, the GPU generates data that requires persistent storage. This data is first transferred to the first cache and then enters the non-volatile memory management controller. In the RMW logic, if the data to be updated is smaller than the ECC codeword size (e.g., 64 bytes less than 128 bytes), the RMW logic reads the 128-byte codeword containing the updated data from the PMEM and obtains the error-free 128-byte data through the short error-correcting code decoder. Subsequently, the RMW logic replaces a portion of the old data with the new data, while leaving the unupdated portion unchanged, forming a new data to be written. Next, the short error-correcting code encoder / decoder encodes this data to be written, generating a new ECC codeword. Finally, through the backend controller and the first media channel, this new ECC codeword is quickly written to the PMEM media particles, completing the data update and persistent storage.
[0128] By organically integrating the Access Management Unit (RMW) logic, short error-correcting code codec, and back-end controller, the non-volatile memory management controller can efficiently handle data updates and persistent storage in GPU servers. The RMW logic not only reduces unnecessary data reads and writes, lowering latency and energy consumption in data operations, but also ensures the atomicity of data updates. The addition of the short error-correcting code codec further improves data reliability and durability; even if data in the media particles is disturbed due to storage characteristics or power outages, errors can be detected and corrected using ECC technology.
[0129] For example, in a scenario where a GPU server cluster is used for large-scale machine learning model training, the GPU generates 64-byte model parameter updates during training. These parameters need to be persistently stored in non-volatile memory units (such as PMEM). The GPU marks this data as persistent tags and transmits it to the HBP controller of the memory unit via the GPU interconnect interface. In the HBP controller, the RMW logic first reads a 128-byte codeword containing the 64-byte update data from the PMEM, and obtains error-free 128-byte data through a short error correction code decoder. The updated data replaces a portion of the old data, while the rest remains unchanged. Subsequently, the short error correction code decoder encodes the updated new data to generate a new ECC codeword. This new ECC codeword is quickly written to the PMEM media particles through the back-end controller and the first media channel, achieving persistent storage of model parameter updates. Even after a power outage, the GPU can read these model parameters from the PMEM, ensure data accuracy through ECC decoding, and thus seamlessly resume model training, improving training efficiency and continuity.
[0130] The above-mentioned structural design has significant benefits for persistent data management within GPU server clusters. It not only accelerates the data persistence process and reduces energy consumption and latency, but also improves data persistence and reliability, providing a solid data storage foundation for large-scale AI model training and inference.
[0131] In the above embodiments of this disclosure, the block storage component includes: a graphics interconnect protocol controller connected to a plurality of graphics processors and configured to transmit second data accessed by any one of the graphics processors; a second cache connected to the graphics interconnect protocol controller and configured to cache the second data; an error correction code encoding / decoding unit connected to the second cache and configured to encode the second data with error correction codes to obtain second encoded data, or to decode persistent data obtained from at least one block storage medium with error correction codes to obtain second data; and at least one second medium channel connected to the error correction code encoding / decoding unit and at least one block storage medium respectively, and configured to transmit the second encoded data or persistent data.
[0132] The aforementioned GPU interconnect protocol controller is a control unit configured to manage the transmission of data between the GPU and block storage components, ensuring that data can communicate efficiently in accordance with the requirements of the GPU interconnect protocol.
[0133] The aforementioned second cache is a storage unit configured to temporarily store data to be persisted received from the GPU or data read from block storage media. It serves as a data cache and exchange unit, reducing the number of times data is directly accessed to persistent storage, thereby improving data access speed and reducing latency.
[0134] The aforementioned error correction code encoding and decoding unit is a processing unit responsible for performing error correction code (ECC) encoding before data is written to persistent storage, and decoding after data is read from persistent storage, so as to ensure the integrity and reliability of data during the storage process, and to correctly recover data even if the medium is slightly damaged.
[0135] The aforementioned second medium channel is a physical or logical connection configured for data transmission between the cache, error correction code encoding / decoding unit, and persistent storage medium, ensuring high-speed and low-latency data transmission.
[0136] The aforementioned block storage medium is the physical carrier for data storage, providing persistent storage functionality. Even after a power failure, data can be preserved. It forms the basis of block storage components and is configured to store persistent data.
[0137] First, the GPU interconnect protocol controller, acting as an intermediate layer, establishes a connection with the GPU and receives data access requests from the GPU. These requests pertain to writing or reading persistent data. Next, the received data or persistent data read requests are temporarily stored in a second cache. This cache provides a buffer for temporary data storage, ensuring that data can be processed through a faster storage unit when being written or read. Then, when data needs to be written to persistent storage, the error correction code encoding / decoding unit encodes it to improve data fault tolerance; when data is read from persistent storage, it is first decoded to restore the original state of the data, ensuring data accuracy and security. Finally, the encoded data is transmitted to the block storage medium for persistent storage through the second media channel, or the decoded data read from the medium is transmitted back to the GPU through the same channel, completing the persistent storage or retrieval process.
[0138] For example, by directly connecting the block storage component to the GPU and using a GPU interconnect protocol controller for data transmission management, not only is the data transmission bandwidth increased, but CPU intervention is reduced, thus lowering read / write latency. The second cache further improves data processing efficiency, avoiding frequent direct access to slower persistent storage. The error correction code encoding / decoding unit ensures data integrity and reliability; even if minor physical damage occurs on the persistent storage medium, data can be recovered through error correction, ensuring the continuity of the training process and data security.
[0139] Through the above design, the GPU can directly access low-latency block storage components, enabling fast read and write of persistent data. This not only resolves the conflict between data persistence and access speed in traditional storage solutions, but also reduces data access latency and improves data transmission efficiency and storage reliability through the collaborative work of the GPU interconnect protocol controller and error correction code encoding / decoding unit. This effectively enhances the computational efficiency and data processing capabilities of GPU server clusters in large model training and inference scenarios.
[0140] For example, in large model training scenarios, after each training round, the GPU needs to persistently store key model parameters to prevent loss of training progress in case of system failure or power outage. Using the solution disclosed herein, the GPU can directly write data to the block storage component. The data is cached in a second cache, encoded by an error correction code encoder / decoder unit, and efficiently transmitted to the block storage medium for persistent storage via a second media channel. When this persistent data needs to be read, it is returned via the same path, decoded to obtain the original data, and directly used for GPU computational tasks without CPU intervention. This significantly accelerates data read / write speeds, reduces the risk of training interruption, and improves training efficiency and data security. This design is particularly important in modern high-performance computing and AI model training because it can significantly improve GPU computational efficiency while ensuring data persistence and security.
[0141] In the above embodiments of this disclosure, the block storage component further includes: a compression unit connected to the second cache, configured to decompress the second data to obtain decompressed data, or to compress decrypted data obtained from the at least one block storage medium to obtain the second data; an encryption / decryption unit connected to the compression unit, configured to encrypt the decompressed data to obtain encrypted data, or to decrypt demodulated data obtained from the at least one block storage medium to obtain the decrypted data; and a modulation / demodulation unit connected to the encryption / decryption unit, configured to modulate the encrypted data to obtain modulated data, or to demodulate the second decoded data obtained from the at least one block storage medium to obtain the demodulated data; wherein the error correction code encoding / decoding unit is connected to the modulation / demodulation unit, configured to encode the modulated data with error correction codes to obtain the second encoded data, or to decode the persistent data with error correction codes to obtain the second decoded data.
[0142] The aforementioned compression unit is a processing component for data compression and decompression, which aims to reduce the amount of data transmitted and improve storage efficiency and transmission speed. In this disclosure, it is connected to a second buffer to compress or decompress the data in the buffer to accommodate subsequent encryption, modulation, and other processing steps.
[0143] The encryption / decryption unit mentioned above is a data security processing module, configured to encrypt or decrypt data to ensure data security and privacy protection during transmission and storage.
[0144] The aforementioned modulation and demodulation unit is a signal processing module that modulates or demodulates data to improve data transmission efficiency and anti-interference capability. In particular, modulation before error correction code encoding can better utilize the protection capability of error correction codes, thereby improving data reliability and transmission speed.
[0145] In the block storage component design of this disclosure, data undergoes multiple processing steps during transmission and storage to improve efficiency, security, and reliability. Data received from the GPU is first compressed by a compression unit, reducing the storage and transmission volume and improving storage space utilization and data transmission speed. The compressed data is then sent to an encryption / decryption unit for encryption, ensuring data security during storage and transmission and preventing data leakage or unauthorized access. The encrypted data is then modulated by a modulation / demodulation unit, improving the data transmission format on the physical medium and effectively enhancing transmission efficiency and anti-interference capabilities. The modulated data is then passed to an error correction code encoding / decoding unit for error correction code encoding, enhancing data robustness. Even if minor damage occurs during storage or transmission, data integrity and accuracy can be restored through error correction. The encoded data is finally written to the block storage medium for persistent storage, while data read from the block storage medium undergoes the reverse process of decoding, demodulation, decryption, and decompression to ensure correct reading and use by the GPU.
[0146] The GPU interconnect protocol controller receives data from the GPU. The data is first cached, and in a second cache, a compression unit compresses the data to reduce its size for more efficient subsequent processing. The compressed data is then encrypted by an encryption / decryption unit to ensure security and privacy. The encrypted data is modulated by a modem unit to improve the format for transmission over the physical medium, and then encoded with error correction codes by an error correction code encoding / decoding unit to enhance transmission reliability and interference resistance. Finally, the data is written to the block storage medium for persistent storage via a second media channel. Every step in this process aims to improve the efficiency of data transmission and storage while ensuring data security and integrity.
[0147] By designing a block storage component that includes a compression unit, an encryption / decryption unit, a modulation / demodulation unit, and an error correction code encoding / decoding unit, embodiments of this disclosure can significantly improve data storage and processing efficiency in GPU server clusters. Data compression reduces the volume of transmission and storage, improving storage space utilization; data encryption ensures data security and prevents unauthorized access; data modulation improves the transmission format and increases transmission efficiency; and error correction code encoding enhances data reliability and robustness. The combined use of these steps not only improves data access speed and reduces latency but also enhances persistent data storage capabilities, ensuring the efficiency of large model training and inference processes and data security.
[0148] For example, during the training of large models, GPUs need to frequently read and write large amounts of data, including model parameters, training data, and checkpoints. When the GPU writes data to the block storage component, the data is first compressed by the compression unit to reduce storage space usage; then encrypted by the encryption / decryption unit to ensure data security; next, the modulation / demodulation unit improves the transmission format to increase transmission efficiency; and finally, it is encoded by error correction codes to enhance the data's anti-interference capability. When reading data, the above process occurs in reverse order. After the data is read from the block storage medium, it undergoes error correction code decoding, demodulation, decryption, and decompression, and is finally provided to the GPU in its original format for computation. This series of processing steps not only ensures the secure and efficient transmission and storage of data, but also significantly improves the performance and stability of GPU server clusters when handling large-scale AI model training and inference tasks by reducing the amount of data transmitted and enhancing data robustness. For example, when a server cluster is processing a deep learning model for image recognition, the large amount of data generated by the GPU is compressed, encrypted, modulated, and encoded before being stored. When this data needs to be read for the next round of training, the data undergoes the above reverse process to accurately provide it to the GPU, ensuring the continuity and efficiency of the training process.
[0149] The block storage component of this design establishes a direct communication link with the GPU via a GPU interconnect protocol controller, thereby accelerating data transmission. Data generated by the GPU (second data) is first transmitted to a second cache via the GPU interconnect protocol controller. Subsequently, the data is compressed by a compression unit to improve transmission efficiency and storage space utilization. The compressed data is then encrypted by an encryption / decryption unit to ensure data security. The encrypted data is then modulated by a modulation / demodulation unit to adapt to the transmission characteristics of the physical link, and subsequently transmitted to the block storage medium for storage via at least one second media channel. Before being written to the block storage medium, an error correction code encoding / decoding unit performs ECC encoding on the data to increase redundancy and improve storage reliability and durability.
[0150] The reading process is the reverse. Data read from the block storage medium is first transmitted through the second media channel to the error correction code encoding / decoding unit for ECC decoding to detect and correct potential errors. The data is then demodulated by the modulation / demodulation unit to restore the original data format. Next, it is decrypted by the encryption / decryption unit to expose the original data to the backend processing of the decryption unit. The decrypted data is then decompressed by the compression unit to restore the original second data, and finally transmitted to the GPU for its use via the GPU interconnect protocol controller.
[0151] For example, during large-scale AI model training, GPUs need to persistently store large amounts of training data in block storage devices. This data requires multiple steps, including compression, encryption, and ECC encoding, to ensure transmission efficiency and data security. The GPU transmits the second data to the GPU interconnect protocol controller in the block storage component via the GPU interconnect protocol. The data is then cached in a second cache and compressed by a compression unit to reduce the data size and improve bandwidth utilization. The compressed data is then encrypted by an encryption / decryption unit to ensure data privacy. Subsequently, the data is modulated by a modulation / demodulation unit to adapt to the transmission requirements of the physical link, and then ECC encoded by an error correction code encoding / decoding unit to increase data redundancy and improve persistence and reliability in non-volatile storage media. Finally, the second encoded data is written to the block storage medium through a second media channel.
[0152] When the GPU needs to access this persistent data, the data is transferred from the block storage medium to the error correction code encoding / decoding unit via the second media channel for ECC decoding to detect and correct potential errors. Subsequently, the data is demodulated by the modulation / demodulation unit to restore the original data format, and then decrypted by the encryption / decryption unit to expose the original data. The decrypted data is decompressed by the compression unit to restore the initial second data, and finally transmitted to the GPU via the GPU interconnect protocol controller for subsequent computational operations.
[0153] Direct GPU connection significantly reduces data transmission latency, greatly improving the storage performance of GPU servers when handling large-scale AI tasks. Simultaneously, by integrating compression, encryption, modulation, and ECC encoding, the block storage component ensures data security, integrity, and persistence, providing a highly efficient and reliable data storage solution for GPU server clusters. This helps accelerate large-scale model training and inference processes, enhancing overall system performance.
[0154] Figure 8 is a schematic diagram of a low-latency non-volatile block storage design for direct connection to a graphics card according to an embodiment of the present disclosure. As shown in Figure 8, the block storage component design for direct connection to a graphics card in this disclosure is illustrated, wherein the persistent medium is the same as that in the video memory component. It should be noted that block storage is different from memory; block storage uses 4KB size access, while memory uses 64B size access. Video memory (a type of memory) and block storage are two modules with different usage methods and each has its own purpose in server architecture. Block storage devices also have the advantage of low read / write latency, but their access method is block storage mode, not memory interface semantics.
[0155] The block storage component in Figure 8 may include a host interface connector, using a PCIe communication module as the physical link, and running a GPU interconnect protocol controller to achieve low-latency data transmission. The multi-core processor subsystem (multi-core CPU subsystem) runs firmware to implement transmission protocols, cache management, and media management (including address mapping). To further improve data transmission bandwidth, data written to and read from the storage device by the GPU is compressed to transmit more data within the same link bandwidth. After the stored data enters the cache, it is first decompressed to restore the data. Then, an encryption unit protects data privacy, ensuring that data leakage does not occur if the storage device is accidentally disassembled. The modulation module then reconstructs the data order and uses Cooperative Error Correction Code (ECC) encoding to improve noise immunity. Because block storage is mostly much longer than the memory access unit size, ECC can use high code length error correction codes to provide stronger error correction capabilities. This objectively reduces the requirement for the raw bit error rate (RBER) of the persistent media. The persistent media management module allocates the physical address for writing data segments and writes them to the persistent media via the media access channel. The reading process involves the media management module addressing the current physical address of the requested data and reading the raw data (including noise) via the media channel. Then, the error correction code is decoded to restore the noise-free, correct data, which is then demodulated to restore the data format. Next, the noise-free data is decrypted to obtain the actual data previously written by the GPU, compressed again to increase transmission bandwidth, and then placed in the data cache. The GPU receives the compressed data, decompresses it, and then uses it for computation.
[0156] In the above embodiments of this disclosure, the block storage component includes: multiple namespaces corresponding to multiple graphics processors, configured to divide the storage space of the block storage component.
[0157] The namespace mentioned above is a mechanism for logically isolating and organizing storage resources in a storage system. It allows different users, applications, or devices to have independent storage spaces, even when sharing physical storage media. Multiple namespaces can correspond to multiple graphics processing units (GPUs), providing each GPU with an independent storage area to achieve efficient storage management and access control.
[0158] The block storage component in this disclosure is designed with multiple namespaces to accommodate the concurrent access requirements of multiple GPUs in a GPU server cluster. Each namespace is directly associated with a GPU, allowing the GPU to access and manage its corresponding storage space independently. This design aims to avoid data access conflicts between different GPUs, improve the utilization efficiency of storage resources, and simplify storage space management and access control.
[0159] For example, in a scenario where GPU servers are used for deep learning model training, each GPU can independently access and store model parameters, intermediate results, or training data. By defining multiple namespaces within the block storage component, each GPU has its own dedicated storage area, allowing for independent data read and write operations. Assuming a GPU cluster has three GPUs, each assigned a different namespace: namespace1, namespace2, and namespace3. When GPU1 needs to store training data, it writes the data to namespace1; similarly, GPU2 and GPU3 write their data to namespace2 and namespace3 respectively. This avoids data access conflicts between GPUs, ensures data isolation and independent management, and improves the storage performance of the GPU server cluster during AI model training.
[0160] Introducing a namespace mechanism into GPU server clusters can significantly improve the overall efficiency and stability of the storage system. The use of multiple namespaces not only provides logical isolation of data but also allows each GPU to access its dedicated storage space efficiently and with low latency, reducing wait times and contention for storage resources. Furthermore, the independence of namespaces facilitates flexible allocation and management of storage space, enabling GPU server clusters to better adapt to multi-task concurrent processing scenarios and accelerate the training and inference processes of large-scale AI models.
[0161] For example, in a large-scale AI model training scenario, suppose there is a server cluster consisting of four GPUs, each responsible for a different model training task, requiring independent reading and writing of training data and model parameters. Through multiple namespaces in the block storage component, we can allocate an independent storage area for each GPU, such as namespace1 to namespace4. When GPU1 processes its task and needs to store training results, it can directly write the data to namespace1 without considering data access conflicts with other GPUs. Similarly, GPU2, GPU3, and GPU4 also perform read and write operations in their respective namespaces. Because each GPU has its own independent namespace, data access becomes more efficient and faster, while ensuring data integrity and security. This promotes efficient storage and management of the GPU server cluster in AI model training, significantly improving the system's concurrent processing capabilities and overall performance.
[0162] Figure 9 is a schematic diagram of a low-latency non-volatile block storage system used in direct connection with multiple graphics processors according to an embodiment of the present disclosure. As shown in Figure 9, a low-latency non-volatile block storage system can present multiple namespaces (NSs) to the outside world. The graphics processor (GPU) achieves storage space separation by accessing each namespace. A namespace is a logical-level block device that presents a physical block device as multiple independent logical block devices, which can isolate the read and write access of each NS. The GPU can read and write NSs through storage access protocols. The GPU can directly read or write namespaces via PCIe links. A data channel is formed by direct connection between the GPU and the non-volatile block storage. The CPU and system memory implement the management path, realizing functions such as file system, block layer, driver, metadata management, control, and recording. The GPU generates and distributes the data to be written or initiates read requests. The data path is direct between the non-volatile block storage and the GPU, while the control path is implemented by the processor (CPU) and system memory. In practice, to reduce latency related to management paths, GPUs can directly manipulate block storage, establishing and maintaining simple and easy-to-use data addressing to accelerate read and write operations. For example, a GPU can treat the entire block storage as a one-dimensional array and write data segments sequentially. The addressing metadata of each data segment is marked according to its starting address and duration within the one-dimensional array. While persistently writing data, the GPU also writes the corresponding metadata to persistent video memory or periodically flushes it to non-volatile block storage. It should be noted that the GPU described above is configured as an application to generate storage data, the file system described above implements metadata management, control, and recording functions, and the namespace described above provides block devices.
[0163] The above embodiments of this disclosure further include: a first bus switch connected to at least one video memory component and a plurality of graphics processors; and at least one second bus switch connected to the first bus switch and a corresponding storage device.
[0164] The aforementioned first bus switch is a hardware component configured in a computer system to connect multiple devices, such as GPUs, video memory, CPUs, and storage devices, and to allow efficient data exchange between them. In this disclosure, the first bus switch acts as a bridge between video memory and the GPU, providing a high-bandwidth, low-latency data transmission path, allowing the GPU to access video memory in a point-to-point manner, thereby forming a unified, scalable, high-bandwidth video memory pool.
[0165] At least one of the aforementioned second bus switches is configured to connect the first bus switch and different block storage devices, enabling a direct connection between the GPU server and the block storage devices. This direct connection method reduces intermediate steps in data transmission, lowers access latency, and improves data transmission efficiency.
[0166] The first bus switch and at least one second bus switch of this disclosure together constitute a highly interconnected architecture between the GPU server cluster and the video memory components and block storage devices. The first bus switch enables high-speed data communication between the video memory components and the GPU through the graphics card interconnect protocol, forming a video memory pool; while the second bus switch enables the GPU to directly access multiple block storage devices through direct connection, reducing data access latency and improving storage performance.
[0167] For example, suppose a GPU server cluster contains four GPU servers, each of which requires access to a high-bandwidth memory pool and low-latency block storage. A first bus switch is responsible for establishing high-speed links between the GPUs and memory components within its servers, forming a high-bandwidth memory pool. Simultaneously, at least one second bus switch connects these servers to one or more block storage devices, enabling direct connections between the GPU servers and block storage devices and constructing a distributed storage pool. In this way, the GPU server cluster can efficiently access and manage large amounts of memory and storage resources, meeting the needs of large-scale AI model training and inference.
[0168] Introducing first and second bus switches into a GPU server cluster not only enhances memory pooling capabilities but also enables direct, high-speed communication between the GPU and block storage devices, significantly improving data access efficiency and storage performance. Specifically, the first bus switch allows the GPU to access memory components with high bandwidth, reducing memory access latency and accelerating the training and inference processes of AI models. The second bus switch allows the GPU to directly access block storage devices, avoiding the intermediate steps of the CPU and system memory, further reducing data access latency and increasing data transfer speed. This enables the GPU to quickly persist training data and model parameters, thereby enhancing the storage and computing capabilities of GPU servers when handling large-scale AI tasks.
[0169] In a scenario where a GPU server cluster is used for distributed training of deep learning models, the cluster consists of multiple GPU servers, each requiring GPUs to access a unified high-bandwidth memory pool. Through a first bus switch (PCIe switch), GPUs can access memory components at high speed, achieving memory pooling. Simultaneously, to accelerate persistent data storage and loading, the GPU servers in the cluster are also directly connected to multiple block storage devices via at least one second bus switch, constructing a distributed non-volatile storage pool. During training, model parameters and intermediate data generated by the GPUs can be directly written to the high-bandwidth memory pool, and can also be quickly written to block storage devices via the second bus switch, achieving persistent data storage. When training data or model parameters need to be loaded, the GPUs can also read data from the block storage devices via the second bus switch, further accelerating data transfer and model training, and improving the efficiency and performance of the GPU server cluster in large-scale AI model training.
[0170] Figure 10 is a schematic diagram of a direct-connect graphics card memory expansion and storage device expansion according to an embodiment of the present disclosure. As shown in Figure 10, the present disclosure can directly connect to a graphics processing unit (GPU) and expand a high-bandwidth memory pool and low-latency non-volatile block storage as needed. Multiple GPUs run a point-to-point graphics interconnect protocol to achieve high-bandwidth communication and pool the memory in each GPU for common access by all GPUs. The memory components shown in Figure 10 achieve point-to-point communication with each GPU through the same graphics interconnect protocol. In other words, the memory components can be regarded as GPUs without computing power, while the controller performs memory management and communication with each GPU through the graphics interconnect protocol. From the perspective of each GPU, the memory components are GPUs with computing modules removed and memory capacity increased. Figure 10 also shows the low-latency non-volatile block storage connected through a first bus switch (PCIe switch). The block storage is accessed by the GPU through a direct connection protocol between the GPU and the storage device. By cascading PCIe switches, the number of low-latency non-volatile block storage devices can be expanded, allowing for flexible expansion according to user needs. The Network Interface Card (NIC) connects to the first bus switch and is configured to enable network connectivity and communication. The direct-access storage protocol in Figure 10 refers to a communication protocol that allows the GPU to directly access storage devices without going through the traditional path of the CPU and system memory. This can reduce data transfer latency and improve data transfer efficiency.
[0171] Figure 11 is a schematic diagram of a distributed storage system constructed by a server cluster using direct-attached block storage pooling according to an embodiment of the present disclosure. Figure 11 shows a design for constructing a distributed storage pool using non-volatile block storage directly connected to the graphics processing units (GPUs) of multiple servers. Distributed management software runs on the processors (CPUs) and system memory of each server, and manages the system memory, CPU, and network interface cards (NICs) of each server in the direct-attached block device pooling management software. The NICs are connected to the CPUs, while numerous low-latency non-volatile block storage devices provide a distributed storage system with strong data consistency protection. The CPUs are presented as a unified storage space for multiple GPUs running in parallel. This unified storage space is then mapped to multiple specific non-volatile block storage components that meet fault domain rules based on its blocks. Taking a three-backup configuration as an example, the logical block i in Figure 11 is actually backed by three physical spaces of the same size (i.e., block i1, block i2, and block i3), each originating from three different servers that meet the fault domain requirements. The fault domain requirement ensures strong data consistency and avoids single points of failure in hardware and software. Data backups or data slices must be placed in different physical areas, such as different racks or different servers. The interconnect structure (fabric) in Figure 11 can have various options, and its function is to achieve high-bandwidth interconnection of multi-mode servers at the cluster level, including but not limited to the current Ethernet solution.
[0172] This disclosure presents a novel video memory component accessed via a GPU interconnect protocol and a block storage component directly connected to the GPU to accelerate the training and inference of large-scale intelligent computing models. The novel video memory component adopts a GPU memory pooling protocol, allowing its memory capacity to be accessed by multiple GPUs with high bandwidth, and provides both volatile and persistent memory semantics. The novel non-volatile block storage also employs a novel persistent medium to achieve low latency, presents multiple namespaces for multiple GPUs to share their storage space, and achieves data path acceleration through direct connection with the GPU. Simultaneously, it implements local file systems, block layers, drivers, and other functions through the CPU execution management path, and further designs distributed non-volatile storage pooling to achieve strongly consistent data storage within a GPU server cluster. GPUs can also directly manipulate the storage space through multiple directly connected non-volatile block storage units, establishing and maintaining simple metadata to access stored data blocks, while the metadata is saved through persistent video memory and other methods.
[0173] This disclosure also provides a server, including: a plurality of graphics processors and a data storage system of any one of the above embodiments.
[0174] This disclosure also provides a server cluster, including: multiple servers as described in the above embodiments, with the multiple servers mutually redundant.
[0175] This disclosure provides a data storage method as shown in FIG12. FIG12 is a flowchart of a data storage method according to an embodiment of this disclosure. As shown in FIG12, the method includes:
[0176] Step S1202: Receive a data access request sent by any one of the multiple graphics processors;
[0177] The above describes the control and communication mechanisms within a GPU server cluster. In GPU-intensive computing environments, multiple GPUs can access shared storage resources. GPUs send read and write requests to a unified storage system (which can be a memory management unit or a dedicated storage controller) through their interfaces to acquire or store data. These "data access requests" can be read requests, write requests, or mixed read / write requests, depending on the GPU's current computing task requirements.
[0178] Step S1204: Based on the data access request, determine the target access object from the video memory pool or at least one block storage component;
[0179] The video memory pool consists of video memory units in at least one video memory component and video memory units in multiple graphics processors.
[0180] The above steps form the core of the storage access logic. Based on the access request sent by the GPU, the system determines where data should be read from or written. The memory pool is a shared resource composed of memory units and memory components across multiple GPU servers, designed to provide high-bandwidth, low-latency memory access. Block storage components, on the other hand, are devices that provide persistent data storage. The system intelligently selects either the memory pool or the block storage component as the target access object based on the nature of the access request (such as data type and access mode) to improve data transfer efficiency and latency.
[0181] Step S1206: Based on the data access request, perform data access on the target access object to obtain the data access result;
[0182] In the steps described above, the system performs the actual read and write operations on the target access object. If the target is the video memory pool, data will be read or written to the video memory at high speed; if the target is a block storage component, data will be read and written via a persistent storage protocol. The system executes the corresponding data access operation based on the type of request sent by the GPU and collects the access results (read data or a signal confirming successful write) to prepare for the next step.
[0183] Step S1208: Send the data access result to the graphics processor.
[0184] The system feeds back the results of data access operations to the GPU that initiated the request. This could involve sending the read data back to the GPU for further processing, or confirming to the GPU that the data has been successfully written. Through an efficient data transfer mechanism, the system ensures that the GPU receives the access results in a timely manner, thereby maintaining the continuity of computing tasks.
[0185] The above steps implement the data access and management process in a GPU server cluster. From the moment a data access request is received from the GPU, the system intelligently schedules the data to determine whether it should be accessed from the GPU memory pool or the block storage component. The GPU memory pool provides high-performance, low-latency memory access, while the block storage component is responsible for persistent data storage. After executing the access operation, the system quickly feeds back the result to the GPU to ensure that GPU computing tasks can be performed efficiently and continuously. This mechanism not only improves the efficiency of data access but also provides the GPU with flexible storage resource selection, enabling it to better match and improve the needs of different computing tasks, especially in scenarios involving large-scale AI model training and inference.
[0186] Through the above steps, a data access request is received from any one of the multiple graphics processing units (GPUs). Based on the data access request, a target access object is determined from the memory pool or at least one block storage component. The memory pool consists of memory units in at least one memory component and memory units in multiple GPUs. Based on the data access request, data access is performed on the target access object to obtain the data access result. The data access result is then sent to the GPU. It is noteworthy that the memory component contains multiple memory units and communicates with multiple GPUs and other memory components via the GPU interconnect protocol. This design allows memory units to be shared by multiple GPUs, thus decoupling the fixed ratio of GPU computing power to memory capacity and bandwidth, achieving the goal of flexible on-demand configuration of memory resources. The communication link between the memory units and multiple GPUs via the GPU interconnect protocol constitutes the memory pool. The concept of the memory pool further expands the access scope of the GPU, enabling multiple GPUs to access shared memory resources in parallel. This not only increases memory capacity but also improves memory bandwidth, thereby accelerating the training and inference process of large models. By designing the graphics memory and block storage components, the efficiency of training and inference for large models can be significantly improved, solving the problem of data storage system bottlenecks limiting the processing efficiency of large intelligent computing models in related technologies. Furthermore, it improves data access speed and ensures persistent data storage. This, in turn, resolves the technical issue of data storage systems affecting the processing efficiency of large models in related technologies.
[0187] In the above embodiments of this disclosure, when the target access object is any video memory component, data access is performed on the target access object based on the data access request to obtain a data access result, including: parsing the data access request to obtain a target tag; when the target tag is a volatile tag, data access is performed on the volatile video memory unit in the video memory component based on the data access request to obtain a data access result; when the target tag is a persistent tag, data access is performed on the non-volatile video memory unit in the video memory component based on the data access request to obtain a data access result.
[0188] In the embodiments of this disclosure, the video memory component is designed to simultaneously support both volatile and non-volatile storage modes, which is achieved by using "target tags" to distinguish different data attributes and access requirements. Target tags are a key innovation, allowing the system to intelligently determine the storage location and access method of data based on its characteristics.
[0189] When the GPU sends a data access request, the request includes a target tag that indicates whether the requested data is volatile or persistent. This mechanism ensures that the system can determine which part of the video memory should store the data based on the request: volatile memory (such as DRAM) or non-volatile memory (such as persistent memory).
[0190] When the target tag is a volatile tag, if the data requested by the GPU is volatile, it means that the data does not need to persist after power is disconnected. This is typically applicable to data accessed frequently in a short period, such as model parameters and intermediate calculation results. The system combines the parsed volatile tag with the data access request to locate the volatile memory unit within the video memory module. Because volatile memory typically has higher bandwidth and lower latency, data access can be completed at extremely high speeds. The GPU's requested data will be directly read from or written to these volatile memory units, ensuring efficient execution of the computational task.
[0191] When the target label is a persistent label, if the data requested by the GPU is persistent, it means that the data must be retained even after a power outage to prevent the training or inference process from being interrupted by an unexpected power failure. This data typically includes model checkpoints, important configuration information, or datasets that need to be stored long-term. Once the system recognizes a persistent label, it directs the data access request to non-volatile memory units (such as PMEM) within the GPU memory. These non-volatile memory units guarantee data persistence; although their read / write speeds are slightly slower than volatile memory, they protect data from loss during power failures, thus avoiding the overhead of retraining the model.
[0192] Through the steps described above, the system can intelligently manage data storage and access, ensuring efficiency and data persistence. The access requests initiated by the GPU not only contain data location information but also implicitly include the data's lifecycle and access characteristics. This allows the memory components to dynamically adjust the data storage location based on the nature of the request. When data is volatile, it can be stored in high-bandwidth, low-latency volatile memory units; when data is persistent, it will be stored in non-volatile memory units, ensuring data integrity even in the event of a power failure. This mechanism significantly improves the computational efficiency and data processing capabilities of GPU server clusters, especially in large model training and inference scenarios. By rationally allocating and managing data storage locations, the system can better meet the storage performance requirements of different applications.
[0193] The above steps take into account the flexibility and efficiency of data storage, and realize intelligent classification and management of data through the target labeling mechanism. This ensures the accuracy and response speed of data access, while also reducing the risk of recalculation due to data loss, thereby improving the performance and reliability of the entire computing system.
[0194] In the above embodiments of this disclosure, when the target access object is any video memory component, the method further includes one of the following: when the data access request is a write request, allocating storage space for the data access request in the video memory component; when the video memory component meets the space reclamation conditions, reclaiming the storage space in the video memory component; when the data access request is a write request, the target tag is a volatile tag, and the remaining space of the volatile video memory unit is less than the amount of data corresponding to the write request, performing data access on the non-volatile video memory unit based on the data access request to obtain the data access result.
[0195] When the data access request is a write request, storage space is allocated in the video memory (VRAM) to accommodate it. When the GPU sends a write request, meaning it intends to store data in VRAM, the system needs to find or reserve sufficient space in the VRAM to store this data. VRAM, especially volatile memory units (such as DRAM), is ideal for short-term, high-frequency data access in GPU computing tasks due to its high bandwidth and low latency. The system intelligently allocates storage space based on the size of the requested data and the current VRAM usage, ensuring that write operations can be completed as quickly as possible while maintaining efficient utilization of VRAM.
[0196] When the memory units meet the space reclamation conditions, the storage space within them is reclaimed. Memory space is not infinite; as GPU computing tasks continue, data in memory units may become unnecessary, such as expired model parameters or intermediate calculation results. To improve memory utilization, the system periodically or based on certain conditions checks memory usage. Once it detects unnecessary data or space that meets reclamation conditions, such as expired data or memory unit idle time reaching a certain threshold, space reclamation is performed. This includes migrating data to non-volatile memory units (such as PMEM) to maintain its persistence, or directly clearing unnecessary data to free up space for new data, thereby ensuring dynamic adjustment and efficient utilization of memory.
[0197] When the data access request is a write request, and the remaining space of the volatile memory unit is less than the amount of data corresponding to the write request, data access is performed on the non-volatile memory unit based on the data access request to obtain the data access result.
[0198] When the amount of data the GPU needs to write exceeds the remaining space in the volatile memory units, the system automatically switches to non-volatile memory units (such as PMEM) for data writing. Although non-volatile memory units may not be as efficient as volatile memory units in terms of bandwidth and latency, they provide data persistence and are better suited for large-capacity data storage. Through this mechanism, even when memory capacity is limited, the GPU can continue its computational tasks, writing data to persistent storage to ensure uninterrupted computation, while also utilizing the large capacity of non-volatile memory units to store more data.
[0199] Through the steps described above, the data storage and access process can be improved based on the real-time demands of the GPU and the actual state of the memory components. By employing precise space management and intelligent volatile / non-volatile memory switching, the system not only improves the GPU's computational efficiency and reduces latency caused by insufficient memory capacity, but also ensures data persistence and integrity, providing a more stable and efficient support environment for GPU computing tasks. This mechanism is particularly effective in handling large-scale AI model training and inference tasks, enhancing data processing speed and overall system resource utilization efficiency.
[0200] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0201] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure.
[0202] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solutions of this disclosure, in essence, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this disclosure.
[0203] According to an embodiment of the present disclosure, a data storage device for implementing the above-described data storage method is also provided. FIG13 is a schematic diagram of a data storage device according to an embodiment of the present disclosure. As shown in FIG13, the device 1300 includes: a receiving module 1302, a determining module 1304, an access module 1306, and a sending module 1308.
[0204] The receiving module 1302 receives a data access request sent by any one of the multiple graphics processors; the determining module 1304 determines a target access object from the video memory pool or at least one block storage component based on the data access request, wherein the video memory pool consists of video memory units in at least one video memory component and video memory units in multiple graphics processors; the access module 1306 performs data access on the target access object based on the data access request and obtains the data access result; and the sending module 1308 sends the data access result to the graphics processor.
[0205] It should be noted that the receiving module 1302, determining module 1304, accessing module 1306, and sending module 1308 correspond to steps S1202 to S1208 in the above embodiments. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above modules can also be part of a device and run in the computer terminal 10 provided in the above embodiments.
[0206] In the above embodiments of this disclosure, the access module is further configured to parse the data access request to obtain the target tag; if the target tag is a volatile tag, to access the volatile memory unit in the video memory component based on the data access request to obtain the data access result; if the target tag is a persistent tag, to access the non-volatile memory unit in the video memory component based on the data access request to obtain the data access result.
[0207] In the embodiments disclosed above, the device further includes: a request module and a recycling module.
[0208] The request module is configured to allocate storage space in the video memory component for the data access request when the data access request is a write request; the reclamation module is configured to reclaim the storage space in the video memory component when the video memory component meets the space reclamation conditions; the access module is further configured to access the non-volatile video memory unit based on the data access request when the data access request is a write request, the target tag is a volatile tag, and the remaining space of the volatile video memory unit is less than the amount of data corresponding to the write request, and obtain the data access result.
[0209] It should be noted that the preferred embodiments involved in the above embodiments of this disclosure are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.
[0210] Embodiments of this disclosure can provide an electronic device, which can be any one of a group of electronic devices. Optionally, in this embodiment, the aforementioned electronic device can also be replaced with a terminal device such as a mobile terminal.
[0211] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0212] In this embodiment, the computer terminal described above can execute the program code in the method.
[0213] Optionally, FIG14 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. As shown in FIG14, the electronic device A may include: one or more (only one is shown in the figure) processors 102, memory 104, memory controller, and peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display.
[0214] The memory can be configured to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0215] It will be understood by those skilled in the art that the structure shown in Figure 14 is merely illustrative, and the electronic device may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. Figure 14 does not limit the structure of the aforementioned electronic device. For example, electronic device A may include more or fewer components (such as a network interface, a display device, etc.) than shown in the figure, or may have a different configuration than that shown in Figure 14.
[0216] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0217] Embodiments of this disclosure also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium may be configured to store program code executed by the method provided in the above embodiments.
[0218] Optionally, in this embodiment, the storage medium may be located in any one of the electronic devices in the group of electronic devices in the computer network, or in any one of the mobile terminals in the group of mobile terminals.
[0219] Embodiments of this disclosure also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.
[0220] Embodiments of this disclosure also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium configured to store a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.
[0221] Embodiments of this disclosure also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.
[0222] In the above embodiments of this disclosure, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0223] In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0224] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0225] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0226] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0227] The above description is only a preferred embodiment of this disclosure. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure.
Claims
1. A data storage system, comprising: At least one video memory component and at least one block storage component; The video memory component includes multiple video memory units and communicates with multiple graphics processors and other video memory components respectively through the graphics card interconnection protocol. The other video memory components are the video memory components other than the video memory component in the at least one video memory component. The video memory units in the at least one video memory component and the video memory units in the multiple graphics processors constitute a video memory pool. The video memory pool is configured to cache the data accessed by the multiple graphics processors. The block storage component communicates with the plurality of graphics processors via a direct connection protocol and is configured to persistently store the data accessed by the plurality of graphics processors.
2. The system according to claim 1, wherein, The video memory component includes: A graphics card interconnect interface is connected to the plurality of graphics processors and the other video memory components respectively. A graphics card interconnect protocol runs on the graphics card interconnect interface. The graphics card interconnect interface is configured to transmit first data accessed by any graphics processor, wherein the first data includes a target tag. Volatile memory units are configured to store data containing volatile tags; Non-volatile memory units are configured to store data containing persistent tags; The controller, connected to the graphics card interconnect interface, the volatile memory unit, and the non-volatile memory unit, is configured to write the first data to the volatile memory unit or the non-volatile memory unit, or read the first data from the volatile memory unit or the non-volatile memory unit, based on the target tag.
3. The system of claim 2, wherein, The controller includes: The first cache, connected to the graphics card interconnect interface, is configured to cache the first data; A volatile memory management unit is connected to the first cache and the volatile video memory unit respectively, and is configured to write the first data to the volatile video memory unit or read the first data from the volatile video memory unit when the target tag is a volatile tag. A non-volatile memory management unit is connected to the first cache and the non-volatile video memory unit respectively, and is configured to write the first data to the non-volatile video memory unit or read the first data from the non-volatile video memory unit when the target tag is a persistent tag.
4. The system of claim 3, wherein, The non-volatile memory management unit includes: A non-volatile memory management controller, connected to the first cache, is configured to encode the first data using short error-correcting codes to obtain first encoded data, or to decode the raw data obtained from the non-volatile video memory unit using short error-correcting codes to obtain the first data. A non-volatile memory management interface is connected to the non-volatile memory management controller and the non-volatile video memory unit, respectively, and is configured to write the first encoded data into the non-volatile video memory unit and read the original data from the non-volatile video memory unit.
5. The system according to claim 4, wherein, The non-volatile memory management controller includes: The access unit is configured to modify and write non-volatile data read from the non-volatile memory unit based on the first data to obtain data to be written, or to read the first decoded data obtained from the non-volatile memory unit to obtain the first data. A short error-correcting code codec, connected to the access unit, is configured to encode the data to be written using short error-correcting codes to obtain the first encoded data, or to decode the original data to obtain the first decoded data. The back-end controller includes at least one first medium channel, which is respectively connected to at least one medium particle in the short error-correcting code codec and the non-volatile memory unit, and is configured to transmit the first encoded data or the original data.
6. The system according to claim 1, wherein, The block storage component includes: A graphics card interconnect protocol controller, connected to the plurality of graphics processors, is configured to transmit second data accessed by any one of the graphics processors; The second cache, connected to the graphics card interconnect protocol controller, is configured to cache the second data; An error correction code encoding / decoding unit, connected to the second cache, is configured to encode the second data with error correction codes to obtain second encoded data, or to decode persistent data obtained from the at least one block storage medium with error correction codes to obtain the second data; At least one second medium channel is connected to the error correction code encoding / decoding unit and the at least one block storage medium, respectively, and is configured to transmit the second encoded data or the persistent data.
7. The system according to claim 6, wherein, The block storage component also includes: A compression unit, connected to the second cache, is configured to decompress the second data to obtain decompressed data, or to compress decrypted data obtained from the at least one block storage medium to obtain the second data; An encryption / decryption unit, connected to the compression unit, is configured to encrypt the decompressed data to obtain encrypted data, or to decrypt the demodulated data obtained from the at least one block storage medium to obtain the decrypted data. A modulation and demodulation unit, connected to the encryption and decryption unit, is configured to modulate the encrypted data to obtain modulated data, or to demodulate the second decoded data obtained from the at least one block storage medium to obtain the demodulated data; The error correction code encoding / decoding unit is connected to the modulation / demodulation unit and is configured to encode the modulated data with error correction codes to obtain the second encoded data, or to decode the persistent data with error correction codes to obtain the second decoded data.
8. The system according to any one of claims 1 to 7, wherein, The block storage component includes: Multiple namespaces, corresponding to the multiple graphics processors, are configured to divide the storage space of the block storage component.
9. The system according to any one of claims 1 to 7, wherein, Also includes: A first bus switch is connected to the at least one video memory component and the plurality of graphics processors; At least one second bus switch is provided, which is connected to the first bus switch and the corresponding storage device.
10. A server comprising: Multiple graphics processors and a data storage system according to any one of claims 1 to 9.
11. A server cluster, comprising: Multiple servers as described in claim 10, with multiple servers mutually redundant.
12. A data storage method, applied to the data storage system according to any one of claims 1 to 9, the method comprising: Receive data access requests from any one of the multiple graphics processors; Based on the data access request, a target access object is determined from the video memory pool or at least one block storage component, wherein the video memory pool is composed of video memory units in at least one video memory component and video memory units in the plurality of graphics processors. Based on the data access request, data access is performed on the target access object to obtain the data access result; The data access result is sent to the graphics processor.
13. The method according to claim 12, wherein, When the target access object is any video memory component, the step of accessing the target access object based on the data access request and obtaining the data access result includes: The data access request is parsed to obtain the target tag; If the target tag is a volatile tag, data access is performed on the volatile memory unit in the video memory component based on the data access request to obtain the data access result; If the target tag is a persistent tag, data access is performed on the non-volatile memory unit in the video memory component based on the data access request to obtain the data access result.
14. The method according to claim 13, wherein, When the target tag is a volatile tag, based on the data access request, data access is performed on the volatile memory unit in the video memory component to obtain the data access result, including: When the data access request is a write request, the data is written to the volatile video memory unit to obtain the data access result; When the data access request is a read request, the data is read from the volatile memory unit to obtain the data access result.
15. The method according to claim 13, wherein, When the target tag is a persistent tag, based on the data access request, data access is performed on the non-volatile memory units in the video memory component to obtain the data access result, including: When the data access request is a write request, the data is written to the non-volatile video memory unit to obtain the data access result; When the data access request is a read request, the data is read from the non-volatile video memory unit to obtain the data access result.
16. The method according to claim 13, wherein, When the target access object is any video memory component, the method further includes one of the following: When the data access request is a write request, storage space is allocated for the data access request in the video memory component; When the video memory component meets the space reclamation conditions, the storage space in the video memory component is reclaimed; When the data access request is a write request, the target tag is a volatile tag, and the remaining space of the volatile memory unit is less than the amount of data corresponding to the write request, the non-volatile memory unit is accessed based on the data access request to obtain the data access result.
17. A computer-readable storage medium comprising a stored executable program, wherein, When the executable program runs, it controls the device containing the storage medium to perform the following steps: Receive data access requests from any one of the multiple graphics processors; Based on the data access request, a target access object is determined from the video memory pool or at least one block storage component, wherein the video memory pool is composed of video memory units in at least one video memory component and video memory units in the plurality of graphics processors. Based on the data access request, data access is performed on the target access object to obtain the data access result; The data access result is sent to the graphics processor.
18. The computer-readable storage medium according to claim 17, wherein, When the target access object is any video memory component, the device containing the storage medium also performs the following steps during the execution of the executable program: The data access request is parsed to obtain the target tag; If the target tag is a volatile tag, data access is performed on the volatile memory unit in the video memory component based on the data access request to obtain the data access result; If the target tag is a persistent tag, data access is performed on the non-volatile memory unit in the video memory component based on the data access request to obtain the data access result.
19. The computer-readable storage medium according to claim 18, wherein, If the target tag is a volatile tag, the device containing the storage medium will also perform the following steps during the execution of the executable program: When the data access request is a write request, the data is written to the volatile video memory unit to obtain the data access result; When the data access request is a read request, the data is read from the volatile memory unit to obtain the data access result.
20. The computer-readable storage medium of claim 18, wherein, When the target access object is any video memory component, the executable program controls the device containing the storage medium to perform the following steps during runtime: When the data access request is a write request, storage space is allocated for the data access request in the video memory component; When the video memory component meets the space reclamation conditions, the storage space in the video memory component is reclaimed; When the data access request is a write request, the target tag is a volatile tag, and the remaining space of the volatile memory unit is less than the amount of data corresponding to the write request, the non-volatile memory unit is accessed based on the data access request to obtain the data access result.