Storage system, cluster node, system, and storage resource scheduling method and apparatus

By introducing a first switch into the storage system, direct connection and resource sharing between the storage controller and the memory are achieved, solving the problems of low resource utilization and low access efficiency caused by the coupling of computing and storage in traditional storage systems, and meeting the high bandwidth and low latency requirements of new distributed applications and artificial intelligence.

WO2026091386A1PCT designated stage Publication Date: 2026-05-07INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2025-03-19
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

In traditional storage systems, the coupling of computing and storage prevents data from directly accessing the storage medium, resulting in low resource utilization, low access efficiency, and an inability to meet the high bandwidth and low latency requirements of new distributed applications and artificial intelligence.

Method used

By introducing a first switch into the storage system, a direct connection between the storage controller and the memory is achieved. The switch is used to allocate and schedule storage resources, enabling storage resource sharing and reducing the interaction between storage controllers and cross-layer memory relocation operations.

Benefits of technology

It improves the utilization and access efficiency of storage resources, supports independent expansion of computing and storage, meets the high bandwidth and low latency requirements of new distributed applications and artificial intelligence, and reduces costs and latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025083484_07052026_PF_FP_ABST
    Figure CN2025083484_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a storage system, a cluster node, a system, and a storage resource scheduling method and apparatus. The storage system comprises a host, a first switch, first storage devices and a plurality of storage controllers. The host is connected to the plurality of storage controllers; the plurality of storage controllers are all directly connected to the first switch; the plurality of storage controllers are used for recognizing a data request sent by the host, so as to determine corresponding storage resources; and the first switch is connected to the first storage devices, and the first switch is used for receiving storage resources of the plurality of storage controllers, and performing allocation and scheduling processing on resources of the first storage devices on the basis of the storage resources of the plurality of storage controllers, so as to realize storage resource sharing of the first storage devices. The first switch comprises a baseboard and a plurality of switch units, wherein the storage controllers are each connected to a switch unit, the switch units are located on switch boards, which are connected by means of the baseboard to realize the connection between the switch units, and the switch units are respectively connected to corresponding first storage devices.
Need to check novelty before this filing date? Find Prior Art

Description

Storage system, cluster node, system, storage resource scheduling method and device

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411546272.1, filed on November 1, 2024, entitled “Storage System, Cluster Node, System, Storage Resource Scheduling Method and Apparatus”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to a storage system, a cluster node, a system, a storage resource scheduling method, and an apparatus. Background Technology

[0004] In the architecture design of storage systems, the relationship between computing and storage is one of the core considerations, and it is usually tightly coupled, that is, storage devices are directly connected to servers.

[0005] Due to the uneven development of hardware resources such as computing and storage, the coupling between computing and storage cannot be scaled independently, thus failing to simultaneously meet the business's computing and storage needs. When computing needs to be scaled up or down, or when data requests are made, data needs to be moved. The connection and access process between storage controllers within the architecture involves multiple control interactions between storage controllers and cross-layer memory migration operations. As a result, data cannot directly access the storage medium, thus wasting a significant amount of computing power, network, and memory channel resources of the storage controllers, further increasing access latency and reducing access efficiency.

[0006] Therefore, how to decouple computation and storage in a storage system so that data can directly access the storage medium and improve resource utilization and access efficiency is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] According to an embodiment of this application, in a first aspect, a storage system is provided, including a host, a first switch, a first memory, and a plurality of storage controllers;

[0008] The host connects to multiple storage controllers;

[0009] Multiple storage controllers are directly connected to the first switch; these multiple storage controllers are used to identify and process data requests sent by the host to determine the corresponding storage resources.

[0010] The first switch is connected to the first memory; the first switch is used to receive storage resources from multiple storage controllers, and to allocate and schedule the resources of the first memory according to the storage resources of the multiple storage controllers to realize the sharing of storage resources of the first memory.

[0011] The first switch includes a base plate and multiple switching units;

[0012] Each storage controller is connected to each switching unit;

[0013] Each switching unit is located on a switching board, and the switching boards are connected to each other via a substrate to achieve the connection between the switching units; and

[0014] Each switching unit is connected to its corresponding first memory.

[0015] Secondly, a cluster node is also provided, comprising at least one of the aforementioned storage systems; and

[0016] When there are multiple storage systems, they are connected by fiber optic links.

[0017] Thirdly, a cluster system is also provided, comprising at least one of the aforementioned cluster nodes; and

[0018] When there are multiple cluster nodes, the cluster nodes are connected through fiber optic links.

[0019] Fourthly, a storage resource scheduling method based on a storage system is also provided, applied to a first switch of the storage system. The storage system includes a host, a first switch, a first memory, and multiple storage controllers; the host is connected to the multiple storage controllers; each of the multiple storage controllers is directly connected to the first switch; the first switch is connected to the first memory; the first switch includes a base plate and multiple switching units; each storage controller is connected to each switching unit; each switching unit is located on a switching board, and the switching boards are connected to each other through the base plate to realize the connection between the switching units; each switching unit is connected to its corresponding first memory; the method includes:

[0020] Receive storage resources sent by multiple storage controllers; wherein the storage resources are identified and determined within the storage controller by data requests sent by the host;

[0021] Acquire resources from the first memory; and

[0022] The storage resources of the first memory are allocated and scheduled according to the storage resources of multiple storage controllers to achieve storage resource sharing of the first memory.

[0023] Fifthly, a storage resource scheduling device based on a storage system is also provided, applied to a first switch of the storage system. The storage system includes a host, a first switch, a first memory, and multiple storage controllers; the host is connected to the multiple storage controllers; each of the multiple storage controllers is directly connected to the first switch; the first switch is connected to the first memory; the first switch includes a base plate and multiple switching units; each storage controller is connected to each switching unit; each switching unit is located on a switching board, and the switching boards are connected to each other through the base plate to realize the connection between the switching units; each switching unit is connected to its corresponding first memory; the device includes:

[0024] The receiving module is used to receive storage resources sent by multiple storage controllers; wherein the storage resources are identified and determined within the storage controller by data requests sent by the host.

[0025] The acquisition module is used to acquire resources of the first memory; and

[0026] The allocation and scheduling module is used to allocate and schedule the resources of the first memory according to the storage resources of multiple storage controllers, so as to realize the sharing of storage resources of the first memory.

[0027] Sixthly, a storage resource scheduling device based on a storage system is also provided, comprising:

[0028] A second memory is used to store computer programs; and

[0029] The second processor is used to implement the steps of the storage resource scheduling method based on the storage system described above when executing computer programs.

[0030] In a seventh aspect, a non-transitory computer-readable storage medium is also provided, on which a computer program is stored, and when the computer program is executed by a second processor, it implements the steps of the storage resource scheduling method based on the storage system as described above.

[0031] Eighthly, a computer program product is also provided, including a computer program / computer-readable instructions that, when executed by a second processor, implement the steps of the storage resource scheduling method based on the storage system described above.

[0032] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description

[0033] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 is a structural diagram of a storage system provided in some embodiments of this application;

[0035] Figure 2 is a schematic diagram of storage resource sharing of the first memory provided in some embodiments of this application;

[0036] Figure 3 is a structural diagram of a first switch provided in some embodiments of this application;

[0037] Figure 4 is a schematic diagram of the structure of the base plate and switching board of the first switch provided in some embodiments of this application;

[0038] Figure 5 is a schematic diagram of the optical interconnect interface between the transmitter and receiver provided in some embodiments of this application;

[0039] Figure 6 is a schematic diagram of communication between the sending end and the receiving end provided in some other embodiments of this application;

[0040] Figure 7 is a schematic diagram of the overall layout of a storage system provided in some other embodiments of this application;

[0041] Figure 8 is a schematic diagram of the structure of a first data storage device provided in some embodiments of this application;

[0042] Figure 9 is a schematic diagram of the structure of a second data storage device provided in some embodiments of this application;

[0043] Figure 10 is a schematic diagram of the structure of a cluster node provided in some embodiments of this application;

[0044] Figure 11 is a flowchart of a storage resource scheduling method based on a storage system provided in some embodiments of this application;

[0045] Figure 12 is a schematic diagram of the internal pooling system in which the first switch is located, according to some embodiments of this application;

[0046] Figure 13 is a schematic diagram of information display under the asset management module provided in some embodiments of this application;

[0047] Figure 14 is a schematic diagram of resource allocation and scheduling provided in some embodiments of this application;

[0048] Figure 15 is a structural diagram of a storage resource scheduling device based on a storage system provided in some embodiments of this application;

[0049] Figure 16 is a structural diagram of a storage resource scheduling device based on a storage system provided in some embodiments of this application;

[0050] Figure 17 is a structural block diagram of a non-transitory computer-readable storage medium provided in some embodiments of this application;

[0051] Figure 18 is a structural block diagram of a computer program product provided in some embodiments of this application. Detailed Implementation

[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0053] The core of this application is to provide a storage system, cluster node, system, storage resource scheduling method and apparatus to solve the problem that the coupling between computing and storage in the storage system causes data to be unable to directly access the storage medium, thus reducing resource utilization and access efficiency.

[0054] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] Storage, computing, and networking are key components of a data center. To fully realize its potential, it's essential to ensure efficient data storage, high computing speeds, and stable network transmission. However, in practice, many data centers adopt general-purpose servers and integrated computing and storage solutions, leading to inefficient storage and wasted computing resources.

[0056] In practical business applications, traditional storage solutions generally fall into two categories: converged storage and compute, and separate storage and compute. Converged storage and compute can uniformly manage and schedule computing, storage, and network resources, and has horizontal scalability. However, when actual business needs for computing and storage differ and their growth rates vary, problems arise with inflexible resource expansion and low utilization. Separate storage and compute, on the other hand, splits storage and computing resources into independent subsystems, offering advantages such as high resource utilization, efficient sharing of storage resources, flexible deployment across multiple scenarios, and network-storage collaboration. Traditional separate storage and compute architectures separate computing and storage resources into independent computing and storage domains, interconnecting them via Ethernet or dedicated storage networks. This enables flexible expansion and efficient sharing of storage resources and has been widely applied in core transaction systems and databases in industries such as finance and telecommunications.

[0057] With the rapid development of data storage technology and the accelerated innovation of new applications, the development speed of hardware resources such as computing and storage is unbalanced, and the difference between the life cycle of computing power and the life cycle of data is becoming increasingly prominent. In order to meet the requirements of new data centers for resource utilization, reliability, performance and efficiency, provide efficient storage support for the development of next-generation information technologies such as cloud computing, big data and high-performance computing, and promote the evolution towards composable disaggregated infrastructure, we need to achieve intensive and high-density high-quality development.

[0058] Current data storage is mainly divided into five categories: Direct Attached Storage (DAS), Storage Area Network (SAN), Network Attached Storage (NAS), cloud storage, and software-defined storage. Regardless of the type, storage systems are limited by hardware and software design, making it difficult to achieve a balance between performance and resource utilization. At the resource level, the ratio of computing to storage resources in traditional storage server architectures is fixed and cannot be independently expanded. In practical applications, when new data applications emerge, enterprises typically adopt the simplest application-local disk coupled server architecture for rapid deployment and trial of new services. Adopting an integrated architecture based on server-deployed storage allows for unified management and scheduling of computing, storage, and network resources. However, due to the uneven development speed of computing and storage hardware resources, the difference between the lifecycle of computing power and the lifecycle of data is widening, resulting in insufficient scalability of traditional storage server architectures. On the one hand, the coupling of computing and storage makes it impossible to independently expand computing and storage resources, failing to simultaneously meet the business's computing and storage needs. On the other hand, due to the coupling of computing and storage, when computing needs to be scaled up or down, data must be moved, thus affecting the performance of storage nodes providing computing services simultaneously. If storage and computation can be separated, the computation layer and storage layer can independently add or remove nodes without interfering with each other.

[0059] At the business application level, traditional storage servers cannot meet the demands of new distributed applications for minimalist and efficient shared storage. The emergence of microservices, serverless computing, and other new distributed applications, which are expanding from stateless to stateful architectures and including containerized components such as databases and message buses, are creating increasing demands for high-concurrency, low-latency shared data access. Meanwhile, applications such as artificial intelligence and machine learning require the collaboration of large amounts of heterogeneous computing power, and even shared memory access. They prioritize high-bandwidth, low-latency access capabilities, requiring only lightweight and convenient shared storage systems, rather than traditional storage systems with complex enterprise characteristics.

[0060] In data-intensive scenarios, traditional storage servers face multiple performance challenges: Traditional storage servers rely on general-purpose central processing units (CPUs) to handle infrastructure layer tasks, resulting in low efficiency. With the slowdown of Moore's Law, the marginal cost of increasing general-purpose CPU performance is rapidly rising. Data shows that the current annual performance growth rate of CPUs (after area normalization) is only about 3%. In 100G networks, CPUs processing network requests consume up to 30% of computing power. Since general-purpose CPUs are not designed specifically for data processing, their energy efficiency is relatively low. Furthermore, with increasing reliability requirements, the demand for data processing computations such as high-proportion erasure coding also adds to the burden on general-purpose CPUs. In traditional storage systems, host data requests require multiple CPU control interactions and cross-layer memory transfers, preventing direct data access to the storage medium, thus wasting significant CPU computing power and network / memory channel resources, and increasing latency. In summary, in data-intensive scenarios, the "data center tax" paid by applications for acquiring data under the traditional CPU-centric server architecture is constantly increasing. The main solution to this problem is to integrate heterogeneous computing power and offload data processing tasks on the input / output (I / O) path. The storage system provided in this application can solve the aforementioned technical problem.

[0061] Figure 1 is a structural diagram of a storage system provided in an embodiment of this application. As shown in Figure 1, the storage system includes a host 1, a first switch 3, a first memory 4, and multiple storage controllers 2.

[0062] Host 1 is connected to multiple storage controllers 2;

[0063] Multiple storage controllers 2 are directly connected to the first switch 3; the multiple storage controllers 2 are used to identify and process the data requests sent by the host 1 to determine the corresponding storage resources.

[0064] The first switch 3 is connected to the first memory 4; the first switch 3 is used to receive storage resources from multiple storage controllers 2, and to allocate and schedule the resources of the first memory 4 according to the storage resources of multiple storage controllers 2 to realize the sharing of storage resources of the first memory 4.

[0065] The first switch 3 includes a base plate 5 and multiple switching units;

[0066] Each storage controller 2 is connected to each switching unit;

[0067] Each switching unit is located on a switching board 6, and the switching boards 6 are connected to each other via a base plate 5 to realize the connection between the switching units;

[0068] Each switching unit is connected to its corresponding first memory 4.

[0069] Specifically, the first switch and storage controller within the storage system can be configured inside the storage server or used as external devices. The first switch is a network device used for forwarding electrical (optical) signals and can access the dedicated electrical signal paths provided by any two network nodes of the switch. The type of switch link is not limited in this application; it can be Ethernet, Compute Express Link (CXL), Peripheral Component Interconnect Express (PCIE) link, or other link types, depending on the actual situation. The type of transmitted signal is not limited; it can be electrical signals, optical signals, etc., and can be configured according to the actual situation. The storage controller can be a component inside the storage server or exist as an independent device, communicating with the storage server or other computing nodes via external connections (Fibre Channel or Ethernet, etc.).

[0070] The host connects to multiple storage controllers, acting as a computing unit. Through the connection of storage controllers and switches, it decouples from the primary storage, remotely integrating the server's local disks and memory. This transforms the storage server into a diskless server and a remote storage pool. This allows subsequent applications to select and configure virtual disks and pooled memory spaces with different performance and capacities according to their needs. When storage resource utilization is low, the storage server can return idle storage resources to storage modules, improving resource utilization through this decoupling method.

[0071] In addition, independent computing and storage enable on-demand expansion and cost savings. It also simplifies server purchase costs by purchasing the host and storage media independently, eliminating the need to distinguish between complex server types such as capacity-based, performance-based, and balanced servers.

[0072] In this embodiment, the storage system has multiple storage controllers. The host is connected to all of these controllers, but they do not interact with each other; each controller is directly connected to the first switch. Conventional storage systems involve multiple storage controllers interacting with each other, each controller corresponding to a memory module. Memory data transfer is handled through multiple forwardings, multiple control interactions between storage controllers, and cross-layer memory movement. Global memory is entirely controlled by the storage controllers. For example, storage controller 1' is connected to SSD pool A1, and storage controller 2' is connected to SSD pool B1. If storage controller 1' wants to access SSD pool B1, it needs to do so through storage controller 2', increasing the host's computing resources and preventing other direct access to the storage medium. This embodiment, however, uses a single-processing mechanism via the first switch to convert global memory into multi-point direct control via the switch protocol, achieving global metadata access encoding.

[0073] Multiple storage controllers are used to identify data requests sent by the host to determine the storage resources required by the host. The specific identification process is not limited; it can involve combining algorithms for identification, or identifying the data requests through encoding and decoding, etc. A first switch is connected to a first memory. After receiving the storage resources identified and processed by the multiple storage controllers, the first switch allocates and schedules the resources of the first memory to achieve resource sharing. It should be noted that the number of first memories in this embodiment is not limited; it can be one or multiple. The first memory can be a virtual disk, cloud storage, or an actual hardware storage component; it is not limited here.

[0074] Regarding the allocation and scheduling method, there are no restrictions. It can be to monitor in real time and make the remaining resources in the first memory available to the host, or it can be to calculate the selection of the remaining resources under multiple first memories based on a certain model algorithm. There are no restrictions here.

[0075] As shown in Figure 1, multiple storage controllers are represented by 1' to N'. There can be one or more first memories, represented by A1' to AN' when multiple first memories exist. In Figure 1, one switching unit corresponds to one first memory, or one switching unit can correspond to multiple first memories. That is, all first memories are connected to only one corresponding switching unit. Memory sharing among multiple first memories within the first switch is achieved through connections between switching units (i.e., connections between multiple switching units are achieved through a substrate). Conventional connections between switches and memories involve each memory being connected to each switch individually, without connections between switches. This increases the number of branches between each memory and each switch, reducing the IO (Input / Output) capability of the shared hard drive and consequently decreasing the transmission rate of a single branch. Furthermore, multiple branches increase data security risks; insufficient bandwidth due to multiple branches between the hard drive and the switch can even cause lag or disconnections during remote control. In this embodiment, the switching units are connected to achieve memory sharing among the first memories connected to each switching unit; not every first memory is connected to all switching units. For example, there are five first memories, A1', A2', A3', A4', and A5', corresponding to two switching units. One switching unit connects to the first memories A1', A2', and A3', while the other switching unit connects to A4' and A5'. The two switching units are interconnected, enabling interactive sharing among the five first memories. A conventional approach would connect each of the five first memories to a single switching unit, resulting in a large number of branches in the first memories.

[0076] Figure 2 is a schematic diagram of storage resource sharing of a first memory according to an embodiment of this application. As shown in Figure 2, taking four storage controllers as an example, multiple storage controllers are located in the storage control layer, and multiple first memories are located in a heterogeneous memory pool, with the actual corresponding capacity layer being hard disks (represented by B1' to B4') for storage. The first memory read and write data flow ensures performance requirements, and the data redundancy control flow ensures redundancy compass and data source, realizing decoupling and pooling of memory and SSD storage resources, improving resource utilization and system performance.

[0077] The first switch includes a base plate and multiple switching units;

[0078] Each storage controller is connected to each switching unit;

[0079] Each switching unit is located on a switching board, and the switching boards are connected to each other via a base plate;

[0080] Each switching unit is connected to the first memory.

[0081] Specifically, Figure 3 is a structural diagram of a first switch provided in an embodiment of this application. As shown in Figure 3, the first switch 3 includes a substrate 5 and multiple switching units. The substrate 5 is typically a copper-clad laminate, which not only provides physical support for electronic components but also provides electrical connections and signal transmission pathways. It is mainly a circuit board. Each switching unit is located on the switching board 6 in Figure 3 (the board where each switching unit is located is indicated by a dashed box in the figure). The switching boards 6 are electrically connected to each other through the substrate 5. The switching board 6 is responsible for data exchange functions and is the core module. It can output high-speed links of various bandwidth types, such as 32 x 16 bandwidth high-speed links, to provide large-scale, highly flexible high-speed link interconnection applications. The substrate 5 mainly carries the memory resource management engine (such as the first controller) and the control platform (such as the substrate 5 management controller).

[0082] The first switch integrates a management unit responsible for resource scheduling and management, device status monitoring, power-on / off management, and overall rack power-on / off management. Compact space utilization is achieved through optimized system architecture design. This first switch supports 32 high-speed I / O interfaces, providing an ultra-high-performance, ultra-large-scale switching network to meet the topology requirements of various storage system applications.

[0083] Referring to Figure 3, each storage controller 2 is connected to each switching unit on the first switch 3, and each switching unit is connected to the first memory 4. In this embodiment, the first memory 4 is located within the storage resource pool. Table 1 shows the specifications of the first switch. As shown in Table 1, the first switch 3 is the core switching unit of the storage system, providing 32 high-speed I / O interfaces in standby mode. The ports support arbitrary uplink and downlink configurations. Storage controllers 2 and the first memory 4 are interconnected with the first switch 3 via high-speed cables, enabling system reconfiguration. In this embodiment, the system topology can be interconnected according to different convergence ratios such as 1:1, 1:2, or 1:3, depending on application requirements. This allows for large-scale, high-speed, and highly flexible data link transmission, meeting the optimal performance requirements of the system. It should be noted that convergence usually refers to the process of merging or unifying data flows or signal flows in a network at a certain point. For example, in network design, multiple network traffic flows may converge onto a single link through routers or switches. In this embodiment, the convergence ratio refers to the ratio of the link bandwidth between the first switch 3 and the storage controller 2 to the link bandwidth between the first switch 3 and the first memory 4.

[0084] Table 1 Specifications of the First Switch

[0085] The Form Factor in Table 1 refers to the physical dimensions. In Table 1, the first switch corresponds to a 19-inch 2U form factor, where U represents the thickness. The power supply is a Power Supply Unit (PSU) that supports redundant mode. The fans are also redundant, where N represents the number of fans.

[0086] This application provides a storage system including a host, a first switch, a first memory, and multiple storage controllers. The host is connected to the multiple storage controllers. Each storage controller is directly connected to the first switch. The multiple storage controllers are used to identify and process data requests sent by the host to determine the corresponding storage resources. The first switch is connected to the first memory. The first switch includes a base plate and multiple switching units. Each storage controller is connected to each switching unit. Each switching unit is located on a switching board, and the switching boards are connected to each other via the base plate to achieve connection between the switching units. Each switching unit is connected to its corresponding first memory. The first switch is used to receive storage resources from the multiple storage controllers and to allocate and schedule the resources of the first memory according to the storage resources of the multiple storage controllers to achieve storage resource sharing. Separating the first memory (medium) and the host (computing unit) means moving the local disk and memory in the storage system away, achieving decoupling between computing and the medium, and allowing for the selection and configuration of memory spaces with different performance and capacity according to application needs. Multiple storage controllers are directly connected to the first switch, without inter-controller interaction. This eliminates the need for multiple controller interactions and cross-layer memory relocation operations. The first switch provides direct access to the first memory, saving storage controller resources in terms of computing power, network, and memory channels, reducing latency, and improving resource utilization. Furthermore, the computing units and the first memory are independent, allowing for on-demand expansion and contraction and independent upgrades, saving costs. This further reduces interaction and usage between storage controllers, enabling cross-controller memory and storage resource sharing, improving memory bandwidth and storage system performance. Finally, each switching unit connects to its respective first memory, enabling connections between first memories within each switching unit within the first switch. This reduces the number of branches for each first memory in each switching unit, preventing data transfer rate degradation. The connection relationship between the first switch's baseboard and multiple switching units is key to hardware resource decoupling and reconfiguration. By combining high-performance, large-scale switching I / O with basic management software, the system enables flexible switching between various storage server topologies and on-demand resource allocation, meeting the memory expansion needs of multiple storage application scenarios.

[0087] In some embodiments, the substrate includes a first controller, a network switching unit, a control platform, and a second controller;

[0088] The two ends of the first controller, the two ends of the network switching unit, and the two ends of the second controller are connected to different switching boards;

[0089] The first controller, network switching unit, and control platform are connected in sequence.

[0090] Specifically, Figure 4 is a schematic diagram of the structure of the base plate and switching board of a first switch provided in an embodiment of this application. As shown in Figure 4, the base plate 5 mainly carries a memory resource management engine (such as a first controller) and a control platform (such as a base plate management controller). The memory resource management engine is mainly responsible for completing the initialization configuration and port configuration of the first switch, and realizing memory resource scheduling management. The control platform mainly performs real-time status detection through the management bus and reports through the management network port. Once the machine malfunctions (such as the first switch malfunctioning, switching link error alarm, etc.), the control platform will transmit the data upward to alarm, and record it in the form of a log by detecting the detailed time scale of the signal, so that users can retrieve and analyze the cause of the malfunction.

[0091] The primary function of the first controller is as a memory resource management engine; it can also be a specific type of controller, which is not limited here. Its main function is to implement memory resource scheduling and management. The network switching unit is mainly used for data transmission between the control platform and the first controller. It can use any interface protocol, such as high-speed serial interface standards (e.g., Serial Gigabit Media Independent Interface, SGMII). The control platform typically refers to a centralized management tool used to manage and monitor various software and hardware systems. It helps administrators better manage and control the system. Here, the control platform can be a Baseboard Management Controller (BMC) or other control and management systems, which is not limited here.

[0092] The two ends of the first controller, the two ends of the network switching unit, and the two ends of the second controller are connected to different switching boards, as shown in Figure 4. The first switch includes multiple switching boards 6, and the different switching boards 6 are connected to each other via a base plate 5. The connection between the first controller and the switching board 6 can be through a PCIe link, an Inter-Integrated Circuit (I2C) link, or a Universal Asynchronous Receiver / Transmitter (UART) link, etc. The connection between the network switching unit and the switching board 6 can be through the SGMII protocol or other links. The connection between the second controller and the switching board 6 can be through the PERSET protocol or the SYS_RST protocol, etc. The second controller can be a programmable logic device, such as a Complex Programmable Logic Device (CPLD) or a Field-Programmable Gate Array (FPGA), etc., which is not limited here. The high-speed interface provided in the switch is used to connect to the storage controller and the first memory.

[0093] In this embodiment, the baseboard within the first switch is connected to the switching units on the switching board, integrating the management unit of the first switch. This unit is responsible for resource scheduling management, device status monitoring, power-on / off management, and overall rack power-on / off management. It provides an ultra-high-performance, ultra-large-scale switching network to meet the topology requirements of storage systems in various application scenarios.

[0094] In some embodiments, each storage controller includes a plurality of first processors;

[0095] Multiple first processors are connected sequentially, and the multiple first processors are in a redundant mode;

[0096] Each first processor is connected to each switching unit.

[0097] Referring to Figure 3, each storage controller 2 includes multiple first processors. These first processors connect via self-testing. This sequential connection only applies to a single storage controller 2; there are no connections between different storage controllers 2. Figure 3 also shows a storage controller 2 containing two first processors, connected in a redundant mode. This redundancy corresponds to a primary / standby configuration, with one primary first processor and one standby first processor. Each first processor is connected to each switching unit, allowing the standby first processor to take over data transmission in case of a failure of the primary first processor.

[0098] The redundant mode presented in this embodiment, which includes multiple first processors and switching units within each storage controller, as well as the connection relationships between the multiple first processors, improves the reliability of the system through primary and backup redundancy.

[0099] In some embodiments, the connection between the host, multiple storage controllers, the first switch, and the first memory is via a fiber optic link;

[0100] The fiber optic link can be either a first protocol fiber optic link or a second protocol fiber optic link.

[0101] It should be noted that the connection between the host, multiple storage controllers, the first switch, and the first memory can be achieved via copper or optical cables. However, copper cable link loss increases with data transmission rate, making long-distance interconnection between resource pools within the storage system difficult, thus limiting resource pool distribution and overall system scale. High-speed copper cable interconnection solutions can only achieve meter-level interconnections. Therefore, this embodiment uses optical fiber links. The type of optical fiber link is not limited; it can be a first-protocol optical fiber link or a second-protocol optical fiber link. Furthermore, the type of optical fiber link between different devices within the storage system is not standardized. However, the type of optical fiber link between two devices is only one: a first-protocol optical fiber link or a second-protocol optical fiber link. The optical fiber link can be a PCIE link or a CXL link, etc., without limitation. That is, the first-protocol optical fiber link can be a PCIE or CXL link, and the second-protocol optical fiber link can be a PCIE or CXL link, as long as the two protocols are different within the same link. By overcoming challenges such as protocol adaptation, clock synchronization locking, and low-speed signal coordination, high-speed optical interconnection technology supporting both PCIE and CXL protocols is achieved, supporting extended link transmission distances greater than 2 meters.

[0102] Fiber optic links utilize optical interconnect technology, a method of transmitting data between components or subsystems within a computer using optical fibers or other optical transmission media. This technology features high bandwidth, high speed, and low latency.

[0103] The optical fiber link provided in this embodiment connects various devices. The optical fiber link has the characteristics of low link loss and long transmission distance. Moreover, it adopts optical interconnection technology, so that the devices do not involve protocol conversion when transmitting signals over long distances, which greatly reduces transmission delay and effectively improves the overall system performance in some application scenarios.

[0104] In some embodiments, when the optical fiber link is a first protocol optical fiber link, the transmitting end and receiving end corresponding to the host, multiple storage controllers, the first switch and the first memory each include an optical module, a second switch and a third controller.

[0105] The optical modules of the transmitting and receiving ends are connected;

[0106] The optical modules of both the transmitting and receiving ends are connected to their respective second switches and third controllers.

[0107] Specifically, each component in the host, multiple storage controllers, the first switch, and the first memory contains a transmitting end for transmitting signals and a receiving end for receiving signals. Both the transmitting and receiving ends include optical modules, a second switch, and a third controller. The optical module is an optoelectronic device primarily used for photoelectric and electro-optical conversion. It plays a crucial role in the connection between the server and the fiber optic network, converting the electrical signals generated by the server into optical signals for transmission via optical fiber. The optical module includes a transmitting section and a receiving section. The transmitting section receives an electrical signal of a certain bit rate, which is processed by an internal drive controller to drive a semiconductor laser (LD) or light-emitting diode (LED) to emit a modulated optical signal at a corresponding rate. It incorporates an automatic optical power control circuit to maintain a stable output optical signal power. The receiving section receives an optical signal of a certain bit rate, which is converted into an electrical signal by a photodetector diode, then amplified by a preamplifier to output an electrical signal of the corresponding bit rate.

[0108] Figure 5 is a schematic diagram of the optical interconnect interface between a transmitter and a receiver provided in an embodiment of this application. As shown in Figure 5, the optical modules 7 of the transmitter and receiver are connected via optical fibers, and each optical module 7 is connected to its own second switch 9 and third controller 8. Here, the second switch 9 is a switch type corresponding to the specific link type of the first protocol optical fiber link, and the third controller 8 is a programmable logic device, such as an FPGA. In Figure 5, the third controller 8 is an FPGA, and the second switch 9 is a PCIe switch.

[0109] The internal structural connections of the transmitting and receiving ends of the various devices provided in this embodiment are transmitted through the transmitting and receiving ends via optical fiber links, thereby improving the data transmission rate and enhancing the performance of the storage system.

[0110] In some embodiments, when the first protocol fiber link is a PCIe protocol link, the corresponding standard for the optical module is not specifically designed for the PCIe protocol. This can lead to incompatibility between some PCIe protocol settings and the standard of conventional optical modules. If the optical module is used directly to build a PCIe optical interconnect path, it will fail. The PCIe interconnect link establishment process involves a link training process. The entire interconnect link must be able to transmit signals correctly during the link training process for the link to be successfully established.

[0111] When the protocol transmitted by the second switch is incompatible with the transmission standard protocol of the optical module and the transmitting end detects a non-compliant receiving end, the second switch at the transmitting end is used to set the first register value to block the detection function of the receiving end.

[0112] Specifically, Figure 6 is a schematic diagram of another communication between a transmitter and a receiver provided in an embodiment of this application. As shown in Figure 6, during the link training process, the transmitter 10 first sends a detection sequence to the receiver 11 to detect the peer device. Since the optical module does not undergo impedance change, the pulse emitted by the transmitter 10 cannot generate a changing current in the optical module, causing the transmitter 10 to fail to detect a compliant receiver 11, thus leading to link establishment failure. This non-compliance could mean that the sequence received by the receiver 11 has no feedback information, or the feedback information cannot be fed back normally, or even that the feedback information is incorrect, etc., which are not limited here. In this embodiment, the second switch of the transmitter 10 supports the configuration of the link training process, and the first register value is set to shield the receiver 11's detection function during the link training process.

[0113] The second switch at the transmitting end provided in this embodiment sets the first register value and uses a shielding function to prevent the receiving end from misidentifying and reporting errors.

[0114] In other embodiments, when the protocol transmitted by the second switch is incompatible with the transmission standard protocol of the optical module and the first protocol fiber optic link is in an electrically idle state, the second switch at the transmitting end is used to set a second register value to reduce the time parameter of the electrically idle state.

[0115] Specifically, the link establishment process involves multiple periods of electrical idle, during which no data signals are transmitted in the interconnect link. When data transmission is needed, the transmitting end sends a corresponding character sequence to wake up the receiving end and continue the link training process. The interconnect link may experience burst transmission states, switching from no data signal transmission to data transmission. The transimpedance amplifier at the optical module receiver needs a period of time to stabilize upon receiving a burst signal. During this period before stabilization, it cannot accurately convert the front-end photocurrent signal into a correct voltage signal, leading to link establishment failure.

[0116] Electrically idle state typically refers to a low-power state of a device or link, during which no data transmission occurs between the transmitting or receiving ends, and the relevant electrical signals are inactive. In this embodiment, by utilizing the feature of the second switch at the transmitting end supporting the configuration of the link training process and setting the value of the second register, the electrically idle state time is significantly shortened.

[0117] The second switch at the transmitting end provided in this embodiment sets a second register value to avoid the problem that noise signals caused by electrical idle state may cause inconsistencies in the working state of the optical modules at the transmitting end and the receiving end, thereby improving the reliability of data transmission.

[0118] In other embodiments, when the protocol transmitted by the second switch is incompatible with the transmission standard protocol of the optical module and the first speed control signal with a speed less than the threshold speed cannot be sent, the third controller at the transmitting end recompiles the second speed control signal with a speed greater than or equal to the threshold speed into the first speed control signal so as to send it to the receiving end through the optical module.

[0119] The third controller at the receiving end is used to interpret the first speed control signal as a second speed control signal after the optical module receives the first speed control signal; wherein, the optical modules at the transmitting end and the receiving end do not contain processors for processing digital signals and processors for clock data for transmitting the second speed control signal with a speed greater than or equal to a threshold speed.

[0120] Specifically, in addition to high-speed data signals, low-speed control signals also exist in the PCIe channel. However, the Digital Signal Processor (DSP) and Clock and Data Recovery (CDR) processor in high-speed optical modules generally only support fixed high-speed signal transmission and cannot transmit low-speed control signals. In this embodiment, the low-speed control signal is a control signal whose first speed control signal is less than a threshold speed. The transmitting end processes the low-speed control signal through a third controller, recompiling it into a high-speed control signal (such as a 600MHz high-speed signal), which is then transmitted to the receiving end through the optical module and fiber optic link. Subsequently, the third controller at the receiving end recompiles (interprets) the high-speed control signal into a low-speed control signal.

[0121] Furthermore, since the DSP and CDR processors in traditional optical modules only support high-speed signals (high-speed signals are second-speed control signals with speeds greater than or equal to a threshold speed), it should be noted that the DSP is a processor for processing digital signals, and the CDR is a processor for processing clock data. In this embodiment, a linear direct-drive optical module is used, eliminating the DSP and CDR controllers, which means that high-speed and low-speed signals are not controlled.

[0122] The optical module provided in this embodiment removes the DSP and CDR controllers, and processes low-speed control signals through a third controller to achieve the function of transmitting low-speed control signals. Furthermore, the internal structural changes of the optical module in this embodiment avoid the loss of lock-up of the DSP and CDR controllers during link rate switching, thus enabling rate switching during the link rate negotiation process.

[0123] In some embodiments, the protocol buses supported by the plurality of first processors within each memory controller are not limited. The plurality of first processors within each memory controller support a first protocol bus and a second protocol bus, wherein each first processor provides a first protocol bus and a second protocol bus to the outside world for use in first memory expansion, respectively.

[0124] Figure 7 is a schematic diagram of the overall layout of another storage system provided in an embodiment of this application. As shown in Figure 7, the storage controller 2 is connected to the first switch 3, and the first switch 3 is connected to the first memory 4. The first memory 4 includes dual in-line memory modules (DIMMs) and SSD hard drives. A BMC monitoring and management module and a device status monitoring module are simultaneously installed within the storage controller 2 to achieve optimal utilization of compact space. A storage controller 2 includes multiple first processors and provides IO expansion modules to meet the topology requirements of various application scenarios.

[0125] Table 2 shows the specifications of the storage controller. As shown in Table 2, taking a storage controller with two primary processors as an example, both support 8 x16 bandwidth PCIe buses, of which 4 x16 are compatible with CXL buses. The interconnection of the two primary processors occupies 2 x16 bandwidth PCIe buses and 2 x16 bandwidth CXL buses, which are used for network card / hard disk expansion and CXL network expansion, respectively, and support the implementation of CXL memory expansion within the storage system.

[0126] Table 2 Specifications of Storage Controllers

[0127] In Table 2, the E3.S and U.2 specifications indicate hard drives with standard form factors.

[0128] The multiple first processors provided in this embodiment support a first protocol bus and a second protocol bus for first memory expansion, meeting the topology requirements of multiple application scenarios and improving the reliability of the storage system.

[0129] In some embodiments, the first memory includes at least a first data storage device and / or a second data storage device;

[0130] The first data storage device is used to support memory module expansion;

[0131] A second data storage device is used to support hard drive expansion.

[0132] As shown in Figure 7, the first memory 4 includes at least a first data storage device and / or a second data storage device to support DIMM (memory module) expansion and SSD hard disk expansion, respectively.

[0133] The multiple data storage devices within the first memory provided in this embodiment enhance the diversity and flexibility of expansion.

[0134] In some embodiments, the first data storage device includes a memory expansion board and a first substrate; wherein the memory expansion board includes a first connector, a plurality of first interfaces and a plurality of memory controllers;

[0135] The first end of each first interface is connected to the first switch;

[0136] The second end of each first interface is connected to a first preset number of memory controllers;

[0137] The memory controller's memory interface is connected to a second preset number of memory modules;

[0138] The communication interfaces of the memory controller are all connected to the first end of the first connector;

[0139] The second end of the first connector is connected to the first substrate.

[0140] Specifically, Figure 8 is a schematic diagram of the structure of a first data storage device provided in an embodiment of this application. As shown in Figure 8, the memory resource pool is interconnected with a high-performance first switch downlink high-speed interface (first interface) via a high-speed connection, including a memory expansion board and a first base plate. The memory expansion board includes multiple first interfaces, a first connector, and multiple memory controllers. The first end of each first interface is connected to the first switch, and the second port of each first interface is connected to a first preset number of memory controllers. In Figure 8, the high-speed link of each first interface can be connected through two memory controllers. Through the memory interface of the memory controller, a second preset number of memory modules are connected. In Figure 8, one memory controller connects to two memory modules to realize remote memory expansion. The communication interfaces of the memory controllers are all connected to the first end of the first connector. The communication interface here can be any type of communication interface, such as an Improved Inter-Integrated Circuit (I3C) interface, etc., which is not limited here. The second end of the first connector is connected to the first base plate. It should be noted that the first substrate includes programmable logic devices and a control platform, which are of the same type as the two devices on the substrate in the above embodiments. The memory expansion board mainly expands memory through a memory controller. The entire device in Figure 8 can be expanded externally with 64 memory modules, providing high-density memory resource expansion. The first substrate mainly carries the control platform and programmable logic devices, used for internal logic control and status monitoring and management of the unit. Information status is reported through the management network.

[0141] Table 3 shows the specifications of the first data storage device. As shown in Table 3, it supports 64 memory modules for expansion. Based on a single memory module capacity of 256GB, the maximum expandable memory capacity of a single machine can reach 16TB.

[0142] Table 3 Specifications of the First Data Storage Device

[0143] The power supplies in Table 3 are provided in a redundant mode.

[0144] The first data storage device provided in this embodiment supports memory module expansion and provides high-density memory resource expansion to achieve remote memory expansion.

[0145] In other embodiments, the second data storage device includes a hard disk expansion board and a second substrate; wherein the hard disk expansion board includes a plurality of second connectors, a plurality of signal conditioning cards and a plurality of second interfaces;

[0146] The third preset number of hard drives are respectively connected to the first end of the second connector;

[0147] The second end of the second connector is connected to the first end of the signal conditioning card;

[0148] The second end of the signal conditioning card is connected to the first end of the corresponding second interface; and the third end of each card is connected to the second substrate.

[0149] The second end of the second interface is connected to the first switch.

[0150] Specifically, Figure 9 is a schematic diagram of a second data storage device provided in an embodiment of this application. As shown in Figure 9, the second data storage device includes a hard disk expansion board and a second base plate. The hard disk expansion board includes multiple second connectors, multiple signal conditioning cards, and multiple second interfaces. A third preset number of hard disks are respectively connected to the first end of the second connectors. Referring to Figure 9, every two hard disks are connected to one second connector, and the second end of each second connector is connected to the first end of a signal conditioning card. The second end of the signal conditioning card is connected to the first end of the corresponding second interface; and the third end of each is connected to the second base plate; the second end of the second interface is connected to a first switch. The signal conditioning card is an electronic device used to process and adjust signals. Its main function is to convert the raw signals acquired from sensors or other signal sources into standard signals suitable for processing by the data acquisition system. These conditioning operations include, but are not limited to, amplification, filtering, isolation, modulation and demodulation. The signal conditioning card can be installed in various data acquisition systems or testing equipment to ensure signal quality and reliability, and improve the overall performance and accuracy of the system.

[0151] It should be noted that the second substrate has the same type of device as the first component in the above embodiment, including programmable logic devices and a control platform. The second substrate mainly carries the control platform and programmable logic devices, and is used for internal logic control, status monitoring and management of the unit. The information status is reported through the management network.

[0152] Table 4 shows the specifications of the second data storage device. As shown in Table 4, it supports 12 memory hard disk expansions and the high-speed interface supports 12 x16 bandwidth high-speed signals.

[0153] Table 4 Specifications of the Second Data Storage Device

[0154] In the second data storage device provided in this embodiment, it is interconnected with a high-performance first switch via a high-speed cable. After the high-speed signal is introduced through the high-speed interface (second interface), it is conditioned and optimized by the signal conditioning card to improve the quality of the high-speed link signal and improve the system reliability. After that, the conditioned and optimized high-speed signal is connected to the SSD hard disk storage module through the high-speed connector to realize SSD storage expansion.

[0155] Furthermore, this application also provides a cluster node, including at least one of the storage systems mentioned in the above embodiments. Figure 10 is a schematic diagram of the structure of a cluster node provided in an embodiment of this application. As shown in Figure 10, it includes at least one of the above-mentioned storage systems.

[0156] When there are multiple storage systems, they are connected by fiber optic links.

[0157] It should be noted that a cluster node may include one storage system or multiple storage systems. In a cluster node consisting of multiple storage systems, the storage systems are connected to each other through fiber optic links to achieve east-west interconnection.

[0158] The cluster's write data flow and metadata control flow ensure performance and data redundancy requirements, while the cluster message control flow meets manageability requirements. Traditional storage systems within cluster nodes use network interconnection for east-west connections, resulting in significant latency and limited bandwidth. This embodiment employs fiber optic links for optical interconnection, guaranteeing high bandwidth, low latency, and high reliability.

[0159] For an introduction to the cluster node provided in this application, please refer to the above-described storage system embodiments. This application will not repeat the details here, as it has the same beneficial effects as the above-described storage system.

[0160] Furthermore, this application also provides a cluster system, including at least one cluster node;

[0161] When there are multiple cluster nodes, the cluster nodes are connected through fiber optic links.

[0162] Specifically, referring to Figure 10, optical interconnection technology between multiple cluster nodes (represented as cluster nodes E1' to EN' in Figure 10) is realized, which has the same connection effect as the optical fiber link within the cluster node in the above embodiment. This embodiment achieves high bandwidth, low latency and high reliability in the transmission between cluster nodes.

[0163] For an introduction to the cluster system provided in this application, please refer to the above-described storage system embodiments. This application will not repeat the details here, as it has the same beneficial effects as the above-described storage system.

[0164] Furthermore, this application also provides a storage resource scheduling method based on a storage system, applied to a first switch of the storage system. The storage system includes a host, a first switch, a first memory, and multiple storage controllers; the host is connected to the multiple storage controllers; the multiple storage controllers are all directly connected to the first switch; the first switch is connected to the first memory; the first switch includes a base plate and multiple switching units; each storage controller is connected to each switching unit; each switching unit is located on a switching board, and the switching boards are connected to each other through the base plate to realize the connection between the switching units; each switching unit is connected to its corresponding first memory; Figure 11 is a flowchart of a storage resource scheduling method based on a storage system provided by an embodiment of this application. As shown in Figure 11, the method includes:

[0165] S11: Receive storage resources sent by multiple storage controllers;

[0166] Storage resources are determined by the data requests sent by the host within the storage controller.

[0167] S12: Obtain the resources of the first memory;

[0168] S13: Allocate and schedule the resources of the first memory according to the storage resources of multiple storage controllers to achieve storage resource sharing of the first memory.

[0169] Specifically, the first switch is used to decouple the resources of the internal pooling system, and provides an interactive interface through the management engine to enable unified management of devices in the resource pool, and complete the identification, management and dynamic allocation of resources.

[0170] Figure 12 is a schematic diagram of an internal pooling system in which a first switch is located, according to an embodiment of this application. As shown in Figure 12, the management engine of the first switch implements the following functions:

[0171] 1. Resource identification and asset information display.

[0172] The resource management engine uses a dedicated communication protocol interface to obtain all storage information from the CXL switch to complete resource identification. Simultaneously, the upper-layer management software provides a network application programming interface (API). In addition to the network API, the resource management engine also provides a visual interface, allowing users to access the management engine through a browser to view the resource list and system topology. Figure 13 is a schematic diagram of information display under an asset management module provided in an embodiment of this application. As shown in Figure 13, the visual interface provides both list display and graphical display methods. The list display mainly includes: a device resource list (device asset information, device partition information, device allocation relationship), a computing unit list (computing unit basic information, computing unit location information, computing unit resource list), and a port resource list (port basic attributes, port configuration information, port reserved resource information).

[0173] Storage resources mainly include: device manufacturer information, device type, device partition information, and resource allocation status. The port resource list mainly includes: basic port attributes, port resource reservation status, and port allocation relationships.

[0174] The management engine acquires and displays information about the storage server controller, primarily including: the location of the storage server controller; the status of resources allocated to the storage server controller; and the health status of the storage server controller's device links. The management engine provides a graphical display and management of resources through a web graphical user interface (WEB GUI). Users can dynamically allocate resources in the interconnect topology view and monitor device links in real time by dragging and dropping elements.

[0175] 2. Equipment partitioning and dynamic allocation information.

[0176] Figure 14 is a schematic diagram of resource allocation and scheduling provided in an embodiment of this application. As shown in Figure 14, the resource management engine is connected to each storage server controller via a network and interconnected with the resource management controller on each storage controller. The resource management software on each storage controller is responsible for monitoring the resource usage of its own storage system, while the resource management system engine is responsible for the resource monitoring and scheduling of the entire system.

[0177] Device partitioning information refers to the allocation of resources. A storage controller can have multiple memory modules connected to it, and all resources under the same storage controller can be arbitrarily divided into several intervals. The specific partitioning can be perceived and monitored by the resource management system, which provides a visual and interactive way to adjust device partitions.

[0178] Dynamic resource allocation refers to the resource management engine's ability to dynamically migrate storage resources from one computing unit to another while the entire operating system remains online. This eliminates the need to restart computing units and ensures uninterrupted operation of ongoing services. The entire process can be monitored in real-time via the management engine's web interface. These features are based on the hot-swappable nature of the CXL device.

[0179] The specific method for identifying and processing storage resources in the storage controller in this embodiment is not limited. It can be based on a specific identification algorithm or implemented through a certain mapping relationship, etc.

[0180] For allocation and scheduling, there are no restrictions. It can monitor in real time to make the remaining resources in the first memory available to the host, or it can be based on a certain model algorithm to calculate the selection of the remaining resources under multiple first memories. There are no restrictions here.

[0181] In some embodiments, resource allocation and scheduling of the first memory based on the storage resources of multiple storage controllers includes:

[0182] Obtain the target storage resources from the target storage controller;

[0183] A mapping relationship between the target storage controller and the first memory is established in advance;

[0184] Determine whether the remaining resources of the first memory corresponding to the target memory controller are greater than the target memory resources based on the mapping relationship;

[0185] If the remaining resources of the first memory corresponding to the target storage controller are greater than or equal to the target storage resources, then the first memory to which the remaining resources are greater than or equal to the target storage resources belong will be scheduled.

[0186] If the remaining resources of the first memory corresponding to the target storage controller are less than the target storage resources, the remaining resources of the reserved free first memory will be scheduled to the target storage controller to facilitate the allocation of target storage resources by the target storage controller.

[0187] Specifically, a mapping relationship is pre-established between the storage controller and the first memory. This can be established between one target storage controller and all first memories, or between one target storage controller and multiple first storage controllers, to facilitate subsequent scheduling directly based on the mapping relationship through the first switch. Before scheduling, it is necessary to determine whether the resources of the first memory corresponding to the target storage controller with which the mapping relationship is established are greater than the target storage resources. Here, the target storage resources are the storage resources required by the target storage controller. If the remaining resources of the current first memory are greater than or equal to the target storage resources, the current first memory to which the remaining resources belong will be scheduled. If the remaining resources of the current first memory are less than the target storage resources, then the remaining resources of the reserved free first memory need to be scheduled. This scheduling can be done by scheduling the entire first memory with remaining resources greater than or equal to the target storage resources, which only requires one scheduling. Alternatively, multiple first memories with remaining resources in other reserved spaces can be scheduled if they are less than the target storage resources. In this case, the scheduling process will involve multiple scheduling operations, but it can make full use of the remaining resources in other reserved spaces.

[0188] The specific process of storage resource allocation and scheduling provided in this embodiment improves the accuracy of scheduling. At the same time, it prioritizes the scheduling of the resources of the first memory corresponding to the target storage controller with a mapping relationship. When the remaining resources of the corresponding first memory are less than the target storage resources, the remaining resources of the reserved idle first memory need to be scheduled, thereby improving the reliability of scheduling.

[0189] In other embodiments, it also includes:

[0190] The interface displays the allocation and scheduling of storage resources and storage resource information; the list corresponding to the storage resource information includes at least a device resource list, a host resource list, and a port resource list.

[0191] Correspondingly, the allocation and scheduling process is implemented through a touch screen interface.

[0192] Specifically, referring to Figure 13, the allocation and scheduling information of storage resources and storage resource information are displayed through an interface, showing each scheduling process in real time. The storage resource information lists mentioned in the above embodiments include at least a device resource list, a host resource list, and a port resource list. Scheduling can be performed via touchscreen on the interface.

[0193] The touchscreen method provided in this embodiment implements the scheduling and allocation process, making the entire scheduling and allocation process more intuitive and simple, facilitating user management and scheduling, and improving the user experience.

[0194] This embodiment provides a storage resource scheduling method based on a storage system. It receives storage resources sent by multiple storage controllers. Storage resources are determined by data requests sent by the host within the storage controller. Resources of a first memory are acquired. The resources of the first memory are allocated and scheduled according to the storage resources of the multiple storage controllers to achieve storage resource sharing of the first memory. By separating the first memory (media) and the host (computing unit), i.e., by distancing the local disk and memory within the storage system, decoupling computation and media, different performance and capacity memory spaces can be configured according to application needs. Multiple storage controllers are directly connected to a first switch, without interaction between them. This eliminates the need for multiple storage controller control interactions and cross-layer memory relocation operations. The first memory can be directly accessed through the first switch, saving storage controller computing power, network, and memory channel resources, reducing latency, and improving resource utilization. Furthermore, the computing unit and the first memory are independent, allowing for on-demand expansion and contraction and independent upgrades, saving costs. Further, it reduces interaction and usage between storage controllers, achieving cross-controller memory and storage resource sharing, improving memory bandwidth and storage system performance. Finally, each switching unit is connected to its respective first memory, realizing the connection between the first memories under each switching unit within the first switch. This reduces the number of branches of the first memory corresponding to each switching unit, thus preventing the data transmission rate reduction problem. The connection relationship between the base plate of the first switch and multiple switching units is the key to achieving hardware resource decoupling and reconstruction. Through the combination of high-performance, large-scale switching I / O and basic management software, flexible switching between various topologies of the storage server system is achieved, and resources are scheduled and allocated on demand to meet the memory expansion needs of multiple storage application scenarios.

[0195] The above describes in detail various embodiments of the storage resource scheduling method based on a storage system. Based on this, this application also discloses a storage resource scheduling device based on a storage system corresponding to the above method, applied to a first switch in a storage system. The storage system includes a host, a first switch, a first memory, and multiple storage controllers. The host is connected to the multiple storage controllers; each of the multiple storage controllers is directly connected to the first switch; the first switch is connected to the first memory; the first switch includes a base plate and multiple switching units; each storage controller is connected to each switching unit; each switching unit is located on a switching board, and the switching boards are connected through the base plate to achieve the connection between the switching units; each switching unit is connected to its corresponding first memory. Figure 15 is a structural diagram of a storage resource scheduling device based on a storage system provided in an embodiment of this application. As shown in Figure 15, the storage resource scheduling device based on a storage system includes:

[0196] The receiving module 12 is used to receive storage resources sent by multiple storage controllers; wherein the storage resources are identified and determined by the data request sent by the host within the storage controller.

[0197] Acquisition module 13 is used to acquire resources of the first memory;

[0198] The allocation and scheduling module 14 is used to allocate and schedule the resources of the first memory according to the storage resources of multiple storage controllers, so as to realize the sharing of storage resources of the first memory.

[0199] Since the embodiments of the device part correspond to the embodiments described above, please refer to the embodiments described in the method part for the embodiments of the device part, and will not be repeated here.

[0200] For a description of the storage resource scheduling device based on the storage system provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above storage resource scheduling method based on the storage system.

[0201] Figure 16 is a structural diagram of a storage resource scheduling device based on a storage system provided in an embodiment of this application. As shown in Figure 16, the device includes:

[0202] The second memory 21 is used to store computer programs;

[0203] The second processor 22 is used to implement the steps of the storage resource scheduling method based on the storage system when executing computer programs.

[0204] The storage resource scheduling device based on the storage system provided in this embodiment may include, but is not limited to, smartphones, tablets, laptops, or desktop computers.

[0205] The second processor 22 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The second processor 22 may be implemented using at least one hardware form selected from Digital Signal Processor (DSP), FPGA, and Programmable Logic Array (PLA). The second processor 22 may also include a main processor and a coprocessor. The main processor, also known as the CPU, is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the second processor 22 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the second processor 22 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.

[0206] The second memory 21 may include one or more non-transitory computer-readable storage media, which may be non-transitory. The second memory 21 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the second memory 21 is used to store at least the following computer program 211, which, after being loaded and executed by the second processor 22, can implement the relevant steps of the storage resource scheduling method based on the storage system disclosed in any of the foregoing embodiments. In addition, the resources stored in the second memory 21 may also include an operating system 212 and data 213, etc., and the storage method may be temporary storage or permanent storage. The operating system 212 may include Windows, Unix, Linux, etc. The data 213 may include, but is not limited to, the data involved in the storage resource scheduling method based on the storage system.

[0207] In some embodiments, the storage resource scheduling device based on the storage system may further include a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27.

[0208] Those skilled in the art will understand that the structure shown in Figure 16 does not constitute a limitation on storage resource scheduling devices based on storage systems and may include more or fewer components than shown.

[0209] The second processor 22 implements the storage resource scheduling method based on the storage system provided in any of the above embodiments by calling instructions stored in the second memory 21.

[0210] It should be noted that the second memory and the second processor in this embodiment are different from the first memory and the first processor in the above embodiments.

[0211] For a description of the storage resource scheduling device based on the storage system provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above storage resource scheduling method based on the storage system.

[0212] Furthermore, this application also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by the second processor 22, it implements the steps of the storage resource scheduling method based on the storage system described above.

[0213] It is understood that if the methods in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in a non-transitory computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0214] For an introduction to the non-transitory computer-readable storage medium provided in this application, please refer to the above method embodiments. This application will not repeat the details here, but it has the same beneficial effects as the above-described storage resource scheduling method based on the storage system.

[0215] Furthermore, this application also provides a computer program product, including a computer program / instructions that, when executed by a second processor, implement the steps of the above-described storage resource scheduling method based on a storage system.

[0216] For an introduction to the computer program product provided in this application, please refer to the above method embodiments. This application will not repeat the details here, but it has the same beneficial effects as the above-described storage resource scheduling method based on the storage system.

[0217] The foregoing has provided a detailed description of a storage system, cluster node, system, storage resource scheduling method, and apparatus provided in this application. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

[0218] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

Claims

1. A storage system, characterized in that, Includes a host, a first switch, a first memory, and multiple memory controllers; The host connects to multiple storage controllers; All of the aforementioned storage controllers are directly connected to the first switch; Multiple storage controllers are configured to identify and process data requests sent by the host to determine the corresponding storage resources; The first switch is connected to the first memory; The first switch is used to receive storage resources from multiple storage controllers and to allocate and schedule the resources of the first memory according to the storage resources of the multiple storage controllers to realize storage resource sharing of the first memory. The first switch includes a base plate and multiple switching units; Each of the storage controllers is connected to each of the switching units; Each of the switching units is located on a switching board, and the switching boards are connected to each other via a substrate to achieve the connection between the switching units; as well as Each of the aforementioned switching units is connected to its corresponding first memory.

2. The storage system according to claim 1, characterized in that, The substrate includes a first controller, a network switching unit, a control platform, and a second controller; The two ends of the first controller, the two ends of the network switching unit, and the two ends of the second controller are connected to different switching boards; and The first controller, the network switching unit, and the control platform are connected in sequence.

3. The storage system according to claim 1, characterized in that, Each of the storage controllers includes a plurality of first processors; The plurality of first processors are sequentially connected, and the plurality of first processors are in a redundant mode; and Each of the first processors is connected to each of the aforementioned switching units.

4. The storage system according to claim 1, characterized in that, The host, the plurality of storage controllers, the first switch, and the first memory are connected via fiber optic links; and The optical fiber link can be either a first protocol optical fiber link or a second protocol optical fiber link.

5. The storage system according to claim 4, characterized in that, When the optical fiber link is a first protocol optical fiber link, the transmitting end and receiving end of the host, the plurality of storage controllers, the first switch and the first memory each include an optical module, a second switch and a third controller; The optical modules of the transmitting end and the receiving end are connected to each other; The optical modules of the transmitting end and the receiving end are connected to the second switch and the third controller of the transmitting end and the receiving end, respectively.

6. The storage system according to claim 5, characterized in that, When the protocol transmitted by the second switch is incompatible with the transmission standard protocol of the optical module and the transmitting end detects a non-compliant receiving end, the second switch of the transmitting end is used to set a first register value to block the detection function of the receiving end.

7. The storage system according to claim 5, characterized in that, When the protocol transmitted by the second switch is incompatible with the transmission standard protocol of the optical module and the first protocol fiber optic link is in an electrically idle state, the second switch at the transmitting end is used to set a second register value to reduce the time parameter of the electrically idle state.

8. The storage system according to claim 5, characterized in that, When the protocol transmitted by the second switch is incompatible with the transmission standard protocol of the optical module and the first speed control signal with a speed less than the threshold cannot be sent, the third controller of the transmitting end will recompile the second speed control signal with a speed greater than or equal to the threshold into the first speed control signal so as to send it to the receiving end through the optical module. The third controller at the receiving end is used to interpret the first speed control signal as the second speed control signal after the optical module receives the first speed control signal; wherein, the optical modules at the transmitting end and the receiving end do not contain a processor for processing digital signals and a processor for clock data for transmitting the second speed control signal at a speed greater than or equal to a threshold speed.

9. The storage system according to claim 3, characterized in that, Each of the first processors within the memory controller supports a first protocol bus and a second protocol bus, wherein each first processor provides the first protocol bus and the second protocol bus to the outside world for first memory expansion, respectively.

10. The storage system according to claim 1, characterized in that, The first memory includes at least a first data storage device and / or a second data storage device; The first data storage device is used to support memory module expansion; and The second data storage device is used to support hard disk expansion.

11. The storage system according to claim 10, characterized in that, The first data storage device includes a memory expansion board and a first base plate; wherein the memory expansion board includes a first connector, a plurality of first interfaces and a plurality of memory controllers; The first end of each of the first interfaces is connected to the first switch; The second end of each of the first interfaces is connected to a first preset number of memory controllers; The memory interface of the memory controller is connected to a second preset number of memory modules; The communication interfaces of the memory controllers are all connected to the first end of the first connector; and The second end of the first connector is connected to the first substrate.

12. The storage system according to claim 10, characterized in that, The second data storage device includes a hard disk expansion board and a second base plate; wherein the hard disk expansion board includes multiple second connectors, multiple signal conditioning cards and multiple second interfaces; The third preset number of hard drives are respectively connected to the first end of each of the second connectors; The second end of each of the second connectors is connected to the first end of each of the signal conditioning cards; The second end of each signal conditioning card is connected to the first end of the corresponding second interface; and the third end of each signal conditioning card is connected to the second substrate; and The second end of the second interface is connected to the first switch.

13. A cluster node, characterized in that, Includes at least one storage system as described in any one of claims 1 to 12; When there are multiple storage systems, the multiple storage systems are connected by fiber optic links.

14. A cluster system, characterized in that, Includes at least one cluster node as described in claim 13; When there are multiple cluster nodes, the multiple cluster nodes are connected by fiber optic links.

15. A storage resource scheduling method based on a storage system, characterized in that, A first switch is applied to a storage system, the storage system including a host, a first switch, a first memory, and multiple storage controllers; the host is connected to the multiple storage controllers; each of the multiple storage controllers is directly connected to the first switch; the first switch is connected to the first memory; the first switch includes a base plate and multiple switching units; each storage controller is connected to each of the switching units; each switching unit is located on a switching board, and the switching boards are connected to each other through the base plate to realize the connection between the switching units; Each of the aforementioned switching units is connected to its corresponding first memory; the method includes: Receive storage resources sent by a plurality of the storage controllers; wherein the storage resources are determined by the identification processing within the storage controller based on data requests sent by the host; Acquire the resources of the first memory; and The storage resources of the first memory are allocated and scheduled according to the storage resources of the multiple storage controllers to achieve storage resource sharing of the first memory.

16. The storage resource scheduling method based on a storage system according to claim 15, characterized in that, The process of allocating and scheduling resources of the first memory based on the storage resources of the multiple storage controllers includes: Obtain the target storage resources from the target storage controller; A mapping relationship between the target storage controller and the first memory is established in advance; Based on the mapping relationship, determine whether the remaining resources of the first memory corresponding to the target storage controller are greater than the target storage resources; In response to the fact that the remaining resources of the first memory corresponding to the target storage controller are greater than or equal to the target storage resources, the first memory to which the remaining resources are greater than or equal to the target storage resources are scheduled; In response to the fact that the remaining resources of the first memory corresponding to the target storage controller are less than the target storage resources, the remaining resources of the reserved free first memory are scheduled to the target storage controller to facilitate the allocation of the target storage resources by the target storage controller.

17. The storage resource scheduling method based on a storage system according to claim 15, characterized in that, The method further includes: The interface displays the allocation and scheduling of storage resources and storage resource information; wherein the list corresponding to the storage resource information includes at least a device resource list, a host resource list, and a port resource list; and Correspondingly, the allocation and scheduling process is implemented through a touch screen interface.

18. A storage resource scheduling device based on a storage system, characterized in that, A first switch is applied to a storage system, the storage system including a host, a first switch, a first memory, and multiple storage controllers; the host is connected to the multiple storage controllers; each of the multiple storage controllers is directly connected to the first switch; the first switch is connected to the first memory; the first switch includes a base plate and multiple switching units; each storage controller is connected to each of the switching units; each switching unit is located on a switching board, and the switching boards are connected to each other through the base plate to realize the connection between the switching units; Each of the switching units is connected to its corresponding first memory; the device includes: A receiving module is configured to receive storage resources sent by a plurality of the storage controllers; wherein the storage resources are determined by the identification processing within the storage controller based on data requests sent by the host; An acquisition module is used to acquire the resources of the first memory; The allocation and scheduling module is used to allocate and schedule the resources of the first memory according to the storage resources of the multiple storage controllers, so as to realize the sharing of storage resources of the first memory.

19. A storage resource scheduling device based on a storage system, characterized in that, include: The second memory is used to store computer programs; as well as The second processor is configured to implement the steps of the storage resource scheduling method based on a storage system as described in any one of claims 15 to 17 when executing the computer program.

20. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores a computer program, which, when executed by a second processor, implements the steps of the storage resource scheduling method based on a storage system as described in any one of claims 15 to 17.

21. A computer program product comprising a computer program or computer-readable instructions, characterized in that, When the computer program or computer-readable instructions are executed by the second processor, they implement the steps of the storage resource scheduling method based on the storage system as described in any one of claims 15 to 17.

Citation Information

Patent Citations

  • Storing method, storing system and controller

    CN101763221A

  • Method for designing multi-control storage system

    CN103152397A

  • Server

    CN104657316A

  • Storage framework as well as initialization method, data storage method and data storage and management apparatus therefor

    CN105739930A

  • Server system, resource scheduling method of server system, chip and chip grain

    CN118210634A