Storage controller, read / write method, and server

By configuring multiple dies in the storage controller to implement different levels of storage access paths, the performance degradation problem caused by cache synchronization between dies in a multi-CPU, multi-die architecture is solved, achieving more efficient I/O performance and resource utilization.

CN121364829BActive Publication Date: 2026-03-24INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In a multi-CPU, multi-die architecture, the performance degradation of storage data I/O is mainly caused by the overhead of inter-die cache synchronization, which cannot be effectively solved by existing software implementations.

Method used

By configuring multiple dies in the storage controller to implement the host interaction layer, cache layer, redundant disk array layer and disk management layer respectively, and communicating in an orderly manner according to a fixed storage access path, the distribution module is avoided. The entire I/O function is implemented in a pipeline manner, reducing the communication overhead between dies and cache synchronization overhead.

Benefits of technology

It improves the overall I/O performance of the storage controller, avoids the communication overhead and cache synchronization problems of the distribution module, and improves CPU resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121364829B_ABST
    Figure CN121364829B_ABST
Patent Text Reader

Abstract

The application provides a storage controller, a read-write method and a server, and can be applied to the technical field of storage. The storage controller comprises at least one multi-core processor, each of the at least one multi-core processor comprising a plurality of processing units connected to each other; the at least one multi-core processor is configured to: sequentially control the plurality of processing units on a storage access path, receive a data transmission instruction and perform data read-write on a disk according to the data transmission instruction; wherein the plurality of processing units on the storage access path are configured to: respectively implement orderly communication in a host interaction layer, a cache layer, a redundant array of independent disks layer and a disk management layer according to a processing sequence of the data transmission instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technology, specifically to a storage controller, a data read / write method, and a server. Background Technology

[0002] Unified Storage (US) is a storage solution that integrates Network-Attached Storage (NAS) and Storage Area Network (SAN) architectures. It supports multiple protocols, such as Fibre Channel (FC), Internet Small Computer System Interface (iSCSI), Network File System (NFS), and Common Internet File System (CIFS), enabling unified management of block-level, file-level, and object-level data access.

[0003] Storage performance is typically closely related to the Central Processing Unit (CPU). With the continuous development of CPU technology, multi-CPU, multi-die architectures have become increasingly common. For example, multiple CPUs and their internal dies can be used to implement data output / input. However, in multi-CPU, multi-die architectures, implementing data I / O via software requires multiple dies within the CPU, which can lead to performance issues caused by cache synchronization between dies. Summary of the Invention

[0004] In view of the above problems, this application provides a storage controller, a read / write method, and a server.

[0005] According to a first aspect of this application, a storage controller is provided, including at least one multi-core processor, each of the at least one multi-core processor including a plurality of interconnected processing units; the at least one multi-core processor is configured to: sequentially control the plurality of processing units located on a storage access path, receive data transmission instructions, and perform data read and write operations on the disk according to the data transmission instructions; wherein the plurality of processing units located on the storage access path are configured to: implement ordered communication through a host interaction layer, a cache layer, a redundant array of disks layer, and a disk management layer, respectively, according to the processing order of the data transmission instructions.

[0006] The second aspect of this application provides a data read / write method applied to a server, the server including the storage controller as described above, the read / write method including: at least one multi-core processor in the storage controller sequentially controls multiple processing units located on the storage access path, receiving data transmission instructions and performing data read / write on the disk according to the data transmission instructions; wherein, the multiple processing units located on the storage access path are configured to: achieve ordered communication according to the processing order of the data transmission instructions, respectively through a host interaction layer, a cache layer, a disk redundant array layer and a disk management layer.

[0007] A third aspect of this application provides a server, comprising: a disk; a storage controller including at least one multi-core processor, each of the at least one multi-core processor including a plurality of interconnected processing units; the at least one multi-core processor sequentially controls the plurality of processing units located on a storage access path, receiving data transmission instructions and performing data read and write operations on the disk according to the data transmission instructions; wherein the plurality of processing units located on the storage access path are configured to: achieve ordered communication through a host interaction layer, a cache layer, a disk redundant array layer and a disk management layer, respectively, according to the processing order of the data transmission instructions.

[0008] In the embodiments of this application, multiple dies are configured to respectively implement the host interaction layer, cache layer, redundant array of disks layer, and disk management layer. These dies communicate in an ordered manner according to a fixed storage access path, enabling the storage controller to implement the entire I / O function through a pipelined approach using multiple dies. On one hand, the storage controller of this application does not require a distribution module, avoiding communication between the dies to which the distribution module belongs and other dies, thereby avoiding cache synchronization overhead caused by acquiring and storing hardware resource information and allocating appropriate I / O to each die based on that information. On the other hand, since the entire I / O function is fixed to be implemented by only a subset of dies, strict isolation between the functions of different dies avoids cache synchronization problems caused by multiple dies jointly implementing the same function, thus reducing the overhead of data cache synchronization caused by communication between modules with the same function. Attached Figure Description

[0009] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0010] Figure 1A An example of cross-CPU communication based on CPU I / O is shown.

[0011] Figure 1B An example of implementing I / O using a virtual node architecture is shown.

[0012] Figure 1C This illustrates the first example of a virtual node architecture implementing I / O through a distribution module.

[0013] Figure 1D This illustrates a second example of how a virtual node architecture implements I / O through a distribution module.

[0014] Figure 1E This illustrates an example of how a virtual node architecture implements I / O through multiple instances.

[0015] Figure 2 A schematic diagram of the architecture of a storage controller according to an embodiment of this application is shown.

[0016] Figure 3 A schematic diagram of the architecture of a multi-core processor according to an embodiment of this application is shown.

[0017] Figure 4 A schematic diagram showing the relationship between three write types and stripes according to embodiments of this application is provided.

[0018] Figure 5 A flowchart illustrating the operation of lowercase types according to an embodiment of this application is shown.

[0019] Figure 6 A flowchart illustrating the operation of a lowercase type according to a specific embodiment of this application is shown.

[0020] Figure 7 A flowchart illustrating the operation of a full-write type according to an embodiment of this application is shown.

[0021] Figure 8 A schematic diagram of a scenario using uppercase types according to an embodiment of this application is shown.

[0022] Figure 9 A flowchart illustrating a read / write method according to an embodiment of this application is shown. Detailed Implementation

[0023] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0026] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0027] Figure 1A An example of cross-CPU communication based on CPU I / O is shown. Figure 1B An example of implementing I / O using a virtual node architecture is shown. Figure 1C This illustrates the first example of a virtual node architecture implementing I / O through a distribution module. Figure 1D This illustrates a second example of how a virtual node architecture implements I / O through a distribution module. Figure 1E This illustrates an example of how a virtual node architecture implements I / O through multiple instances.

[0028] The development of CPUs follows two clear lines: the exponential growth in the number of cores and the fundamental shift in architecture design from single-chip to multi-die integration. This evolution not only reflects the progress of semiconductor technology but also embodies the architectural innovations made to overcome performance bottlenecks. The explosive growth in the number of cores is the most significant feature of CPU development. However, in the early 21st century, when single-core performance stagnated due to power consumption and frequency limitations, multi-core became an option. From the initial explorations of dual-core and quad-core CPUs to today's mainstream server CPUs equipped with 64 or even 96 cores, the number of cores has almost doubled every eighteen months. This growth has directly driven the improvement of parallel computing capabilities, enabling data centers to handle thousands of virtual machines and container tasks simultaneously.

[0029] like Figure 1AAs shown, CPUs 0-2 can perform read / write operations on addresses a0-a3, such as reading a0, reading a1, reading a2, reading a3, and writing a0, writing a1, writing a2, and writing a3. Caches 0-2, corresponding to CPUs 0-2 respectively, are used to store the data read / written from a0-a3 and the addresses a0-a3. Caches 0-2 and memory can share the data read / written from a0-a3 and the addresses a0-a3 via the data bus and address bus. The shared bus allows control signals to be shared among the CPU, memory, and caches. However, simply stacking cores also brings new challenges: memory access latency, cache coherence overhead, and inter-core communication bottlenecks are becoming increasingly prominent, and the parallel limitations revealed by Amdahl's Law are gradually becoming apparent.

[0030] Thus, the multi-die architecture emerged and became the mainstream direction in current CPU design. In a relevant example, a traditional CPU is broken down into multiple functional dies: dies focused on core computation, such as charge-coupled devices (CCDs), and I / O dies that uniformly manage interconnects and external interfaces. This division of labor can significantly improve yield rates and manufacturing flexibility, enabling manufacturers to use more advanced process technologies to produce computing units, while the I / O portion can use mature processes to reduce costs.

[0031] When a storage system used for implementing storage data I / O runs in a multi-die environment, performance degrades by approximately 20% with each additional die (the exact figure varies depending on the CPU manufacturer). This is because deploying storage software in a multi-die environment significantly increases the cross-die cache synchronization overhead of the I / O storage protocol stack, which handles the software I / O portion.

[0032] To address the cross-die caching overhead issue, in one example, a storage system implemented with multiple CPUs and multiple dies can be divided into multiple virtual nodes (vNodes). Each vNode corresponds to a set of CPU and memory resources, which is a unified allocation of resources under a unified storage pool. Associating vNodes with CPUs can reduce CPU scheduling, achieve true "resource pooling" and "load balancing," and improve CPU utilization.

[0033] In the vNode architecture, all hardware resources of the entire storage system (CPU, memory, cache, and ports of all controllers) are integrated into a unified resource pool. The storage system dynamically and on demand allocates resources to each vNode from this pool using an intelligent distribution algorithm. Multiple vNodes can run simultaneously on a single physical controller, and a vNode's resources can come from multiple physical controllers. When a Logical Unit Number (LUN) is created, the storage system assigns it a master vNode and a slave vNode (master-slave design for high availability). All I / O operations of this LUN are handled entirely by its master vNode, and this mapping relationship is completely transparent to the host; the host only sees a unified storage target.

[0034] like Figure 1B As shown, each host I / O data (such as multiple data items of 64MB each) is allocated to multiple vNode nodes, such as vNode0 to vNode31, via an intelligent distribution algorithm. Each vNode node corresponds to one or more LUNs in the resource pool, achieving intelligent balanced allocation of global cache from vNode nodes and intelligent balanced allocation of the underlying global storage pool. This achieves storage from the "controller level" to the "vNode level". In this architecture, if a physical controller fails, it only means that the vNodes running on it need to be migrated. The storage system will evenly and automatically migrate these vNodes to other healthy physical controllers, instead of migrating all LUNs in the failed physical controller A to the healthy physical controller B in the traditional architecture. The advantages of the vNode architecture include: 1. The load brought by failover is shared by all remaining controllers, without the performance bottleneck of a single controller, resulting in seamless business operations and smooth performance; 2. vNode migration is completed in seconds, much faster than traditional controller-level switching; 3. The scope of the failure impact is reduced from the entire controller to a few vNodes, achieving more granular fault isolation. In addition, when the storage system detects that a physical controller is overloaded, it can automatically migrate some vNodes running on it to a controller with a lower load to achieve global load balancing. When a new physical controller is added, the storage system can automatically migrate some vNodes to the new controller to make full use of the new resources and achieve linear and smooth expansion.

[0035] In a vNode architecture, each vNode in a storage system can implement a complete I / O storage protocol stack. A complete I / O storage protocol stack can include: a front-end module (Host layer, HL) as the host layer of the I / O storage protocol stack, responsible for interacting with the host; a cache module (CA) as the cache layer of the I / O storage protocol stack, responsible for storing hot data during the I / O process; a Redundant Array of Inexpensive Disks (RAID) module as the RAID layer of the I / O storage protocol stack, responsible for data segmentation and data verification; and a back-end module (Volume layer, VL) as the disk management layer, responsible for interacting with the back-end disks, serving as the back-end of the storage.

[0036] In the vNode architecture, each physical controller consists of multiple vNodes. Besides the vNode itself, there's a distribution module that allocates resources to each vNode. vNodes can be bound to sockets, meaning they are bound to the entire CPU of each socket; alternatively, in future multi-die environments, vNodes are expected to be bound to dies, resulting in smaller vNodes. To implement vNodes in multi-die environments, dies can be isolated according to vNodes to reduce cache synchronization between dies, thus avoiding synchronization overhead and improving performance.

[0037] However, the core of the vNode solution is a "distribution module" capable of distributing services to allocate I / O to each vNode. This distribution module can be implemented in hardware, such as a shared large card, to handle the "forwarding" of host-side I / O to the vNodes. Figure 1C As shown, the distribution module (shared large card) receives I / O data from hosts 1 to 3 and distributes it to vNode1 to vNode3. Each vNode undertakes the I / O storage protocol stack function after vertical partitioning, such as HL, CA, RAID, and VL.

[0038] When implementing vNode in a multi-CPU, multi-die architecture, the development cost of the hardware distribution module is too high. In the absence of a shared mainframe GPU, a multi-instance solution can be used to achieve a similar vNode solution. Like vNode, the multi-instance solution vertically divides a complete I / O storage protocol stack into several subsystems, achieving business segmentation similar to vNode. The multi-instance solution is a purely software solution; system management and distribution are implemented by separate modules or within a single instance. However, in a multi-die environment, any software implementation requires CPU involvement, inevitably leading to inter-die issues.

[0039] like Figure 1D As shown, a complete I / O storage protocol stack is deployed for each die; that is, both instance 1 and instance 2 include the I / O protocol stack's HL, CA, RAID, and VL. Instance 1 and instance 2 are deployed on die 1 and die 2, respectively. Furthermore, a "distribution module" that implements the distribution service via software is deployed on die 1. For example... Figure 1D The distribution module, located in die1, has remote access to other dies, which leads to performance loss due to cache synchronization between dies.

[0040] like Figure 1E As shown, to avoid the drawbacks of software-implemented distribution modules, each module of the I / O storage protocol stack can be implemented as multiple instances. Instance 1 and Instance 2 deployed in die1 and die2 respectively include HL1~VL1 and HL2~VL2. For example, HL is divided into HL1 and HL2, CA into CA1 and CA2; RAID is divided into RAID1 and RAID2, with RAID1 and RAID2 managing different stripes, such as managing stripes 1-100 and 101-200 respectively; VL is also divided into VL1 and VL2. Among the multiple instances implemented in each module, there is a master instance (such as HL1~VL1), responsible for the response of the entire module. For example, HL1 distributes CA1 or CA2 according to business needs. The master instance CA1 is responsible for managing some common data, while slave instances CA2 and possibly CA3 and CA4 are slave instances. Furthermore, CA1 also distributes data to RAID1, RAID2, etc., according to its own business needs. However, the master-slave instance design results in two issues: firstly, the master and slave instances are aware of each other's existence and interact with each other, weakening the die isolation effect; secondly, it requires significant development work for master-slave instances, as each module in the I / O storage protocol stack needs to be developed for multiple instances, resulting in a huge workload. Furthermore, the distribution of slave instances is handled by HL1, CA1, and RAID1 within the master instance, leading to frequent communication between die1 (where the master instance resides) and die2 (where the slave instances reside), resulting in significant inter-die performance degradation.

[0041] To address this, this application provides a novel I / O storage protocol stack deployment method under a multi-CPU, multi-die architecture, which reduces the performance overhead caused by inter-die communication and inter-die cache synchronization while enabling disk read and write functions based on I / O operations.

[0042] Figure 2 A schematic diagram of the architecture of a storage controller according to an embodiment of this application is shown. Figure 2 The storage controller shown includes a CPU, which includes four dies, such as die0 to die3.

[0043] Data transfer commands can be instructions obtained by the storage controller from the host to perform data read and write operations. The storage controller receives data transfer commands and performs data read and write operations on the disk according to the data transfer commands; it also feeds back the read and write data obtained from the disk to the host.

[0044] A physical storage controller may include at least one multi-core processor, such as a multi-core CPU. Each multi-core processor includes multiple processing units (dies), which can be interconnected within the multi-core processor via an interconnect bus and communicate with each other based on protocols supported by the interconnect bus. Each die within the multi-core processor may include multiple processor cores, such as C1, C2, C3, etc. Multiple processor cores can share resources bound to the die, such as shared caches (e.g., L3 caches) and buses within the die. Each die with independent memory control capabilities can be recognized by the storage system as a non-uniform memory access (NUMA) node, enabling overall memory access efficiency of the die based on the one-to-one relationship between the die and NUMA.

[0045] For example, taking a 24-core CPU as an example, the CPU can include 4 dies, each die including 6 processor cores. The 4 dies can be recognized by the operating system as 4 NUMA nodes. The storage controller can be a single 24-core CPU as described above, or it can be a "multi-CPU multi-die architecture" including multiple CPUs. It is understood that the number of processor cores and dies of the CPU can be determined according to the actual situation and is not limited to the numbers mentioned above.

[0046] Multiple dies within a storage controller (belonging to one or more multi-core CPUs) can be considered as a whole, collectively receiving data transfer instructions from the host and performing data I / O on the disk based on the I / O operations indicated by the data transfer instructions. In a multi-die architecture of the storage controller, each die only undertakes a part of the functions in the entire data I / O process; that is, a single die cannot perform a complete data I / O operation.

[0047] For example, at least one multi-core processor within the storage controller is configured to sequentially control multiple dies on the storage access path, receive data transfer instructions, and perform data read / write operations on the disk according to these instructions. The multiple processing units on the storage access path are configured to communicate in an ordered manner, using a host interaction layer, a cache layer, a redundant array of disks layer, and a disk management layer, respectively, according to the processing order of the data transfer instructions.

[0048] The storage access path is pre-configured by the storage controller based on the functions of each die. The processing order of data transfer commands is as follows: the data transfer command is processed sequentially through the host interaction layer, cache layer, RAID layer, and disk management layer to access the disk (read or write access). Multiple processing units within the storage access path communicate according to the above processing order.

[0049] The host interaction layer serves as the entry point for interaction between the storage system and the host (server), enabling communication with the host to receive data transmission commands. Following the order of command processing, the cache layer, below the host interaction layer, receives pre-processed data transmission commands from the host interaction layer and caches the data. The cache layer can cache data read from the disk and data to be written to the disk. Further, the redundant array layer, below the cache layer, combines disks into a logical array to achieve data redundancy protection and parallel I / O read / write. The redundant array layer can perform parallel read / write operations on the cached data within the cache layer through the logical array. The disk management layer, below the redundant array layer, communicates directly with the underlying disks and manages the disks throughout their lifecycle, acting as a bridge between the redundant array layer and the disks. The host interaction layer, cache layer, redundant array layer, and disk management layer are respectively referred to as HL, CA, RAID, and VL.

[0050] For each data transfer instruction received from the host, the instruction is sequentially transmitted to the corresponding die according to the storage access path to perform its respective function. Since each die communicates sequentially according to a fixed storage access route, and each die only undertakes a part of the function in the entire data I / O process, there is no need to configure a distribution module in the storage controller.

[0051] For example, in the process of receiving data transfer instructions and performing data I / O on the disk based on the I / O operations indicated by the data transfer instructions, HL, CA, RAID, and VL can be implemented by two or more dies, as long as the entire I / O process can be horizontally divided into multiple dies working together, and a single die cannot implement the entire I / O process.

[0052] In one specific embodiment, if the storage controller includes two dies, one die implements one of the four modules and the other die implements the remaining modules of the four modules; or, one die implements two of the four modules and the other die implements the remaining modules of the four modules; if the storage controller includes three dies, one die implements two of the four modules and the remaining two dies implement the remaining two modules of the four modules respectively, and so on, with each of the four modules implemented by a separate die.

[0053] In the embodiments of this application, multiple dies are configured to respectively implement the host interaction layer, cache layer, redundant array of disks layer, and disk management layer. These dies communicate in an ordered manner according to a fixed storage access path, enabling the storage controller to implement the entire I / O function through a pipelined approach using multiple dies. On one hand, the storage controller in this embodiment does not require a distribution module, avoiding communication between the dies to which the distribution module belongs and other dies. This avoids the cache synchronization overhead caused by acquiring and storing hardware resource information and allocating appropriate I / O to each die based on that information. On the other hand, since the entire I / O function is fixed to be implemented by only a subset of dies, strict isolation between the functions of different dies avoids the cache synchronization problem caused by multiple dies jointly implementing the same function, thereby reducing the overhead of data cache synchronization caused by communication between modules with the same function.

[0054] According to an embodiment of this application, a die includes multiple processor cores and a shared cache of the multiple processor cores, such as an L3 cache, which is used to store data received or processed by each die. Figure 3 A schematic diagram of the architecture of a multi-core processor according to an embodiment of this application is shown. Figure 3 As shown, this CPU consists of four dies, each corresponding to one NUMA, such as NUMA0~3 for the four dies. In this CPU architecture, the shared cache of each die is divided with finer granularity; that is, the three processor cores (C1~C3) within each die share an 8MB L3 cache. This architecture of three processor cores + L3 can be understood as a CPU core complex or core cluster (CCX), which is the most basic core module unit. It should be noted that the number of processor cores within a CCX can vary depending on the architecture.

[0055] According to embodiments of this application, the multiple dies of the storage controller can be divided into a first die, a second die, a third die, and a fourth die according to their functions, respectively implementing the host interaction layer, the cache layer, the redundant array of disks layer, and the disk management layer. It is understood that the first die, second die, third die, and fourth die can each be one or more, determined based on the resources consumed by each die in implementing its function. For example, functions that consume more resources can be implemented by multiple dies, such as multiple second dies. It should be clarified that if any of the above four functions is implemented by multiple dies, then the multiple dies implementing that function can be further divided into different sub-functions implementing that function.

[0056] The first die is configured to: receive data transmission instructions and perform protocol conversion on the data transmission instructions; and distribute the write or read instructions obtained after protocol conversion. The second die is configured to: upon receiving a write instruction from the first processing unit, forward the first data corresponding to the write instruction to the third processing unit; and / or, upon receiving a read instruction from the first processing unit, read the second data corresponding to the read instruction from the shared cache of the second processing unit. The third die is configured to: upon receiving the first data, process the first data based on the Redundant Array of Independent Disks (RAID) level to obtain multiple first blocks, and send the multiple first blocks to the fourth processing unit. The fourth die is configured to: write the multiple first blocks to the disk.

[0057] In one embodiment, taking the HL, CA, RAID, and VL included in the I / O storage protocol stack above as an example, the first die is configured to implement the function of HL, the second die is configured to implement the function of CA, the third die is configured to implement the function of RAID, and the fourth die is configured to implement the function of VL.

[0058] For example, the first die primarily performs HL (High-Level Hierarchy) functions, receiving data transfer commands from the host and parsing them into read or write commands based on the configured transport protocol (such as FC, iSCSI, etc.). The read or write command is a protocol format supported by the storage controller and can be determined based on the actual CPU configuration. It can be understood that the read or write command is derived from the specific instructions within the data transfer commands received from the host; that is, when the host has a read / write requirement, the first die receives the corresponding read / write command.

[0059] The second die primarily serves as a cache for the CA (Cardboard Cache). For read commands translated from the first die's protocol, the second data corresponding to the read command can be read from the second die's L3 cache. Conversely, for write commands translated from the first die's protocol, the write command and its corresponding first data can be forwarded to the lower-level third die. It's understandable that if the first data corresponding to a write command is I / O hot data, the second die can also store the first data in its corresponding L3 cache so that when the next read command reads the first data (the currently stored first data is considered the second data for the next read command), it can be directly read from the second die and returned.

[0060] The third die primarily performs RAID functions. For the first data received from the second die and destined for disk, it processes the data according to the RAID level configured on the third die, resulting in multiple first blocks. The RAID level is also known as the redundancy list level, and the RAID level configured on the third die can be RAID 0, RAID 1, RAID 5, RAID 6, etc. The multiple first blocks obtained after processing by the third die can be determined based on the RAID level. For example, multiple first blocks obtained based on RAID 5 processing can include data blocks after data partitioning and a parity block (or simply parity block); multiple first blocks obtained based on RAID 6 processing can include data blocks after data partitioning and two different parity blocks.

[0061] The fourth die performs the VL (Volume Level) function, writing multiple first blocks received from the third die to their corresponding physical addresses on the disk. The third die performs RAID (Range Array of Independent Disks) functionality, managing the disk through striping. A fixed mapping relationship exists between the stripes managed by the third and fourth dies and their respective physical addresses. When the third die transmits multiple first blocks, it carries their respective stripe information so that the fourth die can determine the corresponding physical addresses of each first block based on the aforementioned mapping relationship and perform the write operation.

[0062] In the embodiments of this application, for the multi-die architecture of the storage controller, the I / O storage protocol stack is horizontally partitioned, and the modules that perform some functions after horizontal partitioning are deployed on different dies. The entire storage controller is regarded as a single instance implemented by software. There are no multiple instances within the storage controller, thus avoiding the problem of insufficient isolation between master and slave instances, as well as the cross-die cache synchronization problem of the same modules in the vertical architecture of vNode and multi-instance schemes. Therefore, the multi-die deployment method of the storage controller in the embodiments of this application can at least partially overcome the cross-die cache synchronization problem.

[0063] In one specific embodiment, there are two multi-core processors: the first multi-core processor includes a first die, a second die, and a fourth die; the second multi-core processor includes a third die.

[0064] The CA and RAID modules implemented in the second and third dies, respectively, are key modules in the I / O storage protocol stack, and they perform a large number of tasks. For example, the RAID module accounts for approximately 40%, the CA module accounts for approximately 25%-30%, and HL / VL accounts for approximately 10%. In a multi-CPU, multi-die architecture, to avoid processing efficiency issues caused by improper resource allocation, the first to fourth dies can be deployed on different CPUs according to the amount of tasks performed by each die.

[0065] For example, the third die, which performs the most tasks, can be deployed on a separate CPU, while the first, second, and fourth dies, which perform the remaining three functions, can be deployed on another CPU. If there are other small modules in the I / O storage protocol stack that perform other functions (with very small tasks), the fifth die, which implements these modules, can be deployed on the same CPU as the third die to make full use of CPU resources.

[0066] In this embodiment, each die is deployed according to the workload of each die implementing its function under the multi-core, multi-CPU architecture, which enables the reuse of CPU resources. Furthermore, compared to the vNode and multi-instance schemes where CA and RAID communicate across sockets (across dies), by deploying the second and third dies implementing CA and RAID separately on different CPUs, the cache synchronization overhead of CA and RAID across sockets is reduced, improving the overall I / O performance of the storage controller.

[0067] According to embodiments of this application, each multi-core processor includes multiple dies that communicate across dies based on a first protocol; the second and third dies, and the fourth and third dies, respectively located between the two multi-core processors, all communicate across processors and dies based on a second protocol. The first and second protocols are different. The first protocol may be a Global Memory Interface (GMI) protocol, and the second protocol may be a Socket-Chip Global Memory Interconnect (xGMI) protocol.

[0068] For example, in the multi-CPU, multi-die architecture of this application, the first die transmits read / write instructions across dies to the second die within the same CPU via the GMI protocol; after reading the second data from its cache, the second die can return the second data to the first die via the GMI protocol. Alternatively, the second die can transmit the first data corresponding to a write instruction across CPUs and dies to the third die via the xGMI protocol, and then the third die sends the processed first data into multiple first blocks to the fourth die. It is understood that multiple dies within the same CPU can be interconnected via an on-chip bus. During I / O implementation, if certain dies need to communicate, they can do so via the GMI protocol; other dies that do not need to communicate do not need to communicate.

[0069] In one specific embodiment, the two CPUs may consist of only one first through fourth die.

[0070] Table 1

[0071]

[0072] In another specific embodiment, under a two-CPU multi-die architecture, the multiple dies can be deployed as shown in Table 1 above. As shown in Table 1, the storage controller may include two 24-core multi-core processors, CPU0 and CPU1. Each CPU includes 4 dies, and each die includes 6 processor cores. Each die can be identified as a NUMA, such as CPU0 including dies0-3, and CPU1 including dies4-7. Among them, die2 in CPU0 is the first die responsible for HL function, die0 and die1 are both the second dies responsible for CA function, and die3 is the fourth die responsible for VL function; in CPU1, dies4, dies5, and dies6 are the third dies responsible for RAID function. In addition, a die7 is allocated in CPU1 to handle other functions in the I / O storage protocol stack.

[0073] In the above embodiments, the processor core allocation scheme of the multi-CPU multi-die architecture uses dual-CPU as the boundary to deploy the four main functional modules of the I / O storage protocol stack in different CPUs, which identifies the problem of dual-CPU and reduces the cache synchronization overhead between the same functional modules.

[0074] However, in the above implementation, CA spans die0 and die1, occupying 12 processor cores, while RAID spans die4-6, occupying 18 processor cores. CA and RAID inevitably incur cross-die cache synchronization overhead, resulting in the storage controller's performance still not being optimal.

[0075] Therefore, the embodiments of this application further divide the multiple dies used to implement the same function into horizontal functions, so that the multiple dies can be combined to implement the function. However, the multiple dies are only used to implement a part of the functions under the function, and the operations performed by the multiple dies can be considered as pipeline operations / parallel operations under the function.

[0076] According to an embodiment of this application, there are two second dies; the first second die is configured to: read the second data corresponding to the read instruction from the shared buffer of the second die when a read instruction is received from the first die; the second second die is configured to: forward the first data corresponding to the write instruction to the third die when a write instruction is received from the first die.

[0077] In this embodiment, the two second dies used to implement the CA function are divided into two different dies according to their read and write functions. One die is used to operate according to read instructions, and the other die is used to operate according to write instructions. That is, one die is used to handle cache read operations, and the other die is used to handle the remaining CA functions. The first second die can be called the second die that implements the upper cache (UCA), and the second second die can be called the second die that implements the lower cache (LCA).

[0078] In this embodiment, two second dies, one for UCA and one for LCA, are dedicated to I / O data caching and read / write operations within the memory controller. The UCA shares the L3 cache within the first second die, and the LCA shares the L3 cache within the second second die. For example, the first die can send a read command to the first second die implementing UCA via inter-die communication, allowing it to read second data from the corresponding L3 cache. Similarly, the first die can send the first data corresponding to a write command to the first second die implementing LCA via inter-die communication; the second second die can then send the first data across CPUs and dies to the third die via inter-die communication, or temporarily write the first data into the corresponding L3 cache.

[0079] In one specific embodiment, the first second die can store I / O hot data and multiple data blocks in prefetch mode; the second second die can directly send the first data to the third die, or it can delay sending the first data to the third die through write-back mode, and temporarily write the first data into the L3 cache corresponding to the second second die.

[0080] In the embodiments of this application, the two second dies that implement the original CA function are deployed according to the different functions of reading and writing operations. This can minimize the performance overhead caused by inter-die communication and data cache synchronization between the two second dies. Moreover, the two second dies can implement their respective read and write operations in parallel. Thus, the embodiments of this application can further divide each functional module horizontally on the basis of horizontally dividing the I / O storage protocol stack function, thereby improving the overall I / O performance of the storage controller.

[0081] It is understandable that cache misses are quite common in actual use. Since the second die is responsible for functions other than reading cache, if the first die misses the cache, the second die can handle the read operation for the cache miss through inter-die communication.

[0082] According to an embodiment of this application, the first second die is further configured to: send a read instruction to the second second die if the second data corresponding to the read instruction is not read from the shared cache of the second die; the second second die is further configured to: forward the read instruction to the third die upon receiving the read instruction from the first second die; the third die is further configured to: determine the stripe information to which each data block used to compose the second data belongs upon receiving the read instruction, and send the stripe information to the fourth die; the fourth die is further configured to: read each data block used to compose the second data from the disk according to the stripe information from the third die.

[0083] When the first second die cache misses in the implementation of LCA, inter-die communication exists between the two second dies. The first second die sends a read command to the second second die, which then transmits the read command in Transparent Transmission Mode to the third die. The read command can include the logical block address (LBA) of the second data to be read after it has been divided into data blocks on the disk, and the length of the data to be read. The third die can determine the logical address and data length from the read command, determine the stripe information to which each data block used to compose the second data belongs, and send the stripe information to the fourth die. Thus, the fourth die, which is responsible for managing the disk, can read the data blocks used to compose the second data from the disk.

[0084] The third die can pre-establish a mapping relationship between logical addresses and stripe information based on its configured RAID level and / or managed stripes. Therefore, after parsing the LBA from the read request, the third die determines the stripe information for each data block used to compose the second data based on the aforementioned mapping relationship, and sends this stripe information to the fourth die. The fourth die then locates the physical address on the disk based on the stripe information and performs the read operation. After reading the data blocks used to compose the second data, the fourth die returns them to the third die, which reassembles the data blocks into the second data and feeds it back to the second die. The second data is then fed back to the host via the first die.

[0085] In the embodiments of this application, the corresponding data is first read from the shared cache by the first second die used to implement the read operation. If the cache miss occurs in the first second die, the data is then read from the disk by the second second die, minimizing the overhead of directly reading data from the disk. In addition, the first second die never needs to interact with the third die, and the first second die and the second second die only communicate in the event of a cache miss, reducing the cross-die cache synchronization overhead within the overall storage controller.

[0086] When implementing RAID functionality through multiple dies, the RAID of each processor core within each die is separate from the stripes managed by other RAIDs. However, this partitioning scheme still results in significant data sharing, leading to high cache synchronization overhead when sharing data across dies, especially in scenarios involving frequent I / O and modifications to I / O data, such as reading / writing data blocks or modifying parity blocks.

[0087] Therefore, according to one embodiment of this application, when there are at least two third dies, the plurality of first blocks include a plurality of target data blocks obtained by dividing the first data and a verification block for verifying the first data. The first third die is configured to: upon receiving the first data, divide the first data to obtain a plurality of target data blocks; generate intermediate verification data based on the plurality of target data blocks; and send the plurality of target data blocks to the fourth die so that the target data blocks are written to the disk through the fourth die. The second third die is configured to: generate a verification block for verifying the first data based on the intermediate verification data from the first third die; and send the verification block to the fourth die so that the verification block is written to the disk through the fourth die.

[0088] The third die used to implement RAID functionality can split the first data according to its configured RAID level. Some RAID levels, such as RAID 5 and RAID 6, will also generate corresponding parity blocks after splitting the first data. Therefore, this embodiment can horizontally divide the original RAID functionality into multiple processes that perform different operations, such as separating data processing and parity block generation operations and deploying them on different third dies. The first and second third dies are the data processing branch and the parity branch, respectively, and the communication between the two branches only involves cross-die communication of information required for generating parity blocks.

[0089] For example, the first third die can divide the first data into multiple target data blocks according to the size of each block within the management stripe, and determine the stripe information for each target data block. Subsequently, the divided target data blocks and their associated stripe information can be sent to the fourth die to write the target data blocks. Furthermore, since the generation of the check block depends on the data being written, the first third die can also generate intermediate check data based on the multiple target data blocks and send this intermediate check data to the second third die, which then generates a check block based on the intermediate check data. Similarly, after generating the check block, the second third die sends it to the fourth die, which stores the check block within the same stripe as the target data blocks.

[0090] In the embodiments of this application, intermediate verification data is used as data for cross-die communication between two third dies. The target data block obtained by the first third die is not synchronized to the second third die across dies, and the verification block generated by the second third die is not synchronized to the first third die across dies. Thus, the transmission of intermediate verification data and modified data is only involved between the two third dies of RAID, which reduces the cache synchronization overhead of cross-die shared data and improves the overall I / O performance of the storage controller.

[0091] In one specific embodiment, if the third die is configured for RAID5, the first third die can be deployed as a data processing branch, and the second third die as a parity branch to generate a parity block (P-block). Alternatively, if there are two or more third dies, they can be deployed as data processing branches or as parity branches.

[0092] In another specific embodiment, there are two verification blocks used to verify the first data; and three third dies. The second and third third dies are respectively configured to: generate verification blocks for verifying the first data based on intermediate verification data from the first third die; and send their respective generated verification blocks to the fourth die so that the respective generated verification blocks can be written to disk via the fourth die; wherein the verification blocks generated by the second and third third dies are different.

[0093] Understandably, when generating two different check blocks, two different check branches are deployed on two different third dies, each used to generate different check blocks. The operations performed on the second and third third dies can be similar, but they operate in parallel for different check blocks, and there is no cross-die transfer of check blocks between them.

[0094] For example, the third die is configured as RAID6, with two different parity blocks, both of which are parity blocks, abbreviated as P parity block and Q parity block. In this case, the three third dies in the storage controller can be referred to as RAID-D, RAID-P, and RAID-Q branches, respectively.

[0095] In one specific embodiment, a multi-die architecture with two CPUs may include two second dies and three third dies, and be deployed as shown in Table 2 below.

[0096] The two 24-core multi-core processors in Table 2, their die numbers within the CPUs, and the processor cores are similar to those in Table 1, and can be referred to Table 1. Under the deployment scheme in Table 2, die2 in CPU0 remains the first die responsible for HL functionality, and die3 remains the fourth die responsible for VL functionality. The original CA services in Table 1 are horizontally split into UCA and LCA, and deployed in different second dies, such as die0 and die1; the original RAID services in Table 1 are also horizontally split into different RAID-D, RAID-P, and RAID-Q, used for data processing, generating P parity blocks, and generating Q parity blocks, respectively.

[0097] Table 2

[0098]

[0099] In the embodiments of this application, for RAID levels that require the generation of different parity blocks, data processing and the generation of non-universal parity blocks are implemented through different third dies, so that the second and third third dies only have cross-die shared data (intermediate parity data) with the first third die. The generation of the two parity blocks is completely parallel, and there is no cross-die communication between them, which reduces the cache synchronization overhead of cross-die shared data and improves the overall I / O performance of the storage controller.

[0100] According to an embodiment of this application, the first third die is further configured to process the third data corresponding to the updated write instruction when the first third die receives an updated write instruction from the second die and the second and / or third third die receives intermediate verification data.

[0101] Different third dies are used to implement different RAID functions, and the operations implemented by multiple third dies can be regarded as RAID pipeline operations. In pipelined operations, after the first third die processes the first data and generates intermediate parity data, it sends the intermediate parity data to the second and third third dies. The processing operation of the first third die on the current first data can be considered complete, and the first third die can then continue to receive updated write commands from the second die and process the new third data. Simultaneously, the second and third third dies process the intermediate parity data of the first data; the first, second, and third third dies work in parallel.

[0102] For ease of understanding, the following explanation will use multiple write commands in a sequential order as an example. For instance, the first die receives multiple data transmission commands sequentially over a period of time. These commands are then converted into write commands by the protocol. The second die can then send these write commands to the third die in sequence. The data corresponding to each write command can be called I / O data. Multiple I / O data sequences can be referred to as I / O1, I / O2, I / O3, etc. The data processed by the RAID-D, RAID-P, and RAID-Q branches are simply referred to as I / O1-D, I / O1-P, I / O1-Q, I / O2-D, etc.

[0103] Taking the continuous reception of 8 I / O data as an example, the I / O timing of the third die is shown in Table 3 below.

[0104] For each I / O data item, data processing is first performed in the first third die (RAID-D branch). After data processing is complete, a corresponding PQ parity block generation task is generated and distributed to the second third die (RAID-P branch) and the third third die (RAID-Q branch) for implementation. As shown in Table 3, only at time 1, the first third die processes I / O1-D, while the second and third third dies do not process data. However, from time 2 to 8 thereafter, the second and third third dies process the I / O data from the previous time sequence of the first third die in parallel. For example, at time 2, the first third die processes I / O2-D, and the third third die processes I / O1-P and I / O1-Q in parallel. The same logic applies to other times.

[0105] Table 3

[0106]

[0107] In the embodiments of this application, the RAID function is implemented through a pipeline between the three third dies, so that the CPU resources of each third die and each processor core in each third die are nearly consistent through the pipeline operation, resulting in resource balance and high utilization.

[0108] The following explanation will use two or three third dies to implement RAID functionality as an example to illustrate the three write types of RAID. It is understood that if RAID functionality is implemented using two third dies, the multiple first blocks can consist of only multiple target data blocks and one parity block. The storage controller can include either the second or third third die (the third die capable of implementing this RAID level), as well as the first third die. If RAID functionality is implemented using three third dies, the storage controller can include all three third dies mentioned above.

[0109] To make it easier to understand, the following will use three third-party dies as an example to illustrate the three types of writing.

[0110] Figure 4 A schematic diagram illustrating the relationship between three write types and stripes according to embodiments of this application is shown. Figure 4 As shown, in RAID, each processor core within each third die can manage multiple stripes. Each stripe can be divided into 10 blocks, including 8 data blocks and 2 parity blocks, such as data blocks D0~D7, P parity blocks, and Q parity blocks. Lowercase type refers to a write operation on a single data block within a stripe, such as a lowercase type on data block D0; full write type refers to a write operation on data blocks D0~D7 (all data blocks within the stripe); uppercase type lies between lowercase and full write types. For example, uppercase type can be used for 2 data blocks, such as D0 and D1, and so on, up to 7 data blocks, such as D0~D6.

[0111] According to an embodiment of this application, the first third die is further configured to: read historical data blocks corresponding to target data blocks from the target stripe via the fourth die when the write type of the write instruction is lowercase for a single data block within the stripe, wherein the target stripe is the stripe corresponding to the write instruction; and perform logical operations on the historical data blocks and the target data blocks based on the disk redundancy array level configured for the third die to obtain intermediate verification data.

[0112] In this embodiment, the stripe type can be determined based on the logical address and data length of the write command. It is understood that RAID manages disks through striping. For a write command sent by the second die, the first and third dies can determine the write type and the target stripe corresponding to the write command based on the logical address and data length of the write command (the stripe information mentioned above includes the target stripe information). Therefore, the fourth die can read the corresponding historical data block based on the received target stripe and the correspondence between the stripe and the disk's physical address. Since each target stripe also has positions for corresponding P-parity blocks and Q-parity blocks, the fourth die can also read the historical parity blocks (including P-historical parity blocks and Q-historical parity blocks) of the write command based on the target stripe.

[0113] In this embodiment, since the generation of the verification block requires the use of historical data blocks, and the three third dies are divided into different dies according to their functions, and the first third die is used to process data, after the first third die divides the first data into different target data blocks, it also needs to read the historical data blocks corresponding to the target data blocks from the disk, and calculate the intermediate verification data through the historical data blocks and the target data blocks. The intermediate verification data can be understood as data-related intermediate data used for transmission between the second and third third dies.

[0114] It should be noted that different RAID levels have different logics for calculating parity blocks. Therefore, the logic for determining intermediate parity data based on historical data blocks and target data blocks can be determined according to the RAID level. For example, for RAID6, an XOR operation can be performed on the historical data blocks and the target data blocks to obtain the intermediate parity data.

[0115] In the embodiments of this application, for the first third die used to process data, the third die performs the reading of historical data blocks and the calculation of intermediate verification data related to the data, so that the second and third third dies do not need to share historical data blocks with the first third die. For write modifications, the second and third third dies do not need to consume resources to synchronize cached historical data blocks, but only need to receive the intermediate verification data finally obtained from the write modification, thereby reducing cross-die cache synchronization overhead.

[0116] According to an embodiment of this application, the second and / or third third die is configured to: receive intermediate check data sent by the first third die, and read historical check blocks from the target stripe through the fourth die; and perform logical operations on the historical check blocks and intermediate check data based on the disk redundancy array level to obtain a check block.

[0117] The same stripe stores both data blocks and parity blocks. Therefore, after determining the target stripe for the write instruction, the fourth die can read historical parity blocks from the target stripe containing the target data block. In this embodiment, for a storage controller with two third dies, one of the second or third third dies performs the parity block generation operation; for a storage controller with three third dies, both generate parity blocks, but the method by which they generate parity blocks can be determined according to the disk redundancy array level.

[0118] Figure 5 A flowchart illustrating the operation of lowercase types according to an embodiment of this application is shown. Figure 5 As shown, if the write type is lowercase, the first third die reads and writes the target data block ND to the read cache and reads the historical data block OD to the write cache. Based on the target data block ND and the historical data block OD, the intermediate parity data PD is determined as ND XOR OD. Furthermore, the first third die can also write the target data block ND to the disk corresponding to the target stripe through communication with the fourth die.

[0119] In this embodiment, reading and writing the target data block ND to the read cache is a fetch operation, which means reading the target data block after the first data split from the location of the split cache and writing it to the read cache, such as the Array Processing Unit (APU); at the same time, the historical data block OD can be read from the disk and written to the write cache, such as the Smart Cache Engine (SCE).

[0120] The second and third dies can read the historical parity block OP, calculate the new parity block NP = OP XOR PD, and write the new parity block NP to the disk corresponding to the target stripe through communication with the fourth die.

[0121] The third die reads the historical parity block OQ, calculates the new parity block NQ based on the intermediate parity data OP, and writes the new parity block NQ to the disk corresponding to the target stripe through communication with the fourth die. A complete RAID cycle ends after writing the target data block ND, parity block NP, and parity block NQ.

[0122] For example, for the PQ parity block supported by RAID6, the parity block NP, also known as the P parity block, can be obtained through an XOR operation; the parity block NQ, also known as the Q parity block, can be obtained through more complex operations, such as operations based on the Galois field.

[0123] In the embodiments of this application, for lowercase types with large write amplification, intermediate verification data is transmitted between multiple third dies, so that the second and third third dies do not need to perform the operation of reading historical data blocks or consume resources to synchronize the historical data blocks cached from the first third die, but only need to receive intermediate verification data, thereby reducing the cross-die cache synchronization overhead.

[0124] It should be noted that, to ensure data consistency during write operations on the target stripe, the first third die can request a stripe lock for the target stripe in advance. Before writing the target parity block, it confirms whether the stripe lock has been successfully acquired. If the stripe lock has been acquired, the write operation on the target stripe can proceed. The second and / or third third die writes the parity block; therefore, the stripe lock is released after the second and / or third third die successfully writes the parity block.

[0125] The pipeline implementation of two or three third dies in the embodiments of this application has been described above. For example, the data block process is implemented by the first third die, the P check block process is implemented by the second third die, and the Q check block process is implemented by the third third die. The specific operations of each implementation are described below using three third dies as an example.

[0126] Figure 6 A flowchart illustrating the operation of a lowercase type according to a specific embodiment of this application is shown. Figure 6 As shown, for each target data block, when a RAID write operation is required, the process can jump to the RAID thread to request a control block and reserve resources. The RAID thread can be a fiber-based FC thread, which can request a TaskControl Block (TCB) to track and manage the state of the write command. It also reserves memory for storing data, such as historical data blocks read from disk, historical parity blocks, calculated intermediate parity data, and new parity blocks. In addition to memory, the reserved resources also need to initialize the corresponding I / O structures. The initialized I / O structures (Standard IO) are abbreviated as sio, Stripe Data Element (SDE) is abbreviated as Stripe Data Element (SDE), and I / O Package is abbreviated as IPK. These resources are used to build a complete I / O operation environment.

[0127] After the I / O operation environment is built, the target data block to be written can be obtained from the Volume Group (VG) buffer (Cache Line Buffer, CLB), and historical data blocks can be read from the Volume Level (VL). The VL, which is the function implemented by the fourth die, reads historical data blocks from the disk via the VL and returns the result (e.g., successful read) to the RAID thread. After the data reading is complete, Phase 2 begins, where the data is processed. For example, intermediate parity data can be calculated based on the historical data blocks and the target data block.

[0128] Subsequently, the stripe lock can be tested. If the test indicates that the stripe lock has been acquired, the writing of the target data block can begin, and the P / Q checksum block process can be initiated. If the stripe lock has not been acquired, the process waits for the stripe lock and returns to the step of reading the target data block after the wait is completed. In the data block process, the target data block is written via VL, the disk write is returned to the RAID thread, the return result is processed (e.g., successful write), the memory is released (actually returning the used memory to the memory pool), and the reserved memory objects are unreserved (removing the marker of memory resources that were pre-allocated for this operation but may not have been fully used). At this point, the data I / O process ends.

[0129] The P-check block process and Q-check block process implemented in the second and third third dies are the same or similar. The following will use the P-check block process as an example for explanation.

[0130] After obtaining intermediate parity data from the data block process, the P / Q parity block phase 1 begins. Historical parity blocks are read from the VL (P parity block process reads P parity blocks, Q parity block process reads Q parity blocks), followed by a disk read and return to the RAID thread for processing. If the read is successful, phase 2 begins, calculating a new parity block using the intermediate parity data (the P parity block process calculates the new P parity block using the same method as the P parity block process, and the Q parity block process calculates the new Q parity block using the same method as the Q parity block process). Then, the parity block is written, and the write operation is implemented through the VL. The disk write returns to the RAID thread, which processes the return result, releases and unreserves related memory objects, releases the SDE, and releases the stripe lock. At this point, the P / Q parity block process ends. For processes involving both P and Q parity blocks, the stripe lock can be released by either the P / Q parity block process.

[0131] According to an embodiment of this application, the first third die is further configured to: when the write type of the write instruction is a full write type for stripes, perform logical operations on the segmented target data block to obtain intermediate check data, and send the intermediate check data to the second and / or third third die, so that the second and / or third third die performs logical operations on the intermediate check data according to the disk redundancy array level to obtain a check block.

[0132] For full write types, since the target data blocks to be written constitute a complete stripe, there is no need to read historical data blocks and historical check blocks from the disk. The check block can be directly calculated based on multiple target data blocks. Multiple target data blocks and check blocks can overwrite the data within the original target stripe.

[0133] Figure 7 A flowchart illustrating the operation of a full-write type according to an embodiment of this application is shown. Figure 7 As shown, the first third die requests a stripe lock in advance when it starts processing the first data, and then segments the first data into multiple target data blocks. For each target data block, a fetch operation is performed, reading and writing the data block to the read cache (APU). For example, if n target data blocks are obtained, target data blocks 1 to n can be read and written to the read cache. If all target data blocks 1 to n are successfully read and written to the read cache, they can be written to disk; if any target data block fails to be successfully read and written to the read cache, it can wait for its successful read and write before writing all target data blocks to disk in batches. Figure 7 The verification of whether a stripe lock has been acquired is not shown in the figure. In the specific implementation, the write operation is performed on the target data block if all target data blocks 1 to n have been successfully read and written to the read cache and the stripe lock has been acquired.

[0134] If target data blocks 1 through n are successfully read and written to the read cache in the first third die, intermediate parity data can be calculated based on the n target data blocks. This intermediate parity data is then sent to the second and / or third third dies, whereby the second and / or third dies perform logical operations on the intermediate parity data according to the RAID level to calculate the parity block and write it to the disk. After writing the target data blocks and the parity block, a complete RAID cycle is finished.

[0135] In the embodiments of this application, for the full write type, the transmission of intermediate check data enables the second and third third dies to only receive the intermediate check data and directly use it to generate check blocks, reducing the cross-die cache synchronization overhead.

[0136] According to an embodiment of this application, the first third die is further configured to: when the write type of the write instruction is uppercase for multiple data blocks within the stripe, read target historical data blocks from the target stripe via the fourth die, wherein the target historical data blocks are the remaining historical data blocks within the target stripe excluding the historical data blocks corresponding to the target data blocks; perform logical operations on the target historical data blocks and the target data blocks based on the disk redundancy array level to obtain intermediate check data, and send the intermediate check data to the second and / or third third die, so that the second and / or third third die performs logical operations on the intermediate check data according to the disk redundancy array level to obtain check blocks.

[0137] Figure 8 A schematic diagram illustrating a scenario of uppercase types according to an embodiment of this application is shown. Regarding uppercase types, such as... Figure 8 As shown, the example illustrates how the first data written by a write command is divided into D0' and D1' data blocks. The target stripe corresponds to historical data blocks stored on the disk, such as D0 to D7, where D0 and D1 are the data blocks to be overwritten in this write operation. After obtaining D0' and D1' data blocks, the first third die can read the target historical data blocks, i.e., D2 to D7, from the target stripe. Considering write amplification for uppercase types, D0' and D1' data blocks and the read D2 to D7 data blocks can be combined into a complete stripe, and intermediate parity data is calculated using a full-write method. The second and / or third third die also calculates parity blocks (such as P and Q parity blocks) based on the intermediate parity data using a full-write method and writes them.

[0138] In the embodiments of this application, for uppercase types, the transmission of intermediate check data enables the second and third dies to receive only the intermediate check data and directly use it to generate check blocks, reducing cross-die cache synchronization overhead.

[0139] It should be noted that the operations implemented using the above-mentioned storage controller, or the operations implemented by the server using the storage controller, can be found in the detailed description above, and will not be repeated hereafter.

[0140] Figure 9 A flowchart illustrating a read / write method according to an embodiment of this application is shown. Figure 9As shown, this data read / write method is applied to a server, which includes the aforementioned storage controller. The read / write method includes: operation S910, where at least one multi-core processor within the storage controller sequentially controls multiple processing units located on the storage access path, receives data transmission instructions, and performs data read / write operations on the disk according to the data transmission instructions; wherein, the multiple processing units located on the storage access path are configured to: achieve ordered communication through the host interaction layer, cache layer, disk redundant array layer, and disk management layer, respectively, according to the processing order of the data transmission instructions.

[0141] It is understood that in the embodiments of this application, the server may be equipped with the storage controller, which can be regarded as a storage system implemented by at least one multi-core processor. The server can interface with various clients to implement the storage or computing functions corresponding to each client. The server can act as one or more hosts (corresponding to the clients) to perform read and write computing. The instructions corresponding to the read and write computing can be sent to the storage controller as data transfer instructions, so that the storage controller can read and write to the disk according to the data transfer instructions. The operation implemented by the storage controller can be referred to the above, and will not be repeated here.

[0142] According to an embodiment of this application, upon receiving a data transmission instruction sent by the server application layer, data reading and writing on the disk includes: a first processing unit performing protocol conversion on the data transmission instruction and distributing the resulting write or read instruction; a second processing unit forwarding first data corresponding to the write instruction to a third processing unit upon receiving a write instruction from the first processing unit; and / or, upon receiving a read instruction from the first processing unit, reading second data corresponding to the read instruction from the shared cache of the second processing unit; a third processing unit processing the first data based on the disk redundancy array level upon receiving the first data to obtain multiple first blocks and sending the multiple first blocks to a fourth processing unit; and the fourth processing unit writing the multiple first blocks to the disk.

[0143] According to an embodiment of this application, the storage controller includes two multi-core processors, the first multi-core processor including a first processing unit, a second processing unit and a fourth processing unit; the second multi-core processor includes a third processing unit.

[0144] According to an embodiment of this application, the method further includes: multiple processing units included in each multi-core processor communicate across processing units based on a first protocol; the second and third processing units, and the fourth and third processing units respectively disposed between the two multi-core processors communicate across processors and across processing units based on the second protocol, wherein the first protocol and the second protocol are different.

[0145] According to an embodiment of this application, when there are two second processing units, the method further includes: when the first second processing unit receives a read instruction from the first processing unit, it reads the second data corresponding to the read instruction from the shared cache of the second processing unit; when the second second processing unit receives a write instruction from the first processing unit, it forwards the first data corresponding to the write instruction to the third processing unit.

[0146] According to an embodiment of this application, when there are three third processing units, the first data is processed based on the redundancy array level to obtain multiple first blocks, and the multiple first blocks are sent to the fourth processing unit. This includes: the first third processing unit, upon receiving the first data, segments the first data to obtain multiple target data blocks; generates intermediate verification data based on the multiple target data blocks; sends the multiple target data blocks to the fourth processing unit so that the target data blocks can be written to the disk by the fourth processing unit; the second and third third processing units respectively generate verification blocks for verifying the first data based on the intermediate verification data from the first third processing unit; and send their respective generated verification blocks to the fourth processing unit so that their respective generated verification blocks can be written to the disk by the fourth processing unit, wherein the verification blocks generated by the second and third third processing units are different.

[0147] According to embodiments of this application, this application also provides a server, including: a disk; a storage controller, including at least one multi-core processor, each of the at least one multi-core processor including multiple processing units interconnected with each other; the at least one multi-core processor is configured to: sequentially control the multiple processing units located on the storage access path, receive data transmission instructions, and perform data read and write operations on the disk according to the data transmission instructions; wherein, the multiple processing units located on the storage access path are configured to: implement ordered communication through a host interaction layer, a cache layer, a disk redundant array layer, and a disk management layer, respectively, according to the processing order of the data transmission instructions. The operation implemented by the storage controller can be referred to the above, and will not be repeated here.

[0148] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0150] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0151] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A storage controller, characterized in that, It includes at least one multi-core processor, and each of the at least one multi-core processor includes a plurality of processing units interconnected with each other; The at least one multi-core processor is configured to: sequentially control multiple processing units located on the storage access path, receive data transmission instructions, and perform data read and write operations on the disk according to the data transmission instructions; The multiple processing units on the storage access path are configured to communicate in an orderly manner through the host interaction layer, cache layer, disk redundant array layer and disk management layer, respectively, according to the processing order of the data transmission instructions. The processing unit includes a shared cache for storing data received or processed by the processing unit; the plurality of processing units located on the storage access path include: a first processing unit, two second processing units, a third processing unit, and a fourth processing unit configured to respectively implement the host interaction layer, the cache layer, the redundant array of disks layer, and the disk management layer; The first processing unit is configured to: receive the data transmission instruction and perform protocol conversion on the data transmission instruction; and distribute the write instruction or read instruction obtained after protocol conversion. The first second processing unit is configured to: upon receiving the read instruction from the first processing unit, read second data corresponding to the read instruction from the shared cache of the second processing unit; The second processing unit is configured to: upon receiving the write instruction from the first processing unit, forward the first data corresponding to the write instruction to the third processing unit; The third processing unit is configured to: upon receiving the first data, process the first data based on the disk redundancy array level to obtain multiple first blocks, and send the multiple first blocks to the fourth processing unit; The fourth processing unit is configured to write multiple of the first blocks to the disk.

2. The storage controller according to claim 1, characterized in that, The multi-core processor comprises two units: the first multi-core processor includes the first processing unit, the second processing unit, and the fourth processing unit; the second multi-core processor includes the third processing unit.

3. The storage controller according to claim 2, characterized in that, The multi-core processor includes multiple processing units that communicate across processing units based on a first protocol; The second processing unit and the third processing unit, as well as the fourth processing unit and the third processing unit, which are respectively located between the two multi-core processors, communicate across processors and processing units based on a second protocol. The first protocol and the second protocol are different.

4. The storage controller according to claim 1, characterized in that, The first second processing unit is further configured to: send the read instruction to the second second processing unit if the second data corresponding to the read instruction is not read from the shared cache of the second processing unit; The second second processing unit is further configured to: upon receiving the read instruction from the first second processing unit, forward the read instruction to the third processing unit; The third processing unit is further configured to: upon receiving the read instruction, determine the stripe information to which each data block constituting the second data belongs, and send the stripe information to the fourth processing unit; The fourth processing unit is further configured to read from the disk, based on the stripe information from the third processing unit, the data blocks used to compose the second data.

5. The storage controller according to claim 1, characterized in that, The third processing unit comprises at least two units, and the plurality of first blocks include a plurality of target data blocks obtained by segmenting the first data and a verification block for verifying the first data. The first third processing unit is configured to: upon receiving the first data, segment the first data to obtain multiple target data blocks; Intermediate verification data is generated based on multiple target data blocks; the multiple target data blocks are sent to the fourth processing unit so that the target data blocks are written to the disk by the fourth processing unit; The second third processing unit is configured to: generate a verification block for verifying the first data based on the intermediate verification data from the first third processing unit; and send the verification block to the fourth processing unit so that the verification block is written to the disk by the fourth processing unit.

6. The storage controller according to claim 5, characterized in that, The verification blocks used to verify the first data consist of two parts; the third processing unit consists of three parts. The second and third third processing units are respectively configured to: generate a verification block for verifying the first data based on the intermediate verification data from the first third processing unit; and send the respective generated verification blocks to the fourth processing unit so that the respective generated verification blocks are written to the disk by the fourth processing unit. The second and third verification blocks generated by the third processing unit are different.

7. The storage controller according to claim 5 or 6, characterized in that, The first third processing unit is further configured to process the third data corresponding to the updated write instruction when the first third processing unit receives an updated write instruction from the second processing unit, and the second and / or third third processing units receive the intermediate verification data.

8. The storage controller according to claim 5 or 6, characterized in that, The first third processing unit is further configured to: When the write type of the write instruction is lowercase for a single data block within a stripe, the fourth processing unit reads the historical data block corresponding to the target data block from the target stripe, wherein the target stripe is the stripe corresponding to the write instruction; and Based on the disk redundancy array level configured for the third processing unit, logical operations are performed on the historical data block and the target data block to obtain the intermediate verification data.

9. The storage controller according to claim 8, characterized in that, The second and / or third third processing unit is configured to: receive the intermediate verification data sent by the first third processing unit, and read historical verification blocks from the target strip through the fourth processing unit; Based on the disk redundancy array level, logical operations are performed on the historical check block and the intermediate check data to obtain the check block.

10. The storage controller according to claim 5 or 6, characterized in that, The first third processing unit is further configured to: when the write type of the write instruction is a full write type for stripes, perform logical operations on the segmented target data block to obtain the intermediate verification data, and send the intermediate verification data to the second and / or the third third processing unit, so that the second and / or the third third processing unit performs logical operations on the intermediate verification data according to the disk redundancy array level to obtain the verification block.

11. The storage controller according to claim 5 or 6, characterized in that, The first third processing unit is further configured to: When the write type of the write instruction is uppercase for multiple data blocks within a stripe, the fourth processing unit reads the target historical data block from the target stripe, wherein the target historical data block is the remaining historical data block within the target stripe excluding the historical data block corresponding to the target data block; Based on the disk redundancy array level, logical operations are performed on the target historical data block and the target data block to obtain the intermediate verification data. The intermediate verification data is then sent to the second and / or third third processing unit, so that the second and / or third third processing unit performs logical operations on the intermediate verification data according to the disk redundancy array level to obtain the verification block.

12. A data read / write method, characterized in that, Applied to a server, the server including a storage controller as described in any one of claims 1 to 11, The read / write method includes: The storage controller contains at least one multi-core processor that sequentially controls multiple processing units located on the storage access path, receives data transmission instructions, and performs data read and write operations on the disk according to the data transmission instructions. The multiple processing units along the storage access path are configured to: communicate in an ordered manner through a host interaction layer, a cache layer, a redundant array of disks layer, and a disk management layer, respectively, according to the processing order of the data transmission instructions; receiving the data transmission instructions and performing data read / write operations on the disk according to the data transmission instructions includes: The first processing unit receives the data transmission instruction and performs protocol conversion on the data transmission instruction; it then distributes the write or read instructions obtained after the protocol conversion. When there are two second processing units: when the first second processing unit receives the read instruction from the first processing unit, it reads the second data corresponding to the read instruction from the shared cache of the second processing unit; when the second second processing unit receives the write instruction from the first processing unit, it forwards the first data corresponding to the write instruction to the third processing unit. Upon receiving the first data, the third processing unit processes the first data based on the disk redundancy array level to obtain multiple first blocks, and sends the multiple first blocks to the fourth processing unit. The fourth processing unit writes multiple of the first blocks to the disk.

13. The method according to claim 12, characterized in that, The storage controller includes two multi-core processors. The first multi-core processor includes the first processing unit, the second processing unit, and the fourth processing unit. The second multi-core processor includes the third processing unit.

14. The method according to claim 13, characterized in that, Also includes: The multi-core processor includes multiple processing units that communicate across processing units based on a first protocol; The second processing unit and the third processing unit, as well as the fourth processing unit and the third processing unit, which are respectively located between the two multi-core processors, communicate across processors and processing units based on a second protocol. The first protocol and the second protocol are different.

15. The method according to claim 12, characterized in that, When there are three third processing units, the first data is processed based on the disk redundancy array level to obtain multiple first blocks, and the multiple first blocks are sent to the fourth processing unit, including: Upon receiving the first data, the first third processing unit segments the first data to obtain multiple target data blocks; generates intermediate verification data based on the multiple target data blocks; and sends the multiple target data blocks to the fourth processing unit so that the target data blocks can be written to the disk by the fourth processing unit. The second and third third processing units respectively generate verification blocks for verifying the first data based on the intermediate verification data from the first third processing unit; and send the respective generated verification blocks to the fourth processing unit so that the fourth processing unit can write the respective generated verification blocks to the disk, wherein the verification blocks generated by the second and third third processing units are different.

16. A server, characterized in that, include: disk; A storage controller includes at least one multi-core processor, each of the at least one multi-core processor including a plurality of processing units interconnected with each other; The at least one multi-core processor is configured to: sequentially control multiple processing units located on the storage access path, receive data transmission instructions, and perform data read and write operations on the disk according to the data transmission instructions; The multiple processing units on the storage access path are configured to communicate in an orderly manner through the host interaction layer, cache layer, disk redundant array layer and disk management layer, respectively, according to the processing order of the data transmission instructions. The processing unit includes a shared cache for storing data received or processed by the processing unit; the plurality of processing units located on the storage access path include: a first processing unit, two second processing units, a third processing unit, and a fourth processing unit configured to respectively implement the host interaction layer, the cache layer, the redundant array of disks layer, and the disk management layer; The first processing unit is configured to: receive the data transmission instruction and perform protocol conversion on the data transmission instruction; and distribute the write instruction or read instruction obtained after protocol conversion. The first second processing unit is configured to: upon receiving the read instruction from the first processing unit, read second data corresponding to the read instruction from the shared cache of the second processing unit; The second processing unit is configured to: upon receiving the write instruction from the first processing unit, forward the first data corresponding to the write instruction to the third processing unit; The third processing unit is configured to: upon receiving the first data, process the first data based on the disk redundancy array level to obtain multiple first blocks, and send the multiple first blocks to the fourth processing unit; The fourth processing unit is configured to write multiple of the first blocks to the disk.

Citation Information

Patent Citations

  • Method and system for transmitting data between multi-core processor and disk array

    CN103838517A

  • Task processing method, processor, equipment and readable storage medium

    CN113407352A