Storage backend full sharing extension cabinet management method, system and device and storage medium

By setting up multiple storage management system nodes and dual-port paths in the storage expansion cabinet, unified management of disk and disk port status is achieved, solving the problem of insufficient node sharing in existing technologies, improving system redundancy and reliability, and optimizing path performance.

CN117289866BActive Publication Date: 2026-08-25INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311269289.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2026-08-25
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

In existing storage expansion cabinet management methods, insufficient node sharing of backend SAS, SSD, and HDD disks leads to low performance and redundancy, failing to meet the needs of more control, and insufficient path optimization and redundancy result in low reliability.

Method used

Configure the system to support at least four storage management system nodes, build two paths between each node and the backend disk, achieve unified management of disk and disk port status, and optimize the path through queue depth control and IO path selection to ensure system redundancy and reliability.

Benefits of technology

It improves system redundancy and reliability, avoids drive-forwarding when disk ports fail, makes full use of system resources, and improves data read and write capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117289866B_ABST
    Figure CN117289866B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of storage backend full sharing extension cabinet management method, system, device and storage medium, it is related to storage backend extension cabinet management field.The present application is provided to support at least four storage management system nodes, for each storage management system node is built to the path to rear disk, the unified management of disk and disk port state is realized at CSM end, and it is synchronized to agent end;Each storage management system node is configured to support to build two paths to rear disk, when two paths are formed between storage management system node and rear disk, queue depth control, IO path selection, path management are realized.When one storage management system node is damaged, more storage management system nodes can be provided to share work, and the system has higher redundancy and robustness.Because multiple paths are formed between the iogroup formed by rear disk and storage management system node, path optimization strategy is supported, system resources can be fully utilized, and the ability of data read and write is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of storage back-end expansion cabinet management, and more particularly to a storage back-end fully shared expansion cabinet management method, system, device, and storage medium. Background Technology

[0002] With the development of the big data era, storage applications are becoming more and more widespread and the designs are becoming more and more complex. Currently, storage with one frame and four controllers has emerged. Ordinary expansion cabinets can no longer meet the needs of storage with one frame and four controllers. Four-controller or more-controller back-end shared expansion cabinets are needed to support front-end horizontal expansion.

[0003] In the current expansion cabinet management method, the backend SAS, SSD, and HDD disks are shared by two nodes within the same iogroup. The number of two nodes is insufficient to meet the needs of more controllers, resulting in low performance and redundancy. Furthermore, the storage management system nodes and backend disk ports are in a one-to-one correspondence, and not every storage management system node has a path to both disk ports, causing problems with low redundancy and reliability. Path optimization is also impossible, and the performance of more controllers cannot be maximized without path optimization. Currently, if one disk port node is kicked from a dual controller, drive-forwarding will occur, with only one target candidate node, resulting in low reliability. Summary of the Invention

[0004] To solve the above-mentioned technical problems, or at least partially solve them, the present invention provides a storage backend fully shared expansion cabinet management method, system, device, and storage medium.

[0005] In a first aspect, the present invention provides a method for managing a fully shared expansion cabinet for a storage backend, comprising:

[0006] Configure the system to support at least four storage management system nodes, build a path to the backend disk for each storage management system node, implement unified management of disk and disk port status on the CSM side, and synchronize it to the agent side;

[0007] Each storage management system node is configured to support the construction of two paths to the backend disk. When two paths are formed between the storage management system node and the backend disk, queue depth control, IO path selection, and path management are implemented. Among them, queue depth control includes: the maximum queue depth of the disk is shared by both ports within a storage management system node; IO path selection includes: when both ports are fast paths, IO is sent in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected and IO is sent based on the relationship between the current IO error and the disk port; when the path of either port fails, IO is sent from the other port.

[0008] Furthermore, the number of discovery domains is set to at least four to support at least four storage management system nodes, and the data structure is designed according to the maximum number of storage management system nodes supported so that the data structure supports data records for at least four controllers.

[0009] Furthermore, the process of constructing a path to the backend disk for each storage management system node, achieving unified management of disk and disk port status on the CSM end, and synchronizing it to the agent end includes:

[0010] The CSM terminal receives a target object abstracted from a disk port, carrying a unique disk port identifier; the disk port is discovered independently by the local discovery function of each storage management system, and the discovered disk port is abstracted into a target object.

[0011] The CSM end indexes the target objects of the same disk port reported by different storage management system nodes to the same CSM target object based on the same disk port identifier; then, it maintains the online node bitmap and exclusion status according to the relationship between the target object of the disk port, the storage management system node, and the CSM target object to form a unified view of the disk port, and then reverse-synchronizes the unified view information of the disk port to all local agent nodes.

[0012] The CSM end maps target objects with disk ports carrying the same UUID to dual ports of the same disk, forms a unified disk status view based on the online node bitmap and exclusion status of the dual ports, and synchronizes the unified disk status view to all local agent nodes.

[0013] Furthermore, the unified view of disk status is updated by triggering events, including: the local node fabric event when the storage system is created, the storage node pend and unpend events, and the port sinker event caused by local IO errors.

[0014] Furthermore, the number of disk channels is configured to be equal to the number of disk ports; each channel of the disk is distinguished by a channel index, and an exclusion status is set for channels distinguished by the channel index. Login of ports in the exclusion status is eliminated to achieve preferred login of channels from the target object.

[0015] Furthermore, when performing IO path selection, it includes selection based on load or traffic balancing. Selection based on load or traffic balancing includes: obtaining the number of fast paths, determining whether the current channel is a fast path, if the number of fast paths is greater than 1 and the current channel is not a fast path, then obtaining fast paths, filtering out paths with high IO load or traffic overload from the selected fast paths, determining whether the number of remaining fast paths after filtering is non-zero, if so, polling and selecting channels in the remaining fast paths, and returning the channel index, if the number of remaining fast paths is zero, then the channel remains unchanged.

[0016] Secondly, the present invention provides a storage backend fully shared expansion cabinet management system, implementing the aforementioned storage backend fully shared expansion cabinet management method, including:

[0017] The discovery module implements device access discovery, maintains the discovery command sequence, and is responsible for identifying device categories and types. It responds to login events to complete local discovery, and responds to cluster module requests to cooperate with multi-discovery reporting or update business objects. Local discovery in response to login events is performed independently on each storage management system node. It works with the object module to abstract the discovered disk ports into target objects and reports them to the CSM terminal with a unique disk port identifier. It supports each storage management system node in independently constructing paths to backend disks, enabling unified management of disk and disk port status on the CSM terminal.

[0018] Cluster Module: The functions of the cluster module include actively notifying each agent module, receiving basic events from agent modules, and using generated events to achieve mutual notification between clusters; initiating cluster-wide multi-discovery and maintaining a unified cluster view of business objects; cooperating with the CSM client to achieve unified management of disk and disk port status; and providing command line and event interfaces.

[0019] The proxy module implements local components and interacts with the cluster module through basic events; it maintains local node services, protocol objects and relationships, and keeps them synchronized with the cluster module; it implements local node services and, when local services involve path selection, it works with the IO path module to achieve path selection; the proxy module is configured with the number of disk channels equal to the number of disk ports to support each storage management system node in building two paths to the backend disk.

[0020] IO Path Module: This module implements the protocol command path and is responsible for encapsulating external I / O, internal I / O, and internal business and management requests into commands and sending them to the access devices. It implements functions including queue depth control, channel selection, thread scheduling, flow control, and command error handling. Queue depth control includes: the dual disk ports share the maximum queue depth within a storage management system node; IO path selection includes: when both disk ports are fast paths, IO is issued in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected and IO is issued based on the relationship between the current IO error and the disk port recorded in the ERP processing table; when the path on either disk port fails, IO is issued from the other port.

[0021] Object Module: The object module provides object functionality for the cluster module and the proxy module. It supports the creation of two target objects based on UUID for each backend disk, so that each storage management system node can build two paths to the backend disk.

[0022] Furthermore, the agent module distinguishes each channel of the disk through the channel index, sets the exclusion status of the target object, sets the exclusion of the channel according to the channel index, and selects the target object according to the exclusion status of the channel, so as to realize the preferred login from the target object query of the channel.

[0023] Thirdly, the present invention provides a storage backend fully shared expansion cabinet management device, comprising: at least one processing unit, the processing unit being connected to a storage unit via a bus unit, the storage unit storing a computer program, and the computer program being executed by the processing unit to implement the storage backend fully shared expansion cabinet management method.

[0024] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the storage backend fully shared expansion cabinet management method.

[0025] The technical solutions provided in the embodiments of the present invention have the following advantages compared with the prior art:

[0026] This application supports at least four storage management system (SMS) nodes. For each SMS node, a path to the backend disk is constructed. Horizontal scaling of SMS nodes within an iogroup is supported. Since each SMS node and backend disk are connected via a channel within the iogroup, when one SMS node fails, more SMS nodes can share the workload, resulting in higher system redundancy and robustness. The disk ports of the SMS nodes and backend disks are no longer in a one-to-one correspondence; each SMS node supports two paths to the backend disk. Because the backend disk supports two paths, drive-forwarding does not occur when a disk port is removed. Furthermore, since the iogroup formed by the backend disk and the SMS nodes creates multiple paths, path optimization strategies can be implemented. Path optimization is based on queue depth control; both disk ports share the maximum disk queue depth within a single SMS node. IO path selection includes: when both ports are fast paths, IO is distributed in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected based on the relationship between the current IO error and the disk port; when one port path fails, the IO is distributed on the other port. It can make full use of system resources and improve data reading and writing capabilities.

[0027] In essence, local discovery within each node operates independently. Disk ports are abstracted as target objects, each carrying a unique wwpn identifier, and reported to the CSM. The same disk port target object reported by different nodes is indexed to the same CSM target object on the CSM side based on the same wwpn. Then, the online node bitmap and exclusion status are updated. A unified view of disk ports is maintained on the CSM side and synchronized to all local agent nodes. Additionally, the CSM side maps disk port objects carrying the same UUID to dual ports of the same disk. Based on the online node bitmap and exclusion status of the dual ports, a unified view of disk status is maintained and synchronized to all local agent nodes. This achieves unified management of disk and disk port status on the CSM side and synchronization to the agent side. Attached Figure Description

[0028] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 A schematic diagram of a storage backend fully shared expansion cabinet management system provided in an embodiment of the present invention;

[0031] Figure 2 This is a flowchart illustrating the local discovery process performed by the discovery module in an embodiment of the present invention.

[0032] Figure 3 A flowchart illustrating the local discovery, target object creation, update, and completion processes performed by the cluster module, proxy module, and discovery module provided in this embodiment of the invention.

[0033] Figure 4 This is a flowchart illustrating the target object binding of the proxy module execution channel and the data handover when a channel change occurs, provided in an embodiment of the present invention.

[0034] Figure 5 A flowchart for path selection to achieve traffic balancing and I / O load balancing, provided for embodiments of the present invention;

[0035] Figure 6 A schematic diagram of an object module provided in an embodiment of the present invention;

[0036] Figure 7 This is a schematic diagram of a storage backend fully shared expansion cabinet management device provided in an embodiment of the present invention. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0039] To make the objectives, technical solutions, and advantages of the embodiments of this invention clearer, the English terms involved in this application are explained below. CF: Processing front-end parameter parsing, reading command parameters input by the user, then encapsulating the parameters and sending them to SMI; receiving data returned by SMI and parsing it, then printing the results to standard output. SMI: Shared memory, acting as a bridge connecting CF and PLCL. PLCL: The actual execution process of the command line. Hl: A module in the storage management system responsible for host object management, protocol parsing and processing, and namespace management. PLIF: A module in the storage management system responsible for general function interface abstraction, extended logical ports, and Fabrics information management and querying. CLI: Short for command line issued by the storage system. CSM end: Responsible for forwarding data or information. AGT: Responsible for collecting and acquiring data or information. SNAP: The package name in the storage management system where packaged logs are stored.

[0040] Example 1

[0041] This invention provides a method for managing a fully shared storage backend expansion cabinet, comprising:

[0042] The system is configured to support at least four storage management system nodes. A path to the backend disk is built for each storage management system node. Unified management of disk and disk port status is implemented on the CSM side and synchronized to the agent side. Since each storage management system node in iogroup is connected to the backend disk via a channel, if one storage management system node fails, more storage management system nodes can share the workload, resulting in higher system redundancy and robustness.

[0043] Each storage management system node is configured to support two paths to the backend disk. When two paths exist between the storage management system node and the backend disk, queue depth control, IO path selection, and path management are implemented. Queue depth control includes: both disk ports share the maximum queue depth within a single storage management system node; IO path selection includes: when both disk ports are fast paths, IO is distributed in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected based on the relationship between the current IO error and the disk port recorded in the ERP processing table; when one port's path fails, the other port distributes the IO. The backend disk supports two paths, and the number of disk channels is configured to match the number of disk ports. Because the backend disk supports two paths, drive-forwarding does not occur when a disk port is removed. Furthermore, since multiple paths are formed between the iogroups of the backend disk and the storage management system node, path optimization strategies can be implemented.

[0044] In practice, to support at least four controllers, the number of discovery domains is set to at least four to support at least four storage management system nodes. The data structure is designed according to the maximum number of storage management system nodes supported, so that the data structure can support data records for at least four controllers. For example, when constructing an object, the length of the array recording port or node data in the object's trace is adapted to the expansion requirements of the storage management system nodes in the iogroup.

[0045] To enable each storage management system node to build a path to the backend disks, each storage management system node needs to independently discover disk ports locally, abstracting the discovered disk ports into target objects and reporting them to the CSM (Disk Management System) with a unique disk port identifier. Target objects for the same disk port reported by different storage management system nodes are indexed into the same CSM target object on the CSM based on the same disk port identifier. Then, based on the relationship between the storage management system node and the CSM target object, an online node bitmap and exclusion status are maintained to form a unified view of disk ports on the CSM. This unified view information is then back-synchronized to all local agent nodes. This achieves unified management of disk port status on the CSM and synchronization with the agent nodes.

[0046] The CSM (Card Manager Service) maps target objects with the same UUID to dual ports of the same disk. Based on the online node bitmap and exclusion status of the dual ports, a unified view of disk status is formed and synchronized to all local agent nodes. This unified management of disk port status is achieved on the CSM and synchronized to the agent. Furthermore, mapping target objects with the same UUID to dual ports of the same disk on the CSM enables storage management system nodes to build and debug paths to backend disks. When two paths exist between the storage management system node and the backend disk, queue depth control, IO path selection, and path management are implemented.

[0047] By triggering the unified disk status view through defined events, such as the local node fabric event, storage node pend and unpend events, and port sinker event caused by local IO errors when the storage system is created, the disk status update discovery process between the agent node and the CSM end is triggered. This ensures that the CSM end can update the unified disk status view in real time, provide the latest unified disk status view in a timely manner, and ensure the synchronization of the unified disk status view between nodes.

[0048] Since a disk supports multiple channels, each channel is distinguished by a channel index. An exclusion status is set for channels distinguished by their channel indexes, and logins on ports with this exclusion status are removed to ensure optimal logins for channels from the target object. Specifically, in this application, a disk and a storage management system node may form an iogroup with multiple disk channels. To accommodate this, and to ensure optimal logins for channels from the target object, this application, in addition to setting the target object's sinker flag, also adds an exclusion status setting. Each channel of the disk is distinguished by its channel index. In practice, when querying optimal logins for a target object, the port table is traversed, and logins with an exclusion value of false are set for each channel according to its channel index. This binds available target object logins to the channel, and subsequent initialization of target object logins is performed as needed.

[0049] Queue depth control includes: disk dual-port sharing the maximum queue depth within a storage management system node; disk dual-port sharing a maximum queue depth, with I / O to be distributed from any storage management system node stored in a queue for the dual-port, and the I / O to be distributed in the queue being allocated to the port; it supports both round-robin distribution in the case of dual-port fast path and selective distribution based on disk port load, traffic and fast path conditions.

[0050] IO path selection includes: when both disk ports are fast paths, IO is sent out in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected and IO is sent out based on the relationship between the current IO error and the disk port recorded in the ERP processing table; when the path of either of the disk ports fails, IO is sent out from the other port.

[0051] See Figure 5 As shown, during error retries, when performing IO path selection, the selection is performed based on load or traffic balancing. The selection based on load or traffic balancing includes: obtaining the fast path quantity and determining whether the current channel is a fast path. If the fast path quantity is greater than 1 and the current channel is not a fast path, then a fast path is obtained. Paths with high IO load or traffic overload are filtered out from the selected fast paths. It is then determined whether the remaining fast path quantity after filtering is non-zero. If so, channels in the remaining fast paths are selected in a round-robin fashion, and the channel index is returned. If the remaining fast path quantity is zero, then the channel remains unchanged.

[0052] Example 2

[0053] See Figure 1 As shown, this embodiment of the invention provides a storage backend fully shared expansion cabinet management system, which realizes device access and business applications for shared storage protocol interface devices. Shared device access refers to the process by which the storage management system, as the initiator, discovers, identifies, and configures backend protocol objects reported on a specific link channel through storage protocols (SCSI / NVME), and maintains the functionality of the backend protocol objects. Supported device types include storage disks, chassis, heterogeneous storage, and storage management system nodes. Business applications include: I / O services and management services. I / O services refer to encapsulating external and internal I / O into protocol I / O commands to achieve data access to devices; management services include arbitration, inspection, and management of the status and attributes of business objects. The device supports setting up at least four storage management system nodes. For each storage management system node, a path to the backend disk is constructed. Unified management of disk and disk port status is achieved at the CSM end and synchronized to the agent end. Each storage management system node is configured to support the construction of two paths to the backend disk. When two paths are formed between the storage management system node and the backend disk, queue depth control, IO path selection, and path management are implemented. Among them, queue depth control includes: the maximum queue depth of the disk is shared by both disk ports within a storage management system node; IO path selection includes: when both disk ports are fast paths, IO is sent in a round-robin manner; during error retries, if the fast path condition is met, the path is reselected and IO is sent according to the relationship between the current IO error and the disk port recorded in the ERP processing table; when the path of either of the disk ports fails, IO is sent to the other port.

[0054] The storage backend fully shared expansion cabinet management system includes:

[0055] The discovery module comprises four sub-modules: discovery domain, discovery initiator, discovery ITN, and discovery LUN. Its functions include: implementing device access discovery; maintaining discovery command sequences; identifying device categories and types; responding to login events to complete local discovery; and responding to cluster module requests to cooperate with multi-discovery reporting or update business objects. In this application, to support at least four controllers, at least four discovery domains are set, and the data structure is designed according to the maximum number of storage management system nodes supported, so that the data structure supports data records for at least four controllers. The data structure will be explained later.

[0056] In the specific implementation process, please refer to Figure 2 As shown, the discovery module implements the following process:

[0057] Create or locate the vl login, device, and disco_itn. Enable login and start the disco_itn state machine. Determine if node identification is required; if so, enter the testing phase, execute preset test commands, identify node information, and generate device information at the end of the testing phase. After the testing phase, enter the completion phase. Otherwise, enter the backend phase, execute preset backend commands, identify basic device information, and generate device information at the end of the backend phase, updating the relationship between the itn and the domain. The device information includes devtab type and basic device information. Then, enter the lun_x phase. In the lun_x phase, select the preset command to identify media information based on the devtab type, identify detailed media information of the disk or chassis, and create and update agent_lu, save detailed media information, and update the relationship between agent_lu and the domain. Determine if all media of the disk or chassis has been identified; otherwise, repeat the lun_x phase; if so, enter the completion phase. In the completion phase, configure the itn to stop working, discover the domain and report to the CSM, and log in and register. Local discovery of login events is performed independently on each storage management system node. In conjunction with the object module, the discovered disk ports are abstracted into target objects and reported to the CSM terminal with a unique disk port identifier.

[0058] Cluster Module: The functions of the cluster module include: actively notifying each agent module through em cluster call based on the CSM client; receiving basic events sent by the agent modules; and using generated events to realize mutual notification between clusters; initiating cluster-wide multi-discovery and maintaining a unified cluster view of business objects; cooperating with the CSM client to realize unified management of disk and disk port status; and providing CLI and event interfaces.

[0059] In the specific implementation process, please refer to Figure 3 As shown, the cluster module, CSM terminal, discovery module, and agent module work together to achieve this.

[0060] The cluster module triggers a discovery request and sends it to the agent module via the CSM client; the agent module then uses domain tagging to associate the request.

[0061] When the cluster module enters the invalid phase, it starts the state machine to trigger the agent module domain to perform local discovery; the agent module starts the discovery module's discovery ITN state machine, performs local discovery, and notifies the cluster module through basic events.

[0062] The cluster module submits the agent_lu view to the agent module.

[0063] The cluster module enters the creation phase. The control agent module generates drive agent_lu links and enclosure agent_lu links. The agent module reports the newly added drive agent_lu links and enclosure agent_lu links to the CSM end, and the CSM end synchronizes them to each agent module.

[0064] During the target object creation phase, the cluster module uses the object module to generate the target object, and the control agent module creates drive agent_lu links and enclosure agent_lu links for the target object. The agent module reports the newly added object to the CSM end, and the CSM end synchronizes it to each agent module.

[0065] During the target object update phase, the cluster module controls the agent module to report the target object update to the CSM end, and the CSM end synchronizes it to each agent module.

[0066] During the drive / enclosure update phase, the cluster module controls the agent module to report the drive / enclosure update to the CSM terminal, and the CSM terminal synchronizes it to each agent module.

[0067] During the completion phase, the cluster module updates the online node bitmap and exclusion status.

[0068] The proxy module implements local components and interacts with the cluster module through basic events; it maintains local node services, protocol objects (shared JBOD chassis, SAS disks, storage management system nodes, ports, and logins) and their relationships, and keeps them synchronized with the cluster module; it implements local node services including quorum, scrub, and device initialization.

[0069] When the storage system is created, events including local node fabric events, storage node pend and unpend events, and port sinker events caused by local IO errors trigger a discovery process between the agent node and the CSM (Disk Management System) to update disk status. This ensures that the CSM can update the unified disk status view in real time and provide the latest unified disk status view promptly, ensuring synchronization of the unified disk status view between nodes. For example, in the inactive process after SAS domain discovery is completed, there are two points that trigger service channel reselection: first, the vlun_sm status is PENDING_DISCOVERY, meaning the previous process relied on rediscovery to update the disk vlun object as a whole; second, the target object has a sinker flag, meaning that a port of the disk vlun corresponding to the target object has had an error count, requiring a partial update of a service channel of that disk vlun. The two channel reselection processes change the unified disk status view.

[0070] See Figure 4 As shown, in the specific implementation, the agent module iteratively executes the channel query to select the preferred login from the target object. This application configures the number of disk channels to be equal to the number of disk ports, meaning a disk may have multiple disk channels. To accommodate this, and to implement the channel query to select the preferred login from the target object, this application, in addition to setting the target object's `slander` flag, also adds an exclusion status setting for the target object. Each channel on the disk is distinguished by its channel index. In the specific implementation, when the target object queries for the preferred login, the port table is traversed, and logins with the exclusion set to `false` are set according to the channel index. This binds the available target object logins to the channels, and the target logins are subsequently initialized as needed.

[0071] In the actual implementation process, if there is a replacement between the old and new channels, the agent module will repeatedly execute the data handover according to the channel and initialize the channel.

[0072] IO Path Module: This module implements the protocol command path, responsible for encapsulating external I / O, internal I / O, and internal business and management requests into commands and sending them to the access devices. It implements functions including queue depth control, channel selection, thread scheduling, flow control, and command error handling. Specifically, when the fabric path changes, the IO path module updates the channel information of iopathvlun and initializes the login process. It selects the channel (corresponding to the disk port) for I / O distribution within iopath to fully utilize the performance of the backend dual-port disk and achieve load balancing of I / O across different paths.

[0073] The interface for reselection is vl_iopath_vlun_channel_query_next, which is executed on the corresponding thread of the driver to prevent conflicts. The selection timing includes: dev layer I / O execution, I / O queue rescheduling, instruction retry during ERP error handling, and delayed delivery of TIM layer instructions; in addition, the format task also performs channel reselection when periodically issuing TUR instructions to determine the completion status.

[0074] Channel selection follows these principles: A. When both ports of a disk are fast paths, issue requests via a two-port round-robin system. B. When both ports are fast paths, if one port is overloaded, the port with traffic below the threshold is prioritized for request delivery. C. When retrying an I / O error, if the fast path condition is met, the path is reselected based on the relationship between the current I / O error and the disk port recorded in the ERP processing table. D. When a single-port path fails, the I / O is delivered to the other port. (See also...) Figure 4 As shown, an exemplary process for selecting an IO path for load or traffic balancing includes: obtaining the number of fast paths; determining whether the current channel is a fast path; if the number of fast paths is greater than 1 and the current channel is not a fast path, then the channel is selected as the target fast path channel; otherwise, the channel remains unchanged. Selecting a channel for a fast path includes: obtaining fast paths; filtering out paths with high IO load or traffic overload from the selected fast paths; determining whether the remaining number of fast paths after filtering is non-zero; if so, polling and selecting channels from the remaining fast paths; returning the channel index; if the remaining number of fast paths is zero, then the channel remains unchanged.

[0075] Object Module: This module provides object-related functionalities for the cluster module and proxy module. (See also...) Figure 6 As shown, the object module provides object tables and state machine model implementations for CSM and proxy modules. The object module reuses the existing architecture as a whole, and adapts the data structure to the expansion of iogroup. For example, the array length of the port or node data recorded in the trace is adapted to the expansion requirements of iogroup.

[0076] Example 3

[0077] See Figure 7As shown, this embodiment of the invention provides a storage backend fully shared expansion cabinet management device, including: at least one processing unit, the processing unit being connected to a storage unit via a bus unit, the storage unit serving as a computer-readable storage medium, which can be used to store software programs, computer-executable programs, and modules, such as the software program, computer-executable program, and module corresponding to the storage backend fully shared expansion cabinet management method in this embodiment of the invention. The processing unit implements the aforementioned storage backend fully shared expansion cabinet management method by running the software program, computer-executable program, and module stored in the storage unit, including:

[0078] Configure the system to support at least four storage management system nodes, build a path to the backend disk for each storage management system node, implement unified management of disk and disk port status on the CSM side, and synchronize it to the agent side;

[0079] Each storage management system node is configured to support the construction of two paths to the backend disk. When two paths are formed between the storage management system node and the backend disk, queue depth control, IO path selection, and path management are implemented. Among them, queue depth control includes: the maximum queue depth of the disk is shared by both ports within a storage management system node; IO path selection includes: when both ports are fast paths, IO is distributed in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected and IO is distributed according to the relationship between the current IO error and the disk port recorded in the ERP processing table; when the path of either port of the disk fails, IO is distributed from the other port.

[0080] Of course, the computer program stored in the storage unit of the storage back-end fully shared expansion cabinet management device provided in the embodiments of the present invention is not limited to the operation of the method described above, and can also execute related operations in the storage back-end fully shared expansion cabinet management method provided in any embodiment of the present invention.

[0081] Example 4

[0082] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed, it implements the storage backend fully shared expansion cabinet management method, the method comprising:

[0083] Configure the system to support at least four storage management system nodes, build a path to the backend disk for each storage management system node, implement unified management of disk and disk port status on the CSM side, and synchronize it to the agent side;

[0084] Each storage management system node is configured to support the construction of two paths to the backend disk. When two paths are formed between the storage management system node and the backend disk, queue depth control, IO path selection, and path management are implemented. Among them, queue depth control includes: the maximum queue depth of the disk is shared by both ports within a storage management system node; IO path selection includes: when both ports are fast paths, IO is distributed in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected and IO is distributed according to the relationship between the current IO error and the disk port recorded in the ERP processing table; when the path of either port of the disk fails, IO is distributed from the other port.

[0085] Of course, the computer program stored in the computer-readable storage medium provided in the embodiments of the present invention is not limited to the operation of the method described above, but can also execute related operations in the storage back-end fully shared expansion cabinet management method provided in any embodiment of the present invention.

[0086] In the embodiments provided by this invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, structures, or units, and may be electrical, mechanical, or other forms.

[0087] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0088] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0089] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for managing a fully shared expansion cabinet for a storage backend, characterized in that, include: Configure the system to support at least four storage management system nodes, build a path to the backend disk for each storage management system node, implement unified management of disk and disk port status on the CSM side, and synchronize it to the agent side; Each storage management system node is configured to support the construction of two paths to the backend disk. When two paths are formed between the storage management system node and the backend disk, queue depth control, IO path selection, and path management are implemented. Among them, queue depth control includes: the maximum queue depth of the disk is shared by both ports within a storage management system node; IO path selection includes: when both ports are fast paths, IO is sent in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected and IO is sent based on the relationship between the current IO error and the disk port; when the path of either port fails, IO is sent from the other port.

2. The storage backend fully shared expansion cabinet management method according to claim 1, characterized in that, Set the number of discovery domains to at least four to support at least four storage management system nodes, and design the data structure according to the maximum number of storage management system nodes supported, so that the data structure supports data records of at least four controllers.

3. The storage backend fully shared expansion cabinet management method according to claim 1, characterized in that, The process of constructing a path to the backend disk for each storage management system node, achieving unified management of disk and disk port status on the CSM end, and synchronizing it to the agent end includes: The CSM terminal receives a target object abstracted from a disk port, carrying a unique disk port identifier; the disk port is discovered independently by the local discovery function of each storage management system, and the discovered disk port is abstracted into a target object. The CSM end indexes the target objects of the same disk port reported by different storage management system nodes to the same CSM target object based on the same disk port identifier; then, it maintains the online node bitmap and exclusion status according to the relationship between the target object of the disk port, the storage management system node, and the CSM target object to form a unified view of the disk port, and then reverse-synchronizes the unified view information of the disk port to all local agent nodes. The CSM end maps target objects with disk ports carrying the same UUID to dual ports of the same disk, forms a unified disk status view based on the online node bitmap and exclusion status of the dual ports, and synchronizes the unified disk status view to all local agent nodes.

4. The storage backend fully shared expansion cabinet management method according to claim 3, characterized in that, The unified view of disk status is updated by triggering the settings of events, including: the local node fabric event when the storage system is created, the storage node pend and unpend events, and the port sinker event caused by local IO errors.

5. The storage backend fully shared expansion cabinet management method according to claim 1, characterized in that, Configure the number of disk channels to be equal to the number of disk ports; distinguish each channel of the disk through a channel index, set an exclusion status for channels that are distinguished by the channel index, and remove logins for ports in the exclusion status to achieve preferred logins for channels from the target object.

6. The storage backend fully shared expansion cabinet management method according to claim 1, characterized in that, When performing IO path selection, it includes selection based on load or traffic balancing. Selection based on load or traffic balancing includes: obtaining the number of fast paths, determining whether the current channel is a fast path. If the number of fast paths is greater than 1 and the current channel is not a fast path, then obtain the fast paths, filter out paths with high IO load or traffic overload from the selected fast paths, determine whether the number of remaining fast paths after filtering is non-zero, if so, poll and select channels in the remaining fast paths, and return the channel index. If the number of remaining fast paths is zero, then the channel remains unchanged.

7. A storage backend fully shared expansion cabinet management system, implementing the storage backend fully shared expansion cabinet management method according to any one of claims 1-6, characterized in that, include: The discovery module implements the device access discovery function, maintains the discovery command sequence, and is responsible for device category and type identification; The system responds to login events to complete local discovery, and responds to cluster module requests to cooperate with multiple discovery reporting or update business objects. Local discovery in response to login events is performed independently on each storage management system node. It works with the object module to abstract the discovered disk ports into target objects and reports them to the CSM end with a unique disk port identifier. It supports each storage management system node to independently build a path to the backend disk, and realizes unified management of disk and disk port status on the CSM end. Cluster Module: The functions of the cluster module include actively notifying each agent module, receiving basic events from agent modules, and using generated events to achieve mutual notification between clusters; it is responsible for initiating cluster-wide multi-discovery and maintaining a unified cluster view of business objects. It works in conjunction with the CSM client to achieve unified management of disk and disk port status; it provides command line and event interfaces. The proxy module implements local components and interacts with the cluster module through basic events; it maintains local node services, protocol objects and relationships, and keeps them synchronized with the cluster module; it implements local node services and, when local services involve path selection, it works with the IO path module to achieve path selection; the proxy module is configured with the number of disk channels equal to the number of disk ports to support each storage management system node in building two paths to the backend disk. IO Path Module: This module implements the protocol command path and is responsible for encapsulating external I / O, internal I / O, and internal business and management requests into commands and sending them to the access devices. It implements functions including queue depth control, channel selection, thread scheduling, flow control, and command error handling. Queue depth control includes: the dual disk ports share the maximum queue depth within a storage management system node; IO path selection includes: when both disk ports are fast paths, IO is issued in a round-robin fashion; during error retries, if the fast path condition is met, the path is reselected and IO is issued based on the relationship between the current IO error and the disk port recorded in the ERP processing table; when the path on either disk port fails, IO is issued from the other port. Object Module: The object module provides object functionality for the cluster module and the proxy module. It supports the creation of two target objects based on UUID for each backend disk, so that each storage management system node can build two paths to the backend disk.

8. A storage backend fully shared expansion cabinet management system according to claim 7, characterized in that, The agent module distinguishes each channel of the disk by channel index, sets the exclusion status of the target object, sets the exclusion of the channel according to the channel index, and selects the target object according to the exclusion status of the channel, so as to realize the preferred login from the target object.

9. A storage backend fully shared expansion cabinet management device, characterized in that, include: At least one processing unit is provided, which is connected to a storage unit via a bus unit. The storage unit stores a computer program. When the computer program is executed by the processing unit, it implements the storage backend fully shared expansion cabinet management method as described in any one of claims 1-6.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the storage backend fully shared expansion cabinet management method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Large-scale data storage and delivery system

    CN104903874A

  • Storage control device and storage control device path switching method

    US20070002847A1