Storage system

By introducing a front-end interface to the storage system to manage the session establishment process of multiple connections, the problem of too long session establishment time in the NVMe/TCP protocol is solved, and the access performance and reliability of the storage system are improved.

CN120541010AInactive Publication Date: 2025-08-26HITACHI VANDALA CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411294725.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-21
Filing Date
2024-09-14
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the NVMe/TCP protocol, the session establishment time is too long, resulting in a degradation of the access performance of the storage system. Especially when the number of multiple CPU cores exceeds 100, the possibility of session establishment failure increases.

Method used

A front-end interface (FE I/F) is introduced in the storage system, which manages the session establishment process of multiple connections, and shortens the session establishment time by generating session management information and connection management information.

Benefits of technology

By optimizing session management, the time required for session establishment is reduced, the access performance and reliability of the storage system is improved, and the risk of session establishment failure is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541010A_ABST
    Figure CN120541010A_ABST
Patent Text Reader

Abstract

The invention provides a storage system which shortens the time required for session establishment. The storage system communicates with the host in a session including one or more connections. The storage system comprises a front-end interface, a processor and a storage area. The storage area stores session management information that manages sessions of communication with a host. The front-end interface stores connection management information that manages connections of the session. The front-end interface controls access from the host with reference to the connection management information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a storage system. Background Art

[0002] In recent years, storage area networks (SANs) have become increasingly popular as a means of connecting storage systems and host servers in information systems. In a SAN, network cables, such as optical fibers, connect the storage system and host servers via switches. A SAN enables storage resources to be shared among multiple host servers. The software that runs on a host server and accesses the storage system is called an initiator, while the software that runs in the storage system and accepts storage requests from the initiator and provides access to the storage is called a target.

[0003] SAN types include FC-SANs, which use Fibre Channel (FC), and IP (Internet Protocol)-SANs, which use Ethernet. FC-SANs use dedicated interface modules and FC switches for lossless data forwarding, resulting in high reliability and suitable for mission-critical IT systems. Meanwhile, IP-SANs are based on standard IP protocols and can be easily used without the advanced expertise required for FC-SANs. They ensure reliability by performing retransmission control on communication data at the TCP layer of the upper protocol, and their adoption in mission-critical information systems is increasing. Furthermore, the widespread adoption of 100Gb Ethernet and 200Gb Ethernet is driving broadband growth, raising expectations for IP-SANs.

[0004] As non-volatile memory-based protocols become more common, NVMe / TCP (Non-Volatile Memory Express over Transmission Control Protocol), which is expected to offer even higher performance, is gaining popularity as a replacement for the conventional iSCSI (Internet Small Computer System Interface).

[0005] In iSCSI, a logical connection between an initiator and a target is called a session. A session basically exchanges iSCSI requests / responses over a TCP connection, allowing a host server to access a storage device.

[0006] On the other hand, in NVMe / TCP, in an NVMe association (equivalent to a session in iSCSI, hereinafter referred to as a "session" in this specification except when the protocol needs to be distinguished), which is a logical connection between a host server and a storage system, multiple TCP connections (NVMe / TCP connections, hereinafter referred to as "connections" in this specification except when the protocol needs to be distinguished) can be used to exchange NVMe requests / responses and access the storage system from the host server. As a result, NVMe / TCP can perform storage access with improved IO (Input Output) parallelism, achieving storage access with wide bandwidth and low latency.

[0007] Patent Document 1 discloses a SmartNIC-utilizing storage system that incorporates a SmartNIC in the storage system controller and performs protocol processing via the SmartNIC. The SmartNIC is a network interface device equipped with a CPU and memory, and operates a general-purpose operating system (OS) and OSS (Open Source Software) protocol server.

[0008] Unlike storage controllers, running protocol-related software on SmartNICs can reduce the controller load and improve storage performance. In addition, by changing the software on the SmartNIC without making major changes to the controller, new protocols and communication functions can be supported.

[0009] In protocols such as NVMe / TCP that improve access performance to storage devices by increasing IO parallelism, the number of connections per initiator increases compared to protocols such as iSCSI that access through a single connection. This is because initiators usually run in multiple CPU (Central Processing Unit) cores, and each CPU core shares the storage access processing of multiple connections, so that it is expected that the access performance to the storage device can be improved. Therefore, in order to expect the highest access performance, the same number of connections as the number of CPU cores are established, but in recent years, the number of CPU cores sometimes exceeds 100. In this case, more than 100 connections need to be established in one session establishment in NVMe / TCP.

[0010] Patent Document 1 does not disclose the detailed operations during session establishment. Session management in a storage system must be performed by a controller. Therefore, in iSCSI session establishment, which is based on a single connection, it is natural for the SmartNIC to notify the controller each time a connection is established, and for the connection, or session, to be managed within the controller.

[0011] However, if it is applied to NVMe / TCP, every time multiple connections (NVMe / TCP connections) constituting one session (NVMe association) are established, the SmartNIC notifies the controller, and the multiple connections are managed as sessions within the controller.

[0012] As a result, establishing a session takes time. Furthermore, if the initiator performs sequential connection establishment, i.e., it starts the next connection process after completing the connection process for one connection, session establishment takes even longer. Depending on the initiator's settings and requirements, if session establishment takes time, it may be considered a session establishment failure.

[0013] That is, in a storage protocol in which a single session is composed of a plurality of connections, the time required to establish the session becomes an issue.

[0014] Patent Document 1: Japanese Patent Application Laid-Open No. 2023-142021 Summary of the Invention

[0015] One embodiment of the present invention is a storage system for communicating with a host in a session including one or more connections, comprising a front-end interface, a processor, and a storage area, wherein the storage area stores session management information for managing the communication session with the host, the front-end interface stores connection management information for managing the connection of the session, and the front-end interface controls access from the host with reference to the connection management information.

[0016] According to one aspect of the present invention, the time required to establish a session including multiple connections can be shortened. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is an overall structural diagram of the information system of Example 1.

[0018] Figure 2 This is a structural diagram of the storage control block.

[0019] Figure 3 A configuration example of FE I / F is shown.

[0020] Figure 4 This is an example of a host server configuration diagram.

[0021] Figure 5 This is an example of a management server configuration diagram.

[0022] Figure 6 This shows an example structure of a subsystem management table.

[0023] Figure 7 This shows an example structure of a namespace management table.

[0024] Figure 8 This shows an example of the structure of the controller management table.

[0025] Figure 9 This shows an example of the structure of the table held by the controller.

[0026] Figure 10 This shows an example of the structure of the connection management table.

[0027] Figure 11 This is a timing diagram showing the initialization process.

[0028] Figure 12 This is a flowchart showing the FE I / F during initialization processing.

[0029] Figure 13 A flowchart showing the memory control block during initialization processing.

[0030] Figure 14 This is a sequence diagram showing the management queue connection establishment process.

[0031] Figure 15 A timing diagram showing the IO queue connection establishment process.

[0032] Figure 16A This is a flowchart of the FE I / F connection establishment process.

[0033] Figure 16B This is a flowchart of the FE I / F connection establishment process.

[0034] Figure 17 This is a flowchart of the connection establishment processing of the storage control block.

[0035] Figure 18 This is a timing diagram showing IO access processing.

[0036] Figure 19 This is a flowchart of the IO access processing of the FE I / F.

[0037] Figure 20 This is a flowchart of the IO access processing of the storage control block.

[0038] Figure 21 This is a sequence diagram of the connection disconnection process triggered by the host server.

[0039] Figure 22 This is a flowchart of the connection disconnection process triggered by the host server of the FE I / F.

[0040] Figure 23 This is a flowchart of the connection disconnection process triggered by the host server of the storage control block.

[0041] Figure 24 This is a sequence diagram of the connection disconnection process triggered by the storage system.

[0042] Figure 25 This is a flowchart of the storage system-triggered disconnection process of the FE I / F.

[0043] Figure 26 This is a flowchart of the connection disconnection process triggered by the storage system of the storage control block.

[0044] Figure 27 This is a flowchart showing the connection disconnection process triggered by the host server of the FE I / F according to the second embodiment.

[0045] Figure 28 A structural example of a connection management table according to the fourth embodiment is shown. DETAILED DESCRIPTION

[0046] Example 1

[0047] The following describes the structure of an information system using a general structure implemented by hardware. However, the structure of the present invention is not limited to hardware. Virtualization technology can also be used to implement part or all of the hardware through software to ensure flexibility when the information system is changed. In addition, the number of components such as storage areas, CPUs (Central Processing Units), and buses is described as one unless otherwise specified. However, multiple components can be prepared to achieve redundancy, load distribution, or partitioning to improve convenience and cost-effectiveness. The bus can also be partitioned to facilitate arbitration, or a wide-band bus such as PCIe (Peripheral Component Interconnect-Express) can be used to improve performance.

[0048] Storage, also known as memory, is typically an area for storing information, typically composed of DRAM (Dynamic Random Access Memory). However, it is also possible to optimize storage capacity, access speed, and cost by layering memory using, for example, SRAM (Static Random Access Memory), flash memory, or HDDs. Furthermore, by placing some or all of the storage area remotely and accessing and utilizing it appropriately via a network connected by input / output devices, it is possible to conserve the storage space required for the computer.

[0049] In the following structural description, the computer is based on a general structure such as a CPU, storage area, input / output devices, and bus. Therefore, even if not separately described, the storage area stores programs and data executed by the CPU to control the computer's operations. In addition, commonly used devices in computers can be added to improve convenience. For example, a serial bus can be added, and user interface devices such as a keyboard and display can be added to improve the operator's operability of the information system. Alternatively, a structure that allows operators to access the system remotely via a network can be adopted to improve convenience.

[0050] In a storage system according to one embodiment of the present invention, the FE I / F (Front-end Interface) notifies the processor of the storage control block of the establishment of the initial connection among the multiple connections that constitute a session. The storage control block then generates management information for a new session that includes the notified initial connection. Furthermore, the FE I / F manages information about each connection and manages it in correspondence with the storage control block's session management information. According to one embodiment of the present invention, a storage system equipped with the FE I / F can shorten the time required to establish a session when using a protocol that uses multiple connections to form a single session.

[0051] Figure 1 1 is an overall configuration diagram of the information system of Example 1. The information system includes one or more host servers 200, a storage system 1, a network 30, and a management server 50. The host server 200, the storage system 1, and the management server 50 are interconnected via a network 3.

[0052] Storage system 1 includes one or more storage device units 20 and a storage control device 10. Storage control device 10 includes one or more storage control blocks 100. To improve the availability of storage system 1, storage control device 10 may include multiple storage control blocks 100, each powered by a dedicated power supply. Alternatively, multiple storage control devices 10 may be interconnected via an HCA (Host Channel Adapter) network to achieve improved availability and load balancing performance.

[0053] The storage control device 10 or storage control block 100 is generally also called a storage controller or simply a controller, and provides storage functions. Figure 1 In the embodiment, a redundant configuration is shown as a configuration in which two storage control blocks 100 are held in the storage control device 10. However, a simple configuration may be adopted in which only one storage control block 100 is provided, which also serves as the storage control device 10. In this case, the storage control block 100 becomes a storage controller.

[0054] The storage control block 100 includes a BE I / F (Back-end Interface) 120 and one or more FE I / Fs 110. In this example, the FE I / F 110 is a SmartNIC (Network Interface Card).

[0055] The storage device unit 20 includes one or more PDEVs 21. PDEVs 21 are physical devices, such as HDDs (Hard Disk Drives), other storage devices (non-volatile storage devices), flash memory devices such as SSDs (Solid State Drives), and battery-powered DRAMs (Dynamic Random Access Memory).

[0056] The storage device unit 20 can also include different types of PDEVs 21 to achieve improved fault tolerance and optimized performance and cost due to diversity. Furthermore, a RAID (Redundant Array of Inexpensive Disks) group can be formed using multiple PDEVs 21 of the same type, storing data according to a predetermined RAID level to achieve optimized fault tolerance and capacity based on the requirements.

[0057] The network 3 is a network used for mutual communication between the connected host server 200, the storage system 1, and the management server 50, for example, using a LAN (Local Area Network), but it is also possible to use virtual network technology to logically construct and mix different types of networks to reduce network deployment costs, or use wireless technology to reduce the complexity of cable wiring.

[0058] The host server 200 is a device that connects to the storage system 1 and performs storage access. Specifically, it sends requests to establish / disconnect a connection to the storage system 1, change settings, and input / output stored data (data write request, data read request).

[0059] The management server 50 is a PC (personal computer) or server having a user interface based on a GUI (Graphical User Interface) or a CLI (Command Line Interface), and provides functions for users or operators to control and monitor the storage system 1 .

[0060] Figure 21 is a block diagram of the storage control block 100. The storage control block 100 includes a BE I / F 120, one or more FE I / Fs 110, a CPU (Central Processing Unit) 103, and a storage area 104, which are connected to each other via a bus.

[0061] BE I / F 120 and FE I / F 110 are equivalent to input and output devices in a computer. BE I / F 120 is an interface for communicating with storage device unit 20. FE I / F 110 is a SmartNIC, a programmable network interface, and operates as part of the storage protocol when the host server 200 accesses the storage system 1.

[0062] In this embodiment, NVMe over TCP (Non Volatile Memory Expressover Transmission Control Protocol) is used as an example of a storage protocol, but other storage protocols such as iSCSI (Internet Small Computer System Interface) can also be used to select a storage access method that optimizes the cost and access speed of the information system requirements.

[0063] The storage system 1 processes logical devices (LDEVs) divided from centralized physical devices (PDEVs) such as RAID. Here, several NVMe terms are explained in this manual. The host is the device that uses the storage system 1 (equivalent to the initiator in iSCSI) and is uniquely identified by a host ID and has a name such as host NQN.

[0064] A subsystem provides one or more devices (equivalent to storage servers or iSCSI targets) within storage system 1. A subsystem has one or more controllers and one or more namespaces. A subsystem is uniquely identified by a subsystem ID and has a name such as a subsystem NQN.

[0065] The NVMe controller is the interface used to access the subsystem and is identified by its controller ID. A namespace is a logical device provided by the subsystem and is identified by its namespace ID. A port (Fabricport) is the network interface used to access the controller and is identified by its port ID.

[0066] An association is a logical connection between a host and a subsystem. A host establishes an association with a controller by accessing a port. When a host accesses a subsystem, it can reference one or more namespaces, or logical devices, identified by namespace IDs.

[0067] In NVMe / TCP, for access to a subsystem (association), a management queue connection and any number of IO queue connections are established. The queue connection is established first, followed by the IO queue connection. The namespace of the access destination is specified during IO (Read / Write) access.

[0068] The storage area 104 stores the storage control program group P0 executed by the CPU 103 and the management information managed by the storage control program group P0. The management information includes a subsystem management table T10, a namespace management table T20, and a controller management table T30. Details of the management information and the processing of the storage control block 100 will be described later.

[0069] Figure 3 This figure shows an example of the configuration of the FE I / F 110. In this example, the FE I / F 110 is a SmartNIC. A SmartNIC is a high-performance network card (NIC) that can be programmed (added) with user-desired functions through software or hardware. It is a front-end interface device. For example, a SmartNIC can perform functions at the transport layer and the application layer.

[0070] The following description of SmartNICs applies not only to interface devices whose functions can be programmed through software executed by a processor, but also to interface devices with programmable logic circuit structures such as FPGAs (Field Programmable Gate Arrays). FPGAs can include logic circuits that implement various functions executed by programs and cache memory used for calculations.

[0071] The FE I / F 110 includes a network I / F 111, an internal I / F 112, a CPU 113, and a storage area 114. These are connected to each other via a communication line such as a bus.

[0072] The network I / F 111 is an interface device for communicating with the host server 200. The network I / F 111 functions as a network port (hereinafter referred to as a port) for setting an IP address and performing communication. An IP address is an identifier on the network, and the host server 200 communicates with the FE I / F 110 using the IP address set to the port.

[0073] The internal I / F 112 is an interface device for communicating with the storage control block 100. The internal I / F 112 is connected to the CPU and the like of the storage control block 100 via, for example, PCIe (Peripheral Component Interconnect-Express).

[0074] The CPU 113 controls the operation of the FE I / F 110. The storage area 114 stores programs and data used for controlling the operation of the CPU 113. The storage area 114 stores an interface processing program group P10, a controller holding table T50, and a connection management table T60.

[0075] The interface processing program group P10, executed by the CPU 113, controls the connection for communication between the host server 200 and the storage system 1, and a session consisting of one or more connections. In this embodiment, a TCP / IP (Transmission Control Protocol / Internet Protocol) connection is assumed as the connection type, and an NVMe / TCP association is assumed as the session.

[0076] The interface processing program group P10 configures a TCP port for a listening service that accepts connection requests for each port of the FE I / F 110. Upon receiving a connection request to the listening service, the interface processing program group P10 establishes a TCP connection, and then accepts a session request from the host server to establish a session.

[0077] The interface processing program group P10 includes an OS (Operating System) of the FE I / F 110 , and communicates with the storage control block 100 to perform initialization of the FE I / F 110 , resource management, fault management, and task scheduling.

[0078] The interface handler group P10 receives various read / write requests from the host server 200 and other programs, and processes the block protocols contained in these requests. The interface handler group P10 processes block access protocols such as NVMe / TCP received from the host server 200 and converts them into block access command requests for the storage control block 100. The interface handler group P10 communicates with the storage control block 100 and performs data writing / reading operations on the LDEVs that constitute the subsystem's namespace in response to these various requests.

[0079] Figure 4 2 is an example of a configuration diagram of the host server 200. The host server 200 includes a network I / F 201, a CPU 202, and a storage area 203. These are connected to each other via a communication line such as a bus.

[0080] The network I / F 201 is an interface device for communicating with the storage system 1 and the management server 50. The CPU 202 controls the operation of the host server 200. The storage area 203 stores programs and tables used to control the operation of the CPU 202. The storage area 203 stores an application program P41 and a storage connection program P43. The storage area 203 also stores information used by the programs.

[0081] The application program P41 is executed by the CPU 202 to read and write data to the subsystem namespace provided by the storage system 1 via the storage connection program P43. The storage connection program P43 receives various read / write requests from the application program P41 and reads and writes data to the storage system 1.

[0082] Figure 5 This is an example of a configuration diagram of the management server 50. The management server 50 includes a network I / F 51, a CPU 52, and a storage area 53. These are connected to each other via a communication line such as a bus. The network I / F 51 is an interface device for communicating with the storage system 1 and the host server 200.

[0083] The CPU 52 controls the operation of the management server 50. The storage area 53 stores programs and data used to control the operation of the CPU 52. The storage area 53 stores a management server program P50. The management server program P50 includes a user interface based on a GUI, CLI, or the like, and provides functions for users or operators to control and monitor the storage system 1. Upon receiving control or monitoring instructions for the storage system 1 from the user, the management server program P50 communicates with the storage system 1 to perform control or monitoring.

[0084] The following describes in detail the management information stored in the storage control block 100. While the management information is represented in tables, key-value representations, such as those suitable for improving functionality and performance, such as fault tolerance, can also be used. Furthermore, while tables sometimes use a single field within a single entry to store multiple values, it is also possible to divide tables or entries to store information across multiple tables and entries, thereby achieving standardization that meets performance and functional requirements.

[0085] Figure 6 The following shows an example of the structure of the subsystem management table T10. The subsystem management table T10 associates the subsystem with the subsystem NQN and the port of the FE I / F 110. Figure 6In the illustrated configuration example, subsystem management table T10 includes a subsystem ID column C101, a subsystem NQN column C102, and a port ID column C103. Subsystem ID column C101 stores the identifier of the corresponding subsystem within storage system 1. Subsystem NQN column C102 indicates the NQN (NVMe Qualified Name), the subsystem's identifier in the NVMe / TCP protocol. Port ID column C103 indicates the port identifier used by the host to access the subsystem's FE I / F 110.

[0086] Figure 7 This table shows an example of the structure of the namespace management table T20. Each subsystem provides one or more namespaces to the host. Only one logical device (LDEV) is allocated to each namespace. The namespace management table T20 manages the relationship between them.

[0087] exist Figure 7 In the illustrated example, namespace management table T20 includes a subsystem ID column C201, a namespace ID column C202, and an LDEVID column C203. Each entry represents information about a single namespace. Subsystem ID column C201 indicates the ID of the subsystem that owns each namespace. Namespace ID column C202 indicates the ID that identifies the namespace within the subsystem. Each namespace is identified within storage system 1 by the combination of the subsystem ID and the namespace ID. LDEVID column C203 indicates the ID of the LDEV that constitutes each namespace. Within each subsystem, there is a one-to-one correspondence between namespaces and LDEVs.

[0088] Figure 8 The following table shows an example of the structure of the controller management table T30. The controller management table T30 is an interface for access from the host, and the controller management table T30 manages current association information. When an association is created, a new entry is added, and when the association is completed, the entry is deleted.

[0089] exist Figure 8 In the illustrated configuration example, the controller management table T30 includes a controller ID column C301, a subsystem ID column C302, a port ID column C303, a host NQN column C304, a host ID column C305, a protocol column C306, a scheduled queue number column C307, and a connection number column C308. Each entry indicates information on one current association.

[0090] The controller ID column C301 indicates the ID of the controller accessed by the host in the context. The controller ID identifies the controller within the storage system 1. The subsystem ID column C302 indicates the ID of the subsystem having the controller. The port ID column C303 indicates the ID of the port of the FE I / F 110 that is the access destination in the context.

[0091] The Host NQN column C304 indicates the NQN of the source host in the association. The Host ID column C305 indicates the ID of the source host in the association. The Protocol column C306 indicates the type of communication protocol used in the association. In this example, NVMe / TCP is used. FC-NVMe and iSCSI can also be used as other examples.

[0092] The scheduled queue number column C307 indicates the maximum number of IO queues in the association. Since IO queues are set for each IO queue connection (excluding management queue connections), the scheduled queue number is equivalent to the scheduled IO queue connection number. For example, the scheduled queue number is set to a request value from the host. In addition, a maximum allowable value is pre-set for the scheduled queue number, and the scheduled queue number in the association can be set and registered within a range below the maximum allowable value. The connection number column C308 indicates the current number of IO queue connections in the association. This value is equivalent to the current number of IO queues in the association.

[0093] Next, the management information held in the FE I / F 110 will be described in more detail. Figure 3 As shown, the FE I / F 110 stores the controller retention table T50 and the connection management table T60 in the storage area 114 .

[0094] Figure 9 The following shows an example of the structure of the controller storage table T50. The controller storage table T50 is composed of cache data and additional information from the controller management table T30 stored and managed by the storage control block 100. The controller storage table T50 may contain only information related to the use of the port of the FE I / F 110, or may contain information related to the use of the port of another FE I / F 110.

[0095] When the former is adopted, the number of objects to be managed can be reduced, thereby saving resources and effort required for management. On the other hand, when the latter is adopted, information required for redundancy and exclusion processes in cooperation with multiple FE I / Fs 110 can be confirmed without additional inquiries, thereby simplifying processing. Furthermore, since the controller-held table T50 is a cache of information held by the storage control block 100, it is possible to omit part or all of the controller-held table T50 in order to query the storage control block 100 for the required information, thereby saving computer resources required by the FE I / F 110.

[0096] exist Figure 9 In the illustrated configuration example, the controller retention table T50 includes a controller ID column C501, a subsystem ID column C502, a port ID column C503, a host NQN column C504, a host ID column C505, a protocol column C506, a scheduled queue number column C507, a connection number column C508, and an available namespace ID column C509.

[0097] Each entry represents a current association (session) information. The data in columns C501 to C508 is a cache of the data in columns C301 to C308 of the same name in the controller management table T30, and the data between them is the same. Figure 9 In the cache of the controller management table T30 held in the controller retention table T50 shown, some data may be omitted.

[0098] The Available Namespace ID column C509 indicates the namespaces accessible to the host in the association, that is, the IDs of the LDEVs. One association can access one or more specified namespaces (LDEVs). The information in the Available Namespace ID column C509 is passed from the storage control block 100 to the FE I / F 110.

[0099] Figure 10 This shows an example of the structure of the connection management table T60. The connection management table T60 manages the controller's TCP information. Each entry represents information about a single connection. The connection management table T60 manages both queue connections and IO queue connections. The connection management table T60 manages information about connections through the FE I / F 110 that maintains it, but may not include information about other FE I / Fs 110. The former approach reduces the number of management targets, thereby saving resources and effort required for management. On the other hand, the latter approach simplifies processing by confirming information required for redundancy and exclusion processes for collaboration with multiple FE I / Fs 110 without additional inquiries.

[0100] exist Figure 10 In the illustrated configuration example, the connection management table T60 includes a connection ID column C601 , a controller ID column C602 , a host IP address column C603 , a host port number column C604 , a target IP address column C605 , a target port number column C606 , a queue ID column C607 , and a connection setting column C608 .

[0101] The connection ID column C601 is an identifier of an entry in the connection management table T60. The connection ID column C601 may be omitted. The controller ID column C602 indicates the controller ID associated with the connection.

[0102] The Host IP Address column C603 and Host Port Number column C604 indicate the IP address and TCP port number of the connected host. These values ​​are specified by the host. The Target IP Address column C605 and Target Port Number column C606 indicate the target IP address and TCP port number. These values ​​are set by the storage control block 100.

[0103] The queue ID column C607 indicates the ID of the queue assigned to the connection. "0" is set for the management queue, and integer values ​​greater than "1" are assigned to the IO queues in sequence. The connection setting column C608 indicates the setting information for each connection. In this example, the setting information for the keep-alive function is registered, specifically the keep-alive timeout (KeepAliveTimeout) time. The keep-alive timeout time is the time to wait for the next new request without closing the connection after the completion of a request, and the unit is, for example, milliseconds.

[0104] Requests can also be limited to specific requests such as keep-alive instructions to make processing easier to understand. As another example of the connection setting field C608, the processing priority of the connection relative to other connections can be set to prioritize or lag behind specific connections to use resources appropriate to the request.

[0105] The following describes the processing of the storage control block 100 and the FE I / F 110. In the following description, the order of processing may be changed while maintaining consistency, or previous and subsequent processing may be combined to simplify the processing and reduce the number of communications.

[0106] Parameters used in the processing of the storage control block 100 and the FE I / F 110 are used when information is exchanged between the storage control block 100 and the FE I / F 110. Otherwise, information pre-set in the storage control block 100 and the FE I / F 110 is used. Some or all of the parameters included in the information exchanged via communication may be pre-set in the storage control block 100 and the FE I / F 110 to save on communication and processing. Conversely, additional information may be added during communication to reduce the number of pre-set parameters, making it easier to change settings.

[0107] First, the initialization process of the storage control block 100 and the FE I / F 110 will be described. Figure 11 A timing diagram showing the initialization process, Figure 12 A flowchart showing the FE I / F 110 during initialization processing is shown. Figure 13 A flowchart showing the initialization process of the storage control block 100.

[0108] exist Figure 11-13 Before starting the initialization process shown, the storage control block 100 generates the subsystem management table T10 and the namespace management table T20 in advance by design and operator settings. In addition, the FE I / F 110 is configured in advance for initializing communication with the storage control block 100 by design and initial settings.

[0109] like Figure 11As shown in FIG. 13 , the storage control block 100 starts up the FE I / F 110 ( S21 ). The FE I / F 110 is started up, for example, by supplying power. Figure 11 As shown in FIG. 1 , the activated FE I / F 110 and the storage control block 100 establish communication therebetween ( S11 , S22 ). Figure 12 As shown, the activated FE I / F 110 establishes communication with the storage control block 100 (S11). Figure 13 As shown, the storage control block 100 establishes communication with the activated FE I / F 110 (S22).

[0110] Then, if Figure 11 As shown in FIG13, the storage control block 100 sends a port setting instruction to the FE I / F 110 (S23). The port setting instruction specifies the IP address and TCP port number of each port of the FE I / F 110. Figure 11 As shown in FIG. 12 , the FE I / F 110 receives a port setting instruction from the storage control block 100 and sets the IP address and TCP port number of each port ( S12 ).

[0111] When the port setting is completed, the FE I / F 110 sends a port setting completion notification to the storage control block 100 (S13).

[0112] like Figure 11 As shown in FIG13 , after receiving the port setting completion notification from the FE I / F 110 ( S24 ), the storage control block 100 transmits a port wait start instruction to the FE I / F 110 ( S25 ).

[0113] like Figure 11 As shown in FIG12 , the FE I / F 110 receives a port wait start instruction from the storage control block 100 ( S14 ) and waits for communication from the host server ( S15 ). For example, the FE I / F 110 starts the NVMe / TCP target software. Thereafter, the FE I / F 110 sends a port wait completion notification to the storage control block 100 ( S16 ).

[0114] like Figure 11 As shown in FIG. 13 , the storage control block 100 receives the port wait completion notification from the FE I / F 110 ( S26 ), and waits for an instruction from the operator and a request from the FE I / F 110 ( S27 ).

[0115] Next, we'll explain the connection establishment process between the host server 200 and the storage control block 100. In NVMe / TCP, the host server 200 and storage system 1 establish a management queue connection and an arbitrary number of IO queue connections for access (association) to a single subsystem. The management queue connection is established first, followed by several IO queue connections. The destination namespace (LDEV) is specified during IO (read / write) access.

[0116] Figure 14 This is a sequence diagram showing the management queue connection establishment process. Figure 14 The following shows the processing sequence when no error occurs during the processing. First, the FE I / F 110 receives a connection request from the host server 200 (S31). The connection request includes information stored in the connection management table T60 and information stored in the controller storage table T50.

[0117] For example, the information used in the connection management table T60 can include a queue ID, host IP address, host port number, and connection setting information. The controller ID is omitted when requesting a management queue connection. The queue ID is "0" for a management queue connection. The connection setting information indicates the KeepAliveTimeout period. Furthermore, the FE I / F 110 can obtain the target IP address and target port number from the storage control block 100.

[0118] The information used in the controller-held table T50 may include the subsystem NQN, host NQN, host ID, the requested value of the predetermined number of queues, and the number of connections (current number of I / O queues). In the case of a management queue connection, the number of connections is "0." Furthermore, the port ID is held as a set value within the FE I / F 110.

[0119] Next, the FE I / F 110 adds an entry to the connection management table T60 (S32). At this time, the controller ID is set to a value indicating that no entry has been made (for example, 0xffff). Furthermore, the FE I / F 110 sends a controller addition request to the storage control block (S33).

[0120] The controller addition request includes information stored in the controller management table T30. Specifically, the controller addition request includes the subsystem NQN, port ID, host NQN, host ID, protocol, number of queues, and number of connections. For management queue connections, the number of connections is "0." Furthermore, the protocol indicates the communication protocol between the host server 200 and the FE I / F. In this example, NVMe over TCP is assumed.

[0121] The storage control block 100 that has received the controller addition request adds a new entry to the controller management table T30, secures hardware resources such as a memory area and a CPU core, and sets a controller ID (S34).

[0122] The storage control block 100 searches the subsystem management table T10 for an entry whose subsystem NQN matches the subsystem NQN included in the controller add request, obtains the subsystem ID of this entry, and sets it in the controller management table T30. As described above, the protocol set is NVMe over TCP, and the number of connections set to manage queue connections is 0. The requested value is set as the predetermined queue number. Furthermore, if the requested value exceeds a pre-set maximum allowable value, the maximum allowable value may be set.

[0123] In this specification, retrieving an entry whose value of a specific field (in the above example, subsystem NQN) of an entry in a specific table (in the above example, subsystem management table T10) is consistent with the value assigned with the same name (in the above example, the subsystem NQN included in the controller addition request) as described above, and obtaining the value of the other field (in the above example, subsystem ID) of the entry is simply expressed as using a specific field value to obtain other values ​​from a specific table (in the above example, using subsystem NQN to obtain subsystem ID from subsystem management table T10), or retrieving a specific field value from a specific table to obtain other values ​​(in the above example, retrieving subsystem NQN from subsystem management table T10 to obtain subsystem ID).

[0124] Next, the storage control block 100 searches the namespace management table T20 for an entry whose subsystem ID matches the aforementioned subsystem ID, and obtains the namespace ID of the matching entry as a list of available namespace IDs (S35). Next, the storage control block 100 transmits a controller addition response containing specific information to the FE I / F 110 (S36).

[0125] The controller addition response includes the controller ID, the number of queues set, and a list of available namespace IDs. The namespace ID list is determined as a sequential number starting from 1, and the communication data volume can be reduced by returning the number of namespace IDs.

[0126] The FE I / F 110 that receives the response from the storage control block 100 adds an entry to the controller holding table ( S37 ), and sets the controller ID of the entry added to the controller holding table T50 to the corresponding entry in the connection management table T60 ( S38 ).

[0127] FE I / F 110 prepares the memory area, CPU core, and other hardware resources required for processing in the management queue connection (S39) and returns a connection completion response to the host server 200 (S40). The connection completion response includes the set controller ID, the set number of scheduled queues, and a list of available namespace IDs. In addition, if an error occurs during processing, an error response is sent to the host server 200. In addition, the management queue connection can also be established multiple times. In this case, it is treated as different associations and managed as connections with different controller IDs.

[0128] Figure 15 This is a sequence diagram showing the process of establishing an IO queue connection belonging to the association after the management queue connection is established. Figure 15 The following shows the processing sequence when no error occurs during the processing. First, the FE I / F 110 receives a connection request from the host server 200 (S45). The connection request includes information stored in the connection management table T60 and information stored in the controller storage table T50.

[0129] For example, the information used in the connection management table T60 can include a controller ID, queue ID, host IP address, host port number, and connection setting information. The controller ID uses the value returned when establishing a management queue connection. For I / O queue connections, the queue ID is an integer greater than or equal to "1." The connection setting information indicates the requested keep-alive timeout. Furthermore, the FE I / F 110 can obtain the target IP address and target port number from the storage control block 100.

[0130] The information for the controller storage table T50 can include the controller ID. By searching for an entry in the controller storage table T50 that matches the controller ID indicated by the host, the corresponding association can be determined.

[0131] Next, the FE I / F 110 adds an entry to the connection management table T60 ( S46 ) Next, the FE I / F 110 searches the controller storage table T50 for the designated controller ID ( S47 ).

[0132] If an entry with a matching controller ID exists in the controller retention table T50, the FE I / F 110 further transmits a connection addition request to the storage control block 100 (S48). The connection addition request includes information on the number of connections in addition to the controller ID. The number of connections is set to the value obtained by adding 1 to the number of connections in the entry with the matching controller ID. If the entry does not exist or the current number of connections exceeds the predetermined queue number, an error is reported.

[0133] Upon receiving the connection addition request, the storage control block 100 searches the controller management table T30 for the corresponding entry using the controller ID and updates the connection count of the entry corresponding to the controller ID to the connection count value included in the connection addition request from the FE I / F 110 (S49). The storage control block 100 then returns a normal response (S50). Alternatively, if the entry does not exist or the connection count exceeds the predetermined queue count due to the current I / O queue connection, an error response is returned.

[0134] The FE I / F 110 that has received the normal response from the storage control block 100 updates the number of connections in the controller retention table T50 in a manner consistent with the controller management table T30 ( S51 ).

[0135] FE I / F 110 prepares hardware resources such as memory areas and CPU cores to manage queue connections (S52) and returns a connection completion response to host server 200 (S53). The connection completion response includes the controller ID. In addition, if an error occurs during processing, an error response is sent to host server 200.

[0136] Next, refer to the flowchart for reference Figure 14 15 and the respective processes of the FE I / F 110 and the storage control block 100 in the connection establishment process described in FIG. Figure 14 And the same steps in step 15 are given different symbols. Figure 16A And 16B is a flowchart of the connection establishment process of the FE I / F 110.

[0137] The FE I / F 110 receives a connection request for a management queue connection or an IO queue connection from the host server 200 (S61). The information included in the connection request is as described above.

[0138] Next, the FE I / F 110 adds a corresponding entry to the connection management table T60 (S62). At this point, no controller ID is entered. Next, the FE I / F 110 determines whether the queue ID of the connection request is "0" (S63). If the queue ID is "0," it is a request for a connection to the management queue. If it is a larger integer, it is a request for a connection to the IO queue.

[0139] If the queue ID is "0", that is, the connection request is for a management queue connection (S63: Yes), a controller addition request is sent to the storage control block 100, and a controller addition response including the controller ID is received from the storage control block 100 (S64). The controller addition request and the controller addition response include information such as that in the reference. Figure 14 As stated.

[0140] If the controller addition response is an error response (S65: Yes), the FE I / F 110 returns an error response to the host server 200 (S66). If the controller addition response is not an error response (S65: No), the FE I / F 110 adds an entry to the controller retention table T50 (S67).

[0141] Furthermore, the FE I / F 110 sets the controller ID added to the entry in the controller retention table T50 to the corresponding entry in the connection management table T60 (S68). The FE I / F 110 prepares hardware resources such as memory areas and CPU cores to manage the queue connection (S69), and returns a connection completion response to the host server 200 (S70). The information included in the connection completion response is as shown in FIG. Figure 14 As described.

[0142] In step S63, if the queue ID is an integer greater than "0", that is, the connection request is a request for an IO queue connection (S63: No), the flow enters the connection via connector A. Figure 16B In step S71, the FE I / F 110 searches the controller holding table T50 for the controller ID.

[0143] If the controller ID does not exist (S71: No), the FE I / F 110 returns an error response to the host server 200 (S78). If the controller ID exists in the controller storage table T50 (S71: Yes), the FE I / F 110 determines whether the value obtained by adding 1 to the number of connections indicated by the connection request (current IO queue number) exceeds the predetermined queue number (S72).

[0144] If the value of the number of connections indicated by the connection request plus one exceeds the predetermined number of queues (S72: Yes), the FE I / F 110 returns an error response to the host server 200 (S78). Alternatively, the storage system-triggered disconnection process described later may be executed. If the value of the number of connections indicated by the connection request plus one is less than the predetermined number of queues (S72: No), the FE I / F 110 sends a connection addition notification to the storage control block 100 and receives its response (S73).

[0145] If an error response is received from the storage control block 100 (S74: YES), the FE I / F 110 returns the error response to the host server 200 (S78). If a normal response is received from the storage control block 100 (S74: NO), the FE I / F 110 updates the number of connections in the controller retention table T50 (S75).

[0146] Next, the FE I / F 110 prepares hardware resources such as memory areas and CPU cores for managing the queue connection (S76), and returns a connection completion response to the host server 200 (S77). The information included in the connection completion response is as shown in FIG. Figure 15 As stated.

[0147] Next, the processing of the storage control block 100 will be described. Figure 17 1 is a flowchart of the connection establishment process of the storage control block 100. Figure 17 The symbols in Figure 16B Some of the symbols in are repeated, but Figure 17 and Figure 16B Even if the same symbols are used, they are different steps. The storage control block 100 receives a request from the FE I / F 110 (S75) and determines whether the request is a controller addition request or a connection addition notification (S76). The information included in each addition request is as shown in FIG. Figure 14 and Figure 15 As stated.

[0148] If the received request is a controller addition request (S76: controller addition request), the storage control block 100 searches the subsystem management table T10 for an entry whose subsystem NQN matches the subsystem NQN included in the controller addition request, and obtains the subsystem ID of the entry (S77).

[0149] Next, the storage control block 100 adds a new entry to the controller management table T30, securing hardware resources such as memory areas and CPU cores, and setting the controller ID (S78). The protocol used in this case is assumed to be NVMe over TCP. The storage control block 100 sets the number of connections to "0" and the requested value for the number of queues. However, if the requested value exceeds the maximum allowed value, the maximum allowed value may be set.

[0150] Next, the storage control block 100 searches the namespace management table T20 for the subsystem ID and obtains a list of available namespace IDs (S79). The storage control block 100 then returns a controller addition response to the FE I / F 110 (S80). The controller addition response includes the controller ID, the set number of queues, and a list of available namespace IDs.

[0151] In step S76, if the received request is a connection addition request (S76: Connection Addition Request), the storage control block 100 searches the controller management table T30 for a corresponding entry using the controller ID (S81). If no entry exists (S82: No), the storage control block 100 returns an error response to the FE I / F 110 (S87).

[0152] If an entry exists (S82: YES), the storage control block 100 compares the predetermined queue number of the corresponding entry with the connection number (current IO queue number) indicated by the request from the FE I / F 110 (S83). If the value obtained by adding 1 to the connection number exceeds the predetermined queue number, that is, if the IO queue number exceeds the predetermined number due to the current IO queue connection (S84: YES), the storage control block 100 returns an error response to the FE I / F 110 (S87).

[0153] If the number of IO queues does not exceed the predetermined number due to the current IO queue connection (S84: No), the storage control block 100 updates the connection number of the corresponding entry in the controller management table T30 to the value obtained by adding 1 to the connection number (S85) and returns a normal response to the FEI / F 110 (S86).

[0154] Next, the processing of the storage system 1 in response to IO access (read access or write access) from the host server 200 will be described. Figure 18 This is a timing diagram showing IO access processing.

[0155] The FE I / F 110 receives an IO command from the host server 200 (S91). The IO command is a read command or a write command. The IO command includes the command type (read or write), namespace ID, access destination address, access destination size, and write data in the case of a write command.

[0156] The FE I / F 110 obtains a list of controller IDs and namespace IDs corresponding to the IO command from the connection management table T60 ( S92 ), adds the obtained controller ID, and transfers the IO command to the storage control block 100 ( S93 ).

[0157] Upon receiving the controller ID and IO command from the FE I / F 110, the storage control block 100 obtains the corresponding subsystem ID from the controller management table T30 (S94). Furthermore, the storage control block 100 searches the namespace management table T20 for the combination of the subsystem ID and the namespace ID, and obtains the corresponding LDEV ID (S95).

[0158] Next, the storage control block 100 performs an access permission determination on the access destination address and the access destination size (S96). This determination is a general storage access control that determines user authority and whether other host servers are currently in use.

[0159] Next, the storage control block 100 returns a response to the FE I / F 110 (S97). If the access permission determination is an error, an error response is returned. If the access permission determination is not an error, read data is returned for a read command, and a success response is returned for a write command.

[0160] The FE I / F 110 receives the response from the storage control block 100 and forwards it to the host server 200 ( S98 ).

[0161] Alternatively, the host server 200 may send a write command without write data to the FE I / F 110 and, after receiving a write-permit response, send a write command with write data. In this case, the above sequence is executed twice.

[0162] Next, refer to the flowchart for reference Figure 18 The following description will be made of the respective processes of the FE I / F 110 and the storage control block 100 in the IO access process. Figure 18 The same steps in the steps are given different symbols.

[0163] Figure 19 This is a flowchart of the IO access processing of the FE I / F 110. First, the FE I / F 110 receives an IO command from the host server (S101). The information contained in the IO command is as shown in FIG. Figure 18 Next, the FE I / F 110 obtains a list of controller IDs and namespace IDs corresponding to the IO commands from the connection management table T60 ( S102 ).

[0164] When the namespace ID of the IO command is not included in the acquired namespace ID list ( S103 : No), the FEI / F 110 returns an error response to the host server 200 ( S107 ).

[0165] If the obtained namespace ID list contains the namespace ID of the IO command (S103: Yes), the FE I / F 110 sends the obtained controller ID and IO command to the storage control block 100 (S104). The FE I / F 110 then receives a response from the storage control block 100 (S105) and returns it to the host server 200 (S106). As described above, the response can be an error response or a normal read or write response. A normal read response includes the read data.

[0166] Figure 20 This is a flowchart of the IO access process of the storage control block 100. The storage control block 100 receives an IO command and a controller ID from the FE I / F 110 (S111). Next, the storage control block 100 obtains the corresponding subsystem ID from the controller management table T30 (S112).

[0167] Next, the storage control block 100 searches the namespace management table T20 for a combination of the subsystem ID and the namespace ID, and acquires the corresponding LDEV ID ( S113 ).

[0168] Next, the storage control block 100 determines whether the access destination address and access destination size can be accessed by the IO command (S114). This step is a general storage access control step to determine user authority, whether other hosts are using the address, etc.

[0169] If an access error occurs (S114: No), the storage control block 100 returns an error response to the FE I / F 110 (S118). If access is permitted (S114: Yes), the storage control block 100 determines whether the IO command is a read command or a write command (S115).

[0170] If the IO command is a read command (S115: Read), the storage control block 100 reads data of the access target size from the access target address of the LDEV indicated by the LDEV ID and returns it as an IO result to the FE I / F 110 (S116). If the IO command is a write command (S115: Write), the storage control block 100 writes data of the access target size to the access target address of the LDEV indicated by the LDEV ID and returns a success response as an IO result to the FE I / F 110 (S117).

[0171] The following describes the connection disconnection process triggered by the host server 200. Figure 21 This is a sequence diagram of the connection disconnection process triggered by the host server. The FE I / F 110 receives a connection disconnection request from the host server (S121). Alternatively, the host server 200 may disconnect the TCP connection.

[0172] Next, the FE I / F 110 obtains the controller ID of the connection from the connection management table T60 (S122). The FE I / F 110 sends a disconnection completion notification to the host server 200 to disconnect the connection (S123). If the disconnection of the TCP connection is the trigger, this step is not necessary.

[0173] Next, the FE I / F 110 deletes the corresponding entry in the connection management table T60 to release resources (S124). Next, the FE I / F 110 sends a connection deletion request to the storage control block 100 (S125). The connection deletion request includes the controller ID and the updated number of connections.

[0174] The storage control block 100, which has received the connection deletion request from the FE I / F 110, updates the connection count of the entry with the same controller ID in the controller management table T30 (S126). Here, the connection count is reduced by 1. Thereafter, the storage control block 100 sends a completion notification to the FE I / F 110 (S127).

[0175] The FE I / F 110 that has received the completion notification updates the number of connections (current queue number) in the controller retention table T50 ( S128 ). Here, the value of the number of connections is reduced by 1.

[0176] When the number of connections (current IO queue number) in the controller retention table T50 is 0, the FE I / F 110 sends a controller deletion request to the storage control block 100 (S129). The controller deletion request includes the target controller ID.

[0177] Upon receiving the controller deletion request, the storage control block 100 deletes the entry with the corresponding controller ID from the controller management table T30, releasing resources (S130). The storage control block 100 then sends a completion notification to the FE I / F 110 (S131). Upon receiving the completion notification, the FE I / F 110 deletes the corresponding entry from the controller retention table T50, releasing resources (S132).

[0178] Next, refer to the flowchart for reference Figure 21 The following description will describe the processing of the FEI / F 110 and the storage control block 100 in the connection disconnection process triggered by the host server. Figure 21 The same steps in the steps are given different symbols.

[0179] Figure 22 This is a flowchart of the connection disconnection process triggered by the host server of the FE I / F 110. First, the FE I / F 110 receives a connection disconnection request from the host server (S141). As described above, the disconnection of the TCP connection by the host server 200 can be a trigger.

[0180] Next, the FE I / F 110 obtains the controller ID of the connection from the connection management table T60 (S142). The FE I / F 110 sends a disconnection completion notification to the host server 200 and disconnects the connection (S143). This step is unnecessary when the disconnection of the TCP connection is the trigger.

[0181] Next, the FE I / F 110 deletes the corresponding entry in the connection management table T60 to release resources (S144). Next, the FE I / F 110 sends a connection deletion request to the storage control block 100 and receives a completion notification (S145). The connection deletion request includes the controller ID and the updated number of connections.

[0182] Upon receiving the completion notification, the FE I / F 110 updates the connection number (current queue number) in the controller table T50 ( S128 ). Here, the connection number is reduced by 1. The FE I / F 110 determines whether the connection number (current IO queue number) in the controller table T50 is 0 ( S147 ).

[0183] If the number of connections is greater than 0 (S147: No), this process ends. If the number of connections is 0 (S147: Yes), the FE I / F 110 sends a controller deletion request to the storage control block 100 and receives a completion notification from the storage control block 100 (S148). The controller deletion request includes the target controller ID. The FE I / F 110 then deletes the corresponding entry from the controller retention table T50, releasing resources (S149).

[0184] Figure 23 This is a flowchart of the connection disconnection process triggered by the host server of the storage control block 100. The storage control block 100 receives a connection deletion request from the FE I / F 110 (S151). The connection deletion request indicates the controller ID and the updated number of connections. The storage control block 100 updates the connection number of the entry with the same controller ID in the controller management table T30 (S152). Here, the connection number is decremented by 1. The storage control block 100 then sends a completion notification to the FE I / F 110 (S153).

[0185] The storage control block 100 then receives a controller deletion request (S154). The controller deletion request indicates the controller ID. The storage control block 100 deletes the entry with the matching controller ID from the controller management table T30, releasing resources (S155). The storage control block 100 then sends a completion notification to the FE I / F 110 (S156).

[0186] The following describes the disconnection process triggered by the storage system 1 . Figure 24 This is a timing diagram of the storage system-triggered disconnection process. In response to an operator action or an error within the storage system 1, the storage control block 100 invalidates the entry with the corresponding controller ID in the controller management table T30 (S161). This entry can be invalidated by, for example, setting the predetermined queue count to a negative number, adding an invalidation flag to the entry, or temporarily deleting the entry. Next, the storage control block 100 sends a disconnection request specifying the controller ID to the FE I / F 110 (S162).

[0187] Upon receiving the connection disconnection request, the FE I / F 110 searches the connection management table T60 for a connection with a matching controller ID ( S163 ) and transmits a connection disconnection notification to all matching host servers 200 to disconnect the connection ( S164 ).

[0188] Next, the FE I / F 110 deletes the entry with the same controller ID from the connection management table T60, releasing the resource (S165). Furthermore, the FE I / F 110 deletes the entry with the same controller ID from the controller retention table T50, releasing the resource (S166). The FE I / F 110 then sends a disconnection completion notification to the storage control block 100 (S167).

[0189] Upon receiving the disconnection completion notification from the FE I / F 110 , the storage control block 100 deletes the entry with the same controller ID from the controller management table T30 ( S168 ).

[0190] Next, refer to the flowchart for reference Figure 24 The following description will describe the processing of the FE I / F 110 and the storage control block 100 in the disconnection process triggered by the storage system. Figure 24 The same steps in the steps are given different symbols.

[0191] Figure 25 This is a flowchart of the storage system-triggered disconnection process performed by the FE I / F 110. The FE I / F 110 receives a disconnection request including a controller ID from the storage control block 100 (S171). The FE I / F 110 searches the connection management table T60 for connections with matching controller IDs (S172). The FE I / F 110 sends disconnection notifications to all matching host servers 200 and then disconnects the connections (S173).

[0192] Next, the FE I / F 110 deletes the entry with the same controller ID from the connection management table and releases the resource (S174). Next, the FE I / F 110 deletes the entry with the same controller ID from the controller retention table and releases the resource (S175). The storage control block is notified of the disconnection completion (S176).

[0193] Figure 26 This is a flowchart of the storage system-triggered disconnection process performed by the storage control block 100. In response to an operator action or an error within the storage system 1, the storage control block 100 invalidates the entry with the same controller ID in the controller management table T30 (S181). This invalidation can be accomplished by, for example, setting the predetermined queue number to a negative number, adding an invalid flag to the entry, or temporarily deleting the entry. If the entry has already been deleted, updating the controller management table T30 is omitted.

[0194] Next, the storage control block 100 sends a disconnection request specifying the controller ID to the FE I / F 110 (S182). The storage control block 100 then receives a disconnection completion notification from the FE I / F 110 (S183). The storage control block 100 then deletes the entry with the matching controller ID from the controller management table T30 (S184). If the entry has already been deleted, deletion is omitted.

[0195] Various modifications can be made to the above embodiment. For example, the connection number information can be omitted from the controller management table T30. In this case, communication of information related to the connection number, processing related to updating the connection number to the storage control block 100 in the FE I / F 110, and processing related to the connection number in the storage control block 100 can be omitted between the FE I / F 110 and the storage control block 100. As a result, when establishing an I / O queue connection, communication between the FE I / F 110 and the storage control block 100 is omitted, further shortening the time it takes to establish a connection (association) with the controller.

[0196] In addition to the number of connections, the number of reservation queues (the number of allowed queues) in the controller management table T30 can also be omitted. In this case, during initialization, the storage control block 100 notifies the FE I / F 110 of the number of reservation queues, and the FE I / F 110 subsequently manages the number of reservation queues. As a result, the management of the connection (association) to the controller is performed by the FE I / F 110, which can reduce the involvement of the storage control block 100 and simplify the processing of the FE I / F 110. However, since the storage control block 100 cannot manage the number of reservation queues, if the number of reservation queues is changed midway, an instruction from the storage control block 100 to the FE I / F 110 is required.

[0197] In the above process description, the order of the steps can be changed within the scope of maintaining compatibility, and the previous and subsequent steps can be combined to simplify the process or reduce the number of communications. In addition, if there is no need to release resources immediately due to surplus resources, the release of resources can also be omitted within the scope of maintaining compatibility.

[0198] In this embodiment, the NVMe over TCP communication protocol is used between the host server 200 and the FE I / F 110, and DMA communication is used between the FE I / F 110 and the storage control block 100. However, different communication methods may be used to optimize required computer resources or facilitate implementation.

[0199] For example, additional development and processing can be suppressed by adopting communication methods supported by the host server and storage control block. For example, NVMe over FC or iSCSI can be used between the host server 200 and the FE I / F 110, and UDP (User Datagram Protocol) can be used between the FE I / F 110 and the storage control block 100. Furthermore, when using other communication methods, information specific to that communication method is used. For example, in iSCSI, the iSCSI Qualified Name is used instead of the subsystem NQN.

[0200] The information transmission method is not limited to the above examples. For example, the amount of data transmitted in a single communication can be reduced by segmenting specific information, or error detection can be performed sequentially to improve processing reliability. Information can also be sent multiple times to create redundancy, or mismatches can be checked during communication to improve fault tolerance. Furthermore, different information can be combined (for example, by concatenating the host NQN and host ID as a string) to reduce the number of communications.

[0201] In the above embodiment, the FE I / F 110 exists between the host server 200 and the storage control block 100. Therefore, when information transmission is divided into multiple stages, communication becomes two stages. There are two methods for processing the divided information, and either method can be used.

[0202] In one method, the FE I / F 110 performs multiple communications between the host server 200 and the storage control block 100, combines the divided information, and transmits it to the other party. This simplifies communication. In another method, the FE I / F 110 transmits information received from either the host server 200 or the storage control block 100 sequentially to the other party. This minimizes latency.

[0203] If the information can be determined from other information or is not essential for the communication method, the information can be sent in another format, or only a portion of the information can be sent. For example, if the subsystem NQN can be restored from the table that manages the controller ID by simply transmitting the controller ID, only the controller ID can be sent. This can reduce the amount of communication data and communication processing.

[0204] The response does not necessarily contain all the information in the above examples. For example, if an error occurs, the error is notified. By attaching a request ID to each request, it is possible to distinguish which request the response or data is for. In addition, if the processing content does not change due to disconnection, deletion, etc. of the response content to the communication, the sending of the response can be omitted. By omitting the sending of the response, deadlock caused by waiting for the response can be avoided, the time until the processing is completed can be shortened, and the processing can be simplified. By sending the response, if an error occurs at either end of the communication, the matching of the status of the two ends of the communication can be maintained.

[0205] Example 2

[0206] The following describes embodiment 2 of the present invention. The processing of the FE I / F 110 in the host server-triggered disconnection process in this embodiment differs from that in embodiment 1. In the following description, the description of embodiment 1 applies to configurations not specifically mentioned.

[0207] Figure 27 This flowchart shows the host server-triggered disconnection process of the FE I / F 110. The FE I / F 110 deletes the controller in one disconnection (all disconnection). This simplifies the processing of the FE I / F 110, reduces the load, and improves performance.

[0208] Reference Figure 27 Steps S141 to S143 are the same as those in Example 1. Figure 22 After step S143, the FE I / F 110 sends a disconnection notification to the host server 200 for all entries of the controller ID in the connection management table T60, disconnecting all connections (S191). The FE I / F 110 deletes all entries of the controller ID from the connection management table T60, releasing resources (S192). The subsequent steps S148 and S149 are the same as those of Example 1. Figure 22 The flowchart shown is the same.

[0209] Example 3

[0210] The following describes a third embodiment of the present invention. This embodiment adds a connection ID to the structure of the connection management table T60 in the first embodiment. The information transmitted from the FE I / F 110 to the storage control block 100 includes the connection ID. When the storage control block 100 logs the processing, it includes the connection ID notified from the FE I / F 110 in the log. This makes it easy to track connections that require a long time or connections that have errors. The remaining structure is the same as in the first embodiment.

[0211] Example 4

[0212] The following describes a fourth embodiment of the present invention. This embodiment manages the CPU core responsible for connection in the FE I / F 110. The following describes differences from the first embodiment. The description of the first embodiment applies to configurations not specifically mentioned.

[0213] Figure 28 The following shows an example of the structure of the connection management table T90 of this embodiment. The connection management table T90 includes a responsible core column C609 in addition to the columns C601 to C608 of the connection management table T60 of embodiment 1. The responsible core column C609 indicates the identifier of the CPU core that executes the processing for the corresponding connection.

[0214] The FE I / F 110 includes the responsible core information in its request to the storage control block 100. When the storage control block 100 logs the process, the storage control block 100 includes the responsible core information (including the core ID) notified from the FE I / F 110 in the log. This information may be omitted.

[0215] The storage control block 100 includes the responsible core ID notified from the FE I / F 110 in the response indicating the processing result sent to the FE I / F 110. In the case of "storage control block-triggered shutdown", the storage control block 100 adds an invalid value to the response.

[0216] When the FE I / F 110 processes the response from the storage control block 100 , the FE I / F 110 refers to the included responsible core ID and sets the core to be processed by the FE I / F 110 .

[0217] According to this embodiment, during IO access, the FE I / F 110 can immediately allocate processing to the core responsible for the IO access based on the response from the storage control block 100. This can equalize processing variations between cores, improving cache efficiency and latency. Furthermore, as in Example 3, the information transmitted between the FE I / F 110 and the storage control block 100 can include a connection ID instead of the responsible core ID. The FE I / F 110 can identify the responsible core by referring to the connection management table T90. This structure further simplifies the implementation of both Examples 3 and 4.

[0218] Furthermore, the present invention is not limited to the above-described embodiments and encompasses various variations. For example, the above-described embodiments are examples described in detail to facilitate understanding of the present invention and are not necessarily limited to having all of the described structures. Furthermore, a portion of the structure of one embodiment can be replaced with a structure of another embodiment, and a structure of another embodiment can be added to a structure of one embodiment. Furthermore, with respect to a portion of the structure of each embodiment, other structures can be added, deleted, or substituted.

[0219] Furthermore, the aforementioned structures, functions, processing units, etc. may be partially or entirely implemented in hardware, for example, by designing them using integrated circuits. Furthermore, the aforementioned structures, functions, etc. may be implemented in software by having a processor interpret and execute programs that implement the respective functions. Information such as programs, tables, and files that implement the respective functions may be stored in a storage device such as a memory, a hard disk, or an SSD, or a recording medium such as an IC card or an SD card.

[0220] In addition, the control lines and information lines are those considered necessary for explanation and do not necessarily represent all the control lines and information lines on the product. In fact, it can be assumed that almost all the components are connected to each other.

[0221] Explanation of symbols

[0222] 1 Storage system,

[0223] 103CPU,

[0224] 104 storage areas,

[0225] 110FE I / F,

[0226] 113CPU,

[0227] 114 storage areas,

[0228] 200 host servers,

[0229] T10 subsystem management table,

[0230] T20 namespace management table,

[0231] T30 controller management table,

[0232] T50 controller holds table,

[0233] T60 connection management table.

Claims

1. A storage system that communicates with a host in a session comprising one or more connections, characterized in that: The storage system includes a front-end interface, a processor and a storage area. The storage area stores session management information for managing a communication session with the host. The front-end interface stores connection management information for managing the connection of the session, The front-end interface controls access from the host with reference to the connection management information.

2. The storage system according to claim 1, wherein: The front-end interface stores cached session maintenance information including information contained in the session management information, The front-end interface adds a new entry to the connection management information during the initial connection establishment of the session. The front-end interface sends information related to the initial connection to the processor, the processor adding an entry for a new session including the initial connection to the session management information, The processor sends the information of the new session to the front-end interface, The front-end interface updates the session maintenance information according to the received information of the new session, The front-end interface controls access from the host with reference to the connection management information and the session maintenance information.

3. The storage system according to claim 2, wherein: The session maintenance information manages the predetermined number of connections in the session, The front-end interface returns an error to the host when the connection establishment requested from the host exceeds the predetermined number of connections.

4. The storage system according to claim 1, wherein: In disconnecting the connection from the host to the front-end interface, The front-end interface obtains the corresponding session information from the connection management information. The front-end interface sends a disconnection completion notification to the host, The front-end interface specifies the corresponding session to the processor and sends a disconnection request for the connection. The processor updates the session management information according to the received disconnection request.

5. The storage system according to claim 1, wherein: In disconnecting the connection from the host to the front-end interface, The front-end interface obtains the corresponding session information from the connection management information. The front-end interface sends a disconnection completion notification to the host, When the number of IO connections of the corresponding session is 0, the front-end interface sends a deletion request of the corresponding session to the processor. The processor deletes the information of the corresponding session from the session management information.

6. The storage system according to claim 1, wherein: The front-end interface includes a connection ID managed by the front-end interface in a request to the processor, The processor includes the connection ID in log information.

7. The storage system according to claim 1, wherein: The front-end interface sets the responsible core for the connection, The front-end interface includes information about the responsible core of the connection in a request to the processor, The processor includes the information of the responsible core in the processing result sent to the front-end interface, The front-end interface continues processing the connection through the responsible core.

Citation Information

Patent Citations

  • Network interface, and buffer control method thereof

    JP2023142021A