Adapter storage system and method

By employing a multi-level memory architecture in the storage network, combining on-chip and host memory, and optimizing memory usage, the problem of low memory resource utilization in the storage network is solved, achieving more efficient memory resource management and performance improvement.

CN115225713BActive Publication Date: 2025-10-28AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210208609.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-16
Filing Date
2022-03-04
Publication Date
2025-10-28
Estimated Expiration
2042-03-04

AI Technical Summary

Technical Problem

In existing storage networks, memory resources suffer from high costs, high power consumption, and increased bandwidth requirements when handling incomplete I/O swaps, especially as the number of incomplete I/O swaps increases, making it impossible to effectively utilize network/link resources.

Method used

It adopts a multi-level memory architecture, which combines on-chip memory and host memory. By receiving swap resource indicator requests at the adapter and providing swap resource indicators from different storage devices according to their range, it optimizes memory usage by utilizing fixed operation and host support memory units, thereby reducing the burden on on-chip memory.

Benefits of technology

It effectively reduces the cost and power consumption of memory resources, while improving network/link utilization, thus improving system performance and scalability without increasing the size of on-chip memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115225713B_ABST
    Figure CN115225713B_ABST
Patent Text Reader

Abstract

This application relates to an adapter storage system and method. The system and method relate to a bus adapter for a storage network. The bus adapter includes a context memory comprising a first storage device for non-cacheable interchangeable resource indicators (XRIs) and a second storage device for cacheable XRIs. The bus adapter also includes a host support storage unit configured to use a plurality of cache subunits and provide access to different tiers of memory, either locally present or externally present in a host memory extension, based on at least one of the following: input / output phase, first-in-first-out constraint, region of virtual context address associated with the cacheable XRI, protocol associated with the cacheable XRI, transaction size, or work queue information. The host support storage unit also provides the ability to perform optional fixed operations on the cacheable XRIs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to computer networks, storage networks, and communications. Some embodiments of this disclosure relate to systems and methods for more efficiently storing interchangeable contexts in a network. Background Art

[0002] Over the past few decades, the market for network communication devices has grown exponentially, driven by the use of portable devices and increased connectivity and data transfer between various devices. Digital switching technologies have facilitated the large-scale deployment of affordable, easy-to-use communication networks that include storage networks (e.g., Storage Area Networks (SANs)). Wireless communication can operate according to various standards such as IEEE 802.11x, IEEE 802.11ad, IEEE 802.11ac, IEEE 802.11n, IEEE 802.11ah, IEEE 802.11aj, IEEE 802.16 and 802.16a, Bluetooth, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), and cellular technologies.

[0003] A SAN connects computer data storage devices to servers in a commercial data center. SANs can use the Fibre Channel (FC) protocol, a high-speed data transfer protocol that provides ordered, lossless delivery of raw data blocks. Offloading storage input / output (I / O) switching requires an FC adapter or other type of network adapter to track the state of each I / O switch within a set of switching contexts. Storage switching contexts utilize valuable on-chip memory. Summary of the Invention

[0004] In one aspect, this application relates to a method of communicating in a storage network, the method comprising: receiving a request for a swap resource indicator at an adapter for the storage network; and if the swap resource indicator is in a first range, providing the swap resource indicator from a first storage device in the adapter; if the swap resource indicator is in a second range and stored in a second cache in the adapter, providing the swap resource indicator from the second cache; if the swap resource indicator is in the second range and not stored in the second cache, providing the swap resource indicator from host memory; and if the swap indicator is in the second range and a page address translation entry indicates a remapping operation with a direct context memory address, providing the swap resource indicator from context memory.

[0005] In another aspect, this application relates to a bus adapter for a storage network, comprising: a context memory including a first storage device for non-cacheable interchangeable resource indicators and a second storage device for cacheable interchangeable resource indicators; and a host support storage unit configured to provide fixed operation for the cacheable interchangeable resource indicators based on client requests, each client being fixed based on at least one of the following: an input / output phase, a first-in-last-out constraint, a region of a virtual context address associated with the cacheable interchangeable resource indicator, a protocol associated with the cacheable interchangeable resource indicator, the size of a transaction associated with the cacheable interchangeable resource indicator, or work queue information associated with the cacheable interchangeable resource indicator.

[0006] In another aspect, this application relates to a host bus adapter for a storage network including a host having host memory, the host bus adapter comprising: a context memory including a first storage device for non-cacheable swap resource indicators and a second storage device for cacheable swap resource indicators; and a host support storage unit disposed on the host bus adapter and configured to: provide the swap resource indicator from the first storage device in response to a request if the swap resource indicator is within a first range; provide the swap resource indicator from the second cache in response to the request if the swap resource indicator is within a second range and stored in a second cache in the adapter; provide the swap resource indicator from the host memory in response to the request if the swap resource indicator is within the second range and not stored in the second cache; and provide the swap resource indicator from the context memory if the swap resource indicator is within the second range and a page address translation entry indicates a remapping operation with a direct context memory address. Attached Figure Description

[0007] The various objects, aspects, features, and advantages of the invention will become more apparent and better understood through a detailed description taken in conjunction with the accompanying drawings, in which similar reference characters identify corresponding elements throughout. In the drawings, similar reference numerals generally indicate identical, functionally similar, and / or structurally similar elements.

[0008] Figure 1A It is a block diagram depicting an embodiment of a network environment including one or more access points communicating with one or more wireless devices or stations;

[0009] Figure 1B and 1CThis is a block diagram depicting embodiments of computing devices that can be used in conjunction with the methods and systems described herein;

[0010] Figure 2 It is a block diagram of a network configured to use storage technology according to some embodiments; and

[0011] Figure 3 It is according to some embodiments for use similar to Figure 3 The diagram illustrates the host support storage unit and context memory of the network. Detailed Implementation

[0012] The following IEEE standards, including any draft versions of such standards, are hereby incorporated herein by reference in their entirety and are part of this disclosure for all purposes: IEEE 802.3, IEEE 802.11x, IEEE 802.11ad, IEEE 802.11ah, IEEE 802.11aj, IEEE 802.16 and 802.16a, and IEEE 802.11ac. Furthermore, while this disclosure may refer to aspects of these standards, it is in no way limited by these standards. Some standards may relate to Storage Area Networks (SANs) for connecting computer data storage devices to servers in commercial data centers. SANs may use Fibre Channel (FC) standards / protocols, Small Computer System Interface (SCSI) interface standards / protocols, Asynchronous Transfer Mode (ATM) protocols, and Synchronous Optical Networking Protocol (SONET), all of which are incorporated herein by reference in their entirety.

[0013] For the purpose of reading the descriptions of the various embodiments below, the following descriptions of the paragraphs and their corresponding contents may be helpful. Paragraph A describes a network and computing environment that can use data from a SAN practiced in the embodiments described herein, while Paragraph B describes embodiments of systems and methods for reducing on-chip memory requirements for I / O switching contexts and for on-chip architectures for host-supported storage.

[0014] Some embodiments of the systems and methods may utilize Fibre Channel (FC) or network adapters to track the status of each I / O switch within a set of switching contexts. As bandwidth requirements and latency increase in the system, the number of incomplete I / O switches increases to fill the network / system pipeline to maximize network / link utilization. Therefore, a large number of switching contexts are maintained to offload incomplete I / O switches on the adapter. Some embodiments of the systems and methods reduce the costs associated with memory resources used to offload a large number of incomplete I / O switches (in terms of die size, cost, and power). In some embodiments, on-chip memory usage is extended into host memory without increasing the size of the on-chip memory, and the on-chip memory usage employs a multi-tiered memory architecture for multi-level switch offloading.

[0015] In some embodiments, a first-level pool provides multiple non-cacheable swap contexts that are always on-chip and optimized for performance with low-latency operation, while a second-level pool provides multiple non-cacheable swap contexts that may be on-chip or in host memory for scalability. In some embodiments, multiple contexts may be pinned on-chip to eliminate or reduce cache misses during I / O operations. The tier selection can be done manually or automatically.

[0016] Some embodiments relate to a method of communicating in a storage network. The method includes: receiving a request for a swap resource indicator (XRI) at an adapter for the storage network; providing the swap resource indicator from a first storage device in the adapter if the swap resource indicator is in a first range; and providing the swap resource indicator from a second cache in the adapter if the swap resource indicator is in a second range and stored in a second cache. The method further includes: providing the swap resource indicator from host memory if the swap resource indicator is in a second range and not stored in the second cache.

[0017] In some embodiments, the method uses a pinning operation. The pinning operation may involve one or more of the following: pinning the swap resource indicator in the second cache for a specific input / output swapping phase; pinning the swap resource indicator in the second cache on a first-come, first-served basis if the pinning limit is not met; pinning the swap resource indicator in the second cache based on the address if the request is within a specific programmable virtual context address (VCA) range; pinning the swap resource indicator in the second cache within a specific programmable PIN I / O size range based on the request size; pinning the swap resource indicator in the second cache within a specific XRI range based on the XRI number associated with the request; pinning the swap resource indicator in the second cache based on the Fibre Channel protocol; pinning the swap resource indicator in the second cache based on WQ profile configuration; or pinning the swap resource indicator in the second cache based on a PIN bit booted by the host driver.

[0018] Some embodiments relate to a bus adapter for a storage network. The bus adapter includes a context memory comprising a first storage device for non-cacheable swap resource indicators and a second storage device for cacheable swap resource indicators. The bus adapter further includes a host support storage (HBS) unit configured to provide fixed operation for the cacheable swap resource indicators based on at least one of the following: an input / output phase, a first-in-first-out constraint, a region of the virtual context address associated with the cacheable swap resource indicator, a protocol associated with the cacheable swap resource indicator, the size of the transaction associated with the cacheable swap resource indicator, or work queue information associated with the cacheable swap resource indicator.

[0019] Some embodiments relate to a host bus adapter for a storage network including hosts. The host bus adapter includes a context memory containing a first storage device for non-cacheable swap resource indicators and a second storage device for cacheable swap resource indicators. The host bus adapter also includes a host support storage unit configured to provide swap resource indicators from the first storage device in response to a request if the swap resource indicators are within a first range. The host support storage unit is also configured to provide swap resource indicators from a second cache in the adapter in response to a request if the swap resource indicators are within a second range and stored in a second cache. The host support storage unit is further configured to provide swap resource indicators from host memory in response to a request if the swap resource indicators are within the second range and not stored in the second cache.

[0020] A. Computing and Networking Environment

[0021] Before discussing specific embodiments of this solution, it may be helpful to describe the operating environment and associated system components (e.g., hardware elements) in conjunction with the methods and systems described herein. References Figure 1A This describes an embodiment of a network environment. The network includes or communicates with a SAN, security adapter, or Ethernet aggregation network adapter (CAN). In brief, the network environment includes a wireless communication system comprising one or more access points 106, one or more wireless communication devices 102, and network hardware components 192. The wireless communication device 102 may, for example, include a laptop computer 102, a tablet computer 102, a personal computer 102, a wearable device 102, a vehicle 102 (e.g., a car, drone, smart vehicle, robotic unit, etc.), and / or a cellular phone device 102. Reference Figure 1B and 1C Details of some embodiments of the wireless communication device 102 and / or access point 106 are described in more detail below. In one embodiment, the network environment may be an ad hoc network environment, an infrastructure wireless network environment, a wired network coupled to a wireless network, a subnet environment, etc.

[0022] Access points (APs) 106 are operatively coupled to network hardware 192 via a local area network (LAN) connection. Network hardware 192, which may include routers, gateways, switches, bridges, modems, system controllers, appliances, etc., provides LAN connectivity for the communication system. Each of the access points 106 may have an associated antenna or antenna array to communicate with wireless communication devices in its area. Wireless communication devices may register with a specific access point 106 to receive services from the communication system (e.g., via SU-MIMO or MU-MIMO configuration). For direct connections (i.e., point-to-point communication), some wireless communication devices may communicate directly via an assigned channel and communication protocol. Some of the wireless communication devices 102 may be mobile or relatively stationary relative to access points 106.

[0023] In some embodiments, access point 106 includes means or modules (comprising a combination of hardware and software) that allow wireless communication device 102 to connect to a wired network using Wi-Fi or other standards. Access point 106 may sometimes be referred to as a wireless access point (WAP). Access point 106 may be configured, designed, and / or constructed to operate in a wireless local area network (WLAN). In some embodiments, access point 106 may be connected as a standalone device to a router (e.g., via a wired network). In other embodiments, access point 106 may be a component of a router. Access point 106 can provide access to multiple devices on the network. Access point 106 may, for example, connect to a wired Ethernet connection and use a radio frequency link to provide wireless connectivity for other devices 102 to utilize the wired connection. Access point 106 may be constructed and / or configured to support standards for transmitting and receiving data using one or more radio frequencies. These standards and the frequencies they use may be defined by IEEE (e.g., the IEEE 802.11 standard). Access point 106 can be configured and / or used to support public Internet hotspots, and / or on an internal network to extend the Wi-Fi signal range of the network.

[0024] In some embodiments, access point 106 may be used for a home or building wireless network (e.g., IEEE 802.11, Bluetooth, ZigBee, any other type of radio frequency-based network protocol and / or variations thereof). Each of the wireless communication devices 102 may include a built-in radio and / or be coupled to a radio. Such wireless communication devices 102 and / or access points 106 may operate according to various aspects of the present disclosure presented herein to enhance performance, reduce cost and / or size, and / or enhance broadband applications. Each wireless communication device 102 may have the functionality of a client node seeking access to resources (e.g., data, and connections to networked nodes such as servers) via one or more access points.

[0025] The network connection may include any type and / or form of network, and may include at least one of the following: point-to-point network, broadcast network, telecommunications network, data communication network, computer network. The network topology may be a bus, star, or ring network topology. The network may have any such network topology known to those skilled in the art capable of supporting the operations described herein. In some embodiments, different types of data may be transmitted via different protocols. In other embodiments, the same type of data may be transmitted via different protocols.

[0026] The communication device 102 and access point 106 can be deployed as any type and form of computing device and / or executed on any type and form of computing device, such as a computer, network device or appliance capable of communicating and performing the operations described herein on any type and form of network. Figure 1B and 1CA block diagram of a computing device 100 for implementing an embodiment of a wireless communication device 102 or access point 106 is depicted. Figure 1B and 1C As shown, each computing device 100 includes a central processing unit 121 and a main memory unit 122. For example... Figure 1B As shown, the computing device 100 may include a storage device 128, a mounting device 116, a network interface 118, an I / O controller 123, display devices 124a to 101n, a keyboard 126, and a pointing device 127 (e.g., a mouse). The storage device 128 may include, but is not limited to, an operating system and / or software. Figure 1C As shown, each computing device 100 may also include additional optional components such as memory port 103, bridge 170, one or more input / output devices 130a to 130n (generally referred to as reference numeral 130), and cache memory 140 in communication with central processing unit 121.

[0027] Central processing unit 121 is any logical circuit system that responds to and processes instructions fetched from main memory unit 122. In many embodiments, central processing unit 121 is provided by a microprocessor unit, such as a microprocessor unit manufactured by Intel Corporation of Mountain View, California; a microprocessor unit manufactured by International Business Machines of White Plains, New York; or a microprocessor unit manufactured by Advanced Micro Devices of Sunnyvale, California. Computing device 100 may be based on any of these processors, or any other processor capable of operating as described herein.

[0028] Main memory cell 122 may be one or more memory chips capable of storing data and allowing microprocessor 121 to directly access any memory location, such as any type or variant of static random access memory (SRAM), dynamic random access memory (DRAM), ferroelectric RAM (FRAM), NAND flash memory, NOR flash memory, and solid-state drive (SSD). Main memory 122 may be based on any of the aforementioned memory chips, or any other available memory chip capable of operating as described herein. Figure 1B In the illustrated embodiment, processor 121 communicates with main memory 122 via system bus 150 (described in more detail below). Figure 1CAn embodiment of a computing device 100 is depicted, wherein the processor communicates directly with the main memory 122 via a memory port 103. For example, in Figure 1C In this context, the main memory 122 can be either DRDRAM or DRAM.

[0029] Figure 1C An embodiment is depicted in which the main processor 121 communicates directly with the cache memory 140 via a secondary bus (sometimes referred to as the back-end bus). In other embodiments, the main processor 121 communicates with the cache memory 140 using a system bus 150. The cache memory 140 has a faster response time than the main memory 122 and is provided by, for example, SRAM, BSRAM, or EDRAM. Figure 1C In the illustrated embodiment, processor 121 communicates with various I / O devices 130 via a local system bus 150. Various buses can be used to connect central processing unit 121 to any of the I / O devices 130, such as VESAVL bus, ISA bus, EISA bus, Micro Channel Architecture (MCA) bus, PCI bus, PCI-X bus, PCI-Express bus, or Nubus. In embodiments where the I / O device is a video display 124, processor 121 may communicate with display 124 using an Advanced Graphics Port (AGP). Figure 1C An embodiment of a computer or computer system 100 is depicted, wherein the main processor 121 can communicate directly with the I / O device 130b, for example, via HYPERTRANSPORT, RAPIDIO, or INFINIBAND communication technologies. Figure 1C An embodiment in which a hybrid local bus and direct communication is also described: the processor 121 communicates with I / O device 130a using a local interconnect bus, while simultaneously communicating directly with I / O device 130b.

[0030] The computing device 100 may contain various I / O devices 130a to 130n. Input devices include keyboards, mice, trackpads, trackballs, microphones, dials, touchpads, touchscreens, and drawing tablets. Output devices include video displays, speakers, inkjet printers, laser printers, projectors, and dye-to-sublimation printers. I / O devices can be, for example... Figure 1BThe illustrated I / O controller 123 controls one or more I / O devices, such as a keyboard 126 and a pointing device 127 (e.g., a mouse or optical pen). Additionally, the I / O devices may provide storage and / or mounting media 116 for the computing device 100. In yet another embodiment, the computing device 100 may provide a USB connection (not shown) to receive a handheld USB storage device, such as a USB flash memory driver line manufactured by Twintech Industry, Inc., Los Alamitos, California.

[0031] Refer again Figure 1B The computing device 100 may support any suitable mounting device 116, such as a disk drive, CD-ROM drive, CD-R / RW drive, DVD-ROM drive, flash memory drive, tape drive of various formats, USB device, hard disk drive, network interface, or any other device suitable for installing software and programs. The computing device 100 may further include a storage device, such as one or more hard disk drives or a redundant array of independent disks, for storing the operating system and other related software, as well as application software programs, such as any program or software 120 used to implement the systems and methods described herein (e.g., software 120 configured and / or designed for it). Optionally, any mounting device 116 may also be used as a storage device. Additionally, the operating system and software may run from a bootable medium.

[0032] Furthermore, the computing device 100 may include a network interface 118 to interface with the network 104 via various connections, including but not limited to standard telephone lines, LAN or WAN links (e.g., 802.11, T1, T3, 56kb, X.25, SNA, DECNET), broadband connections (e.g., ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethernet over SONET), wireless connections, or some combination of any or all of the above. Various communication protocols (e.g., TCP / IP, IPX, SPX, NetBIOS, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), RS232, IEEE 802.11, IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, IEEE 802.11ac, IEEE 802.11ad, CDMA, GSM, WiMax, and Direct Asynchronous Connection) can be used to establish the connection. In one embodiment, computing device 100 communicates with other computing devices 100' via any type and / or form of gateway or tunneling protocol, such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS). Network interface 118 may include a built-in network adapter, network interface card, PCMCIA network card, card bus network adapter, wireless network adapter, USB network adapter, modem, or any other means suitable for interfaced to computing device 100 to any type of network capable of communication and performing the operations described herein.

[0033] In some embodiments, computing device 100 may include or be connected to one or more display devices 124a to 124n. Therefore, any of the I / O devices 130a to 130n and / or I / O controller 123 may include any type and / or form of suitable hardware, software, or a combination of hardware and software to support, implement, or provide computing device 100 with access to and use of display devices 124a to 124n. For example, computing device 100 may include any type and / or form of video adapter, video card, driver, and / or library to interface with, communicate with, connect to, or otherwise use display devices 124a to 124n. In one embodiment, a video adapter may include multiple connectors to interface with display devices 124a to 124n. In other embodiments, computing device 100 may include multiple video adapters, each connected to display devices 124a to 124n. In some embodiments, any portion of the operating system of computing device 100 may be configured to use multiple displays 124a to 124n. Those skilled in the art will recognize and understand that the computing device 100 can be configured to have one or more display devices 124a to 124n in various ways and embodiments.

[0034] In a further embodiment, I / O device 130 may serve as a bridge between system bus 150 and external communication buses (such as USB bus, Apple Desktop bus, RS-232 serial connection, SCSI bus, FireWire bus, FireWire 800 bus, Ethernet bus, AppleTalk bus, Gigabit Ethernet bus, Asynchronous Transfer Mode bus, FibreChannel bus, Serial Attached Small Computer System Interface bus, USB connection, or HDMI bus).

[0035] Figure 1B and 1C The computing device or system 100 of the type described herein can operate under the control of an operating system that controls task scheduling and access to system resources. The computing device 100 can run any operating system, such as any version of Microsoft Windows, different versions of Unix and Linux, any version of MAC OS for Macintosh computers, any embedded operating system, any real-time operating system, any open-source operating system, any proprietary operating system, any operating system for mobile computing devices, or any other operating system capable of running on a computing device and performing the operations described herein. Typical operating systems include (but are not limited to): Android, manufactured by Google Inc.; Windows 7 and 8, manufactured by Microsoft Corporation of Redmond, Washington; macOS, manufactured by Apple Computer of Cupertino, California; WebOS, manufactured by Research In Motion (RIM); OS / 2, manufactured by International Business Machines of Armonk, New York; and Linux, a freely available operating system distributed by Caldera Corp. of Salt Lake City, Utah, or any type and / or form of Unix operating system, etc.

[0036] Computer system 100 may be any workstation, telephone, desktop computer, laptop or notebook computer, server, handheld computer, mobile phone or other portable telecommunications device, media playback device, gaming system, mobile computing device, or any other type and / or form of communicative computing, telecommunications, or media device. Computer system 100 has sufficient processor power and memory capacity to perform the operations described herein.

[0037] In some embodiments, computing device 100 may have a different processor, operating system, and input device consistent with the device. For example, in one embodiment, computing device 100 is a smartphone, mobile device, tablet, or personal digital assistant. In yet other embodiments, computing device 100 is an Android-based mobile device, an iPhone smartphone manufactured by Apple Computer, Inc. of Cupertino, California, or a Blackberry or WebOS-based handheld device or smartphone, such as a device manufactured by Research In Motion Limited. Furthermore, computing device 100 may be any workstation, desktop computer, laptop or notebook computer, server, handheld computer, mobile phone, any other computer, or other form of computing or telecommunications device capable of communication and having sufficient processor power and memory capacity to perform the operations described herein. These aspects of the operating environment and components will become apparent in the context of the systems and methods disclosed herein.

[0038] B. Systems and methods for storing contexts in a SAN

[0039] refer to Figure 2 Network 200 includes an initiating host 202, an initiating host bus adapter (HBA) 204, a switching network 206, a target HBA 208, and a target host 210. In some embodiments, network 200 is configured as an FC SAN. In some embodiments, network 200 may utilize link-level flow control and host CPU / software / drivers to limit the number of pending commands or control queue depth. Furthermore, in some embodiments, host CPU / software / drivers are configured to coordinate traffic between connections without impacting performance due to CPU overhead limitations.

[0040] Network 200 can be related to the above text about Figures 1A to 1CThe described computing and communication components are used together. For example, computer system 100 or communication device 102 may be a target host 210 or an initiating host 202. Network 200 is a high-speed network connecting initiating host 202 to target host 210 (e.g., a high-performance storage subsystem). In some embodiments, network 200 may access storage processors (SPs) and storage disk arrays (e.g., target host 210). In some embodiments, network 200 may use the FC protocol to encapsulate SCSI commands into FC frames. When data is transferred between initiating host 202 and target host 210, network 200 may utilize multipathing to provide more than one physical path from initiating host 202 to target host 210 via switching network 206.

[0041] The initiating host 202 includes host memory 220. Host memory 220 is volatile or non-volatile memory. Host memory 220 includes one or more memory chips capable of storing data, such as any type or variation of static random access memory (SRAM), dynamic random access memory (DRAM), ferroelectric RAM (FRAM), NAND flash memory, NOR flash memory, and solid-state drive (SSD), or combinations thereof. Host memory 220 may be based on any of the aforementioned memory chips, or any other available memory chip capable of operating as described herein. Host memory 220 may include or operate using a disk drive, CD-ROM drive, CD-R / RW drive, DVD-ROM drive, flash memory drive, tape drive of various formats, USB device, hard disk drive, network interface, or any other device suitable for data storage. In some embodiments, host memory 220 includes extended context memory 222 for storing Scalable Resource Identifiers or Interchangeable Resource Indicators (XRIs). XRIs provide a unified syntax for abstract structured identifiers. In some embodiments, extended context memory 222 is cacheable. The target host 210 may contain similar components.

[0042] Target host 210 includes target memory 224. Target memory 224 is volatile or non-volatile memory. Target memory 224 includes one or more memory chips capable of storing data, such as any type or variant of static random access memory (SRAM), dynamic random access memory (DRAM), ferroelectric RAM (FRAM), NAND flash memory, NOR flash memory, and solid-state drive (SSD), or combinations thereof. Target memory 224 may be based on any of the aforementioned memory chips, or any other available memory chip capable of operating as described herein. Target memory 224 may include or operate using a disk drive, CD-ROM drive, CD-R / RW drive, DVD-ROM drive, flash memory drive, tape drive of various formats, USB device, hard disk drive, network interface, or any other device suitable for data storage. In some embodiments, target memory 224 includes extended context memory 226 for storing Scalable Resource Identifiers or Interchangeable Resource Indicators (XRIs). XRIs provide a unified syntax for abstract structured identifiers. In some embodiments, extended context memory 226 is cacheable.

[0043] In some embodiments, the initiating HBA 204 and the target HBA 208 are line cards, mezzanine cards, motherboard devices, or other devices and are configured according to SCSI, FC, and Serial Advanced Technology (AT) Attachment (SATA) protocols. The initiating HBA 204 includes transmitter ports 214a to b and receiver ports 216a to b. The number of ports 214a to b and 216a to b can be any number from 1 to N. The target HBA 208 includes transmitter ports 254a to b and receiver ports 256a to b. The number of ports 254a to b and 256a to b can be any number from 1 to N. In some embodiments, the initiating HBA 204 and the target HBA 208 are FC HBAs implemented as ASICs of a fourth-generation multi-function PCI-Express device with 4 x 64 gigabyte Fibre Channel (GFC) full-line-rate connectivity. ASICs can provide host-supported storage (HBS) with service level interface (SLI) versions and protocol offloading for Fibre Channel (FCP) command mode and Fibre Channel protocols (FCPs) such as SCSI, NVMe, and FICON.

[0044] HBA 204 includes an XRI allocation / deallocation server 280, a host interface / direct memory access (DMA) engine unit 282, a context memory 284, a host supported storage (HBS) unit 286, and a protocol offload engine 288. The target HBA 208 may contain similar components. HBA 208 includes an XRI allocation / deallocation server 281, a host interface / direct memory access (DMA) engine unit 283, a context memory 285, a host supported storage (HBS) unit 287, and a protocol offload engine 289.

[0045] Context memory 284 includes a memory cache or storage device 292 for cacheable XRIs and is configured to extend the context memory cache. Context memory 284 also includes on-chip memory 294 for non-cacheable XRIs. In some embodiments, context memory 284 may include one or more memory chips capable of storing data and allowing any storage location, such as any type or variant of static random access memory (SRAM), dynamic random access memory (DRAM), ferroelectric RAM (FRAM), NAND flash memory, NOR flash memory, and solid-state drive (SSD) or combinations thereof. Context memory 285 includes a memory cache or storage device 293 for cacheable XRIs and is configured to extend the context memory cache. Context memory 285 also includes on-chip storage device 295 for non-cacheable XRIs. In some embodiments, context memory 285 may include one or more memory chips capable of storing data and allowing any storage location, such as any type or variant of static random access memory (SRAM), dynamic random access memory (DRAM), ferroelectric RAM (FRAM), NAND flash memory, NOR flash memory, and solid-state drive (SSD) or combinations thereof.

[0046] In some embodiments, the XRI allocation and deallocation (XAD) server 280 on HBA 204 is an on-chip unit. In some embodiments, XAD server 280 allocates cacheable and non-cacheable XRIs based on ranges (e.g., address ranges). In some embodiments, XAD server 280 first optimizes the performance of the Tier 1 pool using storage device 294, and then optimizes the performance of the Tier 2 pool using storage device 292 or extended context memory 222. XAD server 280 can be implemented using software executing on the processor or can be based on logic. The XRI allocation and deallocation (XAD) server 281 of HBA 208 is similar to server 280.

[0047] Host interface / direct memory access (DMA) engine unit 282 provides an interface for direct memory access to host memory 220. In some embodiments, host interface / direct memory access (DMA) engine unit 282 may communicate context data and payload and may include a queue prefetch engine, a PCI interface engine, a payload DMA engine, and a control DMA engine. Host interface / direct memory access (DMA) engine unit 282 may be implemented using software that executes on the processor or may be logic-based. Host interface / direct memory access (DMA) engine unit 283 of HBA 208 is similar to unit 282.

[0048] HBS unit 286 is a cache memory controller extension. HBS unit 286 examines the VCA to determine the access range of memory devices 292 and 294 based on the VCA range (e.g., memory device 294 is used for values ​​below 4000, and memory device 292 is used for values ​​equal to or higher than 4000). Host Support Memory (HBS) unit 286 can be implemented using software executing on the processor or can be based on logic. HBS unit 287 of HBA 208 is similar to unit 286.

[0049] Protocol offload engine 288 offloads the swap. Protocol offload engine 288 can extract the context and update and maintain the context to complete the I / O swap. In some embodiments, protocol offload engine 288 seamlessly handles the swap phase, which includes a completion message phase. Protocol offload engine 288 can be implemented using software executing on the processor or can be based on logic. Protocol offload engine 289 of HBA 208 is similar to protocol offload engine 288.

[0050] Switching network 206 includes receive ports 236a to b, transmit ports 234a to b, transmit ports 244a to b, receive ports 246a to b, and a buffer or crossbar switch 238. The number of ports 234a to b, 244a to b, 236a to b, and 246a to b can be any number from 1 to N. Ports 236a to b communicate with ports 214a to b, and ports 234a to b communicate with ports 216a to b. Ports 236a to b communicate with ports 244a to b via crossbar switch 238, and ports 234a to b communicate with ports 246a to b via crossbar switch 238. Ports 244a to b communicate with ports 256a to b, and ports 246a to b communicate with ports 254a to b. In some embodiments, switching network 206 is a structured switching network, an arbitrated ring network, or a point-to-point network. Ports 236a to b, 234a to b, 244a to b, and 246a to b can be associated with servers, hubs, switches, routers, directives, nodes, or other devices. Switching network 206 may include physical layer, interconnection devices, and translation devices.

[0051] The buffer or crossbar switch 238 may be two unidirectional or bidirectional switches configured to interconnect ports 236a to b with ports 244a to b, and ports 234a to b with ports 246a to b. The crossbar switch 238 may include a switch matrix and buffers and other communication and interface circuitry. Each of ports 214a to b, 216a to b, 234a to b, 236a to b, 244a to b, 246a to b, 254a to b, and 256a to b may have a unique addressable identifier. Each of ports 214a to b, 216a to b, 234a to b, 236a to b, 244a to b, 246a to b, 254a to b, and 256a to b can be associated with a network node, and a pair of ports 214a to b, 216a to b, 234a to b, 236a to b, 244a to b, 246a to b, 254a to b, and 256a to b can be associated with a network link.

[0052] In some embodiments, HBA 204 and initiating host 202 are configured to extend on-chip memory usage into host memory 220 without increasing the size of on-chip memory (e.g., context memory 284). HBA 204 and initiating host 202 are configured to provide a multi-tiered memory structure for multi-tiered swap offloading. A first-tier pool provides several non-cacheable swap contexts that are always stored on-chip in memory 294, which are optimized for performance with low-latency operation in some embodiments. A second-tier pool provides several cacheable swap contexts that may be cached on-chip in memory 292 or in host memory 220 for scalability. Additionally, several contexts may be pinned on-chip to eliminate or reduce cache misses during I / O operations. Tier selection can be done manually or automatically.

[0053] In some embodiments, the host driver or internal adapter agent of the HBA 204 can explicitly request XRI resources in a first-tier pool or a second-tier pool based on service type classification, I / O size, etc. As an optimization according to some embodiments, selection to provide the first-tier pool to link to the second-tier pool can be performed automatically, allowing allocation of second-tier pool resources when first-tier pool resources are exhausted. Therefore, in some embodiments, if the first-tier pool is empty, the requesting agent does not need to make two XRI allocation requests. In some embodiments, a self-identification mechanism with XRI ranges is implemented to automatically bootstrap the XRI version to the appropriate first-tier or second-tier pool.

[0054] In some embodiments, the system and method offload I / O swapping with a multi-level swapping context in memory 284 and host memory 220 to optimize performance and scalability. In some embodiments, on-chip memory 284 is effectively scalable into host memory 220 without increasing the size of on-chip memory 284. Host memory 220 and context memory 284 provide a multi-level memory system that offers performance and scalability, and in some embodiments, the multi-level system can be selected manually or automatically.

[0055] Several Level 2 cacheable contexts can be pinned on-chip to eliminate or reduce cache misses during I / O operations using the HBA 204. Several programmable cache line pinning methods are provided to optimize performance using programmable limits for each set of indices, types, and / or ports and can be executed by the HBS unit 286. Once a cache line is pinned, it cannot be replaced until it is depinned. At the end of each I / O operation, the associated I / O interchange cache line is depinned to free up cache line resources. According to some embodiments, pinning methods that can be enabled or disabled independently include:

[0056] 1. I / O Phase Fixing - Fixing can be requested for specific I / O phases, such as non-sequence end frames.

[0057] 2. First-In Fixing - If the fixing limit is not reached, fixing can be requested on a first-come, first-served basis. For set-associative cache schemes, each set index has a programmable fixing count (e.g., programmed N-way / 2 as the PIN count set index limit), which is the number of cache lines on-chip that can be fixed for each set index.

[0058] 3. Region Fixing - Fixing can be requested if the requested address is within a specific programmable Virtual Context Address (VCA) range. A programmable fixing count is provided to limit the number of cache lines that can be fixed for each region.

[0059] 4. Fixed I / O Size Range - If the requested I / O size is within a specific programmable fixed I / O size range, then a fixed operation can be requested.

[0060] 5. Fixed XRI Range - If the requested XRI number is within a specific programmable XRI range, then a fixed operation can be requested.

[0061] 6. Protocol Fixing - Fixing can be requested based on the FC-44 protocol (such as NVMe or SCSI).

[0062] 7. The work queue (WQ) is always specified by the WQ configuration file.

[0063] 8. Work queue entries (WQE) contain fixed bits indicated by the host driver.

[0064] Pinning is inherently opportunistic. Once the pinning limit is reached, subsequent accesses will be cached instead of pinning. I / O pinning can include a command phase, a transfer preparation phase, a data phase (first sequence, etc.), and a sequence phase.

[0065] XAD server 280 provides hardware offloading for shared pools of XRI numbers to improve host drive performance by reducing or eliminating CPU contention. In some embodiments, XAD server 280 allows any CPU to allocate or deallocate XRI numbers without creating semaphores or spinlocks in some embodiments. Creating spinlocks or coordinating between CPUs for XRI allocation / sharing can severely impact performance. Each XRI pool may be preloaded with a list of valid XRI numbers during initialization. In some embodiments, a simple register access (e.g., via a doorbell) at a single offset in the Peripheral Component Interconnect High Speed ​​(PCIe) Function Base Address Register (BAR) initiates the XRI allocation or deallocation process performed by XAD server 280. In some embodiments, an internal adapter agent may access the XRI pool as needed to allocate and deallocate XRI resources.

[0066] The XRI allocation order can be independently configured for each tier to either LIFO (Last In, First Out) or FIFO (First In, First Out) mode. LIFO XRI configuration enhances tier 2 cache performance by allocating or reusing the context of the most recently exited XRI, thereby improving cache line hit rate. Instance configurations use FIFO mode for the tier 1 pool to facilitate debugging and LIFO mode for the tier 2 pool to improve cache performance.

[0067] HBS unit 286 provides access to an extended data storage device for swap-related contexts supported by host memory 220. In some embodiments, HBS unit 286 supports page address translation (PAT) to translate Virtual Context Memory Addresses (VCAs) into PCIe host physical addresses to facilitate address space expansion. In some embodiments, HBS unit 286 provides an on-chip cache for swap-related contexts. HBS unit 286 maintains several cache lines in on-chip memory 284 to provide fast access by HBS clients. The actual cache line data is stored in context memory 284.

[0068] HBS unit 286 maintains several cache lines in context memory 284 to provide fast access by on-chip HBS clients. The actual cache line data is stored in on-chip context memory 284. One or more cache controllers may be provided in HBS unit 286, each managing a different set of contexts or data structures. Several context or data structure types with programmable strides can be encapsulated into a single cache line to increase the cache hit rate of accesses performed sequentially by various modules.

[0069] In the event of a cache miss, HBS unit 286 evicts one of the cache lines stored in memory 284 and writes it back to host memory 220 (if no free entry is available), and fetches a new cache line to fill the on-chip cache (e.g., storage device 292). Both Least Recently Used (LRU) and pseudo-LRU cache replacement policies and a random cache eviction policy are supported. The cache subsystem uses cyclic redundancy check (CRC) to provide optional protection for the cache lines stored in host memory 220. In some embodiments, each context entry is protected by an embedded CRC, and the cache of HBS unit 286 is decoupled from the L1 cache of the client maintained in the client module for high-performance, ultra-low-latency access to context or data structures.

[0070] In some embodiments, the swap data stored in host memory 220 belongs to a PCIe function, and data for a specific PCIe function must not reside in memory used for another PCIe function. Page address translation (PAT) entries must be programmed with the appropriate host memory page address and PCIe function ID.

[0071] In some embodiments, HBA 208 and target host 210 are configured to extend on-chip memory usage into target memory 224 without increasing the size of on-chip memory (e.g., context memory 285). HBA 208 and target host 210 are configured to provide a multi-level memory architecture for multi-level swap offloading. A first-level pool provides several non-cacheable swap contexts that are always stored on-chip in storage device 295, which in some embodiments is optimized for performance with low-latency operation. A second-level pool provides several cacheable swap contexts that may be cached on-chip in storage device 293 or in target memory 224 for scalability. Additionally, several contexts may be pinned on-chip to eliminate or reduce cache misses during I / O operations. Level selection can be done manually or automatically.

[0072] In some embodiments, the host driver or internal adapter agent of the HBA 208 may explicitly request XRI resources in a first-tier pool or a second-tier pool based on service type classification, I / O size, etc. As an optimization according to some embodiments, selection to provide the first-tier pool to link to the second-tier pool can be performed automatically, allowing allocation of second-tier pool resources when first-tier pool resources are exhausted. Therefore, in some embodiments, if the first-tier pool is empty, the requesting agent does not need to make two XRI allocation requests. In some embodiments, a self-identification mechanism with XRI ranges is implemented to automatically bootstrap the XRI version to the appropriate first-tier or second-tier pool.

[0073] In some embodiments, the system and method offload I / O swapping with a multi-level swapping context in memory 285 and target memory 224 to optimize performance and scalability. In some embodiments, on-chip memory 285 can be effectively scaled into target memory 224 without increasing the size of on-chip memory 285. Target memory 224 and context memory 285 provide a multi-level memory system that offers performance and scalability, and in some embodiments, the multi-level system can be selected manually or automatically.

[0074] Several Level 2 cacheable contexts can be pinned on-chip to eliminate or reduce cache misses during I / O operations using the HBA 208. Several programmable cache line pinning methods are provided to optimize performance using programmable limits for each set of indices, types, and / or ports and can be executed by the HBS unit 287. Once a cache line is pinned, it cannot be replaced until it is depinned. At the end of each I / O operation, the associated I / O interchange cache line is depinned to free up cache line resources. According to some embodiments, pinning methods that can be enabled or disabled independently include:

[0075] 1. I / O Phase Fixing - Fixing can be requested for specific I / O phases, such as non-sequence end frames.

[0076] 2. First-In Fixing - If the fixing limit is not reached, fixing can be requested on a first-come, first-served basis. For set-associative cache schemes, each set index has a programmable fixing count (e.g., programmed N-way / 2 as the PIN count set index limit), which is the number of cache lines on-chip that can be fixed for each set index.

[0077] 3. Region Fixing - Fixing can be requested if the requested address is within a specific programmable Virtual Context Address (VCA) range. A programmable fixing count is provided to limit the number of cache lines that can be fixed for each region.

[0078] 4. Fixed I / O Size Range - If the requested I / O size is within a specific programmable fixed I / O size range, then a fixed operation can be requested.

[0079] 5. Fixed XRI Range - If the requested XRI number is within a specific programmable XRI range, then a fixed operation can be requested.

[0080] 6. Protocol Fixing - Fixing can be requested based on the FC-44 protocol (such as NVMe or SCSI).

[0081] 7. The work queue (WQ) is always specified by the WQ configuration file.

[0082] 8. Work queue entries (WQE) contain fixed bits indicated by the host driver.

[0083] Pinning is inherently opportunistic. Once the pinning limit is reached, subsequent accesses will be cached instead of pinning. I / O pinning can include a command phase, a transfer preparation phase, a data phase (first sequence, etc.), and a sequence phase.

[0084] XAD server 281 provides hardware offloading for shared pools of XRI numbers to improve host driver performance by reducing or eliminating CPU contention. In some embodiments, XAD server 281 allows any CPU to allocate or deallocate XRI numbers without creating semaphores or spinlocks in some embodiments. Creating spinlocks or coordinating between CPUs for XRI allocation / sharing can severely impact performance. Each XRI pool may be preloaded with a list of valid XRI numbers during initialization. In some embodiments, a simple register access (e.g., via a doorbell) at a single offset in the Peripheral Component Interconnect High Speed ​​(PCIe) Function Base Address Register (BAR) initiates the XRI allocation or deallocation process performed by XAD server 281. In some embodiments, an internal adapter agent may access the XRI pool as needed to allocate and deallocate XRI resources.

[0085] The XRI allocation order can be independently configured for each tier to either LIFO (Last In, First Out) or FIFO (First In, First Out) mode. LIFO XRI configuration enhances tier 2 cache performance by allocating or reusing the context of the most recently exited XRI, thereby improving cache line hit rate. Instance configurations use FIFO mode for the tier 1 pool to facilitate debugging and LIFO mode for the tier 2 pool to improve cache performance.

[0086] HBS unit 287 provides access to an extended data storage device for swap-related contexts supported by target memory 224. In some embodiments, HBS unit 287 supports page address translation (PAT) to translate Virtual Context Memory Addresses (VCAs) into PCIe host physical addresses to facilitate address space expansion. In some embodiments, HBS unit 287 provides an on-chip cache for swap-related contexts. HBS unit 287 maintains several cache lines in on-chip memory to provide fast access by HBS clients. The actual cache line data is stored in context memory 285.

[0087] HBS unit 287 maintains several cache lines in context memory 285 to provide fast access by on-chip HBS clients. The actual cache line data is stored in on-chip context memory 285. One or more cache controllers may be provided in HBS unit 287, each managing a different set of contexts or data structures. Several context or data structure types with programmable strides can be encapsulated into a single cache line to increase the cache hit rate of accesses performed sequentially by various modules.

[0088] In the event of a cache miss, HBS unit 287 evicts one of the cache lines stored in memory 285 and writes it back to target memory 224 (if no free entry is available), and fetches a new cache line to fill the on-chip cache (e.g., memory device 293). Both Least Recently Used (LRU) and pseudo-LRU cache replacement policies and a random cache eviction policy are supported. The cache subsystem uses cyclic redundancy check (CRC) to provide optional protection for the cache lines stored in target memory 224. In some embodiments, each context entry is protected by an embedded CRC, and the cache of HBS unit 287 is decoupled from the L1 cache of the client maintained in the client module for high-performance, ultra-low-latency access to context or data structures.

[0089] In some embodiments, the swap data stored in target memory 224 belongs to a PCIe function, and data for a specific PCIe function must not reside in memory used for another PCIe function. In some embodiments, page address translation (PAT) entries must be programmed with the appropriate host memory page address and PCIe function ID.

[0090] The entry client provides HBS hints to preload cache entries before the command / frame arrives at the protocol engine for unloading. To mitigate cache miss penalties, HBS unit 286 or 287 provides retry operations to allow the HBS client to deactivate cache miss requests and move to another command / context without blocking. With retries enabled, upon a read cache miss, HBS unit 286 or 287 returns a miss response to the requesting HBS client and initiates a context fetch DMA via host interface / DMA engine unit 282 or 283 to obtain host-supported data from the on-chip cache before the next retry.

[0091] In some embodiments, a programmable non-cacheable region is provided to provide low-latency memory access to a non-cacheable context, bypassing HBS cache lookups. Furthermore, in some embodiments, the HBS client provides separate programmable, cacheable, and non-cacheable context base addresses to facilitate simultaneous non-blocking multi-level memory operations between cacheable and non-cacheable requests.

[0092] While exemplary embodiments are described and illustrated herein in reference to FC HBAs or storage adapters, it should be understood that the embodiments disclosed herein are not limited thereto, but can be applied in the context of many other adapters with congestion management, such as Ethernet aggregation network adapters (CNAs) or security adapters. According to some embodiments, even though two tier pools are described, three or more tiers may be incorporated to provide fine-grained differentiation for various service types as needed.

[0093] refer to Figure 3 System 300 can support a large number of I / O context structures using a multi-level memory storage method. System 300 can be used in a network HBA with host-supported storage (HBS), where the on-chip cache is similar to that of network 200 ( Figure 2 System 300 can be used in a mix of FC HBA, Ethernet HBA, or heterogeneous clients (FC and Ethernet or others) within the same adapter. In some embodiments, system 300 may be a large HBA Level 2 cache for on-chip Level 1 cache.

[0094] In some embodiments, system 300 provides dynamic structure allocation with seamless memory tier selection, featuring service optimizations to avoid blocking between tiers. In some embodiments, system 300 optimizes cacheable services to allow out-of-order client service whenever possible. In some embodiments, system 300 maximizes parallel client operations on data, provides a parallel client consistency scheme, and is easily scalable in terms of resources. In some embodiments, system 300 seamlessly allocates structures in cached VCA space and performs this with low power consumption (e.g., a low-power envelope of less than 10 watts for HBAs).

[0095] System 300 includes memory 304, crossbar switch 302, cache controller 308, and page address translation (PAT) subunit 306. PAT subunit 306 includes a page translation cache (PTC) cache 332 for PAT table data. Memory 304 includes a first-level storage device 340, a storage device 342 for second-level storage (HBS cacheable data 346 or remapped data), a third-level storage device 344, an HBS cacheable data storage device 346 for second-level and / or third-level cached data, and a PAT table storage device 348. In some embodiments, crossbar switch 302, cache controller 308, and PAT subunit 306 are host (PCIe root complex) supported memory storage units (e.g., similar to HBS unit 286) configured to manage context memory 304 (e.g., as a cache / memory controller). Figure 2 The host-supported memory is accessible via link 352 and can be replaced by any other memory (e.g., memory attached to a PCIeUpstream switch or located on a different PCIe bus).

[0096] Level 1 storage device 340 stores non-cacheable data, and Level 2 storage device 342 stores HBS cacheable data or remapped data. Level 3 storage device 344 stores other data types. In some embodiments, other data types may include Level 3 data, XRI data, or context structure data. HBS cacheable data storage device 346 stores data for Level 2 and / or Level 3 cached data, and PAT table storage device 348 stores PAT table data. In some embodiments, PAT table 348 is a host PAT table.

[0097] Memory 304 is an on-chip memory having multiple ports 312a to n. In some embodiments, memory 304 is similar to context memory 284. Figure 2The number of ports ranges from 2 to N, where N is an integer. In some embodiments, each of ports 312a to n may correspond to a corresponding client among n clients. Memory 304 includes ports coupled to PAT subunit 306 via link 314. Memory 304 may be any of the storage types discussed above.

[0098] Memory 304 is shared among non-cached storage device 340, cached data storage devices 342 and 344, HBS cached data storage device 346, and PAT table 348. Memory 304 exposes the same data to each of the multiple ports 312a to n. In some embodiments, client-to-client request ordering is not enforced by caching, but it is part of the client-to-client protocol. Storage device 342 stores cacheable remapped data not supported by host memory 304, and for said data, the Level 2 VCA address is remapped into context memory 304. The remapped state is part of PTA table 348. In some embodiments, the remapped feature allows for faster access times, similar to non-cached data but within cacheable space and therefore transparent to the client.

[0099] In some embodiments, crossbar switch 302 is a semi-coherent crossbar switch configuration. Crossbar switch 302 is used to connect client ports 310a to n to one of ports 312a to n. Crossbar switch 302 does not need to support crossbar ports (e.g., N to N+M) and provides direct N to N connectivity. In some embodiments, crossbar switch 302 on the client side at the port includes logic for decoding VCA addresses and separating the addresses into independent non-blocking data paths. Crossbar switch 302 includes logic for supporting cache snooping of multiple cache lines every clock cycle. In some embodiments, snooping requests are provided from cache controller 308 via link 322. The crossbar switch snooping interface associated with link 322 allows cache line ownership checks between each client and cache controller 308. The crossbar switch input client logic of crossbar switch 302 also supports out-of-order fill and refresh signals. The output side of crossbar switch 302 is connected to each of ports 312a to n and supports three data paths including non-cacheable, cacheable hit, or cache-missed lines. In some embodiments, the crossbar switch 302 provides three independent (nx3) data paths and offers an optimal solution for non-blocking, unordered, low-latency, and high-bandwidth operation.

[0100] The crossbar switch 302's crossbar switch client queue logic supports deactivation of cache lines for cacheable data. The client queue stores VCA address information after the client completes its task. The client can hit a valid deactivated cache line and directly access cached data (if present), further reducing cache overhead to zero. In some embodiments, the cache controller 308 provides sideband control (e.g., eviction, flushing, invalidation, etc.) to the client to leverage cache architecture-specific temporal and spatial locality.

[0101] The cache controller 308 uses a serial pipeline, where new requests can enter the serial pipeline every clock cycle. If the last stage becomes full, the cache controller 308 can stop the pipeline's four stages (arbitration, lookup read, lookup decode, and controller). In some embodiments, the cache controller 308 processes cacheable requests every clock cycle.

[0102] Given low latency, a large number of cache lines, and very low memory power, a group (way) associative write-back cache structure with Least Recently Used (LRU) or random path selection is used. In some embodiments, the number of paths supported by each cache index is 16 or more. For each cache index, an LRU state is available. In some embodiments, per-cache-index LRU allows for very low placement granularity and improved cache hits.

[0103] In some embodiments, the cache controller 308 supports three cache line states: invalid, valid (clean), and modified. The cache controller 308 supports different cache data sizes (4MB to 512KB), 16 or 8 path selections, and several VCA-to-cache index mapping functions. In some embodiments, these mapping functions can be used to reduce cache corruption through swapping. In some embodiments, cache memory power consumption can be reduced by disabling unused cache memory group memory.

[0104] The cache controller 308 supports several features for handling cache hits and misses using internal and external retries and independent data paths to achieve non-blocking. The cache controller 308 supports locking cache lines via directional (by the client) or autonomous locking or unlocking cacheable structures belonging to the same cache line. In some embodiments, this allows clients to access critical data faster and more predictably. The cache controller 308 also supports several hint ports for incoming VCA requests without on-chip memory access. Hint ports are used by clients aware of lead time for incoming VCA requests. Hints allow the cache controller 308 to fetch cacheable VCA data in advance. In some embodiments, the cache controller 308 uses an on-chip DMA engine (e.g., unit 282). Figure 2 This involves moving large amounts of HBS data (e.g., 256 bytes) across the PCIe link. The cache controller 308 provides one to N management clients. These management clients are controlled by on-chip firmware. The applications for the multiple management clients can include support for or management of each functional group to accelerate processing.

[0105] The cache controller 308 supports I / O virtualization. Each I / O external port service can be assigned one or more PCI functions. Virtual PCI functions can be associated with physical functions. In some embodiments, physical function services and host memory locations are controlled by and dependent on the host I / O MMU. PCI functions can be enabled or disabled at any time, and host memory data locations can change over time based on the PCI function state. Flexible and programmable host address mapping to the cache VCA space is necessary for pages with a minimum size of 4KB. Translated access must also be very low-latency and provide concurrent access to the PAT table 348. Because a translation can contain multiple structures within the same 4KB page size, address translation will have temporal and spatial locality of the translated cache. Due to the exposure of off-chip memory to potential corruption (e.g., rogue functions, hackers, etc.), early HBS structure corruption detection is provided in some embodiments.

[0106] The cacheable and unremapped Tier 2 data is backed by host memory and partitioned for each PCI function. In some embodiments, PCI function X will not share the VCA space with function Y. Upon entry creation, the initial function allocation is stored in each PAT entry in PAT table 348. The function ID is stored in each cache path and used as needed via cache controller 308 (host DMA, PCIe, etc.). In some embodiments, cache controller 308 supports cacheable VCA lookups, but also supports function ID lookups for cache management. In some embodiments, the cache does not distinguish between physical and virtual functions.

[0107] In some embodiments, the PAT page structure reduces the number of table entries and keeps PAT table 348 on-chip. Due to the page structure of the VCA's PAT mapping, the same PAT entries are used for multiple sets of 256-byte structures (up to 16 sets with a 4KB page size). This allows PAT entries to be cached and expects to occupy some space or time locally. PAT subunit 306 contains cache 332 similar to cache 308. In some embodiments, a similar 16-way set association is used without write-back features.

[0108] Related to PCI function virtualization, there are issues concerning HBS memory data protection. In some embodiments, an early protection mechanism is added to the structure of cache line data stored in the host. In some embodiments, for each cache line, the HBA stores and checks a two-byte CRC. In some embodiments, the HBA calculates the cache line CRC when it DMAs data into host memory. In some embodiments, the HBA checks the cache line CRC when it DMAs a cache line from host memory.

[0109] Client structures can be allocated to different memory tiers based on their criticality, rather than all being allocated to the same tier. In some embodiments, the client controls the persistence of structures in the cache and can hold, pin, flush, invalidate, evict, or reuse them, rather than holding them from beginning to end. In some embodiments, system 300 allows up to 262,144 composite structures (256K x CCB). In some embodiments, system 300 reduces on-chip RAM memory and / or reallocates on-chip RAM memory to other resources.

[0110] It should be noted that certain paragraphs of this disclosure may use terms related to devices, bit depth, transmission duration, etc. (e.g., "first" and "second") to identify or distinguish one another or others. These terms are not intended to relate entities (e.g., first device and second device) merely in time or according to sequence, although in some cases such a relationship may be included. These terms also do not limit the number of possible entities (e.g., devices) that can operate in the system or environment.

[0111] It should be understood that the system described above may provide any or more of those components, and these components may be provided on a standalone machine or, in some embodiments, on multiple machines in a distributed system. Furthermore, the system and methods described above may be provided as one or more computer-readable programs or executable instructions embodied in or on one or more articles of art. The articles of art may be floppy disks, hard disks, CD-ROMs, flash memory cards, PROMs, RAMs, ROMs, or magnetic tapes. Generally, the computer-readable program may be implemented in any programming language (e.g., LISP, PERL, C, C++, C#, PROLOG) or in any bytecode language (e.g., JAVA). The software program or executable instructions may be stored as object code on or in one or more articles of art.

[0112] While the foregoing written description of the methods and systems enables those skilled in the art to make and use what is currently considered the best mode of implementation, it should be understood and appreciated that variations, combinations, and equivalents of specific embodiments, methods, and examples exist herein. Therefore, these methods and systems should not be limited to the embodiments, methods, and examples described above, but rather to all embodiments and methods within the scope and spirit of this invention.

Claims

1. A method for communicating in a storage network, wherein data is communicated using interchangeable resource indicators, the method comprising: A request for an interchange resource indicator is received at an adapter for the storage network, the adapter being an electronic device and including an electronic context memory, the electronic context memory including a first storage device and a second cache; If the swap resource indicator is within a first range of swap resource indicators, then the swap resource indicator is provided from the first storage device in the adapter; If the swap resource indicator is within the second range of swap resource indicators and is stored in the second cache, then the swap resource indicator is provided from the second cache in the adapter; If the swap resource indicator is within the second range of swap resource indicators and is not stored in the second cache, then the swap resource indicator is provided from host memory located outside the adapter; and If the swap resource indicator is within the second range of the swap resource indicator and the page address translation entry indicates a remapping operation with a direct context memory address, then a remapping swap resource indicator is provided from the second cache in the electronic context memory.

2. The method according to claim 1, further comprising: For a specific input / output switching phase, the switching resource indicator is fixed in the second cache.

3. The method of claim 2, wherein the specific input / output swapping phase is a non-terminal sequence phase.

4. The method of claim 1, further comprising: If the fixed limit is not reached, the swap resource indicator is fixed in the second cache based on first-come, first-served service.

5. The method of claim 4, wherein the fixed limitation includes a programmable fixed PIN count for each set of indices.

6. The method of claim 1, further comprising: Based on the requested address, the swap resource indicator is fixed in the second cache within the specific programmable virtual context address (VCA) range.

7. The method of claim 1, further comprising: Based on the size of the request within a specific programmable fixed I / O size range, the swap resource indicator is fixed in the second cache.

8. The method of claim 1, further comprising: Based on the swap resource indicator corresponding to the request, within a specific swap resource indicator range, the swap resource indicator is fixed in the second cache.

9. The method of claim 1, further comprising: The interchangeable resource indicator is fixed in the second cache based on Fibre Channel, SCSI, NVMe, or Fibre Channel Connectivity Protocol FICON.

10. The method of claim 1, further comprising: The swap resource indicator is fixed in the second cache based on the work queue WQ configuration file.

11. The method of claim 1, further comprising: The interchangeable resource indicator is fixed in the second cache based on a fixed PIN bit indicated by the host driver.

12. A host bus adapter for a storage network, the storage network including a host having host memory, the host bus adapter comprising: A context memory, the context memory including a first storage device for non-cacheable swap resource indicators and a second cache for cacheable swap resource indicators, wherein the cacheable swap resource indicators include cacheable remapped swap resource indicators not supported by the host memory; and The host supports a storage unit, which is located on the host bus adapter and configured to: In response to a request, the interchange resource indicator is provided from the first storage device if the interchange resource indicator is within a first range of the interchange resource indicator; In response to the request, if the interchange resource indicator is within a second range of the interchange resource indicator and stored in the second cache, the interchange resource indicator is provided from the second cache in the host bus adapter; In response to the request, if the swap resource indicator is within the second range of the swap resource indicator and is not stored in the second cache, the swap resource indicator is provided from the host memory; and If the swap resource indicator is within the second range of the swap resource indicator and the page address translation entry indicates a remapping operation with a direct context memory address, then a remapping swap resource indicator is provided from the second cache in the context memory.

13. The host bus adapter of claim 12, further comprising: A direct memory access engine that supports communication with the host's storage units; A protocol offloading engine that communicates with the host-supported storage unit; and wherein the host-supported storage unit includes a crossbar switch, a page address translation table including a page translation cache, and a cache controller.

14. The host bus adapter of claim 12, wherein the host bus adapter is a converged Ethernet network adapter.

15. The host bus adapter of claim 12, further comprising: A direct memory access engine that supports communication with the host's storage units; and A protocol offloading engine that supports communication with the host's storage units.

16. The host bus adapter of claim 12, wherein the host support storage unit is configured to provide optional fixed operation for the cacheable interchangeable resource indicator based on at least one of an input / output phase, a first-in-last-out constraint, a region of the virtual context address associated with the cacheable interchangeable resource indicator, a protocol associated with the cacheable interchangeable resource indicator, the size of the transaction associated with the cacheable interchangeable resource indicator, or work queue information associated with the cacheable interchangeable resource indicator.

17. The host bus adapter of claim 12, further comprising: Assign and deallocate servers configured to assign interchangeable resource indicators for cacheable and non-cacheable resources based on address ranges.

18. The host bus adapter of claim 12, wherein pool selection automatically provides a first-tier pool to link to a second-tier pool, and allocates second-tier pool resources when the resources of the first-tier pool are exhausted.

19. The host bus adapter of claim 12, wherein pool selection automatically provides a first-level pool to link to a second-level pool, and allocates resources to the second-level pool when the resources of the first-level pool are exhausted, and wherein if the first-level pool is empty, the requesting agent does not need to make two swap resource indicator (XRI) allocation requests.

Citation Information

Patent Citations

  • Methods and apparatus for managing page crossing instructions with different cacheability

    CN104662520A

  • Method and apparatus for pinning memory pages in a multi-level system memory

    CN108139983A

  • N-Port virtualization driver-based application programming interface and split driver implementation

    US20070174851A1