Method of operating a storage system and method of partitioning a hierarchy of storage resources
By monitoring and dynamically adjusting the storage system partitions, the problem of underutilization of top-level resources was solved, achieving efficient partition management and performance optimization of the storage system, improving cache hit rate and reducing I/O management costs.
Patent Information
- Application Number
- CN202110953967.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-04
- Filing Date
- 2021-08-19
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2041-08-19
AI Technical Summary
Existing storage systems fail to effectively utilize top-level storage resources in partition management, resulting in poor performance and efficiency. In particular, it is difficult to optimize partitioning under varying workloads and sudden loads among different clients.
By monitoring client workloads, storage resource partitioning is dynamically adjusted. Based on workload changes and performance estimates, proactive partitioning schemes and prior knowledge are used to optimize partitioning, enabling flexible allocation and reallocation of storage resources.
It improves the overall performance and cache hit rate of the storage system, reduces I/O management costs, and ensures fair use and efficient allocation of storage resources.
Smart Images

Figure CN114385073B_ABST
Abstract
Description
[0001] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 088,447, filed October 6, 2020, entitled "Systems, Methods, and Devices for Disaggregated Storage with Partition Management," and U.S. Patent Application No. 17 / 168,172, filed February 4, 2021, which are incorporated herein by reference. Technical Field
[0002] This disclosure relates generally to data processing, and more specifically to systems, methods, and apparatus for partition management of storage resources. Background Technology
[0003] A storage system can divide storage resources into partitions for use by one or more storage clients.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the invention, and therefore may contain information that does not constitute prior art. Summary of the Invention
[0005] A method of operating a storage system may include: allocating a first partition of a tier of storage resources to a first client, wherein the tier operates at least partially as a storage cache; allocating a second partition of the tier of storage resources to a second client; monitoring the workload of the first client; monitoring the workload of the second client; and reassigning the first partition of the tier of storage resources to the first client based on the monitored workload of the first client and the monitored workload of the second client. The method may further include: reassigning the second partition of the tier of storage resources to the second client based on the monitored workload of the first client and the monitored workload of the second client. The first partition may be reassigned based on the input and / or output (I / O) requirements of the first client's workload. The first partition may be reassigned based on the I / O requirements of the second client's workload. The first partition may be reassigned based on an estimated performance change of the first client's workload. The first partition may be reassigned based on an estimated performance change of the second client's workload. The first partition may be reassigned based on the read / write ratio of the first client's workload. The first partition may be reassigned based on the working set of the first client's workload. The first partition may be reassigned based on the working volume of the first client's workload.
[0006] A system may include: a storage tier comprising storage resources configured to operate at least partially as a storage cache; an application server tier configured to handle I / O requests to the storage tier from a first client and a second client; monitoring logic configured to monitor the workload of the first client and the second client; decision logic configured to determine an adjusted partitioning scheme based on the monitored workload of the first client and the monitored workload of the second client; and partitioning logic configured to allocate a first partition of the storage resource tier to the first client, allocate a second partition of the storage resource tier to the second client, and reassign the first partition of the storage resource tier to the first client based on the adjusted partitioning scheme. The monitoring logic may include an I / O filter. The monitoring logic may include a super-supervisor. The application server tier may include an interface configured to provide a standardized I / O interface to the first client and the second client.
[0007] A method for partitioning storage resources into tiers may include: determining read and write working volumes for a first client and a second client of the tier; determining a workload type based on the read and write working volumes; and partitioning the storage resources into tiers between the first and second clients based on the workload type. The method may further include determining a read ratio based on the read and write working volumes; and determining a workload type based on the read ratio. The step of determining the workload type based on the read ratio may include comparing the read ratio to a threshold. The method may further include determining a working set size based on the workload type; and partitioning the storage resources into tiers between the first and second clients based on the working set size. The method may further include determining a working volume size based on the workload type; and partitioning the storage resources into tiers between the first and second clients based on the working volume size.
[0008] A method for detecting bursts of access to a tier of storage resources may include: monitoring the workload of the tier of storage resources, determining the burstability of the workload, and detecting bursts of access based on the burstability. The step of detecting bursts of access based on burstability may include comparing the burstability with a burst threshold. The step of determining the burstability of the workload may include determining the read intensity of the workload. The step of determining the burstability of the workload may include determining changes in the working volume of the workload and calculating the burstability based on the changes in the working volume of the workload. The step of determining the burstability of the workload may include determining changes in the working set of the workload and calculating the burstability based on the changes in the working set of the workload.
[0009] A method for partitioning storage resources by hierarchy may include: monitoring the workload of the storage resource hierarchy; determining a first cache requirement of a first client based on the monitored workload; determining a second cache requirement of a second client based on the monitored workload; allocating a first partition of the storage resource hierarchy to the first client at least partially based on the first cache requirement; and allocating a second partition of the storage resource hierarchy to the second client at least partially based on the second cache requirement. The first and second partitions may be allocated at least partially proportional to the first and second cache requirements. The step of determining the first cache requirement may include determining a first workload amount of the first client, and the step of determining the second cache requirement may include determining a second workload amount of the second client. The first and second workload amounts may be determined at least partially based on the intensity of the monitored workload. The first workload amount may be determined at least partially based on the working volume size of the first client, and the second workload amount may be determined at least partially based on the working volume size of the second client. The first workload amount may be determined at least partially based on the working set size of the first client, and the second workload amount may be determined at least partially based on the working set size of the second client. The method may further include weighting the first workload amount and weighting the second workload amount. The method may further include sorting the first client and the second client based on a weighted first workload and a weighted second workload. The method may also include allocating a first partition of the storage resource hierarchy to the first client based at least partially on the weighted first workload, and allocating a second partition of the storage resource hierarchy to the second client based at least partially on the weighted second workload.
[0010] A method for partitioning storage resources into tiers may include: determining a first partitioning plan for a first client and a second client for the tier of storage resources; determining a first expected cache hit based on the first partitioning plan; determining a second partitioning plan for the first client and the second client for the tier of storage resources; determining a second expected cache hit based on the second partitioning plan; and selecting one of the first partitioning plan or the second partitioning plan based on the first expected cache hit and the second expected cache hit. The method may further include: determining a first expected hit rate for the first client based on the first partitioning plan; determining a second expected hit rate for the second client based on the first partitioning plan; determining a first expected working volume for the first client based on the first partitioning plan; and determining a second expected working volume for the second client based on the first partitioning plan. The step of determining the first expected cache hit may include determining the first expected hit rate of the first client and a weighted sum of the first expected working volume and the second expected hit rate and the second expected working volume of the second client. The method may further include: determining a third partitioning plan for a first client and a second client at the tier of storage resources; determining a third expected cache hit based on the third partitioning plan; and selecting one of the first partitioning plan, the second partitioning plan, or the third partitioning plan based on the first expected cache hit, the second expected cache hit, and the third expected cache hit.
[0011] A method for determining the expected cache hit rate for a client at a tier of storage resources may include: recording I / O transactions by the client at the tier of storage resources; determining a reuse distance based on the recorded I / O transactions; and determining the expected cache hit rate based on the reuse distance. The step of determining the expected cache hit rate may include: determining a distribution function of the reuse distance based on the recorded I / O transactions, and determining the expected cache hit rate based on the distribution function. The expected cache hit rate may be based on the difference between the reuse distance and the cache size.
[0012] A method for determining a client's intended working volume for a tier of storage resources may include: recording I / O transactions by the client for the tier of storage resources; determining a weighted average of the recorded I / O transactions; and determining the intended working volume based on the weighted average. The weighted average may be weighted exponentially.
[0013] A method for partitioning a storage resource hierarchy may include: allocating a first partition of the storage resource hierarchy to a client, wherein the hierarchy operates at least partially as a storage cache; allocating a second partition of the storage resource hierarchy to the client, wherein the size of the second partition may be larger than the size of the first partition; and passively updating the second partition. The step of passively updating the second partition may include updating the second partition based on one or more I / O transactions of the client's workload. The second partition may be allocated based on partition resizing and content update windows.
[0014] A method for hierarchical prefetching of data for storage resources may include: allocating a first partition of the storage resource's hierarchy to a client, wherein the hierarchy operates at least partially as a storage cache; determining a pattern of I / O request sizes for the client; allocating a second partition of the storage resource's hierarchy to the client, wherein the size of the second partition may be larger than the size of the first partition; and prefetching data from the second partition using a prefetch data size based on the I / O request size pattern. The prefetch data size may include a high I / O size popularity for the client. The method may also include applying a scaling factor to the prefetch data size.
[0015] A method for hierarchically partitioning storage resources may include: allocating a first partition in a first region of the storage resource hierarchy to a first client based on prior knowledge of the characteristics of a first client; and allocating a second partition in a second region of the storage resource hierarchy to a second client based on prior knowledge of the characteristics of a second client. The method may further include adjusting the size of the first region and the size of the second region. The size of the first region and the size of the second region may be adjusted proportionally. The size of the first region may be adjusted based on a first demand for the first region, and the size of the second region may be adjusted based on the demand for the second region. The first region may include a read-intensive region, and the size of the first region may be adjusted based on the working set size. The first region may include a write-intensive region, and the size of the first region may be adjusted based on the working volume size. The first region may include a read-write-intensive region, and the size of the first region may be adjusted based on the working volume size. The characteristics may include at least one of read ratio, working volume size, or working set size. Attached Figure Description
[0016] Throughout the accompanying drawings, for illustrative purposes, the drawings are not necessarily drawn to scale, and elements with similar structures or functions may generally be indicated by the same reference numerals or by parts thereof. The drawings are intended only to facilitate the description of the various embodiments described herein. The drawings do not depict every aspect of the teachings disclosed herein and do not limit the scope of the claims. To prevent obscurity, not all components, connections, etc., are shown, and not all components will have reference numerals. However, the pattern of component configuration can be readily understood from the drawings. The drawings, together with the description, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0017] Figure 1 An embodiment of a storage system architecture based on a disclosed example embodiment is shown.
[0018] Figure 2 An example embodiment of a storage system architecture based on a disclosed example embodiment is shown.
[0019] Figure 3 An embodiment of a storage system software architecture based on a disclosed example embodiment is shown.
[0020] Figure 4 An example embodiment of the inputs and outputs of an interface for file-oriented storage operations is shown, according to a disclosed example embodiment.
[0021] Figure 5 An example embodiment of the inputs and outputs of an interface for object-oriented storage operations, according to a disclosed example embodiment, is shown.
[0022] Figure 6 An example embodiment of a storage layer is shown, based on a disclosed example embodiment.
[0023] Figure 7 An example embodiment of a workflow architecture for a partition manager system is shown, based on a disclosed example embodiment.
[0024] Figure 8 An example embodiment of a time window for a workflow of a partition manager system, based on a disclosed example embodiment, is shown.
[0025] Figure 9 An example embodiment of a hybrid read and write workload is shown, based on a disclosed example embodiment.
[0026] Figure 10 An example embodiment of a hybrid read and write workload is shown, based on a disclosed example embodiment.
[0027] Figure 11 An example embodiment of a burst calculation and burst detection method according to a disclosed example embodiment is shown.
[0028] Figure 12 An example embodiment of a partitioning method for bursty workloads, according to a disclosed example embodiment, is shown.
[0029] Figure 13 An example embodiment of a CDF curve for hit rate estimation operation is shown according to a disclosed example embodiment.
[0030] Figure 14 An example embodiment of a partitioning and content update workflow is shown, based on a disclosed example embodiment.
[0031] Figure 15 An example embodiment of an adaptive prefetch size adjustment method according to a disclosed example embodiment is shown.
[0032] Figure 16 An example embodiment of an adaptive prefetch size adjustment method according to a disclosed example embodiment is shown.
[0033] Figure 17 An example embodiment of a prior knowledge-based zoning method according to a disclosed example embodiment is shown.
[0034] Figure 18 An example embodiment of a prior knowledge-based zoning workflow is shown, based on a disclosed example embodiment.
[0035] Figure 19 An embodiment of a method for operating a storage system according to a disclosed example embodiment is shown. Detailed Implementation
[0036] 1. Introduction
[0037] 1.1 Overview
[0038] A storage system with one or more tiers of storage resources can divide one or more of the tiers into separate partitions, each accessible to one of a plurality of storage clients. For example, storage resources can be divided into different tiers based on access characteristics such as performance, latency, etc. In some embodiments, a partition manager system according to the disclosed example embodiments can periodically and / or dynamically adjust the size of one or more partitions in one or more tiers for one or more storage clients. Partition resizing processing (which may be referred to as repartitioning) can improve or optimize the overall performance of the storage system and / or ensure fairness to storage clients. In some embodiments, a partition manager system according to the disclosed example embodiments can provide automatic repartitioning decision-making and / or operation based on one or more factors such as runtime workload analysis, performance improvement estimates, Quality of Service (QoS), Service Level Agreement (SLA), etc.
[0039] 1.2 Storage workload
[0040] In some embodiments of storage systems used for game streaming, a large percentage of gameplay input and / or output (I / O) workloads can be read-intensive or even read-only. Some embodiments of the partition manager systems and methods according to the disclosed example embodiments can provide different cache repartitioning strategies based on the difference between read-intensive and non-read-intensive workloads.
[0041] In some embodiments of storage systems, workloads from different clients may exhibit different behaviors, and even the same workload for a single client may vary during runtime. Some embodiments of the partition manager system and method according to the disclosed example embodiments may allocate top-level storage resources equally to each client and / or allow clients to freely compete for top-level storage resources based on their client needs. However, in some embodiments, one or both of these techniques may not fully utilize the storage resources including the top-level storage resources. Some embodiments of the partition manager system and method according to the disclosed example embodiments may periodically capture and / or predict I / O changes to adaptively reallocate storage resources, including the top-level resources, by using different partitioning methods for workloads with different burst levels. Depending on the implementation details, this can improve storage cache hit rates and / or reduce I / O management costs.
[0042] In some embodiments of the storage system, the I / O size distribution may differ for different workloads. Some embodiments of the partition manager system and method according to the disclosed example embodiments may use adaptive I / O sizes for first-second-tier prefetch operations. Depending on the implementation details, this can improve overall hit rate and / or reduce I / O latency.
[0043] In some embodiments of the storage system, during the deployment phase, there may not be sufficient client runtime data available for efficient allocation of top-level cache storage to clients. Some embodiments of the partition manager system and method according to the disclosed example embodiments may use prior client knowledge (e.g., a priori knowledge) such as from a client pattern library, vendor-selected hardware and / or software, QoS, SLA, etc., to perform initial partitioning of top-level cache storage to clients. In some embodiments, different clients may be placed in different workload regions based on prior knowledge of one or more factors (e.g., workload read ratio, workload working set size, etc.). In some embodiments, zoning (or region partitioning) can be physical zoning, virtual zoning, and / or any combination thereof. In physical zoning, the granularity may be greater than or equal to the size of the storage device (e.g., grouping multiple storage devices into a region). In virtual zoning, the granularity may be less than or equal to the size of the storage device (e.g., zoning within a storage device).
[0044] 1.3 Partition Management
[0045] Some embodiments of the partition management system and / or method according to the disclosed example embodiments may periodically and / or dynamically repartition one or more tiers (e.g., the top tier) of a storage system that can be shared by multiple clients. Some embodiments may allocate top-level storage resources to clients that need them (e.g., based on dynamic I / O requirements) and / or to clients who can benefit from them the most (e.g., based on estimates of performance improvements). Allocation may be based on, for example, analysis of runtime workload parameters (such as read / write ratio, working set changes, working volume changes, etc.).
[0046] Some embodiments may provide a partition optimization framework that takes into account various factors of some or all storage clients (such as workload changes during recent workload monitoring windows, the weight of one or more clients (e.g., based on QoS, SLA, etc.), the estimated hit rate that can be expected as the partition size increases or decreases).
[0047] Some embodiments provide burst detection methods to determine whether the current workload is bursty. If the current workload is bursty, the system may not have enough time to apply the partitioning optimization framework before the I / O burst arrives. Therefore, for bursty workloads, an aggressive partitioning scheme can be used to quickly partition the top level of storage across clients. In some embodiments, the aggressive partitioning scheme can scale partitions proportionally based on recent demand, for example, determined by the working volume size (for non-read-intensive workloads) or the working set size (for read-intensive workloads).
[0048] Some implementations enable techniques that perform initial partitioning by separating clients into different regions based on prior knowledge during the deployment phase. Depending on the implementation details, region-based partitioning techniques can reduce device-side write amplification, overfeeding, total ownership costs, etc., and can improve latency, throughput, predictability, etc.
[0049] 1.4 Storage Type
[0050] The partition management system, method, device, workflow, etc., according to the disclosed example embodiments can be used with any type of storage (such as direct-attached storage (DAS), storage area network (SAN), decomposed storage, etc.).
[0051] In a DAS system according to a disclosed example embodiment, storage devices (such as solid-state drives (SSDs) and hard disk drives (HDDs)) can be attached to a single server. This configuration can provide relatively high performance for any workload running on that server. The storage device capacity and / or performance can be available to the server, and the capacity and / or performance can be increased (by adding drives to the server) or expanded (by adding additional servers).
[0052] In a SAN system according to a disclosed example embodiment, storage devices may be arranged in a storage array, which may be configured to be available to one or more servers on a network. The SAN system can distribute storage to many servers (e.g., dozens or hundreds of servers), which can improve capacity utilization.
[0053] In some embodiments of cloud computing according to the disclosed example embodiments, the client device may be implemented as a light terminal capable of assigning tasks and / or collecting the results of assigned tasks, while heavy computing tasks may be performed on a remote distributed server cluster. This light terminal / heavy data center architecture may involve highly available storage systems. In some embodiments, storage input and / or output (I / O) may be a bottleneck, for example, in a data center.
[0054] In a distributed storage system according to the disclosed example embodiments, a number of storage devices may be used as a logical storage pool that can be allocated, for example, to any server on the network via a high-performance network architecture. In some embodiments, the distributed storage system may utilize the flexibility of a SAN to provide the performance of local storage, and / or may be dynamically reconfigurable, which allows physical resources to be reconfigured to improve or maximize performance and / or reduce latency.
[0055] 2. Architecture
[0056] Figure 1 An embodiment of a storage system architecture based on a disclosed example embodiment is shown. Figure 1 The architecture shown can represent hardware, software, workflows, and / or any combination thereof.
[0057] Figure 1 The embodiments shown may include an application server layer 104 and a storage layer 106. Storage layer 106 may include any number and / or type of storage resources, which may be configured as a pool with one or more tiers (including those described below). For example, a first tier of storage resources may operate as a storage cache for one or more other tiers of storage resources. One or more tiers may be partitioned into partitions that can be allocated to one or more storage clients 102.
[0058] Application server layer 104 may include any number and / or type of compute and / or I / O resources, which are configured to enable one or more storage clients 102 to access the storage resource pool in storage layer 106.
[0059] Application server layer 104 can be connected to one or more of storage clients 102 via one or more connections 108. Application server layer 104 can be connected to storage layer 106 via one or more connections 110. Connections 108 and 110 can be implemented using any number and / or type of networks, interconnections, etc. (including those described below).
[0060] Application server layer 104 and storage layer 106 may each include logic 140 and 142, which can be used to implement any function performed by the respective layer (such as monitoring system operation, making decisions, performing system operations, performing calculations, etc.). In some embodiments, logic 140 and 142 may implement any of the techniques disclosed herein (such as monitoring the workload of one or more layers of storage layer 106, determining the read intensity of the workload, detecting bursts in I / O access, determining working volumes, determining working sets, determining repartitioning strategies, partitioning one or more layers of storage layer, implementing partitioning strategies, implementing client partitioning techniques, etc.).
[0061] Logics 140 and 142, as well as any methods, techniques, processes, etc., described herein, may be implemented using hardware, software, or any combination thereof. For example, in some embodiments, either logic 140 or 142 may be implemented using the following: combinational logic, sequential logic, one or more timers, counters, registers, state machines, complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), complex instruction set computer (CISC) processors, and / or reduced instruction set computer (RISC) processors, etc.
[0062] against Figure 1 The operations and / or components described in the embodiments shown herein, as well as in all other embodiments described herein, are example operations and / or example components. In some embodiments, some operations and / or components may be omitted and / or other operations and / or components may be included. Furthermore, in some embodiments, the temporal and / or spatial order of operations and / or components may vary. Although some components may be shown as separate components, in some embodiments, some components shown separately may be integrated into a single component, and / or some components shown as a single component may be implemented using multiple components.
[0063] 2.1 Data Center Architecture
[0064] Figure 2 An example embodiment of a storage system architecture based on a disclosed example embodiment is shown. Figure 2 The system shown may include a cloud client layer (CCL, or client cloud layer) 202, an application server layer (ASL) 204, and a storage layer 206. In some embodiments, the storage layer 206 may be implemented as a back-end storage layer (BSL). The client cloud layer 202 and the application server layer 204 may be connected, for example, via one or more networks 208. The application server layer 204 and the storage layer 206 may be connected, for example, via I / O path 210.
[0065] In some embodiments, the client cloud layer 202 may include any number and / or type of cloud client devices 203 (e.g., heterogeneous devices) and may ultimately be connected to the application server layer 204 via one or more of networks 208. The cloud client devices 203 may be connected to the client cloud layer 202 and / or connected between client cloud layers 202 via any number and / or type of network. Examples of cloud client devices 203 may include desktop and / or laptop computers, servers, smartphones, tablets, printers, scanners, internet-connected appliances, etc. The client cloud layer 202 may handle any number of the following conditions and / or processes: I / O congestion, latency and / or timeouts, workload surges, dropped packets, etc.
[0066] One or more networks 208 can be implemented using any number and / or type of network (including local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), etc.) based on any number and / or type of network hardware, software, protocols, etc. Examples of network types may include Ethernet, wireless bandwidth, Fibre Channel, Wi-Fi, Bluetooth, etc.
[0067] In some embodiments, the application server layer 204 may be implemented using any number and / or type of server equipment, including motherboards, switch boards, I / O cards and / or modules, backplanes, midplanes, network interface cards (NICs), etc., configured in one or more chassis, server racks, server rack groups, data rooms, data centers, edge data centers, mobile edge data centers, and / or any combination thereof. The application server layer 204 may handle any number of the following conditions and / or processes: resource allocation, data contention, I / O latency and / or errors, I / O waiting and / or queuing, I / O timeouts, locking (e.g., excessive locking), system instability, central processing unit (CPU) overhead, etc.
[0068] I / O path 210 can be implemented using any number and / or type of connections and combinations thereof (such as NVMe via a structured NVMe (NVMe-oF) etc.), including one or more interconnects and / or protocols (such as PCIe, Compute Express Link (CXL), Advanced Scalable Interface (AXI) etc.), one or more storage connections and / or protocols (such as Serial ATA (SATA), Serial Attached SCSI (SAS), Non-Volatile Memory Express (NVMe) etc.), and one or more network connections and / or protocols (such as Ethernet, Fibre Channel, Wireless Bandwidth etc.).
[0069] Storage tier 208 can be implemented using any number and / or type of storage devices, such as solid-state drives (SSDs), hard disk drives (HDDs), optical disc drives, drives based on any type of persistent memory (such as cross-gridded non-volatile memory with varying bulk resistance), and / or any combination thereof. Such devices can be implemented in any form and / or configuration (e.g., storage devices with form factors such as 3.5-inch, 2.5-inch, 1.8-inch, M.2, etc., and / or configured using any connector such as SATA, SAS, U.2, etc.). Some embodiments can be implemented entirely or partially using server chassis, server racks, data rooms, data centers, edge data centers, mobile edge data centers, and / or any combination thereof, and / or within server chassis, server racks, data rooms, data centers, edge data centers, mobile edge data centers, and / or any combination thereof.
[0070] In some embodiments, the combination of application server layer 204 and storage layer 206 can implement a data center. Therefore, depending on the implementation details, application server layer 204 and storage layer 206 can be considered as a near-field arrangement, while client cloud layer 202 and application server layer 204 can be considered as a far-field arrangement.
[0071] Figure 2 The embodiments shown may also include one or more partition manager applications 212, which may interface between the client cloud layer 202 and the application server layer 204 and / or between the application server layer 204 and the storage layer 206, for example. In some embodiments, the partition manager application may determine how to allocate one or more tiers of storage resources in the storage layer 206 (e.g., the top tier of storage resources) to different clients (e.g., cloud client 203) and / or their applications.
[0072] While the disclosed principles are not limited to any specific implementation details, some example embodiments may be implemented as follows for illustrative purposes. The client cloud layer 202 may be connected to the application server layer 204 via one or more LAN, MAN, or WAN connections utilizing multi-Gb Ethernet connections (e.g., 10Gb / 25Gb / 40Gb / 100Gb Ethernet). The application server layer 204 may employ one or more PCIe interconnects, which may deliver throughput exceeding 8GB / s using PCIe Gen 4. The application server layer 204 may use one or more NVMe-oF connections to connect to the storage layer 206; for example, NVMe storage devices may deliver millions of input and / or output operations per second (IOPS) via one or more NVMe-oF connections. Depending on the implementation details, such embodiments may reduce or eliminate bandwidth limitations and / or bottlenecks.
[0073] 2.2 Software I / O Path Architecture
[0074] Figure 3 An embodiment of a storage system software architecture according to a disclosed example embodiment is shown. For example, Figure 3 The software architecture shown can be compared with Figure 2 The storage system architecture shown is used together, but it can also be used with other systems. Figure 3 The embodiments shown may include a cloud client layer (CCL) 302, an application server layer (ASL) 304, and a storage layer (e.g., BSL) 306.
[0075] In some embodiments, the client cloud layer 302 may include one or more clients 303, each client 303 running one or more operating systems (OS) 316 and one or more applications (APPs) 314. In this example, there may be N clients (where N is a positive integer): client 1, client 2, ..., client N. In some embodiments, client 303 may be modeled as operating system 316 and / or application 314. Depending on the implementation details, client 303 may have one or more operating systems 316 and / or applications 314. Furthermore, each operating system 316 may host any number of applications 314.
[0076] Depending on the implementation details, some or all of the components in the client can be isolated from each other. Examples of clients may include... Figure 2 This includes any individual client 203 shown, as well as one or more virtual machines (VMs) that may be rented to users from cloud service providers. Depending on the implementation details, clients 303 (such as VMs) may have different workload characteristics based on, for example, user applications, and therefore may have different levels of sensitivity to storage I / O speed.
[0077] Each client 303 (also referred to as a storage client) can connect to the application server layer 304 via network connection 318, which can, for example, use any number and / or type of network connection (such as, as mentioned above). Figure 2 This is achieved through a network connection (described in the diagram). For a VM, the network connection can be implemented as a virtual network connection.
[0078] Application server layer 304 may include I / O filter 320, partition manager interface (or interface) 322, and partition decision maker 324. In some embodiments, I / O filter 320 may be implemented using one or more hypervisors (also known as virtual machine monitors) or as part of one or more hypervisors, which may also be used, for example, to implement client VMs. I / O filter 320 may collect I / O-related statistics from the client and provide these statistics to partition decision maker 324. In embodiments including a hypervisor, hypervisor software may be used to host one or more VMs, and the hypervisor software may have built-in I / O filter functions (e.g., "IOFilter()") for processing I / O requests and / or collecting I / O statistics. For illustrative purposes, Figure 3 The I / O filter 320 in the embodiment shown can be shown as a super-supervisor, but other types of I / O filters can be used.
[0079] In some embodiments, partition manager interface 322 may provide a standard I / O interface to one or more of clients 303. Depending on implementation details, as described in more detail below, partition manager interface 322 may provide operation for file-based systems, key-value object storage systems, and / or any other type of storage system and / or combinations thereof. Partition manager interface 322 may be interfaced with storage layer 306 via I / O path 326, which may use any number and / or type of connections (e.g., see reference 322). Figure 2 The connection described in I / O path 210 shown in the figure is used to achieve this.
[0080] The partitioning decision maker 324 can make decisions at various times (e.g., at periodic intervals) to trigger repartitioning of one or more tiers of storage tier 306. For example, if one or more clients are migrated to another tier, the repartitioning can be based on predicted client performance and the corresponding migration overhead.
[0081] Storage layer 306 is operable on one or more tiers 328 of a storage device, which are arranged, for example, as a pool of storage resources that can be allocated among different clients. In this example, there may be M tiers (where M is a positive integer): tier 1, tier 2, ..., tier M. In some embodiments, different tiers may have different characteristics (e.g., a trade-off between access speed, capacity, etc.). Figure 3 In the embodiments shown, all storage devices may be shown as various types of SSDs, but storage resources of any type and / or configuration may be used.
[0082] Storage tier 306 may also include a partitioner 330, which may perform repartitioning operations and / or trigger one or more content updates between tiers 328 (e.g., between the top tier (tier 1) and the second tier (tier 2)) based on partitioning decisions from partitioning decision maker 324. In some embodiments, partitioner 330 of a partition manager system according to a disclosed example embodiment may adjust the allocation of partitions in one or more tiers 328 (e.g., the top tier) in a storage pool for each client 303. Each client 303 has access to one or more corresponding partitions in one or more tiers 328 allocated by partitioner 330.
[0083] although Figure 3 Some components shown may be illustrated as separate components associated with other specific components, but in other embodiments, components may be combined, separated, and / or associated with other components. For example, in some embodiments, partition decision maker 324 and partitioner 330 may be combined into a single component located in application server layer 304, storage layer 306, and / or some other location.
[0084] The partition manager interface 322 enables the partition manager system according to the disclosed example embodiments to provide one or more interfaces and / or functions to the client 303, for example, through an application programming interface (API).
[0085] When an I / O request from a client arrives at the application server layer 304, the Super Supervisor I / O Filter (HIF) 320 distinguishes each I / O stream and performs statistical analysis (such as determining workload changes and / or top-level storage cache hit rates). The I / O request can then proceed to the partition manager interface 322, which provides standard I / O interfaces for various clients.
[0086] Algorithm 1 illustrates some example operations of an embodiment that provides an interface for file-oriented storage operations to client 303 on storage layer 306, according to the disclosed example embodiments. When one of the functions shown in Algorithm 1 is called, the return type provided by the algorithm may be indicated by the ">" symbol.
[0087] Figure 4 An example embodiment of the inputs and outputs of an interface for file-oriented storage operations is shown, according to a disclosed example embodiment. Figure 4 The embodiments shown can be used, for example, with the embodiments shown in Algorithm 1. Figure 4 Each function shown can be called using the corresponding input. Then, as... Figure 4 As shown, the function can respond using the corresponding output.
[0088] Algorithm 1
[0089]
[0090]
[0091] Algorithm 2 illustrates some example operations of an embodiment that provides an interface for object-oriented storage operations on storage layer 306 for client 303, according to the disclosed example embodiments. When one of the functions shown in Algorithm 2 is called, the return type provided by the algorithm may be indicated by the ">" symbol.
[0092] Algorithm 2
[0093]
[0094] Figure 5 An example embodiment of the inputs and outputs of an interface for object-oriented storage operations, according to a disclosed example embodiment, is shown. Figure 5 The embodiments shown can be used, for example, with the embodiments shown in Algorithm 2. Figure 5 Each function shown can be called using the corresponding input. Then, as... Figure 5 As shown, the function can respond with the corresponding output.
[0095] In Algorithm 1 and Figure 4 In the file-oriented embodiment shown, the interface can operate on files (which can be pointed to by paths to files). In Algorithm 2 and Figure 5 In the object-oriented embodiment shown, the interface can operate on key-value pairs (which can be pointed to by a key).
[0096] In some embodiments, the partition manager system according to the disclosed example embodiments may adapt to different storage systems based on, for example, client requests. Furthermore, depending on implementation details, the adaptation process may be transparent to the cloud client layer 302 and / or storage layer 306. In some embodiments, the partition manager interface 322 may be extended to support additional I / O storage types (e.g., custom I / O storage types) based on client requirements for additional storage types.
[0097] After an I / O request passes through the partition manager interface 322, the partition decision maker 324 can analyze the behavior of each client's application and corresponding feedback from one or more of the storage tiers 328 (e.g., from the top tier that can operate as a storage cache). In some embodiments, the partition decision maker 324 can use this analysis and feedback to adjust the partitions of each client's top tier (tier 1) based on factors such as Quality of Service (QoS) and / or the actual workload behavior of each client, for example, to improve or optimize the overall performance of the entire storage system and / or ensure fairness. The partition decision maker 324 can periodically and / or dynamically adjust the storage partition size of one or more of the tiers 328 for each client (e.g., the top tier of the highest performance tier) based on runtime workload analysis results.
[0098] In some embodiments of a storage device having two or more levels, the algorithms and / or operations described herein may be implemented between multiple pairs of levels (e.g., between level 1 and level 2, between level 2 and level 3, between level 1 and level 3, etc.).
[0099] Table 1 illustrates some example implementation details of an example embodiment of the full SSD storage layer 306 according to the disclosed example embodiments. Each layer can be implemented using different types of storage devices, and different types of storage devices can have different characteristics suitable for that layer. The details shown in Table 1 can be, for example, compared with... Figure 3 It is used in conjunction with the embodiment of storage layer 306 shown in the figure.
[0100] In some embodiments, example values, example parameters, etc. may be provided for the purpose of illustrating the principle, but the principle is not limited to these examples, and other values, parameters, etc. may be used.
[0101] Table 1
[0102]
[0103]
[0104] Referring to Table 1, Tier 1 can be implemented using storage devices that provide extremely high performance based on persistent memory (such as cross-grid non-volatile memory with varying bulk resistance). Tier 1 can operate, for example, as a storage cache for one or more of the other tiers. Tier 2 can be implemented using single-cell (SLC) NVMe SSDs, which offer high performance but not as high as Tier 1. Tier 3 can be implemented using multi-cell (MLC) NVMe SSDs, which offer medium-level performance. Tiers 4 and 5 can be implemented using three-cell (TLC) NVMe SSDs and four-cell (QLC) NVMe SSDs, respectively, which offer successively lower performance but higher capacity. Such embodiments can be used, for example, in cloud computing applications. The partition manager system according to the disclosed example embodiments can periodically and / or dynamically adjust the partitions of each client in the top tier, for example, which can operate as a storage cache for a storage system (such as a distributed storage system). In some embodiments, the partition manager system according to the disclosed example embodiments may, for example, apply a similar adjustment method between level 2 and level 3, wherein level 2 may operate as a cache level for level 3, etc.
[0105] Figure 6 An example embodiment of a storage layer according to a disclosed example embodiment is shown. Figure 6 In the embodiment shown, storage layer 606 can be implemented as a layer 628 having a combination of SSDs and HDDs. Table 2 shows examples of such layers. Figure 6 Some example implementation details of using storage layer 606 together are shown in the figure.
[0106] Table 2
[0107]
[0108] Referring to Table 2, Tier 1 can be implemented using SLC NVMe SSDs, which offer extremely high performance and relatively fine granularity. Tier 1 can operate, for example, as a storage cache for one or more of the other tiers. Tier 2 can be implemented using QLC NVMe SSDs, which offer relatively high performance and relatively coarse granularity. Tier 3 can be implemented using HDDs, which offer relatively low performance and relatively coarse granularity, but at a relatively low cost.
[0109] Figure 6The embodiments shown in Table 2 can be used, for example, in a data center for an online game streaming service. The top-level layer (Level 1) can be used to cache runtime game data (such as runtime instances of the game, metadata, the state of game settings, all or part of hot game images (e.g., hot files), etc.). Some embodiments may use a delta-differential storage approach to further save space. The second layer (Level 2) can have a combination of relatively high performance and relatively large capacity to cache hot game images. Therefore, Level 2 can operate as a cache for Level 3. The third layer (Level 3) can have a large capacity and can be used as a repository for game images (e.g., all game images).
[0110] 2.3 Workflow Architecture
[0111] Figure 7 An example embodiment of a workflow architecture for a partition manager system is shown, based on a disclosed example embodiment. Figure 7 The embodiments shown herein may include, for example, those disclosed herein. Figures 1 to 6 The embodiments shown herein are used in any embodiment, and are consistent with those disclosed herein. Figures 1 to 6 Any of the embodiments shown herein may be used together with, and in conjunction with, the embodiments disclosed herein, including Figures 1 to 6 Any embodiment shown may be used in combination, etc. For illustrative purposes, it may be described in the context of a system having a storage layer. Figure 7 The embodiment shown has a top-level storage tier that can operate as a storage cache for one or more other tiers. However, the principle can be applied to any other configuration of tiers and / or caches.
[0112] Figure 7 The embodiments shown may include a workload monitor subsystem (subsystem 1) 702, a burst detector subsystem (subsystem 2) 710, a policy selector subsystem (subsystem 3) 718, and a partition operator subsystem (subsystem 4) 724. The workload monitor subsystem 702 may include a workload status monitor (component 1) 704 and a hit rate monitor component (component 2) 706. The burst detector subsystem 710 may include a burst degree calculator (component 3) 714 and a burst detector (component 4) 716. The policy selector subsystem 718 may include an active solution component (component 5) 720 and an optimal solution component (component 6) 722. The partition operator subsystem 724 may include a partitioner component (component 7) 726.
[0113] Figure 7The workflow shown can operate throughout the loop to periodically and / or dynamically adjust the storage partition size for each client in one or more tiers of the storage layer. For example, the workflow could adjust the storage partition size for each client at the top level (e.g., the highest performance tier in some implementations) based on runtime workload analysis results.
[0114] In some embodiments, Figure 7 The workflow shown can operate on two time windows that run in parallel on the storage system. Figure 8 An example embodiment of a time window for a workflow in a partition manager system, according to a disclosed example embodiment, is shown. Figure 8 As shown, the first window (window 1) can operate on a relatively short cycle during which the workflow can monitor workload behavior and / or cache status information collected by the workload monitor subsystem 702.
[0115] In some embodiments, window 1 may be implemented as a performance monitoring sliding window (PMSW) 708, wherein workflow and / or cache state data may be sampled, acquired, processed, etc., during an epoch that can be moved over time. In some example implementations, window 1 may have an epoch length T1 of approximately 5 minutes, but any other time period may be used. In some embodiments, window 1 may be set to a frequency at which the burst detector subsystem 710 records workflow and / or cache state (e.g., top-level hit rate) data acquired by the workload monitor subsystem 702. Examples of workload behaviors that may be used by the burst detector subsystem 710 may include the working set (WS) size in the most recent epoch, the working volume (WV) size in the most recent epoch, the read ratio (R) in the most recent epoch, etc., which will be described in more detail below. In some embodiments, the partition manager system according to the disclosed example embodiments may retain workflow and / or cache state data for k most recent consecutive or discrete epochs, wherein k can be any number. In some example implementations, k can be 10. In some embodiments, an exponential moving average window (EMAW) technique may be applied to the data for k epochs.
[0116] In some embodiments, window 2 may be implemented as a Partition Adjustment and Content Update Window (PACUW) 712. Window 2 may determine the repartitioning period at which a new repartitioning assessment can be triggered. At the repartitioning boundary, a burst factor calculator component 714 may be triggered by window 2 to determine a method for calculating burst factor (Bd), which may be based on the read ratio (R) to indicate an impending I / O burst, as described in more detail below. In some embodiments, window 2 may have a window length T2 that is larger than a period of window 1. For example, T2 may be approximately one hour, but any other time period may be used. In some embodiments, implementing a window 2 that is longer than window 1 may reduce the cost of operating the partition manager system because repartitioning can be relatively expensive in terms of bandwidth, power consumption, etc., depending on implementation details. In some embodiments, the relative lengths of window 1 (T1) and window 2 (T2) may be selected to balance the trade-off between performance and overhead. For example, more frequent repartitioning may provide improved performance, but depending on implementation details, the increased overhead associated with repartitioning may outweigh the improved performance. Between repartitions, the burst detector subsystem 710 may continue to receive working set size (WS), working volume size (WV), read ratio (R), and / or other data that may be collected by the workload monitor subsystem 702 during window 1 period.
[0117] Table 3 lists some example embodiments of various aspects related to the workflow of a partition manager system according to the disclosed example embodiments. Table 3 lists each aspect and corresponding symbols and corresponding code symbols that may be used, for example, in one or more algorithms disclosed herein. In some embodiments, the aspects listed in Table 3 may be described as follows.
[0118] The working volume size |V| indicates the total amount of data (e.g., in bytes) accessed on other units of a storage device or storage resource. In a memory-related context (e.g., dynamic random access memory (DRAM)), the working volume size may be referred to as a footprint.
[0119] The working set size |S| indicates the total address range (e.g., in bytes) of the data being accessed, which can be, for example, a set of working addresses of a working volume. In some embodiments, a large working set can cover more storage space. If the cache size is greater than or equal to the workload's working set size, the workload's I / O hit rate can be equal to or close to 100% (e.g., using a Least Recently Used (LRU) cache algorithm).
[0120] The read ratio R can be determined, for example, by dividing the read work volume by the total work volume, where the total work volume can be determined by the sum of the read work volume and the write work volume. In some embodiments, the read ratio can range from zero percent to one hundred percent [0%, 100%], where a higher read ratio indicates a more read-intensive workload.
[0121] Window 1 W1 (e.g., the length of PMSW) can indicate the length of the period T1 during which workload behavior and / or cache status information can be recorded before the start of the monitoring period.
[0122] Window 2 W2 (e.g., the length of PACUW) can indicate the repartitioning period T2. This window can trigger the entire repartitioning operation workflow from the burst detector subsystem 710 through the policy selector subsystem 718 to the partition operator subsystem 724.
[0123] Table 3
[0124] Work volume size WV |V| working set size WS |S| Read ratio R R PMSW era W1 <![CDATA[W1]]> PACUW W2 <![CDATA[W2]]>
[0125] Refer again Figure 7 The burst detector subsystem 710 can calculate burst rates (in some embodiments, bursts may also be referred to as spikes) based on, for example, recently determined different read ratios (R) at intervals that can be determined by window 2. Burst rate calculations can be performed, for example, by a burst rate calculator 714. For read-intensive workloads (such as those determined based on read ratios (R)), burst rates can be calculated based on changes in the working set (WS) size. For non-read-intensive workloads, burst rates can be calculated based on changes in the working volume (WV) and / or imprint size.
[0126] After the burstiness calculator 714 determines the burstiness, it can transmit this information to the burst detector 716, which can compare the burstiness with a preset threshold to determine whether the current workload can be represented as bursty or non-burst. The burst / non-burst determination can be used to determine which type of repartitioning strategy to use in the strategy selector subsystem 718.
[0127] If the current workload is bursty, then Figure 7The workflow shown can invoke the proactive solution component 720 to apply a proactive repartitioning strategy, which can, for example, scale the top-level partitions of each client based on demand during one or more recent PMSW periods. For instance, the proactive approach can scale the partitions of each client based on its most recent working set size change (for read-intensive workloads) or working volume size change (for non-read-intensive workloads). In some embodiments, the proactive strategy can be implemented relatively quickly, enabling the system to react rapidly to anticipated I / O bursts.
[0128] However, if the current workload is non-bursting, the system has sufficient time to apply the optimization framework using the optimal solution component 722, which strives to find the globally best partitioning solution (e.g., by solving, improving, and / or maximizing the objective function to best utilize the top-level (e.g., the cache level) as described below). Therefore, for non-bursting workloads, the optimal solution component 722 can be selected to apply a repartitioning strategy based on the optimization framework, which may consider various factors relevant to all managed clients (such as workload changes in one or more recent PMSW periods, estimated hit rates in the event of partition size increases or decreases, weights for each client based on, for example, Quality of Service (QoS) and / or Service Level Agreement (SLA)).
[0129] The partitioner component 726 can perform the actual repartitioning operation based on the resulting partitioning strategy provided by the policy selector subsystem 718, such as at the top level of the storage level. Figure 7 The workflow shown can then loop back to waiting for the next repartitioning operation to be triggered by window 2 in the burst detector subsystem 710.
[0130] 3. Emergency Detection
[0131] Since higher burstiness allows for less time to react and vice versa, the burst detector subsystem 710 according to the disclosed example embodiment can detect whether the current workload of the storage layer is bursty or non-burst, so that the policy selector subsystem 718 can select a repartitioning policy based on burstiness. In some embodiments, a burst (also referred to as a spike) can be represented by a relatively high number of I / O accesses to a relatively large amount of data that can occur within a relatively short period of time.
[0132] In some embodiments, burst detection operations may consider the I / O patterns of all or most clients (e.g., all I / O client accesses to the storage tier during one or more recent PMSW periods) to determine which strategy to use for repartitioning. However, in some other embodiments, burst determination may focus on a relatively small number of clients (e.g., on a per-client basis).
[0133] In some embodiments, burst rate calculation may involve different methods based on the read ratio (R) of recent accesses to the top-level storage by all or most clients. For example, a preset read ratio threshold e can be used to determine whether a recent workload is read-intensive or non-read-intensive. For illustrative purposes, in some example embodiments disclosed herein, e may be set to 95%. However, in other implementations, e may be set to any suitable value, for example, a threshold for determining read-intensive workloads, within the range of [0%, 100%].
[0134] In some embodiments, if the read ratio is in the range [0, e), the workload can be classified as non-read-intensive. In this case, burstability can be calculated using a change in the size of the working volume. In some embodiments, a working volume change (also referred to as imprinting) can include all records for all touched addresses. Because each new write request can be allocated to a new slot if there is a relatively large amount of write I / O, and repeated write requests for the same address can also be allocated to new slots in the cache, the working volume can be used to calculate burstability for non-read-intensive workloads. Therefore, burstability calculation can be based on a change in the working volume size.
[0135] Figure 9 This illustrates an example embodiment of I / O requests for non-read-intensive workloads, based on a disclosed example embodiment. Figure 9 In the embodiment shown, a top-level (e.g., storage cache) 902 can be displayed from time T0 to T5. Table 904 can indicate each read or write operation performed on data chunks A, B, C, and D, and the amount of data in each chunk. Letters can indicate data chunks, and marked letters (e.g., C') can indicate updated data. From time T0 to T3, data A, B, C, and D can be written to the top-level 902. At time T4, two chunks of B and two chunks of A can be evicted and overwritten with data C'. At time T5, one of chunks of D can be read from the top-level 902. Therefore, the storage system can allocate a new slot for each new write request, and repeated write requests to the same address can also be allocated a new slot in the cache (e.g., four new data chunks C at time T4). Thus, the system can calculate burstiness based on changes in the size of the working volume according to the workload.
[0136] Algorithm 3 illustrates some example operations of an embodiment of burst degree calculation according to the disclosed example embodiments. Algorithm 3 can be, for example, by... Figure 7 The burstiness calculator component 714 (component 3 of subsystem 2) shown in the diagram is executed. In the embodiment shown in Algorithm 3, the code for parameter selection for non-read-intensive cases may be located in lines 4 through 7. In Algorithm 3, curWV may represent the current working volume, preWV may represent the previous working volume, curWS may represent the current working set, and preWS may represent the previous working set.
[0137] Algorithm 3
[0138]
[0139] In some embodiments, if the read ratio is within the range [e, 1], the workload can be classified as read-intensive. In this case, burstiness can be calculated using changes in the size of the working set. In some embodiments, the working set can be represented as a deduplicated version of the working volume.
[0140] Figure 10 This illustrates an example embodiment of I / O requests for read-intensive workloads, based on a disclosed example embodiment. Figure 10 In the embodiment shown, the top-level (e.g., storage cache) 1002 can be displayed from time T0 to T5. Table 1004 can indicate each read or write operation performed on data blocks A, B, C, and D, and the amount of data in each block. Data blocks A, B, C, and D can be read from the top-level 1002 from time T0 to T3. Data block C can be read again at time T4. Data block D can be read again at time T5. Figure 10 As shown, in the case of read-intensive workloads, if the requested I / O content is already cached in the top-level 1002 (and still exists in the cache), the system may not allocate a new slot for duplicate I / O requests. Therefore, the system can calculate burstiness based on changes in the size of the workload's working set (e.g., the unique address already accessed).
[0141] In the embodiment shown in Algorithm 3, the code for selecting parameters for reading dense cases can be located in lines 8 through 10.
[0142] Once a parameter (e.g., working volume size or working set size) has been selected for burst computation, it can be used to calculate burstiness. In some embodiments, such as shown in Equation 1, a piecewise function can be used to calculate burstiness. Some of the notations used in Equation 1 can be represented as follows.
[0143] |V curwin| can represent the size of the working volume during the current PMSW period.
[0144] |V prevwin | can indicate the size of the work volume during the previous PMSW period.
[0145] |S curwin | can represent the size of the working set during the current PMSW period.
[0146] |S prevwin | can represent the size of the working set during the previous PMSW period.
[0147] The increment function can be used to track changes in the working set size (Δ(S)) between the current PMSW window and the previous PMSW window. curwin S prevwin )) or changes in the work volume size (Δ(V) curwin V prevwin (In some embodiments, the same method can be used to consider more than two PMSW windows to calculate the burstiness.) The relative difference between two PMSW windows can then be calculated as the burstiness Bd as follows:
[0148]
[0149] The calculated value of Bd can then be compared with a predefined burst threshold Bt. If Bd > Bt, it indicates that the current workload can be considered bursty. The value of Bt can be preset to any suitable value (e.g., in the range of 40% to 60%). In some embodiments, setting the value of Bt to 50% ensures that the start (rising edge) and end (falling edge) of a burst can be captured. For example, if Bt is set to 50%, it indicates that the workload is determined to be bursty if there is a change of more than 50% of a change in working volume size (for non-read-intensive cases) or working set size (for read-intensive cases).
[0150] In some embodiments, the burstiness calculator component (such as, Figure 7 Some conditional relationships and supporting considerations for component 3) shown in the diagram are as follows. For read-intensive workloads, burstability calculation can be based on changes in the working set size of the current workload. For read-intensive workloads, if the requested I / O content is already cached (and still is), the storage system may not allocate new slots for duplicate I / O requests at the top level (e.g., storage cache). Therefore, burstability calculation can be based on changes in the working set (which may include unique addresses that have been accessed).
[0151] For non-read-intensive workloads, burstability calculation can be based on changes in the working volume size of the current workload. If the workload includes a relatively large amount of write I / O, new slots can be allocated at the top level for each new write request, and duplicate write requests for the same address can also be allocated new slots in the cache. Therefore, burstability calculation can be based on changes in the working volume of the workload. In some embodiments, the working volume can be an imprint of all records that may include all contact addresses. Therefore, the working set can be represented as a deduplicated version of the working volume.
[0152] In some embodiments, a burst detector component (such as, Figure 7 Some conditional relationships and supporting considerations for component 4) shown in the diagram are as follows. For bursty workloads, solving the optimization framework may take more time than is available before the burst arrives. Therefore, an active approach can be applied to the repartitioning operation. In some embodiments, the active approach may adjust the partitions for each client based on recent changes in its working set size (for read-intensive workloads) or working volume size (for non-read-intensive workloads). For non-burst workloads, the partition manager system has sufficient time to solve the optimization framework. Therefore, an optimization approach that strives to find the globally best partitioning solution to achieve the best objective function can be applied to the repartitioning operation.
[0153] Figure 11 An example embodiment of a burst calculation and burst detection method according to a disclosed example embodiment is shown. The method may begin at operation 1102, where the read ratio of the current workload is compared to a read ratio threshold e. If the read ratio is greater than or equal to e, the method may proceed to operation 1104, where changes to the working set are used to calculate the burst level Bd. However, if it is determined in operation 1102 that the read ratio is less than e, the method may proceed to operation 1106, where changes to the working volume are used to calculate the burst level Bd. In operations 1108 to 1110, the burst level Bd is compared to a burst threshold Bt. If the burst level is greater than Bt, the method may proceed to operation 1112, where the current workload is determined to be a burst workload. However, if the burst level is less than or equal to Bt in operation 1110, the method may proceed to operation 1114, where the current workload is determined to be a non-burst workload.
[0154] 4. Partitioning methods for sudden workloads
[0155] When burst I / O is identified by a burst detector (such as burst detector subsystem 710), a partition manager system according to a disclosed example embodiment can proactively adjust the partitions for each client at a storage tier (e.g., the top tier that can operate as a storage cache) based on recent changes in its working set size (for read-intensive workloads) or working volume size (for non-read-intensive workloads). In some embodiments, this can result in relatively rapid repartitioning operations that can adapt to upcoming I / O bursts.
[0156] 4.1 Partitioned Workflow for Burst Workloads
[0157] In some embodiments, the partitioning method for bursty workloads according to the disclosed example embodiments may determine the partition size for each client based at least in part on one or more of the following two aspects of the workload: (1) (e.g., PMSW) workload changes during a recent period; and / or (2) cache hit rate in the current period.
[0158] While not limited to any specific implementation details, the partitioning method for bursty workloads described below can be, for example, based on feedback about the workload from the proactive solution component 720 of the policy selector subsystem 718 (e.g., from...). Figure 7 The data received by the workload monitoring subsystem 702 shown in the figure is used.
[0159] In some embodiments, such as workload changes, the required space size (DSS) can be calculated to quickly adapt to the workload change. For example, a client may experience I / O bursts that can involve a relatively large amount of cache space. However, if there is not sufficient space available in the cache for the I / O burst, the client's latency and / or delay will increase. In some embodiments, DSS can be used to reduce or prevent this type of situation.
[0160] In some embodiments, for example, the storage cache hit rate can be used to calculate the guaranteed minimum space size (GMSS) as the contribution rate of each client's cache. In some embodiments, this can be used to reduce or prevent cache slots from being flushed from the cache due to I / O bursts from other clients.
[0161] Figure 12 An example embodiment of a method for partitioning burst workloads according to a disclosed example embodiment is shown. Figure 12 The method shown can be, for example, by Figure 7The proactive solution component 720 shown in the diagram executes the method. The method can begin at operation 1202, where the current period state (lines 3 and 4 in Algorithm 4) can be obtained from PACUW window 712, which can record multiple period information from PMSW window 708 in workload monitor subsystem 702. At operation 1204, the workload for each client can be calculated as described below. At operations 1206 and 1208, the cache upper and lower limits can be calculated for each client based on the workload, respectively, as described below. At operation 1210, the client list can be sorted based on the workload and weight (in some embodiments, the weight may be based on QoS) for each client. At operation 1212, the cache can be partitioned to each client to meet the upper limit for each client, such that partitions for clients with higher weights (e.g., priority) can be allocated first. At operation 1214, the partitioning plan can be sent to the partition operator subsystem 724 for implementation.
[0162] Any type of weighting function can be used for sorting and / or partitioning operations. In some embodiments, any weighting function can be used to combine workload and QoS. In some example implementations, percentage-based QoS can be multiplied by workload. For example, for client i, the weighted workload can be calculated using workload Wi and QoS percentage Qi as follows:
[0163]
[0164] Here, v can be an iterative client in the set of all clients V.
[0165] Algorithm 4 illustrates some example operations of an embodiment of a method for partitioning bursty workloads according to a disclosed example embodiment. Algorithm 4 can be, for example, derived from... Figure 7 The active solution component 720 shown in the diagram is executed. In the embodiment shown in Algorithm 4, the code for obtaining the current epoch state may be located in lines 3 and 4. In Algorithm 4, curEpochStatus may represent the current epoch state.
[0166] Algorithm 4
[0167]
[0168] 4.2 Workload
[0169] For example, from the perspective of cache slot allocation, the workload |W| can be used to represent the actual workload. In some embodiments, as shown in Equation 3, the workload can be determined using a piecewise function based on different workload read ratios.
[0170]
[0171] In some embodiments, for read-intensive workloads, new cache slots may not be allocated for addresses that are repeatedly requested, as they do not modify content. Therefore, as shown in Equation 3, if the workload during the most recent PMSW period is a non-read-intensive workload, the working volume can be used as the workload. However, if the workload during the most recent PMSW period is a read-intensive workload, the working set can be used as the workload.
[0172] At the end of each PMSW period, the partition management system according to the disclosed example embodiment can check the most recent status of all clients of the shared decomposed storage system. Table 4 shows some example values of status information that can be obtained from the workload monitor subsystem 702 and recorded by the PACUW window 712 in the burst detector subsystem 710. The values provided in Table 4 are for illustrative purposes only, and other values may be obtained and / or used.
[0173] As shown in Table 4, either the working set size |S| or the working volume size |V| can be selected as the workload amount |W| that can be used to calculate the allocation for the next period.
[0174] Table 4
[0175]
[0176] In some embodiments, the status information may include additional information such as the contacted file number, the contacted portion of the file, and the percentage of the contacted file size.
[0177] 4.3 Upper and Lower Limits
[0178] For bursty workloads, the partition management system according to the disclosed example embodiments can proactively and / or proportionally allocate top-level storage resources (which can operate as cache space) to each client based on the client's workload in the recent period. Therefore, the system can attempt to dynamically and / or proactively allocate more top-level storage resources to clients that can have more bursty I / O in the recent period. In some embodiments, the system can proactively advance the cache space for the next period to each client's needs based on the client's current period state. To achieve this allocation, cache space can be proportionally reallocated among all clients using an upper limit as described in Equation 4. In Equation 4, |C| new Indicates the new cache size, |C| max This can represent the maximum cache size, |C| cur It can represent the current cache size, |H| cur It can represent the current hit rate.
[0179]
[0180] In some embodiments, for example, the upper limit may be referred to as the required space size (DSS) because the upper limit may be based on a proportional allocation method that approximates or matches the size of the next period.
[0181] In some embodiments, for example, a lower bound can be used to ensure that one or more clients have a guaranteed minimum cache size (GMSS) such that their cached data is not easily flushed out by bursts in the workload. For example, the last hit rate of the cache can be useful in the next period based on the hit rate of the last period. Therefore, this amount can be used as the guaranteed minimum cache space for each client.
[0182] The code used to obtain the upper and lower limits for each client can be found in lines 7 through 10 of Algorithm 4.
[0183] In some embodiments, there may be special cases where the system can assign the minimum of the upper and lower limits if the lower limit is greater than or equal to the upper limit. That is, upper limit = min(lower limit, upper limit) and lower limit = min(lower limit, upper limit). This can happen, for example, when a client has a relatively high number of cache hits in the current period (e.g., the lower limit may be large), while the workload is not considered bursty in the current period (e.g., the upper limit may be small). In this case, because the upper limit can be based on a predicted target, it can have a greater impact on the partition size. Therefore, DSS can override any existing GMSS. (In some embodiments, GMSS may be intended to prevent a client's cached data from being completely flushed out due to bursts from other clients.) This can be seen, for example, in lines 11 through 13 of Algorithm 4.
[0184] 5. Partitioning methods for non-burst workloads
[0185] In some embodiments, for non-burst workloads, the partition manager system according to the disclosed example embodiments may have sufficient time to perform a more extensively optimized repartition that takes into account one or more factors, such as estimated hit rate, workload, weight per client, etc.
[0186] 5.1 Partitioned Workflow for Non-Burst Workloads
[0187] In some embodiments, a repartitioning workflow for non-burst workloads according to the disclosed example embodiments may adjust the partition size for each client (v) to maximize the objective function, which provides a way to evaluate the effectiveness of a given partitioning plan.
[0188] Equation 5 provides an embodiment of an optimization framework that can be used, for example, to repartition one or more levels of a storage layer (e.g., the top level that can operate as a storage cache). Table 5 provides some example meanings of the symbols used in Equation 5. Although the embodiments of the workflow shown in Equation 5 are not limited to any particular implementation details, in some embodiments, it can be used, for example, by... Figure 7 The optimal solution component 722 shown in the figure is implemented.
[0189] Table 5
[0190] maximum
[0191] St:Δ|C| v ∈[-|C| v , |C| max -|C| v ]
[0192] ExpHitRatio(v,|C| V +Δ|C| V Equation 5: ∈ [0%, 100%]
[0193]
[0194]
[0195] If the client has a cache size change Δ|C| in the next PMSW period v (For example, ExpHitRatio(v, |C|) v +Δ|C| v Multiplying this by its expected work volume (e.g., ExpWorkVol(v)), the framework of Equation 5 can be started by calculating a weighted sum of the expected hit rates for each client. As mentioned above in Section 4.1, the term "weight" (W) v This can reflect the QoS, SLA, etc. of each client. It can iterate over the term Δ|C| for a range of possible cache allocation scenarios. v In some embodiments, the granularity of the iteration steps can be set to a relatively coarse-grained amount (e.g., 100 MB), which can reduce the overhead associated with solving the optimization framework.
[0196] Then, the framework of Equation 5 can divide the weighted sum by the total cache size (|C|). max ).
[0197] In some embodiments, the physical meaning of this result can be represented as the expected number of hits (e.g., working volumes in bytes) that can be achieved in the next period based on the implementation of a particular partitioning plan, expressed in terms of cache capacity per byte. Therefore, maximizing the objective function of Equation 5 can improve or optimize the partitioning plan.
[0198] In some embodiments, the expected hit count can be Δ|C| for all clients. v The whole (i.e., Δ|C|) v (List).
[0199] Equation 5 may include a constraint portion (“St”) which may apply the following restrictions: (1) the size of the space per client in the top-level storage (e.g., cache size); (2) the expected hit rate range of [0%, 100%]; (3) all changes (Δ|C| v The sum of the weights of all clients is zero (because the amount of cache space does not change, it can only be redistributed among clients); and (4) the sum of the weights of all clients is 100%.
[0200] In some embodiments, the objective function in Equation 5 can be used to find an incremental list of all clients (e.g., changes in top-level storage) (where the list can be represented as {Δ|C| opt (v∈{V})} is used for the optimal solution. As shown in Equation 6, based on the list {Δ|C| opt (v∈{V})}, the partition manager system according to the disclosed example embodiment can return the final repartition plan for all clients.
[0201]
[0202] 5.2 Hit Rate Estimation
[0203] In some embodiments, in order to estimate the hit rate (ExpHitRatio(v, |C|)) v +Δ|C| v An analytical modeling framework based on reaccess distance can be used to provide online predictions of cache performance for a range of cache configurations and / or replacement strategies (e.g., Least Recently Used (LRU)).
[0204] Algorithm 5 illustrates an example embodiment of an analysis modeling framework based on re-access distance, according to a disclosed example embodiment. The main procedure, located in lines 1 through 6, calls the function CDF.update (where CDF may indicate a cumulative distribution function) to update the CDF curve (which can be implemented in lines 8 through 12). Based on the updated CDF curve, the hit rate can be estimated by inputting a given cache size into the CDF's getProbability function (which can be implemented in line 6). In some embodiments, this determines what percentage of records accessed may have a reuse distance smaller than the cache size.
[0205] The function `runtimeUpdateRecordTable` (which can be implemented in lines 14 through 37) can be a separate function that runs in the background and updates the record table when a new I / O request arrives. In Algorithm 5, `curEpochID` can represent the current epoch ID.
[0206] Algorithm 5
[0207]
[0208]
[0209] In Algorithm 5, the function `runtimeUpdateRecordTable(clientID)` runs in the background during runtime. It first converts incoming I / O requests (from `clientID`) into data buckets, which can have a granularity set to, for example, 5MB or 50MB for statistical analysis purposes. A counter named `curAccessTime` increments by 1 for each new data bucket.
[0210] Algorithm 5 can record the "distance" between the current access to an incoming I / O request data block and its last access. In some embodiments, the distance may represent "Absolute Revisit Distance (ARD)" instead of "Unique Revisit Distance (URD)".
[0211] This can be understood by referring to the revisit interval (RI). The address trajectory T can be a mapping of consecutive integers in ascending order, representing consecutive positions in the trajectory. A tuple (x, τ) can be used to represent all visits where (x, τ) ∈ T, and where x represents an address and τ indicates its repetition count. The first occurrence of address x in the trajectory can be represented by (x, 0). The expression t = T -1 It can represent an inverse function, and t(x, τ) can represent the position of the τth occurrence of address x in the trajectory.
[0212] As shown in Equation 7, RI can be defined only when τ>0, and can represent part of the closed trajectory between the τth occurrence and the (τ-1)th occurrence of x.
[0213]
[0214] Then, URD can be calculated as shown in Equation 8. RI can represent the total number of unique addresses between two occurrences of the same address in the trajectory.
[0215]
[0216] ARD can also be calculated as shown in Equation 9. ARD can represent the total number of positions between two occurrences of the same address in the trajectory.
[0217]
[0218] The function then searches for each new block in the recordTable, which can be implemented, for example, using a hash table, where the key can be the bucketID (block ID) and the value can be a tuple of multiple factors. If no entry is found, it indicates that this is the first occurrence of the data block, and a new entry can be created, with the value of that entry stored in the recordTable. In some embodiments, the current epochID can also be stored in the table, for example, to further facilitate the fading out of older data. The code for this operation can be found on lines 10 through 29 of Algorithm 5.
[0219] Table 6 shows some example values that can be recorded in the recordTable using the function runtimeUpdateRecordTable(clientID). The values provided in Table 6 are for illustrative purposes only, and other values may be obtained and / or used.
[0220] Table 6
[0221]
[0222]
[0223] In Algorithm 5, the function `estimateHitRatio` can be called to estimate the cache hit rate for a given cache size. The `estimateHitRatio` function triggers the function `CDF.update`, which updates the CDF record (e.g., ...). Figure 13 As shown in the figure, a new CDF curve is constructed by combining the maximum reuse distance and the percentage of bins.
[0224] Figure 13An example embodiment of a CDF curve for hit rate estimation operation is shown according to a disclosed example embodiment. Figure 13 The CDF curve shown can help estimate the hit rate for a given clientID and cacheSize. Figure 13 In this model, the X-axis can represent the maximum reuse distance, and the y-axis can represent the percentage of blocks with a maximum reaccess distance lower than a certain distance on the X-axis. Therefore, the hit rate estimation operation can iterate over different cache sizes (maximum reuse distance) to examine the upper limit of the hit rate it can provide for the workload during that period. In some embodiments, a cache hit is possible only if the cache size is greater than the maximum reuse distance. Therefore, a connection can be established between cache size and the theoretically optimal hit rate of the LRU cache algorithm. In some embodiments, the curve can represent the maximum cache hit rate. Therefore, a reduction adjustment can be assigned to it.
[0225] 5.3 Work Volume Estimation
[0226] In some embodiments, averaging techniques can be used to estimate the working volume ExpWorkVol(v). For example, some embodiments may use moving window averaging techniques (such as weighted moving average (WMA) and / or exponentially weighted moving average (EWMA)) to predict the working volume. In some embodiments, the weighted average may be implemented as an average with a multiplicative factor to assign different weights to the data at different locations in the sample window. Thus, mathematically, the weighted moving average may be a convolution of a reference point with a fixed weighting function. In some embodiments, the EWMA technique may involve weights decreasing in an arithmetic progression.
[0227] For illustrative purposes, some embodiments may be described in the context of specific values of the EWMA technique (which may employ a first-order infinite impulse response filter to apply an exponentially decreasing weighting factor) and parameters. However, other techniques and / or values may be obtained and / or used according to this disclosure.
[0228] Table 7 provides some example meanings of the symbols used in Equation 10. While the embodiments of the workflow shown in Equation 10 are not limited to any particular implementation details, in some embodiments, it can be, for example, through... Figure 7 The optimal solution component 722 shown in the figure is implemented.
[0229] Table 7
[0230]
[0231] In some embodiments, as shown in Equation 10, EWMA technology can be used to calculate (e.g., predict) the next period work volume of a client based on the client v's previous period work volume records.
[0232] |V| v,t+1 =α[|V| v,t +(1-α)|V| v,t-1 +(1-α) 2 |V| v,t-2 +(1-α) 3 |V| v,t-3 +…] Equation 10
[0233] Because Equation 10 can be an infinite sum with decreasing terms, the work volume estimation technique according to the disclosed example embodiments can limit the number of terms used in Equation 10 to reduce computational and / or storage overhead (e.g., storage for recording I / O transaction history). For example, the number of previous periods recorded can be limited to the value "k". That is, to approximate the prediction of the work volume, the kth term (e.g., k in Equation 11) can be omitted from Equation 10. max The term following the term. Therefore, equation 10 can be simplified to equation 11.
[0234]
[0235] In some embodiments, a small number of data points of the working volume size may be available during the initialization phase. Therefore, during the initialization phase, the working volume may be calculated, for example, based on the first few observations and / or the average of 4 to 5 PMSW periods.
[0236] 6. Partitioning and content updates
[0237] After the policy selector subsystem 718 determines the partitioning policy for a storage tier (e.g., the top tier that can operate as a storage cache), it can be passed to the partition operator subsystem 724 for implementation. In some embodiments, implementing the partitioning policy may involve eviction and / or fetch operations.
[0238] 6.1 Partitioning and Content Update Workflow
[0239] Figure 14 An example embodiment of a partitioning and content update workflow is shown, based on a disclosed example embodiment. Although Figure 14 The workflow embodiments shown are not limited to any particular implementation details, but in some embodiments, it may be, for example, achieved through... Figure 7 The partitioner component 726 shown is used to implement this. In some embodiments, the terms "cache size", "cache space", "new partition" and "quota" are used interchangeably.
[0240] Figure 14The workflow shown can begin at operation 1402, where the size of the client's new partition is compared to the size of its current partition. In some embodiments, the comparison may result in one of three cases: (Case 1) If the size of the client's new partition is smaller than its previous size, the workflow can proceed to operation 1404, where the workflow can remove one or more cache slots from the client and then shrink the partition to the new size. (Case 2) If the size of the client's new partition is equal to its previous size, the workflow can proceed to operation 1406, where it can bring the partition to its current size. (Case 3) If the size of the client's new partition is larger than its previous size, the workflow can proceed to operation 1408, where the workflow can increase the size of the partition. In case 3, the new space in the client's partition has two options: (Option 1) The workflow can proceed to operation 1410, which can be described as a passive option. In the passive option, the new space in the client's partition is initially left empty. Because new I / O requests can gradually fill the empty space, the new empty space can be gradually filled. (Option 2) The workflow may proceed to operation 1412, which can be described as an active option. The workflow may actively prefetch data from another storage tier (e.g., a second tier if the cache is the top tier) and copy the data to the first tier to fill empty spaces. For example, in some embodiments, the workflow may copy warm data that may not be cached but is frequently used based on I / O history to empty spaces in a new partition on the client. In some embodiments, the workflow may proceed to operation 1414 before filling all new empty spaces, where the prefetch size may be adaptively adjusted as described below.
[0241] After the repartition is complete, the client can continue to use its new quota until the next repartition operation occurs.
[0242] In some embodiments, the decision to use the passive option (Option 1) or the active option (Option 2) may involve one or more trade-offs weighted by system implementation details. For example, the passive option may be simpler to implement and / or involve less overhead because no data will be moved until the data becomes the subject of an I / O request. However, in the case of the passive option, the hit rate of clients with increased cache space may not be fully utilized because blanks may not be filled (e.g., until future I / O bursts fill the blanks). Additionally or alternatively, for example, some of the blank space may not be fully utilized if burst predictions are inaccurate and / or if the client's workload changes during the next period.
[0243] With proactive options, client hit rates can increase because more data becomes available in the cache. Furthermore, proactively filling empty spaces leads to better cache utilization. However, proactive options can incur higher overhead, for example, because the system can perform one or more operations to select (including any computations following the selection) and pass warm data that wasn't previously cached in the first tier. Additionally, there's a risk that some or all of the prefetched data might not be useful in the next period (e.g., not be re-accessed), resulting in further overhead.
[0244] Algorithm 6 illustrates some example operations of an embodiment of a partitioning and content update workflow according to a disclosed example embodiment. Algorithm 6 can be, for example, derived from... Figure 7 The partitioner component 726 shown in the diagram is executed. In the embodiment shown in Algorithm 6, the code for handling case 1 may be located in lines 4 to 6, the code for handling case 2 may be located in lines 7 to 10, and the code for handling case 3 (including options 1 and 2) may be located in lines 11 to 18. In Algorithm 6, curEpoch represents the current period.
[0245] Algorithm 6
[0246]
[0247]
[0248] 6.2 Adaptive Prefetch Size Adjustment
[0249] During a repartitioning operation, if the client's new partition is larger than its previous size and the partition manager system decides to proactively fill the empty space in the new partition (e.g., case 3, option 2 as described above), some embodiments of the partition manager system according to the disclosed example embodiments may provide an adaptive prefetch resizing (APSA) feature.
[0250] In some embodiments, the APSA feature may adjust the prefetch data size, for example, based on the top I / O-size popularity for each client in a recent period. For example, the APSA feature may adaptively and / or automatically change the prefetch data size. Prefetching techniques may improve performance, for example, by fetching data from a slower second-tier storage to a faster first-tier (cached tier) storage before the data is actually needed. In proactive operation options, data may be prefetched to fill empty spaces in newly allocated partitions before clients request data (e.g., by making I / O requests to previously empty cache locations).
[0251] In some embodiments, if most I / O from the client in recent times has been in the high I / O size (range) (e.g., 16KB), it is feasible or likely that the client will follow the same or similar pattern in the future. Therefore, if a partial cache hit occurs in the first tier in the future, pre-warming data from the second tier with a high I / O size granularity can help bypass another lookup in the second tier (which can be time-consuming). A partial cache hit may occur when only a portion of the requested I / O is found in the cache; therefore, a second-tier lookup can be performed to find the remainder of the requested I / O.
[0252] Therefore, by performing prefetching with high I / O size granularity, the partition manager system according to the disclosed example embodiments can help upgrade partial cache hits to full cache hits (e.g., all requested I / O can be cached in the first level).
[0253] Equation 12 illustrates an embodiment of an equation that can be used to implement adaptive prefetch size adjustment according to a disclosed example embodiment. Equation 12 can be used, for example, by... Figure 7 The partitioner component 726 shown in the figure is used for implementation. Table 8 provides some example meanings of the symbols used in Equation 12 and further explains the relationship between cache miss, partial cache hit, and full cache hit.
[0254] In some embodiments, CachedIOReq can be represented as the data found for a new I / O request (cached in the first level), and CachedIOReq can be calculated based on Equation 12 as follows.
[0255] CachedIOReq = {x | x∈IOReq, x∈CachedIO} Equation 12
[0256] Here, IOReq can represent a set of data in a new I / O request, and CachedIO can represent all data currently in the cache hierarchy. The union of the IOReq and CachedIO sets can be the data for that I / O found in the cache.
[0257] In some embodiments, for example, a preset prefetch amplification factor AF can be used to increase the flexibility of APSA features. As an example, the amplification factor can be in the range of [-100%, 100%]. In some embodiments, the prefetch granularity can be further adjusted (e.g., by setting AF to 5% so that it prefetches an additional 5% of data compared to a “high I / O size”). This can be useful, for example, in situations where storage device block and page alignment can be considered. For example, cross-block boundary I / O can trigger more data access than its I / O request size.
[0258] Table 8
[0259] Cache miss N / A |CachedIOReq|=0 Cache hit Partial cache hit |IOReq|>|CachedIOReq| Cache hit Full cache hit |IOReq|≤|CachedIOReq|
[0260] Algorithm 7 illustrates some example operations of an embodiment of the adaptive prefetch size adjustment feature according to the disclosed example embodiments. Algorithm 7 can be, for example, derived from... Figure 7 The partitioner component 726 shown in the diagram is executed. In the embodiment shown in Algorithm 7, the code for implementing the amplification factor AF may be located in lines 6 and 7.
[0261] Algorithm 7
[0262]
[0263] Figure 15 An example embodiment of an adaptive prefetch size adjustment method according to a disclosed example embodiment is shown. Although Figure 15 The embodiments shown are not limited to any particular implementation details, but in some embodiments, it may be, for example, derived from... Figure 7 The partitioner component 726 shown in the figure is used to implement this.
[0264] Figure 15 The method shown can begin at operation 1502, where the method obtains the prefetch amplification factor AF. At operation 1504, the method creates a results list that can be used to store the results of the adaptive prefetch size adjustment method. At operation 1506, the method determines the high I / O data size for each client through client iteration. At operation 1508, the results list is returned.
[0265] 6.3 Data Selection for Filling Empty Cache Space
[0266] During a repartitioning operation, if the client's new partition is larger than its previous size, and the partition manager system decides to proactively fill the empty space in the new partition (e.g., case 3, option 2 as described above), some embodiments of the partition manager system according to the disclosed example embodiments may select warm data to fill the empty space. In some embodiments, the warm data may not be cached, but may be used frequently based on I / O history.
[0267] Algorithm 8 illustrates some example operations of an embodiment of the warm data selection method according to the disclosed example embodiments. Algorithm 8 can be, for example, derived from... Figure 7 The partitioner component 726 shown in the figure is executed.
[0268] Algorithm 8
[0269]
[0270] In the embodiment shown in Algorithm 8, lines 4 through 6 create a candidate list to store all descendingly sorted Global History Records (GHRs) in the second level, which can be sorted, for example, by the access frequency of each data block. Any block size can be used (e.g., 10 MB). In some embodiments, one or more sliding window techniques (e.g., EWMA) can be applied to the global history to record I / O popularity statistics for bins that have been accessed in the most recent period. Depending on the implementation details, this can more accurately reflect warm data access behavior.
[0271] As shown in line 6, by ignoring already cached data (cachedIOList), the method can further iterate each block from hottest to coldest, and (in some embodiments, considering the amplified granularity calculated in Algorithm 7) add data to the padding of the dataToPrefetch (data used for prefetching) list until the empty space has been filled. If any empty space still exists, and the iterated data block is larger than the empty space, the data block can be pruned to fill the empty space to prevent any wasted empty space.
[0272] As shown in line 14, once the empty space is filled, `dataToPrefetch` can be sent to the prefetch function to perform the actual content update from the second level to the top level (cache level). After the content update is complete, the top-level storage can run as a cache using any caching algorithm, such as LRU, Application-Aware Cache Replacement (ACR), Clock with Adaptive Replacement (CAR), etc. In some embodiments, write-through and / or write-back strategies may be supported.
[0273] Figure 16 An example embodiment of an adaptive prefetch size adjustment method according to a disclosed example embodiment is shown. Although Figure 16 The embodiments shown are not limited to any particular implementation details, but in some embodiments, it may be, for example, derived from... Figure 7 The partitioner component 726 shown in the figure is used to implement this.
[0274] Figure 16The method shown can begin at operation 1602, where a candidate list can be created. At operation 1604, the candidate list can be sorted, for example, based on frequency of use. At operation 1606, cached data can be removed from the candidate list. At operation 1608, a loop can begin for each candidate block. As long as the size of the data used for prefetching is less than the blank size, operations 1610, 1612, and 1614 can add data blocks based on the selected granularity level, prune data that may be larger than the blank, and add the results to the list of data to be added. At operation 1608, once the size of the data used for prefetching is greater than or equal to the blank size, the method can proceed to operation 1616, which can interrupt the operation, and then proceed to operation 1618, where the prefetched data is actually written to the cache. The method can terminate at operation 1620.
[0275] 7. Zoning based on prior knowledge
[0276] In some embodiments, during the deployment phase (also referred to as the early or initial phase), the partition manager system according to the disclosed example embodiments may not have sufficient information to assign one or more tiers of storage (e.g., the top tier that can operate as a storage cache) to clients. In these cases, the partition manager system can use prior knowledge of the client's I / O patterns, such as from a client pattern library, vendor-selected hardware and / or software, SLAs, QoS, etc., to create a partitioning plan.
[0277] For example, in a game streaming data center application, there may be prior knowledge of the characteristics of one or more hosted games (such as a first-person shooter game that may involve frequent writes of little random data or a racing game that may involve frequent, large sequential read I / O). As another example, in a general data center, there may be prior knowledge of the workload (such as developers working on virtual machines that may be write-intensive, and 3D video editing on virtual machines that may be CPU-intensive and may have read / write I / O mixed with high working volume size and working set size).
[0278] Such prior knowledge can be used to implement one or more partitioning techniques to separate workloads based on multiple factors, such as read ratios for each workload according to prior knowledge, working set size, etc. Partitions can be physical, virtual, or any combination thereof. In some embodiments, the granularity of physical partitions can be greater than or equal to the size of the storage device (e.g., grouping multiple storage devices into a region). In some embodiments, the granularity of virtual partitions can be less than or equal to the size of the storage device (e.g., partitioning within a storage device).
[0279] Figure 17An example embodiment of a prior knowledge-based zoning method according to a disclosed example embodiment is shown. Although Figure 17 The embodiments shown are not limited to any particular implementation details, but in some embodiments, it may be, for example, derived from... Figure 7 The partitioner component 726 shown in the figure is used to implement this.
[0280] exist Figure 17 In the embodiments shown, as indicated in Table 1702, prior knowledge of the first four clients (Client 1, Client 2, Client 3, and Client 4) may include average read ratio, average working set size, and average working volume size. In some embodiments, the average may be based on a simple historical average. Optionally or additionally, other than […], can be used in the deployment partitioning process. Figure 17 Factors other than those listed in the document.
[0281] exist Figure 17 In the embodiment shown, the top level of storage 1704 may be implemented using an SSD (e.g., SSD01 to SSD18), and the second level of storage 1706 may be implemented using an HDD (e.g., HDD01 to HDD22), but any other number and / or configuration of tiers and / or storage devices may be used.
[0282] In this example, the top level of storage 1704 can have three regions: a read-intensive region for clients with workloads having read ratios in the range [90%, 100%], a read-write mixed region for clients with workloads having read ratios in the range [10%, 90%), and a write-intensive region for clients with workloads having read ratios in the range [0%, 10%). These examples are for illustrative purposes only, and other numbers of regions, other ranges of values, etc., can be used to create regions.
[0283] exist Figure 17 In the embodiments shown, the partition manager system can proportionally adjust the size of each region based on, for example, the working set size for reading dense regions and / or (2) the working volume size for reading / writing mixed regions and / or writing dense regions. In some embodiments, applications such as Figure 17 The proportional area adjustment shown in the figure updates the area size.
[0284] In some embodiments, after deployment is complete, the partition manager system according to the disclosed example embodiments can operate independently in each region using any of the techniques disclosed herein. For example, Figure 7 The entire loop shown can run in each region.
[0285] Algorithm 9 illustrates some example operations of an embodiment of a prior knowledge-based zoning method according to a disclosed example embodiment. Algorithm 9 can be, for example, derived from... Figure 7 The partitioner component 726 shown in the figure is executed.
[0286] Algorithm 9
[0287]
[0288] In the embodiment shown in Algorithm 9, the function ZoneDeployment can take two inputs: preknowledgeTable and topTierStorage, and can use preset values such as zoneReadRatioThresholds (which can be any number of zones) as shown in line 2. In line 3, the function creates a zoneMemberPlan by comparing the read ratio of each client in the preknowledgeTable with the threshold. In lines 4 through 6, it then iterates over the zoneMemberPlan and calculates the demand for each zone based on the working set size of the member clients (for read-intensive zones) or the working volume size (for other zones). In line 7, the function then proportionally adjusts the zone size based on the demand for each zone (e.g., zoneAmount). In line 8, the function assigns clients to zones based on both the zoneMemberPlan and ZoneSizePlan.
[0289] In some embodiments, and depending on the implementation details, the prior knowledge-based partitioning method according to the disclosed example embodiments may: (1) reduce device-side write amplification; (2) reduce over-provisioning (which may be similar to early replacement, since in both cases the host may consume more device); (3) reduce memory in the storage device (which may be relatively expensive); (4) improve latency, throughput and / or predictability; and / or (5) enable the software ecosystem, as multiple stakeholders may benefit from one or more improvements.
[0290] Figure 18 An example embodiment of a prior knowledge-based zoning workflow, according to a disclosed example embodiment, is shown. Although Figure 18 The embodiments shown are not limited to any particular implementation details, but in some embodiments, it may be, for example, derived from... Figure 7 The partitioner component 726 shown in the figure is used to implement this.
[0291] Figure 18The embodiment shown can begin at operation 1802, where a regional membership plan is created by comparing the read ratio of each client in a pre-defined knowledge table with a regional read ratio threshold. At operation 1804, the workflow iterates over the regional membership plan and calculates the demand for each region based on either the working set size of the member clients (for read-intensive regions) or the working volume size (for other regions). At operation 1806, the workflow proportionally adjusts the region size based on the demand for each region. At operation 1808, the workflow assigns clients to regions based on the regional membership plan and the region size plan.
[0292] Figure 19 An embodiment of a method for operating a storage system according to a disclosed example embodiment is shown. At operation 1902, the method may begin. At operation 1904, the method may allocate a first partition of a tier of storage resources to a first client, wherein the tier operates at least partially as a storage cache. At operation 1906, the method may allocate a second partition of the same tier of storage resources to a second client. At operation 1908, the method may monitor the workload of the first client. At operation 1910, the method may monitor the workload of the second client. At operation 1912, the method may reallocate the first partition of the same tier of storage resources to the first client based on the monitored workload of the first client and the monitored workload of the second client. At operation 1914, the method may end.
[0293] The embodiments disclosed above have been described in the context of various implementation details, but the principles of this disclosure are not limited to these or any other specific details. For example, some functions have been described as being implemented by specific components, but in other embodiments, functions may be distributed across different systems and components in different locations and have various user interfaces. Some embodiments have been described as having specific processes, operations, etc., but these terms also include embodiments in which specific processes, steps, etc. may be implemented by multiple processes, operations, etc., or include embodiments in which multiple processes, operations, etc. may be integrated into a single process, step, etc. References to components or elements may refer to only a portion of a component or element. For example, a reference to an integrated circuit may refer to all or only a portion of an integrated circuit, and a reference to a block may refer to an entire block or one or more sub-blocks. Unless otherwise clear from the context, the use of terms such as “first” and “second” in this disclosure and claims is for the purpose of distinguishing what they modify and may not indicate any spatial or temporal order. In some embodiments, “based on” may mean “at least partially based on”. In some embodiments, “allocated” may mean “at least partially allocated”. A reference to a first element may not imply the presence of a second element. For convenience, various organizational aids such as chapter titles may be provided, but the subject matter arranged according to these aids and the principles of this disclosure are not limited or restricted by these organizational aids.
[0294] The principles disclosed herein are independently practical and can be implemented individually, and not every principle can be utilized in every embodiment. However, these principles can also be implemented in various combinations, some of which can synergistically amplify the benefits of individual principles. Therefore, the various details and embodiments described above can be combined to produce additional embodiments based on the inventive principles disclosed in this patent. Since modifications to the arrangement and details of the inventive principles disclosed in this patent can be made without departing from the inventive concept, such changes and modifications are considered to fall within the scope of the appended claims.
Claims
1. A method for operating a storage system, the method comprising: The first partition of the storage resource hierarchy is allocated to the first client, wherein the hierarchy is operated at least in part as a storage cache; Allocate the second partition of the storage resource level to the second client; Monitor the workload of the first client; Monitor the workload of the second client; and Based on the workload of the first client and the workload of the second client, the first partition of the storage resource level is reallocated to the first client. The method further includes: Determine the burstiness of the workload of the first client and the workload of the second client; The burstiness level is used to detect whether access to the storage resource at that level is bursty. The redistribution steps include: If the burstiness exceeds a preset threshold, the workload is determined to be a burst workload, and the first and second partitions are proportionally redistributed to the first and second clients based on the workload of the first client in the most recent period and the workload of the second client in the most recent period.
2. The method according to claim 1, wherein, The first partition is reassigned based on the input and / or output requirements of the workload of the first client.
3. The method according to claim 1, wherein, The first partition was reassigned based on the estimated performance changes of the workload of the first client.
4. The method according to claim 1, wherein, The read / write ratio of the first partition is redistributed based on the workload of the first client.
5. The method according to claim 1, wherein, The working set of the first partition based on the workload of the first client was reallocated.
6. The method according to claim 1, wherein, The first partition's working volume based on the first client's workload was reallocated.
7. The method according to claim 1, wherein, The steps to determine the degree of suddenness include: Determine changes to the workload's work volumes and calculate burstiness based on these changes. Determine the changes in the working set of the workload, and calculate the burstiness based on the changes in the working set of the workload.
8. The method according to claim 1, further comprising: The first client's first cache requirement is determined based on the workload of the first client. as well as The second client's second cache requirement is determined based on the second client's workload. The first partition is allocated to the first client based on the first cache requirement, and The second partition was allocated to the second client based on the second cache requirement.
9. The method according to claim 1, further comprising: Record the input and / or output transactions of the first client's workload; Reuse distance is determined based on the recorded input and / or output transactions; as well as The expected cache hit rate is determined based on the reuse distance.
10. The method according to claim 1, Record the input and / or output transactions of the first client's workload; Determine the weighted average of the input and / or output transactions for the records; and The expected working volume is determined based on a weighted average.
11. The method according to claim 1, wherein, The step of reallocating the first partition includes increasing the size of the first partition, and the method also includes updating the first partition based on one or more input and / or output transactions based on the workload of the first client.
12. The method according to claim 1, wherein, The step of reallocating the first partition includes increasing the size of the first partition, and the method further includes: Determine the pattern for the size of the first client's input and / or output requests; and Use a prefetch size based on the input and / or output request size pattern to prefetch data for the first partition.
13. The method according to any one of claims 1 to 12, wherein, The first partition is assigned to the first client based on prior knowledge of the characteristics of the first client.
14. A method for hierarchically partitioning storage resources, the method comprising: Determine the read and write working volumes of the first client and the second client of the level; The workload type is determined based on reading and writing to the work volume; as well as The storage resources are partitioned at the tiers between the first client and the second client based on workload type. The method further includes: Determine the burstiness of the workload of the first client and the workload of the second client; The burstiness level is used to detect whether access to the storage resource at that level is bursty. The partitioning process includes: Based on the burstiness exceeding a preset threshold, the workload is determined to be a burst workload, and the storage resources are partitioned proportionally between the first client and the second client based on the workload of the first client in the most recent period and the workload of the second client in the most recent period.
15. The method of claim 14, further comprising: The read ratio is determined based on reading and writing to the working volume; as well as The workload type is determined based on the read ratio.
16. The method according to claim 14 or 15, further comprising: Determine the working set size based on the workload type; as well as The storage resources are partitioned between the first and second clients based on the working set size.
17. The method according to claim 14 or 15, further comprising: Determine the work volume size based on the workload type; as well as The storage resources are partitioned between the first and second clients based on the working volume size.
18. A method for hierarchically partitioning storage resources, the method comprising: Determine the first partitioning plan for the first client and the second client at the storage resource tier; The first expected cache hit is determined based on the first partition plan; Determine the second partitioning plan for the first and second clients of the storage resource tier; The second expected cache hit is determined based on the second partitioning plan; as well as The first partitioning plan and the second partitioning plan are selected based on the first expected cache hit and the second expected cache hit. The method further includes: Determine the burstiness of the workload of the first client and the workload of the second client; The burstiness level is used to detect whether access to the storage resource at that level is bursty. The selection process includes: Based on the burstiness not exceeding a preset threshold, the workload is determined to be a non-burst workload, and a partition plan corresponding to the maximum value between the first expected cache hit and the second expected cache hit is selected from the first partition plan and the second partition plan.
19. The method of claim 18, further comprising: The first expected hit rate for the first client is determined based on the first partition plan; The second expected hit rate for the second client is determined based on the first partition plan; The first intended working volume for the first client is determined based on the first partition plan; as well as The second expected working volume for the second client is determined based on the first partition plan.
20. The method according to claim 19, wherein, The steps to determine the first expected cache hit include: determining the first expected hit rate of the first client and the weighted sum of the first expected working volume and the second expected hit rate and the second expected working volume of the second client.
Citation Information
Patent Citations
Region based admission / eviction control in hybrid aggregates
US9354989B1