A performance guarantee system and method
The node-based performance guarantee compliance system addresses the challenge of competing traffic in cloud storage by using an orchestrator to allocate fixed logical bandwidth for storage and application operations, maintaining reliable network performance through separate traffic pathways.
Patent Information
- Application Number
- PCT/IL2024/051234
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-31
- Filing Date
- 2024-12-29
- Publication Date
- 2025-07-03
AI Technical Summary
In cloud storage environments, competing storage and application traffic over a single network interface card (NIC) can lead to unpredictable performance, disrupting the reliability and predictability of network operations, especially during varying workloads.
A node-based performance guarantee compliance system that includes an orchestrator to manage network traffic allocation between storage and application operations, using profiling procedures to analyze bandwidth utilization and allocate fixed logical bandwidth shares, ensuring separate logical pathways for each type of traffic over a shared NIC.
This system maintains consistent and reliable network performance by segregating storage and application traffic, preventing disruptions and ensuring compliance with predefined performance guarantees, even under fluctuating workloads.
Smart Images

Figure IL2024051234_03072025_PF_FP_ABST
Abstract
Description
[0001] A PERFORMANCE GUARANTEE SYSTEM AND METHOD
[0002] FIELD OF THE INVENTION
[0003] The present invention in general relates to orchestration procedures for storage systems, and in particular to the implementation of performance guarantees.
[0004] BACKGROUND OF THE INVENTION
[0005] Providing data storage and data backup capabilities represent a significant concern as current computing systems, whether local, remote or cloud based (such as containers packages, private / public / multi-cloud systems, etc.), require ever more extensive data storage solutions for their proper operation. Usually, such data provision and management are made and offered by designated data centers and traditionally the provision of used or expected to be used data storage is provided by stacking physical data storing components, i.e. hybrid hard disk drive (HHD), hard disk drive (HDD), solid-state drive (SSD), etc. Because the methods by which data is stored and edited on different types of drives are so distinct, a similarly broad variety of network configurations and operating methods have emerged to meet the requirements of different network applications.
[0006] Many of these systems and methods include technical features which - whilst distinct - can serve similar functions in the very specific context in which they are disposed, albeit not functions that are independent of said context. Technical features relating to the storage, transfer, sensing, and management of data are employed in a variety of approaches, systems, methods, and network configurations, which have been developed to address a range of technical problems relating to data storage and network management broadly. Much of these are discussed below, in order to provide a broad overview of the relevant prior for the present invention. An approach well established in the field of data storage is the operation of stacking data storing components to create what is termed “Storage Arrays” (or alternatively “disk arrays”) which are used for different kinds of data, broadly categorized by: block -based storage; file-based storage; object storage, among other data types. Rather than store data on a server, storage arrays use multiple drives in a collection capable of storing a huge amount of data, controlled by a local / central controlling system interfacing via storage network protocols to the server.
[0007] Traditionally, a storage array controlling system provides multiple storage services so as to keep track of storage capacity; the allocation of space to different datasets, the management of sections of data storage capacity known as “volumes the periodic backup operations of the data to facilitate restoration and disaster recovery and the creation of pointin-time copy of the data, known “snapshotting” the identification and tracking of errors; the encryption of data communication to protect the integrity and privacy of data; the compression of data to conserve storage capacity; etc. Services of such type require significant computing capacity, metadata, data storage, accelerators, etc. - thus, such services require the designation of extensive infrastructure and budget capacities and resources.
[0008] Commonly, a storage array is separated from a system server's operability and is configured to implement system and application operations on dedicated hardware, for example a server stack, a storage array stack, one or more hard disk or solid state drive (HDD or SSD) and media input / output (I / O) devices configured to communicate with the servers via the storage stack.
[0009] Another approach well established in the field is the employment of an orchestrator, which is a software module logically situated in the control plane (CP) of a distributed network and is responsible for managing the operations of the data plane (DP), such as provisioning and resource coordination. The DP is the layer within a network architecture responsible for the movement, processing and storage of data through the distributed network, whilst the CP is an associated network layer that responsible for controlling how said data flows through the DP. Positioned in the CP, the orchestrator provides centralized management of the DP network, automating data flow across network nodes in accordance with predefined rules. Such coordination is particularly important in distributed network environments where resources such as computational power, storage and network bandwidth are spread across multiple nodes, often in different physical locations. Another approach well established in the field of data storage is the operation of redundant arrays of independent disks (RAID), which can be operated as a way of storing the same data in different places to protect data in the case of a system failure.
[0010] RAID is a general approach and network configuration that virtualizes data and combines multiple physical disk drive components into one or more logical units. Persons skilled in the art will appreciate that the technical problem RAID operations are employed to address depend on the type of RAID operation undertaken: RAID 0 stripes data across multiple disks to address performance bottlenecks and capacity limitations; RAID 1 mirrors data across two or more disks to address data loss due to disk failure; RAID 2 stripes data at the bit level and uses Hamming code for error correction to address data errors and fault tolerance in high- reliability systems; RAID 3 stripes data at the byte level and uses a dedicated parity disk to address single-disk failure and sequential data access bottlenecks; RAID 4 stripes data at the block level with a dedicated parity disk to address single-disk failure and block-level performance bottlenecks; RAID 5 stripes data and distributes parity information across multiple disks to address single-disk failure and storage efficiency; RAID 6 stripes data with double distributed parity to address multiple disk failures and ensure data integrity; RAID 10 combines mirroring (RAID 1) and striping (RAID 0) to address performance bottlenecks and single-disk failure; RAID 01 mirrors two RAID 0 arrays to address performance bottlenecks and fault tolerance; RAID 50 combines RAID 5 arrays and stripes them using RAID 0 to address the performance and reliability limits of RAID 5; RAID 60 combines RAID 6 arrays and stripes them using RAID 0 to address the performance and redundancy limits of RAID 6; RAID 7 uses an embedded real-time OS and dedicated cache to improve performance and address bottlenecks associated with traditional RAID levels; RAID IE stripes mirrored data across an odd number of disks to address fault tolerance and performance in setups where an odd number of disks are available.
[0011] Another approach well established in the field of data storage is the operation of remote replication, which is the process of copying data to a device at a remote location for data protection or disaster recovery purposes. Remote replication may be either synchronous or asynchronous, the former writes data to the primary and secondary sites at the same time, and the latter at different times. Because asynchronous replication is designed to work over longer distances and requires less bandwidth, it is often considered a better option in the field for the recovery of data after a catastrophic disaster. However, the operation of asynchronous replication also introduces several risks, not least the risk of loss of data during a system outage as said data at the target device isn't synchronized with the source data. Most enterprises today use data storage vendors that include replication software on their high-end and mid-range storage arrays, to partially mitigate this risk.
[0012] Another configuration well established in the field of data storage is software-defined storage (SDS), which enables communality of operation of different hardware. SDS configurations include the abstraction of data storage resources from the underlying physical storage hardware, and thereby are able to provide flexible exploitation of available hardware and data storage resources. Typically commercial off-the-shelf servers run a subset of SDS known as hyper-converged infrastructure HCI, in which the abstractions of both the area network and the underlying storage are implemented virtually in software, rather than physically in hardware.
[0013] Both conventional storage arrays and SDS configurations typically include an integrated “storage stack” - a layered software framework that organizes, manages and facilitates data storage, access and retrieval. Said storage stack provides essential services such as data protection (e.g. backup, redundancy, recovery, etc.); space allocation; data optimization, backup and recovery, among other functions. Due to the broad array of functions required by SDSs, the integrated software stack is typically configured to have a high of reliability, and the efficiency of the code is also conventionally prioritized.
[0014] Another data storage configuration taught in the field is directed attached storage (DAS), which typically provides the direct local services (such as encryption, compression, RAID, etc.) in cases where central storage systems are not needed or desired. Conventionally, DAS configurations will exploit a robust collection of internal storage components, without which the means of operating said services would be insufficient for proper network function. Persons skilled in the art will appreciate that the technical problem DAS network configurations are employed to address is: the provision of data storage services in the absence of centralized data management nodes. DAS is mostly limited to non-critical applications due to an inherent drawback related to the fact that DAS is inherently tied to one host: server communication failure precludes data accessibility, typically limiting DAS to non-critical applications. This is in contrast to the SDS solutions previously described, which are typically accessible by multiple servers over the network; if one server or communication channel fails, other servers can still access the storage. Another approach well established in the field of data storage is the operation of hot spares. Traditionally, hot spares act as standby drives in RAID 1, RAID 5, or RAID 6 volume groups, but they have also been applied to other network management approaches. Generally, if a drive fails, for example in a volume group, some control software will reconstruct data from the failed drive on a hot spare. When a drive fails in a storage array, a hot spare drive can be substituted without requiring a physical swap. Persons skilled in the art will appreciate that the technical problem hot spare configurations are employed to address is: minimizing downtime and ensuring quick recovery from disk failures in RAID and other storage systems. Another approach well established in the field of data storage is the operation of snapshots of data. A snapshot is used to represent the content of a particular part of a data stored on a storage system at a particular point in time.. The source of snapshots are typically base volumes, which are usually referred to as “member volumes” of a “consistency group”. The purpose of a consistency group is to facilitate the capture of simultaneous snapshot images of multiple volumes, thus obtaining copies of a collection of volumes at a particular point in time. In practice, most mid-range and high-end storage arrays create snapshot consistency groups within volumes inside the storage array. Persons skilled in the art will appreciate that the technical problems snapshot operations are employed to address are: loss prevention; data recovery; control of database version; rule compliance and auditing; monitoring of storage dynamics, among other technical problems.
[0015] Obtaining a local snapshot is enabled by a server operating system that includes a logical volume manger (LVM) - a software layer that abstracts physical storage disks into virtualized storage units (logical volumes) - enabling the obtaining of a local snapshot on a single virtualized volume. In distributed storage system, since the volumes are distributed across multiple servers, obtaining or creating a consistency group is not usually possible or supported, producing a number of data integrity risks. LVM works by partitioning the physical volumes (PVs) into physical extents (PEs), which are mapped onto logical extents (LEs) which are then pooled into volume groups (VGs), linked together as logical volumes (LVs). Persons skilled in the art will appreciate that the LVM approach is typically undertaken in order to address the technical problems posed by: inflexible partition sizes; fragmentation of disk space; limited scalability of storage infrastructure; complex mirroring and striping setups; difficulty in taking snapshots; and the efficient management of multi -disk systems.
[0016] Another approach well established in the field of data storage is quality of service (QoS), which is critical to deliver consistent storage performance applications where multiple workloads share a single limited resource by preventing the “noisiest neighbor” from disrupting the performance other applications on the same system. On physical storage arrays, QoS can be set for volumes as limits on data transfer. Unlike storage arrays, the distributed servers of storage stacks mean there isn’t a single point that can enforce QoS. Persons skilled in the art will appreciate that QoS is a general approach in data storage array management, which can be disposed to address a number of different challenges, including but not limited to: predictable performance in shared resources; performance spikes caused by noisy neighbors; difficulty maintaining SLA compliance; resource contention during peak loads; and the need for overprovisioning to avoid performance issues.
[0017] Another approach well established in the field of data storage is disk cloning, which is the process of making a copy of a part (or all) of a hard drive, typically undertaken at a particular point in time whilst hosts continue to access the data. Like QoS, this is an approach which is difficult to operate on shared storage stacks, since the source and target may reside on different physical entities. Persons skilled in the art will appreciate that disk cloning is typically undertaken in order to address the technical problems of: efficient data migration; disaster recovery; consistent system deployment; backup integrity, and the prevention of data loss due to hardware failure.
[0018] Another approach well established in the field of data storage is thick provisioning , where the complete amount of virtual disk storage capacity is pre-allocated on the physical storage when the virtual disk is created, rendering capacity unavailable for use by other volume. Persons skilled in the art will appreciate that the thick provisioning approach is typically undertaken in order to address the technical problems posed by: unpredictable availability of storage capacity; overcommitted storage resources; storage fragmentation, performance degradation, the risks of complex storage management; and the resultant shortages in capacity from said technical problems leading to data loss.
[0019] In contrast to thick provisioning, yet another approach well established in the field of data storage is thin provisioning, where a virtual disk consumes only the space that it needs initially, and grows with time according to increase in demand. Whilst thinly provisioned storage consumes less disk space, it consumes significantly more RAM to store the metadata of the thin allocation. Additionally, thin provisioning consumes much more CPU on the I / O transmissions needed to facilitate intensive random access to translate logical addresses to physical, since it has to navigate through a tree-like data structure. Despite these limitations, thin provisioning is a widely undertaken approach to address a number of different technical problems of the field, persons skilled in the art will appreciate that said technical problems include but are not limited to: the inefficient utilization of storage; high upfront capital costs; difficulty in scaling storage; and the over-allocation of resources.
[0020] Another approach well established in the field of data storage is the Clustered Logical Volume Manager (CLVM), which is a set of clustering extensions to LVM, an approach discussed earlier. These extensions allow a cluster of computers to manage shared storage using LVM by locking access to physical storage while a logical volume is being configured. A single misbehaving node can impact the health of the entire cluster, introducing significant risk for the integrity of data stored on a data storage system. Persons skilled in the art will appreciate that the technical problems the CLVM approach is disposed to address include but are not limited to: uncoordinated access to shared storage introducing storage performance limitations; corruption of stored data from multiple read / write operations; limitations to the scalability of storage environments; low storage availability; inefficient data sharing; and the risks of high complexity in the management and expansion of shared storage.
[0021] Another approach well established in the field of data storage is the deployment of a hardware security module (HSM), which is a physical device that manages digital keys for strong authentication. Persons skilled in the art will appreciate that HSMs are typically deployed in order to address the technical problems posed by: secure key generation and storage; tamper detection and resistance; performance bottlenecks for cryptographic operations; regulatory compliance; controlled key access; secure cryptographic operations; auditing; and logging.
[0022] Another approach well established in the field of data storage is the use of tunneling protocols, which are a communications protocols that allow for the movement of private data from one network to another across a public network, using a process called encapsulation. Persons skilled in the art will appreciate that tunneling protocols are typically operated in order to address the technical problems posed by: secure transmission of data over untrusted networks; bypassing network restrictions and firewalls; ensuring confidentiality and integrity of data in transit; preventing eavesdropping and man-in-the-middle attacks; encapsulating incompatible or sensitive protocols; and reducing exposure to external threats. Other approaches have been taught in the art to address the challenges of secure communication in distributed storage environments, including virtual private networks (VPNs)_; reverse proxies; agent-based models; and secure APIs. VPNs and encrypted tunnels create secure connections, but add latency and require extensive setup. Reverse proxies and API gateways offer controlled access to storage servers by routing external requests through a single entry point, but they also add routing layers that create bottlenecks and increase complexity. Agent-based models, which rely on modules within a network to pull commands from the control software rather than receive them directly, help bypass firewall restrictions but delay orchestration by requiring periodic updates instead of real-time communication. Secure APIs, which rely on authentication protocols, provide direct access to storage resources but can be challenging to scale across large networks due to resource demands.
[0023] Similar to the challenges of security, many approaches have been taught in the art to address chattiness, which is when communication between servers consists of repetitive, non- essential notifications that create unnecessary traffic. This challenge is typically addressed using: traffic filtering; message batching; and rate limiting, which selectively blocks non- essential communications; aggregates multiple smaller messages into fewer transmissions; and restricting the volume of messages over a defined interval, respectively.
[0024] Not unlike solutions to chattiness, many systems and methods have been taught in the art to address the challenge of identification of servers within node-based cloud storage networks, particularly in multi-tenant environments. Conventional means for server identification typically rely on: IP address verification; hostname recognition; and basic authentication protocols such as API keys or token-based systems
[0025] Storage systems may be implemented as on-premises data centers, wherein servers and infrastructure are privately owned and managed, or as networked storage environments, such as those offered via cloud computing service providers. Cloud storage systems may exploit shared resources both for the storage media and for the network infrastructure, which serves to connect the storage system to other systems and clients. In some configurations, storage systems utilize shared networks for general operations, while in others, dedicated networks may be required for each function in order to optimize performance and manage system resources more efficiently. Persons skilled in the art would appreciate that the choice of whether to employ shared or dedicated networks may depend on various technical factors, including but not limited to: workload types, data throughput requirements, latency considerations, as well as scaling requirements and others.
[0026] Node systems are critical components within a networked environment, acting as intermediaries that facilitate communication and data exchange across all components in the network. In the context of a data communication network, a node refers to a distinct component capable of transmitting, receiving, or routing data. In addition to their fundamental role in data handling, node systems often provide quality of service (QoS) monitoring capabilities, ensuring that data flow across the network meets predefined performance metrics. A particular implementation of node systems is cloud storage, which delivers scalable and flexible storage solutions for individuals and organizations. Cloud storage nodes operate in distributed environments and are thus capable of dynamically adjusting to accommodate fluctuating storage demands. Cloud storage systems face several technical challenges which impact performance and reliability.
[0027] Service Level Agreements (SLAs) are contractual commitments established between cloud storage providers and users, defining specific performance metrics such as minimum bandwidth, latency, uptime, or IOPS, along with remedies or penalties for non-compliance. Performance guarantees form the technical foundation that enables providers to deliver the consistent and reliable service levels outlined in SLAs, even in complex environments like multi-tenant systems where resource contention and variable workloads can pose significant challenges.
[0028] A predetermined performance guarantee in cloud storage systems refers to the technical assurance provided by the system to deliver consistent and reliable service levels, such as defined bandwidth, latency, or data throughput. These guarantees enable cloud storage providers to ensure their systems perform reliably, enabling users to select storage solutions tailored to their workloads and optimize configurations based on guaranteed thresholds. By providing predictable performance, these guarantees enhance cost efficiency by helping users allocate resources effectively and avoid over-provisioning.
[0029] In addition to operational reliability and cost management, performance guarantees facilitate effective scalability planning by allowing users to anticipate future storage needs with greater accuracy based on consistent performance levels. They also support the implementation of storage policies addressing redundancy, durability, and access control, ensuring compliance with technical and regulatory requirements. Furthermore, performance guarantees are essential for maintaining a consistent quality of service, particularly in multi -tenant environments, where competing workloads could otherwise compromise system performance.
[0030] Performance guarantees ensure that critical workloads are not disrupted by fluctuations in resource availability or competing traffic. For example, a performance guarantee might ensure that a database application consistently receives sufficient network bandwidth to handle user queries, even when storage traffic is high. In multi-tenant cloud environments, where multiple users share the same infrastructure, performance guarantees become particularly important to mitigate conflicting resource demands and ensure adequate access to shared resources. Without such guarantees, tenants face uncertainty about whether their applications and storage operations will function predictably and reliably, especially during periods of peak demand or under mixed workload conditions.
[0031] A critical technical challenge in node-based cloud storage solutions pertains to ensuring a network performance guarantee. In on-premises environments, multiple network interface cards (NICs) are available which can be dedicated to different tasks, such as separating storage network traffic from application network traffic, thereby facilitating improved performance. In cloud environments all traffic, both storage and application, run through a single NIC. This can lead to issues wherein one type of traffic overwhelms the network and thus affects the performance of other traffic types operating via the same physical network.
[0032] Without proper management, application traffic can consume a disproportionate amount of network bandwidth, leaving insufficient bandwidth for storage operations, such as reading or writing data from disks. Conversely, storage traffic can overwhelm the network and affect the performance of applications. For example, in a database system, an application must perform two tasks: (a) handling user queries, wherein a user sends requests over the network to query data from the database, (b) access storage, wherein the database needs to read or write data to storage systems to fulfil said queries. The system must process these queries and return results over the network. Since both tasks share the same network connection, the performance of one can impact the other. If there are too many user queries, the network may not have enough bandwidth left for efficient storage access, or vice versa.
[0033] There is thus a need in the art for a system that effectively manages network bandwidth in cloud environments. Such a system should address the challenges posed by competing storage and application traffic needs, ensuring that network performance remains reliable and predictable even under varying workloads. SUMMARY OF THE INVENTION
[0034] The following embodiments and aspects thereof are described and illustrated in conjunction with systems, devices and methods which are meant to be exemplary and illustrative and not limiting in scope. In various embodiments, one or more of the abovedescribed problems have been reduced or eliminated, while other embodiments are directed to other advantages or improvements.
[0035] The following embodiments and aspects thereof are described and illustrated in conjunction with systems, devices and methods which are meant to be exemplary and illustrative and not limiting in scope. In various embodiments, one or more of the abovedescribed problems have been reduced or eliminated, while other embodiments are directed to other advantages or improvements.
[0036] According to a first aspect of the invention, a node-based performance guarantee compliance system comprises: (i) at least one node that comprises at least one storage volume; and (ii) at least one orchestrator configured to orchestrate said at least one node, wherein said at least one orchestrator orchestrates the allocation of storage and application network traffic to and from said at least one node in compliance with a predefined performance guarantee, and whereby said allocation is configured in accordance with available node-based network bandwidth.
[0037] According to another aspect of the invention, the at least one orchestrator allocates network resources in accordance with a profiling procedure designated to analyze the network utilization of the at least one storage volume of the at least one node.
[0038] According to another aspect of the invention, the profiling procedure is designated to analyze network bandwidth utilization for the at least one storage volume of the at least one node, wherein the profiling procedure identifies operational bandwidth requirements for storage and application traffic based on historical and real-time data.
[0039] According to another aspect of the invention, based on the results of the profiling procedure the orchestrator is configured to allocate fixed logical bandwidth shares for storage and application operations based on the profiling procedure’s results.
[0040] According to another aspect of the invention, the profiling procedure is conducted periodically or triggered by specific system events to optimize network resource allocation and maintain appropriate bandwidth segmentation between storage and application traffic.
[0041] According to another aspect of the invention the at least one node has a single physical network interface card (NIC) configured to manage both application network traffic and storage network traffic.
[0042] According to another aspect of the invention, the system allocates distinct logical bandwidth within the single NIC to segregate storage and application network traffic. This logical allocation ensures that storage operations, such as read and write tasks, and application operations, such as user queries, are conducted over dedicated logical pathways within the same physical NIC.
[0043] According to another aspect of the invention, the at least one orchestrator is configured to separate storage network traffic and application network traffic by assigning fixed logical bandwidth allocations to a shared physical network connection.
[0044] According to another aspect of the invention, the logical allocation of bandwidth within the NIC is fixed. According to another aspect of the invention, at least one bandwidth control mechanism, such as CGroups component, is configured by the at least one orchestrator to manage storage network traffic, and wherein such bandwidth control mechanism is configured by the at least one orchestrator to manage application network traffic.
[0045] According to another aspect of the invention, a book-keeping process is designated to record the profiling procedure’s results.
[0046] According to another aspect of the invention, the at least one orchestrator is configured to send a command message to the at least one node commanding the latter to allocate a selected spare storage media.
[0047] According to another aspect of the invention, the at least one node is at least two nodes with a connection therebetween, wherein said connection is conducted using a multipath connection.
[0048] According to another aspect of the invention, the at least one orchestrator is configured to provide alerts and manage failure-mitigation processes between said nodes.
[0049] According to another aspect of the invention, at least one logical volume manager (LVM) component is configured to manage the storage volume capacity of the at least one storage volume.
[0050] According to another aspect of the invention, a bandwidth control mechanism, such as CGroups component, is designated to conduct throttling in order to restrict the performance and resources of each storage volume.
[0051] According to another aspect of the invention, a profiling procedure is designated to analyze and pre-calculate shared network resources, in order to allocate the network bandwidth of at least one storage volume and / or storage resources to either an application or storage volume / s.
[0052] According to another aspect of the invention, the at least one orchestrator configures at least one bandwidth limiting component to control the bandwidth allocated to components of the at least one node, with said at least one bandwidth limiting component being configured to determine the bandwidth provided to a certain component of the at least one node.
[0053] According to another aspect of the invention, the at least one limiting component is a Qdisc ingress rate limit and / or a Qdisc root rate limit.
[0054] According to another aspect of the invention, the at least one component of the at least one node is a part of a kernel core of a Linux OS forming a part of the performance guarantee system.
[0055] According to another aspect of the invention, the at least one component is configured to operate in a cloud computing SaaS environment providing software services.
[0056] According to another aspect of the invention, the at least one orchestrator is configured to mitigate the effect of the noisy neighbor phenomenon on a shared network.
[0057] According to another aspect of the invention, the at least one orchestrator provides a latency guarantee by enabling a locked IO operation ratio that ensures the least one node has enough resources to prevent latency above a threshold determined by the at least one orchestrator.
[0058] According to another aspect of the invention, the resources on which the at least one storage volume is hosted is a solid-state drive (SSD) device. According to another aspect of the invention, the resources on which the at least one storage volume is hosted is a storage class memory (SCM) device.
[0059] According to another aspect of the invention, the at least one storage volume is hosted is a random-access memory (RAM) device.
[0060] According to another aspect of the invention, the resources on which the at least one storage volume is hosted is a hard disk drive (HHD) device.
[0061] According to another aspect of the invention, the at least one orchestrator is a cloudbased service (SaaS).
[0062] According to another aspect of the invention, a method for allocating network bandwidth enabling performance guarantee compliance in a node-based communication environment, comprises: (i) providing at least one node that comprises at least one storage volume; (ii) configuring at least one orchestrator configured to orchestrate said at least one node; and (iii) allocating by way of the at least one orchestrator storage and application network traffic to and from said at least one node in compliance with a predefined performance guarantee configured in accordance with the available bandwidth to and from the at least one node.
[0063] BRIEF DESCRIPTION OF THE FIGURES
[0064] Some embodiments of the invention are described herein with reference to the accompanying figures. The description, together with the figures, makes apparent to a person having ordinary skill in the art how some embodiments may be practiced. The figures are for the purpose of illustrative description and no attempt is made to show structural details of an embodiment in more detail than is necessary for a fundamental understanding of the invention. In the Figures:
[0065] FIG. 1 constitutes a schematic illustration of a conventional storage array system.
[0066] FIG. 2 constitutes a schematic illustration of the operations of a storage disk locking and classification system which may be used in conjunction with the performance guarantee system, according to some embodiments of the invention.
[0067] FIG. 3 constitutes a schematic illustration of a controller operating as a declarative engine forming a part of the performance guarantee system, according to some embodiments of the invention.
[0068] FIG. 4 constitutes a schematic illustration of a media slicing procedure forming a part of performance guarantee system, according to some embodiments of the invention.
[0069] FIG. 5 constitutes a schematic illustration of a media slicing procedure forming a part of performance guarantee system, according to some embodiments of the invention.
[0070] FIG. 6 constitutes schematic illustration of basic components of the performance guarantee system, according to some embodiments of the invention.
[0071] FIG. 7 constitutes a schematic illustration of the operation of the performance guarantee system, according to some embodiments of the invention.
[0072] FIG. 8 constitutes a schematic illustration of further components of the performance guarantee system.
[0073] DETAILED DESCRIPTION OF SOME EMBODIMENTS
[0074] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be understood by those skilled in the art that the present invention may be practiced without these specific details. In other instances, well-known methods, procedures, and components, modules, units and / or circuits have not been described in detail so as not to obscure the invention. Some features or elements described with respect to one embodiment may be combined with features or elements described with respect to other embodiments. For the sake of clarity, discussion of same or similar features or elements may not be repeated.
[0075] Although embodiments of the invention are not limited in this regard, discussions utilizing terms such as, for example, “controlling” “processing,” “computing,” “calculating,” “determining,” “establishing”, “analyzing”, “checking”, “setting”, “receiving”, or the like, may refer to operation(s) and / or process(es) of a controller, a computer, a computing platform, a computing system, or other electronic computing device, that manipulates and / or transforms data represented as physical (e.g., electronic) quantities within the computer's registers and / or memories into other data similarly represented as physical quantities within the computer's registers and / or memories or other information non-transitory storage medium that may store instructions to perform operations and / or processes.
[0076] Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Additionally, some of the described method embodiments or elements thereof can occur or be performed simultaneously, at the same point in time, or concurrently.
[0077] The term "Controller" as used herein refers to any type of computing platform or component equipped with a Central Processing Unit (CPU) or microprocessor and capable of supporting multiple input / output (I / O) ports. The term “Node” as used herein refers to any system or device that serves as a connection point within a network, facilitating communication and data exchange. Nodes can take various forms, including servers, which provide resources or services such as data storage, website hosting, or application management, as well as computers, wireless access points, modems, gateways, switches, and routers, each fulfilling a distinct role in enabling connectivity and interaction.
[0078] Reference is now made to FIG. 1, which schematically illustrates a typical data storage system 10. As shown, a data storage system 10 may comprise at least one target server 100 that may further comprise a storage media 104 and be configured to run an operating system (for example, a Linux based operating systems such as Red Hat, Suse, etc.), wherein said operating system is designated to host data accessible over a DP network.
[0079] According to some embodiments, at least one initiator server 102 may be configured to run an operating system (for example, a Linux based operating systems such as Red Hat, Suse, etc.), wherein said operating system is designated to access and be exposed to remote resource / s over the DP network.
[0080] According to some embodiments, at least one orchestrator 106 may be configured to interact with each of said target server / s 100 and / or initiator server / s 102 in order to control the CP of said DP network. According to some embodiments, a designated portion of the storage media 104 forming a part of the target server 100 may be exposed to the DP network, in other words, a designated physical space is reserved and specified in order to contribute a storage space used by the DP network.
[0081] According to some embodiments, orchestrator 106 is configured to utilize the designated portion of the storage media by orchestrating storage stack (SS) components and standard storage stack (SSS) of the operating system embedded within said target / initiator server / s 100 / 102, such that the initiator server 102 is configured to interact with the target server
[0082] 100 via the DP network. According to some embodiments, data storage system 10 may be designated to perform the following steps:
[0083] • Using the operating system installed on target the storage media 104 forming a part of the target server 100 and designated to host data accessible over the DP network, wherein said storage media 104 is used to utilize a persistent storage medium,
[0084] • Using the operating system’s logical volume manager (LVM) in order to split the storage media 104 to multiple partitions,
[0085] • Using the operating system installed on at least one initiator server 102 in order to access and consume the storage media 104’ partition / s over the DP network in order to utilize a remote media’s capacity and performance,
[0086] • Using the orchestrator 106 which is configured to interact with each of said target server / s 100 and initiator server / s 102, wherein said orchestrator 106 is designated to control the CP of said DP network,
[0087] • Using the operating system to merge at least two network paths in order to utilize a single storage media partition as a single network device by creating a multipath component, and thus, enabling enhanced redundancy and efficiency.
[0088] According to some embodiments, the steps disclosed above may further include using a resource management component / s in order to provide dynamic allocation and de-allocation capabilities configured to be conducted by the operating system and affect, for example, on processor cores and / or memory pages, as well as on various types of bandwidths, computations that compete for those resources. According to some embodiments, the objective of the steps disclosed above is to allocate resources so as to optimize responsiveness subject to the finite resources available. According to some embodiments and as disclosed above, at least one initiator server 102 is configured to interact with at least two target servers 100 using a multipath connection. According to some embodiments, a multipath connection may be used to improve and enhance the connection reliability and provide a wider bandwidth.
[0089] According to some embodiments, the coordination between the initiator server / s 102 and the target server / s 100 or vice versa, may be conducted using a local orchestrator 107 component configured to manage the CP and further configured to be physically installed on each server. According to some embodiments, installing a local orchestrator 107 on each server may provide a flexible way of utilizing the data storage system 10 as well as eliminate the need to provide data storage system 10 with access to internal software and processes of a client’s servers. According to some embodiments, the operations and capabilities disclosed in the current specification with regards to orchestrator 106, may also apply to local orchestrator 107 and vice versa.
[0090] According to some embodiments, the communication between the orchestrator 106 and between either a target server 100 or the initiator server 102, may be conducted via the DP network by utilizing a designated software component installed on each of said servers.
[0091] According to some embodiments, the initiator server 102 is configured to utilize a redundant array of independent disks (RAID) storage stack component (SSC) configured to provide data redundancy originated from multiple designated portions of the storage media 104 embedded within multiple target servers 100.
[0092] According to some embodiments, the RAID SSC is further configured to provide data redundancy originated from combined multiple initiator paths originated from the designated portion of the storage media 104 of at least two target servers 100. According to some embodiments, the target servers 100 may be located at different locations, such as, in different rooms, buildings or even countries. In this case, the orchestration procedure conducted by the orchestrator 106 is allocated across different resiliency domains. For example, the orchestrator 106 may consider various parameters regarding cyber security, natural disasters, financial forecasts, etc. and divert data flow accordingly. According to some embodiments, said orchestration procedure conducted by the orchestrator 106 and configured to utilize servers’ allocation, is conducted with a consideration of maintaining acceptable system balance parameters.
[0093] According to some embodiments, the orchestrator 106 may be configured to interact with server / s 100 / 102 using an administration protocol. According to some embodiments, the designated portion of the storage media 104 may be allocated using a logical volume manager (LVM) SSC. According to some embodiments, the storage media 104 may be solid-state drive (SSD) based, storage class memory (SCM) based, random access memory (RAM) based, hard disk drive (HHD) based, etc.
[0094] According to some embodiments, the orchestrator 106 may be a physical controller device or may be a cloud-based service (SaaS) and may be configured to command and arrange data storage and traffic in interconnected servers, regardless whether orchestrator 106 is a physical device or not.
[0095] According to some embodiments, the operations on each server / s 100 / 102 may be implemented, wholly or partially, by a data processing unit (DPU), wherein said DPU may be an acceleration hardware such as an acceleration card, wherein hardware acceleration may be use in order to perform specific functions more efficiently when compared to software running on a general -purpose central processing unit (CPU), and hence, any transformation of data that can be calculated in software running on a generic CPU can also be calculated in custom-made hardware, or in some mix of both.
[0096] According to some embodiments, under a traditional SDS, the storage stack code would rely on proprietary software, requiring separate, independent installation and maintenance, whereas data storage system 10 is configured to rely on an already installed operating system’ capabilities combined with using the orchestrator 106 discussed above.
[0097] According to some embodiments, under a traditional SDS, the control protocol would rely on proprietary software requiring separate, independent installation and maintenance, whereas the data storage system 10 is configured to rely on the operating system capabilities using the orchestrator 106 discussed above.
[0098] Under a traditional SDS, the nodes interconnect would rely on proprietary software, whereas according to some embodiments, data storage system 10 is configured to rely on standard storage protocols using the orchestrator 106 discussed above.
[0099] Under a traditional SDS, the stack model that controls the storage array system 10 uses a single proprietary code and has components interleave both DP and CP, whereas according to some embodiments, data storage system 10 is configured to utilize the operating system to execute a dummy data plane while using the orchestrator 106 disclosed above in order to emulate a CP and execute its actual operations upon (among others), the DP.
[0100] Some advantages of the various embodiments disclosed above may facilitate ultra-high performance when compared to a traditional SDS operations, for example, with regards to the number of nodes in a cluster which under data storage system 10 are expected to be unlimited.
[0101] As noted there are certain drawback and inefficiencies in coupling the together of DP and the CP with the storage when implemented in traditional networks, including SDNs. According to some embodiments, an SDN may be configured to decouple the DP from the storage, and the CP from the storage. According to some embodiments, a storage system may be built with dummy devices to forward and store data, from the CP of the network, which controls how the traffic will flow through the network while SDN is considered to enable much cheaper equipment, agility and limitless performance than other decoupling means, since more data plane resources can be flexibly added-on, such decoupling using an SDN enables scalability with no limitation and higher survivability rate due to a limited impact on the processed cluster. According to some embodiments, a single orchestrator may be provided to provide storage services to huge cluster since, data capacity, bandwidth and IOPS do not impact the CP services utilization. Such decoupled DP may be utilized for various data Services, such as: Protocols; RAID; Encryption; QoS Limits; Space Allocation; Data Reduction and others. Whereas such a decoupled CP may be utilized for various storage services and coordination, such as: Volume Lifecycle (Create, Delete, Resize, etc.); Storage Pool Management; Snapshot Coordination - Consistency Groups; Failure Monitoring; Performance Monitoring; Failure Handling.
[0102] According to some embodiments, data storage system 10 may be designated to perform the following steps to obtain an SDN based CP and CD and storage decoupling:
[0103] Using the storage CP 104 on one of servers 100, to create storage DP; Using the storage CP on server 102, to create storage DP, to connect to said server 100 and consume the exposed drive chunk;
[0104] • Using the storage CP on another server 100, to create storage DP;
[0105] Using the storage CP on server 102, to create storage DPe, to connect to said another server 100 and consume the exposed drive chunk; Using the storage cCP on server 102, to create storage DP, that includes multipath, RAID, encryption, compression, deduplication, LVM and replication services.
[0106] Reference is made to FIG. 2, which schematically illustrates the operations of storage disk locking and classification system 10 when integrated with the system of the current invention, according to some embodiments. Said integration of said system may be designated by the at least one orchestrator to analyze and categorize storage media based on key performance characteristics such as read / write speeds, latency, endurance, and capacity, which compliments the performance guarantee provided by the system and method of the present invention. According to some embodiments of the invention, the disk locking and classification systems 10 is a disk slicing system. Said integration affords to the embodiments equipped therewith a number of benefits relating to the fidelity, reliability, and adaptability of storage operations, benefits that are the product of complimentary mechanism of the performance guarantee system and the disk locking and classification system.
[0107] According to some embodiments, in accordance with the performance characteristics of each storage device determined by said storage disk and locking classification system, the orchestrator can allocate network resources in line with the capabilities of the underlying storage media.
[0108] As shown, in operation 12, a server comprising a flash storage drive is identified by the disk locking and classification system.
[0109] According to some embodiments, the capacity of a flash drive is typically identified by the amount of storage space it offers, usually measured in gigabytes (GB) or terabytes (TB). This capacity may be determined by the number of memory cells within the flash drive and the storage technology used. Flash drives also usually have a controller chip that manages the memory cells and handles read and write operations. Additionally, some storage space is reserved for wear-leveling, error correction, and other overhead functions, which may reduce the available storage capacity slightly. A flash drive's capacity is identified by the total amount of data it can store, which is determined by the type of NAND flash memory, the number of memory cells, and the technology used in its construction, wherein the actual usable capacity may be slightly lower due to formatting and overhead.
[0110] In operation 14, the storage disk is locked to a specific read / write ratio and is then configured to undergo a stress-testing procedure to confirm its performance under various conditions, the
[0111] According to some embodiments, an (Input-Output) IO operation ratio of the storage disk is configured to be locked in a pre-calculated ratio in order to ensure pre-designated performance of a storage disk.
[0112] According to some embodiments, an IO operation ratio lock may be determined by the following operations:
[0113] • Data Collection: Data on the read and write operations performed by the storage device must be collected.. This data is typically collected by monitoring the device's I / O operations using specialized tools or software. The monitoring process spans a specific period, which can range from minutes to days, depending on the level of detail required for the analysis.
[0114] • Counting Reads and Writes: During the data collection period, the tool or software counts the number of read and write operations that occur on the device. Read operations involve retrieving data from the storage, while write operations involve storing data on the storage. The read / write ratio can be calculated using the following formula:
[0115] Read / Write Ratio = (Number of Read Operations) / (Number of Write Operations)
[0116] The formula may be described as a percentage by dividing the number of read operations by the total number of operations (reads + writes) and multiplying by 100:
[0117] Read / Write Ratio (%) = [(Number of Read Operations) / (Number of Read Operations + Number of Write Operations)] x 100
[0118] The read / write ratio describes how often data is being read compared to how often it is being written. A high read / write ratio means that the device is primarily used for reading data, which is common for devices that store and serve data, like hard drives in file servers. Conversely, a high write ratio indicates a device that is frequently written to, like a database server or a device used for constant data logging.
[0119] In operation 16, stress testing is conducted to simulate worst-case scenarios for the storage drive under the predetermined read / write ratio. According to some embodiments, this involves applying a workload designed to push the drive to its operational limits, ensuring it can maintain performance and reliability. This testing assesses the drive's ability to handle intensive read and write operations, heat generation, and other factors that could impact performance during extreme conditions.
[0120] In operation 18, real-time performance metrics are assessed to evaluate how said flash storage device performs under typical operating conditions. Unlike stress testing, which simulates extreme workloads, real-time performance monitoring focuses on understanding the device's behavior during standard operations. This process measures key performance metrics such as read / write speeds, latency, and input / output IOPS. These insights are critical for analyzing performance under normal conditions and optimizing resource allocation within the system.
[0121] According to some embodiments, real-time performance assessment can be conducted by running benchmark tests on the storage media to obtain relevant performance metrics. Most benchmarking tools provide results for sequential and random read / write operations, as well as sequential read / write throughput. These tests can evaluate the storage device's ability to maintain consistent throughput over time, particularly in real-world scenarios. Random access performance can also be assessed, which is especially useful for applications with small random I / O operations, such as databases and virtual machines. Additionally, these tests analyze metrics such as IOPS, latency, and disk usage, while monitoring the device's temperature, as excessive heat can negatively affect performance and longevity.
[0122] If a flash storage disk is nearing capacity, it can lead to performance degradation, hence there is a need to ensure that there is sufficient free space and monitor performance over time. Tracking how performance changes as the storage media ages and fills with data provides valuable insights into its long-term behavior and reliability.
[0123] Assessing the real-time performance of storage media is an ongoing process. Regular monitoring and benchmarking facilitate the identification of any performance issues, enabling corrective actions and storage systems optimization.
[0124] Assessing the real-time performance of a storage volume may be performed by applying a specific pattern designated to cause the drive to perform in a specific way in realtime.
[0125] In operation 20, the storage disk locking and classification system 10 is configured to identify the physical location of the flash storage device in order to provide identification of physical location. In operation 22 the classification gathered while utilizing the steps above may be stored in a catalog, which is a data base containing various data pertaining to the performance of the flash drive.
[0126] According to some embodiments, the disk locking and classification system 10 monitors performance of the storage disks, thereby enabling the performance guarantee system to optimize network allocation in accordance with the storage disks capabilities.
[0127] Reference is now made to FIG. 3, which schematically illustrates a controller operating as a declarative engine forming a part of the disk locking and classification system 10 described in FIG. 2, which may be used in conjunction with the performance guarantee system and method of the present invention.
[0128] According to some embodiments, declarative programming is a programming paradigm in which an operator specifies the desired outcome or end result, leaving the system or engine to determine the steps needed to achieve it. This approach emphasizes defining "what" the operator wants to accomplish, rather than detailing "how" to accomplish it, simplifying the implementation process and allowing the underlying system to handle the operational complexity.
[0129] According to some embodiments, a declarative engine integrated into the disk locking and classification system 10 enables an operator to define desired configurations, policies, or resource allocations in a high-level declarative format. Rather than requiring a detailed, step- by-step procedure for resource allocation, the operator specifies the desired outcome, and the engine automatically interprets and enforces these declarations to achieve the specified result.
[0130] For example, a key characteristic of a declarative engine in the context of cloud storage services is the ability to configure systems through high-level declarations. An operator specifies requirements or configurations using declarative syntax or language, such as defining storage capacity, access controls, replication policies, or other relevant parameters. In cloud storage and data management, declarative approaches are widely employed across various services and tools, streamlining the configuration process by focusing on desired outcomes rather than implementation details.
[0131] According to some embodiments, drive 202 may have a storage volume A which is close to full capacity, whereas other drives 204, 206, 208, 210 and 212 relating to the same or other nodes, which have storage volumes B, C, D, E, and F, respectively, which are have additional available capacity and may be used to store additional data. According to some embodiments, the disk locking and classification system 10 may utilize a declarative engine in order to direct resource allocation to ensure optimal storage disk performance.
[0132] Reference is now made to FIG. 4, which schematically illustrates a media partitioning procedure forming a part of the disk locking and classification system 10, according to some embodiments. As depicted, an allocation enforcement procedure is implemented to ensure that application servers can utilize storage resources without exceeding predefined limits, thereby guaranteeing the performance of the storage disks and therefore the overall system. For instance, a storage disk with a capacity to handle 1,000,000 IOPS may be subjected to requests exceeding this limit, which could result in resources allocation to other disks, potentially impacting their performance and overall system reliability. Without an enforcement mechanism, such overuse could disrupt the performance and reliability of the storage system.
[0133] According to some embodiments, the Logical Volume Management (LVM) component
[0134] 400, a storage device management technology integrated into the Linux kernel, enables users to pool and abstract the physical layout of storage devices for flexible administration. This component is configured to reduce the capacity of storage volume 402, without adversely affecting its performance.
[0135] According to some embodiments, the CGroups component 404, another Linux kernel feature that provides hierarchical management and allocation of system resources, including block I / O, is configured to perform throttling. This throttling results in the creation of a new storage volume 406 with reduced but guaranteed performance. Performance in this context may pertain to metrics such as read / write speeds, IOPS, bandwidth, and other storage-related parameters. According to some embodiments, throttling in the context of cloud storage performed for example, by CGroups component 404 may refer to the intentional control or limitation of the rate at which a user or application can access or interact with a storage service. This control is usually implemented to prevent excessive resource usage, protect the service from abuse, and ensure fair and equitable access for all users.
[0136] According to some embodiments, key aspects of throttling in cloud storage may include:
[0137] • Requesting Rate Limiting: Throttling often involves limiting the number of requests or operations that can be performed within a specific time period. This can include read or write operations, API (Application Programming Interfaces) calls, or other interactions with the storage service.
[0138] • Bandwidth Limitations: Throttling can also be applied to limit the amount of data that can be transferred in and out of the storage service within a given timeframe. This helps prevent users from consuming an excessive amount of network bandwidth.
[0139] • Concurrency Limits: Some cloud storage services impose limits on the number of simultaneous connections or concurrent operations that a user or application can have. This may prevent a single user or application from monopolizing system resources. Dynamic Adjustment: Throttling mechanisms may dynamically adjust based on the overall load on the storage system. During periods of high demand or resource contention, the throttling limits may be more stringent to ensure system stability.
[0140] • Rate-Based Billing: In some cases, throttling can also be tied to billing models. Cloud service providers may charge users based on their usage patterns, and throttling can be a way to manage costs and prevent unexpected spikes in usage.
[0141] • API Throttling: Throttling is commonly applied to APIs provided by cloud storage services. This ensures that applications and developers adhere to usage limits and don't overload the API endpoints.
[0142] According to some embodiments, LVM component 400 and CGroups component 404 facilitate the creation of storage volume 406 which has reduced capacity and fixed performance, limiting the resource allocation to target 408.
[0143] According to some embodiments, these restriction components and steps enable the disk locking and classification system 10 to guarantee the performance of the storage disk, such that the subject performance guarantee system can allocate network bandwidth in line with the performance characteristics of the storage disk.
[0144] According to some embodiments, utilizing the LVM components enables the disk locking and classification system 10 to ensure consistent and reliable performance.
[0145] According to some embodiments, the LVM component’s utilization and / or the CGroups utilization are configured to be conducted separately with regards to input data or output data or with regards to bandwidth or IOPS. According to some embodiments, the allocated network bandwidth and / or storage resources for every storage volume is determined in accordance with predefined performance requirements flow.
[0146] Reference is now made to FIG. 5, which schematically illustrates a specific utilization of the disk locking and classification system 10, according to some embodiments. As shown, drive 500 comprises at least two volumes 502 and 504, wherein application stacks 506 and / or 508 are designated to process the data flow. According to some embodiments, in such an arrangement, there is a possible scenario in which application stack 508 will consume resources / IOPS from both several storage volumes. For example, if application stack 508 is not restricted it can use resources / IOPS from both volumes 502 and 504, hence disrupting their operation.
[0147] For example, drive 500 may be capable of executing 1.5 million IOPS, volume 502 may be restricted to use drive 500 for 0.5 million IOPS and volume 504 may be restricted to use drive 500 for 1 million IOPS, if application stack 508 further utilizes 0.7 million IOPS, the excesses IOPS will be reduced from the resources provided by volumes 502 & 504.
[0148] According to some embodiments, the disk locking and classification system prevents such a scenario by disk locking, thereby facilitating the maintenance of the performance guarantee of the claimed invention.
[0149] According to some embodiments, the media slicing procedure depicted in FIG. 4 is also capable of managing and control the procedure depicted in FIG. 5.
[0150] According to some embodiments, an application component may be a modular and self-contained unit within a software application that performs a specific function or set of related functions. Application components are designed to be modular, reusable, and interchangeable, contributing to the overall structure and functionality of the software system.
[0151] These components can be thought of as building blocks that, when combined, create a complete and functional application.
[0152] According to some embodiments, a designated limiting component / s such as Qdisc ingress rate limit and Qdisc root rate limit are configured to determine the bandwidth provided to a certain application component.
[0153] According to some embodiments, a kernel core component of a Linux OS forming is designated to form a part of the disk locking and classification system 10. According to some embodiments, the various components forming disk locking and classification system 10 are configured to operate in a cloud computing SaaS environment providing software services.
[0154] Reference is now made to FIG. 6, which schematically illustrates the basic components of the performance guarantee system of the subject invention, according to some embodiments. As shown, a conventional network comprises two separated networks, storage network 600 and applications network 602, both of which are connected to media storage volumes 604. In cloud computing, each cloud instance has a single NIC configured to operate both storage and application networks.
[0155] According to some embodiments, provision of a performance guarantee to a system operated in cloud computing environment is challenging, since the network application network may adversely affect the storage network since they share the same underlying hardware (a single NIC). Hence, there is a need to provide a solution that will enable an end to end performance guarantee assurance as further disclosed in FIG. 7.
[0156] Reference is now made to FIG. 7, which schematically illustrates a specific utilization of the subject performance guarantee system, according to some embodiments. According to some embodiments, the network bandwidth provided by the single pipe of Network Component 700, comprising both storage and network allocation traffic, is separated into logical pipes by the performance guarantee system, wherein each logical pipe is allocated either network or storage network traffic.
[0157] According to some embodiments, each logical pipe can be further configured to provide a defined allocation of storage or application network bandwidth to at least one storage volume.
[0158] By way of example, network component 700 may be a network card having a 40 Gbps rate limit and be connected to a storage disk comprising three storage volumes. In accordance with the profiling procedure, the performance guarantee system can separate the single network pipe into logical pipes allocated to either network or storage traffic, thereby enabling specific network bandwidth allocation to each of said storage volumes. For example, storage volume 702 which is designated to handle application operations is allocated 20 Gbps of the application network logical pipe, storage volume 704 which is designated to handle storage operations is allocated 15 Gbps of the storage network logical pipe, and storage volume 706 which is designated to handle storage operations is allocated 5 Gbps of the storage network logical pipe.
[0159] According to some embodiments, dividing the single pipe of a single NIC into multiple logical pipes, each with a dedicated bandwidth allocation, ensures guaranteed network performance for each storage volume. This separation prevents the operations of one volume from impacting the performance of others, thereby maintaining consistent and reliable operation across all volumes.
[0160] According to some embodiments, the splitting steps disclosed above enable the performance guarantee system to conduct media slicing / dividing large network components into smaller, more manageable virtual pieces can that be distributed across multiple storage devices or locations and may be beneficial in many aspects, for example: • Scalability: slicing / dividing large network components into smaller, more manageable virtual pieces allows for more efficient scalability. As the amount of data grows, additional storage resources can be added, and the workload can be distributed across multiple slices, preventing bottlenecks and ensuring better performance.
[0161] • Parallel processing: sliced virtual network components enables parallel processing, where different virtual network component can be processed independently. This can lead to improved performance and faster data retrieval times, especially in distributed computing environments.
[0162] • Fault tolerance: data slicing can enhance fault tolerance. If one storage location or server experiences a failure, the impact is limited to the affected virtual network components, and other virtual network components remain accessible. This may improve the overall resilience of the storage system.
[0163] • Load balancing: sliced virtual network components allows for effective load balancing. For example, different virtual network components can be distributed across servers or storage node to ensure that the workload is evenly distributed, preventing individual components from becoming overloaded.
[0164] • Optimized retrieval: when retrieving specific portions of data, the slicing approach can lead to optimized retrieval times since only the relevant virtual network components need to be accessed, reducing the amount of data transferred and improving overall efficiency. Improved data management: data slicing can enhance data organization and management. And may allow for more granular control over data placement, access permissions, and data lifecycle management.
[0165] • Cost efficiency: cloud storage providers often charge based on usage, and data slicing can contribute to cost efficiency. By distributing data strategically and optimizing storage usage, organizations may potentially reduce costs associated with storage services.
[0166] Reference is now made to FIG. 8, which schematically illustrates a the performance guarantee procedure in respect to NIC 800, forming a part of performance guarantee system, according to some embodiments.
[0167] The at least one applications virtual network component disclosed in FIG. 7 may be incorporated into the Linux kernel and may be designated to be implemented while using external storage services, disk
[0168] As part of the performance guarantee procedure, application virtual network 802 may be designated to create and utilize data path 804 which is in turn designated to utilize at least two paths, for example: (i) qdisc ingress path 805 (wherein qdisc is a mechanism used in the Linux kernel in order to manage network traffic by applying various queuing algorithms and disciplines) configured to force ingress rate limit and (ii) qdisc root rate limit path 806 which is configured to force a root rate limit.
[0169] For example, network component 800, can be set to limit data ingress of 25 Gbps and set to limit data output to 30 Gbps. According to some embodiments, storage media 807 (which may be a physical storage volume, a virtual storage volume, etc.) is designated to transfer data in a similar manner, for example, a linux kernel application (that may be a software program or component that directly interacts with and relies on the operating system's kernel and may perform tasks that require privileged access to system resources or services, as opposed to user-level applications that run in a more restricted environment) such as kernel stack 808 (that may be a part of the kernel space), Cgroup blkio 810 (that may be subsystem controls and monitors access to I / O on block devices by tasks in Cgroups), are designated to apply initiator 812 (which may be a part of a server’s operating system and uses host CPU resources to map the Small Computer System Interface (SCSI) I / O command set to TCP / IP for use by the iSCSI storage system) in order to utilize at least two storage paths.
[0170] According to some embodiments, said at least two storage paths may be designated to allow, for example, 10 Gbps rate limit of data to be transferred to (in) network component 800 and a 15 Gbps rate limit of data to be transferred from (out) said network component 800.
[0171] Although the present invention has been described with reference to specific embodiments, this description is not meant to be construed in a limited sense. Various modifications of the disclosed embodiments, as well as alternative embodiments of the invention will become apparent to persons skilled in the art upon reference to the description of the invention. It is, therefore, contemplated that the appended claims will cover such modifications that fall within the scope of the invention.
Claims
CLAIMS1. Anode-based performance guarantee compliance system, comprising:(i) at least one node that comprises at least one storage volume; and(ii) at least one orchestrator configured to orchestrate said at least one node, wherein said at least one orchestrator orchestrates the allocation of storage and application network traffic to and from said at least one node in compliance with a predefined performance guarantee; and whereby said allocation is configured in accordance with available node-based network bandwidth.
2. The system of claim 1, wherein the at least one orchestrator allocates network resources in accordance with a profiling procedure designated to analyze the network utilization of the at least one storage volume of the at least one node.
3. The system of claim 1, wherein the at least one node has a single physical network interface card (NIC) configured to manage both application network traffic and storage network traffic.
4. The system of claim 1, wherein the at least one orchestrator is configured to separate storage network traffic and application network traffic by assigning fixed logical bandwidth allocations to a shared physical network connection.
5. The system of claim 1, wherein at least one bandwidth control mechanism is designated by the at least one orchestrator to manage storage network traffic and network flow bandwidth.
6. The system of claim 5, wherein at least one bandwidth control mechanism is designated to conduct throttling in order to restrict the performance and resources of each storage volume.
7. The system of claim 1, wherein a profiling procedure is designated to analyze and pre-calculate shared network resources, in order to allocate the network bandwidth of at least one storage volume and / or storage resources to either an application or storage volume / s.
8. The system of claim 1, wherein a book-keeping process is designated to record the profiling procedure’s results.
9. The system of claim 1, wherein the at least one orchestrator is configured to send a command message to the at least one node commanding the latter to allocate a selected spare storage media.
10. The system of claim 1, wherein the at least one node is at least two nodes with a connection therebetween, wherein said connection is conducted using a multipath connection.
11. The system of claim 1, wherein at least one logical volume manager (LVM) component is configured to manage the storage volume capacity of the at least one storage volume.
12. The system of claim 1, wherein the at least one orchestrator configures at least one bandwidth limiting component to control the bandwidth allocated to components of the at least one node, wherein said at least one limiting component is configured to determine the bandwidth provided to a certain component of the at least one node.
13. The system of claim 12, wherein the at least one component is a part of a kernel core of a Linux OS forming a part of the performance guarantee system.
14. The system of claim 12, wherein the at least one component is configured to operate in a cloud computing SaaS environment providing software services.
15. The system of claim 1, wherein the at least one orchestrator is configured a mitigate the effect of the noisy neighbor phenomenon on a shared network.
16. The system of claim 1, wherein the at least one orchestrator provides a latency guarantee by enabling a locked IO operation ratio that ensures the least one node has enough resources to prevent latency above a threshold determined by the at least one orchestrator.
17. The system of claim 1, wherein the resources on which the at least one storage volume is hosted is a solid-state drive (SSD) device.
18. The system of claim 1, wherein the resources on which the at least one storage volume is hosted is a storage class memory (SCM) device.
19. The system of claim 1, wherein the resources on which the at least one storage volume is hosted is a random-access memory (RAM) device.
20. The system of claim 1, wherein the resources on which the at least one storage volume is hosted is a hard disk drive (HHD) device.
21. The system of claim 1, wherein the at least one orchestrator is a cloud-based service (SaaS).
22. A method for allocating network bandwidth enabling storage performance guarantee compliance in a node-based communication environment, comprising the steps:(i) configuring at least one node that comprises at least one storage volume;(ii) configuring at least one orchestrator configured to orchestrate said at least one node,(iii) allocating by way of the at least one orchestrator storage and application network traffic to and from said at least one node in compliance with a predefined performance guarantee in accordance with available bandwidth.
Citation Information
Patent Citations
Bandwidth control method and apparatus for solving service quality degradation caused by traffic overhead in SDN-based communication node
US20190349081A1
Dynamic provisioning of storage in the cloud
US20210409342A1