File system function layering method, system, readable medium, and apparatus using cloud object storage
Patent Information
- Application Number
- CN202311337609.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-05-31
- Filing Date
- 2017-12-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2037-12-29
AI Technical Summary
将遗留应用转换成使用对象接口将是昂贵的并且可能不实际或甚至不可能
[0010]本文描述了各种技术(例如,系统、方法、在非瞬态机器可读存储介质中有形地实施的计算机程序产品等),用于在对象接口上提供文件系统功能的分层。在某些实施例中,文件系统功能可以在云对象接口上分层以提供基于云的存储,同时允许遗留应用预期的功能。例如,可移植操作系统接口(POSIX)接口和语义可以在基于云的存储上分层,同时以与对名称层次结构(name hierarchy)中的数据组织的基于文件的访问一致的方式提供对数据的访问。各种实施例还可以提供数据的存储器映射,使得存储器映射改变被反映在持久存储装置中,同时确保存储器映射改变和写入之间的一致性。例如,通过将ZFS文件系统基于盘的存储转换成ZFS基于云的存储,ZFS文件系统获得了云存储的弹性。
Smart Images

Figure CN117193666B_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application 201780080040.3, filed on December 29, 2017, entitled "Method, System, Readable Medium and Apparatus for Functional Layering of File System Using Cloud Object Storage".
[0002] Cross-reference to related applications
[0003] This PCT application claims the rights and priority of U.S. Provisional Application No. 62 / 443,391, filed January 6, 2017, entitled "FILE SYSTEM HIERARCHIES AND FUNCTIONALITY WITH CLOUD OBJECT STORAGE," and U.S. Non-Provisional Application No. 15 / 610,349, filed March 31, 2017, entitled "FILE SYSTEM HIERARCHIES AND FUNCTIONALITY WITH CLOUD OBJECT STORAGE." The entire contents of the foregoing applications are incorporated herein by reference for all purposes. Technical Field
[0004] This disclosure generally relates to systems and methods for data storage, and more specifically to layering file system functionality on an object interface. Background Technology
[0005] The continuous expansion of the internet, along with the expansion and increasing complexity of computing networks and systems, has led to a surge in the amount of content stored and accessible via the internet. This, in turn, has driven the demand for large and complex data storage systems. As the demand for data storage continues to grow, larger and more complex storage systems are being designed and deployed. Many large-scale data storage systems utilize storage devices that include arrays of physical storage media. These storage devices are capable of storing massive amounts of data. For example, at this time, Oracle's SUN ZFS Storage ZS5-4 equipment can store up to 6.9 PB of data. Moreover, multiple storage equipments can be networked together to form storage pools, which can further increase the volume of data stored.
[0006] Typically, large storage systems like these can include file systems for storing and accessing files. In addition to storing system files (operating system files, device driver files, etc.), the file system also provides storage and access to user data files. If any of these files (system files and / or user files) contains critical data, it is advantageous to employ a backup storage scheme to ensure that this critical data is not lost in the event of a failure of the file storage device.
[0007] Conventional cloud-based storage is object-based and offers elasticity and scalability. However, cloud object storage presents several challenges. It provides an interface based on retrieving and putting entire objects. It offers limited search capabilities and typically suffers from high latency. This limited cloud-based interface does not meet the needs of native file system applications. Converting legacy applications to use the object interface would be expensive and potentially impractical or even impossible. Furthermore, cloud object storage encryption makes the data used to create encryption keys more vulnerable and insecure.
[0008] Therefore, systems and methods are needed to address the aforementioned problems in order to provide a layered approach to file system functionality on an object interface. This disclosure addresses this need, as well as other requirements. Summary of the Invention
[0009] Some embodiments of this disclosure generally relate to systems and methods for data storage, and more specifically to systems and methods for layering file system functionality on an object interface.
[0010] This document describes various techniques (e.g., systems, methods, computer program products tangibly implemented on non-transient machine-readable storage media, etc.) for providing layering of file system functionality on an object interface. In some embodiments, file system functionality can be layered on a cloud object interface to provide cloud-based storage while allowing functionality intended for legacy applications. For example, a Portable Operating System Interface (POSIX) interface and semantics can be layered on cloud-based storage while providing access to data in a manner consistent with file-based access to data organization in a name hierarchy. Various embodiments can also provide memory mapping of data such that memory mapping changes are reflected in persistent storage while ensuring consistency between memory mapping changes and writes. For example, by transforming ZFS file system disk-based storage into ZFS cloud-based storage, the ZFS file system gains the resilience of cloud storage.
[0011] Other areas of application of this disclosure will become clear from the detailed description provided below. It should be understood that the detailed description and specific examples are intended for illustrative purposes only when indicating various embodiments and are not intended to necessarily limit the scope of this disclosure. Attached Figure Description
[0012] A further understanding of the nature and advantages of the embodiments according to this disclosure can be achieved by taking into account the following figures and the remainder of the specification.
[0013] Figure 1 An example storage network that can be used according to certain embodiments of the present disclosure is illustrated.
[0014] Figure 2 An example of a file system that can be executed in a storage environment according to certain embodiments of the present disclosure is illustrated.
[0015] Figures 3A-3F The illustration depicts copy-on-write processing for a file system according to certain embodiments of the present disclosure.
[0016] Figure 4 This is a high-level diagram illustrating an example of a hybrid cloud storage system according to certain embodiments of the present disclosure.
[0017] Figure 5 An example network file system of a hybrid cloud storage system according to certain embodiments of the present disclosure is illustrated.
[0018] Figure 6 This is a diagram illustrating additional aspects of a cloud interface apparatus for a hybrid cloud storage system according to certain embodiments of the present disclosure.
[0019] Figures 7A-7F This is a block diagram illustrating an example method according to certain embodiments of the present disclosure, which targets certain features of COW processing for hybrid cloud storage systems, including data services, snapshots, and cloning.
[0020] Figure 8 This is a high-level diagram illustrating an example of a cloud interface apparatus for processing incremental modifications according to certain embodiments of this disclosure.
[0021] Figure 9 This is a block diagram illustrating an example method according to certain embodiments of the present disclosure, which targets certain characteristics of a hybrid cloud storage system that ensure integrity in the cloud and always-consistent semantics based on an eventually consistent object model.
[0022] Figure 10 This is a high-level diagram illustrating an example of a cloud interface apparatus for processing verification according to certain embodiments of the present disclosure.
[0023] Figure 11 This is a simplified example of a feature of a hybrid cloud storage system according to certain embodiments of the present disclosure.
[0024] Figure 12 This is a simplified example of a feature of a hybrid cloud storage system according to certain embodiments of the present disclosure.
[0025] Figure 13This is a block diagram illustrating an example method according to certain embodiments of the present disclosure, which targets certain features of a hybrid cloud storage system for cache management and cloud latency masking.
[0026] Figure 14 An example network file system of a hybrid cloud storage system that facilitates synchronized mirroring according to certain embodiments of the present disclosure is illustrated.
[0027] Figure 15 This is a block diagram illustrating an example method according to certain embodiments of the present disclosure, which addresses certain features of a hybrid cloud storage system for synchronized mirroring and cloud latency masking.
[0028] Figure 16 A simplified diagram is depicted for implementing a distributed system according to certain embodiments of this disclosure.
[0029] Figure 17 This is a simplified block diagram of one or more components of a system environment according to certain embodiments of the present disclosure, wherein services provided by one or more components of the system can be provided as cloud services through the system environment.
[0030] Figure 18 An exemplary computer system is illustrated, in which various embodiments of the present invention can be implemented.
[0031] In the accompanying drawings, similar parts and / or features may have the same reference labels. Additionally, various parts of the same type can be distinguished by a dashed underline following the reference label and a second label to differentiate similar parts. If only the first reference label is used in the specification, the description applies to any similar part having the same first reference label, regardless of the second reference label. Detailed Implementation
[0032] The following description provides only preferred exemplary embodiments (one or more) and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of preferred exemplary embodiments (one or more) will provide those skilled in the art with an implementation description of the preferred exemplary embodiments for carrying out this disclosure. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this disclosure as set forth in the appended claims.
[0033] As mentioned above, cloud-based storage offers resilience and scalability, but it also presents several challenges. Cloud object storage provides an interface based on retrieving and placing entire objects. It offers limited search capabilities and typically suffers from high latency. This limited cloud-based interface does not meet the needs of native file system applications. Converting legacy applications to use the object interface would be expensive and potentially impractical or even impossible. Therefore, a solution is needed that allows direct access to cloud object storage devices without requiring changes to file system applications, as this is inherently complex and costly.
[0034] The solution should allow for the preservation of native application interfaces without introducing various types of adaptation layers to map data from local storage systems to object storage devices in the cloud. Therefore, according to certain embodiments of this disclosure, file system functionality can be layered on top of the cloud object interface to provide cloud-based storage while allowing the functionality expected by legacy applications. For example, non-cloud-based legacy applications can access data primarily as files and can be configured for POSIX interfaces and semantics. From the perspective of legacy applications, it is expected that the content of files can be modified without rewriting them. Similarly, it is expected that data can be organized in a name hierarchy.
[0035] To accommodate this expectation, some embodiments can layer the POSIX interface and semantics on cloud-based storage devices, while providing data access from the user's perspective in a manner consistent with file-based access to data organization in a name hierarchy. Additionally, some embodiments can provide memory mapping of data, such that memory mapping changes are reflected in persistent storage, while ensuring consistency between memory mapping changes and writes. By transforming disk-based storage of the ZFS file system into cloud-based storage, the ZFS file system gains the resilience of cloud storage. By mapping "disk blocks" to cloud objects, the storage requirements of the ZFS file system are only the "blocks" actually used. The system can always be thinly provisioned, without the risk of exhausting backup storage. Conversely, cloud storage gains ZFS file system semantics and services. Full POSIX semantics, as well as any additional data services (such as compression, encryption, snapshots, etc.) provided by the ZFS file system, can be offered to cloud clients.
[0036] Some implementations can provide the ability to migrate data to and from the cloud, and can offer coexistence of on-premises and cloud data through hybrid cloud storage systems. These systems provide storage elasticity and scalability while layering ZFS file system functionality on cloud storage. By extending the ZFS file system to allow object storage in cloud object repositories, a bridge can be provided to traditional object storage, while retaining all ZFS file system data services and functionality. Bridging the gap between the traditional on-premises file system and the ability to store data in various cloud object repositories helps to significantly improve performance.
[0037] Furthermore, embodiments of the present invention enable the use of traditional ZFS data services in conjunction with hybrid cloud storage. As examples, compression, encryption, deduplication, snapshots, and cloning are each available in some embodiments of the invention and are briefly described below. In this invention, when storage is extended to the cloud, users can continue to seamlessly use all the data services provided by the ZFS file system. For example, an Oracle key manager or equivalent manages keys locally on the user's machine, allowing end-to-end secure encryption using locally managed keys when storing to the cloud. The same commands used for compression, encryption, deduplication, snapshotting, and cloning on disk storage are used for storage to the cloud. Therefore, users continue to benefit from the efficiency and security provided by ZFS compression, encryption, deduplication, snapshots, and cloning.
[0038] Compression is typically enabled because it reduces the resources required to store and send data. Computational resources are consumed during compression and often again during the reverse of this process (decompression). Data compression is subject to space-time complexity trade-offs. For example, a compression scheme might require sufficiently fast, intensive decompression processing to be consumed while decompression is in progress. Designing a data compression scheme involves trade-offs between various factors, including the degree of compression and the computational resources required to compress and decompress the data.
[0039] ZFS encryption enables an end-to-end secure block system with locally stored encryption keys, providing an additional layer of security. ZFS encryption itself does not prevent block theft, but it refuses to deliver message content to interceptors. In the encryption scheme, the intended block is encrypted using an encryption algorithm, generating ciphertext that can only be read during decryption. For technical reasons, the encryption scheme typically uses a pseudo-random encryption key generated by the algorithm. In principle, it is possible to decrypt a message without possessing the key; however, for well-designed encryption schemes, significant computational resources and skill are required. ZFS blocks are encrypted using AES (Advanced Encryption Standard) with key lengths of 128, 192, and 256.
[0040] Block duplication is a specialized data compression technique used to eliminate duplicate copies of redundant data blocks. Block deduplication improves storage device utilization and can also be applied to network data transmission to reduce the number of data blocks that must be sent to storage. In deduplication, unique data blocks are identified and stored during analysis. As analysis continues, other data blocks are compared to stored copies, and whenever a match occurs, the redundant data block is replaced with a small reference to the stored data block. Given that the same data block pattern can occur dozens, hundreds, or even thousands of times, deduplication significantly reduces the number of data blocks that must be stored or transmitted.
[0041] Snapshots stored in the ZFS cloud object repository are created seamlessly within the ZFS system. A snapshot freezes certain data and metadata blocks so that they are not written to when a backup is needed. A tree hierarchy can have many snapshots, and each snapshot is saved until it is deleted. Snapshots can be stored locally or in the cloud object repository. Snapshots are "free" in the ZFS system because they require no additional storage capacity other than the root block to which the snapshot was created. When accessed from a snapshot reference to the root block, the root block and all subsequent blocks starting from the root block are unavailable for copy-on-write operations. In the next iteration after a snapshot is taken—the new root block becomes the active root block.
[0042] Clones are created from snapshots, and unlike snapshots, write-time copy operations are available for block pairs accessed using clone references to the root block. Clones allow for development and troubleshooting on the system without corrupting the active root block and the tree. Clones are linked to snapshots, and snapshots cannot be deleted if clones linked to snapshot blocks still exist. In some cases, clones can be promoted to the active hierarchical tree.
[0043] Various embodiments will now be discussed in more detail with reference to the accompanying drawings, from Figure 1 start.
[0044] Figure 1 An example storage network 100 that can be used to implement certain embodiments of the present disclosure is illustrated. Figure 1 The selection and / or arrangement of hardware devices depicted are shown as examples only and are not intended to be limiting. Figure 1Multiple storage devices 120 are provided and connected via one or more switching circuits 122. The switching circuits 122 can connect the multiple storage devices 120 to multiple I / O servers 136, which in turn can provide access to the multiple storage devices 120 for client devices such as local computer systems 130, computer systems available via network 132, and / or cloud computing systems 134.
[0045] Each I / O server 136 may execute multiple independent file system instances, each responsible for managing a portion of the total storage capacity. These file system instances may include an Oracle ZFS file system, as will be described in more detail below. I / O server 136 may include blade and / or standalone servers, each including host port 124 to communicate with client devices by receiving read and / or write data access requests. Host port 124 may communicate with external interface provider 126, which identifies the correct data storage controller 128 for servicing each I / O request. Data storage controllers 128 may each specifically manage a portion of the data content in one or more of the storage devices 120 described below. Therefore, each data storage controller 128 may access a logical portion of the storage pool and satisfy data requests received from external interface provider 126 by accessing its own data content. Redirection via data storage controller 128 may include redirecting each I / O request from host port 124 to the file system instance (e.g., a ZFS instance) executing on I / O server 136 and responsible for the requested block. For example, this could include redirecting a request from host port 124-1 on one I / O server 136-1 to a ZFS instance on another I / O server 136-n. This redirection could allow access from any host port 124 to any portion of the available storage capacity. The ZFS instance could then issue the necessary direct I / O transactions to any storage device in the storage pool to fulfill the request. Acknowledgments and / or data could then be forwarded back to the client device via the originating host port 124.
[0046] A low-latency memory-mapped network can bind host port 124, any file system instance, and storage device 120 together. This network can be implemented using one or more switching circuits 122 (such as Oracle's Sun Data Center InfiniBand Switch 36) to provide scalable, high-performance clustering. Bus protocols (such as the PCI Express bus) can route signals within the storage network. I / O server 136 and storage device 120 can communicate as peers. Redirected traffic and ZFS memory traffic can both use the same switching architecture.
[0047] In various embodiments, many different configurations of the storage device 120 can be used. Figure 1 In the network. In some embodiments, the Oracle ZFS storage appliance family can be used. The ZFS storage appliance provides storage based on the Oracle Solaris kernel using Oracle's ZFS file system (“ZFS”) as described below. Processing core 114 handles any operations required to implement any selected data protection (e.g., mirroring, RAID-Z, etc.), data reduction (e.g., inline compression, replication, etc.), and any other implemented data services (e.g., remote replication). In some embodiments, the processing core may include 2.8 GHz The Xeon processor has 8x15 cores. The processing cores also handle cached data stored in DRAM and flash memory 112. In some embodiments, the DRAM / flash memory cache may include a 3TB DRAM cache.
[0048] In some configurations, storage unit 120 may include I / O port 116 to receive I / O requests from data storage controller 128. Each storage unit 120 may include an integral rack-mount unit with its own internal redundant power supply and cooling system. Concentrator board 110 or other similar hardware devices may be used to interconnect multiple storage devices. Active components such as memory boards, concentrator board 110, power supply, and cooling devices may be hot-swappable. For example, storage unit 120 may include flash memory 102, non-volatile RAM (NVRAM) 104, hard disk drives 105 in various configurations, RAID arrays 108 with drives, disk drives, etc. These storage units may be designed for high availability with hot-swappable storage cards and internal redundancy for power supply, cooling, and interconnection. In some embodiments, RAM may be made non-volatile by backing it up to dedicated flash memory in the event of a power outage. A mix of flash memory and NVRAM cards may be configurable, and both may use the same connectors and board profiles.
[0049] Although not explicitly shown, each I / O server 136 can perform global management processing or a data storage system manager, which can supervise the operation of the storage system in a pseudo-static "low-touch" manner, intervene when capacity must be reallocated among ZFS instances, for global flash wear leveling, for configuration changes, and / or for fault recovery. This "divide and conquer" strategy of allocating capacity among individual ZFS instances enables high scalability in performance, connectivity, and capacity. Additional performance can be achieved by horizontally adding more I / O servers 136 and then assigning less capacity to each ZFS instance and / or assigning fewer ZFS instances to each I / O server 136. Vertical scaling of performance using faster servers is also possible. Additional host ports can be added by filling available slots in the I / O servers 136 and then adding additional servers. Additional capacity can also be achieved by adding additional storage rigs 120 and allocating new capacity to new or existing ZFS instances.
[0050] Figure 2 The illustrations depict storage environments (including storage environments) according to certain embodiments of the present disclosure. Figure 1 An instance of a sample network file system 200 running in a storage environment. For example, file system 200 may include an Oracle ZFS file system (“ZFS”), which provides very large capacity (128-bit), data integrity, consistently consistent on-disk formatting, self-optimizing performance, and real-time remote replication. Among other things, ZFS differs from traditional file systems in that it does not require a separate volume manager. Instead, the ZFS file system shares a common pool of storage devices and acts as both the volume manager and the file system. Therefore, ZFS has full knowledge of the physical disks and volumes (including their condition, status, and logical arrangement within the volumes, along with all files stored on them). As file system capacity requirements change over time, devices can be added or removed from the pool to dynamically grow and shrink as needed without repartitioning the underlying storage pool.
[0051] In some embodiments, system 200 can interact with application 202 through an operating system. The operating system may include functionality for interacting with a file system, which in turn interfaces with a storage pool. The operating system typically interfaces with file system 200 via system call interface 208. System call interface 208 provides conventional file read, write, open, and close operations, as well as VFS-specific VNODE and VFS operations. System call interface 208 can act as the primary interface for interacting with ZFS, which serves as the file system. This layer resides between data management units (DMUs) 218 and presents a file system abstraction of the files and directories stored therein. System call interface 208 can bridge the gap between the file system interface and the underlying DMU 218 interface.
[0052] In addition to the POSIX layer of system call interface 208, the interface layer of file system 200 can also provide a distributed file system interface 210 for interacting with cluster / cloud computing devices 204. For example, it can provide... The interface provides a file system for computer clusters ranging in size from small workgroup clusters to large-scale multisite clusters. The volume emulator 212 also provides mechanisms for creating logical volumes that can be used as block / character devices. The volume emulator 212 not only allows client systems to distinguish between block and character, but also allows client systems to specify desired block sizes, thereby creating smaller, sparser volumes in a process known as "thin configuration." The volume emulator 212 provides raw access 206 to external devices.
[0053] Below the interface layer is the transaction object layer. This layer provides Intent Log 214, which is configured to record the transaction history for each dataset. This transaction history can be replayed in the event of a system crash. In ZFS, Intent Log 214 stores transaction records of system calls that modify the file system, along with sufficient information, in memory to enable the replay of these system calls. These are stored in memory until the DMU 218 commits them to the storage pool, and they can be discarded or flushed. In the event of power failure and / or disk failure, Intent Log 214 transactions can be replayed to keep the storage pool up-to-date and consistent.
[0054] The transaction object layer also provides an attribute processor 216, which can be used to implement directories within the POSIX layer of the system call interface 208 by performing arbitrary {key, value} associations within objects. The attribute processor 216 may include a module located on top of the DMU 218 and can operate on objects referred to as "ZAP objects" in ZFS. ZAP objects can be used to store dataset characteristics, navigate file system objects, and / or store storage pool characteristics. ZAP objects can take two forms: "microzap" objects and "fatzap" objects. Microzap objects can be lightweight versions of fatzap objects and provide a simple and fast lookup mechanism for a small number of attribute entries. Fatzap objects may be more suitable for ZAP objects containing a large number of attributes, such as larger directories, longer keys, longer values, etc.
[0055] The transaction object layer also provides a dataset and snapshot layer 220, which aggregates DMU objects in the hierarchical namespace and provides mechanisms for describing and managing relationships between the characteristics of the object set. This allows for characteristic inheritance, as well as the enforcement of quotas and reservations in the storage pool. DMU objects can include ZFS file system objects, clone objects, CFS volume objects, and snapshot objects. Therefore, the data and snapshot layer 220 can manage snapshots and clones.
[0056] A snapshot is a read-only copy of a file system or volume. A snapshot is a view of the file system at a specific point in time. ZFS snapshots are as useful as snapshots of some other file systems: by backing up snapshots, you have a consistent, unchanging target for backup programs to use. Snapshots can also be used to recover from recent errors by copying corrupted files from them. Snapshots can be created almost instantaneously, and they do not initially consume additional disk space within the pool. However, as data within the active dataset changes, snapshots consume disk space by continuing to reference old data, thus preventing disk space from being released. Blocks containing old data are only released when a snapshot is deleted. Taking snapshots is a timed operation. The existence of snapshots does not slow down any operations. Deleting a snapshot takes time proportional to the number of blocks that will be released and is very efficient. ZFS snapshots include the following characteristics: they remain unchanged across system reboots; the theoretical maximum number of snapshots is 2. 64They do not use separate backup repositories; they consume disk space directly from the same storage pool as the filesystem or volume that created them; recursive snapshots are created quickly as an atomic operation; and they are created together (all at once) or not at all. The advantage of atomic snapshot operations is that snapshot data is always taken at a consistent time, even in descendant filesystems. Snapshots are not directly accessible, but they can be cloned, backed up, rolled back, etc. Snapshots can be used to "roll back" in time to the point when the snapshot was taken.
[0057] A clone is a writable volume or filesystem whose initial content is identical to the dataset that created it. In ZFS systems, clones are always created from snapshots. Like snapshots, creating a clone is almost instantaneous and initially does not consume additional disk space. Furthermore, clones can be snapshotted. Clones can only be created from snapshots. When a snapshot is cloned, an implicit dependency is created between the clone and the snapshot. Even if the clone is created elsewhere in the dataset hierarchy, the original snapshot cannot be destroyed as long as the clone exists. Clones do not inherit the characteristics of the dataset that created them. A clone initially shares all its disk space with the original snapshot. As changes are made to the clone, it uses more disk space. Clones are useful for branching and for development or troubleshooting—and can be promoted to replace live filesystems. Clones can also be used to copy filesystems across multiple machines.
[0058] The DMU 218 presents a transaction object model built on top of a flat address space presented by the storage pool. The modules described above interact with the DMU 218 via object sets, objects, and transactions, where objects are collections of storage slices, such as data blocks, from the storage pool. Each transaction through the DMU 218 comprises a series of operations submitted as a group to the storage pool. This is the mechanism for maintaining disk consistency in the file system. In other words, the DMU 218 takes instructions from the interface layer and translates them into transaction batches. Instead of requesting data blocks and sending single read / write requests, the DMU 218 can combine these operations into batches of object-based transactions, which can be optimized before any disk activity occurs. Once this is complete, the transaction batches are handed over to the storage pool layer for scheduling and aggregation of the raw I / O transactions required to retrieve / write the requested data blocks. As described below, these transactions are written on a copy-on-write (COW) basis, eliminating the need for transaction journaling.
[0059] The storage pool layer, or simply "storage pool," can be referred to as a storage pool allocator (SPA). The SPA provides a public interface for manipulating storage pool configurations. These interfaces can create, destroy, import, export, and pool various storage media, and manage the namespace of the storage pool. In some embodiments, the SPA may include an adaptive replacement cache (ARC) 222, which acts as the central point for memory management for the SPA. Traditionally, ARC provides a basic Least Recently Used (LRU) object replacement algorithm for cache management. In ZFS, ARC 222 includes a self-adjusting cache that can be adjusted based on I / O workload. Furthermore, ARC 222 defines the Data Virtual Address (DVA) used by the DMU 218. In some embodiments, ARC 222 has the ability to evict memory buffers from the cache due to memory pressure to maintain high throughput.
[0060] The SPA may also include I / O pipes 224 or an "I / O manager" that translates the DVA from ARC 222 into logical locations within each virtual device (VDEV) 226 as described below. I / O pipes 224 drive dynamic striping, compression, checksum capabilities, and data redundancy across active VDEVs. Although not described in... Figure 2 As explicitly shown, I / O pipe 224 may include other modules that the SPA can use to read data from and / or write data to the storage pool. For example, I / O pipe 224 may include, but is not limited to, compression modules, encryption modules, checksum modules, and a metaslab allocator. For instance, a checksum can be used to ensure that data has not been corrupted. In some embodiments, the SPA may use the metaslab allocator to manage the allocation of storage space in the storage pool.
[0061] Compression is a process that typically reduces the size of a data block (which can be referred to interchangeably as a leaf node or data node) by utilizing redundancy within the data block itself. ZFS uses many different compression types. When compression is enabled, less storage can be allocated to each data block. The following compression algorithms can be used: LZ4 – an algorithm added after the creation of the feature flag. It is significantly superior to LZJB. LZJB is the original default compression algorithm used for ZFS. It was created to meet the expectations of compression algorithms suitable for file systems. Specifically, it provides fair compression, high compression speed, high decompression speed, and fast detection of incompressible data. GZIP (1 through 9 are implemented in the classic Lempel-Ziv implementation). It provides high compression, but it often makes IO CPU-bound. ZLE (Zero-Length Encoding) – a very simple algorithm that only compresses zeros. In each of these cases, there is a trade-off between the compression ratio and the amount of latency involved in compressing and decompressing data blocks. Generally, the more data is compressed, the longer the compression and decompression time will be.
[0062] Encryption adds end-to-end security to data blocks by cryptographically encoding them with a key. Only the user with the key can decrypt the data block. When used in a ZFS system, ZFS pools can support a mix of encrypted and unencrypted ZFS datasets (filesystems and ZVOLs). Data encryption is completely transparent to applications and provides a highly flexible system for protecting data at rest, requiring no application changes or qualifications. Furthermore, ZFS encryption randomly generates local encryption keys from a passphrase or AES key, and all keys are stored locally on the client—not in the cloud object repository 404 as in traditional filesystems. When enabled, encryption is transparent to applications and storage in the cloud object repository 404. ZFS makes encrypting and managing data encryption easy. You can have both encrypted and unencrypted filesystems in the same storage pool. You can also use different encryption keys for different systems, and you can manage encryption locally or remotely—but the randomly generated encryption keys always remain local. ZFS encryption is inheritable for descendant filesystems. In CCM and GCM operating modes, AES (Advanced Encryption Standard) with key lengths of 128, 192, and 256 is used to encrypt data.
[0063] Deduplication is the process of identifying data blocks that are already stored as existing data blocks on the file system and pointing to those existing data blocks instead of storing the data blocks again. ZFS provides block-level deduplication because this is the finest granularity that makes sense for general-purpose storage systems. Block-level deduplication also naturally maps to ZFS's 256-bit block checksum, which provides a unique block signature for all blocks in the storage pool, provided the checksum is cryptographically strong (e.g., SHA256). Deduplication is synchronous and is performed when a data block is sent to the cloud object repository 404. If the data blocks are not replicated, enabling deduplication will increase overhead without any benefit. If duplicate data blocks exist, enabling deduplication will save both space and improve performance. The space savings are obvious; the performance improvement is due to the elimination of storage writes when storing duplicate data and the reduced memory footprint due to many applications sharing the same memory pages. Most storage environments contain a mixture of mostly unique data and mostly replicated data. ZFS deduplication is per dataset and can be enabled when it may be helpful.
[0064] In ZFS, a storage pool can consist of a collection of Virtual Devices (VDEVs). In some embodiments, at least a portion of a storage pool can be represented as a self-describing Merkle tree, i.e., a logical tree where both data and metadata are stored through VDEVs of the logical tree. There are two types of virtual devices: physical virtual devices called leaf VDEVs, and logical virtual devices called internal VDEVs. Physical VDEVs can include writable media block devices, such as hard disks or flash drives. Logical VDEVs are conceptual groupings of physical VDEVs. VDEVs can be arranged in a tree, with physical VDEVs existing as leaves of the tree. A storage pool can have a special logical VDEV called the "root VDEV," which is the root of the tree. All direct child nodes (physical or logical) of the root VDEV are referred to as "top-level" VDEVs. Generally, VDEVs implement data replication, mirroring, and architectures (such as RAID-Z and RAID-Z2). Each leaf VDEV represents one or more physical storage devices 228 that actually store the data provided by the file system.
[0065] In some embodiments, file system 200 may include an object-based file system where data and metadata are stored as objects. More specifically, file system 200 may include functionality to store data and corresponding metadata in a storage pool. Requests to perform a specific operation (i.e., a transaction) are forwarded from the operating system to DMU 218 via system call interface 208. DMU 218 directly translates requests to perform operations on objects into requests to perform read or write operations at a physical location within the storage pool (i.e., I / O requests). SPA receives the request from DMU 218 and writes blocks to the storage pool using a COW procedure. COW transactions can be performed for write requests to data on a file. Instead of overwriting existing blocks during a write operation, the write request causes a new fragment to be allocated for the modified data. Therefore, retrieved data blocks and corresponding metadata are never overwritten until a modified version of the data blocks and metadata is committed. Thus, DMU 218 writes all modified data blocks to unused fragments within the storage pool and subsequently writes the corresponding block pointers to unused fragments within the storage pool. To complete the COW transaction, SPA issues an I / O request to reference the modified data blocks.
[0066] Figures 3A-3D The illustration depicts a Copy-on-Write (COW) process for a file system (such as file system 200) according to certain embodiments of this disclosure. For example, the ZFS system described above uses a COW transaction model where all block pointers within the file system can contain a 256-bit checksum of the target block, which is verified when the block is read. As described above, the block containing active data is not overwritten in-place. Instead, a new block is allocated, the modified data is written to the new block, and any metadata blocks referencing it are simply read, reallocated, and rewritten.
[0067] Figure 3A A simplified diagram of a file system storage of data and metadata corresponding to one or more files as a logical tree 300, according to some embodiments, is illustrated. Logical tree 300, and other logical trees described herein, can be self-describing Merkle trees, where data and metadata are stored as blocks of logical tree 300. Root block 302 may represent the root or "superblock" of logical tree 300. Logical tree 300 can traverse files and directories by navigating through each child node 304, 306 of root 302. Each non-leaf node represents a directory or file, such as nodes 308, 310, 312, and 314. In some embodiments, a hash of the values of its child nodes may be assigned to each non-leaf node. Each leaf node 316, 318, 320, 322 represents a data block of a file.
[0068] Figure 3BThe diagram illustrates an example of logical tree 300-1 after the initial phase of a write operation. In this example, data blocks represented by nodes 324 and 326 have been written by file system 200. Instead of overwriting the data in nodes 316 and 318, new data blocks are allocated for nodes 324 and 326. Therefore, after this operation, the old data in nodes 316 and 318 remains in memory along with the new data in nodes 324 and 326.
[0069] Figure 3C The diagram illustrates an example of logical tree 300-2 as a write operation continues. To reference newly written data blocks in nodes 324 and 326, file system 200 determines nodes 308 and 310 that reference old nodes 316 and 318. New nodes 328 and 330 are assigned to reference the new data blocks in nodes 324 and 326. The same process is repeated recursively upwards through the file system hierarchy until every node referencing the changed nodes is reassigned to point to the new node.
[0070] When a pointer block is allocated in a new node in the hierarchy, the address pointer in each node is updated to point to the new location of the allocated child node in memory. Furthermore, each data block includes a checksum, which is calculated using the data block referenced by the address pointer. For example, the checksum in node 328 is calculated using the data block in node 324. This arrangement means that the checksum is stored separately from the data block from which it is calculated. This prevents so-called "ghostwrite," in which new data is never written, but the checksum stored with the data block will indicate that the block is correct. The integrity of the logic tree 300 can be quickly verified by traversing the logic tree 300 and calculating the checksum at each level based on the child nodes.
[0071] To complete the write operation, root 302 can be reallocated and updated. Figure 3D The illustration shows an example of logic tree 300-3 at the end of a write operation. When root 302 is ready to be updated, a new superblock root 336 can be allocated and initialized to point to the newly allocated child nodes 332 and 334. Then, root 336 can become the root of logic tree 300-3 in an atomic operation to finally determine the state of logic tree 300-3.
[0072] A snapshot is a read-only copy of a file system or volume. A snapshot is a view of the file system at a specific point in time. ZFS snapshots are as useful as snapshots of some other file systems: by backing up snapshots, you have a consistent, unchanging target for backup programs to use. Snapshots can also be used to recover from recent errors by copying corrupted files from them. Snapshots can be created almost instantaneously, and they do not initially consume additional disk space within the pool. However, as data within an active dataset changes, snapshots consume disk space by continuing to reference old data, thus preventing disk space from being released. Blocks containing old data are only released when a snapshot is deleted. Taking snapshots is a timed operation. The existence of snapshots does not slow down any operations. Deleting a snapshot takes time proportional to the number of blocks that will be released and is very efficient. ZFS snapshots include the following characteristics: they remain unchanged across system reboots; the theoretical maximum number of snapshots is 2^64; they do not use a separate backup repository; they consume disk space directly from the same storage pool as the file system or volume that created them; recursive snapshots are created quickly as an atomic operation; and they are created together (all at once) or not at all. The advantage of atomic snapshot operations is that snapshot data is always taken at a consistent time, even in descendant file systems. Snapshots are not directly accessible, but they can be cloned, backed up, rolled back, and so on. Snapshots can be used to "roll back" in time to the point when they were taken. Figure 3E An example of a snapshot data service in ZFS is depicted, where snapshots are... Figures 3A-3D The COW process described herein takes place before root block 336 is set as the new live root block. The live root block is the root block from which the next data progression will begin during the COW execution. Snapshot roots and "live" roots are shown. The live root is the root that will be operated on in the next storage operation. All blocks (302-322) pointed to by the snapshot root are set to "read-only," meaning they are placed on a list of blocks that cannot be freed for further use by the storage system until the snapshot is deleted.
[0073] It should be noted that some terms can be used interchangeably throughout the application. For example, leaf node, leaf block, and data block can be the same in some cases, especially when referring to a local tree instance. Similarly, non-leaf node, metadata, and metadata block can be used interchangeably in some cases, especially when referring to a local tree instance. Likewise, root node and root block can be used interchangeably in some instances, especially when referring to a local tree instance. Furthermore, it should be noted that references to leaf node, non-leaf node, and root node can be similarly applied to cloud storage objects corresponding to a cloud version of a logical tree that is at least partially based on a local tree instance.
[0074] When blocks are created, they are given a "birth" time, which represents the iteration or progress of the COW (Cost-of-Work) process that created the block. Figure 3F This idea was demonstrated. Figure 3F In, as shown in 365— Figure 3A The birth time of 19.366 is shown. Figure 3D Birth time 25.367 shows the new tree created by blocks 372, 378, 379, 384, 386, and 392, representing birth time 37 and indicating data transactions on the tree 12 iterations after birth time 25. Therefore, rolling back or backing up to a snapshot will only leave data such as... Figure 3A The blocks are shown. Therefore, using a birth-time hierarchy, the ZFS system can generate and roll back from a snapshot of the root of the tree to any point in the birth time of the entire tree structure. In essence, this allows all new blocks with a birth time after a snapshot to be made available in the storage pool, as long as they are not linked to any other snapshots or metadata blocks.
[0075] A clone is a writable volume or file system whose initial content is identical to the dataset that created it. In ZFS systems, clones are always created from snapshots. Like snapshots, creating a clone is almost instantaneous and initially does not consume additional disk space. Furthermore, clones can be snapshotted. Clones can only be created from snapshots. When a snapshot is cloned, an implicit dependency is created between the clone and the snapshot. Even if the clone is created elsewhere in the dataset hierarchy, the original snapshot cannot be destroyed as long as the clone exists. A clone does not inherit the characteristics of the dataset that created it. A clone initially shares all its disk space with the original snapshot. As changes are made to the clone, it uses more disk space. Clones are useful for branching and for development or troubleshooting—and can be promoted to replace live file systems. Clones can also be used to copy file systems across multiple machines.
[0076] The embodiments described herein can be found above. Figure 1 This is implemented in the system described in -3. For example, the system may include... Figure 1 The system includes various storage devices, switching circuits, and / or one or more processors in the server. Instructions may be stored in one or more memory devices of the system, causing one or more processors to perform various operations that affect the functionality of the file system. The steps of the various methods may be defined by... Figure 1-2 The system's processor, memory devices, interfaces, and / or circuitry execute.
[0077] Now go to Figure 4 , Figure 4This is a high-level diagram illustrating an example of a hybrid cloud storage system 400 according to certain embodiments of the present disclosure. The hybrid cloud storage system 400 can transform a network file system (such as the ZFS file system) into a cloud-enabled file system where the file system's functionality (including file system data services) is layered on a cloud object repository remote from the file system. As depicted in the diagram, the hybrid cloud storage system 400 may include a network file system 200 (also referred to herein as "local file system 200"). The local file system 200 may be communicatively coupled to a cloud object storage 404. In some embodiments, the cloud object storage 404 may be connected to... Figure 2 The cluster / cloud 204 indicated in the diagram corresponds to this. The local file system 200 can be communicatively coupled to the cloud object storage 404 via the cloud interface device 402. The cloud interface device 402 can be used by the local file system 200 as an access point to the cloud object storage 404.
[0078] Hybrid cloud storage system 400 provides a solution to overcome the traditional limitations of cloud object storage. Traditional cloud object protocols are limited by restricted data / object access semantics. Traditionally, cloud object storage has limited interfaces and primitives and is incompatible with POSIX. For example, once an object is written, it cannot be modified afterward; it can only be deleted and replaced with a newly created object. As another example, traditional cloud object storage has namespace limitations, which simplify the namespace and limit it to the top-level container. However, hybrid cloud storage system 400 not only enables data migration to and from cloud object repository 404, but also layers the file system functionality of local file system 200 on top of the cloud object interface to cloud object repository 404 to provide cloud-based storage.
[0079] The local file system 200 can be configured for POSIX interface and semantics. For example, the local file system 200 can provide users with access to data as files, allowing modification of file contents without rewriting the files. The local file system 200 can also provide organization of data in a name hierarchy, as is typical for the ZFS file system. All the functionality of the ZFS file system is available to users of the local file system 200. The cloud interface appliance 402 can allow for layering of file system semantics on top of the cloud object protocol—for example, providing the ability to construct namespaces, create files, create directories, etc.—and such capabilities relative to data migration to and from cloud object storage 404. The cloud interface appliance 402 can facilitate plug-and-play object storage solutions to improve upon the local file system 200 while supporting ZFS file system data services.
[0080] Cloud interface apparatus 402 can be configured to provide an object API (Application Programming Interface). In some embodiments, cloud interface apparatus 402 can be configured to use multiple API transformation profiles. According to some embodiments, API transformation profiles can integrate modules and functions (e.g., data services and modules), POSIX interfaces and semantics, and other components that may not be natively designed to interact with cloud storage. In some embodiments, API transformation profiles can (e.g., via API calls) transform the protocols, formats, and routines of file system 200 to allow interaction with cloud data repository 404. Information used for such integration can be stored in an API transformation data repository that may be co-located with or otherwise communicatively coupled to cloud interface apparatus 402. Cloud interface apparatus 402 can utilize this information to cohesively integrate POSIX interfaces and semantics to interface with cloud data repository 404 while preserving these semantics.
[0081] The hybrid cloud storage system 400 allows the local file system 200 to use the cloud object storage 404 as a "drive". In various cases, file 410 can be stored as a data object with metadata objects, and / or as a data block with associated metadata. The cloud interface device 402 can receive file 410 from and transfer file 410 to the local file system 200. In various embodiments, the local file system 200 can receive and / or send file 410 via NFS (Network File System) protocol, SMB (Server Message Block Protocol), etc. In some embodiments, the cloud interface device 402 can transform file 410 into object 412. The transformation of file 410 may include transforming data blocks and associated metadata and / or data objects associated with metadata objects, either of which may correspond to file 410. In some embodiments, the transformation may include the cloud interface device 402 performing API transformation using multiple API transformation profiles. According to some embodiments, the transformation may include the cloud interface device 402 extracting data and / or metadata from file 410, objects, and / or blocks. The cloud interface device 402 can at least partially transform files 410, objects, and / or blocks into cloud storage objects using extracted data. In some embodiments, the cloud interface device 402 can utilize the extracted data to create corresponding cloud storage objects, wherein the extracted data is embedded in a placement request pointing to the cloud object repository 404. Similarly, utilizing any transformation implemented by the cloud interface device 402 to interface to the cloud data repository 404, in some embodiments, the cloud interface device 402 can reverse the transformation process to interface with local components of the local file system 200.
[0082] The cloud interface device 402 can send and receive objects 412 to and from the cloud object storage 404. In some embodiments, the cloud interface device 402 can send and receive objects 412 via HTTPS or the like. In some embodiments, the cloud interface device 402 can be co-located with the local file system 200. In other embodiments, the cloud interface device 402 can be located remotely from the local file system 200, such as alongside at least some devices facilitating the cloud object storage 404 or at some other communication-coupled site.
[0083] As further disclosed herein, files in local file system 200 can be stored as "disk blocks," with the data objects and metadata corresponding to the files stored as virtual storage blocks of logical tree 300 (e.g., a self-describing Merkle tree, where data and metadata are stored as blocks). Cloud interface equipment 402 can create a mapping 406 directly from each logical block in the data tree 300 to a cloud object 414 in cloud object repository 404. Some embodiments may employ a one-to-one block-to-object mapping. Additionally or alternatively, other embodiments may employ any other suitable block-to-cloud-object ratio, for example, mapping multiple blocks to a single cloud object. In some instances of such embodiments, the entire logical tree of blocks may be mapped to a single cloud object. In other cases, only a portion of the logical tree of blocks may be mapped to a single cloud object.
[0084] In some embodiments, address pointers are updated when a block is converted into a cloud object. When a block is converted into a cloud object using a one-to-one block-to-object conversion scheme, the address pointers can be updated so that non-leaf cloud objects in the hierarchy point to sub-cloud objects in the cloud object repository 404. For example, the address pointers can correspond to the object name and path specification of the sub-cloud object, which may include parameters such as object name and bucket specification. Accordingly, some embodiments can convert blocks of logic tree 300 into cloud objects of logic tree 300A. Using some embodiments, this conversion enables cloud interface equipment 402 to traverse the logic tree 300A of cloud objects using the address pointers of the cloud objects.
[0085] In some embodiments, when a block is converted into a cloud object using a many-to-one block-to-object conversion scheme such that a portion of the logic tree 300 is converted into a cloud object, the address pointers of the cloud objects, including logic 300A, can be similarly updated, but at a coarser level, so that non-leaf cloud objects in the hierarchy point to child cloud objects in the cloud object repository 404. This conversion allows the cloud interface apparatus 402 to traverse the logic tree 300A of cloud objects using the address pointers of the cloud objects in a coarser-grained but faster manner than the traversal facilitated by the one-to-one block-to-object conversion scheme. Furthermore, in some embodiments, the conversion process can be used to update checksums. Checksums for individual cloud objects can be updated and stored separately in the parent cloud object. In conversions employing a many-to-one block-to-object conversion scheme, a single checksum can be computed for cloud objects corresponding to a set of blocks.
[0086] Accordingly, the implementation of mapping 406 allows communication to the cloud object repository 404 via one or more networks, and the interface with the cloud object repository 404 can be object-based rather than block-based. As further disclosed herein, utilizing the cloud interface apparatus 402 between the local file system 200 and the cloud object repository 404, the hybrid cloud storage system 400 can possess characteristics and failure modes different from the traditional ZFS file system. The cloud interface apparatus 402 can translate the file system interface of the local file system 202 on the client side and can be coordinated to the cloud object repository 404 via object protocols to read and write data. Through the cloud interface apparatus 402, the cloud object 414 can remain accessible by the local file system 200 via NFS, SMB, etc.
[0087] Using cloud objects 414 as mappings 406 to logical blocks, sets of cloud objects 414 can be grouped to form drives that host ZFS storage pools as self-contained collections. Drive content can be elastic, allowing cloud objects to be created only for logical blocks that have been allocated. In some embodiments, cloud interface equipment 402 can have the ability to assign variable object sizes (e.g., for different data types) to allow for greater storage flexibility. Data is not limited to a specific byte size. Storage size can be scaled as needed by modifying metadata. In some embodiments, cloud-based pools can be imported on any server. Once imported, the cloud-based pool can behave as local storage, and all ZFS services are supported for the cloud-based pool. Cloud-based pools can be designated as new types of storage pools. However, from a user's perspective, data from a cloud-based pool may appear indistinguishable from that from a local pool.
[0088] Figure 5An example network file system 200-1 of a hybrid cloud storage system 400 according to certain embodiments of this disclosure is illustrated. File system 200-1 may correspond to file system 200, but cloud device management is directly integrated into the ZFS control stack. In addition to the disclosure regarding file system 200, file system 200-1 may also include a cloud interface device 502 that facilitates full utilization of the cloud object repository 404 as the storage medium for file system 200-1. Cloud interface device 502 may facilitate cloud drives, at least in part, by mapping cloud storage to device abstractions.
[0089] In some embodiments, cloud interface device 502 may correspond to one or more VDEVs of another VDEV type within the ZFS file system architecture. ZFS may communicate directly with cloud interface device 502. Cloud interface device 502 may reside at a virtual device layer directly above the driver layer of file system 200-1. Some embodiments of cloud interface device 502 may correspond to an abstraction of a device driver interface within the ZFS architecture. Other components of file system 200-1 may communicate with cloud interface device 502 as if it were another VDEV of another device type (such as VDEV 226). To enable the transfer of a larger amount of information through cloud interface device 502 compared to other VDEVs 226, the interface associated with cloud interface device 502 may be wider to allow more information to be transferred outward through I / O pipe 224 and cloud interface device 502 to cloud object data repository 404.
[0090] In some embodiments, cloud interface device 502 can translate file system interfaces on a client. In some embodiments, to provide full POSIX file system semantics, cloud interface device 502 can translate file system interface requests into object interface requests for cloud object repository 404. In some embodiments, cloud interface device 502 may be able to communicate with cloud object repository 404 via object protocols to read and write data.
[0091] Figure 6 This is a diagram illustrating additional aspects of a cloud interface device 402-1 of a hybrid cloud storage system 400-1 according to certain embodiments of the present disclosure. As indicated in the depicted examples, some embodiments of the cloud interface device 402 may include a virtual storage pool 602 and a cloud interface daemon 604. The virtual storage pool 602 may reside at the kernel of the file system 200-1, and the cloud interface daemon 604 may reside in the user space of the file system 200-1. In various embodiments, the cloud interface daemon 604 may interact with... Figure 5 The cloud interface component corresponding to application 202 and / or cluster / cloud 204 as indicated in the instructions.
[0092] In some embodiments, the virtual storage pool 602 may include at least one cloud interface device 502, an intent log 214-2, and a cache 222-1. (The above refers to...) Figure 1-2 And below, regarding Figure 7, we describe intent log 214-2 and cache 222-1. Cloud interface device 502 can interact with cloud interface daemon 604 to coordinate operations regarding cloud object data repository 404, at least in part based on mapping 406. Cloud interface daemon 604 may include cloud client interface 608 to interface with cloud object data repository 404. In some implementations, cloud client interface 608 may include an endpoint providing Swift / S3 compatibility with cloud object data repository 404. Operations of cloud client interface 608 may be at least in part based on retrieving and placing entire data objects 412 to facilitate read and write access to cloud object data repository 404.
[0093] Go back to reference Figure 4 and 5 In some embodiments, requests to perform one or more transactions for one or more files can be received from the application 202 at the application layer of the file system 200-1 via system call interface 208 of the interface layer of the file system 200-1. The requests may be POSIX compliant and may be translated by one or more components of the file system 200-1 into one or more object interface requests to perform one or more operations on a cloud-based instantiation 300A of a logical tree 300 stored in the cloud object repository 404. For example, in some embodiments, the cloud interface apparatus 402 may translate a POSIX compliant request or an intermediate request resulting from a POSIX compliant request into a corresponding object interface request. In some embodiments, the DMU 218 may translate a POSIX compliant request into an I / O request to perform an I / O operation, and the cloud interface apparatus 402 may translate the I / O request into a corresponding object interface request, thereby coordinating the object interface requests using mapping 406.
[0094] In some cases, a transaction can correspond to an operation that causes a file to be stored locally. In some instances, the file system 200-1 can store data objects and their corresponding metadata in a system storage pool 416 provided by one or more physical storage devices 228 in VDEV 226. A data object can correspond to one or more files. As disclosed above, the data objects and metadata corresponding to one or more files can be stored as a logical tree 300. Therefore, the storage of the logical tree 300 can be localized in the system storage pool 416.
[0095] In further operations according to some embodiments, the file system 200-1 may enable the storage of the data object of the logical tree 300 and its corresponding metadata in the cloud object repository 404. While in some embodiments the logical tree 300 may be initially stored in a local storage pool before being migrated to cloud storage, in other embodiments the logical tree 300 may not be stored in a local storage pool before being stored in the cloud object repository 404. For example, some embodiments may create at least a portion of the logical tree 300 in a cache and then migrate it to the cloud object repository 404. Therefore, it should be appreciated that various embodiments are possible.
[0096] To store the data objects and corresponding metadata of the logical tree 300 in the cloud object repository 404, the cloud interface device 502 can create a mapping 406 from each logical block in the logical tree 300 to the corresponding cloud object 414 in the cloud object repository 404. In some embodiments, the DMU 218 can read data from the system storage pool 416 (e.g., from a local pool of RAIDZ or RAIDZ2) to provide the data to the cloud interface device 502 as the basis for creating the mapping 406. In some embodiments, the cloud interface device 502 can communicate directly or indirectly with another VDEV 226 to read data as the basis for the mapping 406. In some embodiments, the mapping 406 can map objects directly to blocks represented in a physical drive. The mapping 406 can be more granular than mapping file portions to objects; it can be mapped at a lower level. Accordingly, the mapping 406 can be a per-object mapping 406. The mapping 406 can map virtual storage blocks to objects in the cloud object repository 404, such that the logical tree 300 is represented in the cloud object repository 404 as shown in logical tree 300A. When the local file system 200-1 interfaces with data object 416 in the cloud object repository 404, the logic tree 300A conforms to a new device type that the local file system 200-1 can communicate with.
[0097] Mapping 406 can be updated using each I / O operation on cloud object repository 404 or only using write operations. In some embodiments, mapping 406 may include an object directory that indexes all cloud objects 414. Cloud object state can be maintained in indexes, tables, index-organized tables, etc., which can be indexed on a per-object basis. In some embodiments, mapping 406 may include an object directory that indexes only some of the cloud objects 414. For example, such embodiments may only index the cloud object 414 corresponding to the superblock. In some embodiments, the cloud object state of each object relative to each leaf path can be indexed. In some embodiments, the object directory may reference cloud objects 414 via address pointers, which may correspond to object names and path specifications. In some embodiments, the object directory may reference cloud objects 414 via URLs.
[0098] The cloud object state indexed in mapping 406 can be used to route object requests. Using the index, cloud interface device 502 can request cloud objects at least partially based on superblocks. According to a first approach, such a request might require requesting a set of cloud objects associated with a specific superblock so that the entire logical tree 300A is represented by the set of cloud objects transmitted in response to the request. According to a second approach, such a request might require iterative requests for a subset of cloud objects associated with a specific superblock to iteratively traverse the entire logical tree 300A until one or more desired cloud objects are read from the cloud object repository 404. In some embodiments, cloud interface device 502 can selectively use one of these two approaches, at least partially based on the size of the cloud objects representing the various logical trees 300A. For example, cloud interface device 502 can use one approach when the size of a cloud object is less than an aggregate size threshold, and transition to another approach when the size of a cloud object meets or exceeds the aggregate size threshold.
[0099] Some embodiments may employ an alternative approach, where the object catalog indexes cloud objects on a per-object basis and can be used to directly request cloud objects without tree traversal at the cloud level. Some embodiments may maintain a local snapshot of the metadata of the logical tree 300A. Such embodiments can utilize the local snapshot to request cloud objects directly or indirectly. Furthermore, some embodiments may maintain a checksum of the logical tree 300A in the object catalog or a local snapshot, which can be used to verify cloud objects retrieved from the cloud data repository 404.
[0100] Therefore, file system 200-1 can maintain a tree of data and map that tree to cloud object repository 404. The namespace of tree 300A can correspond to the metadata stored within the nodes of tree 300A. File system 200-1 can continue to use a hierarchical tree representation, but map the hierarchical tree representation to the cloud object repository as a way of storing data.
[0101] Refer again Figure 6 To perform I / O operations on the cloud object repository 404, the cloud interface device 502 can send requests to the cloud interface daemon 604. For example, in some implementations, the cloud interface device 502 can send requests to the cloud interface daemon 604 through the transaction object layer and interface layer of the file system 200-1. The requests sent by the cloud interface device 502 can be based at least in part on POSIX-compliant requests received via application 202 and / or at least in part on I / O requests created by DMU 218 (e.g., in response to POSIX-compliant requests), which the cloud interface device 502 can translate into requests to the cloud interface daemon 604.
[0102] In some embodiments, a request sent by the cloud interface device 502 to the cloud interface daemon 604 can be transformed into a retrieve request and a place request for the cloud client interface 608. In some embodiments, the request sent by the cloud interface device 502 can be a retrieve request and a place request; in other embodiments, the cloud interface daemon 604 can transform the request sent by the cloud interface device 502 into a retrieve request and a place request. In any case, in response to a request sent by the cloud interface device 502, the cloud interface daemon 604 can communicate with the cloud object repository 404 via one or more object protocols on the network to perform corresponding I / O operations for the data object 414.
[0103] For example, communication to the cloud object repository 404 may include, at least in part, specifying the storage of data objects and corresponding metadata of the logical tree 300A in the cloud object repository 404 based on mapping 406. In some embodiments, this communication may, for example, specify different object sizes for different data types. Thus, the cloud interface apparatus 402 may specify a certain object size to store some data objects identified as having a certain data type, and may specify different object sizes to store other data objects identified as having different data types.
[0104] Figure 7A This is a block diagram illustrating an example method 700 according to certain embodiments of the present disclosure, which addresses certain features of COW processing in a hybrid cloud storage system 400. According to some embodiments, method 700 may begin as indicated by block 702. However, the teachings of this disclosure can be implemented in various configurations. Accordingly, the order of certain steps, including method 700 and / or other methods disclosed herein, can be shuffled or combined in any suitable manner and may depend on the chosen implementation. Moreover, while the steps may be separated for the purpose of description, it should be understood that some steps may be performed simultaneously or substantially simultaneously.
[0105] As indicated by box 702, one or more POSIX-compliant requests to perform one or more specific operations (i.e., one or more transactions) can be received from application 202. Such operations may correspond to writing and / or modifying data. As indicated by box 704, one or more POSIX-compliant requests may be forwarded from the operating system to DMU 218 via system call interface 208. In various embodiments, transactions implemented through DMU 218 may include a series of operations submitted as a group to one or both of system storage pool 416 and cloud object repository 404. These transactions may be written based on COW.
[0106] As indicated in box 706, DMU 218 can directly translate requests to perform operations on data objects into requests to perform write operations (i.e., I / O requests) on physical locations within system storage pool 416 and / or cloud object repository 404. In some modes, these operations can be performed first on locally stored data objects, and then changes to the data objects can be propagated to the corresponding cloud storage data objects by either DMU 218 instructing the specific changes or by DMU 218 instructing cloud interface equipment 402 to read these changes directly or indirectly using another VDEV 226. In other modes, operations can be performed simultaneously or substantially simultaneously on locally stored data objects and their corresponding cloud storage data objects. In still other modes, operations can be performed only on cloud storage data objects. For example, some implementations may not have a local tree 300 and may only have a cloud-based version of logical tree 300A. Various embodiments can be configured to allow a user to select one or more modes.
[0107] As indicated in box 708, the SPA can receive I / O requests from DMU 218. And, in response to a request, the SPA can initiate a COW procedure to write a data object to system storage pool 416 and / or cloud object repository 404. In the mode where the write of the data object is performed on the locally stored data object before or simultaneously with the write to cloud object repository 404, as indicated in box 710, the above (e.g., given...) Figures 3A-3D The publicly disclosed COW process can continue relative to system storage pool 416.
[0108] As indicated in box 712, cloud interface device 402 can receive I / O requests and identify incremental modifications to the logical tree 300A. Incremental modifications can correspond to a new tree portion generated by COW processing to fulfill a write request. Cloud interface device 402 can translate I / O requests into corresponding object interface requests. Using modified data from I / O requests or from reading changes to data objects in local storage, cloud interface device 502 can coordinate object interface requests using mappings 406 of cloud storage objects 414 in the cloud object repository 404.
[0109] For example, in some embodiments, incremental modifications can be determined at least in part based on changes to locally stored data objects to reflect changes to logic tree 300. In some instances, cloud interface equipment 402 can read logic tree 300 or read at least the changes to logic tree 300 to determine incremental modifications. This data can be passed to cloud interface equipment 402 by another component of the file system, such as DMU 218 or mirror VDEV. In some embodiments, incremental modifications can be transmitted to cloud interface equipment 402. However, in some embodiments, cloud interface equipment 402 can determine incremental modifications at least in part based on write requests analyzed in view of a copy or snapshot of logic tree 300 and / or a snapshot of logic tree 300A. For embodiments where cloud interface equipment 402 uses a snapshot of logic tree 300A, in some embodiments, the snapshot can be retained in mapping 406 or otherwise stored in memory and / or the physical layer.
[0110] For more specific details, please refer to [link / reference]. Figure 7A As indicated in box 714, determining incremental modifications may include creating new leaf nodes (e.g., leaf nodes 324, 326). After the initial phase of the write operation, new data blocks (i.e., leaf nodes) have been allocated in memory, and the data for each write operation has been written to the new data blocks by the cloud interface equipment 402, while the previous data and data blocks are also retained in memory.
[0111] As indicated in box 729, after creating the leaf nodes (data blocks) in box 714, it is determined whether any ZFS data services have been enabled or requested. These include, but are not limited to, compression, encryption, deduplication, snapshots, and cloning. If no such data services are needed—then, as indicated in box 716—new non-leaf nodes can be created (e.g., non-leaf nodes 326, 330). As the write operation continues, cloud interface equipment 402 can determine the non-leaf nodes of previous versions of the referenced nodes. To reference the newly written data, the new non-leaf nodes are assigned to reference the new data blocks in the leaf nodes. As reflected in snapshot 301, the same process can be repeated recursively upwards through the hierarchy of logical tree 300A until each non-leaf node referencing the changed node is reassigned to point to the new node. As pointer blocks are assigned to new nodes in the hierarchy, the address pointers in each node can be updated to point to the new location of the assigned child nodes in memory. As indicated in box 718, to complete the write operation, the root node (e.g., root node 336) can be reassigned and updated. When you are ready to update the root node, you can allocate and initialize a new root node (superblock) to point to the newly allocated child nodes under the new root node.
[0112] As part of the write operation, checksums are used to update the metadata of all parts of the tree involved in the transaction write operation. Each node created includes a checksum, which is calculated using the node referenced by the address pointer. This arrangement means that the checksum is stored separately from the node from which it is calculated. By storing the checksum of each node in a pointer to its parent node instead of in the node itself, each node in the tree contains checksums for all its child nodes. About Figures 3A-3D The example further illustrates this point. By doing so, each tree is automatically self-verified, and there is always a possibility of detecting inconsistencies with read operations.
[0113] If data service is needed at box 729, then proceed to the next box 719. Figure 7B Top box 730. Data services include, but are not limited to, compression, encryption, deduplication (which must be done in that order), snapshots, and cloning.
[0114] Compression is a process that typically reduces the size of a data block (which can be referred to interchangeably as a leaf node or data node) by utilizing redundancy within the data block itself. ZFS uses many different compression types. When compression is enabled, less storage can be allocated to each data block. The following compression algorithms can be used: LZ4 – an algorithm added after the creation of the feature flag. It is significantly superior to LZJB. LZJB is the original default compression algorithm used for ZFS. It was created to meet the expectations of compression algorithms suitable for file systems. Specifically, it provides fair compression, high compression speed, high decompression speed, and fast detection of incompressible data. GZIP (1 through 9 are implemented in the classic Lempel-Ziv implementation). It provides high compression, but it often makes IO CPU-bound. ZLE (Zero-Length Encoding) – a very simple algorithm that only compresses zeros. In each of these cases, there is a trade-off between the compression ratio and the amount of latency involved in compressing and decompressing data blocks. Generally, the more data is compressed, the longer the compression and decompression time will be.
[0115] Encryption adds end-to-end security to data blocks by cryptographically encoding them with a key. Only the user with the key can decrypt the data block. When used in a ZFS system, a ZFS pool can support a mix of encrypted and unencrypted ZFS datasets (filesystems and ZVOLs). Data encryption is completely transparent to applications and provides a highly flexible system for protecting data at rest, requiring no application changes or qualifications. Furthermore, ZFS encryption randomly generates local encryption keys from a passphrase or AES key, and all keys are stored locally on the client—not in the cloud object repository 404 as in traditional filesystems. When enabled, encryption is transparent to applications and storage in the cloud object repository 404. ZFS makes encrypting and managing data encryption easy. You can have both encrypted and unencrypted filesystems in the same storage pool. You can also use different encryption keys for different systems, and you can manage encryption locally or remotely—but the randomly generated encryption keys always remain local. ZFS encryption is inheritable for descendant filesystems. In CCM and GCM operating modes, AES (Advanced Encryption Standard) with key lengths of 128, 192, and 256 is used to encrypt data.
[0116] Deduplication is the process of identifying data blocks that are already stored as existing data blocks on the file system and pointing to those existing data blocks instead of storing the data blocks again. ZFS provides block-level deduplication because this is the finest granularity that makes sense for general-purpose storage systems. Block-level deduplication also naturally maps to ZFS's 256-bit block checksum, which provides a unique block signature for all blocks in the storage pool, provided the checksum is cryptographically strong (e.g., SHA256). Deduplication is synchronous and is performed when a data block is sent to the cloud object repository 404. If the data blocks are not replicated, enabling deduplication will increase overhead without any benefit. If duplicate data blocks exist, enabling deduplication will save both space and improve performance. The space savings are obvious; the performance improvement is due to the elimination of storage writes when storing duplicate data and the reduced memory footprint due to many applications sharing the same memory pages. Most storage environments contain a mixture of mostly unique data and mostly replicated data. ZFS deduplication is per dataset and can be enabled when it may be helpful.
[0117] A snapshot is a read-only copy of a file system or volume. A snapshot is a view of the file system at a specific point in time. ZFS snapshots are as useful as snapshots of some other file systems: by backing up snapshots, you have a consistent, unchanging target for backup programs to use. Snapshots can also be used to recover from recent errors by copying corrupted files from them. Snapshots can be created almost instantaneously, and they do not initially consume additional disk space within the pool. However, as data within an active dataset changes, snapshots consume disk space by continuing to reference old data, thus preventing disk space from being released. Blocks containing old data are only released when a snapshot is deleted. Taking snapshots is a timed operation. The existence of snapshots does not slow down any operations. Deleting a snapshot takes time proportional to the number of blocks that will be released and is very efficient. ZFS snapshots include the following characteristics: they remain unchanged across system reboots; the theoretical maximum number of snapshots is 2^64; they do not use a separate backup repository; they consume disk space directly from the same storage pool as the file system or volume that created them; recursive snapshots are created quickly as an atomic operation; and they are created together (all at once) or not at all. The advantage of atomic snapshot operations is that snapshot data is always taken at a consistent time, even in descendant file systems. Snapshots are not directly accessible, but they can be cloned, backed up, rolled back, and so on. Snapshots can be used to "roll back" in time to the point when they were taken.
[0118] A clone is a writable volume or filesystem whose initial content is identical to the dataset that created it. In ZFS systems, clones are always created from snapshots. Like snapshots, creating a clone is almost instantaneous and initially does not consume additional disk space. Furthermore, clones can be snapshotted. Clones can only be created from snapshots. When a snapshot is cloned, an implicit dependency is created between the clone and the snapshot. Even if the clone is created elsewhere in the dataset hierarchy, the original snapshot cannot be destroyed as long as the clone exists. Clones do not inherit the characteristics of the dataset that created them. A clone initially shares all its disk space with the original snapshot. As changes are made to the clone, it uses more disk space. Clones are useful for branching and for development or troubleshooting—and can be promoted to replace live filesystems. Clones can also be used to copy filesystems across multiple machines.
[0119] Now go back to refer to Figure 7BFlowchart 700-2 illustrates a method for determining data services and applying them to data blocks. At decision box 731, if compression is enabled or requested, the next box is 740. At decision box 732, if encryption is enabled or requested, the next box is 750. At decision box 733, if deduplication is enabled or requested, the next box is 738. At decision box 734, if a snapshot is to be taken, the next box is 775. Box 735 returns to box 716. Figure 7B The diagram also illustrates the required ordering for any requested data service. Compression must be performed first, followed by encryption, deduplication, and snapshotting / cloning. When reading data blocks, the reverse order must be followed.
[0120] As indicated by box 720, the data object and metadata corresponding to the incremental modification can be stored in the cloud object repository 404. In various embodiments, the cloud interface device 402 can create, read, forward, define, and / or otherwise specify the data object and metadata corresponding to the incremental modification. As indicated by box 722, in some embodiments, the storage of the data object and metadata may be caused at least in part by the cloud interface device 502 sending a request to the cloud interface daemon 604. As indicated by box 724, the request sent by the cloud interface device 502 to the cloud interface daemon 604 can be transformed into a placement request for the cloud client interface 608. In some embodiments, the request sent by the cloud interface device 502 may be a placement request; in other embodiments, the cloud interface daemon 604 may transform the request sent by the cloud interface device 502 into a placement request.
[0121] As indicated in box 726, in response to a request sent by cloud interface device 502, cloud interface daemon 604 may communicate with cloud object repository 404 via one or more network object protocols to store the data object and metadata corresponding to the incremental modification as a new cloud object. As indicated in box 728, cloud interface device 402 may update mapping 406 based on the data object and metadata corresponding to the incremental modification stored in cloud object repository 404.
[0122] Now for reference Figure 7C Draw flowchart 700-3 starting from box 740 and ending at box 742. Figure 7C A flowchart depicts the process of compressing data blocks to conserve storage space. Compression is performed as follows: Figure 2This is performed in the transaction object layer at DMU 218, as shown. Compression is typically enabled because it reduces the resources required for storage (cloud object repository 404) and data transmission. Computational resources are consumed in DMU 218 during the compression process and are typically consumed during the reversal of this process (decompression). Data compression is subject to space-time complexity trade-offs. For example, a compression scheme may require sufficiently fast, intensive decompression processing to be consumed while decompression is in progress. The design of a data compression scheme involves trade-offs between various factors, including the degree of compression and the computational resources required to compress and decompress data. At box 742, DMU 218 receives or retrieves the compression type to compress data blocks. ZFS supports many different types of compression, including but not limited to LZ4, LZJB, GZIP, and ZLE. At box 744, DMU 218 uses the compression type to compress the data blocks. At decision box 746, it is determined whether there are more data blocks from the tree hierarchy that need to be compressed from which data will be written to cloud object repository 404. If so—then repeat box 744 until DMU218 has compressed all blocks using the compression type. Once all blocks have been compressed, box 748 returns to... Figure 7B Box 732.
[0123] Figure 7D Flowchart 700-4 illustrates the process of encrypting data blocks upon request or when encryption is enabled. At box 752, the passphrase or AES key is retrieved or provided to the DMU 218. ZFS uses a "wraparound" encryption key system that uses a locally stored passphrase or AES key and then randomly generates an encryption key for encrypting the data block, as shown in box 754. The passphrase or AES key can be stored on any local storage device, including the ARC224. Encryption itself does not prevent the data block from being stolen, but rather prevents the message content from being presented to the interceptor. In the encryption scheme, the intended data block is encrypted using an encryption algorithm, generating ciphertext that can only be read upon decryption. For technical reasons, encryption schemes typically use pseudo-random encryption keys generated by the algorithm. In principle, it is possible to decrypt a message without possessing the key; however, for well-designed encryption schemes, significant computational resources and skill are required. ZFS data blocks are encrypted using AES (Advanced Encryption Standard) with key lengths of 128, 192, and 256. At box 756, the data block is encrypted using a randomly generated encryption key. At decision box 758, it is determined whether more data blocks need to be encrypted, and if so, they are encrypted at box 756, continuing until no more data blocks need encryption. Then, box 759 returns to... Figure 7B Box 733 is used to determine whether a data block requires further data service processing.
[0124] Figure 7E A flowchart 700-5 depicts the deduplication of data blocks in the cloud object repository 404. Data block replication is a specialized data compression technique used to eliminate duplicate copies of duplicate data blocks. Data block deduplication is used to improve storage device utilization and can also be applied to network data transmission to reduce the number of data blocks that must be sent to be stored in COW memory. In deduplication, unique data blocks are identified and stored during analysis. As analysis continues, other data blocks are compared to the stored copies, and whenever a match occurs, the redundant data block is replaced with a small reference to the stored data block. Given that the same data block pattern may occur dozens, hundreds, or even thousands of times, using deduplication significantly reduces the number of data blocks that must be stored or transmitted. This type of deduplication differs from that used for... Figure 7C The standard file compression discussed performs deduplication. This compression identifies short, repeating substrings within individual data blocks. Storage-based deduplication aims to verify large amounts of data and identify identical entire data blocks so that only one copy of it is stored. For example, consider a typical email system that might contain 100 instances of identical 1MB (megabyte) file attachments. If all 100 instances of that attachment were stored, 100MB of storage would be required. With deduplication, only one instance of the attachment is actually stored; subsequent instances will reference the saved copy backward, resulting in a deduplication ratio of approximately 100 to 1. Therefore, block deduplication can reduce the storage space required in the cloud object repository 404 and reduce the pressure on the network that transmits data blocks to the cloud object repository 404.
[0125] exist Figure 7DIn box 762, the first process is to generate names for data blocks using a name generation protocol. In ZFS, this involves using a checksum algorithm (such as SHA256). When a checksum algorithm is performed on a data block, it generates a checksum unique to the content of the data block. Therefore, if any other data block has exactly the same content, the checksum using the same algorithm or name protocol will be exactly the same. This is crucial for ZFS data deduplication. In decision box 764, it is determined whether an existing data block with the same name exists. This can be done in several ways. One is to maintain a local table with the existing name. However, this limits data block deduplication to only deduplicating data blocks originating locally. A table with the existing name can be stored on object repository 404. Tables on object repository 404 can be stored in a client-local data pool or global storage and are available to all clients. The location of the data stored in the object repository will affect the amount of data block compression achieved through data block deduplication. For example, if the table with the existing name is global, then even if there are multiple clients using the cloud object repository 404, only one copy of the 1MB file discussed above needs to be stored on the cloud object repository 404. This is also true even if the 1MB attachment becomes "viral" via email and eventually becomes an attachment for thousands of emails. Utilizing the global existing name table on the cloud object repository 404, the 1MB file will only be stored once on the cloud object repository 404, but can be referenced by thousands of metadata blocks pointing to it. To this end, at box 766, when further storage calculations are performed using the data blocks, data blocks with the same name as the existing name are ignored, and metadata blocks are created using pointers to the existing blocks to determine their storage purpose. Figure 7A The methods in boxes 716 and 718 generate the tree at box 768. Decision box 770 repeats the process on as many data blocks in the tree as possible by returning to box 762. At decision box 773, if a snapshot has already been requested, then at box 774, the next box is... Figure 7F 775 in the middle. If a snapshot has not yet been requested, then the next box at 772 is... Figure 7A Box 720 in the middle.
[0126] Figure 7F Flowchart 700-6 depicts the methods for ZFS snapshots and cloning to a cloud object repository 404. ZFS snapshots are an essential backup mechanism for the ZFS system. A ZFS snapshot is a "photograph" of the ZFS tree hierarchy at a specific point in the tree's "lifecycle." For example... Figure 3F As discussed in the text, for snapshots, time is based on the birth time of the root block (alternating with the root node and superblock), which is expressed in terms of which progression the root block was created in. Progression occurs each time data blocks are generated for storage. Ultimately, a complete tree hierarchy is generated before a snapshot is taken, such as... Figures 3A-3D and Figure 7A As described in boxes 716 and 718. As depicted in box 780, a reference to the root block is stored. The root block and all blocks active in this progression are part of the snapshot because the snapshot references the root block, and the root block points to all blocks through each lower-level block, but they are not stored as duplicate blocks. More precisely, as shown in box 782, all blocks in the tree accessible from the root block are marked "not released," which specifies them as "read-only" in the ZFS syntax. In the normal progression of the ZFS storage COW system, once a block is inactive—in other words, it is not referenced by any higher-level block—it can be released for other storage needs. Snapshot blocks must remain unchanged for the snapshot to be fully usable. When a block is referenced by the root block referenced by a snapshot, the block cannot be "released" until the snapshot is deleted. This makes it possible to back up to the point where the snapshot was made until the snapshot is deleted. Snapshots can be stored in the cloud object repository 404, ARC 222, L2ARC 222-3, or any other local storage device such as system storage device 228. Using the ZFS file system, snapshots can be requested at specific instances as the hierarchical tree progresses, and snapshots can be automatically generated at specific periodic points in time. Old snapshots are not automatically deleted when a new snapshot is created, so even if a new snapshot is created, a backup of the previous snapshot can occur as long as it hasn't been deleted. Therefore, snapshots enable incremental backups because it's not necessary to copy the entire file system; in fact, the entire backup already exists in the cloud object repository 404, as pointed to by the root block referenced by the snapshot. The root block referenced by the snapshot becomes the active root block, and all subsequent root blocks and blocks created after the snapshot's birth time can be released for use by other storage.
[0127] At decision box 784, determine whether a clone has been requested. If no clone has been requested, then at box 799, the snapshot process ends, and then... Figure 7A Box 720. If a clone has been requested (it is created from a snapshot and is shown in box 786), then the clone reference points to the same root block that the snapshot points to. Clones are always generated from snapshots, making it impossible to delete a snapshot before deleting a clone, as shown in box 788. Clones are used for various purposes—multiple developments from the same data store, instantiating new virtual machines, problem solving, etc. For this, a clone must be able to perform a copy-on-write (COW) when accessed from a clone reference. In this respect, clones differ from snapshots—because a COW does not occur from a snapshot. Clones do not require additional storage capacity at the time of generation, but will use more storage as progress is made on the clone tree. Snapshots can be made from clones, just as snapshots can be made from the active tree, and for exactly the same reasons. Clones do not inherit the characteristics of root block data. Clones can eventually be promoted to the active tree. At box 799, clone generation is complete, followed by... Figure 7A Box 720 in the middle.
[0128] Figure 8 This is a high-level diagram illustrating an example of a cloud interface apparatus 402 for processing incremental modifications according to certain embodiments of the present disclosure. In addition to the ability to optimally recover from the cloud (without applying multiple incremental updates), certain embodiments also provide highly efficient always-incremental (also known as incremental forever) backup capabilities. Traditional backup methods in the industry involve pushing copies to the cloud. However, certain embodiments of the present disclosure allow cloud-based copy-on-write file systems where only new data objects are written to cloud storage. Cloud storage that only needs to send new and modified data provides a highly efficient solution for backup and restore from cloud providers. No modification to old data objects is required. These embodiments, together with a consistency model using transaction groups, facilitate cloud storage with permanent incremental backups, where consistency can always be confirmed.
[0129] To illustrate, Figure 8 The cloud interface apparatus 402 is depicted as an initial backup of a local instance of the logical tree 300 that has already been created (e.g., cloud data object 414 corresponding to the backup logical tree 300A). In some embodiments, the cloud interface apparatus 402 may create the backup using an active tree image 301 or a full copy. After the initial creation of a full backup of the local instance of the logical tree 300, multiple modifications 303 can be made to the logical tree 300 using various transactions. Figure 8 The diagram illustrates the storage of data objects and metadata corresponding to incremental modification 303. As an example, incremental modification 303 is depicted as having new leaf nodes 324 and 326; new non-leaf nodes 328, 330, 332, and 334; and a new root node 336.
[0130] The cloud interface appliance 402 can be configured to potentially create an unlimited number of incremental backups. To this end, certain embodiments of the cloud interface appliance 402 can utilize snapshot COW processing (e.g., as described above regarding...). Figure 3E and 7F (As disclosed). Figure 8In the example depicted, image 304 can be used to create incremental modifications. In some embodiments, image 304 may correspond to an active tree image; in some embodiments, image 304 may correspond to a snapshot. Using image 304, cloud interface equipment 402 can cause storage of incremental modification 303A, which may correspond to cloud storage object 414A. According to cloud-based COW processing, a new cloud object 414A is allocated for the data object and metadata corresponding to incremental modification 303, wherein the cloud-based instantiation of incremental modification 303 is indicated as incremental modification 303A. Then, the new root node can be set as the root of the modified logical tree 300A-1 through a storage operation to finally determine the state of the modified logical tree 300A-1. The modified logical tree 300A-1 may correspond to the logical tree 300A modified by incremental modification 303A.
[0131] By incrementally modifying storage 303A, blocks of the logical tree 300A (saved as cloud data object 414) containing active data can be prevented from being overwritten in-place. A new cloud data object 414A can be allocated, and modified / new data and metadata can be written to the cloud data object 414A. Previous versions of the data can be retained, allowing for the maintenance of snapshot versions of the logical tree 300A. In some embodiments, any unchanged data can be shared between the modified logical tree 300A-1 and its associated snapshots.
[0132] Therefore, by utilizing snapshots such as exemplary snapshot 301 and superblocks such as the superblock corresponding to roots 336, 336-1, the cloud object repository 404 can be updated only for a subset of nodes that have changed since the last snapshot. Furthermore, the update of incremental 303A can be hooked to version 300A of the tree in the cloud via the root node, making any part of the entire modified logical tree 300A-1 accessible.
[0133] By sending incremental snapshots to cloud data repository 404, metadata can be preserved within the data instance, allowing traversal of the tree and obtaining a picture of the tree at the time of the snapshot. Simultaneously, this allows for a very concise data representation, retaining only the data of the referenced blocks specified / requested to be retained by the snapshot, making it possible to store only a minimal amount of data in the tree to preserve a point-in-time image of the tree. Each increment can be merged with tree 300A previously stored in cloud data repository 404, ensuring that a complete current data representation always exists after each increment is sent. Each merging of additional increments with tree 300A results in a single data instance, eliminating the need for a full backup operation.
[0134] While unnecessary from a performance perspective, in some embodiments, periodic full backups can be performed to avoid having to aggregate large increments when clients do not expect them. For some implementations, multiple versions of full backups can be maintained if desired. Therefore, some embodiments provide complete control over intervals, snapshots, and backups. Advantageously, some embodiments can be configured to dynamically adjust the increment interval itself. For example, some embodiments can initially operate based on a first interval. In various cases, the interval can be each write operation to the local tree, every half hour, daily, etc. When the churn rate (e.g., a metric monitored by the hybrid cloud storage system 400 that indicates the rate of change of the local tree) exceeds (or decreases to) a certain churn threshold, the hybrid cloud storage system 400 can automatically transition to different intervals. For example, if the churn rate exceeds a first churn threshold, then the hybrid cloud storage system 400 can automatically transition to a larger interval for the increment interval. Similarly, if the churn rate decreases to a first churn threshold or another churn threshold, then the hybrid cloud storage system 400 can transition to a smaller interval for the increment interval. Some embodiments can employ multiple churn thresholds to limit the increment frequency for each tier scheme. Similarly, some embodiments may dynamically adjust the full backup interval based at least in part on such churn thresholds and / or incremental size and / or quantity thresholds, which in some cases may be client-defined.
[0135] One of the main problems with writing objects to the cloud and subsequently reading them back is the lack of guarantee regarding the integrity of the read objects. Cloud storage carries the risk of incumbent degradation (e.g., storage loss, transmission failure, bit decay, etc.). Furthermore, because multiple versions of objects are involved, there is a risk of reading unexpected versions of the object. For example, it might be possible to read a previous version of the expected object, such as reading the most recent version when an object update is not yet complete.
[0136] Traditional object storage architectures rely on data replication, where an initial quorum (i.e., two copies) is met and additional copies (e.g., a third copy) are updated asynchronously. This provides the possibility that clients can receive a copy of the data before all copies are updated, resulting in an inconsistent view of the data. A typical solution for object storage consistency is to ensure all copies are made before the data is available; however, ensuring consistency using this solution is often impractical or unachievable. Object-based methods, where checksums are stored in the object's metadata or references in other places, allow verification of object content but not the correct version of the object. Furthermore, solutions that might use object versioning and perform verification from a single location could, to some extent, undermine the purpose and intent of cloud storage.
[0137] However, certain embodiments of this disclosure can provide a consistency model that ensures guaranteed integrity in the cloud and guarantees consistent semantics according to an eventually consistent object model. For example, some embodiments can provide fault isolation and consistency by using the logical trees disclosed herein. Checksums in the metadata of the self-described Merkle tree can be updated using all transactional write operations to the cloud object repository. As described above, storing checksums separately from the nodes from which they are computed ensures that each tree automatically performs self-verification. At each tree level, the nodes below are referenced by node pointers that include the checksums. Therefore, when an object is read from the cloud, it can be determined whether the object is correct.
[0138] Figure 9 This is a block diagram illustrating an example method 900 for certain features of a hybrid cloud storage system 400 according to certain embodiments of the present disclosure, features that ensure guaranteed integrity in the cloud and consistently consistent semantics according to an eventually consistent object model. According to some embodiments, method 900 may begin as shown in block 902. However, as set forth above, the teachings of this disclosure can be implemented in various configurations such that the order of certain steps of the methods disclosed herein can be shuffled or combined in any suitable manner and may depend on the chosen implementation scheme. Moreover, although the steps may be separated for the purpose of description, it should be understood that certain steps may be performed simultaneously or substantially simultaneously.
[0139] As indicated in box 902, an application 202 may receive a POSIX-compliant request to perform one or more specific operations (i.e., one or more transactions). This operation may correspond to reading or otherwise accessing data. As indicated in box 904, the POSIX-compliant request may be forwarded from the operating system to the DMU 218 via system call interface 208. As indicated in box 906, the DMU 218 may directly translate the request to perform an operation on a data object into a request to perform one or more read operations (i.e., one or more I / O requests) against the cloud object repository 404. As indicated in box 908, the SPA may receive this I / O request(s) from the DMU 218. In response to this request(s), the SPA may initiate the reading of one or more data objects from the cloud object repository 404.
[0140] As indicated by box 910, cloud interface device 402 can receive one or more I / O requests and can send corresponding cloud interface requests (one or more) to cloud object repository 404. In some embodiments, cloud interface device 402 can translate I / O requests into corresponding object interface requests. Cloud interface device 502 can use mapping 406 of cloud storage objects 414 in cloud object repository 404 to coordinate object interface requests. For example, cloud interface device 402 can identify and request all or part of a file stored as a cloud data object according to a logical tree.
[0141] As indicated in box 912, cloud interface device 402 can receive one or more data objects in response to an object interface request. As indicated in box 914, the data objects(s) ...(s))(s)(s)(s)(s)(s))(s)(s)(s)(s)(s))(s)(s)(s)(s))(s)(s)(s)(s))(s)(s)(s)(s))(s)(s)(s)(s))(s)(s)(s)(s))(s)(s)(s))(s)(s
[0142] Figure 10 This is a high-level diagram illustrating an example of a cloud interface apparatus 402 for processing verification according to certain embodiments of the present disclosure. Similarly, in other embodiments, another component of system 400 may perform this verification. However, in Figure 10 In the illustration, cloud interface device 402 is depicted as having accessed cloud object repository 404 to read data stored according to logic tree 300A-1. In the illustrated example, cloud storage device 402 is shown as having accessed cloud object 414A corresponding to logic tree 300A-1. Cloud storage device 402 accesses leaf node 324-2 through the addresses of non-leaf nodes 326-2 and 334-2 and root node 336-2.
[0143] As described herein, when a node is read from the logic tree, a pointer to a higher-level node in the logic tree is used to read that node. This pointer includes a checksum of the data expected to be read, so that the data can be verified using the checksum when retrieved from the cloud. The actual checksum of the data can be calculated and compared with the expected checksum already obtained by the cloud interface appliance 402 and / or I / O pipe 224. In the depicted example, non-leaf node 326-2 includes a checksum for leaf node 324-2.
[0144] In some embodiments, a checksum can be received from a cloud data repository along with one object, and data to be verified using that checksum can be received along with different objects. According to some embodiments, checksums can be received from a cloud data repository along with separate objects before receiving the different object with the data to be verified. Some embodiments may employ iterative object interface request processing, such that in response to a specific object interface request, an object with a checksum is received, and in response to a subsequent object interface request, an object with the data to be verified is received. Additionally, for some embodiments, the subsequent object interface request can be made using addressing information received along with the specific object interface request and directed to the object with the data to be verified. For some embodiments, instead of iterative processing, multiple objects are received from the cloud data repository, and verification can then be performed on these multiple objects. The integrity of the logic tree can be quickly verified by traversing the logic tree and calculating the checksums of child nodes at each level based on the parent node. In alternative embodiments, the cloud interface appliance 402 and / or I / O pipe 224 may obtain the checksum before initiating a read operation and / or may obtain the checksum from another source. For example, in such an alternative embodiment, the checksum may be retained in mapping 406, a snapshot of the logic tree, and / or a local data repository.
[0145] Refer again Figure 9 As indicated in box 916, it can be determined whether one or more data objects are verified by checking (one or more) checksums from (one or more) parent nodes in the logical tree against (one or more) actual checksums of the data objects. As indicated in box 918, if one or more data objects are verified, the reading and / or further processing operations of system 400 can continue because it has been determined that the data is not corrupted and is not an incorrect version. Verification can continue using other objects if necessary until all objects have been checked and verified by checking and summing against the parent object pointer metadata.
[0146] Error conditions can be identified when the actual checksum of (one or more) data objects does not match the expected checksum. This is Figure 10 The situation shown illustrates a case where the actual checksum of object D' (which corresponds to leaf node 324-2) does not match the checksum specified by parent node 326-2 (which corresponds to object C'). In this error condition, Figure 9The processing flow can proceed to box 920. A mismatch between the actual and expected checksums may correspond to situations such as receiving an incorrect, outdated data version, data degradation due to cloud storage failure, bit decay, and / or transmission loss. As indicated in box 920, a remedial action can be initiated. In some embodiments, the remedial action may include republishing one or more cloud interface requests to the cloud object repository 404. In response to each republishing request, the processing flow can return to box 912, where the cloud interface appliance 402 can receive one or more data objects and perform another iteration of the checksum processing.
[0147] Republishing one or more cloud interface requests can correspond to a request for the cloud object repository 404 to make a greater effort to find the correct version of the requested object. In some embodiments, the cloud interface apparatus 402 can iteratively traverse cloud nodes / devices until the correct version is retrieved. Some embodiments can iteratively verify snapshots of previous versions of tree portions that can be stored in the cloud object repository 404. Some embodiments can employ a threshold so that the cloud interface apparatus 402 can take further remedial actions after the threshold is met. The threshold can correspond to the number of times one or more cloud interface requests are republished. Alternatively or additionally, the threshold can correspond to a limitation on the search scope of nodes and / or snapshots. For example, the threshold can control how many nodes and / or snapshots to search before the cloud interface apparatus 402 moves on to different remedial actions.
[0148] Another remedy that cloud interface equipment 402 can employ is to request the correct version of the data from another cloud object repository in a multi-cloud implementation, as indicated in box 924. Certain embodiments of multi-cloud storage are further described herein. Records in the second cloud object repository can be verified to determine whether a copy of the data has been stored in the second cloud object repository. In some embodiments, this record may correspond to another mapping 406 specific to the second cloud object repository. After determining that a copy of the data has been stored in the second cloud object repository, cloud interface equipment 402 can initiate one or more object interface requests to the second cloud object repository to retrieve the desired data.
[0149] Using this embodiment, after the threshold for re-issuing a request to cloud object repository 404 has been met without successfully receiving the correct version, a request for the object of interest can be made to the second cloud object repository. Alternatively, as a first default, some implementations may resort to the second cloud object repository instead of re-issuing a request to cloud object repository 404. In any case, in response to a request to the second cloud object repository, the processing flow can return to box 912, where cloud interface equipment 402 can receive one or more data objects and perform another iteration of the verification processing. If the correct data is retrieved from the second cloud object repository and passes verification, the read and / or further processing operations of system 400 can continue, as indicated in box 918. Furthermore, after the correct data has been received from the second cloud object repository, system 400 can write the correct data to cloud object repository 404, as indicated in box 926. As disclosed with respect to Figure 7, the writing of correct data can be implemented using a COW process to apply incremental modifications.
[0150] In some embodiments, if the correct data cannot be retrieved through remedial processing, an error message and the most recent version of the data can be returned, as indicated in box 928. In some cases, the most recent version of the data may be unretrievable. For example, the metadata that should point to the data may be corrupted, making it impossible to reference and retrieve the most recent version. In these cases, an error message can be returned without the most recent version of the data. However, when the most recent version of the data is retrievable, it can also be returned. Thus, utilizing an end-to-end checksum model, some embodiments can cover end-to-end traversal from the client to the cloud and back again. Some embodiments not only provide checksums of the entire tree and error detection, but also the ability to correct portions of the tree through resynchronization and data cleanup.
[0151] Figure 11 This is a simplified example of a feature of a hybrid cloud storage system 400-2, further illustrating certain embodiments of the present disclosure. The hybrid cloud storage system 400-2 illustrates how a hybrid storage pool 1100 is formed at least partially by a system storage pool 416 and a virtual storage pool 602-1. Flowcharts for read and write operations are illustrated for each of the system storage pool 416 and the virtual storage pool 602-1.
[0152] In some embodiments, ARC 222 may include ARC 222-2. ARC 222-2 and system memory block 228-1 can facilitate read operations of system memory pool 416. As indicated, some embodiments of ARC 222-2 can be implemented using DRAM. As also indicated, some embodiments of system memory block 228-1 may have a SAS / SATA-based implementation. For some embodiments, read operations of system memory pool 416 can be further facilitated by an L2ARC device 222-3 that can extend cache size. For some embodiments, ARC 222 may be referred to as including L2ARC device 222-3. Figure 11 As indicated, some embodiments of the L2ARC device 222-3 may have an SSD-based implementation. System storage block 228-1 and intent log 214-3 can facilitate write operations to system storage pool 416. As indicated, some embodiments of intent log 214-3 may have an SSD-based implementation.
[0153] For the local system, virtual storage pool 602-1 can appear and behave as a logical disk. Similar to system storage pool 416, ARC 222-4 and cloud storage object block 414-1 can facilitate read operations on virtual storage pool 602-1. In some embodiments, ARC 222 can be referenced as including ARC 222-4, although ARC 222-4 may be implemented in some embodiments using one or more separate devices. Figure 11 As indicated, certain embodiments of ARC 222-4 can be implemented using DRAM. In some embodiments, ARC 222-4 can be the same cache as ARC 222-2; in other embodiments, ARC 222-4 can be a separate cache from ARC 222-2 for virtual storage pool 602-1. As indicated, certain embodiments of facilitating cloud storage object block 414-1 can have an HTTP-based implementation.
[0154] In some embodiments, L2ARC device 222-5 may also facilitate read operations of virtual storage pool 602-1. In some embodiments, ARC 222 may be referred to as including L2ARC device 222-5. In some embodiments, L2ARC device 222-5 may be the same cache as L2ARC device 222-3; in other embodiments, L2ARC device 222-5 may be a separate cache device from L2ARC device 222-3 for virtual storage pool 602-1. As indicated, some embodiments of L2ARC device 222-5 may have an SSD-based or HDD-based implementation.
[0155] Intent log 214-4 can facilitate write operations from virtual storage pool 602-1 to cloud storage object block 414-1. In some embodiments, intent log 214-4 may be the same as intent log 214-3; in other embodiments, intent log 214-4 may be separate from and different from intent log 214-3. As indicated, some embodiments of intent log 214-4 may have an SSD-based or HDD-based implementation.
[0156] The transition to cloud storage offers several advantages (e.g., cost, scale, and geographic location); however, cloud storage has some limitations for regular use. Latency is often a significant issue with conventional technologies when application clients are located on-premises rather than coexisting in the cloud. However, the hybrid cloud storage system 400 can eliminate this problem. Some embodiments of the hybrid cloud storage system 400 can provide mirroring features to facilitate performance, migration, and availability.
[0157] By utilizing the hybrid cloud storage system 400, the difference in latency between read operations in system storage pool 416 and read operations in virtual storage pool 602-1 can be minimized. In some embodiments, the ARC and L2ARC devices for both pools can be local implementations. The latency at the ARC and L2ARC devices in both pools can be equivalent or substantially equivalent. For example, the typical latency for an ARM read operation can be 0.01 ms or less, and the typical latency for L2ARC can be 0.10 ms or less. The latency for read operations from cloud storage object block 414-1 can be higher, but the hybrid cloud storage system 400 can intelligently manage both pools to minimize this higher latency.
[0158] Some embodiments can provide low-latency, direct cloud access with file system semantics. Some embodiments can facilitate operation from cloud storage while preserving application semantics. Some embodiments can achieve local storage read performance while maintaining replicated copies of all data in the cloud. To provide these features, some embodiments can leverage native caching devices and fully utilize hybrid storage pool caching algorithms. By providing efficient caching and file system semantics to cloud storage, cloud storage can be used for more than just backup and recovery. Hybrid storage systems can use “live” cloud storage. In other words, by intelligently using native caching devices, the full performance benefits can be delivered to the local system without having to maintain full or multiple copies locally.
[0159] Figure 12This is a simplified example of a feature of a hybrid cloud storage system 400-2, further illustrating certain embodiments of the present disclosure. Utilizing efficient caching devices and algorithms, the hybrid cloud storage system 400-2 can cache at the block level rather than the file level, thus eliminating the need to cache the entire object for access to large objects (e.g., large databases). In various embodiments, the depicted examples may correspond to one or a combination of virtual storage pool 602-1 and system storage pool 416.
[0160] The hybrid cloud storage system 400-2 can employ adaptive I / O staging to capture most of the objects required for system operation. The hybrid cloud storage system 400-2 can be configured with multiple cache devices to provide adaptive I / O staging. In the depicted example, adaptive I / O staging is implemented using ARC 222-5. However, in various embodiments, the hybrid cloud storage system 400-2 can be configured to use multiple ARC 222 and / or L2ARC 222 to provide adaptive I / O staging. While the following description uses ARC 222 as an example, it should be understood that various embodiments can use multiple ARC 222 and one or more L2ARC 222 to provide the disclosed features.
[0161] In some embodiments, ARC 222-5 can be self-tuning, allowing it to adjust based on I / O workload. For example, in an embodiment of the hybrid cloud storage system 400-2 using cloud storage in live mode (not just for backup and migration), ARC 222-5 can also provide a caching algorithm that prioritizes objects. The priority for caching can correspond to Most Recently Used (MRU) objects, Most Frequently Used (MFU) objects, Least Frequently Used (LFU) objects, and Least Recently Used (LRU) objects. For each I / O operation, ARC 222-5 can determine whether the prioritized data objects need to be self-tuned. It should be noted that in some embodiments, L2ARC 225-5 can work with ARC 222-5 to facilitate one or more caches. For example, L2ARC 225-5, which may have higher latency than ARC 222-5, can be used for one or more lower-ranked caches, such as LRU and / or LFU caches. In some embodiments, another component of the hybrid cloud storage system 400-2 may enable caching according to these embodiments. For example, cloud storage appliance 402 may coordinate caching and servicing of read and write requests. Additionally, according to some embodiments, cloud storage appliance 402 may include one or more ARC 222-5, one or more L2ARC 222-5 and / or intent log 214-4.
[0162] For each I / O operation, ARC 222-5 can adjust the staging of one or more objects that were previously staging to some extent. At a minimum, this adjustment may include updating the tracking of accesses to at least one object. Adjustments may include degrading to a lower staging area, eviction, or promotion to a higher staging area. The transition criteria used for promotion and degrading may be different for each transition from the current staging area to another staging area or to eviction. As disclosed herein, ARC 222-5 may have the ability to evict memory buffers from the cache due to memory pressure to maintain high throughput and / or meet utilization thresholds.
[0163] For a given I / O operation, if one or more objects corresponding to the I / O operation have not yet been temporarily stored as MRU objects, then those one or more objects can be temporarily stored as MRU objects. However, if one or more objects corresponding to the I / O operation have already been temporarily stored as MRU objects, then ARC 222-5 can apply transition criteria to those one or more objects to determine whether to transition them to a different temporary storage. If the transition criteria are not met, then no change to the temporary storage is required when serving that I / O operation.
[0164] Figure 13 This is a block diagram illustrating an example method 1300 for cache management and cloud latency masking of certain features of a hybrid cloud storage system 400-3 according to certain embodiments of the present disclosure. According to some embodiments, method 1300 may begin as indicated by block 1302. However, as stated above, certain steps of the methods disclosed herein may be shuffled, combined, and / or performed simultaneously or substantially simultaneously in any suitable manner, and may depend on the chosen implementation scheme.
[0165] As indicated in box 1302, one or more POSIX-compliant requests to perform one or more specific operations (i.e., one or more transactions) can be received from application 202. Such operations may correspond to reading, writing, and / or modifying data. As indicated in box 1304, the one or more POSIX-compliant requests can be forwarded from the operating system to DMU 218 via system call interface 208. As indicated in box 1306, DMU 218 can directly translate requests to perform operations on objects into requests to perform I / O operations on physical locations within the cloud object repository 404. DMU 218 can forward I / O requests to the SPA.
[0166] As indicated in box 1308, the SPA can receive I / O requests from DMU 218. In response to a request, the SPA can initiate an I / O operation. As indicated in box 1310, in the case of a write operation, the SPA can use a COW procedure to initiate writing an object to the cloud object repository 404. For example, the cloud interface apparatus 402 can continue the COW procedure disclosed above (e.g., with reference to FIG. 7) with respect to the cloud object repository 404. As indicated in box 1312, in the case of a read operation, the SPA can initiate a read of an object. ARC 222-5 can be checked to locate one or more requested objects. In some embodiments, as indicated in box 1314, it can be determined whether one or more verified data objects corresponding to one or more I / O requests exist in ARC 222-5. This can include the SPA first determining whether one or more objects corresponding to one or more read requests are retrievable from ARC 222-5. Then, if such one or more objects are retrievable, the object(s) can be verified using one or more checksums from one or more parent nodes in the logical tree. In various embodiments, verification may be performed by ARC 222-5, cloud storage appliance 402, and / or I / O pipe 224. As indicated in box 1316, if one or more data objects pass verification (or, in some embodiments where verification is not employed, in the simple case of a hit), the reading and / or further processing operations of system 400 may continue.
[0167] In some embodiments, as indicated by box 1318, if one or more data objects fail validation (or, in some embodiments without validation, in a simple case of no hit), the SPA may initiate a read of one or more objects from local storage 228. In some embodiments, pointers to one or more objects may be cached and used to read one or more objects from local storage 228. If such pointers are not cached, some implementations may not check local storage 228 to look up one or more objects. If one or more objects are retrievable, they may be validated using one or more checksums from one or more parent nodes in the logical tree. Again, in various embodiments, validation may be performed by ARC 222-5, cloud storage appliance 402, and / or I / O pipe 224. As indicated by box 1320, if one or more objects pass validation, the processing flow may transition to box 1316, and the read and / or further processing operations of system 400 may continue.
[0168] As indicated in box 1320, in the event that one or more data objects fail validation (or, in some embodiments where validation is not employed, in a simple case of no hit), the processing flow may transition to box 1322. As indicated in box 1322, the SPA may initiate the reading of one or more data objects from cloud object repository 404. In various embodiments, the initiation of the read may be performed via one or a combination of DMU 218, ARC 222-5, I / O pipe 224, and / or cloud storage appliance 402. Reading one or more objects from cloud object repository 404 may include previously stated herein (e.g., regarding...). Figure 9 The steps disclosed herein may include one or more of the following: issuing one or more I / O requests to cloud interface equipment 402, sending corresponding cloud interface requests to cloud object repository 404 using mapping 406 of cloud storage object 414, and receiving one or more data objects in response to object interface requests. As indicated in box 1324, it can be determined whether a verified object has been retrieved from cloud object repository 404. Again, this may involve steps previously described herein (e.g., regarding...). Figure 9 The steps disclosed involve determining whether to verify one or more data objects using checksums from one or more parent nodes in the logical tree.
[0169] In the event that the actual checksum of (one or more) data objects does not match the expected checksum, the processing flow can transition to block 1326, where remedial processing can be initiated. This may involve previously stated details in this document (e.g., regarding...). Figure 9 The steps disclosed include remedial processing. However, if the data passes verification, the processing flow can transition to box 1316, and the reading and / or further processing operations of system 400 can continue.
[0170] As indicated in box 1328, cache staging can be adjusted after one or more objects have been retrieved. In some cases, cache staging adjustments may include newly cached one or more objects as MRU objects, as indicated in box 1330. If one or more objects corresponding to a given I / O operation have not yet been staging as MRU objects, then those one or more objects may be newly staging as MRU objects.
[0171] However, in certain situations, when one or more objects corresponding to an I / O operation have already been staging as MRU, MFU, LFU, or LRU objects, ARC 222-5 may apply transition criteria to one or more objects to determine whether to transition the one or more objects to a different staging area, as indicated in box 1332. However, if the transition criteria are not met, then a change in staging may not be necessary when servicing the I / O operation.
[0172] In some embodiments, object caching may be at least in part a function of the object's recency. As indicated in box 1334, in some embodiments, adjustments to cache caching may include updating one or more recency attributes. ARC 222-5 can define recency attributes for one or more new objects to track the recency of accesses to one or more objects. Recency attributes may correspond to time parameters and / or order parameters, where the time parameter (e.g., absolute time, system time, time difference, etc.) indicates the last access time corresponding to one or more objects, and the order parameter indicates the access count corresponding to one or more objects whose recency attributes can be compared with those of other objects.
[0173] In various embodiments, transition criteria may include one or more proximity thresholds defined to qualify an object for transition from its current stage. For example, ARC 222-5 may determine whether one or more objects should be transitioned to LFU or LRU staging (or eviction) based at least in part on the values of proximity attributes assigned to one or more objects. In some embodiments, the proximity threshold may be a dynamic threshold that adjusts based on proximity attributes defined for one or more other objects in staging. For example, when the values of proximity attributes defined for staging objects are sorted in ascending or descending order, the proximity threshold may be a function of the lowest value of any proximity attribute defined for any object that has been staging as an MFU object.
[0174] Additionally or alternatively, in some embodiments, object caching may be at least partially a function of the object's access frequency. As indicated in box 1336, in some embodiments, adjustments to cache caching may include updating one or more frequency attributes. For a given I / O operation, ARC 222-5 may increment frequency attributes defined for one or more objects to track the access frequency of one or more objects. The frequency attribute may indicate the number of accesses over any suitable time period, which may be an absolute time period, an activity-based time period (e.g., a user session, or the time since the last access activity that met a minimum activity threshold), and so on.
[0175] In various embodiments, transition criteria may include one or more frequency thresholds defined to qualify an object for transition from the current level. For example, after a frequency attribute value changes, ARC 222-5 may determine whether one or more objects should be staged as an MFU object (or as another object in a staged process). This determination may be made at least in part based on comparing the updated frequency attribute with a frequency threshold. In some embodiments, the frequency threshold may be a dynamic threshold that adjusts based on frequency attributes defined for other staged objects (e.g., staged as MFU objects or another object in a staged process). For example, when the values of frequency attributes defined for staged objects are sorted in ascending or descending order, the frequency threshold may be a function of the lowest value of any frequency attribute defined for any object that has already been staged as an MFU object.
[0176] As indicated in box 1338, additionally or alternatively, according to some embodiments, adjustments to cache caching may include specifying one or more additional caching attributes. Caching attributes may indicate the operation type. In some embodiments, object caching may be at least partially a function of the operation type. For example, the caching algorithm may employ a distinction between write and read operations, such that only objects accessed via read operations can be caching as MFU objects. In such embodiments, objects referenced by write operations may initially be maintained as MRU objects and subsequently demoted according to LFU caching criteria. Alternatively, in such embodiments, objects referenced by write operations may initially be maintained as MRU objects and subsequently demoted according to LRU caching criteria and then evicted, effectively skipping LFU caching. As another alternative, objects referenced by write operations may initially be maintained as MRU objects and then evicted, effectively skipping both LFU and LRU caching. For such alternatives, ARC 222-5 distinguishes write operations used to commit cloud objects to the cloud object repository 404 as those that are unlikely to be needed for subsequent read operations, thus allowing such potential operations to cause cloud access latency if they occur.
[0177] Temporary attributes can indicate data types. Additionally or alternatively, in some embodiments, the temporarying of an object can be at least partially a function of its data type. For example, some embodiments may give metadata a higher priority than data. This higher priority may include retaining all metadata objects and allowing data objects to be temporaryed. Alternatively, this higher priority may include applying different criteria for temporarying transitions (promotion and / or demotion from current temporarying) to metadata objects and data objects. For example, the thresholds defined for qualifying for demotion for data objects (e.g., recentity, frequency, and / or similar thresholds) may be lower (and therefore easier to meet) than the thresholds defined for qualifying for demotion for metadata objects. Alternatively, other embodiments may give data a higher priority than metadata. In some embodiments, a portion of a cloud object may be defined to be always cached, regardless of usage frequency.
[0178] Temporary attributes can indicate operation characteristics. Additionally or alternatively, in some embodiments, object temporarying can be at least partially a function of read operation characteristics. For example, some embodiments may give higher priority to read operations with size characteristics (such that the size of the object being read meets a size threshold). Additionally or alternatively, higher priority may be given to read operations with sequence characteristics (such that the sequence of objects being read meets a sequence threshold). Thus, for caching, large sequential streaming read operations can be given higher priority than smaller, more isolated read operations. In this way, the higher cloud access latency associated with large sequential streaming read operations is avoided.
[0179] Some embodiments of ARC 222-5 may employ different features, transition criteria, and / or thresholds for each cache. In some embodiments, ARC 222-5 may employ a cache scoring system. Some embodiments may score objects using numerical expressions (e.g., cache scores). A cache score may reflect an object's eligibility relative to any suitable criterion, such as a transition criterion. For example, a given cache score for an object may be a cumulative score based on criteria such as access frequency, access recentity, operation type, data type, operation characteristics, object size, etc. A given object may be scored for each criterion. For example, a relatively large value for a attribute such as a frequency attribute or a recentity attribute may be assigned a larger score. Similarly, scores may be assigned based on other criteria and priorities. Objects may be ranked according to priority using the object's cumulative cache score and the cache scores of other objects stored in the cache. Again, priority can be used to transition objects to different caches and / or evict them.
[0180] ARC 222-5 can adjust staging and the objects stored therein to meet one or more cache utilization and capacity constraints. For example, given a cache device capacity of, say, 1TB DRAM, ARC 222-5 can adjust staging and the objects stored therein to maintain a maximum cache utilization of 80%. Furthermore, some embodiments can adjust staging and the objects stored therein to meet one or more speed constraints. For example, ARC 222-5 can monitor throughput to maintain an acceptable amount of access latency (e.g., average access latency) given local and cloud access to determine whether more or less cache should be used to meet one or more latency tolerances. Given this adjustment, the adjustment of staging and the objects stored therein can include applying different functions and thresholds to each staging area to prioritize objects. ARC 222-5 can use this priority to move stored objects toward eviction to meet adjustment constraints.
[0181] As indicated in box 1340, in some embodiments, cache staging adjustments may include transitioning to different cache modes. Some embodiments may dynamically change the operating mode to achieve load balancing while meeting utilization and latency constraints. An initial or default operating mode may correspond to operating live from the cloud, such that objects are accessed first from the cache and then from the cloud if necessary. Some embodiments of ARC 222-5 may initially (e.g., within a session or time period) cache all objects accessed using I / O operations and then transition to staging when cache utilization meets one or more thresholds. The transition to staging may be incremental to one or more secondary operating modes. For example, staging may initially degrade to MRU and MFU staging and then expand to one or more other staging modes when one or more cache utilization thresholds (which may initially be and be below the maximum cache utilization threshold) are met.
[0182] Given that cache utilization is close to utilization constraints and meets one or more utilization thresholds, certain transitional criteria can be applied incrementally using one or more additional operation patterns. For example, initially, objects corresponding to write operations may not be distinguished. However, when cache utilization is close to utilization constraints, the distinguishing criterion can be applied after one or more utilization thresholds are met.
[0183] As another example, as cache utilization approaches utilization constraints and meets one or more utilization thresholds, the hybrid cloud storage system 400-2 can begin utilizing extended caches with one or more L2ARC devices (which may correspond to one or more low-latency sides) for lower-ranked staging (e.g., LRU and / or LFU staging). As yet another example, as cache utilization approaches utilization constraints and meets one or more utilization thresholds, the hybrid cloud storage system 400-2 can begin utilizing local storage 228 to meet latency tolerances regarding one or more third operating modes. More specifically, instead of evicting large sequential streaming read operations to provision for future low-latency accesses, the expansion of the operation and the access frequency of the corresponding object are sufficient to meet size and frequency thresholds, allowing these objects to be transitioned to local storage 228. In this way, the hybrid cloud storage system 400-2 can keep objects available for local lazy read operations while freeing up cache capacity for other low-latency accesses (which may require other large sequential read operations) and avoiding cloud latency if large sequential read operations are invoked again. This selective utilization of local storage 228 and cloud storage at the object level can also facilitate masking cloud latency while using the cache most of the time, and promote load balancing between cloud storage and local storage 228 to operate within latency tolerance. In various embodiments, this trifurcated storage adjustment can be initiated as a fallback operating mode or as the initial default for certain types of operations with certain characteristics.
[0184] Therefore, some embodiments may modify the caching model and techniques, at least in part, based on the characteristics of object access. Some embodiments may leverage caching features, cloud storage, and local storage 228 to mask latency in cloud-based operations. In such embodiments, most operations can be served from the cache, which typically has a cache hit rate exceeding 90% or more, resulting in local latency most of the time. If any local object is lost or corrupted, a cloud copy of the object can be accessed. For some embodiments, a read from the cloud object repository 404 may only be necessary if there is no cache hit and the read request cannot be served from local storage 228.
[0185] In some embodiments, instead of the ARC checks, local storage checks, and / or cloud object storage checks disclosed above, mapping 406 can be used to identify the location of one or more objects of interest. As described above, mapping 406 may include an object catalog and may maintain object state updated with each I / O operation. Cloud object state may be stored in indexes, tables, index-organized tables, etc., indexed on a per-object basis. Object state may include object cache state. Object cache state may indicate the location of objects in any one or a combination of ARC, L2ARC, adaptive staging, local storage, and / or cloud storage. By utilizing mapping 406, cloud interface device 402 can directly identify the location of one or more objects of interest. In some embodiments, cloud interface device 402 may utilize mapping 406 only if there is a no-hits according to ARC checks.
[0186] In some embodiments, as a supplement to or alternative to caching, intelligent pool management includes maintaining a mirror that is continuously synchronized with the cloud object storage 404. Some embodiments can provide mirroring between local storage and cloud storage, at least in part by supporting the cloud object data repository as a virtual appliance. Mirroring cloud storage using local storage enables local storage read performance while maintaining a replicated copy of all data in the cloud. By using mirroring, the full performance benefits can be delivered to the local system without having to maintain multiple copies locally. If any local data is lost or corrupted, the cloud copy of the data can be accessed. Synchronizing and mirroring the cloud and local appliances can facilitate even higher levels of performance.
[0187] To facilitate this synchronized mirroring, some embodiments may include a mirror VDEV. Figure 14 An example network file system 200-2 facilitating synchronized mirroring of a hybrid cloud storage system 400 according to certain embodiments of this disclosure is illustrated. File system 200-2 may correspond to file system 200-1, but has mirror management directly integrated into the ZFS control stack. In addition to the disclosure regarding file system 200-1, file system 200-2 may also include a mirror VDEV 1402, which facilitates a cloud copy of data but has local access time / speed for reading data.
[0188] In some embodiments, the mirror VDEV 1402 may interface with one or more VDEVs of another VDEV type within the ZFS file system architecture. ZFS may communicate directly with the mirror VDEV 1402, which may reside at a virtual device layer directly above the driver layer of file system 200-2, and in some embodiments, corresponds to an abstraction of a device driver interface within the ZFS architecture. File system 200-2 may create the mirror VDEV 1402 as a funnel for I / O operations. In some embodiments, the mirror VDEV 1402 may be a point through which other components of file system 200-2 can primarily communicate. For example, in some embodiments, communication from the transaction object layer may reach the physical layer via the mirror VDEV 1402. More specifically, communication from DMU 218 may be directed to I / O pipe 224 and the mirror VDEV 1402. In response to such communication, the mirror VDEV 1402 can direct the communication to other VDEVs (such as VDEV 226) and cloud interface device 502. In this way, the mirror VDEV 1402 can coordinate I / O operations for local storage 228 and cloud object storage 404.
[0189] In some embodiments, the mirror VDEV 1402 may coordinate only write operations for local storage 228 and cloud object storage 404, so that read operations do not need to go through the mirror VDEV 1402. In such embodiments, other VDEVs and cloud interface devices 502, such as VDEV 226, may bypass the mirror VDEV 1402 for read operations. In alternative embodiments, the mirror VDEV 1402 may coordinate all I / O operations.
[0190] Advantageously, the mirrored VDEV 1402 can coordinate write operations, enabling each write operation to be performed synchronously via one or more VDEVs 226 for local storage 228 and via cloud interface device 502 for cloud object repository 404. This synchronous mirroring of each I / O operation is performed at the object level rather than the file level. Data copying for each I / O operation allows the hybrid cloud storage system 400 to achieve local storage read performance that masks cloud access latency. By default, the hybrid cloud storage system 400 can read from local storage 228 to avoid incurring cloud latency for the vast majority of read operations. The hybrid cloud storage system 400 only needs to access the cloud object repository 404 to read a cloud copy of the desired data when it determines that local data is lost or corrupted. Such exceptions can be implemented on an object basis to minimize cloud access latency.
[0191] Figure 15This is a block diagram illustrating an example method 1500 for synchronizing mirroring and masking certain features of cloud latency in a hybrid cloud storage system 400 according to certain embodiments of the present disclosure. According to some embodiments, method 1500 may begin as shown in block 1502. However, as stated above, certain steps of the methods disclosed herein may be mixed, combined, and / or performed simultaneously or substantially simultaneously in any suitable manner, and may depend on the chosen implementation scheme.
[0192] As indicated in box 1502, one or more POSIX-compliant requests to perform one or more specific operations (i.e., one or more transactions) can be received from application 202. Such operations may involve writing and / or modifying data. As indicated in box 1504, one or more POSIX-compliant requests can be forwarded from the operating system to DMU 218 via system call interface 208. As indicated in box 1506, DMU 218 can directly translate requests to perform operations on data objects into requests to perform write operations (i.e., I / O requests) targeting physical locations within system storage pool 416 and cloud object repository 404. DMU 218 can forward I / O requests to the SPA.
[0193] The SPA's mirror VDEV 1402 can receive I / O requests (e.g., via ARC 222 and I / O pipe 224). As indicated in box 1508, the mirror VDEV 1402 can initiate writes to data objects with synchronous replication on a per-I / O operation basis. A portion of the mirror VDEV 1402 can point to local storage, and a portion of the mirror VDEV 1402 can point to the cloud interface device 502 of the cloud storage apparatus 402. As indicated in box 1510, the mirror VDEV 1402 can direct a first instance of a write operation to one or more of the VDEVs 226. In some embodiments, as indicated in box 1514, the above (e.g., given...) Figures 3A-3D The publicly disclosed COW process can continue to write the data object to local system storage.
[0194] As indicated in box 1512, mirror VDEV 1402 can direct a second instance of the write operation to cloud interface device 502 of cloud storage appliance 402. In some embodiments, as indicated in box 1516, the COW process disclosed above can continue to write the data object to local system storage. For example, method 1500 can transition to another step in box 712 or method 700.
[0195] With each write operation performed synchronously on local storage 228 and on cloud object repository 404, hybrid cloud storage system 400 can then intelligently coordinate read operations to achieve local storage read performance that masks the latency of cloud access. Hybrid cloud storage system 400 can coordinate such read operations at any appropriate time after the replicated data storage is complete. See again Figure 15 As indicated in box 1518, an application 202 may receive a POSIX-compliant request to perform one or more specific operations (i.e., one or more transactions). Such operations may correspond to reading or otherwise accessing data. As indicated in box 1520, the POSIX-compliant request may be forwarded from the operating system to the DMU 218 via system call interface 208. As indicated in box 1522, the DMU 218 may directly translate the request to perform an operation on a data object into a request to perform one or more read operations on local storage 228 (i.e., one or more I / O requests). The SPA may receive one or more I / O requests from the DMU 218. As indicated in box 1524, in response to the request(s), the SPA may initiate the reading of one or more data objects from local storage 228. In some embodiments, ARC 222 may be checked first to find a cached version of one or more data objects, and if no cached version is found, an attempt may be made to read one or more data objects from local storage 228.
[0196] As indicated in box 1526, it can be determined whether one or more validated data objects exist corresponding to one or more I / O requests. This may include the SPA first determining whether one or more objects corresponding to one or more I / O requests are retrievable from local storage 228. Then, if such one or more objects are retrievable, the object(s) can be validated using checksums from one or more parent nodes(s) in the logical tree. In various embodiments, validation may be performed by one or more of VDEV 226, mirror VDEV 1402, and / or I / O pipeline 224. As indicated in box 1528, if one or more data objects are validated, the reading and / or further processing operations of system 400 can proceed because it has been determined that the data is not corrupted and is not an incorrect version.
[0197] However, if the SPA determines that no validated data object exists corresponding to one or more I / O requests, the process can transition to box 1530. This determination can correspond to the absence of a retrievable data object corresponding to one or more I / O requests, in which case an error condition can be identified. Similarly, this determination can correspond to a mismatch between the actual checksum and the expected checksum of one or more data objects corresponding to one or more I / O requests, in which case an error condition can also be identified. In either case, the SPA can initiate the reading of one or more data objects from the cloud object repository 404, as indicated in box 1530. In various embodiments, the reading can be initiated via one or a combination of the DMU 218, the image VDEV 1402, the I / O pipeline 224, and / or the cloud interface apparatus 402.
[0198] Reading one or more data objects from the cloud object repository 404 can include those previously mentioned in this document (e.g., regarding...). Figure 9 The steps disclosed herein may include one or more of the following: issuing one or more I / O requests to cloud interface equipment 402, sending corresponding cloud interface requests to cloud object repository 404 using mapping 406 of cloud storage object 414, and receiving one or more data objects in response to object interface requests. As indicated in box 1532, it can be determined whether a verified data object has been retrieved from cloud object repository 404. Again, this may involve steps previously described herein (e.g., regarding...). Figure 9 The steps disclosed involve determining whether to verify one or more data objects using checksums from one or more parent nodes in the logical tree.
[0199] If the data passes verification, the processing flow can transition to block 1528, and the reading and / or further processing operations of system 400 can continue. Additionally, as indicated in block 1534, the correct data can be written to local system storage. In various embodiments, the correction processing can be performed by one or a combination of DMU 218, mirror VDEV 1402, I / O pipes 224, and / or cloud interface equipment 402. In some embodiments, the processing flow can transition to block 1514, where the COW process can continue to write the correct data to local system storage.
[0200] However, if the actual checksum of one or more data objects does not match the expected checksum, the processing flow can transition to box 1536, where remedial processing can be initiated. This may involve what was previously stated in this document (e.g., regarding...). Figure 9The steps disclosed include remedial processing. For example, remedial processing may include republishing one or more cloud interface requests, requesting the correct version of data from another cloud object repository, and so on.
[0201] Advantageously, when the amount of data maintained by the hybrid cloud storage system 400 exceeds a certain amount, the system transitions from the mirror mode disclosed above to a mode based on the information provided. Figure 12 and 13 The caching patterns of the disclosed embodiments become more cost-effective. Some embodiments can automate this transition. The hybrid cloud storage system 400 can begin with mirroring technology and continue until one or more thresholds are reached. Such thresholds can be defined based on storage utilization. For example, when the utilization of local storage capacity reaches a threshold percentage (or absolute, relative, etc.), the hybrid cloud storage system 400 can transition to adaptive I / O caching. In this way, the hybrid cloud storage system 400 can balance the load applied to local storage by changing its operating mode. The hybrid cloud storage system 400 can then allocate the most relevant X-volume data (e.g., 10TB, etc.) and offload the remaining data to cloud storage. This load balancing allows for fewer storage devices while accommodating increasing data storage volumes.
[0202] Figure 16 A simplified diagram of a distributed system 1600 for implementing one of the embodiments is shown. In the illustrated embodiment, the distributed system 1600 includes one or more client computing devices 1602, 1604, 1606, and 1608 configured to execute and operate client applications, such as web browsers, proprietary clients (e.g., OracleForms), etc., via one or more networks 1610. A server 1612 may be communicatively coupled to remote client computing devices 1602, 1604, 1606, and 1608 via network 1610.
[0203] In various embodiments, server 1612 may be adapted to run one or more services or software applications provided by one or more components of the system. In some embodiments, these services may be provided as web-based services or cloud services, or provided to users of client computing devices 1602, 1604, 1606, and / or 1608 under a Software as a Service (SaaS) model. Users operating client computing devices 1602, 1604, 1606, and / or 1608 may then use one or more client applications to interact with server 1612 to utilize the services provided by these components.
[0204] In the configuration illustrated in the figures, software components 1618, 1620, and 1622 of system 1600 are shown as being implemented on server 1612. In other embodiments, one or more components of system 1600 and / or the services provided by these components may also be implemented by one or more of client computing devices 1602, 1604, 1606, and / or 1608. A user operating the client computing device can then utilize one or more client applications to use the services provided by these components. These components may be implemented in hardware, firmware, software, or a combination thereof. It should be recognized that various different system configurations are possible and may differ from the distributed system 1600. The embodiment shown in the figures is therefore an example of a distributed system for implementing the embodiment system and is not intended to be limiting.
[0205] Client computing devices 1602, 1604, 1606, and / or 1608 can be portable handheld devices (e.g., Cellular phone Computing tablets, personal digital assistants (PDAs), or wearable devices (e.g., Google) Head-mounted displays (or similar devices) that run Microsoft Windows And / or software for various mobile operating systems (such as iOS, Windows Phone, Android, BlackBerry, Palm OS, etc.), and with internet access, email, and SMS services enabled. Or other communication protocols. The client computing device can be a general-purpose personal computer, for example, including those running various versions of Microsoft... Apple Personal computers and / or laptops running Linux operating systems. Client computing devices can be any commercially available operating system. Workstation computers running UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, such as Google Chrome OS). Alternatively or additionally, client computing devices 1602, 1604, 1606, and 1608 may be any other electronic device capable of communicating via one or more networks 1610, such as thin client computers, internet-enabled gaming systems (e.g., with or without...). The gesture input device is the Microsoft Xbox game console and / or a personal messaging device.
[0206] Although the exemplary distributed system 1600 is shown as having four client computing devices, any number of client computing devices can be supported. Other devices (such as devices with sensors) can interact with the server 1612.
[0207] The one or more networks 1610 in the distributed system 1600 can be any type of network familiar to those skilled in the art, capable of supporting data communication using any variety of commercially available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Message Switching), AppleTalk, etc. By way of example only, the one or more networks 1610 can be a local area network (LAN), such as a LAN based on Ethernet, Token Ring, etc. The one or more networks 1610 can be a wide area network and the Internet. It can include virtual networks, including but not limited to Virtual Private Networks (VPNs), intranets, extranets, Public Switched Telephone Networks (PSTN), infrared networks, wireless networks (e.g., according to the IEEE 802.11 protocol suite), (and / or any other wireless protocol operating network); and / or any combination of these and / or other networks.
[0208] Server 1612 can consist of one or more general-purpose computers, dedicated server computers (as an example, including PC (personal computer) servers), Servers can be configured as servers, mid-range servers, mainframe computers, rack-mounted servers, server farms, server clusters, or any other suitable arrangement and / or combination. In various embodiments, server 1612 may be adapted to run one or more services or software applications described in the foregoing disclosure. For example, server 1612 may correspond to a server used to perform the processes described above according to embodiments of this disclosure.
[0209] Server 1612 can run any of the operating systems discussed above, as well as any commercially available server operating system. Server 1612 can also run various additional server applications and / or middleware applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, etc. Servers, database servers, etc. Exemplary database servers include, but are not limited to, those commercially available database servers from Oracle, Microsoft, Sybase, IBM, etc.
[0210] In some implementations, server 1612 may include one or more applications to analyze and integrate data feeds and / or event updates received from users of client computing devices 1602, 1604, 1606, and 1608. As an example, data feeds and / or event updates may include, but are not limited to, feed, The server 1612 may update or receive real-time updates and continuous data streams from one or more third-party information sources, which may include real-time events related to sensor data applications, financial quotation machines, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, vehicle traffic monitoring, and the like. The server 1612 may also include one or more applications to display the data feeds and / or real-time events via one or more display devices of client computing devices 1602, 1604, 1606, and 1608.
[0211] Distributed system 1600 may also include one or more databases 1614 and 1616. Databases 1614 and 1616 may reside in various locations. As an example, one or more of databases 1614 and 1616 may reside on non-transient storage media local to server 1612 (and / or residing within server 1612). Alternatively, databases 1614 and 1616 may be located remotely from server 1612 and communicate with server 1612 via a network-based connection or a dedicated connection. In one set of embodiments, databases 1614 and 1616 may reside in a storage area network (SAN) familiar to those skilled in the art. Similarly, any necessary files for performing the functions of server 1612 may be appropriately stored locally on server 1612 and / or remotely. In one set of embodiments, databases 1614 and 1616 may include relational databases, such as those provided by Oracle, adapted to store, update, and retrieve data in response to commands in SQL format.
[0212] Figure 17 This is a simplified block diagram of one or more components of a system environment 1700 according to an embodiment of the present disclosure, through which services provided by one or more components of the embodiment system can be provided as cloud services. In the illustrated embodiment, system environment 1700 includes one or more client computing devices 1704, 1706, and 1708 that can be used by a user to interact with a cloud infrastructure system 1702 providing cloud services. The client computing devices can be configured to operate client applications, such as web browsers, proprietary client applications (e.g., Oracle Forms), or some other application, which can be used by the user of the client computing devices to interact with the cloud infrastructure system 1702 to use the services provided by the cloud infrastructure system 1702.
[0213] It should be recognized that the cloud infrastructure system 1702 depicted in the figures may have other components besides those depicted. Furthermore, the embodiment shown in the figures is merely one example of a cloud infrastructure system that can be incorporated into embodiments of the present invention. In some other embodiments, the cloud infrastructure system 1702 may have more or fewer components than shown in the figures, may combine two or more components, or may have different component configurations or arrangements.
[0214] Client computing devices 1704, 1706, and 1708 can be similar to the devices described above for 1602, 1604, 1606, and 1608. While the exemplary system environment 1700 is shown with three client computing devices, any number of client computing devices can be supported. Other devices, such as those with sensors, can interact with the cloud infrastructure system 1702.
[0215] One or more networks 1710 can facilitate data communication and exchange between clients 1704, 1706, and 1708 and cloud infrastructure system 1702. Each network can be any type of network familiar to those skilled in the art that supports data communication using any of a variety of commercially available protocols, including those described above for networks 1610. Cloud infrastructure system 1702 may include one or more computers and / or servers, which may include those computers and / or servers described above for server 1612.
[0216] In some embodiments, services provided by a cloud infrastructure system may include a variety of services available on demand to users of the cloud infrastructure system, such as online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database processing, managed technical support services, etc. Services provided by a cloud infrastructure system can be dynamically scaled to meet the needs of users of the cloud infrastructure system. A specific instantiation of a service provided by a cloud infrastructure system is referred to herein as a "service instance." Generally, any service available to users from a cloud service provider's system via a communication network (such as the Internet) is referred to as a "cloud service." Typically, in a public cloud environment, the servers and systems that constitute the cloud service provider's system are different from the customer's own local servers and systems. For example, a cloud service provider's system may host applications, and users can subscribe to and use applications on demand via a communication network such as the Internet.
[0217] In some examples, services within a computer network cloud infrastructure may include protected computer network access to storage devices, hosted databases, hosted web servers, software applications, or other services provided to users by the cloud provider, or as otherwise known in the art. For example, services may include password-protected access to remote storage devices in the cloud via the Internet. As another example, services may include web-based hosted relational databases and scripting language middleware engines for private use by networked developers. As yet another example, services may include access to email software applications hosted on a cloud provider's website.
[0218] In some embodiments, cloud infrastructure system 1702 may include a suite of application, middleware, and database service products delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. An example of such a cloud infrastructure system is the Oracle public cloud provided by this assignee.
[0219] In various embodiments, cloud infrastructure system 1702 can be adapted to automatically provision, manage, and track customer subscriptions to services provided by cloud infrastructure system 1702. Cloud infrastructure system 1702 can provide cloud services via different deployment models. For example, services can be provided based on a public cloud model, where cloud infrastructure system 1702 is owned by an organization selling cloud services (e.g., owned by Oracle), and the services are available to the general public or businesses in different industries. As another example, services can be provided based on a private cloud model, where cloud infrastructure system 1702 operates only for a single organization and can provide services to one or more entities within that organization. Cloud services can also be provided based on a community cloud model, where cloud infrastructure system 1702 and the services provided by cloud infrastructure system 1702 are shared by several organizations in the relevant community. Cloud services can also be provided based on a hybrid cloud model, which is a combination of two or more different models.
[0220] In some embodiments, the services provided by the cloud infrastructure system 1702 may include one or more services offered under the Software as a Service (SaaS) category, Platform as a Service (PaaS) category, Infrastructure as a Service (IaaS) category, or other service categories that include hybrid services. A customer may subscribe to one or more services provided by the cloud infrastructure system 1702 via a subscription order. The cloud infrastructure system 1702 then performs processing to deliver the services in the customer's subscription order.
[0221] In some embodiments, the services provided by the cloud infrastructure system 1702 may include, but are not limited to, application services, platform services, and infrastructure services. In some examples, application services may be provided by the cloud infrastructure system via a SaaS platform. The SaaS platform may be configured to provide cloud services that fall into the SaaS category. For example, the SaaS platform may provide the ability to build and deliver on-demand application suites on an integrated development and deployment platform. The SaaS platform may manage and control the underlying software and infrastructure used to provide SaaS services. By utilizing the services provided by the SaaS platform, customers can leverage applications running on the cloud infrastructure system. Customers can obtain application services without having to purchase separate licenses and support. A variety of different SaaS services may be provided. Examples include, but are not limited to, services providing solutions for sales performance management, enterprise integration, and business agility for large organizations.
[0222] In some embodiments, platform services may be provided by a cloud infrastructure system via a PaaS platform. The PaaS platform may be configured to provide cloud services that fall into the PaaS category. Examples of platform services may include, but are not limited to, services that enable organizations (such as Oracle) to integrate existing applications on a shared, public architecture and to leverage the shared services provided by the platform to build new applications. The PaaS platform can manage and control the underlying software and infrastructure used to provide PaaS services. Customers can access PaaS services provided by the cloud infrastructure system without having to purchase separate licenses and support. Examples of platform services include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), etc.
[0223] By leveraging services provided by a PaaS platform, customers can employ programming languages and tools supported by the cloud infrastructure system and also control the deployed services. In some embodiments, the platform services provided by the cloud infrastructure system may include database cloud services, middleware cloud services (e.g., Oracle Fusion Middleware Service), and Java cloud services. In one embodiment, the database cloud service may support a shared services deployment model that enables organizations to aggregate database resources and provide database-as-a-service to customers in the form of a database cloud. Within the cloud infrastructure system, the middleware cloud service provides customers with a platform for developing and deploying various business applications, and the Java cloud service provides customers with a platform for deploying Java applications.
[0224] Various infrastructure services can be provided by IaaS platforms within cloud infrastructure systems. Infrastructure services facilitate the management and control of underlying computing resources (such as storage devices, networks, and other basic computing resources) for customers to utilize services provided by SaaS and PaaS platforms.
[0225] In some embodiments, the cloud infrastructure system 1702 may further include infrastructure resources 1730 for providing resources to customers of the cloud infrastructure system for delivering various services. In one embodiment, infrastructure resources 1730 may include a combination of pre-integrated and optimized hardware (such as servers, storage devices, and networking resources) to perform services provided by PaaS and SaaS platforms. In some embodiments, resources in the cloud infrastructure system 1702 may be shared by multiple users and dynamically reallocated as needed. Furthermore, resources may be allocated to users in different time zones. For example, the cloud infrastructure system 1730 may enable a first group of users in a first time zone to utilize the resources of the cloud infrastructure system for a specified number of hours, and then enable the same resources to be reallocated to another group of users located in a different time zone, thereby maximizing resource utilization.
[0226] In some embodiments, multiple internal shared services 1732 may be provided, shared by different components or modules of the cloud infrastructure system 1702 and services provided by the cloud infrastructure system 1702. These internal shared services may include, but are not limited to: security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelisting services, high availability, backup and recovery services, cloud-enabled services, email services, notification services, file transfer services, etc. In some embodiments, the cloud infrastructure system 1702 may provide comprehensive management of cloud services (e.g., SaaS, PaaS, and IaaS services) within the cloud infrastructure system. In one embodiment, cloud management functionality may include the ability to provision, manage, and track customer subscriptions received by the cloud infrastructure system 1702.
[0227] In some embodiments, as illustrated in the figures, cloud management functionality may be provided by one or more modules, such as order management module 1720, order orchestration module 1722, order supply module 1724, order management and monitoring module 1726, and identity management module 1728. These modules may include or be provided using one or more computers and / or servers, which may be general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0228] In exemplary operation 1734, a customer using a client device (such as client device 1704, 1706, or 1708) can interact with cloud infrastructure system 1702 by requesting one or more services provided by cloud infrastructure system 1702 and placing an order to subscribe to one or more services offered by cloud infrastructure system 1702. In some embodiments, the customer can access cloud user interfaces (UIs) (cloud UI 1712, cloud UI 1714, and / or cloud UI 1716) and place subscription orders via these UIs. Order information received by cloud infrastructure system 1702 in response to a customer placing an order may include information identifying the customer and the one or more services offered by cloud infrastructure system 1702 that the customer wishes to subscribe to.
[0229] After a customer places an order, the order information is received via cloud UIs 1712, 1714, and / or 1716. At operation 1736, the order is stored in the order database 1718. The order database 1718 can be one of several databases operated by the cloud infrastructure system 1718 and in conjunction with other system components. At operation 1738, the order information is forwarded to the order management module 1720. In some cases, the order management module 1720 can be configured to perform order-related billing and accounting functions, such as verifying the order and, after verification, reserving the order.
[0230] At operation 1740, information about the order is transmitted to the order orchestration module 1722. The order orchestration module 1722 can use the order information to orchestrate services and resource provision for orders placed by customers. In some cases, the order orchestration module 1722 can use the services of the order provisioning module 1724 to orchestrate resource provision to support subscribed services.
[0231] In some embodiments, the order orchestration module 1722 enables the management of business processes associated with each order and applies business logic to determine whether an order should be made available for provisioning. At operation 1742, upon receiving a new subscription order, the order orchestration module 1722 sends a request to the order provisioning module 1724 to allocate resources and configure those resources required to fulfill the subscription order. The order provisioning module 1724 enables the allocation of resources for the services ordered by the customer. The order provisioning module 1724 provides an abstraction layer between the cloud services provided by the cloud infrastructure system 1700 and the physical implementation layer for providing the resources used to provide the requested services. Therefore, the order orchestration module 1722 can be isolated from implementation details such as whether services and resources are actually provided on demand or pre-provided and allocated / assigned only upon request.
[0232] At operation 1744, once services and resources are supplied, a notification of the supplied services can be sent to customers on client devices 1704, 1706, and / or 1708 via the order provisioning module 1724 of the cloud infrastructure system 1702. At operation 1746, the order management and monitoring module 1726 can manage and track customer subscription orders. In some cases, the order management and monitoring module 1726 can be configured to collect service usage statistics from subscription orders, such as storage usage, data transfer volume, number of users, system uptime, and system downtime.
[0233] In some embodiments, the cloud infrastructure system 1700 may include an identity management module 1728. The identity management module 1728 may be configured to provide identity services, such as access management and authorization services within the cloud infrastructure system 1700. In some embodiments, the identity management module 1728 may control information about customers who wish to utilize services provided by the cloud infrastructure system 1702. Such information may include information authenticating the identities of these customers and information describing what actions these customers are authorized to perform relative to various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). The identity management module 1728 may also include management of descriptive information about each customer and how and by whom this descriptive information is accessed and modified.
[0234] Figure 18 An exemplary computer system 1800 in which various embodiments of the present invention can be implemented is shown. System 1800 can be used to implement any of the computer systems described above. As shown, computer system 1800 includes a processing unit 1804 that communicates with a plurality of peripheral subsystems via a bus subsystem 1802. These peripheral subsystems may include a processing acceleration unit 1806, an I / O subsystem 1808, a storage subsystem 1818, and a communication subsystem 1824. Storage subsystem 1818 includes a tangible computer-readable storage medium 1822 and system memory 1810.
[0235] Bus subsystem 1802 provides a mechanism for allowing various components and subsystems of computer system 1800 to communicate with each other as intended. While bus subsystem 1802 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 1802 can be any of several types of bus architectures, including memory buses or memory controllers, peripheral buses, and local buses using any of the various bus architectures. For example, such architectures may include Industry Standard Architecture (ISA) buses, Microchannel Architecture (MCA) buses, Enhanced ISA (EISA) buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses, which may be implemented as Mezzanine buses manufactured according to the IEEE P1386.1 standard.
[0236] A processing unit 1804, which may be implemented as one or more integrated circuits (e.g., a conventional microprocessor or microcontroller), controls the operation of the computer system 1800. One or more processors may be included in the processing unit 1804. These processors may include single-core or multi-core processors. In some embodiments, the processing unit 1804 may be implemented as one or more independent processing units 1832 and / or 1834, each including a single-core or multi-core processor. In other embodiments, the processing unit 1804 may also be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0237] In various embodiments, processing unit 1804 can execute various programs in response to program code and can maintain multiple concurrently executing programs or processes. At any given time, some or all of the program code to be executed may reside in processor(s) 1804 and / or storage subsystem 1818. With appropriate programming, processor(s) 1804 can provide the various functions described above. Computer system 1800 may additionally include processing acceleration unit 1806, which may include a digital signal processor (DSP), a dedicated processor, etc. In some embodiments, processing acceleration unit 1806 may include an acceleration engine as disclosed herein or work in conjunction with an acceleration engine to improve computer system functionality.
[0238] The I / O subsystem 1808 may include user interface input devices and user interface output devices. User interface input devices may include keyboards, pointing devices such as mice or trackballs, touchpads or touchscreens integrated into a display, scroll wheels, click wheels, dials, buttons, switches, audio input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may include, for example, motion sensing and / or gesture recognition devices, such as those from Microsoft… Motion sensors enable users to control devices such as Microsoft products via a natural user interface using gestures and voice commands. The 360 game controller's input device interacts with it. The user interface input device may also include eye gesture recognition devices, such as detecting eye movements from the user (e.g., "blinking" when taking a photo and / or making a menu selection) and translating the eye gestures into the input device (e.g., Google). Google input in ) Blink detector. Additionally, the user interface input device may include enabling the user to interact with a voice recognition system (e.g., ...) via voice commands. Voice recognition sensing devices for interaction with navigators.
[0239] User interface input devices may also include, but are not limited to, 3D mice, joysticks or pointing sticks, game panels and drawing tablets, as well as audio / video devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. Furthermore, user interface input devices may include, for example, medical imaging input devices such as computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), and medical ultrasound equipment. User interface input devices may also include, for example, audio input devices such as MIDI keyboards, digital musical instruments, etc.
[0240] User interface output devices may include display subsystems, indicator lights, or non-visual displays such as audio output devices, etc. Display subsystems may be cathode ray tubes (CRTs), flat panel devices such as those using liquid crystal displays (LCDs) or plasma displays, projection devices, touchscreens, etc. Generally, the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from computer system 1800 to the user or other computers. For example, user interface output devices may include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, voice output devices, and modems.
[0241] Computer system 1800 may include a storage subsystem 1818 containing software elements, shown as currently residing in system memory 1810. System memory 1810 may store loadable and executable program instructions on processing unit 1804, as well as data generated during the execution of these programs. Depending on the configuration and type of computer system 1800, system memory 1810 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that can be immediately accessed by processing unit 1804 and / or are currently being operated and executed by processing unit 1804. In some implementations, system memory 1810 may include various types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS), containing basic routines that facilitate the transfer of information between elements of computer system 1800 during startup, may typically be stored in ROM. As an example, but not a limitation, system memory 1810 also illustrates application programs 1812, program data 1814, and operating system 1816, which may include client applications, web browsers, middleware applications, relational database management systems (RDBMS), etc. As an example, operating system 1816 may include various versions of Microsoft... Apple and / or Linux operating system, and various commercially available... Or a UNIX-like operating system (including but not limited to various GNU / Linux operating systems, Google...) Operating systems, etc.) and / or such as iOS, Phone OS 10OS and A mobile operating system based on the OS operating system.
[0242] The storage subsystem 1818 may also provide a tangible computer-readable storage medium for storing basic programming and data structures that provide the functionality of some embodiments. Software (programs, code modules, instructions) that provides the above-described functionality when executed by a processor may be stored in the storage subsystem 1818. These software modules or instructions may be executed by the processing unit 1804. The storage subsystem 1818 may also provide a repository for storing data used according to the present invention.
[0243] The storage subsystem 1800 may also include a computer-readable storage medium reader 1820 that can be further connected to the computer-readable storage medium 1822. Together with and optionally in conjunction with the system memory 1810, the computer-readable storage medium 1822 can comprehensively represent a remote, local, fixed, and / or removable storage device plus storage medium for temporarily and / or more persistently containing, storing, transmitting, and retrieving computer-readable information.
[0244] The computer-readable storage medium 1822 containing code or portions thereof may also include any suitable medium known or used in the art, including storage and communication media, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, electrically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or other tangible computer-readable media. This may also include non-tangible computer-readable media such as data signals, data transmissions, or any other medium that can be used to transmit desired information and can be accessed by the computing system 1800.
[0245] As an example, computer-readable storage medium 1822 may include a hard disk drive that reads or writes to a non-removable non-volatile magnetic medium, a disk drive that reads or writes to a removable non-volatile magnetic disk, and a removable non-volatile optical disk (such as a CD-ROM, DVD, etc.). An optical disc drive that reads from or writes to a disk or other optical medium. Computer-readable storage media 1822 may include, but is not limited to, Disk drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVDs, digital audio tapes, and so on. Computer-readable storage media 1822 may also include solid-state drives (SSDs) based on non-volatile memory (such as flash memory-based SSDs, enterprise flash drives, solid-state ROMs, etc.), volatile memory-based SSDs (such as solid-state RAM, dynamic RAM, static RAM), DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs using a combination of DRAM-based and flash memory-based SSDs. Disk drives and their associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for computer system 1800.
[0246] The communication subsystem 1824 provides an interface to other computer systems and networks. The communication subsystem 1824 serves as an interface for receiving data from other systems and sending data from computer system 1800 to other systems. For example, the communication subsystem 1824 enables computer system 1800 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1824 may include radio frequency (RF) transceiver components (e.g., advanced data network technologies using cellular telephone technology, such as 18G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), WiFi (IEEE 802.11 series standards), or other mobile communication technologies, or any combination thereof), GPS receiver components, and / or other components for accessing wireless voice and / or data networks. In some embodiments, as an addition to or alternative to the wireless interface, the communication subsystem 1824 may provide a wired network connection (e.g., Ethernet).
[0247] In some embodiments, the communication subsystem 1824 may also represent one or more users who can use the computer system 1800 to receive input communications in the form of structured and / or unstructured data feeds 1826, event streams 1828, event updates 1830, etc. As an example, the communication subsystem 1824 may be configured to receive data feeds 1826 in real time from users of social networks and / or other communication services, such as... feed, Updates, web feeds such as rich site summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0248] Furthermore, the communication subsystem 1824 can also be configured to receive data in the form of a continuous data stream, which may include event streams 1828 and / or event updates 1830 that are essentially continuous or unbounded real-time events without a clearly defined termination. Examples of applications that generate continuous data may include, for example, sensor data applications, financial quotation machines, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, vehicle traffic monitoring, and so on. The communication subsystem 1824 can also be configured to output structured and / or unstructured data feeds 1826, event streams 1828, event updates 1830, etc., to one or more databases that may communicate with one or more streaming data source computers coupled to the computer system 1800.
[0249] The computer system 1800 can be one of various types, including handheld portable devices (e.g., Cellular phone Computing tablets, PDAs), and wearable devices (e.g., Google). This includes head-mounted displays, PCs, workstations, mainframes, information stations, server racks, or any other data processing systems. Due to the constantly evolving nature of computers and networks, the description of the computer system 1800 depicted in the figures is merely a concrete example. Many other configurations with more or fewer components than the system depicted in the figures are possible. For example, custom hardware may be used and / or specific elements may be implemented using hardware, firmware, software (including applets), or a combination thereof. Additionally, connections to other computing devices, such as network input / output devices, may be employed. Based on the disclosure and teachings provided herein, those skilled in the art will recognize other ways and / or methods for implementing the various embodiments.
[0250] In the foregoing description, numerous specific details have been set forth for purposes of explanation in order to provide a thorough understanding of various embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without some of these specific details. In other instances, well-known structures and devices are illustrated in block diagram form.
[0251] The foregoing description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the foregoing description of the exemplary embodiments will provide those skilled in the art with enabling descriptions for implementing the exemplary embodiments. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the spirit and scope of the invention as set forth in the appended claims.
[0252] Specific details have been set forth in the foregoing description to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that these embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may have been shown as components in block diagram form to avoid obscuring the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may have been shown without unnecessary detail to avoid obscuring the embodiments.
[0253] Furthermore, it should be noted that the various embodiments may be described as processes, depicted as flowcharts, data flow diagrams, structural diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations can be executed in parallel or simultaneously. Additionally, the order of operations can be rearranged. A process terminates when its operations are completed, but there may be other steps not included in the diagram. Processes can correspond to methods, functions, procedures, subroutines, subroutines, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.
[0254] The term "computer-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and various other media capable of storing, containing, or carrying one or more instructions and / or data. A code segment or machine-executable instruction can represent any combination of procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0255] Furthermore, embodiments can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments for performing the necessary tasks can be stored in a machine-readable medium. The processor can then perform the necessary tasks.
[0256] In the foregoing specification, various aspects of the invention have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that the invention is not limited thereto. The various features and aspects disclosed above may be used individually or in combination. Furthermore, embodiments may be used in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Accordingly, this specification and the accompanying drawings should be considered illustrative rather than restrictive.
[0257] Furthermore, for illustrative purposes, the methods have been described in a specific order. It should be understood that, in alternative embodiments, the methods may be performed in a different order than described. It should also be understood that the methods described above can be executed by hardware components or can be implemented as a sequence of machine-executable instructions that can be used to cause a machine (such as a general-purpose or special-purpose processor or logic circuit programmed with instructions) to execute the methods. These machine-executable instructions can be stored on one or more machine-readable media, such as CD-ROMs or other types of optical discs, floppy disks, ROMs, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memory, or other types of machine-readable media suitable for storing electronic instructions. Alternatively, the methods can be executed by a combination of hardware and software.
[0258] Furthermore, the terms in the claims have their ordinary, general meaning unless the patentee defines them explicitly and clearly otherwise. As used in the claims, the indefinite articles “a” or “an” are defined herein as indicating one or more elements introduced into a particular article; the subsequent use of the definite article “the” is not intended to negate that meaning. Moreover, the use of ordinal terms such as “first” or “second” to clarify different elements in the claims is not intended to assign a specific position in a series, or any other order or sequence, to the element to which the ordinal term has been applied.
Claims
1. A computer-implemented method, comprising: Receive one or more requests via the file system for performing a first transaction on a file; Data blocks and corresponding metadata are stored in one or more storage devices of the file system, wherein: The data block corresponds to the file, and The data blocks and the corresponding metadata are stored as logical blocks according to a tree hierarchy structure; This enables the storage of a set of cloud storage objects in a cloud object repository, wherein the set of cloud storage objects includes the corresponding metadata of the tree hierarchy and the data of the data blocks, and the enabling storage includes: enabling a second set of one or more requests to be transmitted to the cloud object repository to specify the storage of the set of cloud storage objects in the cloud object repository, enabling the set of cloud storage objects to retain the relationship between the data blocks and the corresponding metadata, the corresponding metadata and the data of the data blocks according to the tree hierarchy, wherein the set of cloud storage objects is stored in the cloud object repository to correspond to the cloud-based instantiation of the tree hierarchy; Receive, via the file system, a third group of one or more requests for performing a second transaction on the file; Convert at least one I / O request corresponding to one or more requests in the third group into an object interface request; and Transmit the object interface request to enable communication with the cloud object repository to perform at least one I / O operation for at least one subset of the set of cloud storage objects.
2. The computer-implemented method as described in claim 1, wherein, Sending one or more requests from the second group to the cloud object repository includes: Transmit one or more requests from the fourth group to the cloud daemon process; In response to one or more requests from the fourth group, the cloud daemon communicates with the cloud object repository via one or more object protocols on the network to specify the storage of the group of cloud storage objects.
3. The computer-implemented method as described in claim 1, further comprising: Multiple requests compatible with the Portable Operating System Interface (POSIX) are converted into multiple object interface requests for performing operations on the cloud-based instantiation of the tree hierarchy stored in the cloud object repository, wherein the multiple requests include the at least one I / O request, and the multiple object interface requests include the object interface request.
4. The computer-implemented method as described in claim 1, further comprising: The third group of one or more requests are converted into at least one I / O request for performing the at least one I / O operation, wherein transmitting the object interface request includes communicating with the cloud object repository to perform the at least one I / O operation for at least one subset of the group of cloud storage objects.
5. The computer-implemented method as described in claim 4, wherein, The I / O operations correspond to write operations and are performed according to copy-on-write processing for at least one subset of the set of cloud storage objects in the cloud object repository.
6. The computer-implemented method as described in claim 5, wherein, Specifying the storage of the set of cloud storage objects in the cloud object repository includes: Specify a first object size for storing a first subset of the set of cloud storage objects identified as belonging to a first data type; and Specify a second object size for storing a second subset of the set of cloud storage objects identified as belonging to a second data type; The size of the first object is different from the size of the second object, and the data type of the first object is different from the data type of the second object.
7. A computer system, comprising: One or more processors, communicatively coupled to the memory, facilitate the following operations: Process the first group of one or more requests used to perform the first transaction on the file; The data blocks and their corresponding metadata are stored in one or more storage devices of the file system, wherein: The data block corresponds to the file, and The data blocks and the corresponding metadata are stored as logical blocks according to a tree hierarchy structure; This enables the storage of a set of cloud storage objects in a cloud object repository, wherein the set of cloud storage objects includes the corresponding metadata of the tree hierarchy and the data of the data blocks, and the enabling storage includes: enabling a second set of one or more requests to be transmitted to the cloud object repository to specify the storage of the set of cloud storage objects in the cloud object repository, enabling the set of cloud storage objects to retain the relationship between the data blocks and the corresponding metadata, the corresponding metadata and the data of the data blocks according to the tree hierarchy, wherein the set of cloud storage objects is stored in the cloud object repository to correspond to the cloud-based instantiation of the tree hierarchy; Process one or more third groups of requests for performing a second transaction on the file; Convert at least one I / O request corresponding to one or more requests in the third group into an object interface request; and Transmit the object interface request to enable communication with the cloud object repository to perform at least one I / O operation for at least one subset of the set of cloud storage objects.
8. The computer system as claimed in claim 7, wherein, Sending one or more requests from the second group to the cloud object repository includes: Transmit one or more requests from the fourth group to the cloud daemon process; In response to one or more requests from the fourth group, the cloud daemon communicates with the cloud object repository via one or more object protocols on the network to specify the storage of the group of cloud storage objects.
9. The computer system of claim 8, wherein the one or more processors further facilitate the following operations: Multiple requests compatible with the Portable Operating System Interface (POSIX) are transformed into multiple object interface requests for performing operations on the cloud-based instantiation of the tree hierarchy stored in the cloud object repository, wherein... The plurality of requests includes the at least one I / O request, and the plurality of object interface requests includes the object interface request.
10. The computer system of claim 7, wherein the one or more processors further facilitate the following operations: The third group of one or more requests are converted into at least one I / O request for performing the at least one I / O operation, wherein transmitting the object interface request includes communicating with the cloud object repository to perform the at least one I / O operation for at least one subset of the group of cloud storage objects.
11. The computer system of claim 10, wherein, The I / O operations correspond to write operations and are performed according to copy-on-write processing for at least one subset of the set of cloud storage objects in the cloud object repository.
12. The computer system of claim 11, wherein, Specifying the storage of the set of cloud storage objects in the cloud object repository includes: Specify a first object size for storing a first subset of the set of cloud storage objects identified as belonging to a first data type; and Specify a second object size for storing a second subset of the set of cloud storage objects identified as belonging to a second data type; The size of the first object is different from the size of the second object, and the data type of the first object is different from the data type of the second object.
13. The computer system of claim 12, further comprising the file system.
14. One or more non-transient machine-readable media having machine-readable instructions thereon, the machine-readable instructions, when executed by one or more processors, causing the one or more processors to perform operations, the operations including: Process the first group of one or more requests used to perform the first transaction on the file; The data blocks and their corresponding metadata are stored in one or more storage devices of the file system, wherein: The data block corresponds to the file, and The data blocks and the corresponding metadata are stored as logical blocks according to a tree hierarchy structure; This enables the storage of a set of cloud storage objects in a cloud object repository, wherein the set of cloud storage objects includes the corresponding metadata of the tree hierarchy and the data of the data blocks, and the enabling storage includes: enabling a second set of one or more requests to be transmitted to the cloud object repository to specify the storage of the set of cloud storage objects in the cloud object repository, enabling the set of cloud storage objects to retain the relationship between the data blocks and the corresponding metadata, the corresponding metadata and the data of the data blocks according to the tree hierarchy, wherein the set of cloud storage objects is stored in the cloud object repository to correspond to the cloud-based instantiation of the tree hierarchy; Process one or more third groups of requests for performing a second transaction on the file; Convert at least one I / O request corresponding to one or more requests in the third group into an object interface request; and This enables the transmission of the object interface request to communicate with the cloud object repository to perform at least one I / O operation for at least one subset of the set of cloud storage objects.
15. One or more non-transient machine-readable media as claimed in claim 14, wherein, Sending one or more requests from the second group to the cloud object repository includes: This causes one or more requests from the fourth group to be transmitted to the cloud daemon. In response to one or more requests from the fourth group, the cloud daemon communicates with the cloud object repository via one or more object protocols on the network to specify the storage of the group of cloud storage objects.
16. One or more non-transient machine-readable media as claimed in claim 14, wherein, The operation also includes: Multiple requests compatible with the Portable Operating System Interface (POSIX) are converted into multiple object interface requests for performing operations on the cloud-based instantiation of the tree hierarchy stored in the cloud object repository, wherein the multiple requests include the at least one I / O request, and the multiple object interface requests include the object interface request.
17. One or more non-transient machine-readable media as claimed in claim 14, wherein, The operation also includes: The third group of one or more requests are converted into at least one I / O request for performing the at least one I / O operation, wherein transmitting the object interface request includes communicating with the cloud object repository to perform the at least one I / O operation for at least one subset of the group of cloud storage objects.
Citation Information
Patent Citations
Virtual file system integrating multiple cloud storage services and operating method of the same
US20140164449A1
Cloud storage for mobile devices based on user-specified limits for different types of data
US20160192178A1