Source-agnostic and destination-agnostic data movement
The data mobility agent addresses hybrid storage environment challenges by enabling seamless, automated data transfer across diverse systems, enhancing data management and disaster recovery through data virtualization and multiprotocol gateways.
Patent Information
- Application Number
- PCT/US2024/041152
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2026-02-12
AI Technical Summary
Hybrid data storage environments with dissimilar storage facilities pose challenges in data redundancy, disaster recovery, and management due to incompatible storage protocols, requiring complex custom solutions for each cloud provider.
A data mobility agent leveraging data virtualization enables source-agnostic and destination-agnostic data movement, allowing seamless data transfer across various storage systems using a flexible, scalable, and vendor-agnostic system with multiprotocol gateways and connectors.
Facilitates efficient, automated, and extensible data management across diverse storage environments, supporting synchronous and asynchronous data movement, and integrating with central control layers for global orchestration.
Smart Images

Figure US2024041152_12022026_PF_FP_ABST
Abstract
Description
SOURCE- AGNOSTIC AND DESTINATION- AGNOSTIC DATA MOVEMENTTECHNICAL FIELD
[0001] This disclosure relates to the technical field of data movement.BACKGROUND
[0002] Hybrid data storage environments may include storing data at various different types of storage facilities and storage systems that are not necessarily compatible with each other. For example, data storage customers may employ multiple tools to manage multi-cloud data, resulting in solutions that can be complex and difficult to manage. Furthermore, data redundancy and disaster recovery can be challenging to implement in hybrid environments, such as due to the necessity of sharing data to or recovering data from dissimilar storage facilities that may store data using different storage protocols or techniques. For instance, public cloud storage providers typically leave the responsibility of protecting the data up to the customer / user, which may be difficult to manage for large storage environments. Existing solutions to multi-cloud data protection can be complex and may require different solutions for each different storage environment. This can result in users having to implement separate data and cloud management strategies across various different public cloud providers. As a result, managing data movement between different cloud accounts often requires custom-built solutions.SUMMARY
[0003] Some implementations include a computing device that receives, from a client device, an indication of a data type of data to be provided by the client device to the computing device for storage at at least one data node. Exposing, by the computing device, to the client device, a gateway endpoint that is configured for receiving the data type of the data. Storing, by the computing device, in a buffer, the data received via the gateway endpoint. Determining, by the computing device, a first target data node that is a destination for storage of the data. Determining, by the computing device, a first connector configured to communicate with the first target data node. Formatting, by the computing device, the data into a first format compatible with the first target data node. Sending, by the computing device, via the determined first connector, the formatted data to the first target data node.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] The detailed description is set forth with reference to the accompanying figures. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The use of the same reference numbers in different figures indicates similar or identical items or features.
[0005] FIG. 1 illustrates an example architecture of a system that enables source-agnostic and destination-agnostic data movement according to some implementations.
[0006] FIG. 2 is a block diagram illustrating an example logical configuration of the data mobility agent in a system according to some implementations.
[0007] FIG. 3 is a block diagram illustrating an example system including the mobility agent at a client device according to some implementations.
[0008] FIG. 4 is a block diagram illustrating an example of an extended client-to-mobility agent-to-data node mapping according to some implementations.
[0009] FIG. 5 is a block diagram illustrating an example system including a multipath configuration from a client device to a redundant group of mobility agents connected to a plurality of data nodes according to some implementations.
[0010] FIG. 6 is a block diagram illustrating an example logical arrangement of data virtualization according to some implementations.
[0011] FIG. 7 is a block diagram illustrating an example system including example mappings between CDEs and VDEs according to some implementations.
[0012] FIG. 8 is a block diagram illustrating an example system including example mappings between VDEs and DNDGs according to some implementations.
[0013] FIG. 9 is a flow diagram illustrating an example process performed by a data mobility agent according to some implementations.
[0014] FIG. 10 illustrates select example components of the service computing device(s) that may be used to implement at least some of the functionality of the systems described herein.DESCRIPTION OF THE EMBODIMENTS
[0015] Some implementations herein are directed to techniques and arrangements for a data mobility agent that leverages data virtualization to enable formation of indiscriminate connections between downstream data nodes and client data. Further, some examples include a data mobility agent and virtualization architecture that allows users to seamlessly move streams of data and discrete data elements wherever the users desire. The arrangements herein may scale to include all manner of data streams, data sources, and data destinations, provided that these may beinteracted with in a programmatic manner. Further, the architecture herein allows data to be moved synchronously or asynchronously in both a continuous and / or discrete manner. The data mobility agent herein may serve as a foundational element of a data virtualization layer and may provide a backbone of a data layer covering edge data, on-premises data, near-cloud data, and cloud data.
[0016] Some examples include a data mobility agent and virtualization architecture configured to perform source-agnostic and destination-agnostic data movement. Implementations herein provide a uniquely flexible solution that can be adapted to serve any of numerous variations of user data mobility requirements. In some cases, a connector-based data mobility agent enables an extensible and high-performance data mobility facilitator. Further a virtualization architecture in some examples herein may enable arbitrary connections to arbitrary data sources. Together the features provided by these implementations can enable a multiple-in, multiple-out, anywhere-to- anywhere data mobility solution that can be used for creation, at least in part, of a large-scale data mobility fabric.
[0017] In addition, some examples of the data mobility agent herein allow for external command and control functionality that enable incorporation into larger-scale systems. For example, the examples herein are able to integrate with a central control layer to orchestrate agent gateway and mobility functionality on a global scale that can enable unprecedented data management capabilities. Thus, the extensible, scalable, and vendor-agnostic solutions herein provide users with improved control over their own data.
[0018] The technology described herein includes both a data mobility agent and a virtualization architecture. These may be incorporated into a software-as-a-service-capable (SaaS- capable) data mobility and data protection system with features that include providing a multiprotocol gateway (e.g., iSCSI, File, NAS and cloud) able to interface with various storage products and public cloud providers. For instance, implementations herein may allow system users to copy block data from volumes to public clouds. The expansion of capabilities provided by the implementations herein supports a wide range of storage systems and public cloud platforms. Furthermore, a jobs management function included in examples herein allows for a fully automated experience for supporting scheduling and management of a wide array of data movement combinations in any of block, file, and / or object system, both on-premises and in the cloud. Furthermore, an effective user interface included in implementations herein makes data movement to and from public cloud providers easy and fast. Consequently, implementations herein provide a highly extensible system that accommodates current and future data sources and data destinations.
[0019] For discussion purposes, some example implementations are described in the environment of one or more service computing devices in communication with data nodes and client devices for managing data movement. However, implementations herein are not limited to the particular examples provided, and may be extended to other types of computing system architectures, other types of storage environments, other types of client configurations, other types of data, and so forth, as will be apparent to those of skill in the art in light of the disclosure herein.
[0020] FIG. 1 illustrates an example architecture of a system 100 that enables source- agnostic and destination-agnostic data movement according to some implementations. The system 100 includes one or more service computing devices 102 that are able to communicate with, or otherwise coupled to, a plurality of data node computing devices 104 (hereinafter “data nodes 104”), such as through one or more networks 106. Further, the service computing devices 102 are able to communicate over the one or more networks 106 with a plurality of client devices 108, such as 108-1, 108-2, .. . , which may be any of various types of computing devices, as discussed additionally below.
[0021] In some examples, the service computing devices 102 may include one or more servers that may be embodied in any number of ways. For instance, the programs, other functional components, and at least a portion of data storage of the service computing devices 102 may be implemented on at least one server, such as in a cluster of servers, a server farm, a data center, a cloud-hosted computing service, and so forth, although other computer architectures may additionally or alternatively be used. Additional details of the service computing devices 102 are discussed below with respect to FIG. 7.
[0022] In this example, the service computing devices 102 may execute a mobility agent 110 for facilitating data movement. For example, the service computing devices 102 may be configured to provide storage and data management services to a plurality of users 112. As several non-limiting examples, the users 112 may include users performing functions for businesses, enterprises, organizations, governmental entities, academic entities, or the like, and which may include storage of very large quantities of data in some examples. Nevertheless, implementations herein are not limited to any particular use or application for the system 100 and the other systems and arrangements described herein. Furthermore, in some cases, the users 112 may include administrative users that may manage and or configure the service computing devices 102, the mobility agent 110, and / or the data nodes 104.
[0023] In some examples, the data nodes 104 may include any of various different types of storage devices and / or storage systems, and may include cloud storage, cloud-based storage, commercial storage solutions, public cloud, private cloud, on-premises storage solutions,enterprise storage, network attached storage, storage area networks, local storage devices, network storage, or any of various other types of data storage that is accessible over the one or more networks 106 and / or that is accessible over a local network or direct connection to the service computing devices 102. For example, public cloud and other commercial storage services may enable a lower-cost storage solution per gigabyte than local or on-premises storage that may be co-located at the service computing devices 102 in some cases. In some examples, the data nodes 104 may include both commercially available cloud storage as well as private or enterprise storage systems accessible only by an entity associated with the service computing devices 102, or any of various other combinations of storage solutions.
[0024] The one or more networks 106 may include any suitable network, including a wide area network (WAN), such as the Internet; a local area network (LAN), such as an intranet; a wireless network, such as a cellular network, a local wireless network, such as Wi-Fi, and / or short- range wireless communications, such as BLUETOOTH®; a wired network including Fibre Channel, fiber optics, Ethernet, or any other such network, a direct wired connection, or any combination thereof. Accordingly, the one or more networks 106 may include both wired and / or wireless communication technologies. Components used for such communications can depend at least in part upon the type of network, the environment selected, or both. Protocols for communicating over such networks are well known and will not be discussed herein in detail. Accordingly, the service computing devices 102, the data nodes 104, and the client devices 108 are able to communicate over the one or more networks 106 using wired or wireless connections, and combinations thereof.
[0025] Each client device 108 may be any suitable type of computing device such as a desktop, laptop, tablet computing device, mobile device, smart phone, wearable device, terminal, and / or any other type of computing device able to send data over a network. Users 112 may be associated with the client devices 108 such as through a respective user account, user login credentials, or the like. Furthermore, the client devices 108 may be able to communicate with the service computing device(s) 102 through the one or more networks 106, through separate networks, or through any other suitable type of communication connection. Numerous other variations will be apparent to those of skill in the art having the benefit of the disclosure herein.
[0026] Further, each client device 108 may include a respective instance of a client application 114 that may execute on the client device 108, such as for communicating with services on one or more of the service computing devices 102, such as for sending user data for storage on the data nodes 104 and / or for receiving stored data from the data nodes 104 through a data request 118 or the like. In some cases, the client application 114 may include a browser ormay operate through a browser, while in other cases, the client application 114 may include any other type of application having communication functionality enabling communication with the service computing devices 102 over the one or more networks 106. Additionally, in some examples, the client devices 108 may use the client application 114 or another application for communication directly with one or more of the data nodes 104.
[0027] In the system 100, the users 112 may store data to, and receive data from, the service computing device(s) 102 with which their respective client devices 108 are in communication. Accordingly, the service computing devices 102 may provide storage services for the users 112 and respective client devices 108. During steady state operation there may be users 108 periodically communicating with the service computing devices 102, such as for reading or writing data.
[0028] Further, in the case that the user 112 is an administrative user, the client application 114 may enable the performance of administrative functions, such as for communicating with a management application or the like (not shown in FIG. 1) executable as a service on one or more of the service computing device(s) 102 or the data nodes 104. For instance, an administrative user may use the client device 108 for sending management instructions for managing the system 100, as well as for sending management data for storage on the data nodes 104 and / or for retrieving stored management data from the data nodes 104, and so forth.
[0029] At least one service computing device 102 may execute the mobility agent 110 to enable multiple-in, multiple-out, anywhere-to-anywhere data mobility. In the illustrated example, the data mobility agent 110 includes one or more gateways 120 able to communicate with a data virtualization layer 122, which is able to communicate with one or more connectors 124. The illustrated arrangement of gateway(s) 120, data virtualization layer 122, and connector(s) 124 provides a flexible data virtualization scheme that enables an any-to-any data mapping topology. For instance, in this arrangement, streams of data can enter or exit the data mobility agent 110 via gateway- type connections, and discrete data can enter and exit via the connectors 124. An example of a stream of data may be a stream of block traffic coming from a client device 108, or a stream may be a higher-level stream of data, such as a video stream. The connectors 124 may handle discrete data elements such as individual volumes, files, objects, or similar data entities.
[0030] The data virtualization layer 122 allows manipulation of data at a position in between the gateway 120 and the connectors 124 to ensure compatibility and integrity of the data with one or more of the data nodes 104. In the illustrated example, suppose that a data stream 125 is passing from the first client device 108-1 via a buffer 126 associated with the gateway 120 and into a data pipe 128 (e.g., a bus) where the data stream 125 is then divided into discrete data elements 130that can then be provided via the correct connector to a desired data node 104. For example, different ones of the data nodes 104 may communicate with different connectors 124 that are configured for providing data that has been prepared to be compatible for storage by particular ones of the data nodes 104.
[0031] In addition, in some examples, the client devices 108 may also communicate directly with one or more of the connectors 124. For instance, in the illustrated example, the client device 108-2 is connected for communication directly with the connectors 124. Accordingly, the client device 108-2 may format data for a specific type of connector 124 that is configured for communication with one or more corresponding types of data nodes 104. Similarly, one or more of the data nodes 104 may communicate directly with one or more of the gateways 120, depending on the type of connection / data required by the particular data node 104.
[0032] In some cases, the service computing devices 102 may be arranged into one or more groups, clusters, systems, or the like, at a site 154. Similarly, the data nodes 104 may be arranged into one or more groups, clusters, systems, or the like, at a site 156. In some cases a plurality of sites 154 and / or 156 may be geographically dispersed from each other such as for providing data replication, disaster recovery protection, or the like. Further, in some cases, the service computing devices 102 at different sites 154 may be configured for securely communicating with each other, such as for providing a federation of a plurality of sites 154. Additionally, in some examples, the sites 154 and 156 may be the same sites. Numerous other variations will be apparent to those of skill in the art having the benefit of the disclosure herein.
[0033] FIG. 2 is a block diagram illustrating an example logical configuration of the data mobility agent 110 in a system 200 according to some implementations. In some examples, the system 200 may correspond to the system 100 discussed above or any of various other possible computing system architectures, as will be apparent to those of skill in the art having the benefit of the disclosure herein. The system 200 may provide for fully scalable distributed data storage and may enable data replication, data redundancy, disaster recovery, and other data management features.
[0034] In the illustrated example, a controller 202 is in communication with at least the gateway(s) 102, and is able to receive communications from the client devices 108. For instance, the controller 202 may help to route data between a data source and a correct data destination (e.g., a data node 104). In this example, the controller 202 includes or is otherwise associated with a queue 204, a scheduler 206, and a job manager 208. As one nonlimiting example, the controller 202 may correspond to one or more processors of a service computing device 102 that execute code for performing controller operations as described herein. For instance, the controller 202may receive an indication of the type of data to be received from a client device 108, such as via the client application 114, and may select a particular gateway endpoint 210 that is configured to receive that data type for communication with the client device 108. If necessary, the controller 202 may add, to the queue 204, a job related to receiving the data from the client device 108. The controller 202 may employ the scheduler 206 for scheduling the timing at which the data processing job is taken from the queue 204, and the controller 202 may use the job manager 208 for managing the execution of the scheduled job along with any other scheduled data processing jobs.
[0035] Furthermore, for receiving and processing data from a client device 108, the gateway 102 may include a gateway endpoint 210 that is configured to receive data (e.g., streaming data or other types of data) from the client device 108 in accordance with instructions received from the job manager 208 of the controller 202. For example, the gateway 102 can be configured to provide specific gateway endpoints 210, such as for receiving different types of data from different types of data sources. As one example, if the data that will be received is block formatted data, the gateway endpoint 210 may be configured to communicate with block data sources. Thus, the mobility agent 110 may further include gateway endpoints for file system data, gateway endpoints for object data, and so forth, that client devices are able to connect to depending on the type of data to be transferred. Accordingly, there may be one or many gateways 102 included with each mobility agent 110, as well as one or many connectors 124, depending on the specific needs of a given use case.
[0036] The gateway endpoint 210 may receive the data from the client device 108 and add the data to the buffer 126. The buffer 126 may be managed by a buffer manager 212 which, for example, may ensure that the buffer 126 does not overflow or otherwise lose data. A log manager 214 may keep track of data sources and data destinations for received data. Further, an identity verifier 216 may verify the identity of the data sources such as based on credentials of a user 112, or the like.
[0037] The data received in the buffer 126 may be taken from the buffer 126 by a data controller 220 of the data virtualization layer 122, which may provide the data to the data pipe 128. As one example, when the buffer reaches a threshold level, the data controller may extract data from the buffer 126 and provide the data to the data pipe 128. Furthermore, to translate between streams of data and discrete data elements an intelligent data placer 222 and data packager 224 may be included. For example, information from the gateway endpoint 210 may be provided to the data placer 222 to provide information related to the data node 104 that is the target data destination for the received data. The data placer 222 may determine the correct connector 124for sending the received data to the targeted data node 104, and may provide an instruction to the packager 224, which may format or otherwise configure the data in the data pipe 128 according to the requirements of the identified connector 124 and the target destination data node 104. In some examples, the packager 224 may also compress the data, encrypt the data, or the like, depending on the requirements of the target data nodes 104.
[0038] As mentioned above, there may be different types of connectors 124 that correspond to the different types of data nodes 104. In the illustrated example, a first type connector 124-1 corresponds to a first type data node 104-1; a second type connector 124-2 corresponds to a second type data node 104-2; and a third type connector 124-3 corresponds to a third type data node 104-3, and so forth. For instance, if the first type of data nodes 104-1 are provided by a first commercial storage provider (e.g., AMAZON S3®), then the type first connector 124-1 may be configured to interact with one or more application programming interfaces (APIs) that has been exposed by the commercial storage provider for use in interacting with the first type of data nodes 104-1, such as for sending data to the first type of data nodes for storage.
[0039] Additionally, suppose that the second type data nodes 104-2 store data in block format, and that the second type of connector 124-2 is configured to communicate with block storage devices that store data in block format, e.g., as opposed to object data storage, file system data , or the like. Thus, as one example, an API that enables communication with the block storage data nodes 104-2 may be used by the second type connector 124-2 for interacting with the second type data nodes 104-2.
[0040] As one example, suppose that data is being provided from a block storage device, such as from an LUN (logical unit), and is being sent to an object-based storage, e.g., at the first type of data node 104-1. The first type of connector 124-1 may be configured to interface with files or other types of objects in the object store. Accordingly, the streamed data received through the buffer 126 may be converted into one or more files by the packager 224, and the one or more files may be provided to the first type connector 124-1 which may send the one or more files to the first type data node 104-1 for storing the data.
[0041] FIG. 3 is a block diagram illustrating an example system 300 including the mobility agent 110 at a client device 108 according to some implementations. In this example, the client device 108 executes the mobility agent 110, either as a software program, or as a piece of embedded code running on a dedicated hardware peripheral, such as a network interface card (NIC). The mobility agent 110 may include a plurality of connectors (not shown in FIG. 3) configured for connecting to various different types of data nodes 104, similar to the example of FIG. 2 discussed above. In some cases, the mobility agent 110 may receive data from the clientapplication 114, and may send the received data to one or more targeted data nodes 104. In other examples, the client device 108 may receive the data from one or more other client devices 108 or from one or more other computing devices (not shown in FIG. 3). Numerous other variations will be apparent to those of skill in the art having the benefit of the disclosure herein.
[0042] FIG. 4 is a block diagram illustrating an example 400 of an extended client-to-mobility agent-to-data node mapping according to some implementations. In this example, there may be a communication channel between a controller 402 and each respective mobility agent 110, e.g., a first mobility agent 110-1 through an Nth mobility agent 110-N, as indicated by dashed line 404. For instance, the communication channel 404 may enable the controller 402 to manage operations of each of the mobility agents 110-1 through 110-N. In some cases, the controller 402 may be similar to the controller 202 discussed above with respect to FIG. 2, but rather than being specific to any particular mobility agent 110, the single controller 402 may control operations of a plurality of mobility agents 110-1 through 110-N.
[0043] Furthermore, the example of FIG. 4 illustrates that the implementations herein are able to map many client devices 108, such as a first client device 108-1 through an Mth client device 108-M, to many data nodes 104, such as via a cluster of mobility agents 110-1 through 110-N. Each of these mappings can have multiple redundant paths and can span many data nodes 104, such as a first type of data node 104-1 through an Lth type of data node 104-L. In some examples, each different type of data node 104-1 through 104-L may contain a different type of data and may be at any location, e.g., physically near, or in the same system as a client device 108, may be in a remote location or any other arbitrary physical location such as at a near or geographically remote data center etc. the cluster of mobility agents 110-1 through 110-N may execute on a single service computing device 102 or on multiple service computing devices 102 (not shown in FIG. 4).
[0044] FIG. 5 is a block diagram illustrating an example system 500 including a multipath configuration from a client device 108 to a redundant group of mobility agents 110 connected to a plurality of data nodes 104 according to some implementations. In this example, similar to the example of FIG. 4 discussed above, a plurality of mobility agents 110-1 through 110-N are grouped together, such as in a cluster, and controlled by the controller 402. A separate bus 502 connects each mobility agent buffer 126-1 through 126-N to the other mobility agent buffers 126- 1 through 126-N that are included in the group of mobility agents 110-1 through 110-N. The separate bus 502 may provide additional redundancy to the system 500. For example, the bus 502 may be implemented by software. Multiple data mobility agents 110 may communicate asynchronously over a local network using RESTful APIs for ensuring non-blocking dataexchange. Each data node 104 my include an in-built database for local data storage, thereby eliminating the need for physical connections to a central storage system. This arrangement enhances scalability, flexibility, and fault tolerance for ensuring consistent and reliable data management across the system.
[0045] As one example, within the cluster of mobility agents 110-1 through 110-N, a buffering operation may be performed (depending in part on the type of connection needed for the relevant data). In this situation, a buffer mirroring function may be performed across the cluster of mobility agents 110-1 through 110-N to ensure that connections routed through the cluster remain robust regardless of whether one or more of the mobility agents 110 in the cluster fail or are otherwise lost. Additionally, in some examples, the connections between the client device(s) 108 and the mobility agents 110-1 through 110-N may also include redundant connections, as well as the connections between the mobility agents 110-1 through 110-N and the data nodes 104-1 through 104-L.
[0046] FIG. 6 is a block diagram illustrating an example logical arrangement 600 of data virtualization according to some implementations. To enable the above-discussed flexible data mobility and compatibility, implementations herein may employ a data virtualization architecture such as illustrated in FIG. 6. In this example, one or more client data entities (CDEs) 602 are associated with a client device 108. Each CDE 602-1 through 602-X may be a file, object, volume, or the like. Alternatively, a CDE 602 may be a streaming device such as a block storage device, or the like. In some cases, each CDE 602 may map to a respective virtual data entity (VDE) 604 that exists in the data virtualization layer 122 as a high-level construct used to facilitate managing the data mapping, e.g., 604-1 through 604-X in the illustrated example.
[0047] The VDEs 604 may in turn map to a respective data node data group (DNDG) 606 or to multiple DNDGs 606. For example, each DNDG 606 may be composed of one or more data node data entities (DNDEs) 608, which may be, for example, files in a file system, objects in an object storage, or so forth. The flexible data virtualization arrangement illustrated in FIG. 6 can enable arbitrary connections between any data source and data destination, and may help to support the robust general purpose data mobility solution described herein.
[0048] FIG. 7 is a block diagram illustrating an example system 700 including example mappings between CDEs 602 and VDEs 604 according to some implementations. This example demonstrates some of the client connection flexibility afforded by the architecture of FIG. 6. For instance, since a core data element, namely the VDE 604, exists outside of both the client device 108 and the data nodes 104 (not shown in FIG. 7) the CDE 602 can serve as a virtual mapping to the corresponding VDE 604. As such, the VDE 604 can, in some cases, map to CDEs 602 onmultiple client devices. Accordingly, in the illustrated example, the first CDE 602-1 on the first client device 108-1 maps to the first VDE 604-1 at the data virtualization layer. Similarly, the first CDE602-1 on the second client device 108-2 also maps to the first VDE 604-1. Additionally, while not all data types may allow for this type of mapping, the architecture herein is flexible enough to support a large number of such cases.
[0049] FIG. 8 is a block diagram illustrating an example system 800 including example mappings between VDEs 604 and DNDGs 606 according to some implementations. This example illustrates flexibility of the data virtualization layer 122 on the data node side. For instance, an individual VDE 604 may be mapped to one or more DNDGs 606 across different data nodes 104. In this example, the first VDE 604-1 is mapped to both the first DNDG 606-1 on the first data node 104-1, and is also mapped to the 1-Xth DNDG 606-(l-X) on the second data node 104-2. The content of the DNDEs 608 (not shown in FIG. 8) within the DNDGs 606 can be different between the data nodes 104, and the data itself can be in a different format. For instance, suppose that the data stored on first data node 104-1 is files in a file system and that the data stored at the second data node 104-2 is objects in an object storage. This flexible mapping can be exploited to optimize data placement for lowering costs or for taking other factors like carbon footprint into consideration.
[0050] FIG. 9 is a flow diagram 900 illustrating an example process performed by a data mobility agent according to some implementations. The process is illustrated as a collection of blocks in a logical flow diagram, which represents a sequence of operations, some or all of which may be implemented in hardware, software or a combination thereof. In the context of software, the blocks may represent computer-executable instructions stored on one or more computer- readable media that, when executed by one or more processors, program the processors to perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures and the like that perform particular functions or implement particular data types. The order in which the blocks are described should not be construed as a limitation. Any number of the described blocks can be combined in any order and / or in parallel to implement the process, or alternative processes, and not all of the blocks need to be executed. For discussion purposes, the process is described with reference to the environments, frameworks, and systems described in the examples herein, although the process may be implemented in a wide variety of other environments, frameworks, and systems. In FIG. 9, the process 900 may be executed at least in part by a service computing device 102 or other computing device executing one or more data mobility agents.
[0051] At 902, the computing device may receive, from a client device, an indication of a data type of data to be provided by the client device to the computing device for storage at at least one data node.
[0052] At 904, the computing device may expose, to the client device, a gateway endpoint that is configured for receiving the data type of the data.
[0053] At 906, the computing device may store the data received via the gateway endpoint in a buffer.
[0054] At 908, the computing device may determine a target data node that is a destination for storage of the data.
[0055] At 910, the computing device may determine a connector configured to communicate with the target data node.
[0056] At 912, the computing device may format the data into a format compatible with the target data node.
[0057] At 914, the computing device may send, via the determined connector, the formatted data to the target data node.
[0058] The processes described herein are only examples of processes provided for discussion purposes. Numerous other variations will be apparent to those of skill in the art in light of the disclosure herein. Further, while the disclosure herein sets forth several examples of suitable frameworks, architectures and environments for executing the processes, the implementations herein are not limited to the particular examples shown and discussed. Furthermore, this disclosure provides various example implementations, as described and as illustrated in the drawings. However, this disclosure is not limited to the implementations described and illustrated herein, but can extend to other implementations, as would be known or as would become known to those skilled in the art.
[0059] FIG. 10 illustrates select example components of the service computing device(s) 102 that may be used to implement at least some of the functionality of the systems described herein. The service computing device(s) 102 may include one or more servers or other types of computing devices that may be embodied in any number of ways. For instance, in the case of a server, the programs, other functional components, and data may be implemented on a single server, a cluster of servers, a server farm or data center, a cloud-hosted computing service, and so forth, although other computer architectures may additionally or alternatively be used. Multiple service computing devices 102 may be located together or separately, and organized, for example, as virtual servers, server banks, and / or server farms. The described functionality may be provided bythe servers of a single entity or enterprise, or may be provided by the servers and / or services of multiple different entities or enterprises.
[0060] In the illustrated example, the service computing device(s) 102 includes, or may have associated therewith, one or more processors 1002, one or more computer-readable media 1004, and one or more communication interfaces 1006. Each processor 1002 may be a single processing unit or a number of processing units, and may include single or multiple computing units, or multiple processing cores. The processor(s) 1002 can be implemented as one or more central processing units, microprocessors, microcomputers, microcontrollers, graphics processors, system-on-chip processors, digital signal processors, artificial intelligence processors, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. As one example, the processor(s) 1002 may include one or more hardware processors and / or logic circuits of any suitable type specifically programmed or configured to execute the algorithms and processes described herein. The processor(s) 1002 may be configured to fetch and execute computer-readable instructions stored in the computer-readable media 1004, which may program the processor(s) 1002 to perform the functions described herein.
[0061] The computer-readable media 1004 may include volatile and nonvolatile memory and / or removable and non-removable media implemented in any type of technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. For example, the computer-readable media 1004 may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, optical storage, solid state storage, magnetic tape, magnetic disk storage, storage arrays, network attached storage, storage area networks, cloud storage, or any other medium that can be used to store the desired information and that can be accessed by a computing device. Depending on the configuration of the service computing device(s) 102, the computer-readable media 1004 may be a tangible non-transitory medium to the extent that, when mentioned, non-transitory computer-readable media exclude media such as energy, carrier signals, electromagnetic waves, and / or signals per se. In some cases, the computer-readable media 1004 may be at the same location as the service computing device 102, while in other examples, the computer-readable media 1004 may be partially remote from the service computing device 102. For instance, in some cases, the computer-readable media 1004 may include a portion of storage in the data nodes 104 discussed above, e.g., with respect to FIGS. 1-6 and 8.
[0062] The computer-readable media 1004 may be used to store any number of functional components that are executable by the processor(s) 1002. In many implementations, these functional components comprise instructions or programs that are executable by the processor(s)1002 and that, when executed, specifically program the processor(s) 1002 to perform the actions attributed herein to the service computing device 102. Functional components stored in the computer-readable media 1004 may include a data mobility agent program 1008 that may be executed to initiate one or more mobility agents 110 as discussed above. In some specific examples, additional functional components may include a user web application 1010 and a management web application 1012. For example, users may communicate with the service computing device 102 via a user interface provided via the user web application 1010, while administrative users may communicate with the service computing device 102 via a user interface provided by the management web application 124. Alternatively in other examples, the user web application 1010 and the management web application 1012 are not included.
[0063] In addition, the computer-readable media 1004 may store data, data structures, and other information used for performing the functions and services described herein. Such as discussed above with respect to FIG. 2. The service computing device 102 may also include or maintain other functional components and data, which may include programs, drivers, etc., and the data used or generated by the functional components. Further, the service computing device 102 may include many other logical, programmatic, and physical components, of which those described above are merely examples that are related to the discussion herein.
[0064] The one or more communication interfaces 1006 may include one or more software and hardware components for enabling communication with various other devices, such as over the one or more network(s) 106. For example, the communication interface(s) 1006 may enable communication through one or more of a LAN, the Internet, cable networks, cellular networks, wireless networks (e.g., Wi-Fi) and wired networks (e.g., Fibre Channel, fiber optic, Ethernet), direct connections, as well as close-range communications such as BLUETOOTH®, and the like, as additionally enumerated elsewhere herein.
[0065] Various instructions, methods, and techniques described herein may be considered in the general context of computer-executable instructions, such as computer programs and applications stored on computer-readable media, and executed by the processor(s) herein. Generally, the terms program and application may be used interchangeably, and may include instructions, routines, modules, objects, components, data structures, executable code, etc., for performing particular tasks or implementing particular data types. These programs, applications, and the like, may be executed as native code or may be downloaded and executed, such as in a virtual machine or other just-in-time compilation execution environment. Typically, the functionality of the programs and applications may be combined or distributed as desired invarious implementations. An implementation of these programs, applications, and techniques may be stored on computer storage media or transmitted across some form of communication media.
[0066] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.
Claims
CLAIMS1. A system comprising: a computing device able to communicate with a plurality of data nodes and at least one client device, the computing device configured by executable instructions to perform operations comprising: receiving, from the client device, an indication of a data type of data to be provided by the client device to the computing device for storage at at least one of the data nodes; exposing, to the client device, a gateway endpoint that is configured for receiving the data type of the data; storing the data received via the gateway endpoint in a buffer; determining a first target data node that is a destination for storage of the data; determining a first connector configured to communicate with the first target data node; formatting the data into a first format compatible with the first target data node; and sending, via the determined first connector, the formatted data to the first target data node.
2. The system as recited in claim 1, wherein a second target data node is determined to also be a target for storage of the data, the second target node storing data in a second format that is different from the first format compatible with the first target data node, the operations further comprising: determining a second connector configured to communicate with the second target data node; formatting the data into the second format compatible with the second target data node; and sending, via the determined second connector, the second formatted data to the second target data node.
3. The system as recited in claim 1, wherein formatting the data is performed at a virtualization layer provided by the computing device, the virtualization layer including a packager that formats the data into the format compatible with the target data node.
4. The system as recited in claim 1, wherein formatting the data includes dividing the data into a plurality of data elements having a format compatible data stored on the target data node.
5. The system as recited in claim 1, further comprising a controller that receives, from the client device, the indication of a data type of the data, and that determines the gateway endpoint for receiving the data.
6. The system as recited in claim 1, wherein the computing device is another client device that is able to communicate over a network with the client device for receiving the data therefrom, and that is also able to communicate over the network with the target data node at least via the determined connector.
7. The system as recited in claim 1, wherein the client device includes a plurality of client data entities that contain data for storage to at least one target data storage, wherein the computing device generates a plurality of virtual data entities, each virtual data entity corresponding to one of the client data entities for managing data associated with the corresponding client data entity.
8. The system as recited in claim 7, wherein each virtual data entity further corresponds to at least one data node data group that includes one or more data node data entities, the virtual data entity enabling the data to be formatted for storage with the corresponding data node data group.
9. The system as recited in claim 7, wherein a client data entity on a first client device and a client data entity on a second device both map to the same virtual data entity on a virtualization layer provided by the computing device.
10. The system as recited in claim 1, wherein there are a plurality of buffers including the buffer, wherein at least some of the buffers are connected to each other via a bus for communication redundancy.
11. The system as recited in claim 1, further comprising a data pipe configured to receive the data from the first buffer, wherein the data in the data pipe is formatted for compatibility with the target data node.
12. A method comprising: receiving, by a computing device, from a client device, an indication of a data type of data to be provided by the client device to the computing device for storage at at least one data node;exposing, by the computing device, to the client device, a gateway endpoint that is configured for receiving the data type of the data; storing, by the computing device, in a buffer, the data received via the gateway endpoint; determining, by the computing device, a first target data node that is a destination for storage of the data; determining , by the computing device, a first connector configured to communicate with the first target data node; formatting, by the computing device, the data into a first format compatible with the first target data node; and sending, by the computing device, via the determined first connector, the formatted data to the first target data node.
13. The method as recited in claim 12, further comprising: determining a second target data node that is also a target for storage of the data, the second target node storing data in a second format that is different from the first format compatible with the first target data node; determining a second connector configured to communicate with the second target data node; formatting the data into the second format compatible with the second target data node; and sending, via the determined second connector, the second formatted data to the second target data node.
14. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, configure the one or more processors to perform operations comprising: receiving, by the one or more processors, from a client device, an indication of a data type of data to be provided by the client device to the computing device for storage at at least one data node; exposing, by the one or more processors, to the client device, a gateway endpoint that is configured for receiving the data type of the data; storing, by the one or more processors, in a buffer, the data received via the gateway endpoint;determining, by the one or more processors, a first target data node that is a destination for storage of the data; determining, by the one or more processors, a first connector configured to communicate with the first target data node; formatting, by the one or more processors, the data into a first format compatible with the first target data node; and sending, by the one or more processors, via the determined first connector, the formatted data to the first target data node.
15. The one or more non-transitory computer-readable media as recited in claim 14, the operations further comprising: determining a second target data node that is also a target for storage of the data, the second target node storing data in a second format that is different from the first format compatible with the first target data node; determining a second connector configured to communicate with the second target data node; formatting the data into the second format compatible with the second target data node; and sending, via the determined second connector, the second formatted data to the second target data node.
Citation Information
Patent Citations
Information processing apparatus and data transfer method
US20110283037A1
Data transfer requests with data transfer policies
US20160197988A1
Flexible buffer management for optimizing congestion control using radio access network intelligent controller for 5g or other next generation wireless network
US20210377790A1