A dynamic and adaptive cross-cloud data replication method and system
Through the combination of server-free computing and replication functions, the efficient and low latency problems of cross-cloud data replication are solved, and the flexibility and low latency of cross-cloud data synchronization are achieved, thereby reducing operational costs.
Patent Information
- Application Number
- CN202510724951.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Existing cross-cloud data replication solutions are difficult to achieve efficient, low latency and cross-cloud data synchronization. Cloud vendors' proprietary services only support inter-regional replication within the same cloud platform. Open source tools are slow to respond in the face of dynamic load changes, and there are additional storage and data transmission costs.
The server-free computing is adopted, and replication functions are introduced. Through performance analysis modeling and dynamic execution strategy planning, combined with decentralized data block granular scheduling, cross-cloud data replication is achieved to meet the needs of high efficiency and low latency.
It improves the dynamic adaptability and low latency of cross-cloud data replication, reduces operational costs, and enhances the flexibility and reliability of cross-cloud data replication.
Smart Images

Figure CN120238547B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing technology, and in particular to a dynamic and adaptive cross-cloud data replication method and system. Background Art
[0002] Cross-cloud data replication is a key technology used to synchronize and distribute data between different cloud platforms to improve business availability, disaster recovery capabilities, and optimize business processing performance.
[0003] Current data replication solutions are primarily divided into two categories: cloud vendor proprietary services and open source tools. Cloud vendor proprietary services typically only support inter-region replication within the same cloud platform, making it difficult to achieve true cross-cloud data synchronization. They also struggle to meet the requirements of low-latency scenarios, incurring additional storage and data transfer costs and increasing overall operational overhead. Open source tools rely on virtual machine (VM) instances as their core architecture, making them less responsive to dynamic load changes. Therefore, neither cloud vendor proprietary services nor VM-based open source tools can meet the requirements for efficient, low-latency, and cross-cloud data replication. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a dynamic and adaptive cross-cloud data replication method and system that can meet the needs of efficient, low-latency and cross-cloud data replication.
[0005] In a first aspect, an embodiment of the present application provides a dynamic and adaptive cross-cloud data replication method, the method comprising:
[0006] When a new cloud region is detected, the performance measurement module obtains performance data and the completion time required to execute a replication task to adjust a performance model. The performance model is used to determine the completion time required for function instances of different replication functions to execute replication tasks in different cloud regions based on the performance data of different cloud regions. The cloud region is either a cloud platform or a cloud region, and the replication function is the cloud function used to execute the replication task.
[0007] Before executing the target replication task, the strategy planning module determines the execution location and parallel replication strategy associated with the target replication task based on the completion time required for different replication function instances to execute the target replication task in different cloud regions according to the performance model, and triggers one or more replication function instances to run at the execution location according to the parallel replication strategy when executing the target replication task;
[0008] When executing the target replication task, the replication engine uses the function instances of the one or more replication functions to copy the target object data to be replicated in the source storage bucket to the target storage bucket. When there are multiple function instances of the replication function used, the replication engine schedules the idle function instances to perform the replication operation at the data block granularity until all data blocks in the target object data are replicated, and the target replication task is terminated. The source storage bucket and the target storage bucket are located in different cloud platforms.
[0009] A second aspect of an embodiment of the present application provides a dynamic and adaptive cross-cloud data replication system, the system comprising a performance measurement module, a policy planning module, and a replication engine, wherein:
[0010] The performance measurement module is configured to, upon detecting the access of a new cloud region, obtain performance data from the new cloud region to adjust a performance model. The performance model is configured to determine, based on the performance data of different cloud regions, the completion time required for function instances of different replication functions to execute replication tasks in different cloud regions. The cloud region is either a cloud platform or a cloud region, and the replication function is a cloud function used to execute replication tasks.
[0011] The strategy planning module is configured to determine, before executing the target replication task, the completion time required for function instances of different replication functions to execute the target replication task in different cloud regions as determined by the performance model, determine the execution location and parallel replication strategy associated with the target replication task, and, when executing the target replication task, trigger one or more function instances of the replication function to run at the execution location according to the parallel replication strategy;
[0012] The replication engine is used to, when executing the target replication task, copy the target object data to be replicated in the source storage bucket to the target storage bucket by using the function instances of the one or more replication functions, and when there are multiple function instances of the replication function used, schedule idle function instances to perform the replication operation at the granularity of data blocks until all data blocks in the target object data are replicated, terminating the target replication task. The source storage bucket and the target storage bucket are located in different cloud platforms.
[0013] A third aspect of an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the dynamic adaptive cross-cloud data replication method as described in the first aspect.
[0014] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the dynamic adaptive cross-cloud data replication method as described in the first aspect are implemented.
[0015] In a fifth aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the dynamic adaptive cross-cloud data replication method as described in the first aspect are implemented.
[0016] It can be seen from the above technical solution that this application realizes a cross-cloud data replication solution based on server-unaware computing by introducing replication functions, thereby meeting the needs of efficient, low-latency and cross-cloud data replication with the help of on-demand triggering, automatic expansion, event-driven and other characteristics of server-unaware computing; and in order to cope with the uncertainty brought about by using server-unaware computing (i.e. replication functions) to replicate data, this application designs a dynamic execution strategy planning solution based on performance analysis modeling, which reasonably selects the execution location and parallel replication strategy according to the performance of one or more function instances of replication functions in different cloud regions, and designs a distributed replication mechanism based on decentralized data block granularity scheduling, so that function instances of different replication functions can still complete the assigned data block replication tasks in a similar time under performance fluctuations, thereby improving the dynamic adaptability of cross-cloud data replication, thereby further ensuring the low latency and high efficiency of cross-cloud data replication. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A flowchart of an implementation method of a dynamic and adaptive cross-cloud data replication method provided in an embodiment of the present application;
[0019] Figure 2 A schematic diagram of distributed replication based on decentralized data block granularity scheduling provided in an embodiment of the present application;
[0020] Figure 3 A cross-cloud data replication flow chart provided in an embodiment of the present application;
[0021] Figure 4 A module diagram of a dynamic and adaptive cross-cloud data replication system provided in an embodiment of the present application;
[0022] Figure 5 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0024] Cross-cloud data replication is used to synchronize and distribute data between different cloud platforms to improve service availability, disaster recovery, and optimize business processing performance. Its importance is reflected in several aspects:
[0025] First, it can avoid the risk of being locked into a single cloud vendor and enhance enterprise flexibility, enabling them to switch freely between multiple clouds or leverage the advantages of different clouds simultaneously.
[0026] Second, cross-cloud replication can provide greater business continuity, ensuring data availability in the event of a cloud platform failure, thereby reducing downtime and the risk of data loss.
[0027] In addition, by replicating data in cloud data centers in different regions, enterprises can optimize data access speed, reduce latency, and meet compliance requirements, such as data localization regulations.
[0028] Therefore, cross-cloud data replication is not only one of the core capabilities in a distributed computing environment, but also a key means to achieve high availability and high elasticity in the cloud computing era.
[0029] Current data replication solutions are mainly divided into two types: cloud vendors' proprietary services and open source tools, but both types of solutions have limitations.
[0030] While cloud vendors' proprietary services are highly integrated, they typically only support replication between regions within the same cloud platform, making true cross-cloud data synchronization difficult. Furthermore, the replication delay of these services is often close to 15 minutes, which falls short of meeting the requirements of low-latency scenarios. Furthermore, they incur additional storage and data transfer costs, increasing overall operational overhead.
[0031] Open-source tools from cloud vendors, such as Skyplane (an open-source cross-cloud replication solution based on virtual machines), have alleviated some of the limitations of inter-cloud data migration. However, the core architecture of these solutions relies on virtual machine instances, which makes them slow to respond to dynamic load fluctuations. Experiments have shown that virtual machine initialization typically takes tens of seconds, making rapid replication of small files inefficient and making it difficult to elastically scale to respond to sudden traffic spikes.
[0032] Therefore, whether it is a cloud vendor's private service or an open source solution based on virtual machines, there are certain challenges in meeting the needs of efficient, low-latency, and cross-cloud data replication.
[0033] To solve the aforementioned problems, this application considers using server-aware computing that decouples computing and storage to replace virtual machines, so as to achieve cross-cloud data replication that is faster and cheaper than traditional solutions, thereby meeting the needs of efficient, low-latency and cross-cloud data replication.
[0034] It is understandable that cross-cloud data replication based on server-aware computing has significant advantages:
[0035] First, server-aware computing (i.e., cloud functions) is triggered on demand and automatically scales, allowing for flexible adjustment of computing resources as data traffic changes. This avoids the startup delays associated with traditional virtual machine solutions, thereby achieving lower replication latency.
[0036] Secondly, server-aware computing adopts an event-driven model, which can immediately trigger replication tasks when data is generated or updated, improving the real-time performance of data synchronization.
[0037] Thirdly, since server-aware computing platforms are usually managed by cloud vendors, users do not need to maintain the underlying infrastructure, which can significantly reduce operation and maintenance costs.
[0038] In addition, server-aware computing has a fine-grained billing model, paying only for the actual computing and data transmission performed, which is more cost-effective than the continuously running virtual machine solution.
[0039] Therefore, server-aware computing can not only provide higher performance and elasticity in cross-cloud data replication scenarios, but also reduce overall costs. It is an efficient and economical solution.
[0040] However, when server-aware computing is used for cross-cloud data replication, it faces two key issues (i.e., two uncertainties):
[0041] First, the ingress and egress bandwidth of data does not simply depend on the source and destination regions. The characteristics of different cloud platforms and cloud regions (i.e., regions within a cloud platform) will affect the actual network throughput. In other words, there is uncertainty in the bandwidth introduced by different cloud platforms and cloud regions, so multiple factors need to be considered comprehensively.
[0042] Secondly, the effective bandwidth of the server-unaware computing function fluctuates greatly between different instances, and this fluctuation has no fixed pattern, making it difficult to predict and optimize task scheduling in advance. In other words, there are large bandwidth differences between different cloud function instances.
[0043] Based on the above analysis, in response to the problem that traditional data replication solutions are difficult to meet the needs of efficient, low-latency and cross-cloud data replication, the embodiments of the present application provide a dynamically adaptive cross-cloud data replication method and system. By introducing a replication function, a cross-cloud data replication solution based on server-unaware computing is implemented, thereby meeting the needs of efficient, low-latency and cross-cloud data replication. In order to cope with the uncertainty brought about by using server-unaware computing to replicate data, a dynamic execution strategy planning solution based on performance analysis modeling is designed. The execution location and parallel replication strategy are reasonably selected according to the performance of function instances of one or more replication functions in different cloud regions. A distributed replication mechanism based on decentralized data block granularity scheduling is designed, so that function instances of different replication functions can still complete the assigned data block replication tasks in a similar time under performance fluctuations, thereby improving the dynamic adaptability of cross-cloud data replication, thereby further ensuring the low latency and high efficiency of cross-cloud data replication.
[0044] First, to facilitate understanding of the technical solutions provided by this application, the main technical concepts involved in the embodiments of this application are briefly described below.
[0045] Cloud platform: refers to the basic platform used to provide services such as elastic computing, storage, networking, security, and development tools, such as Alibaba Cloud, Amazon Web Services (AWS), and Google Cloud Platform (GCP).
[0046] Cloud data center: It is the physical infrastructure that supports the operation of the cloud platform. Each cloud platform includes multiple cloud data centers.
[0047] Cloud databases are database services deployed on cloud platforms, typically provided as "Database as a Service" (DBaaS), such as AWS's Relational Database Service (RDS) and Alibaba Cloud's PolarDB. Users can create and manage database instances through an application programming interface (API) or console without having to worry about the underlying hardware and operational details.
[0048] Serverless computing: This is a type of computing in which users do not need to manage details such as server deployment, operation and maintenance, and capacity expansion. They only need to focus on the computing model of the business logic code. Its resource scheduling, operating environment, and elastic scaling are completely managed and completed automatically by the cloud platform.
[0049] Cloud Function: Also known as server-insensitive computing function, it refers to a function that runs on a cloud platform and is triggered on demand. It is a representative service of server-insensitive computing.
[0050] A cloud function instance is a process or container instance that actually runs a cloud function code on the cloud platform, that is, the specific execution entity when the function is called.
[0051] Degree of parallelism: Also known as concurrency, it indicates the number of cloud function instances running simultaneously.
[0052] Entity Tag (Etag): refers to a unique identifier generated by the cloud platform for each version of data. When the data changes, the Etag will also be updated.
[0053] Object: It can be understood as a file. In object storage, a unit of data is usually called an object.
[0054] See also Figure 1 FIG. 1 is a flowchart illustrating an implementation of a dynamic adaptive cross-cloud data replication method provided by an embodiment of the present application. The method can be executed by a pre-established dynamic adaptive cross-cloud data replication system. The system mainly includes a performance measurement module, a policy planning module, and a replication engine. It can also include a cloud database and a logging module. The method may include the following steps:
[0055] Step S101: When a new cloud region is detected, the performance measurement module obtains its performance data and the completion time required to execute the replication task to adjust the performance model. The performance model is used to determine the completion time required for function instances of different replication functions to execute the replication task in different cloud regions based on the performance data of different cloud regions. The cloud region is any one of a cloud platform and a cloud area, and the replication function is a cloud function used to execute the replication task.
[0056] In specific implementation, when a cloud platform or user wants to connect a new cloud region or cloud platform (i.e., cloud region) to the system, the performance measurement module will first conduct offline performance measurement on it. That is, by performing a series of benchmark tests and data collection tasks, it will obtain the performance data of the newly connected cloud region (such as the computing power, storage throughput, and network latency of the cloud platform or cloud region) and the completion time required to execute the replication task in the newly connected cloud region (such as the call time, startup delay, and scheduling delay related to function startup, as well as the connection time, single data block transfer time, object size, and data block size related to data transmission).
[0057] The performance measurement module then adjusts the parameters of the performance model based on the acquired performance data and the completion time required to execute the replication task to adapt it to the characteristics of the new environment. That is, the performance model can automatically determine the completion time required to execute the replication task based on the performance data of the newly connected cloud region, so as to assist the policy planning module in reasonably determining the execution location and parallel replication strategy, thereby ensuring the efficient execution of the replication task.
[0058] Optionally, the performance model is a mathematical model designed based on experience, which expresses the completion time required to execute the replication task as a function of multiple variables. Performance data such as computing power, storage throughput, and network latency are all variables within the function. By substituting the performance data obtained by offline performance measurement into the function, the predicted result of the replication time (i.e., the completion time required to execute the replication task) can be obtained.
[0059] In this embodiment, the performance measurement module continuously optimizes the performance analysis process (ie, continuously optimizes the parameters of the performance model) based on offline performance measurement, which can enable the system to adapt to the new environment more quickly, thereby improving the reliability and predictability of data replication.
[0060] Step S102: Before executing the target replication task, the strategy planning module determines the completion time required for function instances of different replication functions to execute the target replication task in different cloud regions based on the performance model, determines the execution location and parallel replication strategy associated with the target replication task, and when executing the target replication task, triggers one or more function instances of the replication function to run at the execution location according to the parallel replication strategy.
[0061] During implementation, before executing the target replication task, the strategy planning module will first analyze the performance data such as computing resources and network conditions of the currently available cloud regions and cloud platforms (i.e., different cloud regions) based on the performance model, with the planning goal of meeting service quality objectives (such as ensuring that the replication task is completed within the specified time and performance requirements), determine the cloud region (i.e., execution location) that is most suitable for executing the target replication task, and determine the optimal parallel replication strategy.
[0062] Step S103: When executing the target replication task, the replication engine uses the function instances of the one or more replication functions to copy the target object data to be replicated in the source bucket to the target bucket. If there are multiple function instances of the replication function used, the replication engine schedules the idle function instances at the data block granularity to perform the replication operation until all data blocks in the target object data are replicated, thereby terminating the target replication task. The source bucket and the target bucket are located in different cloud platforms.
[0063] In specific implementation, the replication engine, as the core component responsible for data transmission in the system, can use one or more function instances of replication functions to copy objects or part of the object's data (that is, the object data to be replicated) from the source bucket to the target bucket to ensure that the data remains synchronized between different cloud regions or cloud platforms.
[0064] Considering that in a distributed replication environment, the execution performance of replicated functions may be affected by factors such as resource allocation and load fluctuations, this application's replication engine uses an autonomous scheduling mechanism (i.e., decentralized data block-granularity scheduling) to dynamically adjust replication strategies, thereby mitigating performance differences between different function instances. This mechanism improves replication efficiency, reduces latency, and provides more stable performance in multi-cloud or cross-region replication scenarios.
[0065] Reference Figure 2 The diagram of distributed replication based on decentralized data block granularity scheduling shown in the figure allocates a fixed set of data blocks to each replicated function instance upon invocation (e.g., uniform allocation). While this approach minimizes scheduling overhead, it only works well when function instance performance varies slightly. Concurrent function instances can have significantly different performance. For example, if replica 1 can replicate four data blocks per second, while replica 2 can only process one data block per second, the optimal allocation (which can be achieved using decentralized data block granularity scheduling) is to allocate five data blocks to replica 1 and three to replica 2, rather than allocating four data blocks to each function instance. Consequently, using traditional scheduling methods such as uniform allocation increases the likelihood of observing one or more extremely slow instances when many function instances with widely varying performance collaborate on distributed replication.
[0066] In order to achieve reasonable scheduling when the performance of the function instance is uncertain, this application designs the replication engine to schedule at the data block granularity instead of scheduling all data blocks at once.
[0067] Specifically, the replication engine adopts a decentralized data block granularity scheduling method, the core idea of which is to assign a data block to an available (i.e., idle) function instance as quickly as possible.
[0068] For example, the replication engine can create a shared external storage pool for storing data blocks. Whenever a function instance of a replication function is available, it will schedule the function instance (or the function instance will take the initiative) to obtain a data block from the pool and perform the replication operation. When all data blocks are copied, the function instance of the last replication function to complete the data block replication will end the entire replication task.
[0069] It's understandable that decentralized block-granular scheduling allows faster function instances to process more blocks, while slower instances process fewer blocks, resulting in more balanced execution times across function instances. Furthermore, this approach requires no interaction between participants and only two external storage accesses per block, achieving the desired effect with minimal scheduling overhead.
[0070] It can be seen from the above technical solution that this application realizes a cross-cloud data replication solution based on server-unaware computing by introducing replication functions, thereby meeting the needs of efficient, low-latency and cross-cloud data replication with the help of on-demand triggering, automatic expansion, event-driven and other characteristics of server-unaware computing; and in order to cope with the uncertainty brought about by using server-unaware computing (i.e. replication functions) to replicate data, this application designs a dynamic execution strategy planning solution based on performance analysis modeling, which reasonably selects the execution location and parallel replication strategy according to the performance of one or more function instances of replication functions in different cloud regions, and designs a distributed replication mechanism based on decentralized data block granularity scheduling, so that function instances of different replication functions can still complete the assigned data block replication tasks in a similar time under performance fluctuations, thereby improving the dynamic adaptability of cross-cloud data replication, thereby further ensuring the low latency and high efficiency of cross-cloud data replication.
[0071] As a possible implementation, the completion time required to execute the replication task is characterized by a set percentile of a distribution of task completion times; and the method further includes:
[0072] The performance measurement module collects performance data at multiple time points for each cloud region and uses a preset statistical method to fit the performance data at multiple time points associated with each cloud region to obtain the sample distribution associated with each cloud region.
[0073] The performance measurement module inputs the sample distribution associated with each cloud region into the performance model to obtain the task completion time distribution associated with each cloud region;
[0074] The performance measurement module sends the set percentile of the task completion time distribution associated with each cloud region to the strategy planning module to assist the strategy planning module in determining the execution location and parallel replication strategy;
[0075] The preset statistical method includes at least one of the following:
[0076] normal distribution;
[0077] The Monte Carlo method is used to generate sample distributions when there are multiple function instances of a replicated function executed in a cloud region, and the number of function instances is less than the set number of instances.
[0078] Extreme value theory is used to generate sample distribution when there are multiple function instances of a replicated function executed in a cloud region, and the number of function instances is not less than the set number of instances.
[0079] In this embodiment, considering that the execution performance of cloud functions is affected by the scheduling and operating environment, and the performance fluctuations of different cloud platforms and cloud regions are large, performance data cannot be represented by only a single deterministic value, but should be constructed through sampling to construct a distribution. Therefore, this application adds a distribution-aware feature to the performance measurement module to reflect the performance fluctuations of cloud functions under different operating conditions.
[0080] Specifically, the performance measurement module collects sufficient sample data (i.e., performance data) and uses statistical methods to fit the sample distribution to provide more accurate performance predictions.
[0081] Typically, these samples can be approximated as a normal distribution, as the calculation of function startup time and data transfer time primarily involves the weighted summation of multiple parameters, and the result of this calculation generally conforms to the central limit theorem. For parallel replication tasks, since the maximum task execution time for all replicated function instances must be estimated, the performance measurement module can use Monte Carlo methods to generate the data distribution when the number of replicated function instances is small. However, when the number of instances is large, to improve computational efficiency, extreme value theory can be used to approximate the maximum distribution to a Gumbel distribution, allowing the performance model to quickly calculate the required replication time distribution (i.e., the task completion time distribution).
[0082] After obtaining the task completion time distribution associated with each cloud region, the performance measurement module uses statistical tools to calculate a set percentile (e.g., the 99th percentile) of the task completion time distribution and sends this percentile to the strategy planning module to support optimization decisions for replication tasks. As a result, the performance model designed in this application can help the performance measurement module reasonably predict replication time, and in turn, help the strategy planning module develop the optimal replication strategy (i.e., the optimal execution location and parallel replication strategy) that meets service quality objectives.
[0083] Optionally, the performance measurement module can periodically measure (e.g., multiple sampling) the performance parameters of different cloud regions (i.e., different cloud platforms and cloud areas) and the completion time required to execute the replication task in the different cloud regions to analyze their distribution characteristics, such as using different statistical methods to fit different sample distributions to the sampled data, and determine the preferred statistical method based on the sample distribution with the highest prediction result accuracy (e.g., the difference between the predicted task completion time distribution and the actual task completion time distribution is the smallest).
[0084] As a possible implementation, the parallel copy strategy associated with the target copy task is characterized by a target parallelism, and the execution position and target parallelism associated with the target copy task are determined by the following steps:
[0085] Starting with a parallelism of 1, iterate through different degrees of parallelism exponentially. Using the performance model, determine the time required for the replication function instance to complete the target replication task in each cloud region at the current degree of parallelism. Each cloud region includes the starting cloud region where the source bucket is located and the ending cloud region where the target bucket is located.
[0086] If the current number of iterations does not reach the set number of iterations and the completion time determined by the performance model meets the set completion time, the iteration is terminated, and the cloud region and parallelism associated with the completion time determined by the performance model are respectively determined as the execution location and target parallelism associated with the target replication task;
[0087] When the current number of iterations reaches the set number of iterations, the cloud region and parallelism associated with the smallest one of the completion times determined by the performance model during the iteration process are respectively determined as the execution location and target parallelism associated with the target replication task.
[0088] In this embodiment, the performance model is taken into account so that the strategy planning module can compare the expected replication time of different replication strategies (i.e., different execution locations and degrees of parallelism), but the goal of the strategy planning module is to generate the lowest-cost replication strategy that meets the service quality goals rather than the fastest strategy, and it is not necessary to accurately calculate the cost of each strategy. This is because when more function instances are called, additional interface calls will be triggered, and additional performance overhead will also be generated, thereby increasing the total function execution time.
[0089] Therefore, the present application designs a dynamic replication strategy planning algorithm, which starts from a single-function strategy (i.e., the parallelism is 1) and iterates different parallelisms in an exponential form. For example, the iteration starts from a parallelism of 1, the parallelism of the next round of iteration is 2, the parallelism of the next round of iteration is 4, and so on. The present application does not impose any specific restrictions on the value of the parallelism that increases exponentially from 1 during the iteration process.
[0090] For each degree of parallelism, the algorithm evaluates strategies that use a replicated function in the originating cloud region and a replicated function in the destination cloud region. Once a strategy that meets the service quality target is found, the algorithm immediately returns that strategy without evaluating the remaining strategies. If no strategy meets the service quality target, the algorithm returns the fastest strategy.
[0091] For example, the algorithm takes the user-defined percentile p as input, and takes the percentile P of the task completion time distribution corresponding to the replication time t not less than the target replication time (i.e., the time when the user wants the replication task to be completed) T, which is not less than the user-defined percentile p as the service quality target, that is, For the service quality goal, when the performance model output meets The strategy that meets the service quality goal is determined based on the replication time t. In addition, since the strategy planning module in this example uses the Monte Carlo method to generate the sample distribution of the parallel replication function, in order to avoid violating the service quality goal due to slow distribution generation, the system limits the sampling time. Therefore, the pseudo code of the algorithm can be expressed as follows:
[0092] Algorithm: Dynamic Replication Strategy Planning
[0093] p: target percentile
[0094] n maz : Maximum concurrency
[0095] 1: function GENERATE_PLAN(obj, SLO e2e , P)
[0096] 2: SLO rep ← SLO e2e -(now- obj.timestamp)
[0097] 3:best_time, best_plan ← ∞, NULL
[0098] 4: for i = 1,2..,n max do
[0099] 5:for loc ∈ {src, dst}do
[0100] 6:time ← T rep (i, loc, obj, p)
[0101] 7:if time <best_time then
[0102] 8:best_time, best_plan←time, plan
[0103] 9:if best_time <SLO rep then
[0104] 10:return best_plan
[0105] 11:return best_plan
[0106] In the pseudocode above, line 2 calculates the target copy completion time, and lines 4-10 cycle through different concurrency levels and execution locations to find a solution that meets the target copy completion time. Once a feasible solution is found, polling is discontinued and the feasible solution is returned directly.
[0107] Optionally, when the parallelism is 1, the completion time required for the function instance of the replication function to execute the target replication task in each cloud region includes: the completion time required for the function instance of a single replication function to execute the target replication task in each cloud region;
[0108] The performance model is used to determine the completion time required for a function instance of a single replication function to execute the target replication task in each cloud region through the following steps:
[0109] When the data volume of the target object data is less than the first set data volume, determining that the function startup time corresponding to the function instance of the single replicated function in each cloud region is zero;
[0110] If the target object data size is not less than the first set data size, determining the function startup time corresponding to the function instance of the single replicated function in each cloud region based on the function call time and the startup delay of the function instance corresponding to the single replicated function in each cloud region;
[0111] The data transmission time corresponding to a single replicated function instance in each cloud region is determined as the sum of the time it takes for a function instance of a single replicated function to establish a data transmission connection in each cloud region and the total time required to transmit the target object data.
[0112] The sum of the function startup time and the data transmission time corresponding to the function instance of the single replication function in each cloud region is determined as the completion time required for the function instance of the single replication function to execute the target replication task in each cloud region.
[0113] In specific implementation, this application splits the replication time of a single replication function (ie, the completion time required for a function instance of a single replication function to execute the target replication task in each cloud region) into function startup time and data transmission time.
[0114] Function startup time depends on the object size. When the object is small, the copy can be handled locally and the startup time is zero. However, when the object is large, a copy function needs to be called, and the startup time is determined by the call time and the startup delay of the function instance.
[0115] Data transmission time is positively correlated with object size and consists primarily of the time it takes to establish a connection and the total time required to transmit the data. When scheduling at the data block granularity, the total time required to transmit data can be expressed as the transmission time of a single data block multiplied by the total number of data blocks in the object.
[0116] Optionally, when the parallelism is n, the completion time required for the function instance of the replication function to execute the target replication task in each cloud region includes: the completion time required for n function instances of the replication function to execute the target replication task in parallel in each cloud region, where n is an integer greater than 1;
[0117] The performance model is used to determine the completion time required for n replication function instances to execute the target replication task in parallel in each cloud region through the following steps:
[0118] Determine the function startup time for each cloud region for the n replicated function instances based on the function call time, function instance startup delay, and scheduling delay for each cloud region.
[0119] The maximum of the data transmission times corresponding to each of the n replicated function instances in each cloud region is determined as the data transmission time corresponding to the entire set of n replicated function instances in each cloud region. The data transmission time corresponding to a single replicated function instance in each cloud region is the sum of the time it takes for the single replicated function instance to establish a data transmission connection in each cloud region and the total time required to transmit the target object data.
[0120] The sum of the function startup time and data transmission time corresponding to the entire n function instances of the replicated function in each cloud region is determined as the completion time required for the n function instances of the replicated function to execute the target replication task in parallel in each cloud region.
[0121] In a specific implementation, considering that multiple function instances of a replication function are used for parallel replication, the estimation of replication time becomes more complicated.
[0122] First, due to the scheduling characteristics of cloud platforms, even though the calls to replicated functions are pipelined, the creation of function instances is still subject to scheduling delays. For example, some cloud platform schedulers may only create new instances every few seconds, so this factor must be taken into account when calculating function startup time.
[0123] In addition, the data transmission time in the parallel case depends on the execution time of the slowest function instance. Since the processing capabilities of different function instances may vary, in order to ensure that the replication task can meet the service quality requirements, a reasonable estimation method is to take the maximum execution time of all instances.
[0124] As a possible implementation, the parallel replication strategy associated with the target replication task is determined by the following steps:
[0125] When the data amount of the target object data is less than the second set data amount, determining the parallel replication strategy as: executing the target replication task through a function instance of a single replication function;
[0126] When the data volume of the target object data is not less than the second set data volume, the parallel replication strategy is determined to be: executing the target replication task in parallel through function instances of multiple replication functions.
[0127] In practice, small objects can be replicated using a single instance of the replication function, while larger objects require multiple instances of the replication function to execute in parallel to avoid timeouts and meet quality of service objectives. Therefore, the measurement and planning module can predetermine the parallel replication strategy based on object size, such as determining whether the degree of parallelism is greater than 1. It then determines the execution location and specific degree of parallelism based on the performance model, thereby improving planning efficiency.
[0128] As a possible implementation, the method further includes:
[0129] The strategy planning module receives an object event notification from the cloud platform, wherein the object event notification includes metadata of an object created or metadata of an object deleted in the cloud platform;
[0130] The strategy planning module determines, based on the object event notification, a cloud function required to trigger the function instance to run, where the cloud function required to trigger the function instance to run includes: at least one of a replication function, a cloud function for performing a logging task, and a cloud function for performing an access control policy update task;
[0131] The function instance of the copy function is used to perform the following steps:
[0132] Downloading object data from a source storage bucket and temporarily storing the object data in local storage according to the set storage environment requirements. The local storage includes any one of the temporary storage of the function execution environment, a distributed storage system, and a shared external storage pool pre-created by the replication engine. The shared external storage pool is used to store object data at data block granularity when there are multiple function instances of the replication function.
[0133] Read data from local storage and perform a data integrity check. If the data integrity check passes, upload the read data to the target storage bucket using the specified transfer method.
[0134] After the data is successfully uploaded to the target storage bucket, a setting operation is triggered, where the setting operation includes at least one of recording a replication log, sending a status notification, and updating object location information in the cloud database.
[0135] For example, referring to Figure 3 The cross-cloud data replication process shown in the figure shows that when an object is created and replicated from one cloud region to another (that is, from cloud region 1 to cloud region 2), the cross-cloud data replication process can be divided into the following four stages:
[0136] ①Cloud notification:
[0137] When an object is created or deleted, the cloud platform generates an object event notification in JSON format, which contains the object's metadata, such as the object's name, storage location, size, creation time, and other key information. The strategy planning module triggers the corresponding cloud function based on this information for subsequent processing, such as data replication, logging, or updating access control policies. Among them, the strategy planning module runs in the listener. When the strategy planning module determines that the cloud function required to trigger the function instance to run includes a replication function, it can make decisions on the execution location and parallel replication strategy, and start the replication function accordingly to run the replication engine therein.
[0138] ②Function call:
[0139] After receiving the cloud notification, one or more function instances of the replicated function will be triggered to run. These function instances can be in the starting cloud region where the source bucket is located (such as Figure 3 The cloud region where the bucket 1 is located can also be run in the destination cloud region where the target bucket is located (such as Figure 3 Bucket 2 (shown) is located in cloud region 2. The specific operation depends on the system's architecture design and optimization strategy. The replication function's primary task is to coordinate the data replication process, including parsing notifications, determining replication policies, and performing access rights verification. Once the function is successfully invoked, the data replication process officially begins.
[0140] ③Data download:
[0141] The function instance of the replication function downloads the object data from the source storage bucket in the Qidian Cloud region and temporarily stores it in local storage. This local storage can be the temporary storage of the function execution environment, a distributed storage system, or a shared external storage pool pre-created by the replication engine to ensure data integrity during transmission. During this stage, depending on the requirements of the configured storage environment, the function instance of the replication function may directly store the data or may require additional data processing before data storage, such as data decompression, format conversion, or encryption and decryption, to meet the needs of different storage environments.
[0142] ④Data upload:
[0143] The function instance of the replication function reads data from local storage and uploads it to the target bucket. During the upload process, the function instance may need to perform data integrity checks, such as calculating hash values to ensure that the data has not been tampered with. In addition, to optimize cross-region transmission performance, the system may adopt set transmission methods such as multi-threaded parallel upload, block transmission, or streaming transmission to reduce latency and improve throughput. When the data is successfully uploaded to the target bucket, the function instance may trigger the relevant system modules to perform set operations, such as recording replication logs, sending status notifications, or performing additional follow-up operations, such as updating object location information in the database, to ensure the traceability and consistency of the entire replication process.
[0144] Through the coordinated efforts of these four phases, data can be efficiently and reliably replicated across different cloud regions, ensuring high availability and data consistency across clouds. This process can be used in a variety of scenarios, such as cross-region disaster recovery, global distribution acceleration, and data lake synchronization, to meet diverse business needs.
[0145] As a possible implementation, the method further includes:
[0146] When executing a distributed replication task, the replication engine verifies whether the entity tag ETag of the object data currently to be replicated matches the ETag provided by the strategy planning module. If the two do not match, the currently executed replication task is terminated. The distributed replication task includes: multiple replication tasks, and the multiple replication tasks are used to copy the object data to be replicated in the same source storage bucket to multiple target storage buckets.
[0147] In this embodiment, in order to avoid replicating inconsistent object parts, the replication engine will verify whether the current ETag matches the ETag provided by the strategy planning module during distributed replication. If a mismatch is found, the current replication task will be terminated.
[0148] As a possible implementation, the method further includes:
[0149] When executing a distributed replication task, the replication engine sets a distributed lock to ensure the serialization of the replication tasks corresponding to each of the multiple target buckets. When the current replication task ends and the lock is not released, the ETag of the latest version of the object data in the corresponding source bucket is compared with the ETag of the most recently replicated object data. If the two do not match, the strategy planning module is called again to replicate the latest version of the object data.
[0150] In this embodiment, to ensure consistency, the present application designs an object-granularity replication task lock and an optimistic replication and verification mechanism for the replication engine.
[0151] Specifically, in order to prevent concurrent upload operations (ie, PUT operations) in distributed replication tasks, this application introduces a distributed lock to the replication engine to ensure the serialization of replication tasks.
[0152] When there is an ongoing replication task, the replication engine tracks the ETag of the latest version of the data in each replication task (ETag is basically equivalent to the content hash value of the object and is automatically generated by the cloud platform).
[0153] Before releasing the lock at the end of the replication task, the replication engine compares the pending ETag (i.e., the ETag of the latest version of the object data in the source bucket) with the ETag of the most recently replicated data. If the two do not match, it means that the latest version of the data has not been replicated. At this time, the replication engine will call the strategy planning module again to ensure that the latest version of the data is replicated.
[0154] As a possible implementation, the method further includes:
[0155] The strategy planning module adjusts the execution location and / or parallel replication strategy associated with the target replication task according to the current system status and real-time load conditions.
[0156] In this implementation, the strategy planning module not only relies on existing performance data to make decisions, but also dynamically adjusts the replication strategy based on the current system state and real-time load. For example, if a cloud region is detected to be highly loaded, the strategy planning module may select an alternate cloud region to execute the replication task and / or optimize task execution efficiency by adjusting the degree of parallelism.
[0157] As a possible implementation, the method further includes:
[0158] The cloud database stores the runtime state during the execution of the cloud function instance to maintain the intermediate state information of the cloud function.
[0159] In this embodiment, considering that cloud functions cannot communicate with each other, a cloud database is used as external storage to manage the status of cloud functions.
[0160] As a possible implementation, the method further includes:
[0161] The logging module records operational data, which is used for troubleshooting and subsequent optimization.
[0162] Based on the above embodiments, this application designs a dynamic and adaptive cross-cloud data replication method, which can cope with performance differences and fluctuations between different cloud computing vendors, different cloud regions, and different function instances, and achieve low-latency and low-cost data replication.
[0163] On the one hand, this method can accurately select cloud regions with better performance and parallelism to execute replication tasks through performance modeling, thereby avoiding slower cloud regions or reducing the impact of slower regions by increasing parallelism.
[0164] On the other hand, this method implements distributed replication with decentralized data block granularity scheduling, effectively alleviating the replication delay caused by performance differences between function instances.
[0165] To verify the performance of this method, the researchers of this application implemented and evaluated a related system prototype, and compared it with Skyplane, which uses virtual machines for replication, and data replication tools provided by cloud computing vendors. The experimental results show that this method can effectively reduce replication time compared to traditional methods.
[0166] It should be noted that for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0167] The present application also provides a dynamic and adaptive cross-cloud data replication system. Figure 4 As shown,
[0168] The system includes a performance measurement module, a strategy planning module and a replication engine, wherein:
[0169] The performance measurement module is configured to, upon detecting the access of a new cloud region, obtain performance data from the new cloud region to adjust a performance model. The performance model is configured to determine, based on the performance data of different cloud regions, the completion time required for function instances of different replication functions to execute replication tasks in different cloud regions. The cloud region is either a cloud platform or a cloud region, and the replication function is a cloud function used to execute replication tasks.
[0170] The strategy planning module is configured to determine, before executing the target replication task, the completion time required for function instances of different replication functions to execute the target replication task in different cloud regions as determined by the performance model, determine the execution location and parallel replication strategy associated with the target replication task, and, when executing the target replication task, trigger one or more function instances of the replication function to run at the execution location according to the parallel replication strategy;
[0171] The replication engine is used to, when executing the target replication task, copy the target object data to be replicated in the source storage bucket to the target storage bucket by using the function instances of the one or more replication functions, and when there are multiple function instances of the replication function used, schedule idle function instances to perform the replication operation at the granularity of data blocks until all data blocks in the target object data are replicated, terminating the target replication task. The source storage bucket and the target storage bucket are located in different cloud platforms.
[0172] Optionally, the completion time required to execute the replication task is characterized by a set percentile of the task completion time distribution; and the performance measurement module is further configured to perform the following steps:
[0173] Performance data at multiple time points is collected for each cloud region. Pre-set statistical methods are used to fit the performance data at multiple time points associated with each cloud region to obtain the sample distribution associated with each cloud region.
[0174] Inputting the sample distribution associated with each cloud region into the performance model to obtain the task completion time distribution associated with each cloud region;
[0175] Send the set percentile of the task completion time distribution associated with each cloud region to the strategy planning module to assist the strategy planning module in determining the execution location and parallel replication strategy;
[0176] The preset statistical method includes at least one of the following:
[0177] normal distribution;
[0178] The Monte Carlo method is used to generate sample distributions when there are multiple function instances of a replicated function executed in a cloud region, and the number of function instances is less than the set number of instances.
[0179] Extreme value theory is used to generate sample distribution when there are multiple function instances of a replicated function executed in a cloud region, and the number of function instances is not less than the set number of instances.
[0180] Optionally, the parallel copy strategy associated with the target copy task is characterized by a target parallelism, and the execution position and target parallelism associated with the target copy task are determined by the following steps:
[0181] Starting with a parallelism of 1, iterate through different degrees of parallelism exponentially. Using the performance model, determine the time required for the replication function instance to complete the target replication task in each cloud region at the current degree of parallelism. Each cloud region includes the starting cloud region where the source bucket is located and the ending cloud region where the target bucket is located.
[0182] If the current number of iterations does not reach the set number of iterations and the completion time determined by the performance model meets the set completion time, the iteration is terminated, and the cloud region and parallelism associated with the completion time determined by the performance model are respectively determined as the execution location and target parallelism associated with the target replication task;
[0183] When the current number of iterations reaches the set number of iterations, the cloud region and parallelism associated with the smallest one of the completion times determined by the performance model during the iteration process are respectively determined as the execution location and target parallelism associated with the target replication task.
[0184] Optionally, when the parallelism is 1, the completion time required for the function instance of the replication function to execute the target replication task in each cloud region includes: the completion time required for the function instance of a single replication function to execute the target replication task in each cloud region;
[0185] The performance model is used to determine the completion time required for a function instance of a single replication function to execute the target replication task in each cloud region through the following steps:
[0186] When the data volume of the target object data is less than the first set data volume, determining that the function startup time corresponding to the function instance of the single replicated function in each cloud region is zero;
[0187] If the target object data size is not less than the first set data size, determining the function startup time corresponding to the function instance of the single replicated function in each cloud region based on the function call time and the startup delay of the function instance corresponding to the single replicated function in each cloud region;
[0188] The data transmission time corresponding to a single replicated function instance in each cloud region is determined as the sum of the time it takes for a function instance of a single replicated function to establish a data transmission connection in each cloud region and the total time required to transmit the target object data.
[0189] The sum of the function startup time and the data transmission time corresponding to the function instance of the single replication function in each cloud region is determined as the completion time required for the function instance of the single replication function to execute the target replication task in each cloud region.
[0190] Optionally, when the parallelism is n, the completion time required for the function instance of the replication function to execute the target replication task in each cloud region includes: the completion time required for n function instances of the replication function to execute the target replication task in parallel in each cloud region, where n is an integer greater than 1;
[0191] The performance model is used to determine the completion time required for n replication function instances to execute the target replication task in parallel in each cloud region through the following steps:
[0192] Determine the function startup time for each cloud region for the n replicated function instances based on the function call time, function instance startup delay, and scheduling delay for each cloud region.
[0193] The maximum of the data transmission times corresponding to each of the n replicated function instances in each cloud region is determined as the data transmission time corresponding to the entire set of n replicated function instances in each cloud region. The data transmission time corresponding to a single replicated function instance in each cloud region is the sum of the time it takes for the single replicated function instance to establish a data transmission connection in each cloud region and the total time required to transmit the target object data.
[0194] The sum of the function startup time and data transmission time corresponding to the entire n function instances of the replicated function in each cloud region is determined as the completion time required for the n function instances of the replicated function to execute the target replication task in parallel in each cloud region.
[0195] Optionally, the parallel replication strategy associated with the target replication task is determined by the following steps:
[0196] When the data amount of the target object data is less than the second set data amount, determining the parallel replication strategy as: executing the target replication task through a function instance of a single replication function;
[0197] When the data volume of the target object data is not less than the second set data volume, the parallel replication strategy is determined to be: executing the target replication task in parallel through function instances of multiple replication functions.
[0198] Optionally, the strategy planning module is further configured to perform the following steps:
[0199] Receiving an object event notification from the cloud platform, the object event notification including metadata of an object created or deleted in the cloud platform;
[0200] Determine, based on the object event notification, a cloud function required to trigger the function instance to run, where the cloud function required to trigger the function instance to run includes: at least one of a replication function, a cloud function for performing a logging task, and a cloud function for performing an access control policy update task;
[0201] The function instance of the copy function is used to perform the following steps:
[0202] Downloading object data from a source storage bucket and temporarily storing the object data in local storage according to the set storage environment requirements. The local storage includes any one of the temporary storage of the function execution environment, a distributed storage system, and a shared external storage pool pre-created by the replication engine. The shared external storage pool is used to store object data at data block granularity when there are multiple function instances of the replication function.
[0203] Read data from local storage and perform a data integrity check. If the data integrity check passes, upload the read data to the target storage bucket using the specified transfer method.
[0204] After the data is successfully uploaded to the target storage bucket, a setting operation is triggered, where the setting operation includes at least one of recording a replication log, sending a status notification, and updating object location information in the cloud database.
[0205] Optionally, the replication engine is further configured to perform at least one of the following steps:
[0206] When executing a distributed replication task, verify whether the entity tag ETag of the object data currently to be replicated matches the ETag provided by the strategy planning module. If the two do not match, terminate the currently executed replication task. The distributed replication task includes: multiple replication tasks, each of which is used to replicate the object data to be replicated in the same source bucket to multiple target buckets;
[0207] When executing a distributed replication task, a distributed lock is set to ensure the serialization of the replication tasks corresponding to each of the multiple target buckets. When the current replication task ends and the lock is not released, the ETag of the latest version of the object data in the corresponding source bucket is compared with the ETag of the most recently replicated object data. If the two do not match, the strategy planning module is called again to replicate the latest version of the object data.
[0208] Optionally, the strategy planning module is further configured to adjust the execution location and / or parallel replication strategy associated with the target replication task according to the current system state and real-time load conditions;
[0209] Optionally, the system further includes at least one of the following:
[0210] The cloud database is used to store the runtime state of the cloud function instance during task execution to maintain the intermediate state information of the cloud function;
[0211] The logging module is used to record operational data, which is used for troubleshooting and subsequent optimization.
[0212] It can be seen from the above technical solution that this application realizes a cross-cloud data replication solution based on server-unaware computing by introducing replication functions, thereby meeting the needs of efficient, low-latency and cross-cloud data replication with the help of on-demand triggering, automatic expansion, event-driven and other characteristics of server-unaware computing; and in order to cope with the uncertainty brought about by using server-unaware computing (i.e. replication functions) to replicate data, this application designs a dynamic execution strategy planning solution based on performance analysis modeling, which reasonably selects the execution location and parallel replication strategy according to the performance of one or more function instances of replication functions in different cloud regions, and designs a distributed replication mechanism based on decentralized data block granularity scheduling, so that function instances of different replication functions can still complete the assigned data block replication tasks in a similar time under performance fluctuations, thereby improving the dynamic adaptability of cross-cloud data replication, thereby further ensuring the low latency and high efficiency of cross-cloud data replication.
[0213] The present application also provides an electronic device, Figure 5 , Figure 5 Schematic diagram of the electronic device proposed in the embodiment of the present application. Figure 5 As shown, the electronic device 100 includes: a memory 110 and a processor 120. The memory 110 and the processor 120 are connected via a bus communication. A computer program is stored in the memory 110. The computer program can be run on the processor 120, thereby implementing the steps in the dynamic adaptive cross-cloud data replication method disclosed in the embodiment of the present application.
[0214] An embodiment of the present application also provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the dynamic adaptive cross-cloud data replication method disclosed in the embodiment of the present application is implemented.
[0215] An embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the dynamic adaptive cross-cloud data replication method disclosed in the embodiment of the present application.
[0216] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0217] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0218] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, systems, devices, storage media, and program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0219] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0220] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1A step that specifies a function in one or more boxes.
[0221] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements that are inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0222] The above is a detailed introduction to a dynamic and adaptive cross-cloud data replication method and system provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A dynamic and adaptive cross-cloud data replication method, characterized in that: The method comprises: When a new cloud region is detected, the performance measurement module obtains performance data and the completion time required to execute a replication task to adjust a performance model. The performance model is used to determine the completion time required for function instances of different replication functions to execute replication tasks in different cloud regions based on the performance data of different cloud regions. The cloud region is either a cloud platform or a cloud region, and the replication function is the cloud function used to execute the replication task. Before executing the target replication task, the strategy planning module determines the execution location and parallel replication strategy associated with the target replication task based on the completion time required for different replication function instances to execute the target replication task in different cloud regions according to the performance model, and triggers one or more replication function instances to run at the execution location according to the parallel replication strategy when executing the target replication task; When executing the target replication task, the replication engine uses the function instances of the one or more replication functions to copy the target object data to be replicated in the source storage bucket to the target storage bucket. When there are multiple function instances of the replication function used, the replication engine schedules the idle function instances to perform the replication operation at the data block granularity until all data blocks in the target object data are replicated, and the target replication task is terminated. The source storage bucket and the target storage bucket are located in different cloud platforms.
2. The method according to claim 1, characterized in that The completion time required to execute the replication task is characterized by using a set percentile of a distribution of task completion times; the method further comprising: The performance measurement module collects performance data at multiple time points for each cloud region and uses a preset statistical method to fit the performance data at multiple time points associated with each cloud region to obtain the sample distribution associated with each cloud region. The performance measurement module inputs the sample distribution associated with each cloud region into the performance model to obtain the task completion time distribution associated with each cloud region; The performance measurement module sends the set percentile of the task completion time distribution associated with each cloud region to the strategy planning module to assist the strategy planning module in determining the execution location and parallel replication strategy; The preset statistical method includes at least one of the following: normal distribution; The Monte Carlo method is used to generate sample distributions when there are multiple function instances of a replicated function executed in a cloud region, and the number of function instances is less than the set number of instances. Extreme value theory is used to generate sample distribution when there are multiple function instances of a replicated function executed in a cloud region, and the number of function instances is not less than the set number of instances.
3. The method according to claim 1, characterized in that The parallel copy strategy associated with the target copy task is characterized by a target parallelism. The execution position and target parallelism associated with the target copy task are determined by the following steps: Starting with a parallelism of 1, iterate through different degrees of parallelism exponentially. Using the performance model, determine the time required for the replication function instance to complete the target replication task in each cloud region at the current degree of parallelism. Each cloud region includes the starting cloud region where the source bucket is located and the ending cloud region where the target bucket is located. If the current number of iterations does not reach the set number of iterations and the completion time determined by the performance model meets the set completion time, the iteration is terminated, and the cloud region and parallelism associated with the completion time determined by the performance model are respectively determined as the execution location and target parallelism associated with the target replication task; When the current number of iterations reaches the set number of iterations, the cloud region and parallelism associated with the smallest one of the completion times determined by the performance model during the iteration process are respectively determined as the execution location and target parallelism associated with the target replication task.
4. The method according to claim 3, characterized in that When the parallelism is 1, the completion time required for the function instance of the replication function to execute the target replication task in each cloud region includes: the completion time required for the function instance of a single replication function to execute the target replication task in each cloud region; The performance model is used to determine the completion time required for a function instance of a single replication function to execute the target replication task in each cloud region through the following steps: When the data volume of the target object data is less than the first set data volume, determining that the function startup time corresponding to the function instance of the single replicated function in each cloud region is zero; If the target object data size is not less than the first set data size, determining the function startup time corresponding to the function instance of the single replicated function in each cloud region based on the function call time and the startup delay of the function instance corresponding to the single replicated function in each cloud region; The data transmission time corresponding to a single replicated function instance in each cloud region is determined as the sum of the time it takes for a function instance of a single replicated function to establish a data transmission connection in each cloud region and the total time required to transmit the target object data. The sum of the function startup time and the data transmission time corresponding to the function instance of the single replication function in each cloud region is determined as the completion time required for the function instance of the single replication function to execute the target replication task in each cloud region.
5. The method according to claim 3, characterized in that The completion time required for a function instance of the replication function to execute the target replication task in each cloud region when the parallelism is n includes: the completion time required for n function instances of the replication function to execute the target replication task in parallel in each cloud region, where n is an integer greater than 1; The performance model is used to determine the completion time required for n replication function instances to execute the target replication task in parallel in each cloud region through the following steps: Determine the function startup time for each cloud region for the n replicated function instances based on the function call time, function instance startup delay, and scheduling delay for each cloud region. The maximum of the data transmission times corresponding to each of the n replicated function instances in each cloud region is determined as the data transmission time corresponding to the entire set of n replicated function instances in each cloud region. The data transmission time corresponding to a single replicated function instance in each cloud region is the sum of the time it takes for the single replicated function instance to establish a data transmission connection in each cloud region and the total time required to transmit the target object data. The sum of the function startup time and data transmission time corresponding to the entire n function instances of the replicated function in each cloud region is determined as the completion time required for the n function instances of the replicated function to execute the target replication task in parallel in each cloud region.
6. The method according to claim 1, characterized in that The parallel replication strategy associated with the target replication task is determined by the following steps: When the data amount of the target object data is less than the second set data amount, determining the parallel replication strategy as: executing the target replication task through a function instance of a single replication function; When the data volume of the target object data is not less than the second set data volume, the parallel replication strategy is determined to be: executing the target replication task in parallel through function instances of multiple replication functions.
7. The method according to claim 1, characterized in that The method further comprises: The strategy planning module receives an object event notification from the cloud platform, wherein the object event notification includes metadata of an object created or metadata of an object deleted in the cloud platform; The strategy planning module determines, based on the object event notification, a cloud function required to trigger the function instance to run, where the cloud function required to trigger the function instance to run includes: at least one of a replication function, a cloud function for performing a logging task, and a cloud function for performing an access control policy update task; The function instance of the copy function is used to perform the following steps: Downloading object data from a source storage bucket and temporarily storing the object data in local storage according to the set storage environment requirements. The local storage includes any one of the temporary storage of the function execution environment, a distributed storage system, and a shared external storage pool pre-created by the replication engine. The shared external storage pool is used to store object data at data block granularity when there are multiple function instances of the replication function. Read data from local storage and perform a data integrity check. If the data integrity check passes, upload the read data to the target storage bucket using the specified transfer method. After the data is successfully uploaded to the target storage bucket, a setting operation is triggered, where the setting operation includes at least one of recording a replication log, sending a status notification, and updating object location information in the cloud database.
8. The method according to claim 1, characterized in that The method further comprises at least one of the following steps: When executing a distributed replication task, the replication engine verifies whether the entity tag ETag of the object data currently to be replicated matches the ETag provided by the strategy planning module. If the two do not match, the currently executed replication task is aborted. The distributed replication task includes: multiple replication tasks, each of which is used to replicate the object data to be replicated in the same source bucket to multiple target buckets; When executing a distributed replication task, the replication engine sets a distributed lock to ensure the serialization of the replication tasks corresponding to each of the multiple target buckets. When the current replication task ends and the lock is not released, the ETag of the latest version of the object data in the corresponding source bucket is compared with the ETag of the most recently replicated object data. If the two do not match, the strategy planning module is called again to replicate the latest version of the object data.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises at least one of the following steps: The strategy planning module adjusts the execution location and / or parallel replication strategy associated with the target replication task according to the current system status and real-time load conditions; The cloud database stores the runtime state of the cloud function instance during task execution to maintain the intermediate state information of the cloud function. The logging module records operational data, which is used for troubleshooting and subsequent optimization.
10. A dynamic and adaptive cross-cloud data replication system, characterized in that: The system includes a performance measurement module, a strategy planning module and a replication engine, wherein: The performance measurement module is configured to, upon detecting the access of a new cloud region, obtain performance data from the new cloud region to adjust a performance model. The performance model is configured to determine, based on the performance data of different cloud regions, the completion time required for function instances of different replication functions to execute replication tasks in different cloud regions. The cloud region is either a cloud platform or a cloud region, and the replication function is a cloud function used to execute replication tasks. The strategy planning module is configured to determine, before executing the target replication task, the completion time required for function instances of different replication functions to execute the target replication task in different cloud regions as determined by the performance model, determine the execution location and parallel replication strategy associated with the target replication task, and, when executing the target replication task, trigger one or more function instances of the replication function to run at the execution location according to the parallel replication strategy; The replication engine is used to, when executing the target replication task, copy the target object data to be replicated in the source storage bucket to the target storage bucket by using the function instances of the one or more replication functions, and when there are multiple function instances of the replication function used, schedule idle function instances to perform the replication operation at the granularity of data blocks until all data blocks in the target object data are replicated, terminating the target replication task. The source storage bucket and the target storage bucket are located in different cloud platforms.
Citation Information
Patent Citations
Cross-cloud online cloud host migration method, migration controller and cloud server
CN111797059A
Disaster recovery methods and system
WO2025050946A1