Method and apparatus for processing data object, and storage medium

By migrating cold data objects based on water level differences in the cloud storage system, the problem of uneven storage capacity is solved, achieving efficient utilization of storage resources and cost reduction.

WO2025241490A1PCT designated stage Publication Date: 2025-11-27HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/138105
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-17
Filing Date
2024-12-10
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

The uneven storage capacity in different regions of the cloud storage system leads to either excess or insufficient storage capacity in some regions, increasing operating costs.

Method used

The system obtains the storage cluster water level of each region through the control system, selects the high water level region (source region) to migrate cold data objects to the low water level region (destination region) to balance storage capacity, releases the storage capacity of the source region and utilizes the idle capacity of the destination region.

Benefits of technology

It effectively reduces the operating costs of cloud storage systems, reduces the need for expansion by balancing storage capacity utilization, and improves resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138105_27112025_PF_FP_ABST
    Figure CN2024138105_27112025_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the field of storage. Disclosed are a method and apparatus for processing a data object, and a storage medium. The method is applied to a control system, wherein the control system is used for performing capacity balancing on a plurality of storage clusters, which are comprised in a cloud storage system, the plurality of storage clusters are distributed in a plurality of regions, and each region comprises at least one storage cluster. The method comprises: acquiring first watermarks of a plurality of first regions; on the basis of the first watermarks of the plurality of first regions, determining a first source region and a first destination region, wherein the first watermark of the first source region is higher than the first watermark of the first destination region; and on the basis of at least one storage cluster in the first source region, migrating a cold data object, which is stored in at least one storage cluster in the first source region, to at least one storage cluster in the first destination region. By means of the present application, the operation cost of a cloud storage system can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device and storage medium for processing data objects

[0001] The present application claims priority to Chinese Patent Application No. 202410637811.6, filed on May 20, 2024, and entitled “Method, device and other equipment for data processing”, the contents of which are incorporated herein by reference in its entirety. In addition, the present application claims priority to Chinese Patent Application No. 202410781902.7, filed on June 17, 2024, and entitled “Method, device and storage medium for processing data objects”, the contents of which are incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the field of storage, and in particular to a method, device and storage medium for processing data objects. BACKGROUND

[0003] Object storage services have the need to store data objects in different regions. In order to meet the needs of object storage services, cloud storage systems deploy storage clusters in multiple regions, and one or more storage clusters are deployed in each region. The storage capacity of each region is planned in advance according to the needs of object storage services, and the storage capacity of a region is equal to the sum of the storage capacities of each storage cluster included in the region.

[0004] However, there is a certain degree of uncertainty in advance planning. For example, there may be a large amount of idle storage capacity in some regions due to small user usage, resulting in excess storage capacity. Alternatively, there may be a large amount of storage capacity used in some regions due to large user usage, resulting in insufficient storage capacity.

[0005] Therefore, some regions in the current cloud storage system have excess storage capacity, and some regions have insufficient storage capacity, which increases the operating cost of the cloud storage system. SUMMARY

[0006] The present application provides a method, device and storage medium for processing data objects to reduce the operating cost of the cloud storage system. The technical solution is as follows:

[0007] In a first aspect, the present application provides a method for processing data objects, the method is applied to a control system, the control system is used for capacity balancing of a plurality of storage clusters included in a cloud storage system, the plurality of storage clusters are distributed in a plurality of regions, each region includes at least one storage cluster. In the method, a first water level of each region is obtained, the first water level of a region is used to indicate the used storage capacity of the region, and the storage capacity of the region includes the storage capacity of each storage cluster in the region. A first source region and a first target region are determined based on the first water level of the plurality of regions, the first water level of the first source region is higher than the first water level of the first target region. Based on at least one storage cluster in the first source region, the cold data objects stored in the at least one storage cluster in the first source region are migrated to at least one storage cluster in the first target region.

[0008] In the method, the first water level of each region is obtained, and the first source region and the first target region are determined based on the first water level of the plurality of regions. The first water level of a region is used to indicate the used storage capacity of the region, the first water level of the first source region is higher than the first water level of the first target region, so the storage capacity of the first source region is largely used, and the first target region has a large amount of idle storage capacity. The cold data objects stored in at least one storage cluster in the first source region are migrated to at least one storage cluster in the first target region, so as to release part of the used storage capacity in the first source region, thereby avoiding the expansion of the first source region, and fully utilizing part of the idle storage capacity of the first target region, thereby reducing the operation cost of the cloud storage system.

[0009] In a possible implementation, the first source region is at least one region with the highest first water level among the plurality of regions, or the first source region is at least one region with a first water level higher than a first threshold among the plurality of regions. Thus, it is ensured that the determined first source region is a region with largely used storage capacity among the plurality of regions, and the migration of the cold data objects in the first source region effectively reduces the operation cost of the cloud system.

[0010] In another possible implementation, the first target region is at least one region of the plurality of regions having a first lowest water level, or the first target region is at least one region of the plurality of regions having a first water level lower than a second threshold, the second threshold being less than or equal to the first threshold. Thus, it is ensured that the determined first target region is a region of the plurality of regions having a large amount of free storage capacity, and migrating the cold data object to the first target region will not cause the used storage capacity of the first target region to be too high, and the used storage capacities of the plurality of regions can be effectively balanced.

[0011] In another possible implementation, the access amount of the data objects included in the plurality of regions from other regions is obtained. The second source region and the second target region are determined based on the access amount, at least one storage cluster in the second source region includes the target data object, and the access amount from the second target region to the target data object exceeds a third threshold. The target data object is stored into at least one storage cluster included in the second target region based on the at least one storage cluster in the second source region.

[0012] Since the access amount from the second target region to the target data object exceeds the third threshold, it indicates that a large number of users in the second target region access the target data object in the second source region, and storing the target data object into at least one storage cluster included in the second target region can enable the users in the second target region to access the target data object in the second target region locally, thereby reducing the latency of accessing the target data object.

[0013] In a possible implementation, the at least one storage cluster in each region includes a plurality of storage buckets, and each storage bucket is configured to store at least one data object. A storage capacity of each storage cluster in the region is obtained, and a capacity of each storage bucket is obtained, which is equal to a sum of data amounts of each data object stored in the storage bucket. A used storage capacity of the region is obtained based on the capacities of the plurality of storage buckets, and a total storage capacity of the region is obtained based on the storage capacities of each storage cluster included in the region. A first water level of the region is obtained based on the used storage capacity of the region and the total storage capacity of the region. In this way, the first water level of the region can be accurately obtained, and the first water level can indicate a proportion of the used storage capacity of the region. The first source region and the first target region are obtained based on the first water level, and the cold data objects in the first source region can be migrated to the first target region, so that the proportions of the used storage capacities of the plurality of regions can be effectively balanced.

[0014] In a possible implementation, a total data amount of the cold data objects stored in each storage bucket included in the at least one storage cluster in the first source region is obtained. At least one first storage bucket with a maximum total data amount of the cold data objects stored in the first storage bucket, or with a total data amount of the cold data objects stored in the first storage bucket greater than a fourth threshold, is selected based on the total data amount of the cold data objects stored in each storage bucket. The cold data objects stored in the at least one first storage bucket are migrated to at least one second storage bucket included in the at least one storage cluster in the first target region, the first storage bucket and the second storage bucket belong to a same user, the first storage bucket is any one of the at least one first storage bucket, and the second storage bucket is any one of the at least one second storage bucket.

[0015] The at least one first storage bucket is the storage bucket with the maximum total data amount of the cold data objects stored in the storage bucket, or with the total data amount of the cold data objects stored in the storage bucket greater than the fourth threshold, so that the total data amount of the cold data objects stored in the at least one first storage bucket is relatively large. The cold data objects stored in the at least one first storage bucket are migrated to the at least one second storage bucket included in the at least one storage cluster in the first target region, so that a large amount of storage capacity of the first source region can be released, and the number of storage buckets that need to be migrated can be reduced.

[0016] In another possible implementation, based on the total data amount of the stored cold data objects of each storage bucket, a plurality of storage buckets with the largest total data amount of the stored cold data objects, or the total data amount of the stored cold data objects greater than a fourth threshold, are selected. The number of cold data objects stored in the plurality of storage buckets is obtained. Based on the number of cold data objects stored in the plurality of storage buckets, at least one first storage bucket with the smallest number of stored cold data objects, or the number of stored cold data objects less than a fifth threshold, is selected from the plurality of storage buckets.

[0017] Therefore, the total data amount of the cold data objects stored in the at least one first storage bucket is large, and the number of the stored cold data objects is small, that is, the data amount of each cold data object stored in the at least one first storage bucket is large. Migrating the cold data objects with large data amount to the first target region not only releases a large amount of storage capacity of the first source region, but also reduces the number of migrations and the migration cost.

[0018] In another possible implementation, based on the remaining survival duration of the cold data objects stored in the first storage bucket, at least one cold data object with the longest remaining survival duration, or the remaining survival duration greater than a sixth threshold, is selected from the first storage bucket. Based on the at least one storage cluster in the first source region, the at least one cold data object is migrated to a second storage bucket included in the at least one storage cluster in the first target region.

[0019] After the cold data objects with short remaining survival duration are migrated to the first target region, the cold data objects may be deleted from the at least one storage cluster in the first target region soon, and the migration benefit is low. The at least one cold data object is the cold data object with the longest remaining survival duration, or the remaining survival duration greater than the sixth threshold, so that the at least one cold data object is migrated to the first target region, and the at least one cold data object is stored in the first target region for a long time, thereby maximizing the benefit of migrating the at least one cold data object.

[0020] In another possible implementation, a migration command is sent to the at least one storage cluster in the first source region. The migration command includes identification information of the at least one first storage bucket and identification information of the first target region. The migration command is used to instruct the at least one storage cluster in the first source region to migrate the cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in the at least one storage cluster in the first target region based on the identification information of the at least one first storage bucket and the identification information of the first target region.

[0021] In another possible implementation, the prediction information of the plurality of regions is obtained. Based on the prediction information of the plurality of regions, the second water levels of the plurality of regions in a first time period are determined, the first time period is after the current time. Based on the first water levels of the plurality of regions, m regions with the highest first water levels, or the first water levels higher than a first threshold, and n regions with the lowest first water levels, or the first water levels lower than a second threshold, are selected from the plurality of regions, m and n are integers greater than 1, and the second threshold is less than or equal to the first threshold. Based on the second water levels of the m regions, at least one region with the highest second water level, or the second water level higher than the first threshold, is selected from the m regions as the first source region. Based on the second water levels of the n regions, at least one region with the lowest second water level, or the second water level lower than the second threshold, is selected from the n regions as the first destination region.

[0022] The first source region is the region with the highest second water level, or the second water level higher than the first threshold, so the storage capacity of the first source region is heavily used at the current time and in the first time period, and therefore, the cold data objects in the first source region are migrated, and no large amount of idle storage capacity will appear in the first time period.

[0023] The first destination region is the region with the lowest second water level, or the second water level lower than the second threshold, so the first destination region has a large amount of idle storage capacity at the current time and in the first time period, and therefore, the cold data objects are migrated to the first destination region, and no storage capacity shortage will appear in the first time period.

[0024] In another possible implementation, for each region, the prediction information of the region includes one or more of the following: an increment of the storage capacity used by the region in each of at least one second time period, or data growth information, the at least one second time period is before the current time, and the data growth information is used to describe an increment of the data objects in the region in the first time period compared to the at least one second time period.

[0025] In a second aspect, a device for processing data objects is provided, which is configured to execute the method in the first aspect or any possible implementation of the first aspect. Specifically, the device includes units configured to execute the method in the first aspect or any possible implementation of the first aspect.

[0026] In a third aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory;

[0027] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method in the first aspect or any possible implementation manner of the first aspect.

[0028] In a fourth aspect, the present application provides a computer program product comprising instructions which, when executed by a computing device cluster, cause the computing device cluster to perform the method in the first aspect or any possible implementation manner of the first aspect.

[0029] In a fifth aspect, the present application provides a computer-readable storage medium comprising computer program instructions which, when executed by a computing device cluster, cause the computing device cluster to perform the method in the first aspect or any possible implementation manner of the first aspect.

[0030] In a sixth aspect, the present application provides a chip comprising a memory and a processor, the memory being configured to store computer instructions, and the processor being configured to invoke and run the computer instructions from the memory to execute the method in the first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0031] FIG. 1 is a structural schematic diagram of a cloud storage system according to an embodiment of the present application;

[0032] FIG. 2 is a structural schematic diagram of a network architecture according to an embodiment of the present application;

[0033] FIG. 3 is a schematic diagram of a migrated data object according to an embodiment of the present application;

[0034] FIG. 4 is a structural schematic diagram of a control system according to an embodiment of the present application;

[0035] FIG. 5 is a flowchart of a method for processing a data object according to an embodiment of the present application;

[0036] FIG. 6 is a flowchart of another method for processing a data object according to an embodiment of the present application;

[0037] FIG. 7 is a structural schematic diagram of an apparatus for processing a data object according to an embodiment of the present application;

[0038] FIG. 8 is a structural schematic diagram of a computing device according to an embodiment of the present application;

[0039] FIG. 9 is a structural schematic diagram of a computing device cluster according to an embodiment of the present application;

[0040] FIG. 10 is a schematic diagram of a cluster structure for processing data objects according to an embodiment of the present application. DETAILED DESCRIPTION

[0041] Referring to FIG. 1, a cloud storage system according to an embodiment of the present application includes a plurality of storage clusters 101, which are distributed in a plurality of regions, each region including at least one storage cluster, and the plurality of storage clusters 101 can communicate with each other.

[0042] In some embodiments, the plurality of storage clusters 101 in the cloud storage system can be used to save data objects of an object storage service, and the storage capacity of the plurality of regions can be a capacity planned based on a demand of the object storage service. The storage capacity of a region includes the storage capacity of at least one storage cluster in the region, that is, the total storage capacity of the region is equal to the cumulative value between the storage capacities of each storage cluster in the region.

[0043] For example, as shown in FIG. 1, the cloud storage system includes storage cluster 101a, storage cluster 101b, storage cluster 101c, storage cluster 101d, storage cluster 101e, and storage cluster 101f, and the plurality of regions includes region 1, region 2, and region 3. Region 1 includes storage cluster 101a and storage cluster 101b, region 2 includes storage cluster 101c and storage cluster 101d, and region 3 includes storage cluster 101e and storage cluster 101f.

[0044] The storage capacity of region 1 includes the storage capacity of storage cluster 101a and the storage capacity of storage cluster 101b, that is, the total storage capacity of region 1 is equal to the cumulative value between the storage capacity of storage cluster 101a and the storage capacity of storage cluster 101b in region 1. The storage capacity of region 2 includes the storage capacity of storage cluster 101c and the storage capacity of storage cluster 101d, that is, the total storage capacity of region 2 is equal to the cumulative value between the storage capacity of storage cluster 101c and the storage capacity of storage cluster 101d in region 2. The storage capacity of region 3 includes the storage capacity of storage cluster 101e and the storage capacity of storage cluster 101f, that is, the total storage capacity of region 3 is equal to the cumulative value between the storage capacity of storage cluster 101e and the storage capacity of storage cluster 101f in region 3.

[0045] For each storage cluster included in the region, the storage cluster includes a storage pool, and the storage pool includes storage resources such as memory, hard disk, cache, and the like. The storage capacity of the storage cluster is equal to the capacity of the storage pool included in the storage cluster.

[0046] For each region in the plurality of regions, the water level of each region is different due to different user usage in each region. The water level of the region is used to indicate the used storage capacity of the region. Generally, the higher the water level of the region, the higher the ratio of the used storage resources of the region to the total storage capacity of the region, and the lower the water level of the region, the lower the ratio of the used storage resources of the region to the total storage capacity of the region.

[0047] If the user usage in the region is smaller, the water level of the region is lower, indicating that the ratio of the used storage capacity of the region to the total storage capacity of the region is lower, and the region has a large amount of idle storage capacity, and the region has a storage capacity surplus.

[0048] If the user usage in the region is larger, the water level of the region is higher, indicating that the ratio of the used storage capacity of the region to the total storage capacity of the region is higher, and the region has less idle storage capacity, and the region has a storage capacity shortage, and needs to be expanded.

[0049] Therefore, the water levels of the plurality of regions are unbalanced, resulting in that some regions have a storage capacity surplus and waste a large amount of storage resources, and some regions have a storage capacity shortage and need to be expanded urgently, increasing the operation cost of the cloud storage system.

[0050] In order to balance the water levels of the regions and reduce the operation cost of the cloud storage system, referring to the network architecture shown in FIG. 2, a control system 102 is added in the embodiment of the present application, the control system 102 can communicate with a plurality of storage clusters 101 included in the cloud storage system, and the control system 102 is used for capacity balancing of the plurality of storage clusters 101 included in the cloud storage system.

[0051] The control system 102 can obtain a first water level of a plurality of regions, and determine a first source region and a first target region from the plurality of regions based on the first water level of the plurality of regions.

[0052] The first source region is at least one region with the highest water level among the plurality of regions, or the first source region is at least one region with a water level higher than a first threshold value among the plurality of regions, so that the storage capacity of the first source region is largely used, and a situation of insufficient storage capacity can occur.

[0053] The first destination region is at least one region with the lowest water level among the plurality of regions, or the first destination region is at least one region with a water level lower than a second threshold value among the plurality of regions, the second threshold value being less than or equal to the first threshold value, so that there is a large amount of free storage capacity in the storage capacity of the first destination region.

[0054] Then, the control system 102 migrates, based on at least one storage cluster 101 in the first source region, a cold data object stored in the at least one storage cluster 101 in the first source region to at least one storage cluster 101 in the first destination region.

[0055] In this way, the hot data object in the first source region is still retained in the first source region, ensuring that the user in the first source region can normally access the hot data object, and the cold data object is migrated from the first source region to the first destination region, which can release a large amount of used storage capacity in the first source region, and can utilize a large amount of free storage capacity in the first destination region, so as to balance the water levels of the plurality of regions as much as possible, and reduce the operating cost of the cloud storage system.

[0056] In some embodiments, the control system 102 can further determine a second source region and a second destination region, at least one storage cluster in the second source region includes a target data object, and the access amount from the second destination region to the target data object exceeds a third threshold value. Based on at least one storage cluster in the second source region, the target data object is stored into at least one storage cluster included in the second destination region. In this way, the user in the second destination region can access the target data object from the storage cluster in the second destination region.

[0057] Referring to FIG. 3, at least one storage cluster 101 in a region includes a plurality of storage buckets, for each storage cluster 101, the storage buckets included in the storage cluster 101 are located in the storage pool of the storage cluster 101, and the storage buckets included in the storage cluster 101 are used to store at least one data object. Optionally, the data object can be an object file or the like, or other forms of data, which are not enumerated and described here.

[0058] For the storage bucket included in the storage cluster 101, the capacity of the storage cluster 101 is equal to the cumulative value of the data amount of each data object saved in the storage bucket, so the capacity of the storage bucket increases with the increase of the data objects stored in the storage bucket. Each data object stored in the storage bucket belongs to the same user, and the storage bucket belongs to the user. The used storage capacity of the storage cluster 101 is equal to the sum of the capacity of each storage bucket included in the storage cluster.

[0059] Referring to FIG. 3, when migrating the cold data objects, the cold data objects stored in at least one first storage bucket included in at least one storage cluster 101 in the first source region can be migrated to at least one second storage bucket included in at least one storage cluster 101 in the first destination region. The first storage bucket and the second storage bucket belong to the same user, the first storage bucket is any one of the at least one first storage bucket, and the second storage bucket is any one of the at least one second storage bucket, that is, the at least one first storage bucket and the at least one second storage bucket are in one-to-one correspondence, and the one-to-one corresponding first storage bucket and the second storage bucket belong to the same user.

[0060] In some embodiments, referring to the control system 102 shown in FIG. 4, the control system 102 can include an operation and maintenance platform 1021, a data lake 1022, and a task queue 1023.

[0061] For each storage cluster 101 in a region, the storage cluster 101 can send at least one data to the control system 102. The operation and maintenance platform 1021 of the control system 102 can save the at least one data sent by the storage cluster 101 in the data lake 1022, so the data lake 1022 can save the data sent by each storage cluster 101 in multiple regions. Based on the data saved in the data lake 1022, the operation and maintenance platform 1021 can obtain the first water level of each region in the multiple regions and the cold data objects stored in at least one storage cluster 101 in each region.

[0062] Then the operation and maintenance platform 1021 can determine the first source region and the first destination region based on the first water level of the multiple regions, generate a migration command, and the migration command is used to instruct to migrate the cold data objects stored in at least one storage cluster 101 in the first source region to at least one storage cluster 101 in the first destination region. The operation and maintenance platform 1021 can generate at least one migration command and save the at least one migration command from the tail of the task queue 1023 to the task queue 1023.

[0063] Then, after all the migration commands are generated, the operation and maintenance platform 1021 takes the migration command from the head of the task queue 1023, and sends the migration command to the at least one storage cluster 1201 in the first source region, so that the at least one storage cluster 101 in the first source region migrates the cold data objects stored by the at least one storage cluster 101 to the at least one storage cluster 101 in the first destination region according to the migration command. For details of the migration of the data objects, refer to the implementation of any of the embodiments below, which will not be described here in detail.

[0064] Referring to FIG. 5, the embodiment of the present application provides a method 500 for processing data objects, which is applied to the control system 102 in the embodiments shown in FIG. 2 or FIG. 4. The method 500 is used to determine a first source region and a first destination region, the first water level of the first source region is higher than the first water level of the first destination region, and migrate the cold data objects stored by the at least one storage cluster in the first source region to the at least one storage cluster in the first destination region. The method 500 includes the following processes.

[0065] Step 501: The control system obtains the first water level of a plurality of regions, and the first water level of a region is used to indicate the used storage capacity of the region.

[0066] In step 501, for each region in the plurality of regions, the control system obtains the storage capacity of each storage cluster in the at least one storage cluster in the region, and obtains the capacity of a plurality of storage buckets included in the at least one storage cluster in the region, and the capacity of the storage bucket is equal to the cumulative value of the data amount of each data object saved by the storage bucket. The used storage capacity of the region is obtained based on the capacity of the plurality of storage buckets, and the total storage capacity of the region is obtained based on the storage capacity of each storage cluster included in the region. The first water level of the region is obtained based on the used storage capacity of the region and the total storage capacity of the region.

[0067] Optionally, in implementation, the control system can obtain the first water level of the region through the operations of 5011-5016 as follows.

[0068] 5011: For each storage cluster in the cloud storage system, the control system receives at least one data sent by the storage cluster, and the at least one data includes one or more of the following: a capacity log, or metadata of each storage bucket included in the storage cluster.

[0069] Optionally, the capacity log comprises identification information of a region to which the storage cluster belongs and storage capacity of the storage cluster, etc. The storage capacity of the storage cluster can be changed, for example, the storage pool of the storage cluster can be expanded or contracted, and the storage capacity of the storage cluster is equal to the capacity of the storage pool of the storage cluster, so that the storage capacity of the storage cluster is changed.

[0070] Optionally, the metadata of the storage bucket comprises one or more of the following: identification information of the storage bucket, identification information of a storage cluster to which the storage bucket belongs, identification information of a region to which the storage bucket belongs, capacity of the storage bucket, or number of data objects stored by the storage bucket, etc.

[0071] In some embodiments, the at least one data further comprises access logs, metadata of each data object stored by the storage cluster, or tenant data on the storage cluster, etc.

[0072] Optionally, the access logs comprise one or more of the following: identification information of a data object accessed by a user in the region, type of access operation for accessing the data object, or access timestamp for accessing the data object, etc.

[0073] Optionally, the metadata of the data object comprises one or more of the following: identification information of the data object, identification information of a storage bucket to which the data object belongs, identification information of a storage cluster to which the data object belongs, identification information of a region to which the data object belongs, data volume of the data object, hotness label of the data object, restriction information of the data object, deletion timestamp of the data object, identification information of a user to which the data object belongs, or type of the data object, etc.

[0074] The hotness label of the data object is used to mark whether the data object is a hot data object or a cold data object.

[0075] The restriction information of the data object is used to indicate whether the data object can be migrated.

[0076] Optionally, the tenant data comprises one or more of the following: identification information of the tenant, identification information of at least one storage bucket belonging to the tenant, or restriction information of the tenant, etc.

[0077] The restriction information of the tenant is used to indicate whether a data object stored in a storage bucket belonging to the tenant can be migrated.

[0078] 5012: The control system saves the received at least one data into a data lake of the control system.

[0079] In some embodiments, the control system can extract useful fields in each of the at least one data from each of the at least one data respectively, and save the useful fields in each of the at least one data in the data lake of the control system.

[0080] For each data, the useful fields in the data are the fields required for migrating the data object. For example, the fields required for migrating the data object include the fields required for obtaining the first water level of the region, and / or the fields required for determining the hotness of the data object saved by the storage cluster in the region.

[0081] For example, the storage capacity of the storage cluster and the identification information of the region included in the capacity log are information required for obtaining the first water level of the region, so the storage capacity of the storage cluster and the identification information of the region included in the capacity log are useful fields.

[0082] For another example, the identification information of the storage bucket, the identification information of the storage cluster to which the storage bucket belongs, the identification information of the region to which the storage bucket belongs, and the capacity of the storage bucket included in the metadata of the storage bucket are the fields required for migrating the data object. Therefore, the identification information of the storage bucket, the identification information of the storage cluster to which the storage bucket belongs, the identification information of the region to which the storage bucket belongs, and the capacity of the storage bucket included in the metadata of the storage bucket are useful fields.

[0083] For another example, the identification information of the data object and the access timestamp included in the access log are information required for calculating the access hotness of the data object, and the access hotness of the data object is information required for determining the hotness of the data object. Therefore, the identification information of the data object and the access timestamp included in the access log are useful fields.

[0084] For another example, the identification information of the data object, the identification information of the storage bucket to which the data object belongs, the identification information of the storage cluster to which the data object belongs, the identification information of the region to which the data object belongs, the data amount of the data object, the hotness mark of the data object, the restriction information of the data object, the deletion timestamp of the data object, the identification information of the user to which the data object belongs, or the type of the data object, and the like included in the metadata of the data object are the fields required for migrating the data object. The identification information of the data object, the identification information of the storage bucket to which the data object belongs, the identification information of the storage cluster to which the data object belongs, the identification information of the region to which the data object belongs, the data amount of the data object, the hotness mark of the data object, the restriction information of the data object, the deletion timestamp of the data object, the identification information of the user to which the data object belongs, or the type of the data object, and the like included in the metadata of the data object are useful fields.

[0085] For another example, the identification information of at least one storage bucket belonging to the tenant, the restriction information of the tenant, and the like included in the tenant data are the fields required for migrating the data object. The identification information of at least one storage bucket belonging to the tenant, the restriction information of the tenant, and the like included in the tenant data are useful fields.

[0086] Since the useful fields in the data are saved to the data lake, the storage resources occupied by the data lake can be reduced.

[0087] The control system can receive at least one data sent by each storage cluster in the cloud storage system according to the operation 5011-5012, save the at least one data sent by each storage cluster to the data lake of the control system, and then obtain the first water level of the plurality of regions according to the following operation.

[0088] 5013: For each region, the control system obtains the capacity log of at least one storage cluster included in the region and the metadata of a plurality of storage buckets in the region from the data lake.

[0089] In 5013, the control system obtains at least one capacity log including the identification information of the region and the metadata of a plurality of storage buckets including the identification information of the region from the data lake. The at least one capacity log includes the capacity log of each storage cluster in the region. For the metadata of the plurality of storage buckets including the identification information of the region, the metadata of the plurality of storage buckets includes the metadata of the storage buckets saved in each storage cluster in the region, that is, the metadata of the plurality of storage buckets is the metadata of the plurality of storage buckets in the region.

[0090] 5014: The control system obtains the total storage capacity of the region based on the capacity log of each storage cluster in the region.

[0091] The capacity log of each storage cluster in the region respectively includes the storage capacity of each storage cluster. In 5014, the control system respectively obtains the storage capacity of each storage cluster from the capacity log of each storage cluster, calculates the cumulative value between the storage capacities of each storage cluster, and obtains the total storage capacity of the region.

[0092] 5015: The control system obtains the used storage capacity of the region based on the metadata of the plurality of storage buckets in the region.

[0093] The metadata of the storage bucket includes the capacity of the storage bucket. In 5015, the control system respectively obtains the capacity of each storage bucket from the metadata of each storage bucket in the region, calculates the cumulative value between the capacities of each storage bucket, and obtains the used storage capacity of the region.

[0094] 5016: The control system obtains the first water level of the region based on the used storage capacity of the region and the total storage capacity of the region.

[0095] In 5016, the control system obtains a ratio between the used storage capacity of the region and the total storage capacity of the region, and obtains the first water level of the region. Alternatively, the control system multiplies the ratio by a specified coefficient, and obtains the first water level of the region.

[0096] For example, it is assumed that the control system obtains a ratio between the used storage capacity of the region and the total storage capacity of the region as 0.7, and obtains the first water level of the region as 0.7. Alternatively, it is assumed that the specified coefficient is 10, and the control system multiplies the ratio 0.7 by the specified coefficient 10, and obtains the first water level of the region as 7.

[0097] For each of the regions other than the region among the plurality of regions, the control system obtains the first water level of each of the regions by the above-mentioned operations 5013-5016, and obtains the first water levels of the plurality of regions.

[0098] Step 502: The control system determines a first source region and a first destination region based on the first water levels of the plurality of regions, the first water level of the first source region being higher than the first water level of the first destination region.

[0099] The first source region is at least one region among the plurality of regions with the highest first water level, or the first source region is at least one region among the plurality of regions with the first water level higher than a first threshold value. Therefore, the first source region is a region with a higher first water level, that is, a ratio between a used storage capacity of the first source region and a total storage capacity of the first source region is higher, the storage capacity of the first source region is largely used, the free storage capacity of the first source region is less, and the first source region has a shortage of storage capacity.

[0100] The first destination region is at least one region among the plurality of regions with the lowest first water level, or the first destination region is at least one region among the plurality of regions with the first water level lower than a second threshold value, the second threshold value being less than or equal to the first threshold value. Therefore, the first destination region is a region with a lower first water level, that is, a ratio between a used storage capacity of the first destination region and a total storage capacity of the first destination region is lower, the free storage capacity of the first destination region is larger, and the first destination region has an excess of storage capacity.

[0101] In step 502, the control system determines that the number of the first source regions is at least one and the number of the first destination regions is at least one based on the first water levels of the plurality of regions. Optionally, in implementation:

[0102] The control system selects at least one region with the highest first water level from the plurality of regions as the first source region based on the first water levels of the plurality of regions, or selects at least one region with the first water level higher than a first threshold as the first source region.

[0103] The control system selects at least one region with the lowest first water level from the plurality of regions as the first destination region based on the first water levels of the plurality of regions, or selects at least one region with the first water level lower than a second threshold as the first destination region.

[0104] In some embodiments, the control system can further predict the second water levels of the plurality of regions in a first time period after the current time. In this way, the control system can determine the first source region with the largely used storage capacity at the current time and in the first time period based on the first water levels and the second water levels of the plurality of regions, that is, the free storage capacity of the first source region is small at the current time and in the first time period, i.e., the first source region is in the situation of insufficient storage capacity at the current time and in the first time period.

[0105] And the control system can determine the first destination region with the largely free storage capacity at the current time and in the first time period based on the first water levels and the second water levels of the plurality of regions, that is, the first destination region is in the situation of excess storage capacity at the current time and in the first time period.

[0106] In this way, the situation of insufficient storage capacity of the first destination region and the situation of excess storage capacity of the first source region in the first time period can be avoided after the cold data objects in the first source region are migrated to the first destination region at the current time.

[0107] Optionally, in implementation, the control system can obtain the first source region with the small free storage capacity at the current time and in the first time period and the first destination region with the large free storage capacity at the current time and in the first time period by the operations of 5021-5025.

[0108] 5021: Obtain the prediction information of the plurality of regions.

[0109] For each region, the prediction information of the region comprises one or more of the following: an increment of storage capacity used by the region in each of the at least one second time period, or data growth information, the data growth information being used to describe an increment of data objects in the region in the first time period compared to the at least one second time period, the at least one second time period being before the current time.

[0110] The length of the first time period and / or the length of the second time period is a specified length, for example, the length of the first time period and the length of the second time period are each one day, one week, half a month, one month or one quarter, or the length of the first time period and the length of the second time period are each x days, x weeks, x months or x quarters, x being an integer greater than 1.

[0111] In some embodiments, for each region, the control system counts a total data amount of data objects stored in the at least one storage cluster included by the region in each second time period. Based on the total data amount of data objects stored in the at least one storage cluster included by the region in each second time period, the increment of storage capacity used by the region in each second time period is obtained.

[0112] For example, for each second time period, the total data amount of data objects stored in the at least one storage cluster included by the region in the second time period is subtracted from the total data amount of data objects stored in the at least one storage cluster included by the region in the previous second time period of the second time period, to obtain an increment of data objects in the region in the second time period, the increment of data objects in the region in the second time period being the increment of storage capacity used by the region in the second time period.

[0113] In some embodiments, for each region, the data growth information can comprise one or more of the following: an increment of data objects in the region in the first time period, a holiday or a season in the first time period, etc.

[0114] Optionally, because of a holiday or a seasonal factor, the amount of user usage in the region increases significantly, and thus the total data amount of data objects saved in the storage cluster included by the region also increases significantly. The first time period is the holiday or the season in which the amount of user usage in the region increases significantly.

[0115] For example, it is possible that in summer, the amount of user usage in the region increases, and the data amount of data objects saved in the storage cluster included by the region is much more than the data amount of data objects saved in the storage cluster included by the region in other seasons, and the first time period is summer.

[0116] For another example, the data objects stored by the storage clusters included in the region are user's trajectory data, there are a large number of users going out during a holiday, so that the data volume of the data objects stored in the storage clusters included in the region increases significantly during the holiday, and the first time period is the holiday.

[0117] In 5021, when the holiday comes or the season comes, the control system takes the holiday or the season as the prediction information.

[0118] Optionally, the increment of the data objects in the region in the first time period can be input by a technician to the control system. For example, when the first time period is a holiday or a season in which the user usage increases significantly, the technician can count the average increment of the data objects in the region in the holiday or the season in the past, obtain the increment of the data objects in the region in the first time period based on the average increment, and then input the increment of the data objects in the region in the first time period to the control system.

[0119] 5022: Based on the prediction information of the plurality of regions, determine the second water level of the plurality of regions in the first time period, the first time period being after the current time.

[0120] In some embodiments, for each region, the prediction information of the region includes an increment of the storage capacity used by the region in each of the at least one second time period.

[0121] The control system can count the increment of the storage capacity used by the region in each of the at least one second time period in the past, obtain the storage capacity used by the region in the first time period based on the increment of the storage capacity used by the region in each of the second time period, and obtain the second water level of the region in the first time period based on the storage capacity used by the region in the first time period and the total storage capacity of the region. Optionally, in implementation:

[0122] The control system calculates an average increment based on increments of the storage capacity used by the region in each second time period, and uses the average increment as the increment of the storage capacity used by the region in the first time period. Alternatively, the control system selects a median from the increments of the storage capacity used by the region in each second time period, and uses the median as the increment of the storage capacity used by the region in the first time period. Based on the storage capacity used by the region at the current time and the increment of the storage capacity used by the region in the first time period, the control system obtains the storage capacity used by the region in the first time period. Based on the storage capacity used by the region in the first time period and the total storage capacity of the region, the control system obtains the second water level of the region in the first time period.

[0123] In some embodiments, the prediction information of the region includes an increment of the data objects in the region in the first time period, and the control system obtains the storage capacity used by the region in the first time period based on the storage capacity used by the region at the current time and the increment of the data objects in the region in the first time period. Based on the storage capacity used by the region in the first time period and the total storage capacity of the region, the control system obtains the second water level of the region in the first time period.

[0124] In some embodiments, the prediction information of the region includes the aforementioned holiday or the aforementioned season, and the control system obtains the increment of the data objects in the region in the first time period based on an average increment of the data objects in the region in the holiday or the season in the past.

[0125] Alternatively, the control system can use the average increment as the increment of the data objects in the region in the first time period, or the control system can multiply the average increment by a specified multiple to obtain the increment of the data objects in the region in the first time period. Based on the storage capacity used by the region at the current time and the increment of the data objects in the region in the first time period, the control system obtains the storage capacity used by the region in the first time period. Based on the storage capacity used by the region in the first time period and the total storage capacity of the region, the control system obtains the second water level of the region in the first time period.

[0126] In addition to the above-mentioned several ways of predicting the storage capacity used by the region in the first time period, there are other ways. For example, the control system can also use a capacity growth prediction algorithm to predict the storage capacity used by the region in the first time period based on the prediction information.

[0127] 5023: based on the first water levels of the plurality of regions, select, from the plurality of regions, m regions with the first water levels being the highest, or the first water levels being higher than a first threshold, and select n regions with the first water levels being the lowest, or the first water levels being lower than a second threshold, m and n are both integers greater than 1.

[0128] 5024: based on the second water levels of the m regions, select, from the m regions, at least one region with the second water level being the highest, or the second water level being higher than the first threshold, as the first source region.

[0129] 5025: based on the second water levels of the n regions, select, from the n regions, at least one region with the second water level being the lowest, or the second water level being lower than the second threshold, as the first destination region.

[0130] In some embodiments, the control system can obtain the first source region and the first destination region based on the second water levels of the plurality of regions. Optionally, the control system can select, from the plurality of regions, at least one region with the second water level being the highest, or the second water level being higher than the first threshold, as the first source region, and select at least one region with the second water level being the lowest, or the second water level being lower than the second threshold, as the first destination region, based on the second water levels of the plurality of regions.

[0131] Step 503: the control system migrates, based on at least one storage cluster in the first source region, cold data objects stored in the at least one storage cluster to at least one storage cluster in the first destination region.

[0132] In step 503, the control system can migrate, based on at least one storage cluster in the first source region, cold data objects stored in the at least one storage cluster to at least one storage cluster in the first destination region, by operations of 5031-5033 as follows.

[0133] 5031: obtain a total data amount of cold data objects stored in each storage bucket included in the at least one storage cluster in the first source region.

[0134] The metadata of the data objects in the plurality of regions, the metadata of the buckets, and the access logs of the storage clusters are included in the data lake of the control system. The control system obtains at least one first source region, for each first source region, the control system obtains, based on the metadata of the data objects in the first source region, the metadata of the buckets, and the access logs of the at least one storage cluster, a total data amount of the cold data objects stored by each bucket included by at least one storage cluster in the first source region. Optionally, in implementation, the control system can be implemented by the following flow of operations (11)-(15).

[0135] (11) The control system obtains, based on the identification information of the first source region, the metadata of the plurality of buckets including the identification information of the first source region from the metadata of the buckets saved in the data lake, i.e., obtains the metadata of the plurality of buckets included by at least one storage cluster in the first source region.

[0136] The metadata of the bucket includes one or more of the following: identification information of the bucket, identification information of the storage cluster to which the bucket belongs, or identification information of the region to which the bucket belongs, etc.

[0137] For each bucket, based on the identification information of the bucket, the control system obtains, from the data lake in the control system, tenant data including the identification information of the bucket, the bucket being a bucket belonging to the tenant, the tenant data further including restriction information of the tenant. When the restriction information of the tenant is used to indicate that the data objects stored in the buckets belonging to the tenant can be migrated, the operation of obtaining the total data amount of the cold data objects stored by the bucket is continued as follows. When the restriction information of the tenant is used to indicate that the data objects stored in the buckets belonging to the tenant cannot be migrated, the operation of obtaining the total data amount of the cold data objects stored by the bucket is stopped.

[0138] (12) For each bucket in the plurality of buckets, the control system obtains, from the metadata of the bucket, the identification information of the bucket and the identification information of the storage cluster to which the bucket belongs.

[0139] (13) The control system obtains, from the metadata of the data objects saved in the data lake, the metadata of at least one data object including the identification information of the first source region, the identification information of the bucket, and the identification information of the storage cluster to which the bucket belongs, i.e., obtains the metadata of at least one data object stored by the bucket.

[0140] The metadata of the data object includes one or more of the following: identification information of the data object, identification information of the storage bucket to which the data object belongs, identification information of the storage cluster to which the data object belongs, identification information of the region to which the data object belongs, data volume of the data object, hotness label of the data object, or limitation information of the data object, etc.

[0141] There can be storage buckets belonging to the same user in different regions, the identification information of the storage buckets belonging to the same user can be the same or different, the identification information of the storage buckets in different storage clusters can also be the same, and the same data object can be saved in different regions. Therefore, for at least one data object including the identification information of the first source region, the identification information of the storage bucket, and the identification information of the storage cluster to which the storage bucket belongs, the at least one data object is a data object saved in the storage bucket, and the accuracy of obtaining the metadata of at least one data object stored in the storage bucket is improved.

[0142] (14) Based on the metadata of the at least one data object, the data volume of each cold data object saved in the storage bucket is obtained.

[0143] In some embodiments, for each data object in the at least one data object, if the metadata of the data object includes the hotness label of the data object, the data volume of the cold data object is obtained from the metadata of the data object when the hotness label is used to mark the data object as a cold data object. In the same way as described above, the data volume of each other cold data object saved in the storage bucket is obtained.

[0144] In some embodiments, for each data object in the at least one data object, if the metadata of the data object does not include the hotness label of the data object, the access frequency of the data object in a target time period is obtained, and the target time period is the time period closest to the current time with a specified time length. Based on the access frequency of the data object, it is determined whether the data object is a cold data object, and if the data object is a cold data object, the data volume of the cold data object is obtained from the metadata of the data object. In the same way as described above, the data volume of each other cold data object saved in the storage bucket is obtained.

[0145] Optionally, the operation of obtaining the access frequency of the data object in the target time period can be: obtaining an access log including the identification information of the data object from the data lake, the access log including access timestamps of different users accessing the data object. Based on the access timestamps of accessing the data object in the target time period included in the access log, the access frequency of the data object in the target time period is obtained.

[0146] Optionally, the operation of determining whether the data object is a cold data object based on the access frequency of the data object can be: when the access frequency of the data object is less than or equal to a first frequency threshold, determining that the data object is a cold data object. When the access frequency of the data object is greater than the first frequency threshold and less than a second frequency threshold, obtaining y access time stamps of y times of accessing the data object from the access log in the recent y times, y is an integer greater than 1, and the second frequency threshold is greater than the first frequency threshold. Obtain y-1 access intervals based on the y access time stamps, and the y-1 access intervals include the interval between any two adjacent access time stamps in the y access time stamps. When all the y-1 access intervals are greater than an interval threshold, it is determined that the data object is a cold data object.

[0147] Optionally, when the access frequency of the data object is greater than or equal to the second frequency threshold, or when there is an access interval less than or equal to the interval threshold in the y-1 access intervals, the data object is determined to be a hot data object.

[0148] (15) Obtaining the total data amount of the cold data objects stored in the storage bucket based on the data amount of each cold data object stored in the storage bucket.

[0149] In (15), the sum of the data amount of each cold data object stored in the storage bucket is calculated to obtain the total data amount of the cold data objects stored in the storage bucket.

[0150] The operations of (12)-(15) are repeated to obtain the total data amount of the cold data objects stored in each other storage bucket included in at least one storage cluster in the first source region.

[0151] 5032: Based on the total data amount of the cold data objects stored in each storage bucket, selecting at least one first storage bucket with the largest total data amount of the stored cold data objects, or the total data amount of the stored cold data objects being greater than a fourth threshold.

[0152] Selecting at least one first storage bucket with the largest total data amount of the stored cold data objects, or the total data amount of the stored cold data objects being greater than a fourth threshold, as the storage bucket to be migrated, which can reduce the number of storage buckets that need to be migrated. In the case of reducing the number of migrated storage buckets, more storage capacity in the first source region can be released.

[0153] In some embodiments, based on the total data amount of the stored cold data objects of each storage bucket, a plurality of storage buckets with the largest total data amount of the stored cold data objects or the total data amount of the stored cold data objects greater than a fourth threshold value are selected. The number of cold data objects stored in each storage bucket in the plurality of storage buckets is obtained. Based on the number of cold data objects stored in the plurality of storage buckets, at least one first storage bucket with the smallest number of cold data objects or the number of cold data objects less than a fifth threshold value is selected from the plurality of storage buckets.

[0154] The at least one first storage bucket with the smallest number of cold data objects or the number of cold data objects less than the fifth threshold value is selected from the plurality of storage buckets as the storage bucket to be migrated, so that the cold data objects in each first storage bucket are cold data objects with larger data amount, and thus more storage capacity in the first source region can be released and the migration cost can be reduced while reducing the number of migrated data objects.

[0155] In 5032, the control system can generate one or more migration commands. In the case of generating one migration command, the generated migration command includes the identification information of the at least one first storage bucket and the identification information of the first destination region. In the case of generating a plurality of migration commands, each generated migration command includes the identification information of part of the at least one first storage bucket and the identification information of the first destination region, and the first destination region in each migration command is different, so that the cold data objects in the first source region can be migrated to different destination regions. The generated migration command is saved to the task queue.

[0156] In some embodiments, for each first storage bucket, the control system obtains the remaining survival duration of the cold data objects stored in the first storage bucket. Based on the remaining survival duration of the cold data objects stored in the first storage bucket, at least one cold data object with the longest remaining survival duration or the remaining survival duration exceeding a sixth threshold value is selected from the first storage bucket, and the at least one cold data object is the cold data object to be migrated. For the migration command including the identification information of the first storage bucket, the migration command further includes the identification information of the at least one cold data object to be migrated. Thus, the migration command can instruct at least one storage cluster in the first source region to migrate the at least one cold data object to be migrated to a second storage bucket included in at least one storage cluster in the first destination region, and the first storage bucket and the second storage bucket belong to the same user.

[0157] Optionally, the remaining survival duration of the cold data objects stored in the first storage bucket can be obtained in the following two ways:

[0158] In the first mode, the control system obtains attribute information of the cold data object, obtains a deletion timestamp of the cold data object based on a deletion timestamp prediction model and the attribute information of the cold data object, and obtains a remaining survival duration of the cold data object based on the deletion timestamp of the cold data object.

[0159] In some embodiments, the attribute information of the cold data object can include one or more of the following: a data volume of the cold data object, a type of the cold data object, user identification information to which the cold data object belongs, or a storage path of the cold data object. The storage path can include identification information of a storage bucket to which the cold data object belongs, identification information of a storage cluster to which the cold data object belongs, and identification information of a region to which the cold data object belongs.

[0160] Optionally, the metadata of the cold data object includes the attribute information of the cold data object.

[0161] In the first mode, the control system obtains metadata of the cold data object, obtains attribute information of the cold data object from the metadata of the cold data object, inputs the attribute information of the cold data object into a deletion timestamp prediction model, receives the attribute information of the cold data object by the deletion timestamp prediction model, and obtains a deletion timestamp of the cold data object based on the attribute information of the cold data object. The control system obtains the deletion timestamp of the cold data object output by the deletion timestamp prediction model, and obtains a remaining survival duration of the cold data object based on a current timestamp and the deletion timestamp of the cold data object.

[0162] Optionally, the deletion timestamp prediction model can be obtained by training an artificial intelligence (AI) model based on a plurality of training samples in advance, each training sample including attribute information of a data object and a deletion timestamp of the data object.

[0163] In the second mode, the control system obtains metadata of the cold data object, the metadata including a deletion timestamp of the cold data object, and obtains a remaining survival duration of the cold data object based on the deletion timestamp of the cold data object.

[0164] Optionally, the metadata of the cold data object can include the deletion timestamp of the cold data object, which can be configured by a user to which the cold data object belongs when the cold data object is created.

[0165] Optionally, if the metadata of the cold data object does not include the deletion timestamp of the cold data object, the control system obtains the deletion timestamp of the cold data object by using the first mode. If the metadata of the cold data object includes the deletion timestamp of the cold data object, the control system obtains the deletion timestamp of the cold data object by using the first mode or the second mode, which is more flexible.

[0166] In some embodiments, for each first storage bucket, the control system obtains metadata of each cold data object stored in the first storage bucket, the metadata of the cold data object comprising restriction information of the cold data object, the restriction information of the data object being used to indicate whether the cold data object can be migrated. Based on the restriction information comprised in the metadata of each cold data object, at least one cold data object that can be migrated is selected from the first storage bucket, the at least one cold data object being the cold data object to be migrated. For a migration command comprising the identification information of the first storage bucket, the migration command further comprises identification information of the at least one cold data object to be migrated.

[0167] The control system obtains at least one first source region, and one or more migration commands corresponding to each first source region can be obtained according to the operations 5031-5032 described above. The one or more migration commands corresponding to each first source region are saved into the task queue.

[0168] 5033: Based on at least one storage cluster in the first source region, migrating the cold data object stored in the at least one first storage bucket to at least one second storage bucket comprised in at least one storage cluster in the first destination region.

[0169] Optionally, the at least one first storage bucket and the at least one second storage bucket correspond to each other, and the first storage bucket and the second storage bucket corresponding to each other belong to the same user.

[0170] In 5033, the control system takes out the migration command from the task queue, and sends the migration command to at least one storage cluster in the first source region based on the identification information of the first source region comprised in the migration command. The migration command comprises identification information of the at least one first storage bucket and identification information of the first destination region, and the migration command is used to instruct the at least one storage cluster in the first source region to migrate the cold data object stored in the at least one first storage bucket to at least one second storage bucket comprised in at least one storage cluster in the first destination region based on the identification information of the at least one first storage bucket and the identification information of the first destination region.

[0171] For each storage cluster in the first source region, for the convenience of description, the storage cluster is referred to as a first source storage cluster, the first source storage cluster receives the migration command, determines the first storage bucket included in the first source storage cluster from at least one first storage bucket corresponding to the identification information of the at least one first storage bucket included in the migration command. Based on the identification information of the first destination region included in the migration command, select one storage cluster from at least one storage cluster in the first destination region as a first destination storage cluster. Migrate the cold data object saved by the first storage bucket included in the first source storage cluster to the second storage bucket included in the first destination storage cluster. The second storage bucket in the first destination storage cluster and the first storage bucket in the first source storage cluster belong to the same user.

[0172] In some embodiments, there is a storage cluster including a second storage bucket belonging to the user in at least one storage cluster in the first destination region, and the first source storage cluster selects the storage cluster as the first destination storage cluster. The first source storage cluster sends the cold data object saved by the first storage bucket included in the first source storage cluster to the first destination storage cluster. The first destination storage cluster saves the cold data object into the second storage bucket. Wherein, after the first destination storage cluster saves the cold data object, the first source storage cluster can delete the cold data object.

[0173] In some embodiments, there is a storage cluster including a second storage bucket belonging to the user in at least one storage cluster in the first destination region, and the first source storage cluster selects the storage cluster as the first destination storage cluster. The first source storage cluster sends the cold data object saved by the first storage bucket included in the first source storage cluster to the first destination storage cluster. The first destination storage cluster saves the cold data object into the second storage bucket. Wherein, after the first destination storage cluster saves the cold data object, the first source storage cluster can delete the cold data object.

[0174] In some embodiments, in the case where the migration command includes the identification information of at least one cold data object to be migrated in the first storage bucket, the first source storage cluster sends the at least one cold data object to be migrated saved by the first storage bucket to the destination storage cluster based on the identification information of the at least one cold data object to be migrated.

[0175] Since the at least one cold data object to be migrated is the one with the longest remaining survival time in the first storage bucket, or the at least one cold data object with a remaining survival time exceeding a sixth threshold, the at least one cold data object to be migrated is stored in the first target storage cluster for a longer time, thereby improving the benefit of migrating the data object. Alternatively, since the at least one cold data object to be migrated is the one that can be migrated in the first storage bucket, the at least one cold data object to be migrated is migrated to the first target storage cluster without migration errors.

[0176] In the embodiments of the present application, the control system can receive the capacity logs and the metadata of the storage buckets sent by each storage cluster included in each region, obtain the storage capacity of each storage cluster based on the capacity logs sent by each storage cluster in the region, and obtain the total storage capacity of the region based on the storage capacity of each storage cluster. Based on the metadata of the storage buckets sent by each storage cluster included in the region, the capacity of the storage buckets in each storage cluster is obtained, and the capacities of the multiple storage buckets in the region are obtained, and the used storage capacity of the region is obtained based on the capacities of the multiple storage buckets. Thus, the first water level of the region is obtained based on the used storage capacity of the region and the total storage capacity of the region, and the control system can obtain the first water levels of the multiple regions in the above manner. In this way, the capacities of the multiple storage clusters included in the cloud storage system can be uniformly scheduled and balanced based on the first water levels of the multiple regions. Since the control system determines the first source region and the first target region based on the first water levels of the multiple regions, the first source region is the one with the largest first water level, or the at least one region with the first water level higher than a first threshold, so that the used storage capacity of the first source region accounts for a higher proportion of the total storage capacity of the first source region, and the idle storage capacity in the first source region is less; and the first target region is the one with the smallest first water level, or the at least one region with the first water level lower than a second threshold, so that the used storage capacity of the first target region accounts for a lower proportion of the total storage capacity of the first target region, and the idle storage capacity in the first target region is more. The cold data objects stored in the at least one storage cluster in the first source region are migrated to the at least one storage cluster in the first target region, so that the used part of the storage capacity of the first source region can be released, and the idle storage capacity in the first target region can be utilized, without the need to expand the first source region, so as to balance the water levels of the first source region and the first target region as much as possible, and reduce the operating cost of the cloud storage system.

[0177] Referring to FIG. 6, the embodiment of the present application provides a method 600 of processing data objects, which is applied to the control system 102 shown in FIG. 2 or FIG. 4. The method 600 is used to determine a second source region and a second destination region, at least one storage cluster in the second source region includes a target data object, and the access amount of the target data object from the second destination region exceeds a third threshold value, and the target data object is stored into at least one storage cluster in the second destination region. The method 600 includes the following processes.

[0178] Step 601: The control system obtains the access amount of data objects included in a plurality of regions from other regions.

[0179] Optionally, for each region in the plurality of regions, for the data objects stored in at least one storage system in the region, the control system can obtain the access amount of the data objects from other regions by the operations of 6011-6014 as follows.

[0180] 6011: For each storage cluster in the cloud storage system, the control system receives the traffic log sent by the storage cluster.

[0181] The traffic log includes one or more of the following: identification information of the home region to which the storage cluster belongs, identification information of the data object accessed by the user to the storage cluster in the source region, an access timestamp, identification information of the storage cluster where the data object is stored, or identification information of the source region, and the like, the source region being the region including the data object to be accessed by the user.

[0182] The user accesses a certain data object to the storage cluster, but the storage cluster does not store the data object, but the data object is stored in the storage cluster of other regions, and the other regions are the source regions of the data object, the storage cluster obtains the data object from the storage cluster of the source region, and sends the data object to the user.

[0183] 6012: The control system saves the received traffic log to the data lake of the control system.

[0184] The control system can receive the traffic log sent by each storage cluster in the cloud storage system according to the operations of 6011-6012 as described above, save the traffic log sent by each storage cluster to the data lake of the control system, and then the control system obtains the access amount of the data objects in the region from other regions according to the following operations.

[0185] 6013: The control system selects one region as a source region and one region as a home region from a plurality of regions, obtains a traffic log including identification information of a data object saved by at least one storage cluster in the source region, identification information of the source region and identification information of the home region from a data lake.

[0186] The obtained traffic log records access time stamps of each user in the home region accessing the data object from the source region.

[0187] 6014: The control system obtains an access amount of the data object from the home region based on the access time stamps of each user in the home region accessing the data object from the source region included in the traffic log.

[0188] The control system repeats the operations of 6013-6014 above, and can obtain access amounts of data objects included in a plurality of regions from other regions.

[0189] Step 602: The control system determines a second source region and a second destination region based on the access amounts of data objects included in a plurality of regions from other regions, at least one storage cluster in the second source region includes a target data object, and an access amount of the target data object from the second destination region exceeds a third threshold value.

[0190] For example, taking the access amount of a data object in the source region from the home region obtained in 6014 as an example, if the access amount exceeds the third threshold value, the data object is taken as the target data object, the source region is taken as the second source region, and the home region is taken as the second destination region. The access amount exceeding the third threshold value indicates that there are a large number of users in the second destination region who need to access the target data object.

[0191] Step 603: The control system stores the target data object into at least one storage cluster included in the second destination region based on at least one storage cluster in the second source region.

[0192] In step 603, the control system generates a storage command, the storage command including identification information of the target data object and identification information of the second destination region. The identification information of the storage cluster storing the target data object is obtained from the traffic log including the identification information of the second source region, the identification information of the second destination region, and the identification information of the target data object. The storage cluster corresponding to the identification information of the storage cluster is taken as the second source storage cluster, and the storage command is sent to the second source storage cluster.

[0193] The second source storage cluster receives the storage command, determines a third storage bucket including the target data object based on the identification information of the target data object. Based on the identification information of the second destination region, one storage cluster is selected from at least one storage cluster in the second destination region as a second destination storage cluster. The target data object stored in the third storage bucket is stored into a fourth storage bucket included in the second destination storage cluster, and the third storage bucket and the fourth storage bucket belong to the same user.

[0194] In some embodiments, when there is a storage cluster including a fourth storage bucket belonging to the user in at least one storage cluster in the second destination region, the second source storage cluster selects the storage cluster as the second destination storage cluster and sends the target data object to the second destination storage cluster. The second destination storage cluster saves the target data object into the fourth storage bucket.

[0195] When there is no destination storage cluster including a fourth storage bucket belonging to the user in at least one storage cluster in the second destination region, the second source storage cluster randomly selects one storage cluster from at least one storage cluster in the second destination region as the second destination storage cluster, or selects a storage cluster with the largest idle capacity as the second destination storage cluster, and sends the target data object to the second destination storage cluster. The second destination storage cluster generates the fourth storage bucket, saves the target data object in the fourth storage bucket, and the generated fourth storage bucket and the third storage bucket on the second source storage cluster belong to the same user.

[0196] In some embodiments, in the second source region, if the target data object is a hot data object, the target data object is still saved in the second source storage cluster after the target data object is saved into the second destination storage cluster. Or, in the second source region, if the target data object is a cold data object, the target data object is deleted from the second source storage cluster after the target data object is saved into the second destination storage cluster.

[0197] In the embodiments of the present application, the at least one storage cluster in the second source region includes a target data object, and the access amount from the second destination region to the target data object exceeds a third threshold, indicating that a large number of users in the second destination region access the target data object. The target data object is stored in the at least one storage cluster included in the second destination region based on the at least one storage cluster in the second source region. In this way, a large number of users in the second destination region can access the target data object from the storage cluster in the second destination region, thereby reducing the latency of the users in the second destination region accessing the target data object.

[0198] Referring to FIG. 7, the embodiments of the present application provide a device 700 for processing data objects, which can be deployed on the control system 102 in the embodiments shown in FIG. 2 or FIG. 4. The device 700 is used to perform capacity balancing on a plurality of storage clusters included in a cloud storage system, and the plurality of storage clusters are distributed in a plurality of regions region, each region including at least one storage cluster, and the device 700 includes:

[0199] The acquisition unit 701 is configured to acquire first water levels of the plurality of regions, and the first water level of a region is used to indicate a used storage capacity of the region, and the storage capacity of the region includes a storage capacity of each storage cluster in the region.

[0200] The processing unit 702 is configured to determine a first source region and a first destination region based on the first water levels of the plurality of regions, and the first water level of the first source region is higher than the first water level of the first destination region.

[0201] The processing unit 702 is further configured to migrate a cold data object stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region based on the at least one storage cluster in the first source region.

[0202] Optionally, the detailed implementation process of the acquisition unit 701 acquiring the first water levels of the plurality of regions can refer to the related contents in step 501 of the method 500 shown in FIG. 5, which will not be described in detail here.

[0203] Optionally, the detailed implementation process of the processing unit 702 determining the first source region and the first destination region based on the first water levels of the plurality of regions can refer to the related contents in step 502 of the method 500 shown in FIG. 5, which will not be described in detail here.

[0204] Optionally, the processing unit 702 migrates the cold data object stored in the at least one storage cluster in the first source region to at least one storage cluster in the first destination region. For details, refer to the related content in step 503 of method 500 shown in FIG. 5, which will not be repeated here.

[0205] Optionally, the first source region is at least one region with the first highest water level among the plurality of regions, or the first source region is at least one region with the first water level higher than a first threshold among the plurality of regions.

[0206] The first destination region is at least one region with the first lowest water level among the plurality of regions, or the first destination region is at least one region with the first water level lower than a second threshold among the plurality of regions, and the second threshold is less than or equal to the first threshold.

[0207] Optionally, the obtaining unit 701 is further configured to obtain an access amount of data objects included in the plurality of regions from other regions.

[0208] The processing unit 702 is further configured to determine a second source region and a second destination region based on the access amount, the at least one storage cluster in the second source region including a target data object, and the access amount from the second destination region to the target data object exceeding a third threshold.

[0209] The processing unit 702 is further configured to store the target data object into at least one storage cluster included in the second destination region based on the at least one storage cluster in the second source region.

[0210] Optionally, the obtaining unit 701 obtains the access amount of data objects included in the plurality of regions from other regions. For details, refer to the related content in step 601 of method 600 shown in FIG. 6, which will not be repeated here.

[0211] Optionally, the processing unit 702 determines the second source region and the second destination region based on the access amount. For details, refer to the related content in step 602 of method 600 shown in FIG. 6, which will not be repeated here.

[0212] Optionally, the processing unit 702 stores the target data object into at least one storage cluster included in the second destination region. For details, refer to the related content in step 603 of method 600 shown in FIG. 6, which will not be repeated here.

[0213] Optionally, each of the at least one storage cluster in the region comprises a plurality of storage buckets, each of the storage buckets is configured to store at least one data object, the obtaining unit 701 is configured to:

[0214] obtain a storage capacity of each of the storage clusters in the region, and a capacity of the plurality of storage buckets, the capacity of the storage bucket being equal to an accumulation of data amounts of each of the data objects stored in the storage bucket;

[0215] obtain a used storage capacity of the region based on the capacity of the plurality of storage buckets, and obtain a total storage capacity of the region based on the storage capacity of each of the storage clusters included in the region;

[0216] obtain the first water level of the region based on the used storage capacity of the region and the total storage capacity of the region.

[0217] Optionally, the obtaining unit 701 obtains the storage capacity of each of the storage clusters in the region, the capacity of the plurality of storage buckets, the used storage capacity of the region, the total storage capacity of the region, and the first water level of the region in details according to the related contents in steps 5011-5016 of the method 500 shown in FIG. 5, which will not be described in details herein.

[0218] Optionally, the processing unit 702 is configured to:

[0219] obtain a total data amount of cold data objects stored in each of the storage buckets included in the at least one storage cluster in the first source region;

[0220] select, based on the total data amount of the cold data objects stored in each of the storage buckets, at least one first storage bucket with a largest total data amount of the cold data objects stored in the first storage bucket, or a total data amount of the cold data objects stored in the first storage bucket being greater than a fourth threshold value;

[0221] migrate, based on the at least one storage cluster in the first source region, the cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in at least one storage cluster in the first destination region, the first storage bucket and the second storage bucket belonging to a same user, the first storage bucket being any one of the at least one first storage bucket, and the second storage bucket being any one of the at least one second storage bucket.

[0222] Optionally, the processing unit 702 obtains the total data amount of the cold data objects stored in each of the storage buckets included in the at least one storage cluster in the first source region in details according to the related contents in 5031 of the method 500 shown in FIG. 5, which will not be described in details herein.

[0223] Optionally, the processing unit 702 selects the at least one first storage bucket based on a total data amount of the cold data objects stored in each storage bucket. For details of the implementation process, refer to 5032 in the method 500 shown in FIG. 5.

[0224] Optionally, the processing unit 702 migrates the cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in at least one storage cluster in the first destination region. For details of the implementation process, refer to 5033 in the method 500 shown in FIG. 5.

[0225] Optionally, the processing unit 702 is configured to:

[0226] select, based on a total data amount of the cold data objects stored in each storage bucket, a plurality of storage buckets with a total data amount of the cold data objects stored therein being the largest, or a total data amount of the cold data objects stored therein being greater than a fourth threshold value.

[0227] obtain a number of the cold data objects stored in the plurality of storage buckets;

[0228] select, based on the number of the cold data objects stored in the plurality of storage buckets, at least one first storage bucket with a number of the cold data objects stored therein being the smallest, or a number of the cold data objects stored therein being less than a fifth threshold value.

[0229] Optionally, the processing unit 702 selects the plurality of storage buckets. For details of the implementation process, refer to 5032 in the method 500 shown in FIG. 5.

[0230] Optionally, the processing unit 702 obtains the number of the cold data objects stored in the plurality of storage buckets. For details of the implementation process, refer to 5032 in the method 500 shown in FIG. 5.

[0231] Optionally, the processing unit 702 selects the at least one first storage bucket based on the number of the cold data objects stored in the plurality of storage buckets. For details of the implementation process, refer to 5032 in the method 500 shown in FIG. 5.

[0232] Optionally, the processing unit 702 is configured to:

[0233] select, based on a remaining survival time length of the cold data objects stored in the first storage bucket, at least one cold data object with a remaining survival time length being the longest, or a remaining survival time length exceeding a sixth threshold value.

[0234] Migrating the at least one cold data object into a second storage bucket included in the at least one storage cluster in the first destination region based on the at least one storage cluster in the first source region.

[0235] Optionally, the processing unit 702 selects the at least one cold data object from the first storage bucket based on a remaining survival time length of the cold data object stored in the first storage bucket. For details, refer to related content in 5032 of method 500 shown in FIG. 5, which will not be described in detail here.

[0236] Optionally, the processing unit 702 migrates the at least one cold data object into a second storage bucket included in the at least one storage cluster in the first destination region. For details, refer to related content in 5033 of method 500 shown in FIG. 5, which will not be described in detail here.

[0237] Optionally, the apparatus 700 further includes a sending unit 703.

[0238] The sending unit 703 is configured to send a migration command to the at least one storage cluster in the first source region. The migration command includes identification information of the at least one first storage bucket and identification information of the first destination region. The migration command is used to instruct the at least one storage cluster in the first source region to migrate a cold data object stored in the at least one first storage bucket into at least one second storage bucket included in the at least one storage cluster in the first destination region based on the identification information of the at least one first storage bucket and the identification information of the first destination region.

[0239] Optionally, the sending unit 703 sends the migration command to the at least one storage cluster in the first source region. For details, refer to related content in 5033 of method 500 shown in FIG. 5, which will not be described in detail here.

[0240] Optionally, the obtaining unit 701 is further configured to obtain prediction information of a plurality of regions; and determine, based on the prediction information of the plurality of regions, a second water level of the plurality of regions in a first time period, the first time period being after a current time.

[0241] The processing unit 702 is configured to:

[0242] Select, based on the first water level of the plurality of regions, m regions with the highest first water level or with the first water level higher than a first threshold value, and n regions with the lowest first water level or with the first water level lower than a second threshold value from the plurality of regions, m and n are both integers greater than 1, and the second threshold value is less than or equal to the first threshold value.

[0243] based on the second water levels of the m regions, selecting, from the m regions, at least one region with the highest second water level, or the second water level higher than a second threshold, as the first source region;

[0244] based on the second water levels of the n regions, selecting, from the n regions, at least one region with the lowest second water level, or the second water level lower than a second threshold, as the first destination region.

[0245] Optionally, the obtaining unit 701 obtains prediction information of the plurality of regions; based on the prediction information of the plurality of regions, the second water levels of the plurality of regions in the first time period are determined. For details, refer to 5021-5022 in method 500 shown in FIG. 5, which will not be described in detail here.

[0246] Optionally, the processing unit 702 selects m regions from the plurality of regions based on the first water levels of the plurality of regions, and selects n regions. For details, refer to the related content in 5023 in method 500 shown in FIG. 5, which will not be described in detail here.

[0247] Optionally, the processing unit 702 selects the first destination region from the n regions based on the second water levels of the n regions. For details, refer to the related content in 5025 in method 500 shown in FIG. 5, which will not be described in detail here.

[0248] Optionally, for each region, the prediction information of the region includes one or more of the following: an increment of storage capacity used by the region in each of at least one second time period, or data growth information, the at least one second time period being located before the current time, the data growth information being used to describe an increment of data objects in the region in the first time period compared to the at least one second time period.

[0249] In the embodiments of the present application, the acquisition unit acquires the first water levels of the plurality of regions, and the processing unit determines the first source region and the first target region based on the first water levels of the plurality of regions. The first water level of a region indicates the used storage capacity of the region, and the first water level of the first source region is higher than that of the first target region, so the storage capacity of the first source region is largely used, and the first target region has a large amount of free storage capacity. The processing unit migrates the cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first target region, so as to release part of the used storage capacity in the first source region, thereby not needing to expand the first source region, and fully utilizing part of the free storage capacity of the first target region, thereby reducing the operation cost of the cloud storage system.

[0250] Referring to FIG. 8, an embodiment of the present application provides a computing device 800. For example, the computing device 800 can be a device in the control system 102 shown in FIGS. 2-4, or the computing device 800 can be a device in the control system in the method 500 shown in FIG. 5 or the method 600 shown in FIG. 6, etc.

[0251] As shown in FIG. 8, the computing device 800 includes a bus 802, a processor 804, a memory 806, and a communication interface 808. The processor 804, the memory 806, and the communication interface 808 communicate through the bus 802. The computing device 800 can be a server or a terminal device. It should be understood that the number of processors and memories in the computing device 800 is not limited by the present application.

[0252] The bus 802 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one line is shown in FIG. 8, but it does not mean that there is only one bus or only one type of bus. The bus 802 can include a path for transmitting information between various components (e.g., the processor 804, the memory 806, the communication interface 808) of the computing device 800.

[0253] The processor 804 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), among other processors.

[0254] The memory 806 can include volatile memory, such as random access memory (RAM), and non-volatile memory, such as read-only memory (ROM), floppy disks, mechanical hard disks, or solid state hard disks, among others.

[0255] Referring to FIG. 8, the memory 806 stores executable program code that the processor 804 executes to implement the functions of the obtaining unit 701, the processing unit 702, and the sending unit 703 in the apparatus 700 shown in FIG. 7, respectively, to implement the method provided by any of the embodiments described above. That is, the memory 806 has instructions for implementing the method provided by any of the embodiments described above. Alternatively,

[0256] The communication interface 808 uses a transceiving module, such as but not limited to a network interface card or a transceiver, to enable communication between the computing device 800 and other devices or communication networks.

[0257] Embodiments of the present application also provide a cluster for processing data objects. The cluster for moving data includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device, such as a desktop computer, a notebook computer, or a smartphone.

[0258] As shown in FIG. 9, the cluster for processing data objects includes at least one computing device 800. The memory 806 in one or more computing devices 800 in the cluster for processing data objects can store the same instructions for implementing the method provided by any of the embodiments described above.

[0259] In some possible implementation, the memory 806 of one or more computing devices 800 in the cluster for processing data objects can also respectively store partial instructions for performing the above-mentioned method for processing data objects. In other words, the combination of one or more computing devices 800 can collectively execute the instructions for performing the method provided by any of the above-mentioned embodiments.

[0260] In some possible implementation, one or more computing devices in the cluster for processing data objects can be connected through a network. The network can be a wide area network, a local area network, or the like. FIG. 10 shows one possible implementation. As shown in FIG. 10, two computing devices 800A and 800B are connected through a network. Specifically, the computing devices are connected to the network through the communication interfaces in the computing devices.

[0261] In some possible implementation, the memory 806 in the computing device 800A stores instructions for performing the functions of the obtaining unit 701 and the processing unit 702 in the embodiment shown in FIG. 7. Meanwhile, the memory 806 in the computing device 800B stores instructions for performing the functions of the sending unit 703 in the embodiment shown in FIG. 7.

[0262] It should be understood that the functions of the computing device 800A shown in FIG. 10 can also be completed by multiple computing devices 800. Similarly, the functions of the computing device 800B can also be completed by multiple computing devices 800.

[0263] The embodiments of the present application also provide another cluster for processing data objects. The connection relationship between the computing devices in the cluster for processing data objects can be similar to the connection mode of the cluster for processing data objects described with reference to FIG. 10. The difference is that the memory 806 in one or more computing devices 800 in the cluster for processing data objects can store the same instructions for performing the method provided by any of the above-mentioned embodiments.

[0264] In some possible implementation, the memory 806 of one or more computing devices 800 in the cluster for processing data objects can also respectively store partial instructions for performing the above-mentioned method for processing data objects. In other words, the combination of one or more computing devices 800 can collectively execute the instructions for performing the method provided by any of the above-mentioned embodiments.

[0265] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to perform the method provided by any of the above-mentioned embodiments.

[0266] The embodiments of the present application further provide a computer readable storage medium. The computer readable storage medium can be any available medium or data storage device that can be used to store instructions that can be executed by a computing device. The computer readable storage medium can be a magnetic-based medium, (e.g., a floppy diskette, a hard disk drive, a magnetic tape), an optical-based medium, (e.g., a compact disc, a DVD, etc.), or a semiconductor-based medium, (e.g., a solid state hard drive), etc. The computer readable storage medium includes instructions that are executable by a computing device to perform the methods described above in connection with any of the embodiments.

[0267] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region.

[0268] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0269] The above is only optional embodiments of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method of processing data objects, characterized by, The method is applied to a control system for capacity balancing of a plurality of storage clusters included in a cloud storage system, the plurality of storage clusters are distributed in a plurality of regions region, each region includes at least one storage cluster, and the method comprises: Obtaining a first water level of a plurality of regions, the first water level of the region is used to indicate the used storage capacity of the region, and the storage capacity of the region includes the storage capacity of each storage cluster in the region; Determine a first source region and a first destination region based on the first water level of the plurality of regions, the first water level of the first source region is higher than the first water level of the first destination region; Based on at least one storage cluster in the first source region, migrate the cold data object stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region.

2. The method of claim 1, wherein: The first source region is at least one region in the plurality of regions with the highest first water level, or the first source region is at least one region in the plurality of regions with a first water level higher than a first threshold value; The first destination region is at least one region in the plurality of regions with the lowest first water level, or the first destination region is at least one region in the plurality of regions with a first water level lower than a second threshold value, the second threshold value is less than or equal to the first threshold value.

3. The method of claim 1 or 2, wherein, The method further comprises: Obtaining the access amount of data objects included in the plurality of regions from other regions; Determine a second source region and a second destination region based on the access amount, at least one storage cluster in the second source region includes a target data object, and the access amount from the second destination region to the target data object exceeds a third threshold value; Based on at least one storage cluster in the second source region, store the target data object into at least one storage cluster included in the second destination region.

4. The method according to any one of claims 1 to 3, characterized in that, Each of the at least one storage cluster in the region includes a plurality of storage buckets, each of which is used to store at least one data object, and the method further comprises: Obtaining the storage capacity of each storage cluster in the region and the capacity of the plurality of storage buckets, the capacity of the storage bucket is equal to the cumulative value of the data amount of each data object saved by the storage bucket; Based on the capacity of the plurality of storage buckets, obtain the used storage capacity of the region, and based on the storage capacity of each storage cluster included in the region, obtain the total storage capacity of the region; The method of obtaining a first water level of a plurality of regions comprises: Based on the used storage capacity of the region and the total storage capacity of the region, a first water level of the region is obtained.

5. The method of claim 4, wherein, The migrating the cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in the at least one storage cluster in the first target region based on the at least one storage cluster in the first source region comprises: obtaining a total data amount of the cold data objects stored in each storage bucket included in the at least one storage cluster in the first source region; based on the total data amount of the cold data objects stored in each storage bucket, selecting at least one first storage bucket with a largest total data amount of the cold data objects stored, or a total data amount of the cold data objects stored being greater than a fourth threshold value; the migrating the cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in the at least one storage cluster in the first target region based on the at least one storage cluster in the first source region comprises:

6. The method of claim 5, wherein, based on the total data amount of the cold data objects stored in each storage bucket, selecting at least one first storage bucket with a largest total data amount of the cold data objects stored, or a total data amount of the cold data objects stored being greater than a fourth threshold value; obtaining a total data amount of the cold data objects stored in each storage bucket included in the at least one storage cluster in the first source region; based on the total data amount of the cold data objects stored in each storage bucket, selecting at least one first storage bucket with a largest total data amount of the cold data objects stored, or a total data amount of the cold data objects stored being greater than a fourth threshold value; the migrating the cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in the at least one storage cluster in the first target region based on the at least one storage cluster in the first source region comprises:

7. The method of claim 5 or 6, wherein, based on the remaining survival time length of the cold data objects stored in the first storage bucket, selecting at least one cold data object with a longest remaining survival time length, or a remaining survival time length exceeding a sixth threshold value from the first storage bucket; the migrating the at least one cold data object to the second storage bucket included in the at least one storage cluster in the first target region based on the at least one storage cluster in the first source region comprises: the migrating the cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in the at least one storage cluster in the first target region based on the at least one storage cluster in the first source region comprises:

8. The method of claim 5 or 6, wherein, ​ sending a migration command to at least one storage cluster in the first source region, the migration command comprising identification information of the at least one first storage bucket and identification information of the first destination region, the migration command being used to instruct the at least one storage cluster in the first source region to migrate cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in at least one storage cluster in the first destination region based on the identification information of the at least one first storage bucket and the identification information of the first destination region.

9. The method according to any one of claims 1 to 8, wherein, The method further comprises: obtaining prediction information of the plurality of regions; determining second water levels of the plurality of regions in a first time period, the first time period being after a current time, based on the prediction information of the plurality of regions; The determining of the first source region and the first destination region based on the first water levels of the plurality of regions comprises: selecting, from the plurality of regions, m regions with the highest first water levels or higher than a first threshold, and n regions with the lowest first water levels or lower than a second threshold, m and n being integers greater than 1, the second threshold being less than or equal to the first threshold, based on the first water levels of the plurality of regions; selecting, from the m regions, at least one region with the highest second water level or higher than the first threshold as the first source region, based on the second water levels of the m regions; selecting, from the n regions, at least one region with the lowest second water level or lower than the second threshold as the first destination region, based on the second water levels of the n regions.

10. The method of claim 9, wherein, For each region, the prediction information of the region comprises one or more of: an increment of storage capacity used by the region in each of at least one second time period, or data growth information, the at least one second time period being before a current time, the data growth information being used to describe an increment of data objects in the region in the first time period compared to the at least one second time period.

11. An apparatus for processing data objects, characterized by The apparatus is used for capacity balancing of a plurality of storage clusters included in a cloud storage system, the plurality of storage clusters being distributed in a plurality of regions, each region comprising at least one storage cluster, the apparatus comprising: an obtaining unit, configured to obtain first water levels of the plurality of regions, the first water level of a region being used to indicate a used storage capacity of the region, the storage capacity of the region comprising storage capacities of each storage cluster in the region; determine a first source region and a first destination region based on the first water levels of the plurality of regions, the first water level of the first source region being higher than the first water level of the first destination region; determine a first source region and a first destination region based on the first water levels of the plurality of regions, the first water level of the first source region being higher than the first water level of the first destination region; 12. The apparatus of claim 11, wherein the first source region is at least one region of the plurality of regions with the highest first water level, or the first source region is at least one region of the plurality of regions with the first water level higher than a first threshold value; the first destination region is at least one region of the plurality of regions with the lowest first water level, or the first destination region is at least one region of the plurality of regions with the first water level lower than a second threshold value, the second threshold value being less than or equal to the first threshold value.

13. The apparatus of claim 11 or 12, wherein the obtaining unit is further configured to obtain an access amount of data objects included in the plurality of regions from other regions; the processing unit is further configured to determine a second source region and a second destination region based on the access amount, at least one storage cluster of the second source region including a target data object, and an access amount from the second destination region to the target data object exceeding a third threshold value; the processing unit is further configured to store the target data object to at least one storage cluster included in the second destination region based on at least one storage cluster of the second source region.

14. The apparatus of any one of claims 11-13, wherein, each of the at least one storage cluster in the region includes a plurality of storage buckets, each of the storage buckets being configured to store at least one data object, and the obtaining unit is configured to: obtain a storage capacity of each of the at least one storage cluster in the region, and a capacity of the plurality of storage buckets, the capacity of the storage bucket being equal to an accumulated value of data amounts of each of the data objects stored in the storage bucket; obtain a used storage capacity of the region based on the capacity of the plurality of storage buckets, and obtain a total storage capacity of the region based on the storage capacity of each of the at least one storage cluster included in the region; obtain the first water level of the region based on the used storage capacity of the region and the total storage capacity of the region.

15. The apparatus of claim 14, wherein, the processing unit is configured to: obtain a total data amount of cold data objects stored in each of the storage buckets included in the at least one storage cluster of the first source region; select, based on the total data amount of the cold data objects stored in each of the storage buckets, at least one first storage bucket with the largest total data amount of the cold data objects stored in the at least one first storage bucket, or with a total data amount of the cold data objects stored in the at least one first storage bucket greater than a fourth threshold value; migrate, based on the at least one storage cluster in the first source region, the cold data objects stored in the at least one first storage bucket to at least one second storage bucket included in at least one storage cluster in the first destination region, the first storage bucket and the second storage bucket belonging to a same user, the first storage bucket being any one of the at least one first storage bucket, and the second storage bucket being any one of the at least one second storage bucket.

16. The apparatus of claim 15, wherein, The processing unit is configured to: select, based on the total data amount of the cold data objects stored in each of the storage buckets, a plurality of storage buckets with the largest total data amount of the cold data objects stored in the plurality of storage buckets, or with a total data amount of the cold data objects stored in the plurality of storage buckets greater than a fourth threshold value; obtain a number of the cold data objects stored in the plurality of storage buckets; select, based on the number of the cold data objects stored in the plurality of storage buckets, at least one first storage bucket with the smallest number of the cold data objects stored in the at least one first storage bucket, or with a number of the cold data objects stored in the at least one first storage bucket less than a fifth threshold value.

17. The apparatus of claim 15 or 16, wherein, The processing unit is configured to: select, based on the remaining survival time length of the cold data objects stored in the first storage bucket, at least one cold data object with the longest remaining survival time length, or with a remaining survival time length exceeding a sixth threshold value; migrate, based on the at least one storage cluster in the first source region, the at least one cold data object to the second storage bucket included in at least one storage cluster in the first destination region.

18. The apparatus of claim 15 or 16, wherein, The apparatus further includes a sending unit. The sending unit is configured to send, to the at least one storage cluster in the first source region, a migration command including identification information of the at least one first storage bucket and identification information of the first destination region, the migration command being used to instruct the at least one storage cluster in the first source region to migrate, based on the identification information of the at least one first storage bucket and the identification information of the first destination region, the cold data objects stored in the at least one first storage bucket to the at least one second storage bucket included in at least one storage cluster in the first destination region.

19. The apparatus of any one of claims 11-18, wherein, The obtaining unit is further configured to obtain prediction information of the plurality of regions, and determine, based on the prediction information of the plurality of regions, a second water level of the plurality of regions in a first time period, the first time period being after a current time. The processing unit is configured to: selecting, from the plurality of regions, m regions with the highest first water level, or with the first water level higher than a first threshold, and n regions with the lowest first water level, or with the first water level lower than a second threshold, m and n are integers greater than 1, the second threshold is less than or equal to the first threshold; selecting, from the m regions, at least one region with the highest second water level, or with the second water level higher than the first threshold, as the first source region, based on the second water level of the m regions; selecting, from the n regions, at least one region with the lowest second water level, or with the second water level lower than the second threshold, as the first destination region, based on the second water level of the n regions.

20. The apparatus of claim 19, wherein, For each region, the prediction information of the region comprises one or more of: an increment of storage capacity used by the region in each of at least one second time period, or data growth information, the at least one second time period being before the current time, the data growth information describing an increment of data objects in the region in the first time period compared to the at least one second time period.

21. A cluster of computing devices, characterized in that, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of any one of claims 1-10.

22. A computer-readable storage medium, characterized in that, comprising computer program instructions which, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method of any one of claims 1-10.

23. A computer program product comprising instructions, characterized in that, the instructions, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Data processing method, device and system and storage medium

    CN111562889A

  • Method, system and device for controlling water level of buffer pool and medium

    CN112631521A

  • Data storage system, data storage method, readable medium, and electronic equipment

    CN113901024A

  • Data storage method and distributed storage processing system

    CN118012341A

  • Methods and systems for garbage collection and compaction for key-value engines

    US20240020231A1