Method and device for processing data object and storage medium

By migrating cold data objects in the cloud storage system, the problem of uneven storage capacity was solved, achieving efficient utilization of storage resources and cost reduction.

CN120994301APending Publication Date: 2025-11-21HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410781902.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2024-06-17
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

The uneven storage capacity in different regions of the cloud storage system leads to either excess or insufficient storage capacity in some regions, increasing operating costs.

Method used

The system obtains the storage cluster water level of each region through the control system, determines the high water level region (source region) and the low water level region (destination region), migrates cold data objects from the source region to the destination region to balance storage capacity, releases storage capacity in the source region and utilizes the idle capacity in the destination region.

Benefits of technology

It effectively reduces the operating costs of cloud storage systems, reduces the need for expansion, and improves resource utilization by balancing storage capacity utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994301A_ABST
    Figure CN120994301A_ABST
Patent Text Reader

Abstract

The invention discloses a data object processing method and device and a storage medium, and belongs to the field of storage. The method is applied to a control system, the control system is used for carrying out capacity balancing on a plurality of storage clusters included in a cloud storage system, the plurality of storage clusters are distributed in a plurality of regions, and each region comprises at least one storage cluster. The method comprises the following steps: acquiring first water levels of a plurality of regions; determining a first source region and a first destination region based on a first water level of the plurality of regions, the first water level of the first source region being higher than the first water level of the first destination region; and migrating the cold data object stored in the at least one storage cluster in the first source region to the at least one storage cluster in the first target region based on the at least one storage cluster in the first source region. The operation cost of the cloud storage system can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202410637811.6, filed on May 20, 2024, entitled "A method, apparatus and other equipment for data processing", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of storage, and in particular to a method, apparatus and storage medium for processing data objects. Background Technology

[0003] Object storage services require storing data objects across different regions. To meet these needs, cloud storage systems deploy storage clusters across multiple regions, with one or more clusters deployed in each region. The storage capacity of each region is planned in advance based on the requirements of the object storage service; the total storage capacity of a region equals the sum of the storage capacities of each storage cluster within that region.

[0004] However, pre-planning involves certain uncertainties. Some regions may have low user traffic, resulting in a large amount of unused storage capacity and thus overcapacity. Conversely, some regions may have high user traffic, leading to insufficient storage capacity as a result.

[0005] Therefore, some regions in the current cloud storage system have excess storage capacity, while others have insufficient storage capacity, which increases the operating costs of the cloud storage system. Summary of the Invention

[0006] This application provides a method, apparatus, and storage medium for processing data objects to reduce the operating costs of cloud storage systems. The technical solution is as follows:

[0007] In a first aspect, this application provides a method for processing data objects. The method is applied to a control system for capacity balancing of multiple storage clusters within a cloud storage system. These storage clusters are distributed across multiple regions, each region including at least one storage cluster. In the method, a first water level is obtained for each of the multiple regions. This first water level indicates the storage capacity already used in the region, and the storage capacity of the region includes the storage capacity of each storage cluster within the region. A first source region and a first destination region are determined based on the first water levels of the multiple regions, where the first water level of the first source region is higher than the first water level of the first destination region. Based on at least one storage cluster in the first source region, cold data objects stored in at least one storage cluster in the first source region are migrated to at least one storage cluster in the first destination region.

[0008] The process involves obtaining the first water level of multiple regions and determining the first source region and the first destination region based on these water levels. The first water level of a region indicates the storage capacity already used. Since the first water level of the first source region is higher than that of the first destination region, the storage capacity of the first source region is heavily utilized, while the first destination region has a significant amount of free storage capacity. Migrating cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region frees up some of the used storage capacity in the first source region, eliminating the need to expand the first source region and fully utilizing the free storage capacity of the first destination region, thus reducing the operating costs of the cloud storage system.

[0009] In one possible implementation, the first source region is at least one region among multiple regions with the highest first water level, or the first source region is at least one region among multiple regions whose first water level is above a first threshold. This ensures that the determined first source region is a region among multiple regions where storage capacity is heavily used, and migrating cold data objects in the first source region effectively reduces the operating costs of the cloud system.

[0010] In another possible implementation, the first destination region is at least one region with the lowest first water level among multiple regions, or the first destination region is at least one region whose first water level is lower than a second threshold, where the second threshold is less than or equal to the first threshold. This ensures that the determined first destination region is one with a large amount of free storage capacity among multiple regions. Migrating cold data objects to the first destination region avoids using an excessively high proportion of the used storage capacity in the first destination region, effectively balancing the used storage capacity across multiple regions.

[0011] In another possible implementation, the access volume of data objects from other regions within multiple regions is obtained. A second source region and a second destination region are determined based on the access volume. At least one storage cluster in the second source region contains the target data object, and the access volume from the second destination region for the target data object exceeds a third threshold. Based on at least one storage cluster in the second source region, the target data object is stored in at least one storage cluster included in the second destination region.

[0012] Since the number of accesses to the target data object from the second destination region exceeds the third threshold, it indicates that a large number of users in the second destination region are accessing the target data object in the second source region. Storing the target data object in at least one storage cluster included in the second destination region can allow users in the second destination region to access the target data object in the second destination region from the nearest location, thereby reducing the latency of accessing the target data object.

[0013] In another possible implementation, at least one storage cluster in each region includes multiple buckets, each bucket storing at least one data object. The storage capacity of each storage cluster in the region, as well as the capacities of the multiple buckets, are obtained. The bucket capacity is equal to the sum of the data amounts of each data object stored in the bucket. Based on the capacities of the multiple buckets, the used storage capacity of the region is obtained, and based on the storage capacity of each storage cluster included in the region, the total storage capacity of the region is obtained. Based on the used storage capacity and the total storage capacity of the region, a first water level for the region is obtained. This allows for an accurate determination of the first water level, which indicates the percentage of used storage capacity in the region. The first source region and the first destination region are derived based on the first water level. Migrating cold data objects from the first source region to the first destination region effectively balances the percentage of used storage capacity across multiple regions.

[0014] In another possible implementation, the total amount of cold data objects stored in each bucket of at least one storage cluster in the first source region is obtained. Based on the total amount of cold data objects stored in each bucket, at least one first bucket with the largest total amount of stored cold data objects, or whose total amount of stored cold data objects is greater than a fourth threshold, is selected. Based on at least one storage cluster in the first source region, the cold data objects stored in at least one first bucket are migrated to at least one second bucket of at least one storage cluster in the first destination region, wherein the first and second buckets belong to the same user, and the first bucket is any one of at least one first bucket, and the second bucket is any one of at least one second bucket.

[0015] At least one first bucket is selected because it has the largest total amount of cold data objects stored therein, or the total amount of cold data objects stored therein is greater than the fourth threshold. Therefore, at least one first bucket stores a large amount of cold data objects. Migrating the cold data objects stored in at least one first bucket to at least one second bucket in at least one storage cluster within the first destination region not only frees up a significant amount of storage capacity in the first source region but also reduces the number of buckets that need to be migrated.

[0016] In another possible implementation, based on the total amount of cold data objects stored in each bucket, select multiple buckets with the largest total amount of cold data objects, or buckets with a total amount of cold data objects greater than a fourth threshold. Obtain the number of cold data objects stored in multiple buckets. Based on the number of cold data objects stored in multiple buckets, select at least one first bucket from the multiple buckets with the smallest number of cold data objects, or buckets with a number of cold data objects less than a fifth threshold.

[0017] Therefore, at least one first bucket stores a large total amount of cold data objects, but the number of individual cold data objects is relatively small. In other words, at least one first bucket stores a large amount of data for each cold data object. Migrating these large amounts of cold data objects to the first destination region not only frees up a significant amount of storage capacity in the first source region but also reduces the number of migrations and lowers migration costs.

[0018] In another possible implementation, based on the remaining lifespan of the cold data objects stored in the first bucket, at least one cold data object with the longest remaining lifespan, or whose remaining lifespan exceeds a sixth threshold, is selected from the first bucket. Based on at least one storage cluster in the first source region, at least one cold data object is migrated to a second bucket included in at least one storage cluster in the first destination region.

[0019] Migrating cold data objects with short remaining lifespans to the first destination region may result in their rapid deletion from at least one storage cluster within that region, leading to minimal migration benefits. Conversely, migrating at least one cold data object to the first destination region ensures its continued presence there for a longer period, maximizing the benefits of this migration.

[0020] In another possible implementation, a migration command is sent to at least one storage cluster in the first source region. The migration command includes identification information of at least one first bucket and identification information of the first destination region. The migration command is used to instruct at least one storage cluster in the first source region to migrate cold data objects stored in at least one first bucket to at least one second bucket included in at least one storage cluster in the first destination region based on the identification information of at least one first bucket and the identification information of the first destination region.

[0021] In another possible implementation, prediction information for multiple regions is obtained. Based on the prediction information for multiple regions, the second water level of multiple regions within a first time period, which is after the current time, is determined. Based on the first water levels of the multiple regions, m regions with the highest first water level, or whose first water level is higher than a first threshold, and n regions with the lowest first water level, or whose first water level is lower than a second threshold, are selected from the multiple regions, where m and n are both integers greater than 1, and the second threshold is less than or equal to the first threshold. Based on the second water levels of the m regions, at least one region with the highest second water level, or whose second water level is higher than the first threshold, is selected as the first source region. Based on the second water levels of the n regions, at least one region with the lowest second water level, or whose second water level is lower than the second threshold, is selected as the first destination region.

[0022] The first source region is the region with the highest second water level, or the region with the second water level higher than the first threshold. Therefore, the storage capacity of the first source region is heavily used both currently and during the first time period. So migrating cold data objects in the first source region will not result in a large amount of idle storage capacity during the first time period.

[0023] The first destination region is selected as the region with the lowest second water level, or the region with the second water level below the second threshold. Therefore, the first destination region has a large amount of free storage capacity both at present and during the first time period. So, migrating cold data objects to the first destination region will not result in insufficient storage capacity during the first time period.

[0024] In another possible implementation, for each region, the prediction information for the region includes one or more of the following: the increment of storage capacity used by the region in each of at least one second time period, or data growth information, where at least one second time period is prior to the current time, and the data growth information is used to describe the increment of data objects in the region in the first time period compared to at least one second time period.

[0025] Secondly, this application provides an apparatus for processing data objects, used to perform the method in the first aspect or any possible implementation thereof. Specifically, the apparatus includes units for performing the method in the first aspect or any possible implementation thereof.

[0026] Thirdly, this application provides a computing device cluster, the computing device cluster including at least one computing device, each computing device including a processor and a memory;

[0027] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of the first aspect or any possible implementation thereof.

[0028] Fourthly, this application provides a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method of the first aspect or any possible implementation thereof.

[0029] Fifthly, this application provides a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, perform the method of the first aspect or any possible implementation thereof.

[0030] In a sixth aspect, this application provides a chip including a memory and a processor, wherein the memory is used to store computer instructions, and the processor is used to retrieve and execute the computer instructions from the memory to perform the method in the first aspect or any possible implementation of the first aspect. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the structure of a cloud storage system provided in an embodiment of this application;

[0032] Figure 2 This is a schematic diagram of a network architecture provided in an embodiment of this application;

[0033] Figure 3 This is a schematic diagram of a migration data object provided in an embodiment of this application;

[0034] Figure 4 This is a schematic diagram of the structure of a control system provided in an embodiment of this application;

[0035] Figure 5 This is a flowchart of a method for processing data objects provided in an embodiment of this application;

[0036] Figure 6 This is a flowchart of another method for processing data objects provided in an embodiment of this application;

[0037] Figure 7 This is a schematic diagram of a device structure for processing data objects provided in an embodiment of this application;

[0038] Figure 8This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0039] Figure 9 This is a schematic diagram of a cluster structure for processing data objects provided in an embodiment of this application;

[0040] Figure 10 This is a schematic diagram of a cluster structure for processing data objects provided in an embodiment of this application. Detailed Implementation

[0041] See Figure 1 This application provides a cloud storage system, which includes multiple storage clusters 101 distributed in multiple regions. Each region includes at least one storage cluster, and the multiple storage clusters 101 can communicate with each other.

[0042] In some embodiments, multiple storage clusters 101 in the cloud storage system can be used to store data objects for object storage services, and the storage capacity of multiple regions can be planned based on the demands of object storage services. The storage capacity of a region includes the storage capacity of at least one storage cluster in the region; that is, the total storage capacity of a region is equal to the sum of the storage capacities of each storage cluster in the region.

[0043] For example, such as Figure 1 As shown, the cloud storage system includes storage clusters 101a, 101b, 101c, 101d, 101e, and 101f, and multiple regions including region1, region2, and region3. Region1 includes storage clusters 101a and 101b, region2 includes storage clusters 101c and 101d, and region3 includes storage clusters 101e and 101f.

[0044] The storage capacity of region 1 includes the storage capacity of storage cluster 101a and storage cluster 101b. That is, the total storage capacity of region 1 is equal to the sum of the storage capacities of storage clusters 101a and 101b within region 1. The storage capacity of region 2 includes the storage capacity of storage clusters 101c and 101d. That is, the total storage capacity of region 2 is equal to the sum of the storage capacities of storage clusters 101c and 101d within region 2. The storage capacity of region 3 includes the storage capacity of storage clusters 101e and 101f. That is, the total storage capacity of region 3 is equal to the sum of the storage capacities of storage clusters 101e and 101f within region 3.

[0045] For each storage cluster included in a region, there is a storage pool that includes storage resources such as memory, hard disks, and / or cache. The storage capacity of the storage cluster is equal to the capacity of the storage pools it includes.

[0046] For each of the multiple regions, the water level varies due to different user usage within each region. The region water level indicates the storage capacity currently in use within that region. Generally, a higher water level indicates a higher ratio of used storage resources to total storage capacity, and vice versa.

[0047] If the user usage in a region is low, the region's water level will be low, indicating that the ratio of the region's used storage capacity to its total storage capacity is low, the region has a lot of free storage capacity, and the region has a storage capacity surplus.

[0048] The higher the user usage in a region, the higher the region's water level becomes, the higher the ratio of the region's used storage capacity to its total storage capacity, and the less free storage capacity the region has. When a region experiences insufficient storage capacity, it needs to be expanded.

[0049] Therefore, the uneven distribution of water levels across multiple regions has resulted in some regions having excess storage capacity, wasting a large amount of storage resources, while others have insufficient storage capacity and urgently need expansion, increasing the operating costs of the cloud storage system.

[0050] To balance water levels across regions and reduce the operating costs of the cloud storage system, see [link / reference]. Figure 2 The network architecture shown in this application embodiment adds a control system 102. The control system 102 can communicate with multiple storage clusters 101 included in the cloud storage system. The control system 102 is used to perform capacity balancing on the multiple storage clusters 101 included in the cloud storage system.

[0051] The control system 102 can acquire the first water level of multiple regions, and based on the first water level of multiple regions, determine the first source region and the first destination region from the multiple regions.

[0052] The first source region is at least one region with the highest first water level among multiple regions, or the first source region is at least one region with the first water level higher than the first threshold among multiple regions. Therefore, the storage capacity of the first source region is heavily used, which may lead to insufficient storage capacity.

[0053] The first destination region is at least one of the regions with the lowest water level among multiple regions, or the first destination region is at least one of the multiple regions with a first water level lower than a second threshold, where the second threshold is less than or equal to the first threshold. Therefore, there is a large amount of free storage capacity in the storage capacity of the first destination region.

[0054] Then, the control system 102 migrates the cold data objects stored in at least one storage cluster 101 in the first source region to at least one storage cluster 101 in the first destination region based on at least one storage cluster 101 in the first source region.

[0055] This approach retains the hot data objects in the first source region, ensuring that users in the first source region can access them normally. Meanwhile, cold data objects are migrated from the first source region to the first destination region. This frees up a large amount of used storage capacity in the first source region and utilizes a large amount of idle storage capacity in the first destination region, thereby balancing the storage levels of multiple regions as much as possible and reducing the operating costs of the cloud storage system.

[0056] In some embodiments, the control system 102 may further determine a second source region and a second destination region, wherein at least one storage cluster in the second source region includes a target data object, and the number of accesses to the target data object from the second destination region exceeds a third threshold. Based on at least one storage cluster in the second source region, the target data object is stored in at least one storage cluster included in the second destination region. This allows users in the second destination region to access the target data object from the nearest storage cluster in the second destination region.

[0057] See Figure 3 At least one storage cluster 101 in the region includes multiple storage buckets. For each storage cluster 101, the storage buckets included in the storage cluster 101 are located in the storage pool of the storage cluster 101, and the storage buckets included in the storage cluster 101 are used to store at least one data object. Optionally, the data object can be an object file, or other forms of data, which will not be listed here.

[0058] For the storage buckets included in the storage cluster 101, the capacity of bucket 101 is equal to the sum of the data amounts of each data object stored in the bucket. Therefore, the capacity of a bucket increases as the number of data objects stored in the bucket increases. Each data object stored in a bucket belongs to the same user, and the bucket belongs to that user. The used storage capacity of the storage cluster 101 is equal to the sum of the capacities of each storage bucket included in the storage cluster.

[0059] See Figure 3 When migrating cold data objects, cold data objects stored in at least one first storage bucket included in at least one storage cluster 101 in the first source region can be migrated to at least one second storage bucket included in at least one storage cluster 101 in the first destination region. The first storage bucket and the second storage bucket belong to the same user, and the first storage bucket is any one of at least one first storage bucket, and the second storage bucket is any one of at least one second storage bucket. That is, at least one first storage bucket and at least one second storage bucket correspond one-to-one, and the one-to-one corresponding first storage bucket and second storage bucket belong to the same user.

[0060] In some embodiments, see Figure 4 The control system 102 shown may include an operation and maintenance platform 1021, a data lake 1022, and a task queue 1023.

[0061] For each storage cluster 101 in a region, the storage cluster 101 can send at least one piece of data to the control system 102. The operation and maintenance platform 1021 of the control system 102 can store the at least one piece of data sent by the storage cluster 101 in the data lake 1022, so the data lake 1022 may store data sent by various storage clusters 101 in multiple regions. Based on the data stored in the data lake 1022, the operation and maintenance platform 1021 can obtain the first water level of each region in multiple regions and the cold data objects stored in at least one storage cluster 101 in each region.

[0062] Then, the operation and maintenance platform 1021 can determine the first source region and the first destination region based on the first water level of multiple regions, and generate a migration command. The migration command is used to instruct the migration of cold data objects stored in at least one storage cluster 101 in the first source region to at least one storage cluster 101 in the first destination region. The operation and maintenance platform 1021 can generate at least one migration command and save at least one migration command from the tail of the task queue 1023 to the task queue 1023.

[0063] Then, after generating all migration commands, the operation and maintenance platform 1021 retrieves the migration commands from the head of the task queue 1023 and sends them to at least one storage cluster 1201 in the first source region. This migration command, along with the migration of cold data objects stored in at least one storage cluster 101 in the first source region to at least one storage cluster 101 in the first destination region, is executed according to the instructions in the migration command. For a detailed description of the data object migration process, please refer to any of the following embodiments; it will not be detailed here.

[0064] See Figure 5 This application provides a method 500 for processing data objects, which is applied to... Figure 2 or Figure 4 The control system 102 in the illustrated embodiment. The method 500 is used to determine a first source region and a first destination region, wherein a first water level in the first source region is higher than a first water level in the first destination region, and to migrate cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region. The method 500 includes the following steps.

[0065] Step 501: The control system acquires the first water level of multiple regions, which is used to indicate the storage capacity that has been used in the region.

[0066] In step 501, for each of the multiple regions, the control system obtains the storage capacity of each storage cluster in at least one storage cluster within that region, and also obtains the capacity of multiple storage buckets included in the at least one storage cluster within that region. The capacity of a storage bucket is equal to the cumulative amount of data for each data object stored in that bucket. Based on the capacities of the multiple storage buckets, the system obtains the storage capacity already used in that region, and based on the storage capacity of each storage cluster included in that region, it obtains the total storage capacity of that region. Based on the storage capacity already used in that region and the total storage capacity of that region, the system obtains the first water level for that region.

[0067] Alternatively, in implementation, the control system can obtain the first water level of the region through the following operations 5011-5016.

[0068] 5011: For each storage cluster in the cloud storage system, the control system receives at least one piece of data from the storage cluster, the at least one piece of data including one or more of the following: capacity logs, or metadata of each bucket included in the storage cluster.

[0069] Optionally, the capacity log includes the identifier information of the region to which the storage cluster belongs and the storage capacity of the storage cluster. The storage capacity of the storage cluster may change; for example, the storage pool of the storage cluster may be expanded or shrunk, and the storage capacity of the storage cluster may be equal to the capacity of the storage pool, thus changing the storage capacity of the storage cluster.

[0070] Optionally, the metadata of a storage bucket may include one or more of the following: the bucket's identification information, the identification information of the storage cluster to which the bucket belongs, the identification information of the region to which the bucket belongs, the capacity of the bucket, or the number of data objects stored in the bucket.

[0071] In some embodiments, at least one type of data may also include access logs, metadata of each data object stored in the storage cluster, or tenant data on the storage cluster.

[0072] Optionally, the access log includes one or more of the following: the identification information of the data object accessed by the user in the region, the type of access operation to the data object, or the access timestamp of the data object.

[0073] Optionally, the metadata of a data object includes one or more of the following: the identification information of the data object, the identification information of the bucket to which the data object belongs, the identification information of the storage cluster to which the data object belongs, the identification information of the region to which the data object belongs, the data volume of the data object, the hot / cold flag of the data object, the restriction information of the data object, the deletion timestamp of the deleted data object, the user identification information to which the data object belongs, or the type of the data object, etc.

[0074] The hot / cold tag for a data object is used to mark whether a data object is a hot or cold data object.

[0075] The data object's constraint information is used to indicate whether the data object can be migrated.

[0076] Optionally, tenant data may include one or more of the following: tenant identification information, identification information of at least one bucket belonging to the tenant, or tenant restriction information, etc.

[0077] Tenant restriction information is used to indicate whether data objects stored in a tenant's bucket can be migrated.

[0078] 5012: The control system saves at least one piece of received data to the control system's data lake.

[0079] In some embodiments, the control system may extract useful fields from each data point in at least one set of data, and store the useful fields in each data point in the control system's data lake.

[0080] For each data set, the useful fields within that data are those required to migrate the data object. For example, fields required to migrate a data object might include those needed to obtain the first water level of the region, and / or those needed to determine the hot or cold status of data objects stored in the storage cluster within the region.

[0081] For example, the storage capacity of the storage cluster and the region identification information included in the capacity log are the information needed to obtain the first water level of the region, so the storage capacity of the storage cluster and the region identification information included in the capacity log are useful fields.

[0082] For example, the metadata of a bucket includes the bucket's identifier, the identifier of the storage cluster to which the bucket belongs, the identifier of the region to which the bucket belongs, and the bucket's capacity. These are fields required for migrating data objects. Therefore, the bucket's metadata, including the bucket's identifier, the identifier of the storage cluster to which the bucket belongs, the identifier of the region to which the bucket belongs, and the bucket's capacity, are useful fields.

[0083] For example, the identification information and access timestamps of data objects included in the access log are information needed to calculate the access popularity of data objects, which in turn is needed to determine the access frequency of data objects. Therefore, the identification information and access timestamps of data objects included in the access log are useful fields.

[0084] For example, the metadata of a data object includes fields such as the data object's identifier, the identifier of the bucket it belongs to, the identifier of the storage cluster it belongs to, the identifier of the region it belongs to, the data size of the data object, the cold / hot flag of the data object, the data object's limitations, the deletion timestamp of a deleted data object, the user identifier of the data object, or the data object's type. These are fields required for migrating the data object. The metadata of a data object includes fields such as the data object's identifier, the identifier of the bucket it belongs to, the identifier of the storage cluster it belongs to, the identifier of the region it belongs to, the data size of the data object, the cold / hot flag of the data object, the data object's limitations, the deletion timestamp of a deleted data object, the user identifier of the data object, or the data object's type. These are useful fields.

[0085] For example, tenant data includes fields such as the identifier of at least one bucket belonging to the tenant and the tenant's restriction information, which are fields required for migrating data objects. These fields are useful.

[0086] By storing useful fields from the data in the data lake, the storage resources consumed by the data lake can be reduced.

[0087] The control system can receive at least one data sent by each storage cluster in the cloud storage system according to the operations described in 5011-5012 above, save at least one data sent by each storage cluster to the data lake of the control system, and then the control system obtains the first water level of multiple regions according to the following operations.

[0088] 5013: For each region, the control system retrieves the capacity logs of at least one storage cluster included in the region and the metadata of multiple storage buckets in the region from the data lake.

[0089] In 5013, the control system retrieves at least one capacity log containing the region's identification information, and metadata for multiple buckets containing the region's identification information, from the data lake. The at least one capacity log includes the capacity log for each storage cluster within the region. The metadata for the multiple buckets containing the region's identification information includes the metadata of the buckets stored in each storage cluster within the region; that is, the metadata for the multiple buckets is the metadata of the multiple buckets within the region.

[0090] 5014: The control system obtains the total storage capacity of the region based on the capacity logs of each storage cluster in the region.

[0091] The capacity logs for each storage cluster in this region include the storage capacity of each storage cluster. In 5014, the control system retrieves the storage capacity of each storage cluster from its capacity logs, calculates the sum of the storage capacities of each storage cluster, and obtains the total storage capacity of the region.

[0092] 5015: The control system obtains the used storage capacity of the region based on the metadata of multiple storage buckets in the region.

[0093] The metadata of a storage bucket includes its capacity. In 5015, the control system obtains the capacity of each storage bucket from the metadata of each storage bucket in the region, calculates the cumulative value between the capacities of each storage bucket, and obtains the storage capacity used in the region.

[0094] 5016: The control system obtains the first water level of the region based on the storage capacity used in the region and the total storage capacity of the region.

[0095] In 5016, the control system obtains the ratio between the used storage capacity of the region and the total storage capacity of the region to determine the first water level of the region. Alternatively, the control system multiplies this ratio by a specified coefficient to obtain the first water level of the region.

[0096] For example, suppose the control system obtains a ratio of 0.7 between the used storage capacity of a region and the total storage capacity of the region, and the resulting first water level for that region is 0.7. Alternatively, suppose a specified coefficient is 10, and multiplying the ratio 0.7 by the specified coefficient 10, the resulting first water level for that region is 7.

[0097] For each of the other regions in the multiple regions, the first water level of each of the other regions is obtained by performing the operations described in 5013-5016 above, thus obtaining the first water level of the multiple regions.

[0098] Step 502: The control system determines the first source region and the first destination region based on the first water level of multiple regions, wherein the first water level of the first source region is higher than the first water level of the first destination region.

[0099] The first source region is at least one region with the highest first water level among multiple regions, or, the first source region is at least one region with a first water level higher than a first threshold among multiple regions. Therefore, the first source region is a region with a high first water level, meaning that the ratio between the used storage capacity of the first source region and the total storage capacity of the first source region is high, the storage capacity of the first source region is heavily used, the free storage capacity of the first source region is low, and the first source region is experiencing a storage capacity shortage.

[0100] The first destination region is at least one region among multiple regions with the lowest first water level, or, the first destination region is at least one region among multiple regions whose first water level is lower than a second threshold, where the second threshold is less than or equal to the first threshold. Therefore, the first destination region is a region with a low first water level, meaning the ratio of used storage capacity to total storage capacity in the first destination region is low, the free storage capacity in the first destination region is large, and the first destination region exhibits a storage capacity surplus.

[0101] In step 502, the control system determines, based on the first water levels of multiple regions, at least one first source region and at least one first destination region. Optionally, in implementation:

[0102] The control system selects at least one region with the highest first water level from multiple regions as the first source region, or selects at least one region with a first water level higher than a first threshold as the first source region.

[0103] Based on the first water level of multiple regions, at least one region with the lowest first water level is selected as the first target region, or at least one region with a first water level lower than a second threshold is selected as the first target region.

[0104] In some embodiments, the control system can also predict the second water levels of multiple regions within a first time period, which is after the current time. This allows the control system to determine, based on the first and second water levels of multiple regions, the first source region whose storage capacity is heavily used during the current time and the first time period. In other words, the first source region has relatively low free storage capacity during both the current time and the first time period, indicating insufficient storage capacity during both periods.

[0105] Furthermore, based on the first and second water levels of multiple regions, the first destination region with a large amount of free storage capacity is identified in both the current time and the first time period. In other words, the first destination region has a surplus of storage capacity in both the current time and the first time period.

[0106] This avoids situations where the first destination region has insufficient storage capacity and the first source region has excessive storage capacity during the first time period after cold data objects in the first source region are migrated to the first destination region.

[0107] Optionally, during implementation, the control system can press operations 5021-5025 to determine the first source region with smaller free storage capacity in both the current time and the first time period, and the first destination region with larger free storage capacity in both the current time and the first time period.

[0108] 5021: Obtain prediction information for multiple regions.

[0109] For each region, the region's prediction information includes one or more of the following: the increment of storage capacity used by the region in each of at least one second time period, or data growth information, which describes the increment of data objects in the region in a first time period compared to at least one second time period, where at least one second time period is prior to the current time.

[0110] The duration of the first time period and / or the duration of the second time period are specified durations, such as the duration of the first time period and the duration of the second time period being one day, one week, half a month, one month, or one quarter, etc., or the duration of the first time period and the duration of the second time period being x days, x weeks, x months, or x quarters, etc., where x is an integer greater than 1.

[0111] In some embodiments, for each region, the control system calculates the total amount of data objects stored in at least one storage cluster included in the region during each second time period. Based on the total amount of data objects stored in at least one storage cluster included in the region during each second time period, the increment of storage capacity used by the region during each second time period is obtained.

[0112] For example, for each second time period, the total amount of data objects stored in at least one storage cluster included in the region within the second time period is subtracted from the total amount of data objects stored in at least one storage cluster included in the region within the previous second time period to obtain the increment of data objects in the region within the second time period. The increment of data objects in the region within the second time period is the increment of storage capacity used by the region within the second time period.

[0113] In some embodiments, for each region, data growth information may include one or more of the following: the increment of data objects in the region within a first time period, holidays or a certain season within the first time period, etc.

[0114] Optionally, due to factors such as holidays or seasonality, user usage within the region may increase significantly, leading to a substantial increase in the total amount of data objects stored in the storage cluster encompassing the region. The first time period is the holiday or the season of significant user usage increase.

[0115] For example, during the summer, user activity within a region may increase, resulting in a significantly larger volume of data objects being stored in the region's storage cluster compared to other seasons, with the summer being the first time period.

[0116] For example, the data objects stored in the storage cluster of a region are user trajectory data. During holidays, a large number of users travel, which causes a significant increase in the amount of data stored in the storage cluster of the region during the holiday period. The first time period is the holiday.

[0117] In 5021, when a holiday or season arrives, the control system uses that holiday or season as forecast information.

[0118] Optionally, the increment of data objects in the region within the first time period may be input to the control system by technicians. For example, if the first time period is a holiday or a season with a significant increase in user usage, technicians can calculate the average increment of data objects in the region during the past holidays or seasons, obtain the increment of data objects in the region within the first time period based on the average increment, and then input the increment of data objects in the region within the first time period into the control system.

[0119] 5022: Based on the forecast information of multiple regions, determine the second water level of multiple regions within the first time period, which is after the current time.

[0120] In some embodiments, for each region, the prediction information for the region includes the increment of storage capacity used by the region in each of the second time periods within at least one second time period.

[0121] The control system can statistically analyze the increase in storage capacity used by the region within each of the past at least one second time period. Based on the increase in storage capacity used by the region within each second time period, it can obtain the storage capacity used by the region within the first time period. Based on the storage capacity used by the region within the first time period and the total storage capacity of the region, it can obtain the second water level of the region within the first time period. Optionally, in implementation:

[0122] The control system calculates the average increment based on the increase in storage capacity used by the region within each second time period, and uses this average increment as the increase in storage capacity used by the region within the first time period. Alternatively, it selects the median from the increases in storage capacity used by the region within each second time period and uses this median as the increase in storage capacity used by the region within the first time period. Based on the storage capacity used by the region in the current time period and the increase in storage capacity used by the region within the first time period, the storage capacity used by the region within the first time period is obtained. Based on the storage capacity used by the region within the first time period and the total storage capacity of the region, the second water level of the region within the first time period is obtained.

[0123] In some embodiments, the region's prediction information includes the increment of data objects in the region within a first time period. The control system determines the storage capacity used by the region within the first time period based on the storage capacity used by the region in the current time and the increment of data objects in the region within the first time period. Based on the storage capacity used by the region within the first time period and the total storage capacity of the region, a second water level of the region within the first time period is obtained.

[0124] In some embodiments, the region's forecast information includes the aforementioned holidays or seasons. The control system calculates the average increment of data objects in the region during the past holidays or seasons, and obtains the increment of data objects in the region within a first time period based on the average increment.

[0125] Optionally, the control system may use this average increment as the increment of data objects in the region within the first time period; alternatively, the control system may amplify the average increment by a specified factor to obtain the increment of data objects in the region within the first time period. Based on the storage capacity used by the region in the current time period and the increment of data objects in the region within the first time period, the storage capacity used by the region within the first time period is obtained. Based on the storage capacity used by the region within the first time period and the total storage capacity of the region, the second water level of the region within the first time period is obtained.

[0126] Besides the methods listed above for predicting the storage capacity used by a region in the first time period, there are other methods. For example, based on the prediction information, a capacity growth prediction algorithm can be used to predict the storage capacity used by a region in the first time period.

[0127] 5023: Based on the first water level of multiple regions, select m regions with the highest first water level, or with the first water level higher than a first threshold, and select n regions with the lowest first water level, or with the first water level lower than a second threshold, where m and n are both integers greater than 1.

[0128] 5024: Based on the second water level of m regions, select the region with the highest second water level or at least the region with a second water level higher than a first threshold as the first source region.

[0129] 5025: Based on the second water level of n regions, select the region with the lowest second water level or at least the region with a second water level lower than a second threshold as the first target region.

[0130] In some embodiments, the control system may determine a first source region and a first destination region based on the second water levels of multiple regions. Optionally, the control system may select at least one region with the highest second water level or a second water level higher than a first threshold as the first source region, and select at least one region with the lowest second water level or a second water level lower than a second threshold as the first destination region.

[0131] Step 503: The control system migrates the cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region.

[0132] In step 503, the control system can migrate cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region based on the following operations 5031-5033.

[0133] 5031: Obtain the total amount of cold data objects stored in each bucket of at least one storage cluster in the first source region.

[0134] The control system's data lake includes metadata of data objects in multiple regions, metadata of storage buckets, and access logs of storage clusters. The control system obtains at least one first source region. For each first source region, the control system, based on the metadata of data objects in the first source region, the metadata of storage buckets, and the access logs of at least one storage cluster, obtains the total amount of cold data objects stored in each storage bucket of the at least one storage cluster in the first source region. Optionally, in implementation, the control system can be implemented through the following process (11)-(15).

[0135] (11): The control system obtains the metadata of multiple storage buckets, including the identification information of the first source region, from the metadata of each storage bucket stored in the data lake based on the identification information of the first source region, that is, obtains the metadata of multiple storage buckets included in at least one storage cluster in the first source region.

[0136] The metadata of a storage bucket includes one or more of the following: the bucket's identification information, the identification information of the storage cluster to which the bucket belongs, or the identification information of the region to which the bucket belongs, etc.

[0137] For each bucket, based on its identification information, tenant data, including the bucket's identification information, is retrieved from the data lake in the control system. This tenant data includes the tenant's restriction information. Only when the tenant's restriction information indicates that data objects stored in the tenant's bucket can be migrated, the process continues to obtain the total amount of cold data objects stored in that bucket. If the tenant's restriction information indicates that data objects stored in the tenant's bucket cannot be migrated, the process of obtaining the total amount of cold data objects stored in that bucket is stopped.

[0138] (12): For each of the multiple storage buckets, obtain the identification information of the storage bucket and the identification information of the storage cluster to which the storage bucket belongs from the metadata of the storage bucket.

[0139] (13): Obtain the metadata of at least one data object from the metadata of each data object stored in the data lake, including the identification information of the first source region, the identification information of the storage bucket, and the identification information of the storage cluster to which the storage bucket belongs, and obtain the metadata of at least one data object stored in the storage bucket.

[0140] The metadata of a data object includes one or more of the following: the identification information of the data object, the identification information of the bucket to which the data object belongs, the identification information of the storage cluster to which the data object belongs, the identification information of the region to which the data object belongs, the data volume of the data object, the cold / hot flag of the data object, or the limitation information of the data object, etc.

[0141] In different regions, there may be buckets belonging to the same user. The identification information of buckets belonging to the same user may be the same or different. The identification information of buckets in different storage clusters may also be the same, and the same data object may be stored in multiple different regions. Therefore, for at least one data object that includes the identification information of the first source region, the identification information of the bucket, and the identification information of the storage cluster to which the bucket belongs, the at least one data object is a data object stored in the bucket, thus improving the accuracy of obtaining the metadata of at least one data object stored in the bucket.

[0142] (14): Based on the metadata of the at least one data object, obtain the amount of data for each cold data object stored in the bucket.

[0143] In some embodiments, for each of the at least one data object, if the metadata of the data object includes a hot / cold tag for the data object, then when the hot / cold tag is used to mark the data object as a cold data object, the data volume of the cold data object is obtained from the metadata of the data object. The data volume of each other cold data object stored in the bucket is obtained in the same manner as described above.

[0144] In some embodiments, for each of the at least one data object, if the metadata of the data object does not include a hot / cold flag for the data object, the access frequency of the data object within a target time period is obtained. The target time period is the time period most recent to the present with a specified duration. Based on the access frequency of the data object, it is determined whether the data object is a cold data object. If the data object is a cold data object, the data volume of the cold data object is obtained from the metadata of the data object. The data volume of each other cold data object stored in the bucket is obtained in the same manner as described above.

[0145] Optionally, obtaining the access frequency of the data object within the target time period can be achieved by: retrieving access logs from the data lake, including the identification information of the data object, and the access logs including access timestamps of different users accessing the data object. Based on the access timestamps of users accessing the data object within the target time period included in the access logs, the access frequency of the data object within the target time period is obtained.

[0146] Optionally, determining whether a data object is a cold data object based on its access frequency can be as follows: If the access frequency of the data object is less than or equal to a first frequency threshold, the data object is determined to be a cold data object. If the access frequency of the data object is greater than the first frequency threshold but less than a second frequency threshold, obtain the y most recent access timestamps of the data object from the access log, where y is an integer greater than 1, and the second frequency threshold is greater than the first frequency threshold. Based on the y access timestamps, obtain y-1 access intervals, where y-1 access intervals include the interval between any two adjacent access timestamps. If all y-1 access intervals are greater than an interval threshold, the data object is determined to be a cold data object.

[0147] Optionally, a data object is determined to be a hot data object when the access frequency of the data object is greater than or equal to the second frequency threshold, or when there is an access interval less than or equal to the interval threshold among the y-1 access intervals.

[0148] (15): Based on the amount of data of each cold data object stored in the bucket, obtain the total amount of data of the cold data objects stored in the bucket.

[0149] In (15), the sum of the data volume of each cold data object stored in the bucket is calculated to obtain the total data volume of the cold data objects stored in the bucket.

[0150] Repeat the above operations (12)-(15) to obtain the total amount of cold data objects stored in each of the other buckets in at least one storage cluster in the first source region.

[0151] 5032: Based on the total amount of cold data objects stored in each bucket, select at least one first bucket with the largest total amount of cold data objects stored, or at least one first bucket with the total amount of cold data objects stored being greater than the fourth threshold.

[0152] Select the first bucket with the largest total amount of stored cold data objects, or at least the first bucket with a total amount of stored cold data objects greater than the fourth threshold, as the bucket to be migrated. This can reduce the number of buckets that need to be migrated, and free up more storage capacity in the first source region.

[0153] In some embodiments, based on the total amount of cold data objects stored in each bucket, multiple buckets with the largest total amount of stored cold data objects, or whose total amount of stored cold data objects is greater than a fourth threshold, are selected. The number of cold data objects stored in each of these multiple buckets is obtained. Based on the number of cold data objects stored in the multiple buckets, at least one first bucket with the smallest number of stored cold data objects, or whose number of stored cold data objects is less than a fifth threshold, is selected from the multiple buckets.

[0154] Select the bucket with the smallest number of stored cold data objects from multiple buckets, or at least the bucket with the number of stored cold data objects less than the fifth threshold, as the bucket to be migrated. This ensures that the cold data objects in each first bucket are cold data objects with a large amount of data. In this way, more storage capacity in the first source region can be released while reducing the number of times data objects are migrated, thus reducing migration costs.

[0155] In the 5032, the control system can generate one or more migration commands. When generating a single migration command, it includes at least one identifier for a first bucket and an identifier for a first destination region. When generating multiple migration commands, each command includes identifiers for at least a portion of the first bucket and an identifier for a first destination region. The first destination region is different for each migration command, allowing cold data objects from the first source region to be migrated to different destination regions. The generated migration commands are saved to a task queue.

[0156] In some embodiments, for each first storage bucket, the control system obtains the remaining lifetime of cold data objects stored in the first storage bucket. Based on the remaining lifetime of the cold data objects stored in the first storage bucket, at least one cold data object with the longest remaining lifetime, or whose remaining lifetime exceeds a sixth threshold, is selected from the first storage bucket. This at least one cold data object is the cold data object to be migrated. For a migration command that includes the identification information of the first storage bucket, the migration command also includes the identification information of the at least one cold data object to be migrated. Thus, the migration command can instruct at least one storage cluster in the first source region to migrate the at least one cold data object to be migrated to a second storage bucket included in at least one storage cluster in the first destination region, wherein the first storage bucket and the second storage bucket belong to the same user.

[0157] Optionally, the remaining lifespan of cold data objects stored in the first bucket can be obtained in the following two ways:

[0158] Method 1: The control system obtains the attribute information of cold data objects, and based on the deletion timestamp prediction model and the attribute information of cold data objects, obtains the deletion timestamp of cold data objects, and based on the deletion timestamp of cold data objects, obtains the remaining survival time of cold data objects.

[0159] In some embodiments, the attribute information of a cold data object may include one or more of the following: the data volume, type, user identification information, or storage path to which the cold data object belongs. The storage path may include the identification information of the storage bucket to which the cold data object belongs, the identification information of the storage cluster to which it belongs, and the identification information of the region to which it belongs.

[0160] Optionally, the metadata of a cold data object includes the attribute information of the cold data object.

[0161] In Method 1, the control system acquires the metadata of the cold data object, retrieves its attribute information from the metadata, and inputs this attribute information into the deletion timestamp prediction model. The deletion timestamp prediction model receives the attribute information and infers the deletion timestamp of the cold data object based on this information. The control system then acquires the deletion timestamp of the cold data object output by the deletion timestamp prediction model and, based on the current timestamp and the deletion timestamp, calculates the remaining lifespan of the cold data object.

[0162] Optionally, the deletion timestamp prediction model can be obtained by training an artificial intelligence (AI) model in advance based on multiple training samples. Each training sample includes attribute information of a data object and the deletion timestamp of deleting the data object.

[0163] Method 2: The control system obtains the metadata of the cold data object, which includes the deletion timestamp of the cold data object. Based on the deletion timestamp of the cold data object, the remaining lifespan of the cold data object is obtained.

[0164] Optionally, the metadata of a cold data object may include a deletion timestamp of the cold data object, which may have been configured by the user to whom the cold data object belongs when the cold data object was created.

[0165] Optionally, if the metadata of a cold data object does not include its deletion timestamp, the control system uses method 1 to obtain the deletion timestamp. If the metadata of a cold data object includes its deletion timestamp, the control system uses either method 1 or method 2 to obtain it, offering greater flexibility.

[0166] In some embodiments, for each first storage bucket, the control system acquires metadata for each cold data object stored in the first storage bucket. The metadata of the cold data object includes restriction information for the cold data object, which indicates whether the cold data object can be migrated. Based on the restriction information included in the metadata of each cold data object, at least one cold data object that can be migrated is selected from the first storage bucket, and the at least one cold data object is the cold data object to be migrated. For a migration command that includes the identification information of the first storage bucket, the migration command also includes the identification information of the at least one cold data object to be migrated.

[0167] The control system obtains at least one first source region. Following the operations described in 5031-5032 above, it can obtain one or more migration commands corresponding to each first source region and save one or more migration commands corresponding to each first source region to the task queue.

[0168] 5033: Based on at least one storage cluster in a first source region, migrate cold data objects stored in at least one first storage bucket to at least one second storage bucket included in at least one storage cluster in a first destination region.

[0169] Optionally, the at least one first storage bucket and the at least one second storage bucket are in one-to-one correspondence, and the one-to-one correspondence between the first storage bucket and the second storage bucket belongs to the same user.

[0170] In 5033, the control system retrieves a migration command from the task queue and sends the migration command to at least one storage cluster in the first source region based on the identification information of the first source region included in the migration command. The migration command includes the identification information of at least one first storage bucket and the identification information of the first destination region. The migration command is used to instruct at least one storage cluster in the first source region to migrate the cold data objects stored in at least one first storage bucket to at least one second storage bucket included in at least one storage cluster in the first destination region based on the identification information of at least one first storage bucket and the identification information of the first destination region.

[0171] For each storage cluster in the first source region (hereinafter referred to as the first source storage cluster for ease of explanation), the first source storage cluster receives a migration command and determines the first storage buckets included in the first source storage cluster from at least one first storage bucket corresponding to the identification information of at least one first storage bucket included in the migration command. Based on the identification information of the first destination region included in the migration command, a storage cluster is selected from at least one storage cluster in the first destination region as the first destination storage cluster. Cold data objects stored in the first storage buckets included in the first source storage cluster are migrated to second storage buckets included in the first destination storage cluster. The second storage bucket in the first destination storage cluster and the first storage bucket in the first source cluster belong to the same user.

[0172] In some embodiments, at least one storage cluster in the first destination region contains a storage cluster that includes a second storage bucket belonging to the user. The first source storage cluster selects this storage cluster as the first destination storage cluster and sends the cold data object stored in the first storage bucket included in the first source storage cluster to the first destination storage cluster. The first destination storage cluster then saves the cold data object to the second storage bucket. After the first destination storage cluster saves the cold data object, the first source storage cluster may delete the cold data object.

[0173] If at least one storage cluster in the first destination region does not contain a second bucket belonging to a user, the first source storage cluster randomly selects one from at least one storage cluster in the first destination region as the first destination storage cluster, or selects the storage cluster with the largest free capacity as the first destination storage cluster. The cold data object stored in the first bucket included in the first source storage cluster is then sent to the first destination storage cluster. The first destination storage cluster creates a second bucket and stores the cold data object in the second bucket. The created second bucket belongs to the same user as the first bucket on the first source storage cluster. After the first destination storage cluster stores the cold data object, the first source storage cluster may delete the cold data object.

[0174] In some embodiments, where the migration command includes identification information of at least one cold data object to be migrated in the first bucket, the first source storage cluster sends the at least one cold data object to be migrated stored in the first bucket to the destination storage cluster based on the identification information of the at least one cold data object to be migrated.

[0175] Since at least one cold data object to be migrated is the one with the longest remaining lifespan in the first bucket, or at least one cold data object with a remaining lifespan exceeding the sixth threshold, migrating at least one cold data object to the first destination storage cluster will result in it being stored in the first destination storage cluster for a longer period, thus increasing the benefits of migrating data objects. Alternatively, since at least one cold data object to be migrated is a data object in the first bucket that can be migrated, migrating at least one cold data object to the first destination storage cluster will not result in a migration error.

[0176] In this embodiment, the control system can receive capacity logs and bucket metadata from each storage cluster within each of multiple regions. Based on the capacity logs from each storage cluster in the region, the storage capacity of each storage cluster is obtained, and the total storage capacity of the region is obtained based on the storage capacity of each storage cluster. Based on the bucket metadata from each storage cluster within the region, the capacity of the buckets in each storage cluster is obtained, resulting in the total capacity of multiple buckets in the region. Based on the capacity of these multiple buckets, the used storage capacity of the region is obtained. Thus, based on the used storage capacity and the total storage capacity of the region, a first water level for the region is obtained. The control system can obtain the first water level for multiple regions in the above manner. This allows for unified scheduling and balancing of the capacity of multiple storage clusters within the cloud storage system based on the first water levels of multiple regions. Because the control system determines the first source region and the first destination region based on the first water level of multiple regions, the first source region is either the region with the highest first water level or at least one region with a first water level above a first threshold. Therefore, the ratio of used storage capacity in the first source region to its total storage capacity is relatively high, and the free storage capacity in the first source region is relatively low. Conversely, the first destination region is either the region with the lowest first water level or at least one region with a first water level below a second threshold. Therefore, the ratio of used storage space in the first destination region to its total storage capacity is relatively low, and the free storage capacity in the first destination region is relatively high. Migrating cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region can release some of the used storage capacity in the first source region and utilize the free storage capacity in the first destination region without expanding the capacity of the first source region. This approach aims to balance the water levels of the first source region and the first destination region, thereby reducing the operating costs of the cloud storage system.

[0177] See Figure 6 This application provides a method 600 for processing data objects, which is applied to... Figure 2 or Figure 4The control system 102 is shown. The method 600 is used to determine a second source region and a second destination region, where at least one storage cluster in the second source region includes a target data object, and the number of accesses to the target data object from the second destination region exceeds a third threshold, and the target data object is stored in at least one storage cluster in the second destination region. The method 600 includes the following steps.

[0178] Step 601: Control the system to obtain the access volume of data objects from other regions, including multiple regions.

[0179] Optionally, for each of the multiple regions, for a data object stored in at least one storage system in the region, the control system can obtain the access volume of the data object from other regions through the operations described in 6011-6014.

[0180] 6011: For each storage cluster in the cloud storage system, the control system receives the traffic logs sent by that storage cluster.

[0181] Traffic logs include one or more of the following: the identifier of the region to which the storage cluster belongs, the identifier of the data object accessed by the user through the storage cluster to the storage cluster in the source region, the access timestamp, the identifier of the storage cluster storing the data object, or the identifier of the source region, etc. The source region is the region that includes the data object that the user needs to access.

[0182] A user accesses a data object from the storage cluster, but the storage cluster does not store the data object. However, the data object is stored in the storage cluster of another region, which is the source region of the data object. The storage cluster retrieves the data object from the storage cluster of the source region and sends the data object to the user.

[0183] 6012: The control system saves the received traffic logs to the control system's data lake.

[0184] The control system can receive traffic logs sent by each storage cluster in the cloud storage system according to the operations described in 6011-6012 above, save the traffic logs sent by each storage cluster to the data lake of the control system, and then the control system can obtain the access volume of data objects in the region from other regions according to the following operations.

[0185] 6013: The control system selects one region as the source region and one region as the home region from multiple regions. For data objects stored in at least one storage cluster in the source region, it retrieves traffic logs from the data lake that include the identification information of the data object, the identification information of the source region, and the identification information of the home region.

[0186] The acquired traffic logs record the access timestamps of each user within the home region when accessing the data object from the source region.

[0187] 6014: The control system obtains the number of accesses to the data object from the home region based on the access timestamps of each user in the home region who accessed the data object from the source region, as included in the traffic log.

[0188] The control system repeats the operations described in 6013-6014 above to obtain the access volume of data objects from other regions included in multiple regions.

[0189] Step 602: The control system determines a second source region and a second destination region based on the access volume of data objects from other regions included in multiple regions. At least one storage cluster in the second source region includes the target data object, and the access volume of the target data object from the second destination region exceeds a third threshold.

[0190] For example, taking the access volume of a data object in the source region obtained by 6014 from its home region as an example, if the access volume exceeds the third threshold, then the data object is designated as the target data object, the source region is designated as the second source region, and the home region is designated as the second destination region. The access volume exceeding the third threshold indicates that a large number of users in the second destination region need to access the target data object.

[0191] Step 603: The control system stores the target data object into at least one storage cluster included in the second destination region based on at least one storage cluster in the second source region.

[0192] In step 603, the control system generates a storage command, which includes the identification information of the target data object and the identification information of the second destination region. The identification information of the storage cluster storing the target data object is obtained from the traffic log, which includes the identification information of the second source region, the identification information of the second destination region, and the identification information of the target data object. The storage cluster corresponding to the identification information of the storage cluster is designated as the second source storage cluster, and the storage command is sent to the second source storage cluster.

[0193] The second source storage cluster receives the storage command and determines a third storage bucket containing the target data object based on the target data object's identification information. Based on the identification information of the second destination region, it selects a storage cluster from at least one storage cluster within the second destination region as the second destination storage cluster. The target data object stored in the third storage bucket is then stored in a fourth storage bucket included in the second destination storage cluster; the third and fourth storage buckets belong to the same user.

[0194] In some embodiments, when at least one storage cluster in the second destination region contains a storage cluster that includes a fourth bucket belonging to the user, the second source storage cluster selects that storage cluster as the second destination storage cluster and sends the target data object to the second destination storage cluster. The second destination storage cluster then saves the target data object to the fourth bucket.

[0195] If at least one storage cluster in the second destination region does not contain a destination storage cluster that includes a fourth bucket belonging to the user, the second source storage cluster randomly selects a storage cluster from at least one storage cluster in the second destination region as the second destination storage cluster, or selects the storage cluster with the largest free capacity as the second destination storage cluster, and sends the target data object to the second destination storage cluster. The second destination storage cluster generates a fourth bucket and stores the target data object in the fourth bucket. The generated fourth bucket belongs to the same user as the third bucket on the second source storage cluster.

[0196] In some embodiments, if the target data object is a hot data object in the second source region, it remains in the second source storage cluster after being saved to the second destination storage cluster. Alternatively, if the target data object is a cold data object in the second source region, it is deleted from the second source storage cluster after being saved to the second destination storage cluster.

[0197] In this embodiment, at least one storage cluster in the second source region includes the target data object. If the access volume to the target data object from the second destination region exceeds a third threshold, it indicates that a large number of users in the second destination region are accessing the target data object. The target data object is stored in at least one storage cluster included in the second destination region based on the at least one storage cluster in the second source region. This allows a large number of users in the second destination region to access the target data object from the nearest storage cluster in the second destination region, reducing the latency for users accessing the target data object in the second destination region.

[0198] See Figure 7 This application provides an apparatus 700 for processing data objects, which can be deployed in... Figure 2 or Figure 4 The control system 102 in the illustrated embodiment. The device 700 is used to perform capacity balancing on multiple storage clusters included in the cloud storage system. The multiple storage clusters are distributed in multiple regions, and each region includes at least one storage cluster. The device 700 includes:

[0199] The acquisition unit 701 is used to acquire the first water level of multiple regions. The first water level of a region is used to indicate the storage capacity that has been used in the region. The storage capacity of a region includes the storage capacity of each storage cluster in the region.

[0200] Processing unit 702 is used to determine a first source region and a first destination region based on the first water level of multiple regions, wherein the first water level of the first source region is higher than the first water level of the first destination region;

[0201] The processing unit 702 is also configured to migrate cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region, based on at least one storage cluster in the first source region.

[0202] Optionally, for details on how unit 701 acquires the first water level of multiple regions, please refer to [link to relevant documentation]. Figure 5 The relevant content in step 501 of method 500 shown will not be described in detail here.

[0203] Optionally, for details on how processing unit 702 determines the first source region and the first destination region based on the first water level of multiple regions, please refer to [link to relevant documentation]. Figure 5 The relevant content in step 502 of method 500 shown will not be described in detail here.

[0204] Optionally, for details of the process of processing unit 702 migrating cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region, please refer to [link to detailed implementation process]. Figure 5 The relevant content in step 503 of method 500 shown will not be described in detail here.

[0205] Optionally, the first source region is at least one region with the highest first water level among a plurality of regions, or the first source region is at least one region among a plurality of regions with a first water level higher than a first threshold.

[0206] The first target region is at least one region among multiple regions with the lowest first water level, or the first target region is at least one region among multiple regions with a first water level lower than a second threshold, wherein the second threshold is less than or equal to the first threshold.

[0207] Optionally, the acquisition unit 701 is also used to acquire the access volume of data objects included in multiple regions from other regions;

[0208] Processing unit 702 is further configured to determine a second source region and a second destination region based on access volume, wherein at least one storage cluster in the second source region includes a target data object, and access volume from the second destination region for the target data object exceeds a third threshold.

[0209] The processing unit 702 is also configured to store the target data object into at least one storage cluster included in the second destination region based on at least one storage cluster in the second source region.

[0210] Optionally, for details on how unit 701 retrieves the access volume of data objects from other regions, please refer to [link to relevant documentation]. Figure 6 The relevant content in step 601 of method 600 shown will not be described in detail here.

[0211] Optionally, for details of how processing unit 702 determines the second source region and the second destination region based on access volume, please refer to [link to relevant documentation]. Figure 6 The relevant content in step 602 of method 600 shown will not be described in detail here.

[0212] Optionally, for details of the processing unit 702 storing the target data object in at least one storage cluster included in the second destination region, please refer to [link to relevant documentation]. Figure 6The relevant content in step 603 of method 600 shown will not be described in detail here.

[0213] Optionally, at least one storage cluster in each region includes multiple storage buckets, each storage bucket being used to store at least one data object, and the acquisition unit 701 is used for:

[0214] Get the storage capacity of each storage cluster in the region, as well as the capacity of multiple storage buckets. The capacity of a storage bucket is equal to the sum of the amount of data of each data object stored in the storage bucket.

[0215] The storage capacity used by a region can be obtained based on the capacity of multiple storage buckets, and the total storage capacity of a region can be obtained based on the storage capacity of each storage cluster included in the region.

[0216] The first water level of a region is obtained based on the region's used storage capacity and the region's total storage capacity.

[0217] Optionally, the detailed implementation process of obtaining the storage capacity of each storage cluster in the region, the capacity of multiple storage buckets, the storage capacity already used in the region, the total storage capacity of the region, and the first water level of the region by the acquisition unit 701 is detailed in [link to relevant documentation]. Figure 5 The relevant content in steps 5011-5016 of method 500 shown will not be described in detail here.

[0218] Optionally, the processing unit 702 is used for:

[0219] Obtain the total amount of cold data objects stored in each bucket of at least one storage cluster in the first source region;

[0220] Based on the total amount of cold data objects stored in each bucket, select at least one first bucket with the largest total amount of cold data objects stored, or at least one first bucket with the total amount of cold data objects stored being greater than the fourth threshold.

[0221] Based on at least one storage cluster in the first source region, cold data objects stored in at least one first storage bucket are migrated to at least one second storage bucket included in at least one storage cluster in the first destination region. The first storage bucket and the second storage bucket belong to the same user, and the first storage bucket is any one of at least one first storage bucket, and the second storage bucket is any one of at least one second storage bucket.

[0222] Optionally, for details on how processing unit 702 obtains the total amount of cold data objects stored in each bucket of at least one storage cluster in the first source region, see [link to detailed implementation]. Figure 5 The relevant content in 5031 of method 500 shown will not be described in detail here.

[0223] Optionally, the processing unit 702 selects at least one first storage bucket based on the total amount of cold data objects stored in each storage bucket. For a detailed implementation process, please refer to [link to relevant documentation]. Figure 5 The relevant content in 5032 of method 500 shown will not be described in detail here.

[0224] Optionally, the processing unit 702 migrates cold data objects stored in at least one first storage bucket to at least one second storage bucket included in at least one storage cluster in the first destination region. For detailed implementation, see [link to relevant documentation]. Figure 5 The relevant content in 5033 of method 500 shown will not be explained in detail here.

[0225] Optionally, the processing unit 702 is used for:

[0226] Based on the total amount of cold data objects stored in each bucket, select the buckets with the largest total amount of cold data objects, or the buckets with the total amount of cold data objects stored that is greater than the fourth threshold.

[0227] Get the number of cold data objects stored in multiple buckets;

[0228] Based on the number of cold data objects stored in multiple buckets, select the bucket with the smallest number of stored cold data objects, or at least one first bucket whose number of stored cold data objects is less than a fifth threshold.

[0229] Optionally, for details on how processing unit 702 selects multiple storage buckets, please refer to [link / reference]. Figure 5 The relevant content in 5032 of method 500 shown will not be described in detail here.

[0230] Optionally, for details on how processing unit 702 obtains the number of cold data objects stored in multiple buckets, please refer to [link to relevant documentation]. Figure 5 The relevant content in 5032 of method 500 shown will not be described in detail here.

[0231] Optionally, the processing unit 702 selects at least one first storage bucket from multiple storage buckets based on the number of cold data objects stored in the multiple storage buckets. For details on the implementation process, please refer to [link to relevant documentation]. Figure 5 The relevant content in 5032 of method 500 shown will not be described in detail here.

[0232] Optionally, the processing unit 702 is used for:

[0233] Based on the remaining lifespan of the cold data objects stored in the first storage bucket, select the cold data object with the longest remaining lifespan from the first storage bucket, or at least one cold data object whose remaining lifespan exceeds the sixth threshold.

[0234] Based on at least one storage cluster in the first source region, migrate at least one cold data object to a second storage bucket included in at least one storage cluster in the first destination region.

[0235] Optionally, the processing unit 702 selects at least one cold data object from the first storage bucket based on the remaining lifespan of the cold data objects stored in the first storage bucket. For details on this process, please refer to [link to relevant documentation]. Figure 5 The relevant content in 5032 of method 500 shown will not be described in detail here.

[0236] Optionally, for details of the processing unit 702 migrating at least one cold data object to a second bucket within at least one storage cluster in the first destination region, see [link to detailed implementation]. Figure 5 The relevant content in 5033 of method 500 shown will not be explained in detail here.

[0237] Optionally, the device 700 further includes a transmitting unit 703;

[0238] The sending unit 703 is used to send a migration command to at least one storage cluster in the first source region. The migration command includes identification information of at least one first storage bucket and identification information of the first destination region. The migration command is used to instruct at least one storage cluster in the first source region to migrate cold data objects stored in at least one first storage bucket to at least one second storage bucket included in at least one storage cluster in the first destination region based on the identification information of at least one first storage bucket and the identification information of the first destination region.

[0239] Optionally, for details of the sending unit 703 sending migration commands to at least one storage cluster in the first source region, please refer to [link to relevant documentation]. Figure 5 The relevant content in 5033 of method 500 shown will not be explained in detail here.

[0240] Optionally, the acquisition unit 701 is also used to acquire prediction information for multiple regions; based on the prediction information for multiple regions, to determine the second water level of multiple regions within a first time period, wherein the first time period is after the current time.

[0241] Processing unit 702 is used for:

[0242] Based on the first water level of multiple regions, select m regions with the highest first water level, or with the first water level higher than the first threshold, and select n regions with the lowest first water level, or with the first water level lower than the second threshold, where m and n are both integers greater than 1, and the second threshold is less than or equal to the first threshold.

[0243] Based on the second water level of m regions, select at least one region from the m regions that has the highest second water level, or whose second water level is higher than the first threshold, as the first source region.

[0244] Based on the second water level of n regions, select at least one region from the n regions that has the lowest second water level, or whose second water level is lower than the second threshold, as the first target region.

[0245] Optionally, the acquisition unit 701 acquires prediction information for multiple regions; the detailed implementation process of determining the second water level of multiple regions within the first time period based on the prediction information of multiple regions is described in [reference needed]. Figure 5 The relevant content in methods 5021-5022 of the illustrated method 500 will not be described in detail here.

[0246] Optionally, processing unit 702 selects m regions from multiple regions based on the first water level of multiple regions, and the detailed implementation process for selecting n regions can be found in [link to relevant documentation]. Figure 5 The relevant content in 5023 of method 500 shown will not be explained in detail here.

[0247] Optionally, the detailed implementation process of the processing unit 702 selecting the first target region from the n regions based on the second water level of the n regions can be found in [link to relevant documentation]. Figure 5 The relevant content in 5025 of method 500 shown will not be described in detail here.

[0248] Optionally, for each region, the region's prediction information includes one or more of the following: the increment of storage capacity used by the region in each of at least one second time period, or data growth information, where at least one second time period is prior to the current time, and the data growth information is used to describe the increment of data objects in the region in the first time period compared to at least one second time period.

[0249] In this embodiment, the acquisition unit acquires the first water level of multiple regions, and the processing unit determines the first source region and the first destination region based on the first water level of the multiple regions. The first water level of a region indicates the storage capacity that has been used in the region. The first water level of the first source region is higher than the first water level of the first destination region, so the storage capacity of the first source region is heavily used, while the first destination region still has a large amount of free storage capacity. The processing unit migrates cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region. This releases some of the used storage capacity in the first source region, eliminating the need to expand the first source region and making full use of the free storage capacity in the first destination region, thereby reducing the operating cost of the cloud storage system.

[0250] See Figure 8 This application provides a computing device 800. For example, the computing device 800 may be... Figures 2-4 The device in the control system 102 shown, or the computing device 800, may be... Figure 5 Method 500 or shown Figure 6 The equipment in the control system of method 600 shown.

[0251] like Figure 8 As shown, the computing device 800 includes a bus 802, a processor 804, a memory 806, and a communication interface 808. The processor 804, the memory 806, and the communication interface 808 communicate with each other via the bus 802. The computing device 800 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 800.

[0252] The 802 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 8 The bus 802 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 802 may include a path for transmitting information between various components of the computing device 800 (e.g., processor 804, memory 806, communication interface 808).

[0253] Processor 804 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0254] Memory 806 may include volatile memory, such as random access memory (RAM). Memory 806 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0255] See Figure 8 The memory 806 stores executable program code, and the processor 804 executes the executable program code to implement the following respectively. Figure 7 The functions of the acquisition unit 701, processing unit 702, and sending unit 703 in the illustrated device 700 are used to implement the method provided in any of the above embodiments. That is, the memory 806 stores instructions for executing the method provided in any of the above embodiments. Alternatively,

[0256] The communication interface 808 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 800 and other devices or communication networks.

[0257] This application also provides a cluster for processing data objects. The data migration cluster includes at least one computing device. This computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0258] like Figure 9 As shown, the cluster for processing data objects includes at least one computing device 800. The memory 806 of one or more computing devices 800 in the cluster may store the same instructions for performing the methods provided in any of the above embodiments.

[0259] In some possible implementations, the memory 806 of one or more computing devices 800 in the cluster processing object data may also store partial instructions for executing the methods for processing data objects described above. In other words, a combination of one or more computing devices 800 can jointly execute instructions for performing the methods provided in any of the above embodiments.

[0260] In some possible implementations, one or more computing devices in a cluster that processes data objects can be connected via a network. This network can be a wide area network (WAN), a local area network (LAN), or similar. Figure 10 One possible implementation is shown. For example... Figure 10 As shown, the two computing devices 800A and 800B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.

[0261] In some possible implementations, the memory 806 in the computing device 800A stores the execution of, for example... Figure 7 The instructions for the acquisition unit 701 and processing unit 702 in the illustrated embodiment are shown. Meanwhile, the memory 806 in the computing device 800B stores instructions for performing operations such as... Figure 7 Instructions for the function of the sending unit 703 in the illustrated embodiment.

[0262] It should be understood that Figure 10 The functions of the computing device 800A shown can also be performed by multiple computing devices 800. Similarly, the functions of the computing device 800B can also be performed by multiple computing devices 800.

[0263] This application also provides another type of cluster for processing data objects. The connection relationships between computing devices in a cluster for processing data objects can be similarly referenced. Figure 10 The connection method of the cluster for processing data objects differs in that the memory 806 of one or more computing devices 800 in the cluster for processing data objects can store the same instructions for executing the methods provided in any of the above embodiments.

[0264] In some possible implementations, the memory 806 of one or more computing devices 800 in the cluster processing data objects may also store partial instructions for executing the methods provided in any of the above embodiments. In other words, a combination of one or more computing devices 800 can jointly execute instructions for performing the methods provided in any of the above embodiments.

[0265] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the methods provided in any of the above embodiments.

[0266] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform the method provided in any of the above embodiments.

[0267] All information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are authorized by the user or fully authorized by all parties, and the collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0268] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0269] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A method for processing data objects, characterized in that, The method is applied to a control system for capacity balancing of multiple storage clusters in a cloud storage system, the multiple storage clusters being distributed across multiple regions, each region including at least one storage cluster, the method comprising: Obtain the first water level of multiple regions, the first water level of the regions being used to indicate the storage capacity that the regions have used, the storage capacity of the regions including the storage capacity of each storage cluster in the regions; A first source region and a first destination region are determined based on the first water level of the plurality of regions, wherein the first water level of the first source region is higher than the first water level of the first destination region; Based on at least one storage cluster in the first source region, cold data objects stored in at least one storage cluster in the first source region are migrated to at least one storage cluster in the first destination region.

2. The method as described in claim 1, characterized in that, The first source region is at least one region among the plurality of regions with the highest first water level, or the first source region is at least one region among the plurality of regions with a first water level higher than a first threshold. The first target region is at least one region among the plurality of regions with the lowest first water level, or the first target region is at least one region among the plurality of regions with a first water level lower than a second threshold, wherein the second threshold is less than or equal to the first threshold.

3. The method as described in claim 1 or 2, characterized in that, The method further includes: Obtain the access volume of data objects from other regions included in the multiple regions; Based on the access volume, a second source region and a second destination region are determined, wherein at least one storage cluster in the second source region includes a target data object, and the access volume from the second destination region for the target data object exceeds a third threshold. Based on at least one storage cluster in the second source region, the target data object is stored in at least one storage cluster included in the second destination region.

4. The method according to any one of claims 1-3, characterized in that, At least one storage cluster in each region includes multiple storage buckets, each storage bucket being used to store at least one data object, and the method further includes: Obtain the storage capacity of each storage cluster in the region, as well as the capacity of the multiple storage buckets. The capacity of each storage bucket is equal to the sum of the data amounts of each data object stored in the storage bucket. The storage capacity already used in the region is obtained based on the capacity of the multiple storage buckets, and the total storage capacity of the region is obtained based on the storage capacity of each storage cluster included in the region; The process of obtaining the first water level of multiple regions includes: The first water level of the region is obtained based on the storage capacity already used in the region and the total storage capacity of the region.

5. The method as described in claim 4, characterized in that, The step of migrating cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region, based on at least one storage cluster in the first source region, includes: Obtain the total amount of cold data objects stored in each bucket of at least one storage cluster in the first source region; Based on the total amount of cold data objects stored in each bucket, select at least one first bucket with the largest total amount of cold data objects stored, or at least one first bucket with the total amount of cold data objects stored being greater than the fourth threshold. Based on at least one storage cluster in the first source region, cold data objects stored in the at least one first storage bucket are migrated to at least one second storage bucket included in the at least one storage cluster in the first destination region. The first storage bucket and the second storage bucket belong to the same user, and the first storage bucket is any one of the at least one first storage bucket, and the second storage bucket is any one of the at least one second storage bucket.

6. The method as described in claim 5, characterized in that, The step of selecting at least one first storage bucket based on the total amount of cold data objects stored in each storage bucket, either having the largest total amount of cold data objects stored, or having a total amount of cold data objects stored that is greater than a fourth threshold, includes: Based on the total amount of cold data objects stored in each bucket, select the buckets with the largest total amount of cold data objects, or the buckets with the total amount of cold data objects stored that is greater than the fourth threshold. Obtain the number of cold data objects stored in the multiple storage buckets; Based on the number of cold data objects stored in the plurality of storage buckets, select the first storage bucket that has the smallest number of stored cold data objects, or at least the first storage bucket whose number of stored cold data objects is less than a fifth threshold.

7. The method as described in claim 5 or 6, characterized in that, The migration of cold data objects stored in at least one first storage bucket to at least one second storage bucket within at least one storage cluster in the first destination region, based on at least one storage cluster in the first source region, includes: Based on the remaining lifespan of the cold data objects stored in the first storage bucket, select the cold data object with the longest remaining lifespan from the first storage bucket, or at least one cold data object whose remaining lifespan exceeds the sixth threshold. Based on at least one storage cluster in the first source region, the at least one cold data object is migrated to the second storage bucket included in at least one storage cluster in the first destination region.

8. The method as described in claim 5 or 6, characterized in that, The migration of cold data objects stored in at least one first storage bucket to at least one second storage bucket within at least one storage cluster in the first destination region, based on at least one storage cluster in the first source region, includes: A migration command is sent to at least one storage cluster in the first source region. The migration command includes the identification information of the at least one first storage bucket and the identification information of the first destination region. The migration command is used to instruct at least one storage cluster in the first source region to migrate the cold data objects stored in the at least one first storage bucket to the at least one second storage bucket included in the at least one storage cluster in the first destination region based on the identification information of the at least one first storage bucket and the identification information of the first destination region.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: Obtain prediction information for the multiple regions; Based on the predicted information of the multiple regions, the second water level of the multiple regions is determined within a first time period, which is after the current time. The determination of the first source region and the first destination region based on the first water level of the plurality of regions includes: Based on the first water level of the plurality of regions, select m regions with the highest first water level, or with the first water level higher than a first threshold, and select n regions with the lowest first water level, or with the first water level lower than a second threshold, where m and n are both integers greater than 1, and the second threshold is less than or equal to the first threshold. Based on the second water level of the m regions, select at least one region from the m regions that has the highest second water level, or whose second water level is higher than the first threshold, as the first source region. Based on the second water level of the n regions, at least one region with the lowest second water level or whose second water level is lower than the second threshold is selected as the first target region.

10. The method as described in claim 9, characterized in that, For each region, the prediction information for the region includes one or more of the following: the increment of storage capacity used by the region in each of at least one second time period, or data growth information, wherein the at least one second time period is located before the current time, and the data growth information is used to describe the increment of data objects in the region in the first time period compared to the at least one second time period.

11. An apparatus for processing data objects, characterized in that, The apparatus is used to perform capacity balancing on multiple storage clusters included in a cloud storage system. These multiple storage clusters are distributed across multiple regions, with each region including at least one storage cluster. The apparatus includes: An acquisition unit is used to acquire the first water level of multiple regions, wherein the first water level of a region is used to indicate the storage capacity that the region has been used, and the storage capacity of the region includes the storage capacity of each storage cluster in the region; The processing unit is configured to determine a first source region and a first destination region based on the first water level of the plurality of regions, wherein the first water level of the first source region is higher than the first water level of the first destination region; The processing unit is further configured to migrate cold data objects stored in at least one storage cluster in the first source region to at least one storage cluster in the first destination region, based on at least one storage cluster in the first source region.

12. The apparatus as claimed in claim 11, characterized in that, The first source region is at least one region among the plurality of regions with the highest first water level, or the first source region is at least one region among the plurality of regions with a first water level higher than a first threshold. The first target region is at least one region among the plurality of regions with the lowest first water level, or the first target region is at least one region among the plurality of regions with a first water level lower than a second threshold, wherein the second threshold is less than or equal to the first threshold.

13. The apparatus as claimed in claim 11 or 12, characterized in that, The acquisition unit is also used to acquire the access volume of data objects included in the multiple regions from other regions; The processing unit is further configured to determine a second source region and a second destination region based on the access volume, wherein at least one storage cluster in the second source region includes a target data object, and the access volume from the second destination region for the target data object exceeds a third threshold. The processing unit is further configured to store the target data object into at least one storage cluster included in the second destination region, based on at least one storage cluster in the second source region.

14. The apparatus according to any one of claims 11-13, characterized in that, At least one storage cluster in each region includes multiple storage buckets, each storage bucket being used to store at least one data object, and the acquisition unit is used for: Obtain the storage capacity of each storage cluster in the region, as well as the capacity of the multiple storage buckets. The capacity of each storage bucket is equal to the sum of the data amounts of each data object stored in the storage bucket. The storage capacity already used in the region is obtained based on the capacity of the multiple storage buckets, and the total storage capacity of the region is obtained based on the storage capacity of each storage cluster included in the region; The first water level of the region is obtained based on the storage capacity already used in the region and the total storage capacity of the region.

15. The apparatus as claimed in claim 14, characterized in that, The processing unit is used for: Obtain the total amount of cold data objects stored in each bucket of at least one storage cluster in the first source region; Based on the total amount of cold data objects stored in each bucket, select at least one first bucket with the largest total amount of cold data objects stored, or at least one first bucket with the total amount of cold data objects stored being greater than the fourth threshold. Based on at least one storage cluster in the first source region, cold data objects stored in the at least one first storage bucket are migrated to at least one second storage bucket included in the at least one storage cluster in the first destination region. The first storage bucket and the second storage bucket belong to the same user, and the first storage bucket is any one of the at least one first storage bucket, and the second storage bucket is any one of the at least one second storage bucket.

16. The apparatus as claimed in claim 15, characterized in that, The processing unit is used for: Based on the total amount of cold data objects stored in each bucket, select the buckets with the largest total amount of cold data objects, or the buckets with the total amount of cold data objects stored that is greater than the fourth threshold. Obtain the number of cold data objects stored in the multiple storage buckets; Based on the number of cold data objects stored in the plurality of storage buckets, select the first storage bucket that has the smallest number of stored cold data objects, or at least the first storage bucket whose number of stored cold data objects is less than a fifth threshold.

17. The apparatus as claimed in claim 15 or 16, characterized in that, The processing unit is used for: Based on the remaining lifespan of the cold data objects stored in the first storage bucket, select the cold data object with the longest remaining lifespan from the first storage bucket, or at least one cold data object whose remaining lifespan exceeds the sixth threshold. Based on at least one storage cluster in the first source region, the at least one cold data object is migrated to the second storage bucket included in at least one storage cluster in the first destination region.

18. The apparatus as claimed in claim 15 or 16, characterized in that, The device also includes a transmitting unit; The sending unit is configured to send a migration command to at least one storage cluster in the first source region. The migration command includes the identification information of the at least one first storage bucket and the identification information of the first destination region. The migration command is configured to instruct at least one storage cluster in the first source region to migrate the cold data objects stored in the at least one first storage bucket to the at least one second storage bucket included in the at least one storage cluster in the first destination region based on the identification information of the at least one first storage bucket and the identification information of the first destination region.

19. The apparatus according to any one of claims 11-18, characterized in that, The acquisition unit is further configured to acquire prediction information of the plurality of regions; and based on the prediction information of the plurality of regions, determine the second water level of the plurality of regions within a first time period, wherein the first time period is after the current time. The processing unit is used for: Based on the first water level of the plurality of regions, select m regions with the highest first water level, or with the first water level higher than a first threshold, and select n regions with the lowest first water level, or with the first water level lower than a second threshold, where m and n are both integers greater than 1, and the second threshold is less than or equal to the first threshold. Based on the second water level of the m regions, select at least one region from the m regions that has the highest second water level, or whose second water level is higher than the first threshold, as the first source region. Based on the second water level of the n regions, at least one region with the lowest second water level or whose second water level is lower than the second threshold is selected as the first target region.

20. The apparatus as claimed in claim 19, characterized in that, For each region, the prediction information for the region includes one or more of the following: the increment of storage capacity used by the region in each of at least one second time period, or data growth information, wherein the at least one second time period is located before the current time, and the data growth information is used to describe the increment of data objects in the region in the first time period compared to the at least one second time period.

21. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-10.

22. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1-10.

23. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1-10.