Task scheduling method and device, storage medium and program product

By using an autonomous decision-making approach to manage target nodes, the high cost of manual operation and maintenance and scheduling conflicts in private clouds have been resolved. This has enabled automated data recycling task scheduling under the expansion and contraction of storage clusters, improving the stability and efficiency of cloud storage services.

CN122111580APending Publication Date: 2026-05-29ALIBABA CLOUD COMPUTING CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ALIBABA CLOUD COMPUTING CO LTD
Filing Date
2024-11-29
Publication Date
2026-05-29

Smart Images

  • Figure CN122111580A_ABST
    Figure CN122111580A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a task scheduling method, device, storage medium and program product. In the task scheduling method, upon triggering of a target event, recovery configuration information of at least one available zone in a target region is obtained, a target storage cluster is selected from storage clusters deployed in the at least one available zone according to the recovery configuration information of the at least one available zone and a target cluster selection strategy, and a management and control node deployed in the target storage cluster is taken as a target management and control node. Based on this implementation, in the case that at least one storage cluster in any available zone has a management and control node deployed therein, the target management and control node for scheduling a data recovery task can be autonomously determined according to the recovery configuration information of the at least one available zone, thereby reducing the dependence on operation and maintenance personnel and effectively reducing the operation and maintenance cost of a distributed data recovery system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud network technology, and in particular to a task scheduling method, device, storage medium and program product. Background Technology

[0002] A private cloud, also known as a private cloud, is a cloud computing model where cloud service infrastructure is dedicated to a single customer or organization. Private cloud clusters are typically deployed in the user's own data center or designated location, providing better data security. The infrastructure for private cloud services can be deployed across multiple Availability Zones (AZs) within a Region, with one or more storage clusters deployed in each AZ. Cloud storage services are deployed within the Region, and the storage clusters within the Region can store snapshot data on the cloud storage services. A distributed data reclamation system deployed within the Region then reclaims the snapshot data from the cloud storage services.

[0003] In some private cloud application scenarios, operations and maintenance personnel typically manually configure control nodes in the distributed data reclamation system to schedule data reclamation tasks. This manual operation needs to be performed on-site in the user's private cloud deployment, which is costly. Therefore, a new solution is needed. Summary of the Invention

[0004] This application provides a task scheduling method, device, storage medium, and program product to reduce the operation and maintenance costs of a distributed data recycling system.

[0005] This application provides a task scheduling method for determining a target management node in a distributed data reclamation system. The distributed data reclamation system is used to execute data reclamation tasks corresponding to a target region. The target region includes at least one availability zone. At least one storage cluster in any availability zone is equipped with a management node and worker nodes from the distributed data reclamation system. The management node deployed in any storage cluster is used to schedule worker nodes in different storage clusters within the same availability zone to execute data reclamation tasks. The method includes: responding to a triggering operation of a target event; obtaining reclamation configuration information of at least one availability zone in the target region, wherein the reclamation configuration information of any availability zone is determined based on the configuration information of the management node and / or worker nodes deployed in the availability zone by the distributed data reclamation system; selecting a target storage cluster from the storage clusters deployed in the at least one availability zone based on the configuration information of the at least one availability zone and a target cluster selection strategy; and designating a management node deployed in the target storage cluster as a target management node, wherein the target management node is used to schedule worker nodes in its respective availability zone to execute data reclamation tasks corresponding to the target region.

[0006] Optionally, based on the configuration information of the at least one availability zone and the target cluster selection strategy, a target storage cluster is selected from the storage clusters deployed in the at least one availability zone, including: analyzing the reclamation parallelism of the at least one availability zone based on the reclamation configuration information of the at least one availability zone, wherein the reclamation parallelism of any availability zone is used to describe the number of data reclamation tasks that the management nodes and worker nodes in the availability zone can process in parallel at the same time; selecting the availability zone with a higher reclamation parallelism as the target availability zone based on the analyzed reclamation parallelism of the at least one availability zone; and selecting the target storage cluster from the at least one storage clusters deployed in the target availability zone.

[0007] Optionally, obtaining the reclamation configuration information of at least one availability zone in the target area includes: obtaining the availability zones to which all management nodes in the distributed data reclamation system belong in the target area and the storage clusters in which they reside, as the deployment location information of all management nodes in the distributed data reclamation system; and / or, obtaining the availability zones to which all worker nodes in the distributed data reclamation system belong in the target area, as the deployment location information of all worker nodes in the distributed data reclamation system; using the at least one availability zone as a clustering object, performing clustering based on the deployment location information of all management nodes and / or all worker nodes in the distributed data reclamation system in the target area, to obtain the number of management nodes and / or the number of worker nodes contained in each of the at least one availability zone, as the reclamation configuration information of the at least one availability zone.

[0008] Optionally, based on the reclamation configuration information of the at least one availability zone, the reclamation parallelism of the at least one availability zone is analyzed, including: determining the reclamation parallelism of each of the at least one availability zones based on the number of management nodes and / or the number of worker nodes contained in each of the at least one availability zone.

[0009] Optionally, determining the recycling parallelism of each of the at least one availability zones based on the number of management nodes and / or worker nodes contained in each availability zone includes: determining the recycling parallelism of each of the at least one availability zones based on the number of worker nodes contained in each availability zone, wherein the recycling parallelism of any availability zone is positively correlated with the number of worker nodes contained in the availability zone; or, determining the initial recycling parallelism of each of the at least one availability zones based on the number of worker nodes contained in each availability zone, and correcting the initial recycling parallelism of each of the at least one availability zones based on the number of management nodes contained in each availability zone to obtain the recycling parallelism of each of the at least one availability zones.

[0010] Optionally, selecting the target storage cluster from at least one storage cluster deployed in the target availability zone includes: obtaining version information of each of the at least one storage cluster deployed in the target availability zone; and selecting the storage cluster with the higher version as the target storage cluster based on the version information of each of the at least one storage cluster.

[0011] Optionally, the execution entity of the method is a central node independent of the distributed data recycling system, or any management node in the distributed data recycling system; the management node deployed in the target storage cluster is designated as the target management node, including: if the execution entity of the method is any management node in the distributed data recycling system, the management node determines whether it is located in the target storage cluster; if so, the management node determines itself as the target management node.

[0012] This application embodiment also provides a distributed data reclamation system, which is used to execute data reclamation tasks corresponding to a target area. The target area includes at least one availability zone. At least one storage cluster in any availability zone is equipped with a management node and worker nodes of the distributed data reclamation system. The management node deployed in any storage cluster is used to schedule worker nodes in different storage clusters in the same availability zone. Each management node in the distributed data reclamation system is used to: respond to the triggering operation of a target event; obtain reclamation configuration information of at least one availability zone in the target area, wherein the reclamation configuration information of any availability zone is determined based on the configuration information of the management node and / or worker nodes deployed in the availability zone by the distributed data reclamation system; select a target storage cluster from the storage clusters deployed in the at least one availability zone based on the configuration information of the at least one availability zone and a target cluster selection strategy; and designate the management node deployed in the target storage cluster as the target management node, wherein the target management node is used to schedule worker nodes in its respective availability zone to execute the data reclamation task corresponding to the target area.

[0013] This application also provides an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to perform the steps in the method provided in this application.

[0014] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the method provided in this application.

[0015] This application also provides a computer program product, including: a computer program / instructions, which, when executed by a processor, can implement the steps in the method provided in this application.

[0016] In this embodiment, upon triggering a target event, reclamation configuration information for at least one availability zone in the target area can be obtained. Based on this reclamation configuration information and the target cluster selection strategy, a target storage cluster is selected from the storage clusters deployed in the at least one availability zone, and the management node deployed in the target storage cluster is designated as the target management node. Based on this implementation, when a management node is deployed in at least one storage cluster in any availability zone, the target management node for scheduling data reclamation tasks can be autonomously determined based on the reclamation configuration information of the at least one availability zone. This reduces the risk of data reclamation task conflicts and, compared to manually selecting the target management node, reduces reliance on maintenance personnel, effectively lowering the operation and maintenance costs of the distributed data reclamation system. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0018] Figure 1 A schematic diagram of the deployment architecture of the target area provided in an exemplary embodiment of this application;

[0019] Figure 2 A schematic diagram of the deployment architecture of a distributed data recycling system provided in an exemplary embodiment of this application;

[0020] Figure 3 A flowchart illustrating a task scheduling method provided in an exemplary embodiment of this application;

[0021] Figure 4 A schematic diagram illustrating the deployment of a control node in different versions of a storage cluster within a single site, provided as an exemplary embodiment of this application;

[0022] Figure 5 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” used in the embodiments of this invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. “Multiple” generally includes at least two, but does not exclude the inclusion of at least one.

[0025] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0026] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system that includes said element.

[0027] In cloud computing technology, a region is a geographically independent location. Each region contains one or more physical data centers. Different regions can provide geographical redundancy for data, enabling disaster recovery and data backup across regions. Within a region, to further improve system availability and fault tolerance, multiple Availability Zones (AZs) are created. Each AZ is a collection of one or more physical data centers with independent power supply, cooling systems, and network connectivity. Even if one AZ in a region fails, other AZs can continue to operate normally, thus ensuring service continuity and stability.

[0028] An Availability Zone (AZ) may include one or more storage clusters (e.g., block storage clusters). Each storage cluster refers to a distributed system composed of multiple storage nodes, used to provide highly reliable and high-performance data storage services. New storage nodes can be added to an AZ as storage demand increases to increase storage capacity and performance. Snapshot data of the storage clusters can be stored on cloud storage services within the zone, such as Object Storage Service (OSS). A distributed data reclamation system can be deployed within the zone to reclaim snapshot data from the cloud storage service. This distributed data reclamation system may include a master node and worker nodes. The master node initiates and schedules snapshot data reclamation tasks, while the worker nodes reclaim snapshot data from the cloud storage service according to the master node's scheduling instructions.

[0029] In one deployment mode of a distributed data reclamation system, worker nodes are deployed at the Availability Zone (AZ) level, while control nodes are deployed at the storage cluster level. A storage cluster can have one control node or multiple control nodes supporting primary / standby failover. A control node deployed in any storage cluster within an AZ can schedule worker nodes deployed in different storage clusters within that AZ. The distributed data reclamation system operates at the region level; that is, only one control node in a storage cluster within a region can be active at any given time. A control node deployed in a specific storage cluster within an AZ can schedule worker nodes deployed in different storage clusters within that AZ to execute the data reclamation tasks corresponding to that region.

[0030] In a private cloud deployment scenario, multiple storage clusters within an Availability Zone (AZ) deploy management nodes for a distributed data reclamation system. For example, when different storage clusters within the same AZ use different versions of storage architecture, their workflows differ, making it impossible to uniformly deploy management nodes for the distributed data reclamation system across different clusters. This results in multiple storage clusters with different architectures having management nodes deployed for the distributed data reclamation system. Similarly, when storage clusters are being removed or expanded, it cannot be guaranteed that only one storage cluster in a region has a management node deployed. When performing snapshot data reclamation tasks, management nodes deployed in multiple storage clusters within the same AZ can concurrently schedule worker nodes within that AZ to execute the corresponding data reclamation tasks for that region. On one hand, this concurrent scheduling of multiple management nodes can lead to scheduling conflicts for worker nodes within the AZ. These conflicts can prevent snapshot data in the cloud storage service from being reclaimed in a timely manner, resulting in excessively high data levels and impacting normal service. On the other hand, data reclamation is at the region level. If multiple management nodes in a region are performing data reclamation scheduling tasks, it can lead to risks of data inconsistency or duplicate processing of the data to be reclaimed within that region.

[0031] In a traditional approach, operations personnel typically manually select a storage cluster from an Availability Zone (AZ) within a region and deploy a scheduled task on the control node of that cluster. This task triggers the control node to schedule worker nodes within its AZ to execute data reclamation tasks for the entire region. Apart from the control node with the scheduled task deployed, other control nodes in the region do not initiate data reclamation tasks. This new approach, based on manual deployment by operations personnel, allows a single control node in a storage cluster to schedule worker nodes within its AZ to execute data reclamation tasks for the entire region, mitigating the reclamation conflicts caused by multiple control nodes concurrently launching data reclamation tasks.

[0032] However, in private cloud application scenarios, manual operation and maintenance (O&M) requires on-site operation within the user's private cloud deployment, which is costly. Furthermore, this manual O&M cannot flexibly and efficiently handle scenarios involving storage cluster expansion and contraction. For example, in a storage cluster expansion scenario, a new storage cluster is added to a single Availability Zone (AZ). If this new storage cluster has a management node deployed, and that management node can initiate data reclamation tasks, multiple management nodes within an AZ will still concurrently initiate data reclamation tasks. In a storage cluster contraction scenario, if the storage cluster containing the management node capable of initiating data reclamation tasks within an AZ is deleted, the worker nodes within that AZ will be unable to perform data reclamation operations, leading to excessively high cloud storage service levels and impacting normal service. Therefore, if manual O&M is not performed promptly after each storage cluster expansion or contraction, the cloud storage service will malfunction.

[0033] To address the aforementioned technical problems, a solution is provided in some embodiments of this application. The technical solutions provided by each embodiment of this application are described in detail below with reference to the accompanying drawings.

[0034] Figure 1 This is a schematic diagram of the structure of a distributed data recycling system 10 provided in an exemplary embodiment of this application, as shown below. Figure 1 As shown, the distributed data recycling system 10 can be deployed in a target area 11, which includes at least one availability zone (e.g., Figure 1 The first availability zone 12 and the second availability zone 13 are shown, along with a cloud storage service 14. Each availability zone may deploy one or more storage clusters, which can be block storage clusters, file storage clusters, or object storage clusters; this embodiment does not impose any restrictions. The cloud storage service 14 is used to store snapshot data generated by the storage clusters in the at least one availability zone, and the distributed data reclamation system 10 is used to reclaim snapshot data from the cloud storage service 14.

[0035] In this system, at least one storage cluster in any availability zone has a control node deployed in the distributed data reclamation system 10, such as... Figure 1 As shown, in the first availability zone 12, a first management node 101 and a second management node 102 can be deployed in the first storage cluster 121, and a third management node 103 can be deployed in the second storage cluster 122. In the second availability zone 13, a fourth management node 104 can be deployed in the third storage cluster 131, and a fifth management node 105 can be deployed in the fourth storage cluster 132. The management node deployed in any storage cluster can be deployed on a storage node within that storage cluster, or it can be deployed on an independent computer node within the storage cluster; this embodiment does not impose any restrictions.

[0036] In this embodiment, one or more worker nodes can be deployed in any availability zone, and these worker nodes can be distributed across different storage clusters within the availability zone. Worker nodes deployed in any storage cluster can be deployed on storage nodes within that cluster, or on independent computer nodes within the storage cluster; this embodiment does not impose any limitations. For example, as... Figure 1 As shown, in the first availability zone 12, a first worker node 201 and a second worker node 202 can be deployed in the first storage cluster 121, and a third worker node 203 and a fourth worker node 204 can be deployed in the second storage cluster 122. In the second availability zone 13, a fifth worker node 205 can be deployed in the third storage cluster 131, and a sixth worker node 206 and a seventh worker node 207 can be deployed in the fourth storage cluster 132. Figure 1 The illustrated control nodes and working nodes can form a structure like this: Figure 2 The distributed data recycling system 10 shown is illustrated.

[0037] In the distributed data reclamation system 10, a management node deployed in any storage cluster is used to schedule worker nodes in different storage clusters within the same availability zone. For example, Figure 2 As shown, the first management node 101, the second management node 102, and the third management node 103 in the first availability zone 12 can schedule all worker nodes located in the first storage cluster 121 and the second storage cluster 122 in the first availability zone 12, namely, the first worker node 201, the second worker node 202, the third worker node 203, and the fourth worker node 204. The fourth management node 104 and the fifth management node 105 in the second availability zone 13 can schedule all worker nodes located in the third storage cluster 131 and the fourth storage cluster 132 in the second availability zone 13, namely, the fifth worker node 205, the sixth worker node 206, and the seventh worker node 207.

[0038] exist Figure 2 In the distributed data reclamation system 10 shown, any control node can be used to execute a task scheduling method to determine the target control node in the distributed data reclamation system 10. This target control node is used to schedule worker nodes deployed in different storage clusters within its respective Availability Zone (AZ) to execute the data reclamation task corresponding to the target area 11. The task scheduling method executed by any control node will be illustrated below with reference to the accompanying drawings.

[0039] Figure 3 This is a flowchart illustrating a task scheduling method provided in an exemplary embodiment of this application. The method may include, for example: Figure 3 The steps shown are as follows:

[0040] Step 301: Respond to the triggering operation of the target event and obtain the recycling configuration information of at least one availability zone in the target area. The recycling configuration information of any availability zone is determined based on the configuration information of the management node and / or worker node deployed in the availability zone by the distributed data recycling system.

[0041] Step 302: Based on the configuration information of the at least one availability zone and the target cluster selection strategy, select the target storage cluster from the storage clusters deployed in the at least one availability zone.

[0042] Step 303: The management node deployed in the target storage cluster is designated as the target management node. The target management node is used to schedule the worker nodes in its availability zone to execute the data reclamation task corresponding to the target area.

[0043] In step 301, the target event can be an event that detects an external input operation, a system state change event, or a timer expiration event, etc., and this embodiment does not limit this. The external input operation can be caused by user interaction or interaction with an external system, such as a user clicking a specified button or an external system initiating a call through a specified interface. The system state change event is triggered when the internal state of the system changes; for example, a system state change event could be an event where the amount of data to be recycled in target area 11 exceeds a set threshold. The timer expiration event is triggered when a preset time point is reached or when a time interval ends. For example, the execution cycle of the data recycling task can be set in target area 11, and the timer expiration event is triggered when the execution cycle arrives.

[0044] The reclamation configuration information of any availability zone in the target area refers to all relevant configuration information in that availability zone to enable the data reclamation task to be executed efficiently and reliably. The reclamation configuration information of any availability zone is determined based on the configuration information of the management nodes and / or worker nodes deployed in that availability zone by the distributed data reclamation system 10.

[0045] The configuration information of any node in the distributed data recycling system 10, such as a control node or a worker node, may include, but is not limited to, at least one of the following: deployment location information, specification information, and primary / backup deployment information. Deployment location information refers to the availability zone and / or storage cluster to which the node belongs in the target area 11. Specifications refer to hardware and software configuration parameters related to the node's performance and processing capabilities, such as the processor type and speed, memory size, storage capacity, and operating system version of the server hosting the node. Primary / backup deployment information refers to the configuration information related to the node's system high availability and fault tolerance. In primary / backup deployment mode, a backup node can be configured for the primary node. This backup node is used to take over the functions of the primary node in the event of a failure, ensuring service continuity and reliability.

[0046] In some optional embodiments, the configuration information of any control node can be stored on the operation and maintenance management platform corresponding to the storage cluster where the control node resides. When obtaining the configuration information of the control node, it can be queried through the open interface provided by the operation and maintenance management platform. Similarly, the configuration information of any worker node can be stored on the operation and maintenance management platform corresponding to the storage cluster where the worker node resides. When obtaining the configuration information of the worker node, it can be queried through the open interface provided by the operation and maintenance management platform.

[0047] After obtaining the reclamation configuration information of at least one availability zone based on the above implementation method, in step 302, a target storage cluster can be selected from the storage clusters deployed in at least one availability zone according to the reclamation configuration information of at least one availability zone and the target cluster selection strategy. The target cluster selection strategy can be preset in the execution subject of this embodiment, or it can be dynamically issued to the execution subject; this embodiment does not impose any restrictions. In some embodiments, when multiple management nodes in the target area 11 respectively execute the method provided in this embodiment to determine whether they are target management nodes, the multiple management nodes can execute the same target cluster selection strategy to ensure the consistency and reliability of the selection results. Here, "any management node itself" refers to the entity of that management node and not other management nodes besides that management node.

[0048] Optionally, the target cluster selection strategy may be to first select a target availability zone that meets the set conditions from at least one availability zone in the target region 11, and then select the target storage cluster from the target availability zone. The following will provide an example description.

[0049] In some optional embodiments, when selecting a target availability zone from at least one availability zone in target region 11, the management nodes included in the at least one availability zone can be determined based on the deployment location information of the management nodes in the distributed data reclamation system 10, and the processing capability score of the management nodes included in the at least one availability zone can be determined based on the specification information of the management nodes in the distributed data reclamation system 10. The processing capability score of any management node can be determined based on the specification information of the management node, such as processor type, memory size, storage capacity, and operating system version. Each specification information can correspond to a processing capability score; for example, the higher the number of processor cores, the higher the processing capability score; the larger the memory, the higher the processing capability score, and so on. For any management node, the processing capability scores corresponding to the various specification information of the management node can be weighted and summed according to a set weight to obtain the processing capability score of the management node. Based on the processing capability scores of the management nodes included in each of the at least one availability zone, a target availability zone can be selected from the at least one availability zone. For example, for any availability zone, the average processing capacity score of that availability zone can be calculated based on the processing capacity scores of the control nodes in that availability zone, and the availability zone with the higher average processing capacity score can be selected as the target availability zone from among the at least one availability zones. Alternatively, the availability zone containing the control node with the highest processing capacity score can be selected as the target availability zone from among the at least one availability zones.

[0050] In some alternative embodiments, when selecting a target availability zone from at least one availability zone in target region 11, the recycling parallelism of the at least one availability zone can be analyzed based on its recycling configuration information. Then, based on the analyzed recycling parallelism, the availability zone with the higher recycling parallelism is selected as the target availability zone. Here, recycling parallelism describes the number of data recycling tasks that the management nodes and worker nodes in the availability zone can process in parallel within the same timeframe. Recycling parallelism is associated with the number and specifications of worker nodes configured in the availability zone for parallel execution of data recycling tasks.

[0051] Based on this, in some optional embodiments, when obtaining the configuration information of the control nodes in the distributed data reclamation system 10, the availability zones to which all control nodes in the distributed data reclamation system 10 belong in the target area 11 and the storage clusters located in their respective availability zones can be obtained as the deployment location information of all control nodes in the distributed data reclamation system 10. Optionally, when obtaining the configuration information of all worker nodes in the distributed data reclamation system 10 in the target availability zone, the availability zones to which the worker nodes in the distributed data reclamation system 10 belong in the target area 11 can be obtained as the deployment location information of all worker nodes in the distributed data reclamation system 10. Using the at least one availability zone as a clustering object, clustering is performed according to the deployment location information of all control nodes and / or all worker nodes in the distributed data reclamation system 10 in the target area to obtain the number of control nodes and / or the number of worker nodes contained in each of the at least one availability zones, which serves as the reclamation configuration information of the at least one availability zone.

[0052] For example, with Figure 1 The first availability zone 12 and the second availability zone 13 shown are clustering objects. The deployment location information of all management nodes and all worker nodes in the distributed data reclamation system 10 is clustered, and the reclamation configuration information of the first availability zone is as follows: the first availability zone 12 has three management nodes and four worker nodes of the distributed data reclamation system 10. The reclamation configuration information of the second availability zone 13 is as follows: the second availability zone 13 has two management nodes and three worker nodes of the distributed data reclamation system 10.

[0053] In some optional embodiments, when analyzing the reclamation parallelism of the at least one availability zone based on its reclamation configuration information, the reclamation parallelism of each of the at least one availability zone can be determined according to the number of management nodes and / or worker nodes contained in each of the at least one availability zone. Specifically, the reclamation parallelism of any availability zone is positively correlated with the number of worker nodes it contains. That is, the more worker nodes an availability zone contains, the higher its reclamation parallelism. For example, based on this positive correlation, the reclamation parallelism of the first availability zone 12 is higher than that of the second availability zone 13. This approach facilitates selecting availability zones with higher parallelism to execute the data reclamation task corresponding to the target area 11, thereby improving the execution efficiency of the data reclamation task.

[0054] Besides being related to the number of worker nodes, the data reclamation parallelism of an availability zone can also be associated with its redundancy configuration. In the same availability zone, the more standby nodes configured for the management node, the higher the fault tolerance and the shorter the fault recovery time when the management node fails. Based on this, in some optional embodiments, after determining the initial reclamation parallelism of each of the at least one availability zone based on the number of worker nodes it contains, the initial reclamation parallelism can be further modified based on the number of management nodes it contains to obtain the final reclamation parallelism for each of the at least one availability zone. This approach facilitates selecting availability zones with higher parallelism and stability to execute the data reclamation task corresponding to target region 11, thereby ensuring the execution efficiency and stability of the data reclamation task.

[0055] After determining the target availability zone from at least one availability zone in target region 11 based on the above implementation method, a target storage cluster can be selected from at least one storage cluster deployed in the target availability zone. In some optional embodiments, version information of each of the at least one storage cluster deployed in the target availability zone can be obtained, and a storage cluster with a higher version can be selected as the target storage cluster based on the version information of each of the at least one storage cluster. Optionally, the version information of any storage cluster can be obtained through an open interface provided by the storage cluster's operation and maintenance management platform. In some embodiments, the version information of the storage cluster can be included in the name of the storage cluster. Based on this, the name of the storage cluster can be parsed to obtain the version information of the storage cluster. For example, keywords / words related to the version number in the name of the storage cluster can be parsed, and the version number can be determined based on the keywords / words related to the version number. As another example, the correspondence between the cluster name and the version number can be queried based on the name of the storage cluster, and the version number of the storage cluster can be determined based on the query result, which will not be elaborated further.

[0056] Optionally, when the version information of the at least one storage cluster is the same, a storage cluster can be randomly selected from the at least one storage cluster as the target storage cluster. In some other optional embodiments, the historical failure rate of each of the at least one storage cluster deployed in the target availability zone can be obtained, and the storage cluster with the lower historical failure rate can be selected from the at least one storage cluster as the target storage cluster. The historical failure rate of any storage cluster can be obtained by analyzing the log data of that storage cluster over a historical period, which will not be elaborated here.

[0057] After determining the target storage cluster based on the above implementation method, in step 303, the management node deployed in the target storage cluster can be used as the target management node for initiating the data reclamation task corresponding to the target area 11. When the executing entity of this method is any management node in the distributed data reclamation system 10, the management node can determine whether it is located in the target storage cluster. If so, the management node can determine itself as the target management node. The target management node can schedule the worker nodes deployed in the target availability zone to execute the data reclamation task corresponding to the target area 11.

[0058] In this embodiment, upon triggering a target event, reclamation configuration information for at least one availability zone in the target area can be obtained. Based on this reclamation configuration information and the target cluster selection strategy, a target storage cluster is selected from the storage clusters deployed in the at least one availability zone, and the management node deployed in the target storage cluster is designated as the target management node. Based on this implementation, when a management node is deployed in at least one storage cluster in any availability zone, the target management node for scheduling data reclamation tasks can be autonomously determined based on the reclamation configuration information of the at least one availability zone. This reduces the risk of data reclamation task conflicts and, compared to manually selecting the target management node, reduces reliance on maintenance personnel, effectively lowering the operation and maintenance costs of the distributed data reclamation system.

[0059] For any management node in the distributed data reclamation system 10, the method provided in the above embodiments can be executed to determine the target management node. For different management nodes in the distributed data reclamation system 10, the configuration information of all management nodes and / or all worker nodes in the distributed data reclamation system obtained by different management nodes is consistent, and different management nodes can adopt the same target cluster selection strategy as described above. Therefore, the target storage clusters elected by different management nodes are also consistent. Based on this, any management node can determine whether it is deployed in the elected target storage cluster. If the management node determines that it is deployed in the elected target storage cluster, it can directly designate itself as the target management node and schedule worker nodes in the target availability zone to perform data reclamation tasks. Conversely, if any management node determines that it is not deployed in the elected target storage cluster, it can remain silent and exit the decision-making process. In the primary / standby deployment mode, if the management node determines that it is deployed in the elected target storage cluster, it can further determine whether it is configured as the primary management node in the target storage cluster. If so, it can directly schedule the worker nodes in the target availability zone to perform data reclamation tasks. Otherwise, it can remain silent and exit the decision-making process.

[0060] In some application scenarios, if at least one availability zone in target area 11 contains a storage cluster with a management node, then based on the method provided in this embodiment, the management node in that storage cluster can automatically trigger data reclamation tasks without manually deploying a scheduled task to trigger the data reclamation task. In other application scenarios, if at least one availability zone in target area 11 contains multiple storage clusters with management nodes, then based on the method provided in this embodiment, the management nodes in the multiple storage clusters can autonomously elect a target management node, also without manually deploying a scheduled task to trigger the data reclamation task.

[0061] Furthermore, when the storage cluster is expanded or reduced in size, the target cluster selection strategy can be executed based on the configuration information of the control node and / or worker nodes after the expansion or reduction. Thus, even if the information of the storage cluster changes, it can still autonomously and accurately determine a consistent target control node.

[0062] It is worth noting that in some other alternative embodiments, Figure 3 The execution entity of the task scheduling method shown can be a central node independent of the distributed data recycling system 10. This central node can be deployed at the storage cluster level, or at the availability zone level, or at the region level; this embodiment does not impose any restrictions. The central node can execute... Figure 3 The corresponding task scheduling method identifies the target control node and sends a start command to it, causing the target control node to initiate the data reclamation task. In this implementation, having the central node execute the task scheduling method helps ensure the consistency of the determined target control node.

[0063] Figure 4 The application scenarios of private clouds further illustrate the scenario of deploying multiple control nodes in the same Availability Zone (AZ), such as... Figure 4As shown, the target areas of the private cloud include site A and site B, each corresponding to an Availability Zone (AZ). Site A includes two storage clusters using a V1 architecture and one using a V2 architecture. Both the V1 and V2 storage clusters have management nodes deployed. In site A, both management nodes may initiate data reclamation tasks, potentially leading to conflicting reclamation tasks. Based on the method provided in this application, any management node can obtain the site to which the management nodes and worker nodes in the target area belong, as well as the storage cluster within that site, thus obtaining configuration information. Then, AZs can be used as clustering objects to cluster the configuration information, resulting in an AZ list. The number of storage clusters using V1 and V2 architectures under each AZ, as well as the number of management and worker nodes in each AZ, can be counted. In the clustering results, AZs can be sorted by the string in their names, and the storage clusters within each AZ can also be sorted by the string in their cluster names, forming a cluster list.

[0064] Based on the clustering results, the AZ list can be traversed to select the target AZ. Specifically, since the control nodes are scheduled on an AZ-by-AZ basis, to ensure high parallelism in data reclamation when scheduling worker nodes to perform data reclamation tasks, the AZ with the highest number of worker nodes can be selected as the target AZ when traversing the AZ list. Since the V2 architecture is superior to the V1 architecture, and higher-version architectures typically have higher system stability, after determining the target AZ, a storage cluster using the V2 architecture within the target AZ can be selected as the target storage cluster. If no storage cluster using the V2 architecture exists in the target AZ, the target storage cluster can be selected according to its order of arrangement within the target AZ, or randomly selected. After executing the above method, any control node can determine whether its own storage cluster is a selected target storage cluster; if so, it initiates the data reclamation task. Furthermore, in the case of deploying multiple control nodes in a single site in a private cloud, the decision-making method provided in this embodiment selects the control node for scheduling data reclamation tasks, reducing the risk of data reclamation task conflicts in the case of multiple control nodes at a lower cost.

[0065] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 301 to 304 can be device A; or the execution subject of steps 301 and 302 can be device A, and the execution subject of step 303 can be device B; and so on.

[0066] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 301, 302, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0067] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0068] Figure 5 This illustration shows a structural diagram of an electronic device provided in an exemplary embodiment of this application. The electronic device can be used to deploy a control node in a distributed data reclamation system, or it can be used to deploy a central node independent of the distributed data reclamation system. The control node or central node deployed on the electronic device is used to execute a task scheduling method to determine the target control node in the distributed data reclamation system for scheduling data reclamation tasks. The distributed data reclamation system is deployed in a target area, which includes at least one availability zone. At least one storage cluster in any availability zone deploys the control node and worker nodes of the distributed data reclamation system. The control node deployed in any storage cluster is used to schedule worker nodes in different storage clusters within the same availability zone.

[0069] like Figure 5 As shown, the electronic device includes: a memory 501, a processor 502, and a communication component 503.

[0070] Memory 501 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device.

[0071] Processor 502, coupled to memory 501, is used to execute computer programs in memory 501 for: responding to a triggering operation of a target event; obtaining reclamation configuration information of at least one availability zone in the target area, wherein the reclamation configuration information of any availability zone is determined based on the configuration information of management nodes and / or worker nodes deployed in the availability zone by the distributed data reclamation system; selecting a target storage cluster from the storage clusters deployed in the at least one availability zone based on the configuration information of the at least one availability zone and a target cluster selection strategy; and designating a management node deployed in the target storage cluster as a target management node, wherein the target management node is used to schedule worker nodes in its respective availability zone to execute the data reclamation task corresponding to the target area.

[0072] Optionally, when the processor 502 selects a target storage cluster from the storage clusters deployed in the at least one availability zone based on the configuration information of the at least one availability zone and the target cluster selection strategy, it specifically performs the following: Analyzes the recycling parallelism of the at least one availability zone based on the recycling configuration information of the at least one availability zone, where the recycling parallelism of any availability zone describes the number of data recycling tasks that the management nodes and worker nodes in the availability zone can process in parallel at the same time; selects the availability zone with a higher recycling parallelism as the target availability zone based on the analyzed recycling parallelism of the at least one availability zone; and selects the target storage cluster from the at least one storage cluster deployed in the target availability zone.

[0073] Optionally, when the processor 502 obtains the reclamation configuration information of at least one availability zone in the target area, it is specifically used to: obtain the availability zones to which all the control nodes in the distributed data reclamation system belong in the target area and the storage clusters in which they are located, as the deployment location information of all the control nodes in the distributed data reclamation system; and / or, obtain the availability zones to which all the worker nodes in the distributed data reclamation system belong in the target area, as the deployment location information of all the worker nodes in the distributed data reclamation system; using the at least one availability zone as a clustering object, clustering is performed according to the deployment location information of all the control nodes and / or all the worker nodes in the distributed data reclamation system in the target area to obtain the number of control nodes and / or the number of worker nodes contained in each of the at least one availability zone, as the reclamation configuration information of the at least one availability zone.

[0074] Optionally, when the processor 502 analyzes the recycling parallelism of the at least one availability zone based on the recycling configuration information of the at least one availability zone, it is specifically used to: determine the recycling parallelism of each of the at least one availability zone based on the number of management nodes and / or the number of worker nodes contained in each of the at least one availability zone.

[0075] Optionally, when determining the reclamation parallelism of the at least one availability zone based on the number of management nodes and / or worker nodes contained in each of the at least one availability zones, the processor 502 specifically performs the following: determining the reclamation parallelism of the at least one availability zone based on the number of worker nodes contained in each of the at least one availability zone, wherein the reclamation parallelism of any availability zone is positively correlated with the number of worker nodes contained in the availability zone; or, determining the initial reclamation parallelism of the at least one availability zone based on the number of worker nodes contained in each of the at least one availability zone, and correcting the initial reclamation parallelism of the at least one availability zone based on the number of management nodes contained in each of the at least one availability zone to obtain the reclamation parallelism of the at least one availability zone.

[0076] Optionally, when the processor 502 selects the target storage cluster from at least one storage cluster deployed in the target availability zone, it is specifically configured to: obtain version information of each of the at least one storage cluster deployed in the target availability zone; and select the storage cluster with the higher version as the target storage cluster based on the version information of each of the at least one storage cluster.

[0077] Optionally, the execution subject of the method is a central node independent of the distributed data recycling system, or any management node in the distributed data recycling system; when the processor 502 uses the management node deployed in the target storage cluster as the target management node, it is specifically configured to: if the execution subject of the method is any management node in the distributed data recycling system, then the management node determines whether it is located in the target storage cluster; if so, then it determines itself as the target management node.

[0078] Furthermore, such as Figure 5 As shown, the electronic device also includes other components such as a power supply component 504, a display component 505, and an audio component 506. Figure 5 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 5 The components shown. Figure 5In this embodiment, the components within the dashed boxes are optional, not mandatory, and their specific requirements depend on the product form of the electronic device. The electronic device in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device, or a server-side device such as a conventional server, cloud server, or server array. If the electronic device in this embodiment is a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 5 The components within the dashed box; if the electronic device in this embodiment is implemented as a conventional server, cloud server, or server array, etc., it may be omitted. Figure 5 The component within the dashed box.

[0079] The memory 501 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0080] The communication component 503 is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as Wi-Fi, 2G (e.g., Global System for Mobile Communications (GSM)), 3G (e.g., Wideband Code Division Multiple Access (WCDMA), 4G (e.g., Long Term Evolution (LTE)), 4G+ (e.g., LTE-Advanced (LTE-A)), or 5G (5th Generation Mobile Communication Technology), or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component may be implemented based on Near Field Communication (NFC), Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.

[0081] The power supply component 504 is used to provide power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.

[0082] The display component includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation.

[0083] An audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals may be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0084] In this embodiment, upon triggering a target event, the reclamation configuration information of at least one availability zone in the target region is obtained. Based on the reclamation configuration information of the at least one availability zone and the target cluster selection strategy, a target storage cluster is selected from the storage clusters deployed in the at least one availability zone, and the management node deployed in the target storage cluster is designated as the target management node. Based on this implementation, when a management node is deployed in at least one storage cluster in any availability zone, the target management node for scheduling data reclamation tasks can be autonomously determined based on the reclamation configuration information of the at least one availability zone. This reduces the risk of data reclamation task conflicts and, compared to manually selecting the target management node, reduces reliance on maintenance personnel, effectively lowering the operation and maintenance costs of the distributed data reclamation system.

[0085] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be performed by an electronic device in the above method embodiments.

[0086] This application also provides a computer program product, including: a computer program / instructions, which, when executed by a processor, can implement the steps in the method provided in this application.

[0087] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM (Compact Disc Read-Only Memory), optical storage, etc.) containing computer-usable program code.

[0088] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0089] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0090] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0091] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.

[0092] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0093] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0094] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes said element.

[0095] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A task scheduling method for determining a target control node in a distributed data recycling system, wherein the distributed data recycling system is used to execute data recycling tasks corresponding to a target area, characterized in that, The target area includes at least one availability zone, and at least one storage cluster in any availability zone deploys a control node and worker nodes in the distributed data reclamation system. The control node deployed in any storage cluster is used to schedule worker nodes in different storage clusters within the same availability zone to perform data reclamation tasks; the method includes: In response to the triggering operation of the target event, the recycling configuration information of at least one availability zone in the target area is obtained. The recycling configuration information of any availability zone is determined based on the configuration information of the management node and / or worker node deployed in the availability zone by the distributed data recycling system. Based on the configuration information of the at least one availability zone and the target cluster selection strategy, a target storage cluster is selected from the storage clusters deployed in the at least one availability zone. The control node deployed in the target storage cluster is designated as the target control node. The target control node is used to schedule the worker nodes in its respective availability zone to perform the data reclamation task corresponding to the target zone.

2. The method according to claim 1, characterized in that, Based on the configuration information of the at least one availability zone and the target cluster selection strategy, a target storage cluster is selected from the storage clusters deployed in the at least one availability zone, including: Based on the reclamation configuration information of the at least one availability zone, analyze the reclamation parallelism of the at least one availability zone. The reclamation parallelism of any availability zone is used to describe the number of data reclamation tasks that the management nodes and worker nodes in the availability zone can process in parallel at the same time. Based on the recovery parallelism of the at least one availability zone obtained from the analysis, select the availability zone with higher recovery parallelism as the target availability zone from the at least one availability zone; Select the target storage cluster from at least one storage cluster deployed in the target availability zone.

3. The method according to claim 1, characterized in that, Obtain the reclamation configuration information for at least one availability zone in the target region, including: Obtain the availability zones to which all management nodes in the distributed data recycling system belong in the target area, as well as the storage clusters within those availability zones, as the deployment location information of all management nodes in the distributed data recycling system; and / or, obtain the availability zones to which all worker nodes in the distributed data recycling system belong in the target area, as the deployment location information of all worker nodes in the distributed data recycling system. Using the at least one availability zone as the clustering object, clustering is performed based on the deployment location information of all management nodes and / or all worker nodes in the distributed data reclamation system in the target area to obtain the number of management nodes and / or worker nodes contained in each of the at least one availability zone, which serves as the reclamation configuration information for the at least one availability zone.

4. The method according to claim 2, characterized in that, Based on the reclamation configuration information of the at least one availability zone, analyze the reclamation parallelism corresponding to the at least one availability zone, including: The reclamation parallelism of each of the at least one availability zones is determined based on the number of management nodes and / or worker nodes contained in each availability zone.

5. The method according to claim 4, characterized in that, Based on the number of management nodes and / or worker nodes contained in each of the at least one availability zone, determine the reclamation parallelism of each of the at least one availability zone, including: The reclamation parallelism of each of the at least one availability zones is determined based on the number of worker nodes contained in that availability zone, wherein the reclamation parallelism of any availability zone is positively correlated with the number of worker nodes contained in that availability zone; or... Based on the number of working nodes contained in each of the at least one availability zones, the initial reclamation parallelism of each of the at least one availability zones is determined, and based on the number of management nodes contained in each of the at least one availability zones, the initial reclamation parallelism of each of the at least one availability zones is corrected to obtain the reclamation parallelism of each of the at least one availability zones.

6. The method according to any one of claims 2-5, characterized in that, Selecting the target storage cluster from at least one storage cluster deployed in the target availability zone includes: Obtain the version information of at least one storage cluster deployed in the target availability zone; Based on the version information of each of the at least one storage cluster, select the storage cluster with the higher version as the target storage cluster from the at least one storage cluster.

7. The method according to any one of claims 1-4, characterized in that, The execution entity of the method is either the central node of the distributed data recycling system or any control node in the distributed data recycling system. The control node deployed in the target storage cluster is designated as the target control node, including: If the execution entity of the method is any management node in the distributed data recycling system, then the management node determines whether it is located in the target storage cluster. If so, the control node determines itself as the target control node.

8. A distributed data reclamation system, wherein the distributed data reclamation system is used to execute data reclamation tasks corresponding to a target area, the target area including at least one availability zone, and at least one storage cluster in any availability zone deploys a control node and worker nodes in the distributed data reclamation system, wherein the control node deployed in any storage cluster is used to schedule worker nodes in different storage clusters in the same availability zone, characterized in that, Any management node in the distributed data reclamation system is used to: respond to the triggering operation of the target event, obtain the reclamation configuration information of at least one availability zone in the target area, wherein the reclamation configuration information of any availability zone is determined according to the configuration information of the management node and / or worker node deployed in the availability zone by the distributed data reclamation system; and select the target storage cluster from the storage clusters deployed in the at least one availability zone according to the configuration information of the at least one availability zone and the target cluster selection strategy. The control node deployed in the target storage cluster is designated as the target control node. The target control node is used to schedule the worker nodes in its respective availability zone to perform the data reclamation task corresponding to the target zone.

9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store one or more computer instructions; The processor is configured to execute one or more computer instructions for performing the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it is able to perform the steps of the method described in any one of claims 1-7.

11. A computer program product, characterized in that, include: A computer program / instruction that, when executed by a processor, enables the implementation of the steps in the method described in any one of claims 1-7.