A Method for Implementing a Recycle Bin for HDFS Multi-Tenants

By creating a tenant directory in HDFS and setting quotas and custom recycling strategies, the problem of the existing HDFS recycling bin lacking support for multi-tenant scenarios is solved, and customized recycling bin management and effective control of recycling bin space for each tenant is achieved.

CN118193460BActive Publication Date: 2025-06-20CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410304231.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-06-20
Estimated Expiration
2044-03-18

AI Technical Summary

Technical Problem

The existing HDFS recycling bin setup and management methods lack support for multi-tenant scenarios, and cannot meet the customized recycling bin management needs of various business users. Moreover, the garbage recycling strategy of the recycling bin is only based on time and cannot be cleaned according to the quota threshold, resulting in the recycling bin occupying a large amount of storage space and affecting the normal operation of business operations.

Method used

A method for realizing HDFS multi-tenant recycling bin is proposed. By creating a tenant directory in HDFS and setting the quota and custom recycling strategy for the tenant directory, the customized recycling bin management of each tenant is realized. The specific steps include setting directory quotas according to the tenant's business needs, creating an HDFS tenant directory, and cleaning the recycling bin based on quota and custom policies in the emptier thread.

Benefits of technology

It realizes the customized recycling bin management strategy for each tenant, sets file retention time and quota thresholds according to the tenant's business needs, ensures the limitations on the use of recycling bin space, avoids the problem of recycling bin occupancy of a large amount of storage space, and improves the operation and maintenance efficiency of HDFS clusters and the utilization rate of storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118193460B_ABST
    Figure CN118193460B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for implementing a multi-tenant recycle bin in HDFS. This method includes: determining the directory quota of a tenant according to the tenant's business data volume, business process, and directory usage; creating an HDFS tenant directory and setting the tenant directory as " / tenant / tenant name"; setting the quota and Xattr parameters of the HDFS tenant directory; and the emptier thread cleaning the tenant recycle bin according to the quota and Xattr parameters of the HDFS tenant directory. By creating a tenant directory and setting quota parameters for the tenant directory, this method flexibly limits the usage amount of the entire tenant space and the number of files, restricts the tenant's use of HDFS, and urges the tenant to clean up in a timely manner. The tenant can set the cleaning time of its own recycle bin and the automatic cleaning policy based on the quota threshold calculation ratio through this method, meeting the on-demand recycling requirement for space in the multi-tenant shared Hadoop scenario. In addition, the recycle bin directory in this method is a sub-directory of the tenant directory, which is convenient for administrators to manage and operate, and realizes the timely recycling of storage resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of HDFS recycle bin management methods, and particularly relates to a method for implementing a multi-tenant recycle bin in HDFS. Background Art

[0002] Hadoop Distribute File System (HDFS) is a core project of Hadoop. As a distributed file system, it bears the storage of the entire Hadoop ecosystem, provides scalable, high-throughput, and highly reliable data storage services for upper-layer applications and users in the Hadoop ecosystem, and is suitable for distributed reading and writing of large-scale data, especially scenarios with more reads than writes.

[0003] HDFS adopts a master-slave architecture, consisting of a more core NameNode and DataNode. Among them, the NameNode serves as the master node, mainly responsible for managing the namespace of HDFS, resolving client requests, and controlling client access to HDFS. The DataNode is a slave node, mainly used for physical storage, and its data volume depends on the cluster scale.

[0004] The NameNode (NN) is the master node of HDFS, and its main functions are as follows:

[0005] Responsible for managing the namespace, cluster information, and data blocks of HDFS;

[0006] Maintaining the file directory tree of the entire HDFS, the meta-information of the file directory, and the list of data blocks corresponding to each file;

[0007] Receiving operation requests from clients;

[0008] Managing the mapping relationships between files and data blocks, and between data blocks and DataNodes.

[0009] The DataNode is responsible for processing read and write requests from the file system client, and creating, replicating, and deleting data blocks under the unified scheduling of the NameNode.

[0010] By default, the recycle bin function of HDFS is not enabled, and data will be permanently deleted after being deleted. However, for the HDFS in the online production environment, it is essential to enable the recycle bin function. This function is similar to the recycle bin design of the Linux system. HDFS will create a dedicated recycle bin directory " / user / username / .Trash" for each user. When a user deletes a file, the deleted data is temporarily moved to the recycle bin directory and is not completely cleared. When the data is accidentally deleted, it can be recovered from the recycle bin in a timely manner. The recycle bin-related configuration items are shown in the following table:

[0011]

[0012] As Figure 1 shown, when the NameNode starts, a daemon thread emptier will be started. Every fs.trash.checkpoint.interval, a checkpoint of the current time will be created in the " / user / username / .Trash" directory. The recently deleted data will be placed in this checkpoint, and the expired checkpoints will be deleted every fs.trash.interval period.

[0013] Adding and modifying the following two property values in core-site.xml can enable the Trash function.

[0014]

[0015] For example, set fs.trash.interval = 1440 and fs.trash.checkpoint.interval = 60; then after the file is deleted, it will first enter the Current directory in the recycle bin, and then every 60 minutes (the value corresponding to fs.trash.checkpoint.interval), the subdirectories in the Current directory will be moved to a directory at the same level as Current named after the current timestamp as a checkpoint directory; at the same time, the checkpoint directory 1440 minutes (the value corresponding to fs.trash.interval, 1440 minutes is 24 hours) ago will be cleared.

[0016] The garbage collection policy can be customized by the user and is specified through the parameter fs.trash.classname. HDFS uses TrashPolicyDefault by default.

[0017] With the exponential growth of data volume and the continuous improvement of the Hadoop ecosystem, more and more enterprises choose Hadoop as the basic component of the data warehouse and set up more and more relatively complex application scenarios on the big data cluster. Each business user needs to use HDFS as the underlying distributed storage to save various files and data such as business data, result sets, configuration files, and external data.

[0018] However, the configuration values related to the HDFS recycle bin are global and do not support the multi-tenant scenario. The time policies for the files and folders deleted by each business account to be temporarily stored in the recycle bin are the same, and by default, there is no limit on the storage space used by each tenant, which cannot meet the business and management needs of multi-tenants.

[0019] In addition, when the HDFS recycle bin performs garbage collection, it can only judge whether a file is recycled based on time. In cases where a large number of files are deleted, or a large number of small files are deleted, or the TTL period is unreasonably set to be relatively large, etc., it is unable to provide an additional guarantee, such as recycling according to quota thresholds. In this way, it is very likely that the recycle bin occupies a large amount of storage space of the tenant but cannot be recycled in time, which even affects the normal operation of business jobs. Summary of the Invention

[0020] In order to overcome the above-mentioned defects existing in the existing HDFS recycle bin setting and management methods, the present invention proposes a new method for implementing the HDFS multi-tenant recycle bin.

[0021] For the multi-tenant usage scenario of HDFS in actual production, the present invention provides a solution for implementing the HDFS multi-tenant recycle bin. The HDFS recycle bin using this solution can configure optional tenant directory-level parameters to implement customized recycle bin management strategies for each tenant. In the present invention, by default, the tenant recycle bin directory is moved from " / user / username / .Trash" to " / tenant / tenantname / .Trash" under the HDFS tenant directory, and a quota limit for the tenant directory is set. Different tenant directories can set custom recycle bin file retention times, and can selectively set quota-related thresholds, and adopt a deletion strategy of judging whether checkpoint files need to be deleted according to the thresholds. By forcibly configuring the tenant directory quota, the use of the recycle bin space is indirectly restricted. The tenant can independently determine the file retention time and independently determine the file cleaning time according to its own storage quota usage.

[0022] When implementing the tenant directory, the present invention adds tenant attribute information in Xattr. In order to distinguish it from other ordinary directories in HDFS and better manage the tenant directory, in the present invention, " / user / username" is modified to " / tenant / tenantname", and the tenant can only access its own tenant directory, that is, " / tenant / tenantname". The above design facilitates the management of the recycle bins of users in the superuser group and tenant users under the same HDFS cluster at the same time. The recycle bin directory of the superuser group is still under " / user / username / .Trash", and it is recycled according to the original HDFS configuration policy logic without quota restrictions; while the tenant user recycle bin is under the " / tenant / tenantname / .Trash" directory, and it is processed through the customized TenantsTrashPolicy recycle strategy.

[0023] Specifically, the present invention provides a method for implementing the HDFS multi-tenant recycle bin, as Figure 3 shown, this method includes:

[0024] S1. Determine the directory quota of the tenant according to the tenant's business data volume, business process, and directory usage situation;

[0025] S2. Create an HDFS tenant directory and set the tenant directory to " / tenant / tenant name";

[0026] S3. Set the quota and Xattr parameters of the HDFS tenant directory;

[0027] S4. The emptier thread cleans the tenant recycle bin according to the quota and Xattr parameters of the HDFS tenant directory.

[0028] Furthermore, before creating the HDFS tenant directory in the HDFS multi-tenant recycle bin implementation method of the present invention, first set fs.trash.interval and fs.trash.checkpoint.interval according to the tenant's business data volume, business process, and directory usage situation, and enable the recycle bin function of HDFS.

[0029] Furthermore, when creating the HDFS tenant directory in step S2 of the HDFS multi-tenant recycle bin implementation method of the present invention, the number of files and storage space value of the tenant directory need to be provided to set the name quota and space quota, and the following parameter values can also be selectively provided: user.trash.ttl value, user.trash.quota.pct value, user.trash.spacequota.pct value.

[0030] Furthermore, the setting of the quota of the HDFS tenant directory in step S3 of the HDFS multi-tenant recycle bin implementation method of the present invention includes:

[0031] S311. Set the name quota: The name quota refers to the maximum number of files and directories in the root directory tree, that is, recursively calculate the number of files and directories under the subdirectory. The setting command is as follows:

[0032] hdfs dfsadmin - setQuota name quota size / tenant / tenant name;

[0033] S312. Set the space quota: The space quota refers to the total size limit of all files in a single directory, and the size of file replicas is also calculated. The setting command is as follows:

[0034] hdfs dfsadmin - setSpaceQuota space quota size / tenant / tenant name.

[0035] Furthermore, the setting of the Xattr of the HDFS tenant directory in step S3 of the HDFS multi-tenant recycle bin implementation method of the present invention includes:

[0036] S321. Set time-related Xattr;

[0037] S322. Set quota-related Xattr.

[0038] Furthermore, the setting of time-related Xattr described in step S321 of the HDFS multi-tenant recycle bin implementation method of the present invention includes:

[0039] (1) If the user.trash.ttl value is provided when creating the tenant directory, set the Xattr of the " / tenant / tenant name" directory: user.trash.ttl, indicating the retention time of the recycle bin files of this tenant;

[0040] (2) If the user.trash.ttl value provided when creating the tenant directory is -1, the recycle bin files of this tenant will not be cleaned by the NameNode background thread emptier. When there is no specified recycle bin quota threshold parameter, it will be independently maintained by this tenant. If a recycle bin quota threshold parameter is specified, the cleaning time and cleaning method will be determined according to the comparison result with this parameter;

[0041] (3) If the user.trash.ttl value provided when creating the tenant directory is a positive integer and this value is greater than or equal to fs.trash.checkpoint.interval, after the file is deleted, it will first enter the recycle bin directory (i.e.,.Trash directory) under the " / tenant / tenant name" directory of this tenant, and then every fs.trash.checkpoint.interval minutes, the subdirectories under the Current directory will be moved to a directory named with the current timestamp at the same level as Current as a checkpoint directory, and at the same time, the checkpoint directory before the interval of user.trash.ttl minutes will be cleared;

[0042] (4) If the user.trash.ttl value provided when creating the tenant directory is a positive integer and this value is less than fs.trash.checkpoint.interval, then clean up at intervals of fs.trash.checkpoint.interval;

[0043] (5) If the user.trash.ttl value is not provided when creating the tenant directory, do not set the Xattr of this tenant directory. At this time, the default fs.trash.interval of HDFS will be used.

[0044] Further, the setting of Xattrs related to quotas described in step S322 of the method for implementing the HDFS multi-tenant recycle bin of the present invention includes: If the values of user.trash.quota.pct and user.trash.spacequota.pct are provided when creating a tenant directory, set the Xattrs of the " / tenant / tenant name" directory: user.trash.quota.pct and user.trash.spacequota.pct, which respectively represent the percentage of the total number of currently used files in the tenant directory to the total quota of files and the percentage of the actual storage space currently used in the tenant directory to the total quota of storage space; If the storage water level of the current tenant directory has reached the file quantity quota percentage or the space quantity quota percentage, several checkpoint directories in the recycle bin will be deleted in the fs.trash.checkpoint.interval cycle until the water levels of the tenant file quantity and space size are reduced below the above two thresholds. If emptying the recycle bin directory still cannot reduce it below the threshold, this cleaning operation will be exited.

[0045] Further, in the method for implementing the HDFS multi-tenant recycle bin of the present invention, first, it is judged whether to execute the timeout deletion logic according to the setting of user.trash.ttl, and then it is judged whether to execute the quota threshold ratio calculation strategy deletion logic according to the setting of the quota threshold parameters (user.trash.quota.pct, user.trash.spacequota.pct);

[0046] If the values of user.trash.quota.pct and user.trash.spacequota.pct are provided simultaneously when creating a tenant directory, first calculate and process the user.trash.spacequota.pct logic, and then calculate and process the user.trash.quota.pct logic.

[0047] Further, in step S4 of the method for implementing the HDFS multi-tenant recycle bin of the present invention, the emptier thread cleans the tenant recycle bin according to the quota and Xattr parameters of the HDFS tenant directory, including:

[0048] S41. The emptier thread creates a checkpoint directory and deletes the expired checkpoint directory every fs.trash.checkpoint.interval minutes, and at the same time calculates the usage of the storage space under the current tenant directory;

[0049] S42. Determine whether the user.trash.ttl value is set in the tenant directory. If it is not set, the emptier thread deletes files that were created more than fs.trash.interval minutes ago; if the set user.trash.ttl value is -1, the emptier thread does not execute the deletion logic based on the time setting; if the set user.trash.ttl value is a positive integer, the emptier thread deletes files that were created more than user.trash.ttl minutes ago;

[0050] S43. Determine whether the user.trash.spacequota.pct value is set in the tenant directory. If the user.trash.spacequota.pct value is set, if the actual space occupancy / space quota size in the current tenant directory < user.trash.spacequota.pct, this cleaning step is completed; if the actual space occupancy / space quota size in the current tenant directory ≥ user.trash.spacequota.pct, then calculate the space usage of the checkpoint folder based on the creation time. If deleting the earliest created checkpoint can reduce it below user.trash.spacequota.pct, then only delete the earliest created checkpoint directory; if multiple checkpoint directories need to be deleted, then delete multiple checkpoint directories in the order from the oldest to the newest creation time; if emptying the tenant recycle bin just cannot or also cannot reduce it below user.trash.spacequota.pct, then empty the tenant recycle bin, and this cleaning step is completed;

[0051] S44. Determine whether the user.trash.quota.pct value is set in the tenant directory. If the user.trash.quota.pct value is set, if the actual number of files and folders / name quota size in the current tenant directory < user.trash.quota.pct, this cleaning step is completed; if the actual number of files and folders / name quota size in the current tenant directory ≥ user.trash.quota.pct, then calculate the space usage of the checkpoint folder based on the creation time. If deleting the earliest created checkpoint can reduce it below user.trash.quota.pct, then only delete the earliest created checkpoint directory; if multiple checkpoint directories need to be deleted, then delete multiple checkpoint directories in the order from the oldest to the newest creation time; if emptying the tenant recycle bin just cannot or also cannot reduce it below user.trash.quota.pct, then empty the tenant recycle bin, and this cleaning step is completed.

[0052] Furthermore, in the method for implementing the HDFS multi-tenant recycle bin of the present invention, the tenant directory is set to " / tenant / tenant name", and the tenant recycle bin directory is set to " / tenant / tenant name / .Trash". The tenant can only access its own tenant directory and perform cleaning through a customized recycle bin policy logic. At the same time, the recycle bin directory of the HDFS administrator account (i.e., the superuser in dfs.permissions.superusergroup) remains " / user / user name / .Trash", and it still performs recycling according to the default configuration policy logic of HDFS without quota restrictions.

[0053] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above-mentioned method for implementing the HDFS multi-tenant recycle bin are realized.

[0054] In summary, the method for implementing the HDFS multi-tenant recycle bin of the present invention has the following advantages:

[0055] (1) Through the method of the present invention, each tenant can set its own recycle bin cleaning time, and based on the customized retention policy, multi-tenant on-demand configuration is achieved, meeting the on-demand recycling requirement of space in the multi-tenant shared Hadoop scenario.

[0056] (2) This method restricts the entire tenant space usage and the number of files by setting quotas for the tenant directory, constrains the tenant's use of HDFS, urges the tenant to clean up in a timely manner, and indirectly limits the occupation of the recycle bin space. Moreover, the soft limit method adopted by the present invention is flexible.

[0057] (3) This method realizes a customized automatic cleaning policy from multiple dimensions such as the number of files, space usage, and quota ratio threshold by setting quota parameters.

[0058] (4) In this method, the recycle bin directory is a sub-directory of the tenant directory, which is convenient for administrators to manage, operate, and control permissions. This method simplifies the operation and maintenance of the HDFS cluster, recovers storage resources in a timely manner, and greatly reduces the waste of storage resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required to be used in the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments recorded in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0060] Figure 1It is a flowchart for Hadoop to perform snapshots and cleanups by default based on org.apache.hadoop.fs.TrashPolicyDefault. (To use the trash can function, it is necessary to set fs.trash.interval to a non-zero value and start the NameNode background thread emptier to perform the cleanup operation).

[0061] Figure 2 This is the overall implementation flowchart of the method of the present invention.

[0062] Figure 3 This is the implementation process schematic diagram of the method of the present invention.

[0063] Figure 4 This is the flowchart for the method of the present invention to perform cleanup for quota threshold parameters.

[0064] Figure 5 This is the schematic diagram of the application example of the method of the present invention (above the timeline in the figure is Tenant 1, and below the timeline is Tenant 2). Detailed implementation manners

[0065] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0066] Meanwhile, it should be understood that the protection scope of the present invention is not limited to the specific implementation manners described below; it should also be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific implementation manners, rather than limiting the protection scope of the present invention.

[0067] Embodiment: A method for implementing a multi-tenant trash can in HDFS

[0068] As Figure 2 shown, the implementation process of this method is as follows:

[0069] 1. Before opening the tenant directory

[0070] Before opening the tenant directory, it is necessary to determine the quota of the tenant directory according to the business data volume, business process, and directory usage. For example, the quota of Tenant A's directory is determined as follows:

[0071] HDFS Directory Number of Files Storage Space (GB) hdfs: / / namespace / tenant / Tenant A 10000 300

[0072] 2. Open the HDFS tenant directory

[0073] The customized tenant recycle bin function can only be used by enabling the recycle bin function, that is, reasonably setting fs.trash.interval and fs.trash.checkpoint.interval.

[0074] For each additional Hadoop cluster tenant, an HDFS tenant directory needs to be opened to meet the storage needs of the tenant. The tenant directory is set to " / tenant / tenant name". When opening the tenant directory, optional values of user.trash.ttl, user.trash.quota.pct, user.trash.spacequota.pct, and the required number of files and storage space values of the tenant directory need to be provided.

[0075] 2.1 Quota setting implementation

[0076] 2.1.1 Automatic

[0077] When creating a tenant directory, call the interface to automatically set according to the provided information.

[0078] 2.1.2 Manual

[0079] The super administrator user can also manually set through the following command:

[0080] Set name quota: The name quota refers to the maximum number of files and directories in the root directory tree, that is, recursively calculate the number of files and directories in the subdirectories:

[0081] hdfs dfsadmin - setQuota 10000 / tenant / tenant A

[0082] Set space quota: The space quota refers to the total size limit of all files in a single directory, and the size of file replicas is also included:

[0083] hdfs dfsadmin - setSpaceQuota 322122547200 / tenant / tenant A

[0084] View directory quota status:

[0085] hdfs dfs - count - q / tenant / userA

[0086]

[0087] 2.2 Set Xattr of HDFS tenant directory

[0088] 2.2.1 Time - related Xattr

[0089] If the value of user.trash.ttl is provided when creating a tenant directory, the Xattr of the directory " / tenant / tenant name" will be set: user.trash.ttl, which represents the retention time of the recycle bin files of this tenant. If the value of user.trash.ttl provided when creating the tenant directory is -1, the tenant's recycle bin files will not be cleaned by the NameNode background thread emptier. Without specifying the recycle bin quota threshold parameter, it will be maintained by the tenant itself; if the recycle bin quota threshold parameter is specified, it will be processed according to 2.2.2. When the value of user.trash.ttl is set to a positive integer, it is recommended that this value be greater than or equal to fs.trash.checkpoint.interval. At this time, when a file is deleted, it will first enter the Current directory in the recycle bin directory under the tenant directory, and then the subdirectories in the Current directory will be moved to a directory named with the current timestamp at the same level as Current every fs.trash.checkpoint.interval minutes as a checkpoint directory; at the same time, the checkpoint directory older than user.trash.ttl minutes will be cleared. If it is less than fs.trash.checkpoint.interval, it will be cleaned at intervals of fs.trash.checkpoint.interval.

[0090] If the value of user.trash.ttl is not provided when creating a tenant directory, the Xattr of this tenant directory will not be set, and the default fs.trash.interval of HDFS will be used at this time.

[0091] 2.2.2 Quota-related Xattr

[0092] If the values of user.trash.quota.pct and user.trash.spacequota.pct are provided when creating a tenant directory, the Xattrs of the directory " / tenant / tenant name" will be set: user.trash.quota.pct and user.trash.spacequota.pct, which respectively represent the percentage of the total number of files currently used in the tenant directory to the total quota of files and the percentage of the actual storage space currently used in the tenant directory to the total quota of storage space. When the percentage of the directory and file quantity quota or the percentage of the space quota quantity is reached, several checkpoint directories in the recycle bin will be deleted in the fs.trash.checkpoint.interval cycle until the tenant's file quantity and space size levels are reduced below these two thresholds, or if emptying the recycle bin directory still cannot reduce it below the threshold, this cleaning operation will exit.

[0093] First, determine whether to execute the timeout deletion logic based on the configuration of user.trash.ttl. Then, determine whether to execute the quota threshold ratio calculation policy deletion logic based on the setting of the quota threshold parameter. user.trash.quota.pct and user.trash.spacequota.pct are two optional configuration items that can be provided by the tenant simultaneously, or only one of them can be provided, or neither can be provided. If both parameters are provided, the logic of user.trash.spacequota.pct will be calculated and processed first, and then the logic of user.trash.quota.pct will be calculated and processed. If neither is provided or configured, there is no need to calculate the quota ratio.

[0094] If the value of user.trash.spacequota.pct provided when creating the tenant directory is 0.8 and the value of user.trash.quota.pct is 0.9, when the NameNode background thread emptier creates a checkpoint file and deletes files at regular intervals of fs.trash.checkpoint.interval minutes, it will calculate the storage space usage in the current tenant directory.

[0095] As Figure 4 shown, if user.trash.spacequota.pct is configured, the current tenant space usage will be calculated when performing checkpoint deletion checks. When the actual space occupancy / space quota size < user.trash.spacequota.pct (0.8), this cleanup is completed. If the actual space occupancy / space quota size >= user.trash.spacequota.pct (0.8), then it is necessary to calculate the space usage of the checkpoint folder according to the creation time. Deleting the earliest created checkpoint can reduce it to below 0.8, then only the earliest created checkpoint directory is deleted; if multiple checkpoint directories need to be deleted, multiple checkpoint directories will be deleted in the order from the oldest to the newest creation time; if emptying all the recycle bins still cannot reduce it to within 0.8, the tenant recycle bin will be emptied and this cleanup is completed.

[0096] If user.trash.quota.pct is configured, the usage of the current tenant's space file count will be calculated during checkpoint deletion checks. When the actual number of files and folders / name quota size < user.trash.quota.pct (0.9), this cleanup is completed. When the actual number of files and folders / name quota size >= user.trash.quota.pct (0.9), the usage of the checkpoint folder space needs to be calculated based on the creation time, and deleting the earliest created checkpoint can reduce it below 0.9, then only the earliest created checkpoint directory is deleted; if multiple checkpoint directories need to be deleted, multiple checkpoint directories are deleted in order from the oldest to the newest creation time; if emptying the entire trash can just meets the requirement or still cannot reduce it below 0.9, the tenant's trash can will be emptied and this cleanup is completed.

[0097] In short, the previous fs.trash.interval for the entire cluster was replaced by the tenant-level user.trash.ttl when implementing the tenant-level trash can function. And optional configuration items user.trash.quota.pct and user.trash.spacequota.pct are added to implement custom automatic garbage collection for quota threshold ratio calculation. The former can periodically clear expired files in the tenant's trash can, and the latter will try to ensure that the tenant usage is below an expected level.

[0098] 3. Implementation of Custom TrashPolicy

[0099] Specify the custom garbage collection policy implementation class as org.apache.hadoop.fs.TenantsTrashPolicy through the parameter fs.trash.classname.

[0100] Adding the following property values in core-site.xml can enable the multi-tenant custom Trash function.

[0101]

[0102] It is logically different from the HDFS default implementation class org.apache.hadoop.fs.TrashPolicyDefault, which is reflected in:

[0103] For the HDFS administrator account (i.e., the superuser in dfs.permissions.superusergroup), such as the hdfs user, the trash can directory remains " / user / hdfs / .Trash";

[0104] For business tenants, the recycle bin directory is set to " / tenant / tenant name / .Trash" (" / tenant / tenant name" is the business customized client directory);

[0105] The processing logic of the emptier thread for deleting files in the recycle bin of the administrator account remains unchanged;

[0106] The deletion logic of the emptier thread for tenants is to traverse each tenant account directory under the " / tenant" directory, read the setting of user.trash.ttl in its Xattr, and the processing logic of the emptier thread is the same as that introduced in Section 2.2.

[0107] Since the user.trash.ttl setting at the tenant level is supported, tenants can keep files for a longer time, which also means that tenant directories require more storage space.

[0108] Through the tenant directory quota set in Step 2.1, the entire tenant folder (client directory) is restricted, which means that there is a flexible soft limit on the recycle bin folder. Tenants can decide on their own when to clean up files and the recycle bin directory under the directory according to the directory usage situation, including manual cleaning.

[0109] In addition, by supporting the user.trash.quota.pct and user.trash.spacequota.pct at the tenant level, a customized automatic garbage cleaning policy calculated according to the quota threshold ratio is implemented.

[0110] As an operation example, Figure 5 it describes the file change situation in the recycle bin from 10:00:00 to 13:00:00 on March 10, 2023, when the administrator set user.trash.ttl = 120 for "Tenant 1" and user.trash.ttl = 100 for "Tenant 2" in the HDFS cluster with fs.trash.checkpoint.interval = 60 set. The green ones represent newly added snapshots, and the red ones represent the snapshots deleted this time. Above the timeline is Tenant 1, and below the timeline is Tenant 2. It can be seen that through this method, multi-tenant customized management of the recycle bin can be well achieved.

[0111] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Any person skilled in the relevant art can, without departing from the technical solution of the present invention, make some changes or modifications using the above-disclosed technical content to obtain equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A method for implementing a multi-tenant HDFS recycle bin, characterized in that: The method comprises: S1. Determine the tenant's directory quota based on the tenant's business data volume, business process, and directory usage; S2. Create an HDFS tenant directory and set the tenant directory to " / tenant / tenant name"; S3. Set the quota and Xattr parameters of the HDFS tenant directory; Set quotas for HDFS tenant directories, including: S311. Set name quota: The name quota refers to the maximum number of files and directories in the root directory tree, that is, the number of files and directories in the subdirectories recursively calculated; S312. Set space quota: The space quota refers to the total size limit of all files in a single directory, including the size of file copies; Set the Xattr parameters of the HDFS tenant directory, including: S321. Set time-related Xattr, including: (1) If the user.trash.ttl value is provided when creating the tenant directory, set the Xattr of the " / tenant / tenant name" directory: user.trash.ttl, which indicates the file retention time of the tenant's recycle bin; (2) If the user.trash.ttl value provided when creating a tenant directory is -1, the tenant's recycle bin files will not be cleaned up by the NameNode background thread emptier. If the recycle bin quota threshold parameter is not specified, the tenant will maintain it independently. If the recycle bin quota threshold parameter is specified, the cleaning time and method will be determined based on the comparison result with the recycle bin quota threshold parameter; (3) If the user.trash.ttl value provided when creating a tenant directory is a positive integer, and the value is greater than or equal to fs.trash.checkpoint.interval, after the file is deleted, it will first enter the Current directory under the Recycle Bin directory under the tenant directory, and then periodically move the subdirectories under the Current directory to a directory named with the current timestamp at the same level as the Current directory every fs.trash.checkpoint.interval minutes as a checkpoint directory. At the same time, the checkpoint directory before the interval user.trash.ttl minutes will be cleared; (4) If the user.trash.ttl value provided when creating a tenant directory is a positive integer and the value is less than fs.trash.checkpoint.interval, the directory will be cleaned up at the interval of fs.trash.checkpoint.interval; (5) If the user.trash.ttl value is not provided when creating a tenant directory, the Xattr of the tenant directory will not be set, and the HDFS default fs.trash.interval will be used; S322 sets the quota related Xattr, the quota includes the name quota and space quota; S4. The emptier thread cleans up the tenant recycle bin according to the quota and Xattr parameters of the HDFS tenant directory.

2. The method for implementing a multi-tenant HDFS recycle bin according to claim 1, characterized in that: Before creating an HDFS tenant directory, first set fs.trash.interval and fs.trash.checkpoint.interval according to the tenant's business data volume, business process, and directory usage, and enable the HDFS recycle bin function.

3. The method for implementing a multi-tenant HDFS recycle bin according to claim 1, characterized in that: When creating an HDFS tenant directory in step S2, it is necessary to provide the number of files and storage space values ​​of the tenant directory for setting the name quota and space quota, and also provide the following parameter values: user.trash.ttl value, user.trash.quota.pct value, user.trash.spacequota.pct value, wherein the user.trash.quota.pct value and user.trash.spacequota.pct value respectively represent the percentage of the total number of files currently used in the tenant directory to the total number of files in the quota, and the percentage of the actual storage space currently used in the tenant directory to the total quota storage space.

4. The HDFS multi-tenant recycle bin implementation method according to claim 1, characterized in that: The command for setting the name quota in step S311 is as follows: hdfs dfsadmin -setQuota name quota size / tenant / tenant name; The space quota setting described in step S312 is set by the following command: hdfs dfsadmin -setSpaceQuota space quota size / tenant / tenant name.

5. The method for implementing a multi-tenant HDFS recycle bin according to claim 1, characterized in that: The quota-related Xattr setting described in step S322 includes: if the user.trash.quota.pct value and the user.trash.spacequota.pct value are provided when creating the tenant directory, then the Xattr of the " / tenant / tenant name" directory is set: user.trash.quota.pct and user.trash.spacequota.pct, which respectively represent the percentage of the total number of files currently used in the tenant directory to the total number of files in the quota, and the percentage of the actual storage space currently used in the tenant directory to the total quota storage space; if the current tenant directory storage water level has reached the file number quota percentage or the space number quota percentage, several checkpoint directories in the recycle bin will be deleted during the fs.trash.checkpoint.interval period until the tenant file number and space size water levels are reduced to below the above two thresholds. If emptying the recycle bin directory still cannot reduce the water level below the threshold, the current cleaning operation is exited.

6. The method for implementing a multi-tenant HDFS recycle bin according to claim 5, characterized in that: In this method, firstly, the timeout deletion logic is determined based on the user.trash.ttl setting, and then the quota threshold parameter setting is determined based on whether the quota threshold ratio calculation strategy deletion logic is executed; If the values of user.trash.quota.pct and user.trash.spacequota.pct are provided simultaneously when creating a tenant directory, the logic of user.trash.spacequota.pct is calculated and processed first, and then the logic of user.trash.quota.pct is calculated and processed.

7. The method for implementing a multi-tenant HDFS recycle bin according to claim 5, characterized in that: The emptier thread described in step S4 cleans the tenant recycle bin according to the quota and Xattr parameters of the HDFS tenant directory, including: S41. The emptier thread creates a checkpoint directory every fs.trash.checkpoint.interval minutes and deletes the expired checkpoint directory, and at the same time calculates the usage of the storage space under the current tenant directory; S42. Determine whether the user.trash.ttl value is set for the tenant directory. If it is not set, the emptier thread deletes the files that were created fs.trash.interval minutes ago; if the set user.trash.ttl value is -1, the emptier thread does not execute the deletion logic according to the time setting; if the set user.trash.ttl value is a positive integer, the emptier thread deletes the files that were created user.trash.ttl minutes ago; S43. Determine whether the user.trash.spacequota.pct value is set for the tenant directory. If the user.trash.spacequota.pct value is set, if the actual space occupancy / space quota size under the current tenant directory < user.trash.spacequota.pct, this cleaning step is completed; if the actual space occupancy / space quota size under the current tenant directory ≥ user.trash.spacequota.pct, then calculate the space usage of the checkpoint folder according to the creation time. If deleting the earliest created checkpoint can reduce it below user.trash.spacequota.pct, then only delete the earliest created checkpoint directory; if multiple checkpoint directories need to be deleted, then delete multiple checkpoint directories in the order from the oldest to the newest; if emptying the tenant recycle bin just cannot or also cannot reduce it below user.trash.spacequota.pct, then the tenant recycle bin will be emptied and this cleaning step is completed; S44. Determine whether the user.trash.quota.pct value is set for the tenant directory. If the user.trash.quota.pct value is set, if the actual number / name quota size of files and folders in the current tenant directory < user.trash.quota.pct, this step of cleaning is completed; if the actual number / name quota size of files and folders in the current tenant directory ≥ user.trash.quota.pct, then calculate the space usage of the checkpoint folder according to the creation time. If deleting the earliest created checkpoint can reduce it below user.trash.quota.pct, only delete the earliest created checkpoint directory; if multiple checkpoint directories need to be deleted, delete multiple checkpoint directories in order from the oldest to the newest according to the creation time; if emptying the tenant recycle bin just cannot reduce it below user.trash.quota.pct either, then empty the tenant recycle bin, and this step of cleaning is completed.

8. The method for implementing a multi-tenant HDFS recycle bin according to claim 1, characterized in that: In this method, the tenant directory is set to " / tenant / tenant name", the tenant recycle bin directory is set to " / tenant / tenant name / .Trash", and the tenant can only access its own tenant directory and perform cleaning through the customized recycle bin policy logic; at the same time, the recycle bin directory of the HDFS administrator account remains " / user / user name / .Trash", and it still performs recycling according to the default configuration policy logic of HDFS and has no quota limit.

Citation Information

Patent Citations

  • Automatic HDFS (Hadoop Distributed File System) file cleaning cleaning method and device and storage medium

    CN112800010A

  • Data cleaning method and device, electronic equipment and computer readable storage medium

    CN112925745A

  • Storage platform service method and device, equipment and storage medium

    CN116910015A