A processing method, device, terminal and storage medium for storing objects

By accessing and expiration judgments on the metadata of large objects, the inefficiency problem caused by the massive number of small objects one by one is solved, and efficient small object expiration clearance and data processing performance improvements are achieved.

CN115878027BActive Publication Date: 2025-05-30ZHEJIANG UNIVIEW TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210908976.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-05-30
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

In the prior art, a large number of small objects are judged one by one, resulting in low data processing efficiency and accompanied by a large number of useless lists, which greatly reduces the data processing performance of distributed storage clusters.

Method used

By obtaining the object name of the large object numbered in time, accessing the metadata of the large object in the large object queue in sequence, based on the size relationship between the file expiration time and the writing time of the first small object in the large object and the writing time of the last small object, the expiration judgment of the large object as a whole is realized, reducing the useless list of a large number of small objects, and using the numbered time sequence arrangement relationship between different large objects to make corresponding expiration judgments.

Benefits of technology

Without adding additional overhead, the expired clearance of small objects is achieved, data processing efficiency is improved, and data processing performance of distributed storage clusters is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115878027B_ABST
    Figure CN115878027B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of object storage technology, and provides a method, device, terminal and storage medium for processing stored objects. The method includes: obtaining the object names of each large object in the large object queue, accessing the metadata of the large objects in the large object queue in sequence based on the object names, and obtaining a first start time and a first end time from the metadata of the first large object; obtaining the magnitude relationship between the set file expiration time and the first start time and the first end time; based on the magnitude relationship, determining whether there are expired small objects in the small objects stored in the first large object and whether there are expired small objects in the second large object having a chronological arrangement relationship with the number of the first large object, to obtain a determination result; the determination result is used to indicate whether to perform an expired file clearing operation. This solution can improve data processing efficiency and ensure the data processing performance of the distributed storage cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of object storage, and particularly relates to a method, device, terminal, and storage medium for processing stored objects. Background Art

[0002] In recent years, the birth of distributed storage has enabled the wide application of cloud storage technology. Among them, object storage is applied, and specifically, data uploaded through an object storage gateway needs to be stored in a distributed storage cluster in the form of objects. To save storage costs, the life cycle process of the object storage gateway provides a way to automatically delete expired objects, enabling the storage cluster to regularly release storage space to meet the requirements of continuous writing.

[0003] However, with the expansion of the data scale, small objects that need to be stored formed by a large number of small files will cause a decline in the performance of distributed storage. To solve the problem of storing a large number of small objects, the industry often adopts the method of merging small objects, merging multiple small objects into one large object for storage, greatly reducing the amount of metadata of small objects, and at the same time reducing frequent disk input and output, thereby improving the overall storage performance.

[0004] As the merging of small objects progresses, when judging and deleting the expiration of tens of thousands of small objects, it is necessary to list and traverse each small object in the large object respectively to implement the validity judgment of the small object and effectively eliminate the expired storage objects. However, in this process, judging the expiration of a large number of small objects one by one results in low data processing efficiency and will also be accompanied by a large number of useless listings, greatly reducing the data processing performance of the distributed storage cluster. Summary of the Invention

[0005] Embodiments of this application provide a method, device, terminal, and storage medium for processing stored objects to solve the problem in the prior art that judging the expiration of a large number of small objects one by one results in low data processing efficiency and will also be accompanied by a large number of useless listings, greatly reducing the data processing performance of the distributed storage cluster.

[0006] The first aspect of the embodiments of this application provides a method for processing stored objects, including:

[0007] Obtain the object names of each large object in the large object queue, the object names corresponding to the numbers of the large objects, and the numbers of each of the large objects are arranged in sequence according to time;

[0008] Access the metadata of the large objects in the large object queue in sequence based on the object name, and obtain the first start time and the first end time from the metadata of the first large object; wherein, the metadata of each large object includes a start time and an end time, the start time corresponds to the write time of the first small object, and the end time corresponds to the write time of the last small object;

[0009] Obtain the size relationship between the set file expiration time and the first start time and the first end time;

[0010] Based on the size relationship, determine whether there are expired small objects in the small objects stored in the first large object and whether there are expired small objects in the second large object whose number has a sequential arrangement relationship with the number of the first large object, and obtain a judgment result; the judgment result is used to indicate whether to perform the expired file clearing operation.

[0011] A second aspect of the embodiments of the present application provides a processing device for storing objects, including:

[0012] A first acquisition module, configured to acquire the object names of the large objects in the large object queue, the object name corresponding to the number of the large object, and the numbers of the large objects are arranged in sequence;

[0013] A second acquisition module, configured to access the metadata of the large objects in the large object queue in sequence based on the object name, and obtain the first start time and the first end time from the metadata of the first large object; wherein, the metadata of each large object includes a start time and an end time, the start time corresponds to the write time of the first small object, and the end time corresponds to the write time of the last small object;

[0014] A second acquisition module, configured to obtain the size relationship between the set file expiration time and the first start time and the first end time;

[0015] A judgment module, configured to determine whether there are expired small objects in the small objects stored in the first large object and whether there are expired small objects in the second large object whose number has a sequential arrangement relationship with the number of the first large object based on the size relationship, and obtain a judgment result; the judgment result is used to indicate whether to perform the expired file clearing operation.

[0016] A third aspect of the embodiments of the present application provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the method described in the first aspect are implemented.

[0017] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0018] In a fifth aspect of the present application, a computer program product is provided. When the computer program product runs on a terminal, the terminal is caused to execute the steps of the method described in the first aspect above.

[0019] As can be seen from the above, in the embodiments of the present application, by obtaining the object names of large objects whose numbers are sorted according to time, accessing the metadata of the large objects in the large object queue in sequence, and based on the relationship between the set file expiration time and the write time of the first small object and the write time of the last small object in the current large object, the expiration judgment of the entire large object is realized, reducing the useless enumeration of a large number of small objects. At the same time, by using the sequential arrangement relationship of the numbers between different large objects, the corresponding expiration judgment of other large objects having a sequential arrangement relationship with the current large object is realized, reducing the useless enumeration of other large objects. Without adding additional overhead, the expiration clearing of small objects is realized, improving the data processing efficiency and ensuring the data processing performance of the distributed storage cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 is the flowchart of a method for processing storage objects provided by the embodiments of the present application Figure 1 ;

[0022] Figure 2 is the flowchart of a method for processing storage objects provided by the embodiments of the present application Figure 2 ;

[0023] Figure 3 is a schematic diagram of a management object provided by the embodiments of the present application;

[0024] Figure 4 is a schematic diagram of hole processing provided by the embodiments of the present application;

[0025] Figure 5 is the structural diagram of a device for processing storage objects provided by the embodiments of the present application;

[0026] Figure 6 is the structural diagram of a terminal provided by the embodiments of the present application. Detailed implementation manners

[0027] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0028] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0029] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0030] It should be further understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0031] As used in this specification and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0032] It should be understood that the magnitude of the sequence numbers of the steps in this embodiment does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0033] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.

[0034] Refer to Figure 1 , Figure 1 which is the flowchart of a method for processing stored objects provided by an embodiment of the present application Figure 1 . AsFigure 1 As shown in the figure, a method for processing storage objects, the method comprising the following steps:

[0035] Step 101, obtaining the object names of each large object in the large object queue.

[0036] The object name corresponds to the number of the large object, and the numbers of each of the large objects are arranged in chronological order.

[0037] Each of the large objects corresponds to a storage space, and the storage space is used to write small objects. Herein, a small object generally refers to a file with a size less than 10M.

[0038] The numbers of each of the large objects are arranged in chronological order. Specifically, the numbers of each of the large objects are arranged in corresponding order according to the chronological order of the application times of the large objects. The later the application time of a large object, the larger the numerical value of its number. The large object numbers meet the chronological requirements, and can ensure that when listing a certain large object and all the small objects on this large object are not expired, all the large objects after the number of this large object cannot be expired.

[0039] When assigning numbers to the large objects, it is possible to directly assign numbers according to the order that can reflect the chronological relationship. Therefore, in the large object queue, the application time of the last large object is the latest large object. Based on the number of this large object, the object names of each of the large objects can be constructed according to the number assignment rule of the large object and the composition structure of the large object name.

[0040] By assigning numbers to the large objects that can reflect the chronological order of their application times, the object names of each of the large objects are constructed again. Through these object names, the metadata of the corresponding large objects can be directly accessed, rather than accessing the database through the index method to obtain the object name for the original data access.

[0041] In the embodiment of the present application, the object life cycle management of the large objects is specifically carried out.

[0042] Before clearing the expiration of the small objects in the large objects, it is necessary to first construct the large object queue.

[0043] Specifically, the large objects need to be stored in the object storage buckets (buckets) divided in the distributed storage cluster, and several large objects stored in an object storage bucket form a large object queue according to the time order.

[0044] Corresponding to this process, as shown in Figure 2 Before obtaining the object names of each of the large objects in the large object queue, it includes:

[0045] Step 21, based on the small object to be stored, when it is found that the currently available large object is full, a target object name is constructed based on the serial number of the currently available large object;

[0046] Step 22: Apply to the object cluster for the target large object corresponding to the target object name.

[0047] The target object name includes the serial number of the target large object, and the serial number of the target large object and the serial number of the currently available large object are arranged in chronological order.

[0048] Step 23: write the small object into the storage space corresponding to the target large object.

[0049] Among them, in the embodiment of the present application, when a small object needs to be uploaded, the specific upload process is as follows:

[0050] First, the user initiates a file upload request to the object storage gateway. After receiving the request, the object storage gateway determines whether the uploaded file is a small object. If it is a small object, it pre-allocates N management objects under each bucket and performs allocation and management records of large objects.

[0051] The concept of management object is introduced here. Each bucket has a corresponding management object. When a small object needs to be uploaded, it will first find one of the management objects. The management object determines which large object the small object should be written into and allocates a large object number (ID) for applying for and listing large objects to ensure the orderliness of the IDs assigned to the objects. The smaller the value, the earlier the ID is assigned.

[0052] When a small object is uploaded, it can be mapped to one of the management objects of the bucket through hashing.

[0053] The management object queries whether the currently available large object is full. If full, it constructs the large object name through the recorded ID and applies for a new large object from the object cluster. Each time a new large object is applied, its corresponding ID value increases. Then the small objects can be written into the data area of ​​the allocated large object in sequence.

[0054] The large object metadata records the list of small objects written therein, including information such as the offset and length of the small objects in the data area. At the same time, for the implementation of the embodiments of the present application, the large object metadata also records the following field information: first_time represents the time when the first small object is written, last_time represents the time when the last small object is written, tags_num represents the number of small objects with tag attributes among all small objects, in_hole_queue represents whether it is in the task queue of the hole processing thread, and use_size represents the amount of storage space already used. Among them, the hole processing thread is used to process the data of large objects with storage holes to realize the release of storage space.

[0055] Secondly, when writing a small object into a large object, the metadata information of the large object can be updated.

[0056] Among them, if the current small object is the first small object written into the large object, then update first_time to the current time; if it is the last small object written into the large object, update last_time to the current time; if the small object contains a tag, the tags_num attribute is incremented; in_hole_queue is set to the default value false, which means at this time that the large object is not placed in the task queue of the hole processing thread; use_size is the sum of the sizes of all small objects.

[0057] Among them, the object name format of the large object can also be set. For example: bucket_id–management object serial number - number, that is, it includes the number of the storage bucket where the large object is located, the number of the management object, and the number of the large object itself. Each time a large object is allocated, the number of the large object will increase.

[0058] Step 102, based on the object name, sequentially access the metadata of the large objects in the large object queue, and obtain the first start time and the first end time from the metadata of the first large object.

[0059] Among them, the metadata of each large object includes a start time and an end time. The start time corresponds to the writing time of the first small object, and the end time corresponds to the writing time of the last small object.

[0060] Step 103, obtain the size relationship between the set file expiration time and the first start time and the first end time.

[0061] This size relationship specifically includes: the set file expiration time is earlier than the first start time, the set file expiration time is later than the first end time, and the set file expiration time is within the time range corresponding to the first start time and the first end time.

[0062] Setting the file expiration time can be a set time value, which can be set according to requirements.

[0063] Step 104: Based on the size relationship, determine whether there are expired small objects among the small objects stored in the first largest object and whether there are expired small objects among the second largest object that has a sequential arrangement relationship with the number of the first largest object, and obtain a judgment result.

[0064] This judgment result is used to indicate whether to perform the expired file cleaning operation.

[0065] Based on the relationship between the set file expiration time and the write time of the first small object and the write time of the last small object in the current large object, the expiration judgment of the entire large object can be realized, reducing the useless listing of a large number of small objects. At the same time, using the sequential arrangement relationship between the numbers of different large objects, the corresponding expiration judgment of other large objects that have a sequential arrangement relationship with the current large object is realized, reducing the useless listing of other large objects, further reducing the useless listing of the small objects stored in other large objects, and without the need to introduce additional sequential indexes. Using the data record content of the metadata, the expiration cleaning of small objects is realized, avoiding additional performance overhead, and ensuring the data processing performance of the distributed storage cluster.

[0066] In a specific implementation manner in this process, the determining whether there are expired small objects among the small objects stored in the first largest object and whether there are expired small objects among the second largest object that has a sequential arrangement relationship with the number of the first largest object, and obtaining a judgment result includes:

[0067] If the size relationship is that the set file expiration time is earlier than the first start time, then determine that there are no expired small objects among the small objects stored in the first largest object;

[0068] And according to the sequential arrangement relationship between the numbers, obtain at least one of the second largest objects whose corresponding time of the number is equal to or later than the corresponding time of the number of the first largest object;

[0069] Determine that there are no such expired small objects among the second largest objects.

[0070] In this process, it is necessary to perform expiration judgment on large objects and small objects:

[0071] For a certain large object, the write time range of the small objects therein can be quickly obtained by reading the two data of first_time and last_time on the large object metadata. Assuming that in the expiration judgment rule, the set file expiration time is min_time, then there are the following three cases:

[0072] 1. min_time < first_time ≤ last_time;

[0073] 2. first_time ≤ min_time ≤ last_time;

[0074] 3. first_time ≤ last_time < min_time.

[0075] Among them, in the first case, all small objects in the large object have not expired and do not need to be processed. In this case, large objects with serial numbers sorted after the serial number of the current large object do not need to be traversed anymore.

[0076] Combined with Figure 3 As shown, in a specific example, for example, the management object with ID 3 is assigned to manage large object 1, large object 2, and large object 3.

[0077] If there are two expiration determination rules configured currently, the expiration time of the file is set to 3 days in rule one and 7 days in rule two. The current time is 18:00 on March 9, 2022. According to the minimum expiration time in the rules, which is 3, subtracting 3 from the current time gives 18:00 on March 6, 2022. That is, small objects before this time point may be expired, that is, all files in large object 1 and some files in large object 2 will be expired small objects. Read large object 1 and large object 2 in sequence, traverse the small objects in them, and determine whether they meet the expiration determination conditions. When traversing small objects in large object 2 with a time greater than 18:00 on March 6, 2022, small objects after this small object cannot be expired. Since large object 3 is after large object 2, its time must be greater than the last_time in large object 2, and the small objects in large object 3 have not expired. Files on other large objects with serial numbers after large object 2 also cannot be expired.

[0078] Differently, in another specific implementation, based on the size relationship, determining whether there are expired small objects in the small objects stored in the first large object and whether there are expired small objects in the second large object having a chronological arrangement relationship with the serial number of the first large object, and obtaining a judgment result, including:

[0079] If the size relationship is that the set file expiration time is within the time range corresponding to the first start time and the first end time, then in the writing order of the small objects, sequentially determine whether the writing time of the small objects in the first large object is later than the set file expiration time;

[0080] When determining the first small object with a write time not later than the expiration time of the set file from the first large object, determine the first small object and the second small object in the first large object with a write time earlier than the first small object as expired small objects;

[0081] And, according to the chronological arrangement relationship between the numbers, obtain at least one of the second large objects whose corresponding time of the number is earlier than the corresponding time of the number of the first large object;

[0082] Judge that all small objects in the second large object are expired small objects.

[0083] This process corresponds to the second case above, first_time ≤ min_time ≤ last_time.

[0084] At this time, some small objects in the large object may be expired. In this case, traverse the small objects in sequence. Since the small objects are written into the storage space of the large object in the order of arrival, the small object list also has chronological order. When traversing to a small object that does not meet the requirements of the set file expiration time, the traversal can be stopped. The subsequent small objects cannot be expired. Directly determine the first small object with a write time not later than the expiration time of the set file and the second small object in the first large object with a write time earlier than the first small object as expired small objects. And at the same time, in this case, the small objects in the large object with a number sorting before the number of the current large object do not need to be traversed anymore, and it can be directly judged that all small objects in these large objects are expired small objects.

[0085] Specifically, in an optional embodiment corresponding thereto, the determining whether there are expired small objects in the small objects stored in the first large object and determining whether there are expired small objects in the second large object having a chronological arrangement relationship with the number of the first large object based on the size relationship, and obtaining a judgment result includes:

[0086] If the size relationship is that the expiration time of the set file is within the time range corresponding to the first start time and the first end time, then in accordance with the write order of the small objects, sequentially judge whether the write time of the small objects in the first large object is later than the expiration time of the set file;

[0087] When determining the third small object in the first large object with a write time later than the expiration time of the set file, determine the third small object and the fourth small object in the first large object with a write time later than the third small object as non-expired small objects;

[0088] And according to the chronological arrangement relationship between the numbers, obtain at least one of the second large objects whose corresponding time of the number is later than the corresponding time of the number of the first large object;

[0089] Determine that there is no such expired small object in the second largest object.

[0090] This process also corresponds to the second case mentioned above, where first_time ≤ min_time ≤ last_time.

[0091] At this time, some small objects in the large object may have expired. In this case, traverse the small objects in sequence. Since the small objects are written into the storage space of the large object in the order of arrival, the list of small objects also has a time sequence. When traversing to a small object that does not meet the requirements of the set file expiration time, the traversal can be stopped. The subsequent small objects cannot expire. Directly determine the third smallest object whose write time is later than the set file expiration time and the fourth smallest object in the first largest object whose write time is later than the third smallest object as non-expired small objects. And at the same time, in this case, the small objects in the large objects whose number sorting is after the number of the current large object do not need to be traversed anymore, and it can be directly determined that there are no expired small objects in these large objects.

[0092] Further, in another specific embodiment, the metadata further includes: the statistical number of small objects with tag information; based on the size relationship, determining whether there is an expired small object in the small objects stored in the first largest object and determining whether there is an expired small object in the second largest object having a time sequence arrangement relationship with the number of the first largest object, and obtaining a judgment result, including:

[0093] If the size relationship is that the set file expiration time is later than the first end time, determine whether the statistical number is 0;

[0094] When the statistical number is not 0, according to the write order of the small objects, detect the target small object with the tag information of the expired tag from the first largest object, and determine the target small object as the expired small object;

[0095] When the statistical number is 0, then determine that there is no such expired small object in the second largest object.

[0096] This process corresponds to the third case mentioned above, where first_time ≤ last_time < min_time.

[0097] At this time, all small objects in the large object may have expired. Here, in addition to the set file expiration time, other possible expiration determination conditions need to be further considered, such as the tag information of the small objects. When the expiration determination rule is set to require both the time requirement and the tag requirement to be met, when all small objects in the large object are within the set file expiration time, it is necessary to further determine whether other features meet the conditions.

[0098] Specifically, it is necessary to perform expiration matching in combination with tag information.

[0099] Since the tag attributes of small objects belong to the metadata of small objects, if tags are configured in the expiration determination rule, it is inevitable to read the metadata of each small object in sequence for judgment. If no tags are configured in the small objects, a lot of useless reads will occur. In the embodiment of the present application, for this rule, by accessing the tags_num attribute of the large object, the number of small objects with tags on the large object can be quickly obtained. If the number is 0, there is no need to access each small object in sequence. If it is not 0, when performing object enumeration, if small objects with tag attributes whose quantity is the same as this value are enumerated, the subsequent small objects do not need to be enumerated continuously, and the judgment can be terminated in advance to avoid useless enumeration of small objects.

[0100] Furthermore, expiration matching can also be performed in combination with the prefix of small objects.

[0101] After small objects are merged and stored in a large object, there may be hundreds or thousands of small objects on a large object. If prefix matching is performed on the object names of each small object by enumeration, it will be very time-consuming. Based on this, the embodiment of the present application adopts the prefix tree algorithm. The prefix tree is a classic algorithm for quickly finding strings that match the prefix in the case of massive data. The operations of addition and query are independent of the data volume and only related to the length of the string being operated on. By reading the list of small objects in the metadata on the large object, reading each object name, constructing a prefix tree in memory, and then quickly matching the set prefix in the expiration determination rule through the prefix tree to find the small objects that meet the prefix rule.

[0102] After obtaining the judgment result, it is possible to determine whether to perform the operation of clearing expired files.

[0103] When performing the operation of clearing expired files, specifically, it can be:

[0104] If the judgment result is that all small objects on the large object are expired files and it is determined that all small objects need to be deleted, then the entire large object is deleted. If only some small objects in the large object are expired, then the small objects and the metadata information of these small objects on the large object are deleted, and at the same time, the metadata information on the large object is updated according to the following rules:

[0105] Update first_time to the write time of the current first file after some small objects are deleted; or update last_time to the write time of the current last file after some small objects are deleted; or if there are x small objects with tag information among the expired small objects, then the tag_nums in the metadata of the large object need to be updated to tag_nums minus x.

[0106] Among them, further, the metadata further includes the used storage amount; if the foregoing judgment result indicates to perform the expired file deletion operation, after performing the expired file deletion operation, it further includes:

[0107] Based on the used storage amount, obtain a hole large object with a storage utilization rate less than a threshold from the large object queue, and at least one remaining small object is stored in the storage space corresponding to the hole large object;

[0108] Apply to the object cluster for a replacement large object with the same number as the hole large object, and the storage amount of the space corresponding to the replacement large object is equal to the total file amount of the remaining small objects;

[0109] In accordance with the writing order of each of the remaining small objects in the hole large object, migrate the remaining small objects to the storage space corresponding to the replacement large object respectively;

[0110] Delete the hole large object.

[0111] In this process, in the case of multiple small objects whose expired small objects in a large object have discontinuous storage addresses, for example, the expired rule configured with prefix or label information in the expired judgment rule causes only the small objects matching the rule to be deleted when deleting small objects in the large object, so a large object hole phenomenon will occur. The solution process of the present invention for this phenomenon is as follows:

[0112] During the process of deleting or batch deleting small objects, the use_size field in the large object record will be dynamically maintained. When deleting a file, the size of the deleted file needs to be subtracted from this field. After the expired small objects are deleted, if it is determined that the ratio of the value of use_size in the large object to the actual storage space size of the current large object is less than a pre-set ratio threshold, it is judged that the large object needs to be hole processed.

[0113] The large object needs to be added to the task queue of the hole processing thread. If in_hole_queue in the large object metadata is false, it means it has not been queued yet. At this time, set its value to true, and then write it into the task queue object of the hole processing thread in the form of a group of key-value, where the key is the current timestamp, ensuring that the first written is processed first and the newly queued is processed later, and reserve the deletion time of the index information for the newly queued object. If in_hole_queue is true, it means it has been queued, and no additional operation needs to be repeated.

[0114] Combined with Figure 4As shown in the figure, small objects File1, File2, File3, File4, and File5 are stored in large object_1. Expired small objects File3 and File4 are deleted. At this time, the space utilization rate in large object_1 is less than the threshold, so large object_1 needs to be processed for holes. The large object_1 after deleting the expired small objects is added to the task queue for hole processing for hole processing.

[0115] A timed thread can be set in the background to read the elements in the task queue object during the specified time period and perform hole processing on them in sequence. The processing method is to extract the information of the remaining small objects and rewrite them into a new large object with the same name in the writing order of the small objects, ensuring that the order of the small objects after rewriting into the large object still meets the timing requirements. The same name of the large object is to ensure that the subsequent expiration judgment and object traversal of the large object still meet the timing requirements. The storage space size of the new large object is the sum of the sizes of the remaining small objects, and the newly written large object in_hole_queue is reset to false. If all the small objects on the large object to be processed have been deleted, the large object is directly deleted to release space. After the hole processing is completed, the large object is removed from the task queue object.

[0116] In the embodiment of the present application, by obtaining the object names of large objects sorted according to time by number, accessing the metadata of large objects in the large object queue in sequence, and based on the relationship between the set file expiration time and the writing time of the first small object and the writing time of the last small object in the current large object, the expiration judgment of the large object as a whole is realized, reducing the useless enumeration of a large number of small objects. At the same time, by using the sequential arrangement relationship of the numbers between different large objects, the corresponding expiration judgment of other large objects with a sequential arrangement relationship for the current large object is realized, reducing the useless enumeration of other large objects, and no additional timing index needs to be introduced. Without increasing additional overhead, using the data record content of the metadata, the expiration clearing of small objects is realized, avoiding additional performance overhead, and ensuring the data processing performance of the distributed storage cluster.

[0117] See Figure 5 , Figure 5 is a structural diagram of a storage object processing device provided by an embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the present application are shown.

[0118] The storage object processing device 500 includes:

[0119] A first acquisition module 501, configured to acquire the object names of each large object in the large object queue, where the object name corresponds to the number of the large object, and the numbers between the large objects are arranged in sequence according to time series;

[0120] A second acquisition module 502, configured to sequentially access the metadata of large objects in the large object queue based on the object name, and acquire a first start time and a first end time from the metadata of the first large object; wherein, the metadata of each large object includes a start time and an end time, the start time corresponds to the write time of the first small object, and the end time corresponds to the write time of the last small object;

[0121] A third acquisition module 503, configured to acquire the magnitude relationship between the set file expiration time and the first start time and the first end time.

[0122] A judgment module 504, configured to judge whether there are expired small objects in the small objects stored in the first large object and whether there are expired small objects in the second large object having a sequential arrangement relationship with the number of the first large object based on the magnitude relationship, and obtain a judgment result; the judgment result is used to indicate whether to perform an expired file clearing operation.

[0123] Wherein, the device further includes:

[0124] A file writing module, configured to:

[0125] Based on the small object to be stored, when it is queried that the currently available large object is full, construct a target object name based on the number of the currently available large object;

[0126] Apply to the object cluster for a target large object corresponding to the target object name; the target object name includes the number of the target large object, and the numbers of the target large object and the currently available large object are arranged sequentially;

[0127] Write the small object into the storage space corresponding to the target large object.

[0128] Wherein, the judgment module 504 is specifically configured to:

[0129] If the magnitude relationship is that the set file expiration time is earlier than the first start time, judge that there are no expired small objects in the small objects stored in the first large object;

[0130] And, according to the sequential arrangement relationship between the numbers, acquire at least one of the second large objects whose corresponding time of the number is equal to or later than the corresponding time of the number of the first large object;

[0131] Judge that there are no expired small objects in the second large object.

[0132] Wherein, the judgment module 504 is specifically configured to:

[0133] If the size relationship is such that the set file expiration time is within the time range corresponding to the first start time and the first end time, then in accordance with the writing order of the small objects, sequentially determine whether the writing time of the small objects in the first large object is later than the set file expiration time;

[0134] When determining the first small object in the first large object whose writing time is not later than the set file expiration time, determine the first small object and the second small objects in the first large object whose writing time is earlier than the first small object as expired small objects;

[0135] And, according to the timing arrangement relationship between the numbers, obtain at least one second large object whose corresponding time of the number is earlier than the corresponding time of the number of the first large object;

[0136] Judge that all the small objects in the second large object are expired small objects.

[0137] Among them, the judging module 504 is specifically used for:

[0138] If the size relationship is such that the set file expiration time is within the time range corresponding to the first start time and the first end time, then in accordance with the writing order of the small objects, sequentially determine whether the writing time of the small objects in the first large object is later than the set file expiration time;

[0139] When determining the third small object in the first large object whose writing time is later than the set file expiration time, determine the third small object and the fourth small objects in the first large object whose writing time is later than the third small object as non-expired small objects;

[0140] And according to the timing arrangement relationship between the numbers, obtain at least one second large object whose corresponding time of the number is later than the corresponding time of the number of the first large object;

[0141] Judge that there are no such expired small objects in the second large object.

[0142] Among them, the metadata further includes: the statistical number of small objects with label information; the judging module 504 is specifically used for:

[0143] If the size relationship is such that the set file expiration time is later than the first end time, then judge whether the statistical number is 0;

[0144] When the statistical number is not 0, in accordance with the writing order of the small objects, detect from the first large object the target small object whose label information is the expired label, and determine the target small object as the expired small object;

[0145] When the statistical count is 0, it is determined that there is no such expired small object in the second largest object.

[0146] Among them, the metadata further includes the used storage capacity; the apparatus further includes:

[0147] A storage release module, configured to:

[0148] If the judgment result indicates to perform the expired file clearing operation, based on the used storage capacity, obtain a hollow large object with a storage utilization rate less than a threshold from the large object queue, where at least one remaining small object is stored in the storage space corresponding to the hollow large object;

[0149] Apply to the object cluster for a replacement large object with the same number as the hollow large object, where the storage capacity of the space corresponding to the replacement large object is equal to the total file amount of the remaining small objects;

[0150] Migrate the remaining small objects to the storage space corresponding to the replacement large object respectively according to the writing order of each of the remaining small objects in the hollow large object;

[0151] Delete the hollow large object.

[0152] The storage object processing apparatus provided by the embodiments of the present application can implement each process of the embodiments of the above storage object processing method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0153] Figure 6 It is a structural diagram of a terminal provided by an embodiment of the present application. As shown in this figure, the terminal 6 of this embodiment includes: at least one processor 60 ( Figure 6 only one is shown), a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the steps in any of the above method embodiments are implemented.

[0154] The terminal 6 may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 6 this is only an example of the terminal 6 and does not constitute a limitation on the terminal 6. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the terminal may further include input / output devices, network access devices, a bus, etc.

[0155] The processor 60 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0156] The memory 61 may be an internal storage unit of the terminal 6, such as the hard disk or memory of the terminal 6. The memory 61 may also be an external storage device of the terminal 6, such as a plug-in hard disk equipped on the terminal 6, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 61 may also include both the internal storage unit of the terminal 6 and the external storage device. The memory 61 is used to store the computer program and other programs and data required by the terminal. The memory 61 may also be used to temporarily store data that has been output or is to be output.

[0157] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated units may be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0158] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0159] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0160] In the embodiments provided in this application, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0161] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0162] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0163] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0164] All or part of the processes in the above-described embodiment methods of this application can also be implemented through a computer program product. When the computer program product runs on a terminal, the terminal can be made to implement the steps in the above-described method embodiments when executed.

[0165] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for processing stored objects, characterized in that, comprising: Obtaining the object names of each large object in the large object queue, where the object name corresponds to the number of the large object, and the numbers of each of the large objects are arranged in sequence according to time series; Based on the object name, sequentially accessing the metadata of the large objects in the large object queue, and obtaining a first start time and a first end time from the metadata of the first large object; wherein, the metadata of each of the large objects includes a start time and an end time, the start time corresponds to the write time of the first small object, and the end time corresponds to the write time of the last small object; Obtaining the size relationship between the set file expiration time and the first start time and the first end time; Based on the size relationship, judging whether there are expired small objects in the small objects stored in the first large object and judging whether there are expired small objects in the second large object having a time series arrangement relationship with the number of the first large object, to obtain a judgment result; the judgment result is used to indicate whether to perform an expired file deletion operation.

2. The method according to claim 1, characterized in that, before obtaining the object names of each large object in the large object queue, further comprising: Based on the small object to be stored, when it is queried that the currently available large object is full, constructing a target object name based on the number of the currently available large object; Applying to the object cluster for the target large object corresponding to the target object name; the target object name includes the number of the target large object, and the number of the target large object is arranged in sequence according to time series with the number of the currently available large object; Writing the small object into the storage space corresponding to the target large object.

3. The method according to claim 1, characterized in that, the judging whether there are expired small objects in the small objects stored in the first large object and judging whether there are expired small objects in the second large object having a time series arrangement relationship with the number of the first large object, to obtain a judgment result based on the size relationship, includes: If the size relationship is that the set file expiration time is earlier than the first start time, judging that there are no expired small objects in the small objects stored in the first large object; and, according to the time series arrangement relationship between the numbers, obtaining at least one of the second large objects whose corresponding time of the number is equal to or later than the corresponding time of the number of the first large object; Judging that there are no such expired small objects in the second large object.

4. The method according to claim 1, characterized in that, the judging whether there are expired small objects in the small objects stored in the first large object and judging whether there are expired small objects in the second large object having a time series arrangement relationship with the number of the first large object, to obtain a judgment result based on the size relationship, includes: If the size relationship is that the set file expiration time is within the time range corresponding to the first start time and the first end time, sequentially judging whether the write time of the small objects in the first large object is later than the set file expiration time according to the write order of the small objects; When determining the first small object whose write time in the first large object is not later than the expiration time of the set file, determine the first small object and the second small object in the first large object whose write time is earlier than that of the first small object as expired small objects; And, according to the sequential arrangement relationship between the numbers, obtain at least one of the second large objects whose corresponding time of the number is earlier than the corresponding time of the number of the first large object; Judge that all the small objects in the second large object are expired small objects.

5. The method according to claim 1, wherein, Based on the size relationship, judging whether there are expired small objects in the small objects stored in the first large object and judging whether there are expired small objects in the second large objects having a sequential arrangement relationship with the number of the first large object, and obtaining a judgment result, including: If the size relationship is that the expiration time of the set file is within the time range corresponding to the first start time and the first end time, then in accordance with the write order of the small objects, sequentially judge whether the write time of the small objects in the first large object is later than the expiration time of the set file; When determining the third small object in the first large object whose write time is later than the expiration time of the set file, determine the third small object and the fourth small object in the first large object whose write time is later than that of the third small object as non-expired small objects; And, according to the sequential arrangement relationship between the numbers, obtain at least one of the second large objects whose corresponding time of the number is later than the corresponding time of the number of the first large object; Judge that there are no such expired small objects in the second large object.

6. The method according to claim 1, wherein, The metadata further includes: the statistical number of small objects with tag information; based on the size relationship, judging whether there are expired small objects in the small objects stored in the first large object and judging whether there are expired small objects in the second large objects having a sequential arrangement relationship with the number of the first large object, and obtaining a judgment result, including: If the size relationship is that the expiration time of the set file is later than the first end time, then judge whether the statistical number is 0; When the statistical number is not 0, detect the target small object with the expired tag information from the first large object in accordance with the write order of the small objects, and determine the target small object as an expired small object; When the statistical number is 0, then judge that there are no such expired small objects in the second large object.

7. The method according to claim 1, wherein, The metadata further includes the used storage capacity; if the judgment result indicates to perform the expired file cleaning operation, then after performing the expired file cleaning operation, it further includes: Based on the used storage capacity, obtain a hollow large object with a storage utilization rate less than the threshold from the large object queue, and at least one remaining small object is stored in the storage space corresponding to the hollow large object; Apply to the object cluster for a replacement large object with the same number as the hollow large object, and the storage capacity corresponding to the replacement large object is equal to the total file amount of the remaining small objects; Migrate the remaining small objects to the storage space corresponding to the replacement large object respectively according to the writing order of each of the remaining small objects in the hole large object; Delete the hole large object.

8. A processing device for storage objects, characterized in that, it includes: A first acquisition module, configured to acquire the object names of each large object in the large object queue, where the object name corresponds to the number of the large object, and the numbers of each of the large objects are arranged in sequence according to time series; A second acquisition module, configured to access the metadata of the large object in the large object queue in sequence based on the object name, and acquire a first start time and a first end time from the metadata of the first large object; wherein, the metadata of each large object includes a start time and an end time, the start time corresponds to the writing time of the first small object, and the end time corresponds to the writing time of the last small object; A third acquisition module, configured to acquire the magnitude relationship between the set file expiration time and the first start time and the first end time; A judgment module, configured to judge whether there are expired small objects in the small objects stored in the first large object and whether there are expired small objects in the second large object having a time series arrangement relationship with the number of the first large object based on the magnitude relationship, and obtain a judgment result; the judgment result is used to indicate whether to perform an expired file cleaning operation.

9. A terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Object deleting method and device and data processing method and device

    CN111966867A

  • Object storage data storage management method, device and equipment

    CN114153392A