Distributed log increment reconstruction method and device based on duplicate removal and active elimination
By adopting distributed log incremental reconstruction methods of deduplication and active phase-out in the DAOS storage system, the problem that DAOS cannot perform incremental reconstruction during failure is solved, and efficient data reconstruction and storage space saving is achieved.
Patent Information
- Application Number
- CN202510124056.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-13
AI Technical Summary
DAOS is based on a multi-version architecture. If GC and aggregation occur during a failure, incremental reconstruction cannot be performed, resulting in full reconstruction only after a long-term failure, which is inefficient and time-consuming. At the same time, log-based methods such as Ceph have problems such as large log volume, large performance impact and lack of active elimination mechanisms.
A distributed log incremental reconstruction method based on deduplication and active phase-out is proposed. By generating a global hash table when the storage system is started, recording distributed logs and deduplication processing, persisting them to disk, and selecting incremental or full reconstruction based on log integrity during failure recovery.
It realizes normal GC and aggregation during a long-term failure of the DAOS storage pool. In the event of failure recovery, incremental reconstruction can be achieved by recording distributed logs, which improves data reconstruction efficiency, shortens the reconstruction time, and reduces log writing and storage space usage.
Smart Images

Figure CN119988098A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of distributed log incremental reconstruction, and in particular, relates to a distributed log incremental reconstruction method, device, computer-readable storage medium, and electronic device based on deduplication and active elimination. Background Art
[0002] Currently, Ceph's aggregation and garbage collection (GC) insert a new data (entry) into the distributed RocksDB. Each time an entry is inserted, a new log (log) will be added to the corresponding PostgreSQL (PG). During fault recovery, incremental reconstruction is achieved based on the comparison of the PG log.
[0003] The distributed storage of competitors is based on an append-write architecture, where newly inserted entries for aggregation and GC are newly written data. Each disk knows which objects are available, and whenever an object is degraded, it switches to a new object, so incremental reconstruction can be achieved by restoring the missing objects on the disk.
[0004] DAOS is based on a multi-version architecture, and its incremental reconstruction is also based on versions. Aggregation and GC modify the data of the old version. As long as aggregation and GC exist during the failure, the failed pool will not be able to perform incremental reconstruction.
[0005] The above scheme has the following defects and shortcomings: (1) Both Ceph and other vendors' append-write architectures can achieve incremental reconstruction, but DAOS is based on a multi-version architecture. If GC and aggregation occur during a failure, incremental reconstruction cannot be performed. Incremental reconstruction can be achieved by prohibiting GC and aggregation during a pool failure, but GC and aggregation cannot be performed at the same time, which will result in garbage collection and continuous capacity increase. Therefore, after a long-term failure, only full reconstruction can be performed, which is inefficient and time-consuming.
[0006] (2) For log-based methods such as Ceph, each change corresponds to a log, which requires a large amount of logs to be recorded and has a significant impact on the performance of the original process.
[0007] (3) Log-based methods such as Ceph do not have an effective active elimination mechanism, and the storage space on the disk required to record the log is relatively large. Summary of the invention
[0008] In order to address the above problems, this application proposes a new distributed log incremental reconstruction method based on deduplication and active elimination.
[0009] This application method is designed for distributed logs. If the logs are only recorded locally, when the local disk fails, the user data recovery needs to be incrementally reconstructed, and the logs required for data recovery cannot be read. Like metadata, recording logs in multiple copies can achieve the same reliability as user data storage.
[0010] In order to achieve the above objectives, this application provides the following technical solutions: A first aspect of the present application provides a distributed log incremental reconstruction method based on deduplication and active elimination, the method comprising: When the storage system starts a GC or aggregation process, it first checks whether the process contains operations that modify metadata (such as inserting, deleting, and other operations). If it does, it uses a deduplication method to record distributed logs and persists the contents of the distributed logs to disk. Determine whether the distributed log is recorded successfully. If successful, continue to execute the original GC or aggregation process. If failed, terminate the metadata modification of the GC or aggregation process and report an error. When a pool fails and needs to be restored, check whether the write-before log during the failure period is complete. If it is complete, choose to perform incremental reconstruction; if it is incomplete, all objects are fully reconstructed; Scan all disks in the fault pool to obtain the full set of objects on the faulty disk that need to be reconstructed and restored; read the distributed log data within the fault time interval; fully reconstruct the objects that record the distributed logs, and incrementally reconstruct other objects.
[0011] Optionally, the method of the present application also includes: if it is checked that the GC or aggregation process does not contain an operation to modify metadata (such as only a query operation), the distributed log is not recorded and the original GC or aggregation process is directly executed.
[0012] Optionally, in the method of the present application, the checking of whether the write-before log during the failure period is complete includes: checking whether the effective time interval of the distributed log completely includes the time interval of the failure. If it does not include it, it means that the data recorded in the distributed log is incomplete, and all objects need to be fully reconstructed at this time; if it does include it, it means that the data recorded in the distributed log is complete, and incremental reconstruction is performed based on the distributed log data.
[0013] Optionally, in the method of the present application, the incomplete data recorded in the distributed log is caused by the following reasons: the function of recording distributed logs has been turned off; there is too much data in the recorded log that is eliminated due to insufficient space; other failures in the storage system cause log recording to fail.
[0014] Optionally, in the method of the present application, the adopting a deduplication method to record distributed logs and persisting the content of the distributed logs to a disk includes: (1) When the storage system is started, each subsystem (such as a process) generates its own global hash_table (hash table); when processing GC or aggregation requests in batches, a cache linked list is generated. When encountering a situation where distributed logs need to be recorded (there are operations to modify metadata in the GC or aggregation process), the object and shard information corresponding to each request is converted into the distributed log information that needs to be recorded; (2) Compare each newly added distributed log with the data in the cache linked list. If the data is duplicate, the corresponding request is returned directly without recording the distributed log. If the data is not duplicate, proceed to the next step. (3) The murmur64 algorithm is used to deduplicate the distributed log information and hash_table. If the data is duplicate, the corresponding request is returned directly without recording the distributed log. If the data is not duplicate, the distributed log data is inserted into the cache linked list. (4) Persist the cache list data in batches to disk; (5) The successfully downloaded data is then inserted into the current hash_table for subsequent deduplication of distributed log data.
[0015] Optionally, when recording distributed logs, the method of the present application caches part of the log data persisted on the disk into the memory in chronological order. Since the hash_table corresponds to a version number, the repeated distributed log data will only be persisted on the disk once during the validity period of the memory data, thereby removing most of the repeated data and saving disk space.
[0016] Optionally, the method of the present application also includes an active elimination process, the steps of which are as follows: (1) The normal writing process of the distributed log automatically checks whether the current writing volume exceeds the specification limit. If it exceeds the limit, it triggers elimination and deletes the data with the oldest version number under this process; (2) Determine whether the flushing of the distributed log is successful. If the return value is full, it triggers elimination and deletes the data with the oldest version number under this process; (3) The actual write volume of the distributed log on each disk is counted, and the capacity is periodically checked to see if it exceeds the threshold. If the threshold is exceeded, active elimination is triggered to delete the data with the oldest version number under each process; (4) When the pool version number changes, the stable pool version number is actively pushed to all processes, and the fault version numbers of all disks are periodically obtained. The smaller value of the pool version number and the disk fault version number is used as the trim_epoch data version number that can be eliminated (distributed logs are recorded based on the epoch version number, and trim_epoch represents the maximum version number that can be eliminated in the log). Distributed log data that is less than trim_epoch is actively deleted; (5) Supports manual triggering of elimination. Operation and maintenance personnel can delete the data with the oldest version number under each process by executing the corresponding command of the interface.
[0017] A second aspect of the present application provides a distributed log incremental reconstruction device based on deduplication and active elimination, the device comprising: Logging and storage module: When the storage system starts a GC or aggregation process, it first checks whether the process contains operations to modify metadata. If so, it uses a deduplication method to record distributed logs and persists the contents of the distributed logs to disk. Task (GC or aggregation) process execution module: used to determine whether the distributed log is recorded successfully. If successful, the original GC or aggregation process will continue to be executed; if failed, the metadata modification of the GC or aggregation process will be terminated and an error will be reported; Log check module: when a pool fails and needs to be restored, it checks whether the write-before log during the failure period is complete. If it is complete, incremental reconstruction is performed; if it is incomplete, all objects are fully reconstructed. Incremental reconstruction module: used to scan all disks in the fault pool and obtain the full set of objects on the faulty disk that need to be reconstructed and restored; read the distributed log data within the fault time interval; fully reconstruct the objects that record the distributed logs, and incrementally reconstruct other objects.
[0018] The device implements the steps of the aforementioned distributed log incremental reconstruction method based on deduplication and active elimination when running.
[0019] A third aspect of the present application provides an electronic device, comprising: a memory and a processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the aforementioned distributed log incremental reconstruction method based on deduplication and active elimination.
[0020] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the aforementioned distributed log incremental reconstruction method based on deduplication and active elimination are implemented.
[0021] In summary, this application proposes a new distributed log incremental reconstruction method based on deduplication and active elimination, which has the following advantages: (1) This method is based on a multi-version storage architecture. GC and aggregation can be performed normally during a long-term storage pool failure. During failure recovery, incremental reconstruction can be achieved by recording distributed logs. The data reconstruction efficiency is high, which greatly shortens the reconstruction time.
[0022] (2) The use of a log deduplication mechanism can greatly reduce the amount of log writing and reduce the impact on GC and aggregation processes.
[0023] (3) By adopting multiple active elimination mechanisms for log data on the disk, the storage space on the disk is greatly saved.
[0024] Other features and advantages of the present application will be described in the following description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the techniques indicated in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0026] Figure 1 This is a schematic diagram of the implementation principle of incremental reconstruction based on distributed write-ahead log in the method of this application.
[0027] Figure 2 This is a diagram of the deduplication principle of the distributed write-ahead log in the method of this application.
[0028] Figure 3 This is a schematic diagram of the active elimination mechanism of the distributed write-ahead log in the present application method.
[0029] Figure 4 This is the overall implementation flow chart of the distributed log incremental reconstruction method based on deduplication and active elimination in this application.
[0030] Figure 5 This is a composition structure diagram of the distributed log incremental reconstruction device based on deduplication and active elimination in this application.
[0031] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0033] As used herein, the term “including” and its variations are open inclusions, ie, “including but not limited to”; the term “based on” means “based at least in part on”; and the term “one embodiment” means “at least one embodiment”.
[0034] It should be noted that the modifications of "one" and "plurality" mentioned in the present application are illustrative rather than restrictive, and those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".
[0035] Terminology explanation: Ceph: is an open source distributed storage system that provides a variety of storage services such as object storage, block storage, and file systems. It has the characteristics of high performance, high availability, and scalability.
[0036] RocksDB: An LSM-tree architecture engine developed by Facebook based on LevelDB that provides key-value storage and read-write capabilities.
[0037] DAOS: is an open source distributed asynchronous object storage system designed for large-scale distributed non-volatile memory (NVM), leveraging next-generation NVM technologies such as SCM (Storage-Class Memory) and NVMe (Non-Volatile Memory express). DAOS provides a key-value storage interface and features transactional non-blocking I / O, advanced data protection, self-healing, end-to-end data integrity, fine-grained data control, and elastic storage, aiming to optimize performance and cost.
[0038] Figure 4 The overall implementation process of the distributed log incremental reconstruction method based on deduplication and active elimination provided by the present application is shown, including the following steps: When the storage system starts a GC or aggregation process, it first checks whether the process contains operations that modify metadata (such as inserting, deleting, and other operations). If it does, it uses a deduplication method to record distributed logs and persists the contents of the distributed logs to disk. Determine whether the distributed log is recorded successfully. If successful, continue to execute the original GC or aggregation process. If failed, terminate the metadata modification of the GC or aggregation process and report an error. When a pool fails and needs to be restored, check whether the write-before log during the failure period is complete. If it is complete, choose to perform incremental reconstruction; if it is incomplete, all objects are fully reconstructed; Scan all disks in the fault pool to obtain the full set of objects on the faulty disk that need to be reconstructed and restored; read the distributed log data within the fault time interval; fully reconstruct the objects that record the distributed logs, and incrementally reconstruct other objects.
[0039] Specifically, in order to achieve the purpose of incremental reconstruction, this application designs two processing strategies: deduplication and active elimination.
[0040] At present, in a multi-version storage system, after the GC and aggregation processes modify the old version of the data, incremental reconstruction cannot be performed during fault recovery. This solution records which objects have undergone GC and aggregation during the fault period (that is, objects that have modified the old version of the data) by recording distributed logs. After a disk failure, when recovering data, a full reconstruction is performed on the small number of objects that have recorded distributed logs (that is, all the data of these objects needs to be reconstructed), while other object data that needs to be reconstructed and recovered can only be incrementally reconstructed (that is, only the last version of the data after the failure needs to be reconstructed, and the data before the failure is complete and does not need to be reconstructed), thereby achieving the purpose of overall incremental reconstruction. Incremental reconstruction can greatly shorten the time for data reconstruction and reduce the waste of resources such as bandwidth caused by unnecessary data reconstruction and recovery.
[0041] Furthermore, in the solution of the present application, a deduplication and active elimination mechanism is adopted in the process of continuously recording distributed logs in the storage system.
[0042] First, since the GC and aggregation processes often operate on the same objects, there will be a lot of duplication when these objects are converted into distributed log data. In view of this feature, when recording distributed logs, this solution caches part of the log data persisted on the disk into the memory in chronological order. During the validity period of the memory data, duplicate data is only persisted on the disk once. This deduplication solution can greatly reduce the number of data flushes to the disk, reduce the impact of recording distributed logs on the performance of the original GC and aggregation processes, and greatly save disk space.
[0043] Secondly, since distributed logs need to be continuously written to the disk, the occupied disk capacity will become larger and larger. For this reason, this solution adopts a mechanism for actively eliminating log data to ensure that old log data that is no longer needed can be deleted in a timely manner. The active elimination solution can actively identify which distributed log data on the disk is no longer needed for subsequent fault reconstruction, so as to delete and process them in a timely manner. At the same time, it also ensures that the maximum capacity occupied by distributed log data will not exceed the planned capacity of the system. When the capacity threshold is exceeded, the elimination of the oldest log data will be triggered. In addition, this solution also supports human intervention, and operation and maintenance personnel can specify the deletion of some log data. When the storage system is running, a variety of active elimination mechanisms are combined to ensure that distributed logs can be continuously written to the disk, while the occupied capacity is controllable, and unnecessary data can be automatically identified and actively deleted.
[0044] In order to better understand the technical solution of the present application, further explanation is given with reference to embodiments of the following scenarios.
[0045] 1. Incremental reconstruction of distributed write-ahead logs based on deduplication, such as Figure 1 As shown, including: 1. When the storage system is performing GC and aggregation processes, check whether there are any operations that modify metadata.
[0046] 1.1 If the metadata is not modified (for example, only query operations are performed), there is no need to record the distributed write-before log, and the original GC and aggregation processes can be executed directly.
[0047] 1.2 If there is an operation to modify metadata, you need to record the distributed write-ahead log (hereinafter referred to as log) first. When recording the log, use the following method: Figure 2 The deduplication method shown in the figure and described in part (ii) persists the log content to the disk. If the log record is successful, the metadata is modified (that is, the current GC or aggregation process continues); if the log record fails, the metadata modification operation is terminated (that is, the current GC or aggregation process is terminated).
[0048] 2. When the pool fails and recovers, determine whether incremental reconstruction is possible.
[0049] 2.1 Since the log is recorded continuously when the storage system is running, check whether the effective time interval of the log completely includes the time interval of the failure. If it does not include it, it means that the data recorded in the log is incomplete (it may be that the function of recording logs has been turned off, or there is too much data in the recorded log that is triggered to be eliminated due to insufficient space, or there are other failures in the storage system that cause log recording failure, etc.), and only full reconstruction can be performed at this time; if it can be included, incremental reconstruction is performed based on the log data.
[0050] 3. Perform incremental reconstruction.
[0051] 3.1 Scan all disks in the fault pool to obtain the complete set of objects on the faulty disk that need to be reconstructed and restored.
[0052] 3.2 Read the log data within the fault time interval.
[0053] 3.3 The object that records log data is fully reconstructed (all data of the object needs to be reconstructed from other disks), and other objects are incrementally reconstructed (all version data of the object before the failure is complete, and only the incremental data during the failure period needs to be reconstructed).
[0054] (ii) Deduplication scheme for recording distributed write-ahead logs, such as Figure 2 As shown, including: 1. When the storage system is started, each subsystem (such as a process) generates its own global hash_table. When batch processing GC and aggregation requests, a cache linked list is generated. When encountering a situation where a distributed write-before log needs to be recorded first (there is an operation to modify metadata), the corresponding object and shard information is converted into the log information that needs to be recorded.
[0055] 2. Compare each newly added log with the data in the cache list. If it is duplicate data, the corresponding request can be returned directly without recording the log; if it is not duplicate, proceed to step 3.
[0056] 3. Use the murmur64 algorithm to quickly deduplicate the log information and hash_table. If the data is duplicate, the corresponding request can be returned directly without logging; if it is not duplicate, the data is added to the cache list.
[0057] 4. Persist the cache list data in batches to disk.
[0058] 5. Insert the successfully downloaded data into the current hash_table for subsequent log data deduplication. Since hash_table has a corresponding version number, the same log data will only be downloaded once during this period of time, effectively removing most of the duplicate data and saving disk space. Only a small amount of new data will be downloaded, so deduplication also greatly reduces the impact on GC and aggregation performance.
[0059] (III) Active elimination scheme for distributed write-ahead logs, such as Figure 3 As shown, including: 1. The normal log writing process will check whether the current writing volume exceeds the specification limit. If it exceeds the limit, it will trigger elimination and delete the oldest data in this process.
[0060] 2. Determine whether the log flushing is successful. If the return value is full, trigger the elimination of the oldest version number of the data in this process.
[0061] 3. Perform capacity statistics on the actual write volume of the log on each disk, and periodically check whether the capacity exceeds the threshold. If the threshold is exceeded, the oldest version number of the data under each process will be actively eliminated.
[0062] 4. When the pool version number changes, the stable pool version number is actively pushed to all processes, and the fault version numbers of all disks are periodically obtained. The smaller of the two is used to obtain the data version number trim_epoch that can be eliminated. Since the data before trim_epoch will not be used in subsequent disk failure recovery and is no longer needed, log data smaller than trim_epoch is actively eliminated.
[0063] 5. Supports manual triggering to eliminate the oldest log data under each process. When the operation and maintenance personnel want to manually trigger the elimination of data in a storage container, they can execute the corresponding command of the interface.
[0064] In summary, multiple active elimination mechanisms ensure that log data that is not needed for reconstruction can be eliminated in a timely manner, thereby ensuring that the space on the disk will not be filled up, greatly saving storage space on the disk.
[0065] Figure 5 The present application shows a distributed log incremental reconstruction device based on deduplication and active elimination, the device comprising: Logging and storage module: When the storage system starts a GC or aggregation process, it first checks whether the process contains operations to modify metadata. If so, it uses a deduplication method to record distributed logs and persists the contents of the distributed logs to disk. Task (GC or aggregation) process execution module: used to determine whether the distributed log is recorded successfully. If successful, the original GC or aggregation process will continue to be executed; if failed, the metadata modification of the GC or aggregation process will be terminated and an error will be reported; Log check module: when a pool fails and needs to be restored, it checks whether the write-before log during the failure period is complete. If it is complete, incremental reconstruction is performed; if it is incomplete, all objects are fully reconstructed. Incremental reconstruction module: used to scan all disks in the fault pool and obtain the full set of objects on the faulty disk that need to be reconstructed and restored; read the distributed log data within the fault time interval; fully reconstruct the objects that record the distributed logs, and incrementally reconstruct other objects.
[0066] When the above device is running, the steps of the distributed log incremental reconstruction method based on deduplication and active elimination disclosed in this application are implemented.
[0067] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the device, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the specified function or operation, or can be realized by a combination of special hardware and computer instructions.
[0068] like Figure 6 As shown, an embodiment of the present application also discloses an electronic device, including: a processor 310, a communication interface 320, a memory 330 for storing a processor executable computer program, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 runs the executable computer program to implement the steps of the above-mentioned distributed log incremental reconstruction method based on deduplication and active elimination.
[0069] It is understandable that, in addition to the memory and the processor, the electronic device may also include an input device such as a keyboard, an output device such as a display, and other communication modules. The input device, the output device, and other communication modules communicate with the processor via an I / O interface (i.e., an input / output interface).
[0070] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0071] Furthermore, the present application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the various steps of the distributed log incremental reconstruction method based on deduplication and active elimination disclosed in the present application.
[0072] In the context of the present application, a computer-readable storage medium may be a tangible medium, more specific examples of which would include a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0073] In particular, according to an embodiment of the present application, the process described in the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the distributed log incremental reconstruction method based on deduplication and active elimination disclosed in the present application. When the computer program is executed by a processing device, the above functions defined in the method of the embodiment of the present application are executed.
[0074] Although several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present application. The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by a specific combination of the above technical features, but also should cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above public concept.
[0075] Those skilled in the art should also understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A distributed log incremental reconstruction method based on deduplication and active elimination, characterized in that: The method comprises: When the storage system starts a GC or aggregation process, it first checks whether the process contains operations to modify metadata. If so, it uses a deduplication method to record distributed logs and persists the contents of the distributed logs to disk. Determine whether the distributed log is recorded successfully. If successful, continue to execute the original GC or aggregation process. If failed, terminate the metadata modification of the GC or aggregation process and report an error. When a pool fails and needs to be restored, check whether the write-before log during the failure period is complete. If it is complete, choose to perform incremental reconstruction; if it is incomplete, all objects are fully reconstructed; Scan all disks in the fault pool to obtain the full set of objects on the faulty disk that need to be reconstructed and restored; read the distributed log data within the fault time interval; fully reconstruct the objects that record the distributed logs, and incrementally reconstruct other objects.
2. The method according to claim 1, characterized in that The method further includes: if it is checked that the GC or aggregation process does not include an operation of modifying metadata, the distributed log is not recorded, and the original GC or aggregation process is directly executed.
3. The method according to claim 1, characterized in that: The checking of whether the write-before log during the failure period is complete includes: checking whether the effective time interval of the distributed log completely includes the time interval of the failure. If not, it means that the data recorded in the distributed log is incomplete, and all objects need to be fully reconstructed; if it can be included, it means that the data recorded in the distributed log is complete, and incremental reconstruction is performed based on the distributed log data.
4. The method according to claim 3, characterized in that The incomplete data recorded in the distributed log is caused by the following reasons: the function of recording distributed logs has been turned off; there is too much data in the recorded log that is triggered to be eliminated due to insufficient space; Other failures occurred in the storage system, causing logging failure.
5. The method according to claim 1, characterized in that The method of recording distributed logs by using a deduplication method and persisting the content of the distributed logs on a disk includes: (1) When the storage system is started, each subsystem generates its own global hash_table; when processing GC or aggregate requests in batches, a cache linked list is generated. When encountering a situation where distributed logs need to be recorded, the object and shard information corresponding to each request is converted into the distributed log information that needs to be recorded; (2) Compare each newly added distributed log with the data in the cache linked list. If the data is duplicate, the corresponding request is returned directly without recording the distributed log. If the data is not duplicate, proceed to the next step. (3) The murmur64 algorithm is used to deduplicate the distributed log information and hash_table. If the data is duplicate, the corresponding request is returned directly without recording the distributed log. If the data is not duplicate, the distributed log data is inserted into the cache linked list. (4) Persist the cache list data in batches to disk; (5) The successfully downloaded data is then inserted into the current hash_table for subsequent deduplication of distributed log data.
6. The method according to claim 5, characterized in that When recording distributed logs, the method caches part of the log data persisted on the disk into the memory in chronological order. Since the hash_table corresponds to a version number, the repeated distributed log data will only be persisted on the disk once during the validity period of the memory data, thereby removing most of the repeated data and saving disk space.
7. The method according to claim 1, characterized in that The method also includes an active elimination process, the steps of which are as follows: (1) The normal writing process of the distributed log automatically checks whether the current writing volume exceeds the specification limit. If it exceeds the limit, it triggers elimination and deletes the data with the oldest version number under this process; (2) Determine whether the flushing of distributed logs is successful. If the return value is full, it triggers elimination and deletes the data with the oldest version number under this process; (3) The actual write volume of the distributed log on each disk is counted, and the capacity is periodically checked to see if it exceeds the threshold. If the threshold is exceeded, active elimination is triggered to delete the data with the oldest version number under each process; (4) When the pool version number changes, the stable pool version number is actively pushed to all processes, and the fault version numbers of all disks are periodically obtained. The smaller value of the pool version number and the disk fault version number is taken as the trim_epoch of the data that can be eliminated, and the distributed log data that is less than trim_epoch is actively deleted; (5) Supports manual triggering of elimination. Operation and maintenance personnel can delete the data with the oldest version number under each process by executing the corresponding command of the interface.
8. A distributed log incremental reconstruction device based on deduplication and active elimination, characterized in that: The device comprises: Logging and storage module: When the storage system starts a GC or aggregation process, it first checks whether the process contains operations to modify metadata. If so, it uses a deduplication method to record distributed logs and persists the contents of the distributed logs to disk. Task process execution module: used to determine whether the distributed log is recorded successfully. If successful, the current GC or aggregation process will continue to be executed; if failed, the modification of metadata by the current GC or aggregation process will be terminated and an error will be reported. Log check module: when a pool fails and needs to be restored, it checks whether the write-before log during the failure period is complete. If it is complete, incremental reconstruction is performed; if it is incomplete, all objects are fully reconstructed. Incremental reconstruction module: used to scan all disks in the fault pool and obtain the full set of objects on the faulty disk that need to be reconstructed and restored; read the distributed log data within the fault time interval; fully reconstruct the objects that record the distributed logs, and incrementally reconstruct other objects.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the distributed log incremental reconstruction method based on deduplication and active elimination are implemented as described in any one of claims 1 to 7.
10. An electronic device, characterized in that: include: Memory and processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the distributed log incremental reconstruction method based on deduplication and active elimination as described in any one of claims 1-7.