Reducing diary logs in storage system
By identifying and replacing the target segments in the journal log, new segments arranged in order of write addresses are generated, and the problem of log taking up storage space and low recovery efficiency is solved, achieving rapid data recovery.
Patent Information
- Application Number
- CN202410811859.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2024-06-21
- Publication Date
- 2025-07-18
AI Technical Summary
Diary logs continue to expand with the increase in the number of writes in the storage system, occupying storage space, resulting in reduced data recovery efficiency and being vulnerable to corruption. It is difficult for the prior art to effectively reduce the size of the diary log without losing useful information.
By identifying the target segments in the journal log, determining and recording a subset of the journal entries for the most recent write operation, generating a new segment to replace the target segment, and arranging the new segments in the order of write addresses, reducing the size of the journal log and retaining useful information.
It realizes reducing the size of the journal log without losing useful information, improving data recovery speed and efficiency, and ensuring that data can be recovered quickly when damaged.
Smart Images

Figure CN120336078A_ABST
Abstract
Description
Background Art
[0001] A computing device may include components such as a processor, a memory, a cache system, and a storage device. The storage device may include a hard disk drive that uses magnetic media to store and retrieve data blocks. Some storage systems may transfer data between different locations or devices. For example, certain systems may transfer and store copies of important data for archival and recovery purposes. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Some embodiments are described with reference to the following drawings.
[0003] Figure 1 is a schematic diagram of an example storage system according to some embodiments.
[0004] Figure 2 is an illustration of an example process according to some embodiments.
[0005] Figures 3A - 3D is an illustration of an example operation according to some embodiments.
[0006] Figures 4A - 4B is an illustration of an example data structure according to some embodiments.
[0007] Figure 5 is an illustration of an example process according to some embodiments.
[0008] Figures 6A - 6C is an illustration of an example operation according to some embodiments.
[0009] Figure 7 is a schematic diagram of an example computing device according to some embodiments.
[0010] Figure 8 is an illustration of an example machine-readable medium storing instructions according to some embodiments.
[0011] Figure 9 is an illustration of an example process according to some embodiments.
[0012] Throughout the drawings, like reference numerals refer to similar but not necessarily identical elements. The drawings are not necessarily to scale, and the sizes of some parts may be exaggerated to more clearly illustrate the examples shown. Additionally, the drawings provide examples and / or embodiments consistent with the specification; however, the description is not limited to the examples and / or embodiments provided in the drawings. DETAILED DESCRIPTION
[0013] In this disclosure, the use of the terms "a", "an", or "the" is also intended to include the plural forms unless the context clearly dictates otherwise. Additionally, the terms "comprise", "comprising", "have", or "having" when used in this disclosure specify the presence of the stated elements but do not preclude the presence or addition of other elements.
[0014] In some examples, a computing system may persistently store data in one or more storage devices. For example, a server may store a data set on a local storage array or may store a backup copy of the data set in a remote backup device. In some examples, the backup copy may be stored in a different form than the data set. For example, the backup copy may include a deduplicated representation of the data set. As used herein, a "storage system" may include a storage device or an array of storage devices. The storage system may also include one or more storage controllers that manage access to the storage device(s). A "data unit" may refer to any portion of data that can be individually identified within the storage system. In some cases, a data unit may refer to a data block, a collection of data blocks, or any other portion of data. In some examples, the storage system may store data units in a persistent storage device. The persistent storage device may be implemented using one or more persistent (e.g., non-volatile) storage devices such as one or more disk-based storage devices (e.g., hard disk drives (HDDs)), one or more solid-state devices (SSDs) (such as flash memory devices, etc.), or a combination thereof. As used herein, a "controller" may refer to a hardware processing circuit that may include any one or a combination of a microprocessor, a core of a multi-core microprocessor, a microcontroller, a programmable integrated circuit, a programmable gate array, a digital signal processor, or other hardware processing circuits. Alternatively, a "controller" may refer to a combination of a hardware processing circuit and machine-readable instructions (software and / or firmware) executable on the hardware processing circuit.
[0015] In some examples, a data set may be stored on a block-based storage system. As used herein, a "block-based" storage system may refer to a system that stores data in the form of data blocks (also referred to herein as "block level"). In some examples, the block level may be the level at which a block-based storage device (e.g., a hard disk drive (HDD), a solid-state drive (SSD), etc.) or a virtual volume may store data. A block-based storage device may receive data blocks that make up a data set as a stream of data blocks.
[0016] In some examples, a journal log can provide continuous data protection (CDP) for a storage system. The journal log can be implemented as a linked sequence of segments, where each segment includes multiple entries storing copies or details of block-level writes performed on the storage system. For example, each entry can record a data block written to a storage volume, as well as the (multiple) storage addresses (e.g., offset and length in the storage volume) where the data block was written. In this way, the journal log can form a history of all data written to the storage system. Additionally, the journal log can include data markers (referred to herein as "checkpoints") indicating or representing various points in time. If the stored data is corrupted (e.g., by a malware attack), the entries of the journal log can be read and then used to reconstruct the stored data because it existed at the time points represented by the checkpoints. However, as the number of writes increases over time, the number of entries in the journal log also increases. Consequently, the size of the journal log can also increase over time and can fill up the space available for storing the journal log. Therefore, the amount of data that the journal log can protect can be limited by the maximum storage space available for the journal log.
[0017] According to some embodiments of the present disclosure, a controller can perform operations to reduce the size of the journal log. The journal reduction operations can include identifying suitable target segments in the journal log. For example, the target segments can be identified based on age (e.g., not too old and not too new), checkpoint duration (e.g., not too short and not too long), etc. The journal reduction operations can also include: determining a set of storage locations modified by the write operations recorded in the target segments, and identifying the most recent write operations for each modified storage location. Additionally, the journal reduction operations can include: identifying a subset of journal entries in the target segments that record the identified most recent write operations, and then generating a new segment that includes only the identified subset of journal entries. Then, the generated new segment can replace the target segment in the journal log. In this way, journal entries recording superseded writes (i.e., writes performed for a given location and superseded by subsequent writes for the same location) can be deleted from the journal log. Thus, the size of the journal log can be reduced without losing useful information. Additionally, in some embodiments, generating the new segment can include sorting the subset of journal entries according to the storage addresses (i.e., in the order of the storage locations changed by the writes recorded in the journal entries). Therefore, if the new segment is used to restore or recover the original data, the entries in the segment can be read or written according to the address order of the storage volume. In this way, the journal reduction operations can provide faster restoration or recovery of the original data. The various aspects of the disclosed technology are further discussed below with reference to Figures 1 - 9 further discussion.
[0018] Figure 1 —— Example Storage System
[0019] Figure 1 FIG. 2 illustrates an example storage system 100 in accordance with some embodiments, which includes a computing device 110 and a storage device 140. The computing device 110 may include a storage engine 120 to generate write operations and / or transmit write operations to the storage device 140. For example, the storage engine 120 may receive an input data stream (“input”), and in response may send block-based writes to the storage device 140. The input may specify file system operations (e.g., add a new file, delete an existing directory, move an existing file, etc.). The write operations may cause the stored data 125 (e.g., data and / or metadata blocks) to be written to a specified address or location in a particular volume 145 of the storage device 140.
[0020] In some embodiments, a journaling engine 130 may generate or update a journal 135, and may store some or all of the journal 135 in the storage device 140. The journal 135 may include a plurality of segments 137. Each segment 137 may store a specified size or number of journal entries. In some embodiments, each journal entry may record information about different write operations performed by the storage engine 120. For example, each entry of the journal 135 may record the data blocks written in the operation, as well as the (multiple) storage addresses (e.g., offsets and lengths in the storage volume 145) to which the data blocks were written. In some embodiments, the journal 135 may be implemented as a sequence of linked segments, where each segment includes a plurality of entries. Additionally, the journal 135 may include checkpoints indicating various points in time. Referring below to Figure 2 、 Figures 3A - 3D and Figures 4A - 4B describes an example process for generating the journal 135.
[0021] In some embodiments, when some entries of the journal 135 reach a maximum age, those entries may be removed from the journal 135. Additionally, the removed entries may be used to generate a mirror volume 155. The mirror volume 155 may be a copy of the storage volume 145 as it existed when those removed entries were added to the journal 135. For example, the writes recorded in the removed entries may be applied (e.g., executed) to the mirror volume 155 in the order in which they were recorded.
[0022] In some embodiments, journal log 135 can be used to reconstruct storage volume 145 as it existed at the time point represented by the checkpoint. For example, in the case where storage volume 145 is corrupted or lost (e.g., due to a device failure or malware attack), the writes recorded in journal entries prior to a specific checkpoint can be executed in the order in which they were recorded to reconstruct storage volume 145 as it existed at the time represented by that specific checkpoint. In some examples, such a reconstruction operation can include performing previous writes (e.g., those recorded in journal entries prior to a specific checkpoint) on mirror volume 155.
[0023] In some embodiments, journal log engine 130 can perform operations to reduce the size of journal log 135. For example, journal log engine 130 can identify a target segment 137 in journal log 135, and can identify the locations in storage volume 145 that were modified by the write operations recorded in the target segment. Additionally, journal log engine 130 can determine the most recent write operation for each modified location, and can identify a subset of journal entries in target segment 137 that record the identified most recent write operations. Then, journal log engine 130 can generate a new segment 137 that includes only the identified subset of journal entries, and can insert the new segment 137 into the journal log to replace target segment 137. In this way, journal log engine 130 can delete the journal entries that record the superseded writes, thus reducing the size of journal log 135 without losing useful information. An example process for reducing the size of journal log 135 is described below with reference to Figure 5 and Figures 6A - 6C describes an example process for reducing the size of journal log 135.
[0024] In some embodiments, journal log engine 130 can sort the journal entries in new segment 137 according to storage addresses. Thus, if new segment 137 is used to restore or recover the original data, the entries in new segment 137 can be read and written in the address order of storage volume 145. In this way, the journal reduction operation can provide faster restoration or recovery of the original data.
[0025] In some embodiments, the storage engine 120 and / or the journaling engine 130 may be implemented via hardware (e.g., electronic circuitry) or a combination of hardware and programming (e.g., including at least one processor and instructions executable by the at least one processor). In embodiments using executable instructions, such instructions may be stored on a machine-readable storage medium (e.g., storage device 140), stored in hardware (e.g., circuitry), and so on. The storage device 140 may include one or more non-transitory storage media, such as a hard disk drive (HDD), a solid-state drive (SSD), an optical disk, etc., or a combination thereof. Additionally, in some embodiments, the storage device 140 may include one or more block-based storage devices.
[0026] In some embodiments, the computing device 110 may be a physical computing device (e.g., a server, a device, a desktop, etc.). For example, the computing device 110 may include a controller, a memory, and a persistent storage device ( Figure 1 not shown). The controller may be implemented via hardware (e.g., electronic circuitry) or a combination of hardware and a program (e.g., including at least one processor and instructions executable by the at least one processor and stored on at least one machine-readable storage medium). The memory may be implemented in a semiconductor memory such as random access memory (RAM). The persistent storage device may include one or more non-transitory storage media, such as a hard disk drive (HDD), a solid-state drive (SSD), an optical disk, etc., or a combination thereof. In other embodiments, the computing device 110 may be one or more virtual computing devices (e.g., virtual machines, containers, etc.). For example, the storage engine 120 may be implemented in a first virtual machine, while the journaling engine 130 may be implemented in a second virtual machine. Additionally, in such an example, the first virtual machine may include a volume 145, while the second virtual machine may include a journal 135.
[0027] Note that while Figure 1 one example is shown, the embodiments are not limited thereto. For example, it is contemplated that the storage system 100 may include any number of computing devices 110 and / or storage devices 140. In another example, it is contemplated that the functions of the storage engine 120 and / or the journaling engine 130 may be included in a single engine or software, included in any other engine or software of the storage system 100, included in an external system or device ( Figure 1 not shown), included in a separate virtual computing device, or included in any combination thereof. Additionally, it is contemplated that the storage system 100 may include additional devices and / or components, fewer components, different components, different arrangements, etc. Other combinations and / or variations are also possible.
[0028] Figure 2, Figures 3A - 3D and Figures 4A - 4B —— An example process for generating a journal log
[0029] Figure 2 FIG. 200 shows an example process for generating a journal log according to some embodiments. In some examples, process 200 may be performed by some or all of storage system 100 ( Figure 1 shown in). Process 200 may be implemented in hardware or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions may be stored in a non-transitory computer-readable medium, such as an optical, semiconductor, or magnetic storage device. The machine-readable instructions may be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc. For illustrative purposes, the details of process 200 may be described below with reference to Figures 3A - 3D and Figures 4A - 4B which illustrate some example embodiments. However, other embodiments are possible.
[0030] Process 200 may begin at decision block 210, which may include determining whether a write command has been received. In the affirmative determination (“yes”), process 200 may continue at block 220, which may include inserting a copy of the write command into the journal log. Block 230 may include executing the write command to store data at a storage address. After block 230, or after a negative determination (“no”) at decision block 210, process 200 may continue at decision block 240, which may include determining whether a journal timer has expired. If it is determined at decision block 240 that the journal timer has not expired (“no”), then process 200 may return to decision block 210 (i.e., determine again whether a write command has been received).
[0031] For example, referring to Figure 3A , a controller (e.g., one or more processing engines included in computing device 110 shown in Figure 1 ) receives a first command to perform a data write 340 (e.g., write a set of one or more data blocks) to an address location [2:3] in storage volume 330. The controller executes the first command to perform data write 340 to address location [2:3]. Additionally, the controller creates a first entry 320 in journal segment 310 (e.g., a portion of the journal log) to store information about data write 340. For example, first entry 320 may include a copy of the data block being written, the write location (i.e., address location [2:3]), and any other information about data write 340.
[0032] Now referring to Figure 3B, the controller executes a second command to perform a data write 341 to address locations [6:7], and also creates a second entry 321 in the journal segment 310 to store information about the data write 341. Additionally, the controller executes a third command to perform a data write 342 to address locations [12:14], and also creates a third entry 322 in the journal segment 310 to store information about the data write 342. In some embodiments, each new entry is added to the journal segment 310 in the order of receipt (i.e., the order in which the write commands represented by the entries are received).
[0033] Now refer to Figure 3C , the controller executes a fourth command to perform a data write 343 to address locations [2:7], and also creates a fourth entry 323 in the journal segment 310 to store information about the data write 343. As Figure 3C shown, the data write 343 replaces (i.e., overwrites) the data write 340 and the data write 341.
[0034] Now refer to Figure 3D , the controller executes a fifth command to perform a data write 344 to address locations [12:13], and also creates a fifth entry 324 in the journal segment 310 to store information about the data write 344. As Figure 3D shown, the data write 344 overwrites a portion of the data write 342 that occupies the address locations [12:13]. However, a portion of the data write 342B (i.e., the portion of the data write 342 that is not overwritten by the data write 344) remains at the address location
[14] . Additionally, the controller may determine that the journal segment 310 has been filled (i.e., by storing the entries 320 - 324) to its maximum size, and in response, may initiate a new journal segment ( Figure 3D not shown in) to store any additional journal entries (i.e., records about subsequent write commands).
[0035] Refer again to Figure 2 , if it is determined at the decision block 240 that the journal timer has expired ("Yes"), then the process 200 may continue at block 250, including inserting a checkpoint into the journal log. Additionally, block 260 may include resetting the journal timer. After block 260, the process 200 may return to the decision block 210 (i.e., again determine whether a write command has been received). In some embodiments, the journal timer may be a cyclic timer that indicates (i.e., upon expiration) the desired time period between checkpoints in the journal log (e.g., two seconds, five seconds, etc.).
[0036] For example, refer to Figure 4A, the controller generates a journal log 400 as a linked chain of journal segments 410, 411, 412, 413, 414. Each journal segment can store a specified size or number of journal entries. Additionally, each journal segment can include a link (or links) to its immediate neighbors in the chain. For example, segment 411 can include a first link to the previous segment 410 and a second link to the next segment 412. As shown in Figure 4A , segment 410 can include a first checkpoint 420A created at a first point in time. When the journal timer expires, the controller can determine that the time elapsed since the creation of the first checkpoint 420A has reached the desired time interval between checkpoints. Accordingly, the controller generates a second checkpoint 420B in response to the expiration of the journal timer.
[0037] Now referring to Figure 4B , in some embodiments, the spacing between checkpoints 420 in the journal can increase as those checkpoints get older. For example, the controller can delete a specified number or percentage of checkpoints 420 of a particular age (e.g., delete 40% of the checkpoints 420 that are one day old, delete 50% of the remaining checkpoints 420 that are two days old, and so on). In this way, the time period represented between two consecutive checkpoints can gradually increase as those checkpoints age.
[0038] Figure 5 and Figures 6A - 6C —— Example Processes for Reducing Journal Logs
[0039] Figure 5 illustrates an example process 500 for reducing a journal log according to some embodiments. In some examples, process 500 can be performed by some or all of the storage system 100 (shown in Figure 1 ). Process 500 can be implemented in hardware, or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions can be stored in a non-transitory computer-readable medium, such as an optical, semiconductor, or magnetic storage device. The machine-readable instructions can be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, and so on. For illustrative purposes, the details of process 500 can be described below with reference to Figures 6A - 6C , which shows some example embodiments. However, other embodiments are possible.
[0040] Process 500 can begin at block 510, including detecting a trigger event for a journal reduction operation. For example, referring to Figure 1, the journaling engine 130 detects conditions or events that trigger the execution of a reduction operation on the journal 135. In a first example, the triggering event can be the determination that the system workload (e.g., the current workload of the storage system 100) is below a specified threshold, indicating that the storage system has available processing bandwidth to perform the reduction operation. In a second example, the triggering event can be the determination that the current size of the journal 135 has exceeded the maximum storage allocation. Additionally, in other examples, the triggering event can be a user command to initiate the reduction operation, a scheduled initiation of the reduction operation, a periodic timer expiration, etc. Other examples or combinations of triggering events are possible.
[0041] Referring again to Figure 5 , block 520 can include selecting a set of target segments in the journal. For example, referring to Figure 1 , the journaling engine 130 can analyze the journal 135 to identify a subset of segments 137 that are suitable candidates for the reduction operation (referred to herein as "target segments"). In some examples, the target segments 137 can meet a specified age range (e.g., segments older than a minimum age and / or segments newer than a maximum age). In another example, the target segments can be between two checkpoints separated by a specified interval range (e.g., between two consecutive checkpoints separated by a time interval greater than a minimum and / or less than a maximum). In yet another example, the target segments can include entries representing at least a minimum amount of overwrite operations (e.g., multiple writes to a single storage location). In yet another example, the target segments can include all of the segments 137 included in the journal 135. Other examples of target segments or combinations are possible.
[0042] Referring again to Figure 5 , block 530 can include: identifying a set of storage locations modified by the write operations recorded in the selected set of target segments. Block 540 can include: identifying, in the selected set of target segments, the most recent write operation for each storage location in the set of storage locations. For example, referring to Figure 6A , a diagram 600 is shown that illustrates data writes recorded in journal segment 310 (discussed above with reference to Figures 3A - 3D ). The diagram 600 can illustrate aspects of blocks 530 and 540 performed by a processing engine (e.g., Figure 1 the journaling engine 130 shown in Figure 6AAs shown, illustration 600 illustrates that data write 343 (recorded in entry 323) is the most recent write to address location [2:7], data write 344 (recorded in entry 324) is the most recent write to address location [12:13], and partial data write 342B (recorded in entry 322) is the most recent write to address location
[14] . Note that data writes 340 and 341 (recorded in entries 320 and 321) have been replaced (i.e., overwritten) by data write 343 (recorded in entry 323). Also note that while Figure 6A shows an illustrative example in which the most recent data writes are identified within a single journal segment 310, the implementation is not limited thereto. Instead, it is contemplated that identifying the most recent data writes (i.e., during blocks 530 and 540) can be performed across a set of multiple target segments.
[0043] Referring again to Figure 5 , block 550 may include arranging the set of most recent writes in the order of the write addresses. Block 560 may include generating a new segment that includes the set of most recent writes arranged in the order of the write addresses. For example, referring to Figure 6B , the processing engine determines that the set of most recent writes in the current target segment is data writes 343, 344, and 342B. In addition, the processing engine determines that the journal log information for the set of most recent data writes is stored in entry 323 (recording data write 343), entry 324 (recording data write 344), and partial entry 322B (recording partial data write). In addition, the processing engine sorts the entries of the most recent writes (e.g., entry 323, entry 324, and partial entry 322B) according to the order of the addresses of the most recent writes, and then generates a new segment 315 that includes only these sorted entries.
[0044] Referring again to Figure 5 , block 570 may include replacing the selected set of target segments in the journal log with the new segment. Decision block 580 may include determining whether there are more target segments in the journal log to be reduced. On an affirmative determination (“yes”), process 500 may return to 520 (i.e., select another set of target segments in the journal log). Otherwise, on a negative determination (“no”) at decision block 580, process 500 may be completed. For example, referring to Figure 6C , the processing engine replaces a set of multiple target segments 411, 412, 413 with the new segment 415. In some examples, the processing engine generates links between the previous segment 410 and the new segment 415 and between the new segment 415 and the next segment 414. In this way, the size of the journal log can be reduced and can continue to be used as a linked chain of segments.
[0045] In some embodiments, the reduced journal log can be used to reconstruct a storage volume or device as it existed at the time point represented by checkpoint 420B. For example, in the case where a storage volume is corrupted or lost, the writes recorded in the new segment 415 can be executed to reconstruct the storage volume as it existed at the time represented by checkpoint 420B. In some embodiments, the entries in the new segment 415 are arranged in the order of the write addresses, and thus the same entries are read and executed in the same order in which they were written to restore the storage volume. In this way, a journal log that includes a reduced portion or segment (e.g., new segment 415) can provide faster restoration or recovery of the data in the storage volume.
[0046] In some embodiments, older entries can be removed from the journal log (e.g., when a maximum age is reached), and the writes recorded in the removed entries can be executed to generate a mirrored volume (e.g., Figure 1 mirrored volume 155 as shown). The mirrored volume can represent the storage volume that existed when those removed entries were added to the journal log. In some embodiments, any removed entries included in the reduced segment (e.g., new segment 415) are arranged in the order of the write addresses, and thus such removed entries are read and executed in the same order in which they were written to generate the mirrored volume. In this way, a journal log that includes a reduced segment can provide faster generation of the mirrored volume.
[0047] Figure 7 — Exemplary Computing Device
[0048] Figure 7 A schematic diagram of an exemplary computing device 700 is shown. In some examples, the computing device 700 can generally correspond to some or all of the computing device 110 ( Figure 1 as shown), and the computing device 110 can be separate from the storage device 140. As shown, the computing device 700 can include a hardware processor 702 and a machine-readable storage device 705 that includes instructions 710 - 760. The machine-readable storage device 705 can be a non-transitory medium. The instructions 710 - 760 can be executed by the hardware processor 702 or by a processing engine included in the hardware processor 702.
[0049] Instructions 710 can be executed to detect a trigger event for a reduction operation on a journal log that records writes to a storage volume. For example, referring to Figure 1 , the journal log engine 130 detects conditions or events that trigger the execution of a reduction operation on the journal log 135. The trigger event can be based on the workload of the system 100, the size of the journal log 135, a user command, a schedule, a periodic timer, and so on.
[0050] Instruction 720 can be executed to select a target set of segments in a journal log in response to detection of a trigger event, where the target set of segments includes a plurality of journal entries. For example, referring to Figure 1 , the journal log engine 130 can identify the target set of segments from segment 137 of journal log 135. The target segments can be selected based on segment age, time interval, overwrite amount, and the like.
[0051] Instruction 730 can be executed to determine a set of storage locations modified by write operations recorded in the target set of segments. Instruction 740 can be executed to identify a subset of journal entries in the target set of segments, where each journal entry in the subset of journal entries records a most recent write operation that was recorded for a different storage location in the set of storage locations. For example, referring to Figure 6A , the controller determines that data write 343 is the most recent write to address location [2:7], data write 344 is the most recent write to address location [12:13], and partial data write 342B is the most recent write to address location
[14] . In some embodiments, the controller can read a storage data structure that records the most recent write operation for each address location in a storage volume.
[0052] Instruction 750 can be executed to generate a new segment that includes the identified subset of journal entries. For example, referring to 6B, the controller sorts the most recently written journal entries (e.g., entry 323, entry 324, and partial entry 322B) according to the order of the most recently written addresses, and then generates a new segment 315 that includes only these sorted entries.
[0053] Instruction 760 can be executed to replace the target set of segments in the journal log with the generated new segment. For example, referring to Figure 6C , the controller replaces target segments 411, 412, 413 with new segment 415. In some examples, the controller generates links between the previous segment 410 and the new segment 415, and between the new segment 415 and the subsequent segment 414.
[0054] Figure 8 ——Example Machine-Readable Medium
[0055] Figure 8 Illustrates a machine-readable medium 800 storing instructions 810 - 860 according to some embodiments. The instructions 810 - 860 can be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc., that can be separate from the storage device. The machine-readable medium 800 can be a non-transitory storage medium, such as an optical, semiconductor, or magnetic storage medium.
[0056] Instruction 810 can be executed to detect a trigger event for a reduction operation on a journal log, where the journal log records writes to a storage volume. Instruction 820 can be executed to select a target segment set in the journal log in response to the detection of the trigger event, where the target segment set includes a plurality of journal entries. Instruction 830 can be executed to determine a set of storage locations modified by write operations recorded in the target segment set.
[0057] Instruction 840 can be executed to identify a subset of journal entries in the target segment set, where each journal entry in the subset of journal entries records a most recent write operation that was recorded for a different storage location in the set of storage locations. Instruction 850 can be executed to generate a new segment including the identified subset of journal entries. Instruction 860 can be executed to replace the target segment set in the journal log with the generated new segment.
[0058] Figure 9 ——Example process for reducing a journal log
[0059] Figure 9 Illustrates an example process 900 for reducing a journal log according to some embodiments. In some examples, process 900 can be performed by some or all of the storage system 100 ( Figure 1 shown in). Process 900 can be implemented in hardware, or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions can be stored in a non-transitory computer-readable medium, such as an optical, semiconductor, or magnetic storage device. The machine-readable instructions can be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc.
[0060] Block 910 can include: detecting, by a controller, a trigger event for a reduction operation on a journal log, where the journal log records writes to a storage volume. Block 920 can include: selecting, by the controller, a target segment set in the journal log in response to the detection of the trigger event, where the target segment set includes a plurality of journal entries. Block 930 can include: determining, by the controller, a set of storage locations modified by write operations recorded in the target segment set.
[0061] Block 940 can include: identifying, by the controller, a subset of journal entries in the target segment set, where each journal entry in the subset of journal entries records a most recent write operation that was recorded for a different storage location in the set of storage locations. Block 950 can include: generating, by the controller, a new segment including the identified subset of journal entries. Block 960 can include: replacing, by the controller, the target segment set in the journal log with the generated new segment.
[0062] According to some embodiments described herein, a controller may perform operations to reduce the size of a journal log. The journal reduction operations may include: identifying a suitable target segment in the journal log, determining a set of storage locations modified by write operations recorded in the target segment, and identifying the most recent write operation for each modified storage location. Additionally, the journal reduction operations may include: identifying a subset of journal entries in the target segment that record the identified most recent write operations, generating a new segment that includes only the identified subset of journal entries, and replacing the target segment with the new segment. In this way, journal entries that record superseded writes may be deleted from the journal log, thereby reducing the size of the journal log without losing useful information. Additionally, in some embodiments, the subset of journal entries may be reordered according to the storage addresses in the new segment. Thus, if the new segment is used to restore or recover the original data, the entries in the segment may be read or written in the address order of the storage volume. In this way, the journal reduction operations may provide faster restoration or recovery of the original data.
[0063] Note that while Figures 1 - 9 various examples are shown, embodiments are not limited thereto. For example, referring to Figure 1 , it is contemplated that system 100 may include additional devices and / or components, fewer components, different components, different arrangements, and so on. In another example, it is contemplated that the functionality of journal log engine 130 described above may be included in any other engine or software of system 100. Other combinations and / or variations are also possible.
[0064] Data and instructions are stored in respective storage devices, which are implemented as one or more computer-readable or machine-readable storage media. The storage media include different forms of non-transitory memory, including semiconductor storage devices such as dynamic or static random access memory (DRAM or SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory; magnetic disks such as fixed disks, floppy disks, and removable disks; other magnetic media including magnetic tape; optical media such as compact discs (CDs) or digital video discs (DVDs); or other types of storage devices.
[0065] Note that the instructions discussed above can be provided on a computer-readable or machine-readable storage medium, or alternatively, can be provided on multiple computer-readable or machine-readable storage media distributed in a large system that may have multiple nodes. Such a computer-readable or machine-readable storage medium is considered to be part of an article (or manufacture). An article or manufacture can refer to any single component or multiple components that are manufactured. One or more storage media can be located in a machine that runs the machine-readable instructions, or at a remote site from which the machine-readable instructions can be downloaded over a network for execution.
[0066] In the foregoing description, numerous details have been set forth to provide an understanding of the subject matter disclosed herein. However, the embodiments may be practiced without some of these details. Other embodiments may include modifications and variations based on the details discussed above. The appended claims are intended to cover such modifications and variations.
Claims
1. A computing system, comprising: a processor; and a machine-readable storage device storing instructions executable by the processor to: detect a trigger event for a reduction operation on a journal log, where the journal log records writes to a storage volume; in response to detection of the trigger event, select a target segment set in the journal log, where the target segment set includes a plurality of journal entries; determine a set of storage locations modified by write operations recorded in the target segment set; identify a subset of journal entries in the target segment set, where each journal entry in the subset of journal entries records a most recent write operation that was recorded for a different storage location in the set of storage locations; generate a new segment including the identified subset of journal entries; and replace the target segment set in the journal log with the generated new segment.
2. The computing system according to claim 1, the computing system including instructions executable by the processor to: determine a storage order of the set of storage locations; and in the new segment, arrange the identified subset of journal entries based on the storage order of the set of storage locations.
3. The computing system according to claim 1, where the target segment set is between two checkpoints of the journal log.
4. The computing system according to claim 3, the computing system including instructions executed by the processor to: select the target segment set in response to determining that the two checkpoints are separated by more than a minimum time interval and less than a maximum time interval.
5. The computing system according to claim 1, the computing system including instructions executable by the processor to: select the target segment set in response to determining that the target segment set is older than a minimum age and newer than a maximum age.
6. The computing system according to claim 1, the computing system including instructions executable by the processor to: select the target segment set in response to determining that the target segment set includes entries representing at least a minimum amount of overwrite operations.
7. The computing system according to claim 1, where the trigger event is determining that the available processing bandwidth of the computing device exceeds a threshold level.
8. The computing system according to claim 1, where the trigger event is determining that the journal log has been filled to a predefined level.
9. A non-transitory machine-readable medium storing instructions that, when executed, cause a processor to: detect a trigger event for a reduction operation on a journal log, where the journal log records writes to a storage volume; in response to detection of the trigger event, select a target segment set in the journal log, where the target segment set includes a plurality of journal entries; determine a set of storage locations modified by write operations recorded in the target segment set; identify a subset of journal entries in the target segment set, where each journal entry in the subset of journal entries records a most recent write operation that was recorded for a different storage location in the set of storage locations; Generate a new segment that includes the identified subset of the diary entries; and Replace the target segment set in the diary log with the generated new segment.
10. The non-transitory machine-readable medium according to claim 9, the non-transitory machine-readable medium comprising instructions that, when executed, cause the processor to: Determine the storage order of the set of storage locations; and In the new segment, arrange the identified subset of the diary entries based on the storage order of the set of storage locations.
11. The non-transitory machine-readable medium according to claim 9, the non-transitory machine-readable medium comprising instructions that, when executed, cause the processor to: Select the target segment set in response to determining that two checkpoints are separated by more than a minimum time interval and less than a maximum time interval, wherein the target segment set is located between the two checkpoints of the diary log.
12. The non-transitory machine-readable medium according to claim 9, the non-transitory machine-readable medium comprising instructions that, when executed, cause the processor to: Select the target segment set in response to determining that the target segment set is older than a minimum age and newer than a maximum age.
13. The non-transitory machine-readable medium according to claim 9, the non-transitory machine-readable medium comprising instructions that, when executed, cause the processor to: Select the target segment set in response to determining that the target segment set includes entries representing at least a minimum amount of overwrite operations.
14. The non-transitory machine-readable medium according to claim 9, wherein the detection of the trigger event is based on at least one of system workload, size of the diary log, user command, scheduled event, and periodic timer.
15. A method, comprising: Detecting, by a controller, a trigger event for a reduction operation on a diary log, wherein the diary log records writes to a storage volume; In response to the detection of the trigger event, selecting, by the controller, a target segment set in the diary log, wherein the target segment set includes a plurality of diary entries; Determining, by the controller, a set of storage locations modified by write operations recorded in the target segment set; Identifying, by the controller, a subset of the diary entries in the target segment set, wherein each diary entry in the subset of the diary entries records a most recent write operation that was recorded for a different storage location in the set of storage locations; Generating, by the controller, a new segment that includes the identified subset of the diary entries; and Replacing, by the controller, the target segment set in the diary log with the generated new segment.
16. The method according to claim 15, comprising: Determining the storage order of the set of storage locations; and In the new segment, arranging the identified subset of the diary entries based on the storage order of the set of storage locations.
17. The method according to claim 15, comprising: Selecting the set of target segments in response to determining that two checkpoints are separated by more than a minimum time interval and less than a maximum time interval, wherein the set of target segments is located between the two checkpoints of the journal log.
18. The method according to claim 15, comprising: Selecting the set of target segments in response to determining that the set of target segments is older than a minimum age and newer than a maximum age.
19. The method according to claim 15, comprising: Selecting the set of target segments in response to determining that the set of target segments includes entries representing at least a minimum amount of overwrite operations.
20. The method according to claim 15, wherein the detection of the trigger event is based on at least one of system workload, size of the journal log, user command, scheduled event, and periodic timer.