WAL log space optimization method and device
By aggregating and flashing the logs of multiple IO data, the problems of wasted WAL log space and increasing disk erasing times are solved, and the WAL log space optimization and disk service life are achieved.
Patent Information
- Application Number
- CN202510071214.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-03
AI Technical Summary
Due to the secondary write amplification problem in the database system, WAL logs are wasted seriously, the effective information occupies low, and there is a problem of increasing the number of disk erases, which affects the service life of the disk.
The logs of multiple IO data are aggregated by adding writes to form multiple aggregated logs, and placed them in the brushing queue in turn to perform brushing processing, reducing the number of disk writes, and sharing the pressure of reading and writing disks.
Optimize the WAL log space through log aggregation, reduce the number of disk writes, increase the effective data share of the log disk, reduce the WAL log space, extend the disk service life, and ensure read and write efficiency in high concurrency scenarios.
Smart Images

Figure CN120086197A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of log space optimization, and more particularly, to a method and device for optimizing WAL log space. Background Art
[0002] The WAL log space is used in a database system to ensure data consistency and durability. It records all data modification operations to ensure that data can be restored through the log after a system crash and prevent data loss. Before a transaction is committed, the modifications are first written to the WAL log to ensure atomicity and consistency.
[0003] For a storage system using volatile memory, to ensure fast recovery of IO data and data consistency during exceptions, a commonly used method is to write the WAL log first. The WAL log ensures crash consistency through two writes. In addition to the relatively lagging update of the metadata itself being written to disk for each piece of IO data, there is also an append write log being written to disk, which itself has the problem of secondary write amplification. Since the minimum unit of log disk write is generally 4k, while the actual log of a single piece of IO data only uses dozens to hundreds of bytes, this means that most of the 4k log space is wasted and the occupancy of valid information is very low. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and device for optimizing WAL log space to improve the above problems. To achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0005] In a first aspect, the present application provides a method for optimizing WAL log space, including:
[0006] Aggregating the logs of multiple pieces of IO data by means of append writing to obtain multiple aggregated logs;
[0007] Sequentially putting the multiple aggregated logs into a disk write queue;
[0008] According to the disk write queue, performing disk write processing on the multiple aggregated logs to obtain the disk write result of the associated metadata of the aggregated logs, where the disk write processing includes sequentially performing disk write of the aggregated logs and disk write of the metadata;
[0009] Performing data operations according to the disk write result to obtain the WAL log space optimization result.
[0010] In a second aspect, the present application further provides a device for optimizing WAL log space, including:
[0011] An aggregation module, configured to aggregate the logs of multiple pieces of IO data by means of append writing to obtain multiple aggregated logs;
[0012] A queuing module for sequentially placing multiple said aggregated logs into a disk flushing queue;
[0013] A disk flushing module for performing disk flushing on multiple said aggregated logs according to the disk flushing queue to obtain a disk flushing result of the associated metadata of the aggregated logs, where the disk flushing process includes sequentially performing aggregated log disk flushing and metadata disk flushing;
[0014] An operation module for performing data operations according to the disk flushing result to obtain a WAL log space optimization result.
[0015] The beneficial effects of the present invention are as follows: The present invention optimizes the WAL log space through log aggregation, reduces the number of disk writes to share the disk read / write pressure, increases the proportion of valid data in the log disk, reduces the occupancy of the WAL log space, and reduces the number of disk erasure writes to extend the disk service life. At the same time, through the mutual perception of aggregation and disk flushing, not only the read / write efficiency in the scenario of high pressure and multiple concurrencies is guaranteed, but also the read / write latency of single concurrency and tail IO data can still be ensured, adapting to various different IO models, and ensuring that the aggregated logs are immediately flushed to disk without idling.
[0016] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification or can be understood by implementing the embodiments of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic flowchart of the WAL log space optimization method described in the embodiments of the present invention;
[0019] Figure 2 It is a flowchart of aggregated log disk flushing in the embodiments of the present invention;
[0020] Figure 3 It is a schematic structural diagram of aggregated logs in the embodiments of the present invention. Detailed Embodiments
[0021] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0022] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for differential description and cannot be construed as indicating or implying relative importance.
[0023] Embodiment 1:
[0024] This embodiment provides a method for optimizing the WAL log space.
[0025] See Figure 1 , which shows that this method includes step S100, step S200, step S300 and step S400.
[0026] Step S100: Aggregate the logs of multiple IO data by means of append writing to obtain multiple aggregated logs;
[0027] The step S100 includes:
[0028] Step S101: Obtain the log of each IO data;
[0029] Step S102: When detecting the log corresponding to the submission of the first IO data, allocate a memory of a preset size in the WAL log space to generate a corresponding memory buffer;
[0030] Step S103: Serialize the logs of each IO data into the memory buffer in sequence, and determine whether the preset first aggregation condition or the second aggregation condition is met. If so, obtain an aggregated log and perform the aggregation of the next aggregated log. Otherwise, wait for the log of the next IO data to be serialized into the memory buffer. The first aggregation condition is that there is no remaining space in the memory buffer, and the second aggregation condition is the preset aggregation length, the preset number of aggregations or the preset timeout period.
[0031] In this embodiment, log aggregation is attached to the log submission of the IO process, mainly used to aggregate the logs of multiple IO data on a buffer in an append-write manner, where the buffer represents a memory buffer.
[0032] When the log of the first IO data in each aggregation is submitted, 4k (rounded up to the nearest 4k) of space will be pre-allocated from the log space on the disk to generate the corresponding buffer. After each IO data serializes the current log to the buffer, it is necessary to determine whether to wait for subsequent IO data to append and use this space and serialize it to the buffer. After the condition for triggering the completion of aggregation is met, the aggregated log is obtained and placed in the disk write queue.
[0033] Step S200: Sequentially place the multiple aggregated logs into the disk write queue;
[0034] As Figure 2 shown, it is a schematic diagram of the disk write queue for the aggregated log. Among them, IO 1 to IO 8 are the IO data to be written to the disk in sequence, and the disk write queue is a first-in-first-out disk write queue.
[0035] Before the step S200, it further includes:
[0036] Step A100: Define the header information of each log in the aggregated log, and the header information of each log includes a log sequence number, length information, and type information;
[0037] Step A200: Define a bitmap field through the header information of each log, and the bitmap field is used to mark the validity of each log in the aggregated log;
[0038] Step A300: Define the header information of each aggregated log, and the header information of the aggregated log includes a bitmap field, the starting point of valid logs, and the starting position for reading the disk when replaying the logs.
[0039] In this embodiment, since the WAL log space writes logs first, when the associated metadata update fails after the logs are first written to the disk, the entire IO data should be considered failed and invalid this time, and the logs that have been written to the disk also need to be marked as invalid.
[0040] As Figure 3 shown is the composition of the aggregated log. Among them, HEADER represents the header information of the aggregated log, header represents the header information of the log in the aggregated log, TAIL represents the tail information of the aggregated log, and IO 1log and IO 2log respectively represent the logs corresponding to the first and second IO data in the aggregated log. And HEADER actually represents the global information of the aggregated log.
[0041] Add a bitmap field to the HEADER of each aggregated log. Different bits in the bitmap field are used to identify whether the corresponding log is valid, and by default, all are valid. If the metadata update fails for the IO data, the corresponding bit in the HEADER of the aggregated log is set to invalid. After aggregating the last piece of IO data, if it is found that there has been a failed IO data, the log needs to be rolled back, that is, the buffer of the aggregated log after modifying the bitmap information is written to disk again to overwrite the original log.
[0042] Step S300: According to the disk write queue, perform disk write processing on multiple aggregated logs to obtain the disk write result of the associated metadata of the aggregated logs. The disk write processing includes sequentially performing disk write of the aggregated logs and disk write of the metadata;
[0043] In this embodiment, when performing disk write processing, the disk write of the aggregated logs and the disk write of the metadata are performed sequentially. That is, first write the aggregated logs to disk, then update the associated metadata corresponding to the aggregated logs, and then write the associated metadata to disk. After writing the associated metadata to disk, then trim the logs corresponding to the written associated metadata.
[0044] In the step S300, performing disk write of the aggregated logs includes:
[0045] Step S301: Obtain the marking parameters of the disk write queue. The marking parameters include a first mark and a second mark. The first mark indicates that the disk write queue is not empty, and the second mark indicates that the disk write queue is empty;
[0046] Step S302: If the marking parameter is the first mark, then according to the disk write queue, sequentially perform disk write processing on multiple aggregated logs;
[0047] Step S303: If the marking parameter is the second mark, then take out the unaggregated logs in the memory buffer for disk write processing.
[0048] In this embodiment, after the log of each IO data is serialized to the buffer, if the disk write queue is empty, there is no need to wait for subsequent IO data, but directly put the buffer of the aggregated log into the disk write queue for disk write. This operation can ensure that the IO data can be completed in a timely manner in a single-concurrency scenario.
[0049] If the disk write queue is not empty, then determine whether the buffer of the aggregated log has reached the condition of aggregation completion. If any of the aggregation length, aggregation count, or timeout has reached the standard, put it into the disk write queue and wait for disk write. Then sequentially take out the aggregated logs that have completed aggregation from the disk write queue.
[0050] When the queue is empty and there is no aggregated log, the unaggregated bufer r is taken out for disk flushing. Through this operation, it can be ensured that the last IO data of the write service can be completed in a timely manner.
[0051] Step S400: Perform data operations according to the disk flushing result to obtain the WAL log space optimization result.
[0052] The step S400 includes:
[0053] Step S401: Judge the disk flushing result of the associated metadata of the aggregated log;
[0054] Step S402: If the disk flushing result is successful, trim the logs that are no longer needed, release the WAL log space, and update the header information of the aggregated log. The logs that are no longer needed refer to the logs for which the metadata disk flushing has been completed;
[0055] In this embodiment, after the associated metadata is flushed to the disk metadata partition, the logs that are no longer needed need to be trimmed to release the WAL log space. The associated metadata of multiple IO data corresponding in the aggregated log cannot ensure that they are flushed to the disk metadata partition simultaneously. Therefore, the log sequence number of each log is recorded in the aggregated log.
[0056] When trimming the logs that are no longer needed for the IO data whose metadata disk flushing is completed first, only the log sequence number in the global information is updated, that is, the valid log starting point of the playback log. After the metadata disk flushing of the last IO data in the aggregated log is completed, the starting position of reading the disk during the playback log in the global information also needs to be updated additionally.
[0057] Step S403: If the disk flushing result is a memory update failure, perform a log rollback operation. The log rollback operation includes modifying the invalid bitmap field and re-flushing the aggregated log with the modified bitmap field to overwrite the original flushed aggregated log.
[0058] In this embodiment, after the aggregated log is written to disk, the memory update of the associated metadata may fail due to concurrent conflicts or other reasons. At this time, the log needs to be rolled back. If the memory update of the associated metadata is successful, the log written to disk is reliable and there is no need to roll back the log.
[0059] When performing the disk flushing process, if the system crashes, a log playback operation is performed;
[0060] The log playback operation includes:
[0061] Step B100: Obtain the starting position of reading the disk during the playback log;
[0062] Step B200: Read the aggregated logs to be replayed from the starting position of disk reading when playing back the playback log;
[0063] Step B300: Parse the logs in each of the aggregated logs to be replayed through the header information of the aggregated logs to be replayed, to obtain valid logs;
[0064] Step B400: Obtain the valid log starting point of the aggregated logs to be replayed, and identify the first valid log through the valid log starting point;
[0065] Step B500: Starting from the first valid log, replay the valid logs in sequence, and restore the logs to the state before system crash.
[0066] In this embodiment, after the system crashes, it is necessary to replay the valid logs of the uncropped part to ensure data consistency before and after the crash. When replaying, read from the disk starting from the starting position of disk reading in the header information of the aggregated log, and parse the aggregated log with a granularity of 4K. After parsing the HEADER, obtain the number of logs included in this aggregated log, then parse each log in sequence, identify the first valid log through the valid log starting point, and then start normal replay from here.
[0067] In summary, the present invention reduces the number of disk writes through log aggregation to share the disk read and write pressure, increases the proportion of valid data in the log disk to reduce the occupancy of the WAL log space, and reduces the number of disk erasure writes to improve the disk service life. At the same time, through the mutual perception of aggregation and disk flushing, not only the read and write efficiency in the scenario of high pressure and multi-concurrency is guaranteed, but also the read and write latency of single-concurrency and tail IO data can still be guaranteed.
[0068] The present invention uses bitmap fields to mark the validity of each log in the aggregated log, and realizes log rollback through position reflushing of the bitmap. It also accurately crops the expired logs in the aggregated log through log sequence numbers, and accurately identifies the valid starting log when replaying the logs after the system crashes.
[0069] Embodiment 2:
[0070] This embodiment provides a WAL log space optimization device, and the device includes:
[0071] An aggregation module, configured to aggregate the logs of multiple IO data in an append-write manner to obtain multiple aggregated logs;
[0072] A queuing module, configured to sequentially put the multiple aggregated logs into a disk flushing queue;
[0073] A disk writing module, configured to perform disk writing on multiple pieces of the aggregated logs according to the disk writing queue, so as to obtain a disk writing result of the associated metadata of the aggregated logs, where the disk writing process includes sequentially performing aggregated log disk writing and metadata disk writing;
[0074] An operation module, configured to perform data operations according to the disk writing result to obtain a WAL log space optimization result.
[0075] The aggregation module includes:
[0076] A first acquisition unit, configured to acquire the log of each piece of IO data;
[0077] A generation unit, configured to allocate a preset size of memory in the WAL log space and generate a corresponding memory buffer when detecting the log corresponding to the submission of the first piece of IO data;
[0078] An aggregation unit, configured to sequentially serialize the logs of each piece of IO data into the memory buffer, and determine whether a preset first aggregation condition or a second aggregation condition is met. If so, obtain an aggregated log and perform the aggregation of the next aggregated log. Otherwise, wait for the log of the next piece of IO data to be serialized into the memory buffer. The first aggregation condition is that there is no remaining space in the memory buffer, and the second aggregation condition is a preset aggregation length, a preset number of aggregated items, or a preset timeout period.
[0079] The disk writing module includes:
[0080] A second acquisition unit, configured to acquire the marking parameters of the disk writing queue, where the marking parameters include a first mark and a second mark. The first mark indicates that the disk writing queue is not empty, and the second mark indicates that the disk writing queue is empty;
[0081] A first disk writing unit, configured to perform disk writing on multiple pieces of the aggregated logs sequentially according to the disk writing queue if the marking parameter is the first mark;
[0082] A second disk writing unit, configured to perform disk writing on the unaggregated logs in the memory buffer if the marking parameter is the second mark.
[0083] The operation module includes:
[0084] A third acquisition unit, configured to judge the disk writing result of the associated metadata of the aggregated logs;
[0085] A trimming unit, configured to trim the logs that are no longer needed, release the WAL log space, and update the header information of the aggregated log if the disk writing result is successful in disk writing. The logs that are no longer needed indicate the logs for which metadata disk writing has been completed;
[0086] A rollback unit, configured to perform a log rollback operation if the disk brushing result is a memory update failure.
[0087] It should be noted that, regarding the devices in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0088] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0089] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A WAL log space optimization method, characterized in that: include: Aggregate multiple IO data logs by appending writes to obtain multiple aggregate logs. Putting the plurality of aggregated logs into a disk flushing queue in sequence; According to the disk flushing queue, the plurality of aggregate logs are subjected to disk flushing processing to obtain disk flushing results of metadata associated with the aggregate logs, wherein the disk flushing processing includes sequentially performing aggregate log disk flushing and metadata disk flushing; Data operations are performed according to the disk flushing results to obtain WAL log space optimization results.
2. The WAL log space optimization method according to claim 1, characterized in that ,The logs of multiple IO data are aggregated by appending writes to obtain multiple aggregate logs, including: Get the log of each IO data; When the first IO data is detected to be submitted to the corresponding log, a preset size of memory is allocated in the WAL log space to generate a corresponding memory buffer; Serialize the log of each IO data into the memory buffer in turn, and determine whether the preset first aggregation condition or the second aggregation condition is met. If so, obtain an aggregate log and perform aggregation of the next aggregate log. Otherwise, wait for the next IO data log to be serialized into the memory buffer. The first aggregation condition is that there is no remaining space in the memory buffer, and the second aggregation condition is a preset aggregation length, a preset number of aggregations, or a preset timeout.
3. The WAL log space optimization method according to claim 1, characterized in that , before putting the plurality of aggregate logs into the disk flushing queue in sequence, it also includes: Defining header information of each log in the aggregated log, wherein the header information of each log includes a log sequence number, length information, and type information; A bitmap field is defined by the header information of each log, where the bitmap field is used to mark the validity of each log in the aggregate log; The header information of each aggregate log is defined, wherein the header information of the aggregate log includes a bitmap field, a valid log starting point, and a starting position for reading a disk when playing back a log.
4. The WAL log space optimization method according to claim 2, characterized in that ,The aggregate log flushing includes: Obtaining a mark parameter of the disk flushing queue, the mark parameter including a first mark and a second mark, the first mark indicating that the disk flushing queue is not empty, and the second mark indicating that the disk flushing queue is empty; If the marking parameter is the first marking, the plurality of aggregate logs are sequentially flushed to disk according to the flushing queue; If the marking parameter is the second marking, the unaggregated logs in the memory buffer are taken out for disk flushing.
5. The WAL log space optimization method according to claim 3, characterized in that , performing data operations according to the disk flushing results to obtain WAL log space optimization results, including: Determine the flushing result of the metadata associated with the aggregate log; If the flushing result is successful, the logs that are no longer needed are pruned to release the WAL log space, and the header information of the aggregate log is updated. The logs that are no longer needed indicate the logs for which the metadata flushing has been completed. If the flushing result is a memory update failure, a log rollback operation is performed, which includes modifying the invalid bitmap field and re-flushing the aggregate log with the modified bitmap field to overwrite the original flushed aggregate log.
6. The WAL log space optimization method according to any one of claim 3, characterized in that ,When the disk flushing process is performed, if the system crashes, a log playback operation is performed; The log playback operation includes: Obtaining the starting position of disk reading when replaying the log; Read the aggregate log to be played back through the starting position of the disk read when playing back the log; Parsing each log in the aggregated log that needs to be replayed through the header information of the aggregated log that needs to be replayed to obtain a valid log; Obtaining a valid log starting point of the aggregate log to be replayed, and identifying a first valid log through the valid log starting point; Starting from the first valid log, replay the valid logs in sequence to restore the logs to the state before the system crash.
7. A WAL log space optimization device, characterized in that: include: Aggregation module, used to aggregate multiple IO data logs by appending writes to obtain multiple aggregate logs; A queuing module, used to sequentially put the plurality of aggregated logs into a disk flushing queue; A disk flushing module, used for performing disk flushing processing on the plurality of the aggregate logs according to the disk flushing queue to obtain a disk flushing result of metadata associated with the aggregate log, wherein the disk flushing processing includes sequentially performing aggregate log disk flushing and metadata disk flushing; The operation module is used to perform data operations according to the disk flushing result to obtain the WAL log space optimization result.
8. The WAL log space optimization device according to claim 7, characterized in that: The aggregation module includes: A first acquisition unit is used to acquire a log of each IO data; A generation unit is used to allocate a preset size of memory in the WAL log space and generate a corresponding memory buffer when detecting that the first IO data is submitted to the corresponding log; The aggregation unit is used to serialize the log of each IO data into the memory buffer in turn, and determine whether the preset first aggregation condition or the second aggregation condition is met. If so, an aggregation log is obtained and the next aggregation log is aggregated; otherwise, the next IO data log is waited to be serialized into the memory buffer. The first aggregation condition is that there is no remaining space in the memory buffer, and the second aggregation condition is a preset aggregation length, a preset number of aggregations, or a preset timeout.
9. The WAL log space optimization device according to claim 8, characterized in that: The brushing module comprises: A second acquisition unit is used to acquire a mark parameter of the disk flushing queue, wherein the mark parameter includes a first mark and a second mark, wherein the first mark indicates that the disk flushing queue is not empty, and the second mark indicates that the disk flushing queue is empty; A first disk flushing unit, configured to, if the marking parameter is a first mark, sequentially flush the plurality of aggregate logs according to the disk flushing queue; The second disk flushing unit is used for taking out the unaggregated logs in the memory buffer and performing disk flushing processing if the marking parameter is the second marking.
10. The WAL log space optimization device according to claim 7, characterized in that: The operation module includes: A third acquisition unit, used to determine the flushing result of the metadata associated with the aggregate log; A pruning unit, configured to prune the logs that are no longer needed, release the WAL log space, and update the header information of the aggregate log if the flushing result is successful. The logs that are no longer needed represent logs that have completed the metadata flushing. The rollback unit is used to perform a log rollback operation if the flashing result is a memory update failure.