Log data compression method and device, electronic equipment and medium

By splitting, merging and compressing the write-ahead log, the problem of excessive memory usage on the master node is solved, and efficient data synchronization and memory management are achieved.

CN120687423APending Publication Date: 2025-09-23TSINGHUA UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510837655.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In data synchronization based on master and slave nodes, the master node has a problem of excessive memory usage. Especially when the slave node does not actively pull data, the pre-write log cache will continue to accumulate, resulting in excessive memory usage.

Method used

By segmenting the write-ahead log, multiple data change records are grouped according to data pages. When the preset merge rules are met, they are merged and compressed to generate compressed and merged page data, and data is synchronized through version identification.

Benefits of technology

Effectively control the space occupied by the write-ahead log, reduce the memory usage of the master node, improve data synchronization efficiency, and reduce the computing resource consumption of the master node.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687423A_ABST
    Figure CN120687423A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a log data compression method and device, electronic equipment and a medium, and the method comprises the steps: responding to a data change request, and writing a plurality of data change records into a pre-writing log; one data change record corresponds to a single data page; based on the data page of each data change record, segmenting the pre-writing log to obtain a plurality of page group logs; the page group log comprises a plurality of page logs, and each page log corresponds to the same data page or a page set with logic association; and when the log attribute of any page group log meets the preset merging rule, merging and compressing the page logs in the page group log to generate compressed and merged page data, thereby effectively solving the problem of overlarge memory occupation of a main node in data synchronization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of log data processing, and in particular to a log data compression method, device, electronic device, and medium. Background Art

[0002] Regarding the current data synchronization technology based on master-slave nodes, in the technology, the master node stores the data change operations in a pre-write log, and when it responds to the data synchronization request of the slave node, it sends the cached pre-log to the slave node, so as to achieve data synchronization between the master and slave nodes. It can be seen that in the process of data synchronization between the above-mentioned master and slave nodes, the master node will not actively push the changed data to the slave node, but will accumulate it in the pre-write log cache, and only transmit the changed data when the slave node needs to pull data (respond to the data synchronization request). The problem brought about by this is that the data synchronization between the master and slave nodes is actually driven by the data synchronization request of the slave node, so it is possible that the master node will accumulate too much log data, which will lead to excessive memory usage.

[0003] Therefore, how to solve the problem of excessive memory usage of the master node in a data synchronization solution based on synchronous pull has become a technical problem that technical personnel in this field urgently need to solve. Summary of the Invention

[0004] Based on the above problems, in order to solve the problem of excessive memory usage of the master node in a data synchronization solution based on synchronous pull, the present application provides a log data compression method, device, electronic device and medium.

[0005] The embodiments of this application disclose the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a log data compression method, applied to a master node for performing data synchronization, the method comprising:

[0007] In response to a data change request, write multiple data change records into a write-ahead log; one data change record corresponds to a single data page;

[0008] Based on the data pages of each data change record, the write-ahead log is segmented to obtain a plurality of page group logs; the page group log includes a plurality of page logs, and each of the page logs corresponds to the same data page or a set of logically associated pages;

[0009] When the log attribute of any of the page group logs meets the preset merging rule, the page logs in the page group log are merged and compressed to generate compressed and merged page data.

[0010] In a possible implementation, the preset merging rules include: a first rule, a second rule, and a third rule; the first rule includes: the log accumulation amount of the page group log reaches a page size threshold; the second rule includes: the accumulation time of the page group log reaches a timer threshold; the third rule includes: the remaining capacity of the cache area of ​​the write-ahead log is lower than a preset ratio value;

[0011] When the log attribute of any of the page group logs satisfies the preset merging rule, merging and compressing the page logs in the page group log to generate compressed and merged page data, including:

[0012] When the log attribute of the page group log satisfies at least one of the first rule, the second rule, and the third rule, the page logs in the page group log are merged and compressed to generate the compressed and merged page data.

[0013] In a possible implementation, the master node includes: a first preset network device, and after generating the compressed and merged page data, the method further includes:

[0014] For the compressed and merged page data, a separately corresponding first data version identifier is recorded, so that the slave node performs data synchronization through the first data version identifier, the second data version identifier of the internal data of the slave node, the first preset network device, and the second preset network device set in the slave node;

[0015] The second preset network device is used to interact with the first preset network device so that the slave node can pull the memory data of the master node.

[0016] In one possible implementation, the step of synchronizing data by the slave node based on the first data version identifier, the second data version identifier of the internal data of the slave node, and the second preset network device in the slave node includes:

[0017] When the second data version identifier is smaller than the first data version identifier, the slave node pulls the compressed and merged page data and the unmerged data change record;

[0018] When the second data version identifier is greater than the first data version identifier, the slave node only pulls data change records in the write-ahead log whose data version identifier is greater than the second data version identifier;

[0019] After pulling the above data, the slave node sets the second data version identifier as the identifier of the latest data change record.

[0020] In a possible implementation, the method further includes:

[0021] When the response frequency of the master node to the data change request is greater than a preset threshold, data synchronization of the slave node is completed through the replica page of the page data and the incremental log of the write-ahead log;

[0022] When the response frequency of the master node to the data change request is not greater than the preset threshold, the data synchronization of the slave node is completed based on the incremental log stream of the write-ahead log.

[0023] In a possible implementation, before splitting the write-ahead log, the method further includes:

[0024] If all the data change records in the write-ahead log correspond to the same data page, the splitting process step for the write-ahead log is not performed.

[0025] In a second aspect, an embodiment of the present application provides a log data compression device, applied to a master node, comprising:

[0026] a log writing module, configured to write a plurality of data change records into a write-ahead log in response to a data change request; wherein one data change record corresponds to a single data page;

[0027] a segmentation processing module for segmenting the write-ahead log based on the data pages corresponding to the plurality of data change records to obtain a plurality of page group logs; the page group logs comprising a plurality of page logs, each of which corresponds to the same data page or a set of logically associated pages;

[0028] The merging and compressing module is used to merge and compress the page logs in any page group log when the log attribute of the page group log meets the preset merging rule, so as to generate compressed and merged page data.

[0029] In a possible implementation, the preset merging rules include: a first rule, a second rule, and a third rule; the first rule includes: the log accumulation amount of the page group log reaches a page size threshold; the second rule includes: the accumulation time of the page group log reaches a timer threshold; the third rule includes: the remaining capacity of the cache area of ​​the write-ahead log is lower than a preset ratio value;

[0030] The merging and compression module is specifically used to:

[0031] When the log attribute of the page group log satisfies at least one of the first rule, the second rule, and the third rule, the page logs in the page group log are merged and compressed to generate the compressed and merged page data.

[0032] In a third aspect, an embodiment of the present application provides an electronic device, the device comprising: a processor, a memory, and a system bus;

[0033] The processor and the memory are connected via the system bus;

[0034] The memory is used to store one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by the processor, the processor is caused to perform any possible log data compression method in the first aspect.

[0035] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any possible log data compression method in the first aspect.

[0036] Compared with the prior art, the present application has the following beneficial effects: the embodiment of the present application provides a log data compression method, device, electronic device and medium, in which the method first responds to a data change request and writes a plurality of data change records corresponding to the data change request into a pre-write log. Subsequently, because the storage unit of the database is a data page, and each data change record has a corresponding data page, the pre-write log can be segmented based on the data page corresponding to each data change record to obtain a plurality of page group logs. In the page group log, since each page log corresponds to the same data page or a set of logically associated pages, when the log attribute of any page group log in the plurality of page group logs meets the preset merge rule, the page logs in the page group log are merged and compressed, and the plurality of page logs for the same data page can be merged into a series of operation results in one page. Since the memory occupied by the results of a series of operations on a page is only one page, the space occupied by the write-ahead log can be effectively controlled. Even if multiple data change operations are performed on a data page in the pre-write log, the page log corresponding to each data change operation will eventually be compressed and merged into the content of one page data, effectively solving the problem of excessive memory usage on the master node during data synchronization. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0038] Figure 1 A schematic diagram of implementing data synchronization based on pulling from a node is provided in an embodiment of the present application;

[0039] Figure 2 A flow chart of a log data compression method provided in an embodiment of the present application;

[0040] Figure 3 A schematic diagram of another log compression processing flow provided in an embodiment of the present application;

[0041] Figure 4 A flowchart of a method for pulling data from a response node provided in an embodiment of the present application;

[0042] Figure 5 A schematic diagram of the structure of a log data compression device provided in an embodiment of the present application;

[0043] Figure 6 A schematic diagram of the structure of a log data compression electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] To make the objectives, technical solutions, and advantages of this application more clearly understood, the application is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be noted that the embodiments described in the embodiments of this application are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0045] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by people with ordinary skills in the field to which this application belongs. The words "first", "second" and similar terms used in the embodiments of this application do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "include" or "comprise" mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, but do not exclude other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0046] See also Figure 1 , this figure is a schematic diagram of a method of realizing data synchronization based on pulling from a slave node provided by an embodiment of the present application. As can be seen from the figure, before the data changes fall into the storage engine, the master node needs to write the data changes into the pre-write log in advance. The master node will not actively send the pre-write log to the slave node. Only when the slave node sends a data synchronization request to the master node will the master node send the cached pre-write log to the slave node. The problem caused by this is that if the slave node does not send a data synchronization request to the master node, the pre-log cache in the master node will continue to accumulate. When data change requests are frequent, since each data change operation will occupy a page log, in a scenario with frequent write operations, the memory of the master node is prone to high occupancy.

[0047] In order to solve the above problems, the embodiments of the present application provide a log data compression method, device, electronic device and medium. In the method, first, in response to a data change request, multiple data change records corresponding to the data change request are written into a pre-write log. Then, because the storage unit of the database is a data page, and each data change record has a corresponding data page, the pre-write log can be segmented based on the data page corresponding to each data change record to obtain multiple page group logs. In the page group log, since each page log corresponds to the same data page or a set of logically associated pages. Therefore, when the log attribute of any page group log in the multiple page group logs meets the preset merge rule, the page logs in the page group log are merged and compressed, and multiple page logs for the same data page can be merged into a series of operation results in one page. Since the memory occupied by the results of a series of operations on a page is only one page, the space occupied by the write-ahead log can be effectively controlled. Even if multiple data change operations are performed on a data page in the pre-write log, the page log corresponding to each data change operation will eventually be compressed and merged into the content of one page data, effectively solving the problem of excessive memory usage on the master node during data synchronization.

[0048] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0049] See also Figure 2 , which is a flow chart of a log data compression method provided by an embodiment of the present application, specifically comprising the following steps:

[0050] S101: In response to a data change request, write multiple data change records into a write-ahead log; one data change record corresponds to a single data page.

[0051] Write-Ahead Logging (WAL) is a technology used to ensure data consistency and persistence, and is widely used in database management systems and file systems. Before making changes to a database or file system, WAL records these changes in a write-ahead log, which then applies the changes to the actual data files.

[0052] Therefore, when the master node responds to a data change request from an external client, it needs to write the data change record contained in the data change request into the write-ahead log so that the data can be synchronized to the slave node through the write-ahead log later.

[0053] The data change request is used to represent a specific execution operation on any page of data in the database, such as insert, update, and delete, etc. The data change request contains the page identifier of the target page to be changed, such as the page path, page number, etc.

[0054] Corresponding to a data change request, a data change record is a detailed record of the data change operation. It records the detailed information when the data on the target page is changed, such as the operation method, initial value, and end value. Each data change record corresponds to a single data page or data object, ensuring the accuracy and traceability of subsequent data synchronization processes.

[0055] It's important to note that in database transaction processing scenarios, when a transaction commits, the batch of data modification operations it contains generates multiple corresponding change records, which constitute the specific modifications to the database state. To ensure transaction atomicity—that is, to ensure that all operations in the batch are either successfully written to the final storage or rolled back and discarded—all change records corresponding to each transaction are marked with a unified version identifier, namely the Log Sequence Number (LSN).

[0056] As a globally unique, incremental identifier, LSN not only uniquely identifies the set of all change records for the transaction, but also plays a key role in the database's recovery and synchronization mechanisms: during the data writing process, the LSN will only be marked as valid after all associated change records have been persistently stored (or confirmed discarded according to transaction rules). In the data synchronization scenario between master and slave nodes, the slave node will record the currently synchronized LSN version in real time, and the master node will determine which unsynchronized change records (i.e., the part with an LSN greater than the current record version of the slave node) need to be transmitted to the slave node by comparing the version number. This LSN-based identification and tracking mechanism not only ensures the integrity of transaction execution, but also provides a precise version anchor for the data synchronization process, allowing master and slave nodes to efficiently and reliably maintain data consistency, ensuring that in complex distributed environments, all change records can be correctly applied or rolled back according to transaction boundaries, avoiding data inconsistencies caused by partial synchronization.

[0057] In one possible implementation, the method of writing data change records into the pre-write log can be synchronous writing, asynchronous writing, batch writing, sequential writing and other writing methods, which are determined according to actual needs and scenarios. This embodiment does not limit the writing method of data change records.

[0058] S102: Segmenting the write-ahead log based on the data pages of each data change record to obtain multiple page group logs; the page group log includes multiple page logs, and each page log corresponds to the same data page or a set of logically associated pages.

[0059] In a database, data is stored in the form of pages. Accordingly, each data change record is presented in the form of a data page. After the data change record is written to the write-ahead log, because each data change record is presented in the form of a data page, and the data pages involved in the data change process are often non-contiguous, there will often be multiple data change records for the same data page in the write-ahead log. These data change records each occupy the memory of a page log. Therefore, in write-intensive data change scenarios, the accumulation of write-ahead logs often occupies a large amount of memory, causing the master node to have excessive memory pressure.

[0060] Therefore, after multiple data change records are written to the write-ahead log in the previous step, in order to prevent multiple data change records for the same data page from occupying independent page storage space respectively, it is necessary to segment the write-ahead log based on the data page corresponding to each data change record to obtain multiple page group logs. In the page group log, each page group log contains multiple page logs (which can be understood as data change records). The data change operations represented by these page logs all correspond to the same data page in the database, or there are logically associated page sets. For example, page group log A is the data change operation log corresponding to page6 page, and page log B is the data change operation log corresponding to page7 page. If page5 and page6 are page sets with logical associations, then page log A can also be the data change operation log corresponding to page5 and page6.

[0061] In particular, if there is a data page in the write-ahead log that corresponds to only one data change record, the data change record is established as a single independent page group log to facilitate subsequent expansion.

[0062] The write-ahead log segmentation process in this step forms the basis for subsequent log data compression. By segmenting the write-ahead log based on the data pages corresponding to each data change record, data change records belonging to the same data page are grouped together into the same page group log. This allows subsequent log compression to be performed on a per-page group log basis, preventing multiple data change records for the same data page from occupying excessive memory, thereby reducing master node memory usage.

[0063] The data page corresponding to a data change record can be directly retrieved from the write-ahead log entry. In the write-ahead log, each data change record includes at least the following key fields: transaction ID, operation type, page ID, modification content, and log sequence number. By extracting the page ID corresponding to each data change record in the write-ahead log, the corresponding data page can be retrieved.

[0064] On the other hand, the operation of segmenting the write-ahead log can use a hash table or a dictionary to quickly segment and group the data pages corresponding to each data change record. This embodiment does not limit the specific segmentation processing method.

[0065] In particular, in one possible implementation, if all data change records in the write-ahead log correspond to the same data page, the write-ahead log splitting step is not performed. Instead, if the write-ahead log meets the compression and merging rules, the page logs corresponding to all data change records in the write-ahead log are directly compressed and merged, thereby improving processing efficiency.

[0066] S103: When the log attribute of any of the page group logs meets the preset merging rule, the page logs in the page group log are merged and compressed to generate compressed and merged page data.

[0067] After the write-ahead log is segmented, the resulting page group logs are monitored in real time. When the log attributes of any page group log meet the pre-set merge rules, the page logs in the page group log are merged and compressed to obtain compressed and merged page data, which is then stored in the cache for the page data.

[0068] For the complete write-ahead log splitting and merging and compression process, please refer to Figure 3, this figure is a schematic diagram of another log compression processing flow provided by an embodiment of the present application. In the figure, the page log cache is used to indicate the situation where a data page corresponds to only one data change record. As can be seen from the figure, when the page group log meets the preset merging rules, the page logs in the page group log will be compressed and merged into a single page data, so that multiple page logs that occupy a single page memory are compressed and merged in the page data cache that only occupies one page memory, thereby achieving the effect of reducing memory usage.

[0069] In the embodiment of the present application, the preset merge rules are used to represent the prerequisites for page group logs to be able to perform compression and merging. The preset merge rules include a first rule, a second rule, and a third rule. When the log attributes of the page group log meet at least one of the three rules, the page logs in the page group log are merged and compressed.

[0070] Among them, the first rule is a rule set for the log accumulation amount of the page group log. A page size threshold is set in the first rule. The page size threshold is used to monitor the log accumulation amount of the page group log in real time. When the log accumulation amount in the page group log reaches the page size threshold set by the preset merge rule, it is determined that the log accumulation amount in the page group log has reached the maximum tolerance range of the page data. At this time, the page group log can be compressed and merged, and the multiple page logs inside it can be compressed and merged into one page data, so as to reduce the memory occupation of multiple page logs.

[0071] The second rule is set for the accumulation time of the page group log. In order to prevent the replay of the slave node log, this application sets a timer threshold based on the data loss risk tolerated by the business. When the page group log is not merged and the cache survival time is greater than the set timer threshold, the second rule will be triggered, and the page logs in the page group log will be compressed and merged.

[0072] The third rule is a rule set for the cache capacity of the write-ahead log. In order to prevent the cache capacity from being exhausted, when the remaining cache capacity of the write-ahead log is lower than a preset ratio (e.g., 10%, 20%), a compression and merging operation of the page group logs is automatically triggered, thereby alleviating the cache capacity of the write-ahead log. In one possible implementation, when the third rule triggers a compression and merging operation on multiple page group logs, the multiple page group logs can be compressed and merged in sequence according to the ranking of their memory usage, with priority given to compressing and merging page log groups with many page logs and large memory usage. When the remaining cache capacity of the write-ahead log is greater than the preset ratio, the compression and merging operation on the page group logs can be stopped.

[0073] In this way, the embodiment of the present application can effectively compress and merge multiple page logs belonging to the same data page in the pre-log into a single page data through the segmentation processing of the pre-write log and the merging and compression based on the page log group, so that its memory occupation is converted from multiple pages to only a single page, thereby achieving the effect of effectively reducing memory occupation.

[0074] In particular, when merging and compressing the page logs in the page group log, the merging and compression algorithm that may be involved needs to be closely integrated with the page-level modification characteristics of the database pre-write log to achieve efficient integration of multiple operation records on the same page or page group. First, based on the structural characteristics of page operations, a "difference merging algorithm" can be used, that is, for multiple modification logs of the same page, the differentiated data segments of each operation on the page are extracted, and the key modification areas are identified by analyzing the operation type (such as addition, deletion, and modification). Only the differentiated content that finally takes effect is retained or merged into overlay records in the order of operations to avoid repeated storage of redundant data on the same page. For example, for multiple write operations on the same page, the subsequent write operation can directly overwrite the corresponding data block of the previous operation, and only the final page status and version identifier are recorded, thereby reducing the log volume.

[0075] Secondly, the "Incremental Encoding Merge Algorithm" is designed to handle cross-page related operations within a page group. By identifying inter-page dependencies (such as foreign key relationships and index references), it collaboratively compresses multiple related page operation logs. This algorithm extracts common data prefixes or shared data blocks across pages within a page group, normalizes duplicate data using dictionary encoding or hash mapping, and only records incremental changes relative to a common base version. Furthermore, the "Page Snapshot Merge Algorithm" targets write-intensive scenarios. When log records for a page reach a certain number (reaching the page size or a set capacity threshold), it consolidates all modification logs for that page to generate a final snapshot of the page's state, discarding any intermediate log records. The snapshot retains only the page version number, modification timestamp, key operation metadata (such as LSN), and final data content, thereby converting dispersed log operations into compact page-level snapshot data, significantly reducing memory usage. For write-sparse scenarios, the algorithm automatically degenerates to a lightweight merge mode, performing simple concatenation or differential recording of a small number of logs, ensuring adaptive optimization in different scenarios.

[0076] The above are the types of merge compression algorithms used in different situations when performing merge compression processing in the embodiments of the present application. The embodiments of the present application do not limit the algorithm used to perform merge compression processing.

[0077] As can be seen from the foregoing description, in a data synchronization scheme based on pulling from a slave node, the master node will not actively send the write-ahead log to the slave node, but will wait for the data pull instruction from the slave node. Similarly, in the embodiment of the present application, after completing the compression and merging processing of the page group log, the master node needs to record the data version identifier of the compressed and merged page data to mark the modified version corresponding to the page data. When the master node responds to the data pull instruction of the slave node, it is necessary to compare the data version identifier of the slave node with the data version identifier of the current page data cache to determine the page data that needs to be fed back to the slave node. Next, this process will be introduced in conjunction with the drawings of specific embodiments.

[0078] See also Figure 4 , which is a flow chart of a method for pulling data from a response node provided in an embodiment of the present application, and the method specifically includes the following steps:

[0079] S201: Recording a first data version identifier corresponding to the compressed and merged page data, so that the slave node can synchronize data through the first data version identifier, the second data version identifier of the internal data of the slave node, the first preset network device, and the second preset network device set in the slave node;

[0080] S202: Determine whether the second data version identifier is smaller than the first data version identifier;

[0081] S203: When the second data version identifier is smaller than the first data version identifier, the slave node pulls the compressed and merged page data and the unmerged data change records;

[0082] S204: When the second data version identifier is greater than the first data version identifier, the slave node only pulls data change records in the write-ahead log whose data version identifier is greater than the second data version identifier.

[0083] After the master node completes the compression and merging of the page log group to obtain the page data, in order to indicate the changed version of the page data, it is necessary to record a separate corresponding first data version identifier and store it in a data structure (such as a page information table) that can be directly accessed in the master node memory. Among them, the first data version identifier can be a timestamp, an increasing serial number, a hash value, etc. The first data version identifier is strongly associated with the compressed and merged page data, and is used to accurately characterize the change timing of the page.

[0084] In a possible implementation, the first data version identifier can be generated by incrementing a serial number, and a counter is globally incremented after each compression and operation to generate a version identifier for each page data after compression and merging.

[0085] When the slave node performs data synchronization, it interacts with the first preset network device (also a preconfigured RDMA network card device or NPU network card device) in the master node through the second preset network device (preconfigured RDMA network card device or NPU network card device). The master node uses the direct memory access capability of its internal first preset network device to actively read the page information table in the master node memory, obtain the first data version identifier of the target page, and at the same time call the second data version identifier maintained by itself (that is, the latest data version currently synchronized by the slave node) for version comparison. Throughout the process, the first preset network device (RDMA / NPU network card device) set in the master node cooperates with the slave node CPU to undertake the core operations of data retrieval, version comparison and network transmission. There is no need for the master node CPU to intervene in parsing or responding, and data interaction is completed only through the memory access interface at the hardware level. The master node only needs to write the version identifier and storage address of the compressed page into a fixed data structure when processing transactions, and is completely in a passive response state, thereby saving the computing resources of the master node. During the entire data synchronization process, the network card device in the slave node mainly focuses on memory data access and data transmission, but is not sufficient for complex operations. In this way, the slave node can obtain the memory data of the master node through the built-in network device without consuming the CPU resources of the master node, and then perform subsequent analysis and processing. Its pressure load is much less than that of the master node used for reading and writing.

[0086] This design deeply integrates the version identification system with hardware acceleration capabilities: First, the data version identifier serves as the "identity anchor" of the compressed page, ensuring transaction integrity (identifying the atomic boundaries of merge operations) while providing a precise synchronization index for slave nodes. RDMA / NPU network card devices, through direct hardware-level access mechanisms, enable slave nodes to independently complete version comparisons and data pulls without intervention from the master node CPU. This ultimately achieves a highly efficient architecture of "lightweight recording on the master node and hardware-accelerated synchronization on the slave nodes," significantly reducing master node resource consumption while improving data synchronization efficiency.

[0087] Among them, the second data version identifier is used to represent the latest version identifier of all data pages in the slave node (for example, PageID=0x1A3B, Version=1000). Specifically, during the data synchronization process, the slave node will decide the data synchronization method based on the identifier size comparison between the second data version identifier and the first data version identifier. When the second data version identifier is smaller than the first data version identifier, it means that the data version identifier in the slave node lags behind the version identifier of the master node. At this time, the slave node needs to pull all the page cache data in the master node and all the change records of the corresponding pages, that is, send the complete compressed and merged page data to the slave node, so as to ensure that the slave node can quickly update the latest version of the master node.

[0088] On the contrary, when the second data version identifier is greater than the first data version identifier, the slave node needs to pull the incremental log with a higher version in the pre-write log. Therefore, when the second data version identifier of the slave node is greater than the first data version identifier of the page data cache in the master node, it is necessary to pull the data change record in the pre-write log with a data version identifier greater than the second data version identifier (i.e., the incremental log of the pre-write log) to synchronize the latest data in the master node.

[0089] In particular, in one possible implementation, in order to ensure the effective implementation of data synchronization while saving bandwidth and network resources, different data transmission methods can be used during data synchronization from the node based on two different scenarios: write-intensive scenarios and write-sparse scenarios.

[0090] Among them, when the response frequency of the master node to the data change request is greater than a preset threshold, the master node is in a write-intensive scenario. In this scenario, since the master node needs to process a large number of data change requests, the slave node may not be able to synchronize data in time, resulting in log accumulation. Therefore, in order to address this problem, the master node needs to regularly generate a replica page of the page data and record all change operations since the replica page was generated (i.e., the incremental log of the pre-write log) to prevent network congestion.

[0091] Accordingly, when the master node's response frequency to the data change request is no greater than the preset threshold, the master node is in a write-sparse scenario. In this scenario, the master node only processes a small number of data change requests, and the master and slave nodes have high requirements for data real-time performance. Therefore, in order to ensure the consistency of data synchronization between the master and slave nodes, the slave node needs to pull the change data corresponding to each data change record in real time to form an incremental log stream, thereby achieving high real-time performance and consistency of data synchronization, while also saving bandwidth.

[0092] In addition, in one possible implementation, different log merging methods can be used for write-sparse scenarios and write-intensive scenarios. For example, for write-sparse scenarios, when the system usually writes logs at a lower frequency (such as monitoring logs and periodic business logs), the log data will gradually accumulate in multiple small files. Although the capacity of each of these small files is not large, the large number of them will result in a large number of files needing to be scanned during queries, increasing IO overhead and query time. In this scenario, idle merging is used to merge these small files into larger log files when the system is idle, reducing the total number of files and optimizing the scanning efficiency during subsequent queries.

[0093] In contrast, in write-intensive scenarios, log data is continuously written at an extremely high frequency, generating a large number of small files in a short period of time. If not merged promptly, these files will quickly accumulate, increasing the pressure on metadata management at the storage layer and exponentially increasing the number of files that need to be traversed during queries, ultimately affecting data retrieval efficiency. In this scenario, a scheduled log merge strategy can be adopted. Unlike off-peak merging, which relies on the system's real-time load, scheduled merging is clearly planned and triggers the merge task at a set time point, regardless of whether the current system is busy or not. Essentially, this is to prevent the infinite fragmentation of log files due to high-frequency writes, thereby avoiding a sharp drop in subsequent query performance. Scheduled merging regularly consolidates log files generated within a certain period of time through regular task execution. For example, all small log files for the day are merged into one or several large files at dawn each day. This not only reduces the number of files, but also allows subsequent queries to quickly locate target files by time slice, significantly improving scanning speed.

[0094] The embodiment of the present application provides a log data compression method, in which, first, in response to a data change request, multiple data change records corresponding to the data change request are written into a pre-write log. Subsequently, because the storage unit of the database is a data page, and each data change record has a corresponding data page, the pre-write log can be segmented based on the data page corresponding to each data change record to obtain multiple page group logs. In the page group log, since each page log corresponds to the same data page or a set of logically related pages. Therefore, when the log attribute of any page group log in the multiple page group logs meets the preset merge rule, the page logs in the page group log are merged and compressed, and multiple page logs for the same data page can be merged into a series of operation results in one page. Since the memory occupied by a series of operation results in one page is only one page, the space occupied by the pre-write log can be effectively controlled. Even if multiple data change operations are performed on a data page in the pre-write log, the page log corresponding to each data change operation will eventually be compressed and merged into the content of one page data, effectively solving the problem of excessive memory usage of the master node in data synchronization.

[0095] A log data compression device provided in an embodiment of the present application is introduced below. The log data compression device described below and the log data compression method described above can refer to each other.

[0096] See also Figure 5 , which is a schematic diagram of the structure of a log data compression device provided by an embodiment of the present application, specifically including the following modules:

[0097] The log writing module 100 is configured to write a plurality of data change records into a write-ahead log in response to a data change request; one data change record corresponds to a single data page;

[0098] A segmentation processing module 200 is configured to segment the write-ahead log based on the data pages corresponding to the plurality of data change records to obtain a plurality of page group logs; the page group logs include a plurality of page logs, each of which corresponds to the same data page or a set of logically associated pages;

[0099] The merging and compressing module 300 is configured to merge and compress the page logs in any page group log when the log attributes of the page group log meet a preset merging rule, so as to generate compressed and merged page data.

[0100] In a possible implementation, the preset merging rules include: a first rule, a second rule, and a third rule; the first rule includes: the log accumulation amount of the page group log reaches a page size threshold; the second rule includes: the accumulation time of the page group log reaches a timer threshold; the third rule includes: the remaining capacity of the cache area of ​​the write-ahead log is lower than a preset ratio value;

[0101] The merging and compression module 300 is specifically configured to:

[0102] When the log attribute of the page group log satisfies at least one of the first rule, the second rule, and the third rule, the page logs in the page group log are merged and compressed to generate the compressed and merged page data.

[0103] See also Figure 6 , which is a structural diagram of a log data compression electronic device provided by an embodiment of the present application, including:

[0104] Memory 11, for storing computer programs;

[0105] The processor 12 is configured to implement the steps of a log data compression method described in any of the above method embodiments when executing the computer program.

[0106] In this embodiment, the device may be a vehicle-mounted computer, a PC (Personal Computer), or a terminal device such as a smart phone, a tablet computer, a PDA, or a portable computer.

[0107] The device may include a memory 11 , a processor 12 , and a bus 13 .

[0108] The memory 11 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the device, such as the hard disk of the device. In other embodiments, the memory 11 may also be an external storage device of the device, such as a plug-in hard disk equipped on the device, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. Furthermore, the memory 11 may also include both an internal storage unit of the device and an external storage device. The memory 11 can not only be used to store application software installed on the device and various types of data, such as program code for executing a fault prediction method, but can also be used to temporarily store data that has been output or is to be output. In some embodiments, the processor 12 may be a central processing unit (CPU).

[0109] In some embodiments, the processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, used to run the program code stored in the memory 11 or process data, such as the program code for executing the log data compression method.

[0110] The bus 13 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0111] Furthermore, the device may also include a network interface 14, which may optionally include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the device and other electronic devices.

[0112] Optionally, the device may further include a user interface 15, which may include a display and an input unit such as a keyboard. The optional user interface 15 may also include a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a display screen or display unit, and is used to display information processed in the device and to display a visual user interface.

[0113] Figure 6 Only the device with components 11-15 is shown, and it will be understood by those skilled in the art that Figure 6 The structure shown does not constitute a limitation of the device, and may include fewer or more components than shown, or combine certain components, or arrange the components differently.

[0114] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, an embodiment of the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the log data compression method described in any of the above embodiments.

[0115] The computer-readable media of the embodiments of the present application include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0116] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the log data compression method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0117] It should be noted that the various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for methods, devices, electronic devices and media, since they are basically similar to the method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiments. The methods, devices, electronic devices and media described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement them without expending any creative effort.

[0118] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A log data compression method, characterized in that: Applied to a master node for performing data synchronization, the method comprises: In response to a data change request, write multiple data change records into a write-ahead log; one data change record corresponds to a single data page; Based on the data pages of each data change record, the write-ahead log is segmented to obtain a plurality of page group logs; the page group log includes a plurality of page logs, and each of the page logs corresponds to the same data page or a set of logically associated pages; When the log attribute of any of the page group logs meets the preset merging rule, the page logs in the page group log are merged and compressed to generate compressed and merged page data.

2. The method according to claim 1, characterized in that The preset merging rules include: a first rule, a second rule, and a third rule; the first rule includes: the log accumulation amount of the page group log reaches a page size threshold; the second rule includes: the accumulation time of the page group log reaches a timer threshold; the third rule includes: the remaining capacity of the cache area of ​​the write-ahead log is lower than a preset ratio value; When the log attribute of any of the page group logs satisfies the preset merging rule, merging and compressing the page logs in the page group log to generate compressed and merged page data, including: When the log attribute of the page group log satisfies at least one of the first rule, the second rule, and the third rule, the page logs in the page group log are merged and compressed to generate the compressed and merged page data.

3. The method according to claim 1, characterized in that The master node includes: a first preset network device. After generating the compressed and merged page data, the method further includes: For the compressed and merged page data, a separately corresponding first data version identifier is recorded, so that the slave node performs data synchronization through the first data version identifier, the second data version identifier of the internal data of the slave node, the first preset network device, and the second preset network device set in the slave node; The second preset network device is used to interact with the first preset network device so that the slave node can pull the memory data of the master node.

4. The method according to claim 3, characterized in that The step of synchronizing data by the slave node based on the first data version identifier, the second data version identifier of the internal data of the slave node, and the second preset network device in the slave node includes: When the second data version identifier is smaller than the first data version identifier, the slave node pulls the compressed and merged page data and the unmerged data change record; When the second data version identifier is greater than the first data version identifier, the slave node only pulls data change records in the write-ahead log whose data version identifier is greater than the second data version identifier; After pulling the above data, the slave node sets the second data version identifier as the identifier of the latest data change record.

5. The method according to claim 3, characterized in that The method further comprises: When the response frequency of the master node to the data change request is greater than a preset threshold, data synchronization of the slave node is completed through the replica page of the page data and the incremental log of the write-ahead log; When the response frequency of the master node to the data change request is not greater than the preset threshold, the data synchronization of the slave node is completed based on the incremental log stream of the write-ahead log.

6. The method according to claim 1, characterized in that Before the write-ahead log is segmented, the method further includes: If all the data change records in the write-ahead log correspond to the same data page, the splitting process step for the write-ahead log is not performed.

7. A log data compression device, characterized in that: Applied to a master node, the device comprises: a log writing module, configured to write a plurality of data change records into a write-ahead log in response to a data change request; wherein one data change record corresponds to a single data page; a segmentation processing module for segmenting the write-ahead log based on the data pages corresponding to the plurality of data change records to obtain a plurality of page group logs; the page group logs comprising a plurality of page logs, each of which corresponds to the same data page or a set of logically associated pages; The merging and compressing module is used to merge and compress the page logs in any page group log when the log attribute of the page group log meets the preset merging rule, so as to generate compressed and merged page data.

8. The device according to claim 7, characterized in that The preset merging rules include: a first rule, a second rule, and a third rule; the first rule includes: the log accumulation amount of the page group log reaches a page size threshold; the second rule includes: the accumulation time of the page group log reaches a timer threshold; the third rule includes: the remaining capacity of the cache area of ​​the write-ahead log is lower than a preset ratio value; The merging and compression module is specifically used to: When the log attribute of the page group log satisfies at least one of the first rule, the second rule, and the third rule, the page logs in the page group log are merged and compressed to generate the compressed and merged page data.

9. An electronic device, characterized in that: The device includes: a processor, a memory and a system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the log data compression method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the log data compression method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Database transaction processing method and device and electronic equipment

    CN115145697A

  • Rapid data file merging method and system for time sequence database

    CN116561120A

  • Data processing method and device and computing equipment

    CN118535638A

  • Data page reworking method and device, equipment, medium and product

    CN118838882A

  • Method and means for archiving in a transaction management system

    EP0625752A2