Transaction log processing method of database, storage medium and equipment
By storing the transaction logs separately into assembled linked lists and page data link lists, only other data except page mirror data are synchronized, and the data page mirror data of the same table are merged for disk, the problem of poor synchronous flow replication performance caused by large transaction log data is solved, and synchronization efficiency and database security and stability are improved.
Patent Information
- Application Number
- CN202311865076.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
AI Technical Summary
In the database, the large amount of data in the transaction log results in a low replication performance of synchronous streams, affecting the data consistency and synchronization efficiency of the primary and secondary databases.
The transaction log is stored separately as an assembly linked list and a page data linked list. The page data linked list is stored in the global page data linked list. Only other data except the page mirror data are synchronized, and the data page mirror data of the same table is merged for disk dropping, reducing synchronization amount and disk IO pressure.
It improves the synchronization efficiency of the primary and secondary databases, ensures the security and stability of the database, and reduces the disk IO pressure.
Smart Images

Figure CN120277153A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of databases, and in particular, to a method for processing transaction logs of a database, a storage medium, and a device. Background Art
[0002] During the use of a database, transaction logs are recorded. An important function of the transaction logs is to construct a primary and standby database cluster. Briefly speaking, after the primary and standby database cluster is constructed, the primary database provides read and write services externally and continuously generates transaction logs. The transaction logs are transmitted to the standby database through streaming replication, so that the standby database is almost identical to the primary database.
[0003] In order to keep the data of the primary database and the standby database strictly consistent, the synchronous streaming replication mode is usually adopted. In the synchronous streaming replication mode, after the primary database sends a transaction log to the standby database, it needs to wait for the standby database to store this transaction log locally or replay it successfully before the primary database can continue to execute. Therefore, the amount of data of the transaction logs that the primary database needs to transmit to the standby database is an important factor affecting the performance of synchronous streaming replication. Summary of the Invention
[0004] An object of the present invention is to provide a method for processing transaction logs of a database, a storage medium, and a device that can reduce the amount of data of the transaction logs to be transmitted.
[0005] In particular, the present invention provides a method for processing transaction logs of a database. The transaction logs include an assembled linked list and a page data linked list separately stored in a cache. The page data linked list is a linked list composed of page mirror data in the transaction logs, and the assembled linked list is a linked list composed of other data in the transaction logs except for the page mirror data. The page data linked list is stored in a preset global page data linked list, and the transaction log processing method includes:
[0006] Obtaining a trigger event for flushing a data page in the cache to disk;
[0007] Obtaining a data page to be flushed in the cache;
[0008] Obtaining page mirror data corresponding to the data page to be flushed in the global page data linked list, denoted as basic page mirror data;
[0009] Obtaining co-table data pages, where the co-table data pages are data pages that need to be flushed and belong to the same table as the data page to be flushed;
[0010] Obtaining page mirror data corresponding to the co-table data pages in the global page data linked list;
[0011] Assemble the page mirror data corresponding to all the same-table data pages behind the base page mirror data, denoted as combined page mirror data;
[0012] Generate a combined data storage file on the disk;
[0013] Store the combined page mirror data into the combined data storage file;
[0014] Store the data pages to be disk-written into the disk.
[0015] Optionally, the file name of the combined data storage file has at least the table location information of the table to which the data pages to be disk-written belong.
[0016] Optionally, before the step of generating a combined data storage file on the disk, it includes:
[0017] Obtain the block numbers and log sequence numbers corresponding to all the page mirror data in the combined page mirror data;
[0018] Obtain the maximum block number and the minimum block number among all the block numbers corresponding to the page mirror data;
[0019] Obtain the maximum log sequence number and the minimum log sequence number among all the log sequence numbers corresponding to the page mirror data;
[0020] The step of generating a combined data storage file on the disk includes:
[0021] Generate a combined data storage file whose file name records the table location information of the table to which the data pages to be disk-written belong, the maximum block number, the minimum block number, the maximum log sequence number, and the minimum log sequence number.
[0022] Optionally, after the step of obtaining the block numbers and log sequence numbers corresponding to all the page mirror data in the combined page mirror data, it includes:
[0023] Sort all the page mirror data in the combined page mirror data in ascending order of the block numbers;
[0024] Check whether there are page mirror data with the same block number in the combined page mirror data. If so, sort the page mirror data with the same block number in ascending order of the log sequence numbers.
[0025] Optionally, after the step of generating a combined data storage file on the disk, it includes:
[0026] Add a file header to the combined data storage file, where the content of the file header includes the number of page mirror data in the combined page mirror data, an array of block numbers, an array of log sequence numbers, and an array of starting addresses. Among them, the block numbers in the array of block numbers store the block numbers of all page mirror data in the combined page mirror data, the log sequence numbers in the array of log sequence numbers store the log sequence numbers of all page mirror data in the combined page mirror data, and the starting addresses in the array of starting addresses store the starting addresses of all page mirror data in the combined page mirror data.
[0027] Optionally, before the step of obtaining the same-table data pages, it includes:
[0028] Check whether there are such same-table data pages. If so, execute the step of obtaining the same-table data pages;
[0029] If not, generate a single storage file on the disk;
[0030] Store the basic page mirror data into the single storage file;
[0031] Individually store the data pages to be disk-written into the disk.
[0032] Optionally, before the step of obtaining the page mirror data corresponding to the same-table data pages in the global page data linked list, it includes:
[0033] Check whether there is corresponding page mirror data for all the same-table data pages in the global page data linked list. If so, execute the step of obtaining the page mirror data corresponding to the same-table data pages in the global page data linked list;
[0034] If not, generate a single storage file on the disk;
[0035] Store the basic page mirror data into the single storage file;
[0036] Store the data pages to be disk-written into the disk.
[0037] Optionally, the step of storing the data pages to be disk-written into the disk includes:
[0038] Store the data pages to be disk-written and all the same-table data pages into the disk.
[0039] In another aspect, the present invention also provides a machine-readable storage medium, on which a machine-executable program is stored. When the machine-executable program is executed by a processor, it implements the transaction log processing method of the database according to any one of the above.
[0040] In another aspect, the present invention also provides a computer device, including a memory, a processor, and a machine-executable program stored on the memory and running on the processor, and when the processor executes the machine-executable program, it implements the transaction log processing method of the database according to any one of the above.
[0041] In the transaction log processing method, storage medium and device of the present invention, the transaction log with page mirror data is separately stored in the cache in the form of an assembled linked list and a page data linked list, and the page data linked list is stored in a preset global page data linked list, that is, the page mirror data in the transaction log is separated and separately stored in the global page data linked list. And, in the subsequent disk writing process, when a trigger event for disk writing of the data page in the cache is obtained, a data page to be disk-written in the cache is obtained, the page mirror data corresponding to the data page to be disk-written in the global page data linked list is obtained, denoted as the basic page mirror data, the same-table data page is obtained, the page mirror data corresponding to the same-table data page in the global page data linked list is obtained, and the page mirror data corresponding to all the same-table data pages is assembled behind the basic page mirror data, denoted as the combined page mirror data. A combined data storage file is generated on the disk, the combined page mirror data is stored in the combined data storage file, and the data page to be disk-written and all the same-table data pages are stored in the disk. On the one hand, the page mirror data of the transaction log and other data except the page mirror data can be separately disk-written, that is, they are not disk-written as a whole. Therefore, when this database synchronizes the transaction log to other standby databases, it can only synchronize the data except the page mirror data, thereby reducing the data synchronization amount between the master and standby databases and helping to improve the synchronization efficiency. And the page mirror data is still stored on the disk, which can also ensure the security and stability of this database. On the other hand, by merging the page mirror data corresponding to the data pages to be disk-written belonging to the same table and then performing disk writing, the number of disk writing times of the page mirror data can be reduced, thereby reducing the pressure on disk I / O (Input / Output).
[0042] Those skilled in the art will become more clear about the above and other objects, advantages and features of the present invention according to the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Some specific embodiments of the present invention will be described in detail hereinafter with reference to the accompanying drawings in an exemplary but non-limiting manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:
[0044] Figure 1 is a schematic diagram of a transaction log in the prior art;
[0045] Figure 2 is the first schematic diagram of a transaction log according to an embodiment of the present invention;
[0046] Figure 3 is the second schematic diagram of a transaction log according to an embodiment of the present invention;
[0047] Figure 4 is the schematic diagram of the page data chain of a transaction log according to an embodiment of the present invention;
[0048] Figure 5 is the schematic diagram of the page mirror data content of a transaction log according to an embodiment of the present invention;
[0049] Figure 6 is the third schematic diagram of a transaction log according to an embodiment of the present invention;
[0050] Figure 7 is the schematic diagram of the global page data linked list according to an embodiment of the present invention;
[0051] Figure 8 is the schematic flowchart of the transaction log processing method according to an embodiment of the present invention;
[0052] Figure 9 is the schematic diagram of the combined page mirror data according to an embodiment of the present invention;
[0053] Figure 10 is the schematic flowchart of the transaction log processing method according to another embodiment of the present invention;
[0054] Figure 11 is the schematic flowchart of the transaction log processing method according to yet another embodiment of the present invention;
[0055] Figure 12 is the schematic flowchart of the transaction log processing method according to yet another embodiment of the present invention;
[0056] Figure 13 is the schematic diagram of the combined page mirror data according to another embodiment of the present invention;
[0057] Figure 14 is the schematic diagram of a machine-readable storage medium according to an embodiment of the present invention;
[0058] Figure 15 is the schematic diagram of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0059] Those skilled in the art should understand that the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. This part of the embodiments is intended to explain the technical principles of the present invention, rather than to limit the protection scope of the present invention. Based on the embodiments provided by the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts should still fall within the protection scope of the present invention.
[0060] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices.
[0061] The flowchart provided by the present invention is not intended to indicate that the operations of the method will be executed in any specific order, or that all operations of the method are included in every case. In addition, the method may include additional operations. Within the scope of the technical concept provided by the method of this embodiment, additional changes can be made to the above method.
[0062] To facilitate the understanding of this solution, the operating mode of the database in the prior art will be described. First, during the process of modifying the database data, the table stored on the disk is not directly modified. Instead, the data in the table is extracted into the cache in units of data pages (the data to be modified is located in the data pages), modified in the cache, and then written to disk later.
[0063] Furthermore, when modifying the data pages in the cache, transaction logs will be generated. As Figure 1 shown, the existing transaction logs generally include two parts: a header and a data area. The data area records various modification operations on the data pages. At the same time, there is also a kind of data in the data area that is page mirror data. The page mirror data is the data formed after recording the complete page information of a data page in the transaction log. In the prior art, usually after the log redo point (the log redo point is a log sequence number. Before this log sequence number, the transaction logs have all been written to disk and the information reflected by the written logs and all the data in the actual database on the disk is the same), if the data of a data page in the cache is modified for the first time, then the entire page information of the data page after the first modification will be recorded in the transaction log, thereby generating page mirror data.
[0064] During the process of data page disk write, it is possible that a crash occurs during the disk write of a data page, resulting in a situation where the disk-written data page contains a mixture of old and new data. Therefore, by generating page mirror data, the page information recorded in the page mirror data can be used to replace the broken page with the mixture of old and new data, and then the subsequent operations recorded in the transaction log can be replayed to complete the recovery of the data page. Therefore, page mirror data is important data for ensuring the security and stability of the database where it is located.
[0065] Therefore, page mirror data is important data for ensuring the security and stability of the database where it is located. However, during the synchronization process of the primary and standby databases, even if the page mirror data in the transaction log of the primary database is not synchronized to the standby database, it will not affect the data consistency between the primary and standby databases. That is to say, if it is possible to separate the page mirror data from the transaction log of the primary database on the basis of ensuring that the page mirror data of the primary database can be recorded, the amount of data that the primary database needs to synchronize to the standby database can be reduced on the premise of ensuring the security and stability of the primary database.
[0066] As Figures 2 to 6 shown, in an embodiment of the present application, the transaction log includes an assembled linked list and a page data linked list separately stored in a cache. The page data linked list is a linked list composed of page mirror data in the transaction log, the assembled linked list is a linked list composed of other data in the transaction log except for the page mirror data, and the page data linked list is stored in a preset global page data linked list.
[0067] Specifically, in a database, the length of a transaction log is fixed. As the modification operations continue to increase, the capacity limit of a transaction log will be reached or it will be filled. At this time, the log information generated later needs to be written into the next transaction log. For a filled transaction log, the complete content of the transaction log is determined.
[0068] During the generation process of the transaction log in this embodiment, for a transaction log with page mirror data in its content, when all the content of a transaction log is determined, the transaction log is assembled into a structure composed of an assembled linked list and a page data linked list, and then the assembled linked list and the page data linked list are separately stored in the cache.
[0069] Referring to Figures 2 to 6 shown, specifically, taking a transaction log with two page mirror data nodes in its complete content as an example, Figure 2 is an example of the linked list form of the original transaction log with page mirror data, Figure 3 is Figure 2An example of the linked list form after removing page mirror data from the linked list in []. Here, next indicates pointing to the next node; len represents the len function, which is used to calculate the length of the corresponding data; head data represents the header data of the transaction log; page data represents all page mirror data in the transaction log; tupledata and main data represent other data in the transaction log except page mirror data. Figure 4 A simplified example of the linked list of two page mirror data nodes in the transaction log. wal1 page data1 and wal1 pagedata2 are just simple notations used to represent two different page mirror data nodes belonging to the same transaction log, and do not represent specific content. Figure 5 An example of the specific content of a page mirror data node. RelFileNode represents the location information of the table to which the data page corresponding to this page mirror data belongs on the disk, BlockNumber represents the block number of the data page corresponding to this page mirror data in the table to which it belongs. Therefore, RelFileNode and BlockNumber can determine the location of the data page corresponding to the page mirror data on the disk. lsn represents the log sequence number of the transaction log to which the page mirror data belongs. page data is the complete information of the data page recorded by this page mirror data node, and next indicates pointing to the next node. Figure 6 An example of the finally formed transaction log.
[0070] During the assembly process of the transaction log, first remove all page mirror data in the transaction log, as Figures 2 to 3 shown. Then form a page data linked list with the page mirror data in the order in the transaction log, as Figure 4 shown. Then, connect the page data linked list behind the linked list composed of other data in the transaction log, so as to assemble the transaction log into a structure composed of an assembled linked list and a page data linked list connected after the assembled linked list, as Figure 6 shown. That is to say, Figure 5 can be regarded as Figure 4 a schematic diagram of the specific content of a page mirror data node in []. Figure 6 The assembled linked list in [] represents Figure 3 the linked list of the transaction log in [], that is, the linked list composed of other data in the transaction log except page mirror data.
[0071] Then, during the process of writing the transaction log into the cache, store the assembled linked list and the page data linked list separately. And, the page data linked list is stored in the global page data linked list. That is to say, the global page data linked list is a linked list formed by connecting page data linked lists to each other.
[0072] Refer to Figure 7As shown, it is a simple illustration of the global page data linked list storing two page data linked lists. page_data_list_local represents the page data linked list of a transaction log, and page_data_list_global represents the global page data linked list. Among them, wal1 and wal2 are only used to indicate that the two page data linked lists belong to different transaction logs.
[0073] As Figure 8 shown, in this embodiment, the transaction log processing method generally includes:
[0074] Step S101, obtaining a trigger event for flushing the data pages in the cache to disk. Referring to the foregoing, after the data pages are extracted into the cache, they are continuously modified in the cache and finally need to be flushed to disk. That is, it is detected that the data pages in the cache need to be flushed to disk.
[0075] It should be noted that for the database, it has a trigger event for flushing the data in the cache. Exemplarily, the trigger event can be an event whose time interval from the previous flush point reaches a preset interval. In this embodiment, the data in the cache includes data pages, page mirror data of the transaction log, and data of the assembled linked list part of the transaction log. After the database starts to flush, it is necessary to first flush the data of the assembled linked list of the transaction log formed between the previous flush point and the current flush point in the cache. Because the assembled linked list contains other data except the page mirror data, that is, it includes the operations on the data pages.
[0076] After the assembled linked list of the transaction log formed between the previous flush point and the current flush point is flushed to disk, for the data pages in the cache, if the operation logs corresponding to a data page from the previous flush point to the current flush point have all been flushed, in other words, if the operations recorded in the data of the assembled linked list of the transaction log that has been flushed in this disk flushing work already include all the operations performed on a data page from the previous flush point to the current flush point, then this data page can be flushed in this disk flushing work.
[0077] Therefore, the trigger event for flushing the data pages is that the assembled linked list part of the transaction log that needs to be flushed in this disk flushing work has all been flushed to disk.
[0078] In other words, this step may include: obtaining a trigger event for flushing the data in the cache; flushing the assembled linked list part of the transaction log in the cache; detecting that the assembled linked list part of the transaction log in the cache has been flushed. That is, it is determined that a trigger event for flushing the data pages in the cache is obtained when it is detected that the assembled linked list part of the transaction log in the cache has been flushed.
[0079] Step S102: Obtain a data page to be flushed from the cache. Start flushing the data pages in the cache, and obtain a data page to be flushed from the data pages in the cache that need to be flushed. As described above, the data pages in the cache that need to be flushed are the data pages for which the operation logs corresponding to the period from the previous flush point to the current flush point have all been flushed.
[0080] Step S103: Obtain the page mirror data corresponding to the data page to be flushed in the global page data linked list, denoted as the basic page mirror data. Specifically, after obtaining a data page to be flushed, if the data page to be flushed has corresponding page mirror data in the global page data linked list, then obtain its corresponding page mirror data.
[0081] Exemplarily, denote the data page to be flushed as page_cur, the table it belongs to in the disk as RelFileNode_cur, and the block number corresponding to its belonging table as BlockNumber_cur. Refer to Figure 5 As shown, since the page mirror data records RelFileNode and BlockNumber, it is possible to compare RelFileNode_cur and BlockNumber_cur with the RelFileNode and BlockNumber recorded in the page mirror data to confirm whether there is corresponding page mirror data in the global page data linked list for the data page to be flushed.
[0082] Step S104: Obtain the same-table data pages. The same-table data pages are the data pages that need to be flushed and belong to the same table as the data page to be flushed. Specifically, after obtaining a data page to be flushed, in this flushing operation, if there are data pages that belong to the same table as the data page to be flushed among the other data pages that need to be flushed, then obtain the data pages that belong to the same table as the data page to be flushed.
[0083] As described above, it is possible to check whether the RelFileNode of the other data pages that need to be flushed is the same as RelFileNode_cur. If they are the same, it means they belong to the same table as the data page to be flushed.
[0084] Step S105: Obtain the page mirror data corresponding to the same-table data pages in the global page data linked list. Specifically, if the same-table data pages also have corresponding page mirror data in the global page data linked list, obtain the corresponding page mirror data.
[0085] Step S106: Assemble the page mirror data corresponding to all the same-table data pages behind the basic page mirror data, denoted as the combined page mirror data. Specifically, if the page mirror data of the same-table data pages is obtained, then assemble the obtained page mirror data of the same-table data pages behind the basic page mirror data.
[0086] Refer to Figure 9 As shown, it is an example of combined page mirror data. The page data corresponding to page_cur is the page mirror data corresponding to the data page to be written to disk, that is, the basic page mirror data. The page data corresponding to the same-table data page is the page mirror data of the same-table data page. Among them, len represents the length information of each page mirror data.
[0087] Step S107, generate a combined data storage file on the disk. Specifically, that is to generate a storage file on the disk for storing the combined page mirror data. Preferably, the file name of the combined data storage file has at least the table location information of the data page to be written to disk.
[0088] In one embodiment, after obtaining the page mirror data corresponding to the data page to be written to disk, read the log sequence number of the transaction log to which the page mirror data corresponding to the data page to be written to disk belongs, denoted as lsn_cur. When generating the storage file, generate the file name of the storage file with the table location information of the data page to be written to disk, the block number of the data page to be written to disk in the table to which it belongs, and the log sequence number of the transaction log to which the page mirror data corresponding to the data page to be written to disk belongs. Exemplarily, the file name is RelFileNode_BlockNumber_lsn. In the file name, RelFileNode is the table location information of the data page to be written to disk, that is, RelFileNode_cur in the previous text. BlockNumber in the file name is the block number of the data page to be written to disk in the table to which it belongs, that is, BlockNumber_cur in the previous text. lsn in the file name is the log sequence number of the transaction log to which the page mirror data corresponding to the data page to be written to disk belongs, that is, lsn_cur in the previous text.
[0089] Step S108, store the combined page mirror data into the combined data storage file. After the combined data storage file is generated, the combined page mirror data can be stored in the combined data storage file.
[0090] Step S109, write the data page to be written to disk to the disk. After the combined page mirror data is written to disk, the data page to be written to disk can be written to disk.
[0091] In the solution of this embodiment, the transaction log with page mirror data is separately stored in the cache in the form of an assembled linked list and a page data linked list, and the page data linked list is stored in a preset global page data linked list. That is to say, the page mirror data in the transaction log is separated and separately stored in the global page data linked list. Moreover, in the subsequent disk writing process, when a trigger event for disk writing of the data page in the cache is obtained, a data page to be disk-written in the cache is obtained, the page mirror data corresponding to the data page to be disk-written in the global page data linked list is obtained, which is recorded as the basic page mirror data, the same-table data page is obtained, and the page mirror data corresponding to the same-table data page in the global page data linked list is obtained. All the page mirror data corresponding to the same-table data pages is assembled behind the basic page mirror data, which is recorded as the combined page mirror data. A combined data storage file is generated on the disk, the combined page mirror data is stored in the combined data storage file, and the data page to be disk-written is stored in the disk. On the one hand, the page mirror data of the transaction log and other data except the page mirror data can be disk-written separately, that is, they are not disk-written as a whole. Therefore, when this database synchronizes the transaction log to other standby databases, it can only synchronize the data other than the page mirror data, thereby reducing the data synchronization volume between the master and standby databases and helping to improve the synchronization efficiency. Moreover, the page mirror data is still stored on the disk, which can also ensure the security and stability of this database.
[0092] On the other hand, by merging the page mirror data corresponding to the data pages to be disk-written that belong to the same table and then performing disk writing, the number of disk writing times of the page mirror data can be reduced, thereby reducing the pressure on disk I / O (Input / Output).
[0093] Furthermore, by generating at least the table location information of the table to which the data page to be disk-written belongs in the file name of the storage file, it is convenient to search for the page mirror data on the disk subsequently.
[0094] Such as Figure 10 shown, in one embodiment, the step of storing the data page to be disk-written in the disk includes: storing the data page to be disk-written and all the same-table data pages in the disk.
[0095] The transaction log processing method generally includes:
[0096] Step S201, obtaining a trigger event for disk writing of the data page in the cache.
[0097] Step S202, obtaining a data page to be disk-written in the cache.
[0098] Step S203: Detect whether there is page mirror data corresponding to the data page to be written to disk in the global page data linked list. If so, execute Step S204; if not, execute Step S213. Exemplarily, denote the data page to be written to disk as page_cur, the table it belongs to in the disk as RelFileNode_cur, and the block number corresponding to its belonging table as BlockNumber_cur. Refer to Figure 5 As shown, since the RelFileNode and BlockNumber are recorded in the page mirror data, it is possible to compare RelFileNode_cur and BlockNumber_cur with the RelFileNode and BlockNumber recorded in the page mirror data to confirm whether there is page mirror data corresponding to the data page to be written to disk in the global page data linked list.
[0099] Step S204: Obtain the page mirror data corresponding to the data page to be written to disk in the global page data linked list, denoted as the basic page mirror data.
[0100] Step S205: Check whether there are data pages in the same table. If so, execute Step S206; if not, execute Step S214. The data pages in the same table are the data pages that need to be written to disk and belong to the same table as the data page to be written to disk. Referring to the foregoing, it is possible to check whether the RelFileNode of other data pages that need to be written to disk is the same as RelFileNode_cur. If they are the same, it means that they belong to the same table as the data page to be written to disk.
[0101] Step S206: Obtain the data pages in the same table.
[0102] Step S207: Check whether there is corresponding page mirror data in the global page data linked list for all the data pages in the same table. If so, execute Step S208; if not, execute Step S217. Specifically, the checking method refers to the method of detecting whether there is page mirror data corresponding to the data page to be written to disk in the global page data linked list, that is, using RelFileNode and BlockNumber for confirmation.
[0103] Step S208: Obtain the page mirror data corresponding to the data pages in the same table in the global page data linked list.
[0104] Step S209: Assemble the page mirror data corresponding to all the data pages in the same table behind the basic page mirror data, denoted as the combined page mirror data.
[0105] Step S210: Generate a combined data storage file in the disk.
[0106] Step S211: Store the combined page mirror data into the combined data storage file.
[0107] Step S212: Store the data pages to be disk-written and all data pages of the same table into the disk. After the combined page mirror data is successfully written to the disk, the data pages to be disk-written and all data pages of the same table are written to the disk together.
[0108] Step S213: Store the data pages to be disk-written into the disk. If there is no corresponding page mirror data for the data pages to be disk-written in the global page data linked list, only the data pages to be disk-written need to be written to the disk.
[0109] Step S214: Generate a single storage file in the disk. If there are no data pages of the same table, only the page mirror data corresponding to the data pages to be disk-written needs to be written to the disk, so a single storage file is generated. It should be noted that the name of the storage file can also be generated using the table location information of the table to which the data pages to be disk-written belong, the block number of the data pages to be disk-written in the table, and the log sequence number of the transaction log to which the page mirror data corresponding to the data pages to be disk-written belongs.
[0110] Step S215: Store the basic page mirror data into the single storage file. After the single storage file is generated, the basic page mirror data corresponding to the data pages to be disk-written is stored in the single storage file.
[0111] Step S216: Store the data pages to be disk-written into the disk. After the basic page mirror data is stored in the single storage file, the data pages to be disk-written are stored in the disk.
[0112] Step S217: Generate a single storage file in the disk. If there are data pages of the same table but no corresponding page mirror data for the data pages of the same table, only the page mirror data corresponding to the data pages to be disk-written needs to be written to the disk, so a single storage file is generated. It should be noted that the name of the storage file can also be generated using the table location information of the table to which the data pages to be disk-written belong, the block number of the data pages to be disk-written in the table, and the log sequence number of the transaction log to which the page mirror data corresponding to the data pages to be disk-written belongs.
[0113] Step S218: Store the basic page mirror data into the single storage file. After the single storage file is generated, the basic page mirror data corresponding to the data pages to be disk-written is stored in the single storage file.
[0114] Step S219: Store the data pages to be disk-written into the disk. After the basic page mirror data is stored in the single storage file, the data pages to be disk-written are stored in the disk.
[0115] In the solution of this embodiment, in the case where there are data pages of the same table, the data pages to be disk-written and the data pages of the same table are written to the disk together, reducing the number of times the data pages are written to the disk.
[0116] Such as Figure 11As shown, in one embodiment, the file name of the combined data storage file has at least the table location information of the table to which the data pages to be written to disk belong. Before the step of generating the combined data storage file on the disk, it includes: obtaining the block numbers and log sequence numbers corresponding to all page mirror data in the combined page mirror data, obtaining the maximum block number and the minimum block number among all the block numbers corresponding to the page mirror data, and obtaining the maximum log sequence number and the minimum log sequence number among all the log sequence numbers corresponding to the page mirror data.
[0117] The transaction log processing method generally includes:
[0118] Step S301, obtaining the block numbers and log sequence numbers corresponding to all page mirror data in the combined page mirror data. Specifically, as shown in Figure 9, that is, obtaining the BlockNumber and lsn of all page mirror data in the combined page mirror data.
[0119] Step S302, obtaining the maximum block number and the minimum block number among all the block numbers corresponding to the page mirror data. Specifically, that is, obtaining the maximum value and the minimum value among the block numbers corresponding to all page mirror data in the combined page mirror data. They can be respectively denoted as minBlockNumber and maxBlockNumber.
[0120] Step S303, obtaining the maximum log sequence number and the minimum log sequence number among all the log sequence numbers corresponding to the page mirror data. Specifically, that is, obtaining the maximum value and the minimum value of the log sequence numbers corresponding to the transaction logs to which all page mirror data in the combined page mirror data belong. They can be respectively denoted as minlsn and maxlsn.
[0121] Step S304, generating a combined data storage file whose file name records the table location information of the table to which the data pages to be written to disk belong, the maximum block number, the minimum block number, the maximum log sequence number, and the minimum log sequence number.
[0122] Specifically, the file name of the combined storage file can directly consist of the table location information of the table to which the data pages to be written to disk belong, the maximum block number, the minimum block number, the maximum log sequence number, and the minimum log sequence number. Exemplarily, the file name format of the combined storage file is:
[0123] RelFileNode_minBlockNumber_minlsn.maxBlockNumber_maxlsn;
[0124] RelFileNode is the table location information of the table to which the data pages to be written to disk belong, minBlockNumber is the minimum block number, minlsn is the minimum log sequence number, maxBlockNumber is the maximum block number, and maxlsn is the maximum log sequence number.
[0125] Further, by generating a combined data storage file with the table position information, maximum block number, minimum block number, maximum log sequence number, and minimum log sequence number of the data pages to be written to disk recorded in the file name, it is convenient to search for the combined data storage file with multiple page mirror data subsequently.
[0126] Refer to Figure 12 As shown, further, in this embodiment, after the step of obtaining the block numbers and log sequence numbers corresponding to all page mirror data in the combined page mirror data, it includes:
[0127] Step S401, sort all page mirror data in the combined page mirror data in ascending order of block numbers. Specifically, that is, before storing into the combined data storage file, first sort all page mirror data in the combined page mirror data from smallest to largest block number.
[0128] Step S402, if there is page mirror data with the same block number in the combined page mirror data, sort the page mirror data with the same block number in ascending order of log sequence numbers. Specifically, after sorting all page mirror data in the combined page mirror data from smallest to largest block number, if there is page mirror data with the same block number, sort the page mirror data with the same block number from smallest to largest log sequence number.
[0129] Specifically, although usually page mirror data is generated when the data of a data page is first modified after the log replay point in a database, some databases also have other modes of recording page mirror data, such as recording at regular intervals. In this case, there may be page mirror data with the same block number.
[0130] By sorting all page mirror data in the combined page mirror data in ascending order of block numbers and then sorting the page mirror data with the same block number in ascending order of log sequence numbers, the page mirror data stored in the combined data storage file is arranged in a certain order, which is convenient for subsequent search and extraction.
[0131] Further, in this embodiment, after the step of generating the combined data storage file on the disk, it includes:
[0132] Step S305, add a file header to the combined data storage file. The file header content includes the number of page mirror data in the combined page mirror data, block number array, log sequence number array, and starting address array. Among them, the block number array stores the block numbers of all page mirror data in the combined page mirror data, the log sequence number array stores the log sequence numbers of all page mirror data in the combined page mirror data, and the starting address array stores the starting addresses of all page mirror data in the combined page mirror data.
[0133] Refer to Figure 13As shown in the figure, N in the file header is the number of page mirror data in the combined page mirror data. BlockNumber[N] represents the block number array, lsn[N] represents the log sequence number array, and ptr[N] represents the starting address array. In addition, the block numbers of each page mirror data in the block number array are arranged in the order of the page mirror data in the combined page mirror data, and the log sequence numbers of each page mirror data in the log sequence number array are arranged in the order of the page mirror data in the combined page mirror data.
[0134] Furthermore, as Figure 13 shown, there are three page mirror data in the combined page mirror data. The number of page mirror data in the combined page mirror data is 3. There are three block numbers in the block number array, three log sequence numbers in the log sequence number array, and three starting addresses in the starting address array. Referring to Figure 13 shown, exemplarily, for example, the three starting addresses are ptr[0], ptr[1], and ptr[2]. Then ptr[0] is the starting address of the first page mirror data, ptr[1] is the starting address of the second page mirror data, and ptr[2] is the starting address of the third page mirror data.
[0135] In the solution of this embodiment, by adding a file header to the combined data storage file, the search for the entire combined page mirror data can be completed using the information in the file header, making the search more convenient. Moreover, when the data is found, the position of the page mirror data in the combined page mirror data can be quickly found directly according to the starting address information in the file header.
[0136] This embodiment also provides a machine-readable storage medium and a computer device. Figure 14 is a schematic diagram of a machine-readable storage medium 10 according to an embodiment of the present invention. Figure 15 is a schematic diagram of a computer device 20 according to an embodiment of the present invention.
[0137] The machine-readable storage medium 10 stores a machine-executable program 11 thereon. When the machine-executable program 11 is executed by a processor, the transaction log processing method of the database in any of the above embodiments is implemented.
[0138] The computer device 20 may include a memory 210, a processor 220, and a machine-executable program 11 stored on the memory 210 and running on the processor 220. When the processor 220 executes the machine-executable program 11, the transaction log processing method of the database in any of the above embodiments is implemented.
[0139] For the description of this embodiment, the machine-readable storage medium 10 can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the machine-readable storage medium 10 can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0140] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system.
[0141] The computer device 20 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. In some examples, the computer device 20 can be a cloud computing node. The computer device 20 can be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. The computer device 20 can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0142] The computer device 20 can include a processor 220 suitable for executing stored instructions and a memory 210 that provides temporary storage space for the operation of the instructions during operation. The processor 220 can be a single-core processor, a multi-core processor, a computing cluster, or any other number of other configurations. The memory 210 can include a random access memory (RAM), a read-only memory, a flash memory, or any other suitable storage system.
[0143] The processor 220 can be connected to an I / O interface (input / output interface) suitable for connecting the computer device 20 to one or more I / O devices (input / output devices) through a system interconnection (such as PCI, PCI-Express, etc.). The I / O devices can include, for example, a keyboard and a pointing device, where the pointing device can include a touchpad or a touch screen, etc. The I / O devices can be built-in components of the computer device 20, or can be devices externally connected to the computing device.
[0144] The processor 220 can also be linked to a display interface suitable for connecting the computer device 20 to a display device through a system interconnection. The display device can include a display screen as a built-in component of the computer device 20. The display device can also include a computer monitor, a television, a projector, etc. externally connected to the computer device 20. In addition, a network interface controller (NIC) can be suitable for connecting the computer device 20 to a network through a system interconnection. In some embodiments, the NIC can use any suitable interface or protocol (such as Internet Small Computer System Interface, etc.) to transmit data. The network can be a cellular network, a radio network, a wide area network (WAN), a local area network (LAN), or the Internet, etc. A remote device can be connected to the computing device through the network.
[0145] The solution of the present invention is applicable to relational databases, and particularly applicable to the KingbaseES database (abbreviated as KES database), enriching the functions of the database and improving the efficiency of the database.
[0146] At this point, those skilled in the art should recognize that although multiple exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications that conform to the principles of the present invention can still be directly determined or derived from the content disclosed in the present invention without departing from the spirit and scope of the present invention. Therefore, the scope of the present invention should be understood and determined to cover all these other variations or modifications.
Claims
1. A method for processing transaction logs of a database, wherein, The transaction log includes an assembled linked list and a page data linked list separately stored in a cache. The page data linked list is a linked list composed of page mirror data in the transaction log, and the assembled linked list is a linked list composed of other data in the transaction log except for the page mirror data. The page data linked list is stored in a preset global page data linked list. And the transaction log processing method includes: Obtaining a trigger event for flushing a data page in the cache to disk; Obtaining a data page to be flushed in the cache; Obtaining page mirror data corresponding to the data page to be flushed in the global page data linked list, denoted as basic page mirror data; Obtaining co-table data pages, where the co-table data pages are data pages that need to be flushed and belong to the same table as the data page to be flushed; Obtaining page mirror data corresponding to the co-table data pages in the global page data linked list; Assembling page mirror data corresponding to all the co-table data pages behind the basic page mirror data, denoted as combined page mirror data; Generating a combined data storage file on the disk; Storing the combined page mirror data into the combined data storage file; Storing the data page to be flushed into the disk.
2. The method for processing a transaction log of a database according to claim 1, wherein, The file name of the combined data storage file has at least table location information of the table to which the data page to be flushed belongs.
3. The method for processing a transaction log of a database according to claim 2, wherein, Before the step of generating the combined data storage file on the disk includes: Obtaining block numbers and log sequence numbers corresponding to all page mirror data in the combined page mirror data; Obtaining the maximum block number and the minimum block number among block numbers corresponding to all page mirror data; Obtaining the maximum log sequence number and the minimum log sequence number among log sequence numbers corresponding to all page mirror data; The step of generating the combined data storage file on the disk includes: Generating a combined data storage file with the table location information of the table to which the data page to be flushed belongs, the maximum block number, the minimum block number, the maximum log sequence number, and the minimum log sequence number recorded in the file name.
4. The method for processing a transaction log of a database according to claim 3, wherein, After the step of obtaining block numbers and log sequence numbers corresponding to all page mirror data in the combined page mirror data includes: Sorting all page mirror data in the combined page mirror data in ascending order of the block numbers; Checking whether there are page mirror data with the same block number in the combined page mirror data. If so, sorting the page mirror data with the same block number in ascending order of the log sequence numbers.
5. The method for processing a transaction log of a database according to claim 3, wherein, After the step of generating the combined data storage file on the disk includes: Adding a file header to the combined data storage file, where the file header content includes the number of page mirror data in the combined page mirror data, a block number array, a log sequence number array, and a head address array. Among them, the block number array stores block numbers of all page mirror data in the combined page mirror data, the log sequence number array stores log sequence numbers of all page mirror data in the combined page mirror data, and the head address array stores head addresses of all page mirror data in the combined page mirror data.
6. The method for processing a transaction log of a database according to claim 1, wherein, Before the step of obtaining co-table data pages includes: Checking whether there are co-table data pages. If so, executing the step of obtaining co-table data pages; Otherwise, generate a single storage file in the disk; Store the base page mirror data into the single storage file; Individually store the data pages to be disk-written into the disk.
7. The method for processing a transaction log of a database according to claim 1, wherein, Before the step of obtaining the page mirror data corresponding to the same-table data pages in the global page data linked list, it includes: Check whether there is corresponding page mirror data for all the same-table data pages in the global page data linked list. If so, execute the step of obtaining the page mirror data corresponding to the same-table data pages in the global page data linked list; Otherwise, generate a single storage file in the disk; Store the base page mirror data into the single storage file; Store the data pages to be disk-written into the disk.
8. The method for processing a transaction log of a database according to claim 1, wherein, The step of storing the data pages to be disk-written into the disk includes: Store the data pages to be disk-written and all the same-table data pages into the disk.
9. A machine-readable storage medium, on which a machine-executable program is stored, and when the machine-executable program is executed by a processor, it implements the transaction log processing method of the database according to any one of claims 1 to 8.
10. A computer device, including a memory, a processor, and a machine-executable program stored on the memory and running on the processor, and when the processor executes the machine-executable program, it implements the transaction log processing method of the database according to any one of claims 1 to 8.