Database processing method, storage medium and equipment

By adding file headers to the page mirror file and optimizing its filling timing, the problem of host and backup stream replication performance bottlenecks and inefficient search of page mirror file in synchronous stream replication mode is solved, and efficient database processing methods are realized, reducing disk IO overhead.

CN120277109APending Publication Date: 2025-07-08CETC JINCANG (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311865426.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the synchronous stream replication mode, stream replication between the host and the standby machine is easily a performance bottleneck, and the search efficiency of page mirror files in the prior art is inefficient, and frequent update of file header content leads to excessive extra disk IO overhead.

Method used

Add file headers to the page mirror file to store search data that identifies various location attributes in the storage data area, and build a file header data structure in the database memory, fill search data according to preset conditions, optimize the filling timing of file headers, and reduce disk IO overhead.

Benefits of technology

It improves the search efficiency of page mirror file content, reduces the disk IO overhead introduced by file header content filling, optimizes the file header filling timing, and improves the stability and performance of the database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277109A_ABST
    Figure CN120277109A_ABST
Patent Text Reader

Abstract

The invention relates to a database technology, in particular to a database processing method, a storage medium and equipment. A page mirror image file is pre-created in the database, the page mirror image file comprises a file header and a storage data area, and the file header is used for storing multiple pieces of search data identifying multiple position attributes of the storage data area. The processing method comprises the steps that after a page mirror image file is created, a file header data structure is constructed in a memory; according to various position attributes of the storage data area, filling each piece of search data in the file header data structure; and under the condition that the database meets a preset filling condition, filling the search data in the file header data structure into the to-be-filled file header. According to the database processing method, the content in the file header is filled only under the condition that the preset filling condition is met, so that the problem of excessive extra disk IO overhead caused by frequent updating of the content in the file header is avoided, and the disk IO overhead introduced by filling of the content of the file header is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to database technology, and particularly to a method for processing a database, a storage medium, and a device. Background Art

[0002] Currently, in the synchronous streaming replication mode, after the host sends XLOG logs to the standby, it is necessary to wait for at least one standby to store the XLOG logs locally or replay them successfully before the host can continue to execute. In the case where the streaming replication mode of the primary-standby database cluster is synchronous and the host generates a large number of XLOG logs due to heavy business pressure, the streaming replication between the host and the standby is likely to become a performance bottleneck of the database. Since the page mirror data in the XLOG logs stores the page data of the database pages and the data volume is large, separating the page mirror data in the XLOG logs of the database and continuously storing it in the page mirror file can greatly reduce the data volume of the XLOG logs.

[0003] The inventors recognized that when searching for a page mirror data from a page mirror file, it is only possible to traverse from the first page mirror data area of the page mirror file until the target page mirror data is found or until the last page mirror data area of this page mirror file (if the target page mirror data is not found in the current page mirror file), and such search efficiency is very low. Therefore, adding a file header to the page mirror file to store relevant search data can greatly improve the search efficiency of the page mirror file content.

[0004] However, due to the large time difference in the generation of each part of the content in the file header, it is very difficult to complete all fillings at one time. For example, the log sequence number corresponding to the first page mirror data area in the page mirror file can be determined when the first page mirror data area is written to disk, but the log sequence number corresponding to the last page mirror data area in the page mirror file needs to wait until the last page mirror data area is written to disk to be determined. If the content in the file header is updated frequently, additional disk I / O overhead will be introduced.

[0005] The above information disclosed in this background art is only used to increase the understanding of the background art of the present application. Therefore, it may include prior art that is not known to those of ordinary skill in the art. Summary of the Invention

[0006] In view of the above problems, a method for processing a database, a storage medium, and a device that overcome the above problems or at least partially solve the above problems are proposed.

[0007] An object of the present invention is to provide a method for processing a database that can reduce disk I / O overhead.

[0008] A further object of the present invention is to optimize the filling timing of the file header of the page mirror file, further reducing the disk I / O overhead.

[0009] In particular, the present invention provides a method for processing a database. The database pre-creates a page mirror file, which includes a file header and a storage data area. The storage data area includes a valid data area storing a plurality of page mirror data areas and an invalid data area for writing page mirror data to be written. The page mirror data area is used to store page mirror data separated from the transaction log of the database, and the file header is used to store a plurality of search data identifying various position attributes of the storage data area. And the processing method includes:

[0010] After the page mirror file is created, a file header data structure corresponding to the file header to be filled in the page mirror file is constructed in the memory of the database;

[0011] According to various position attributes of the storage data area, each search data is filled in the file header data structure;

[0012] When the database reaches a preset filling condition, the search data in the file header data structure is filled into the file header to be filled.

[0013] Optionally, the preset filling condition includes that the page mirror data area to be written to disk in the page mirror file includes one or more of the first page mirror data area in the storage data area, switching the page mirror file, and the database being idle; and

[0014] Each preset filling condition is correspondingly set with at least one search data.

[0015] Optionally, the plurality of search data includes a maximum sequence number for identifying the log sequence number of the last page mirror data area record in the storage data area and a length value for identifying the sum of the length of the file header and the length of the valid data area in the storage data area; and

[0016] When the database reaches a preset filling condition, filling the search data in the file header data structure into the file header to be filled includes:

[0017] When switching the page mirror file, the maximum sequence number and the length value in the file header data structure are filled into the file header to be filled.

[0018] Optionally, filling each search data in the file header data structure according to various position attributes of the page mirror data areas in the storage data area includes:

[0019] According to the lengths of the respective page mirror data areas in the storage data area, the length of the valid data area is cumulatively calculated;

[0020] Update the length value in the file header data structure according to the length of the file header and the cumulative length of the valid data area obtained.

[0021] Optionally, multiple search data includes the minimum sequence number for identifying the log sequence number of the first page mirror data area record in the storage data area; and

[0022] When the database reaches the preset filling condition, filling the search data in the file header data structure into the file header to be filled includes:

[0023] When the page mirror data area to be written to disk in the page mirror file includes the first page mirror data area in the storage data area, connect the file header data structure and the first page mirror data area to form a continuous data area;

[0024] Write the continuous data area including the file header data structure and the first page mirror data area to the page mirror file to fill the minimum sequence number in the file header data structure into the file header to be filled.

[0025] Optionally, the storage data area is pre-divided into multiple search areas, and each search area contains multiple page mirror data areas;

[0026] Multiple search data includes a sequence number array and a starting address array. The sequence number array includes search sequence numbers for identifying the log sequence numbers of the first page mirror data area records in each search area in the storage data area, and the starting address array includes search starting addresses for identifying the starting positions of each search area in the page mirror file in the storage data area; and

[0027] When the database reaches the preset filling condition, filling the search data in the file header data structure into the file header to be filled includes:

[0028] When the database is idle, fill the sequence number array and the starting address array in the file header data structure into the file header to be filled.

[0029] Optionally, multiple search data includes the maximum length of the sequence number array, and the maximum length of the sequence number array is used to identify the number of search areas in the storage data area; and

[0030] When the database reaches the preset filling condition, filling the search data in the file header data structure into the file header to be filled includes:

[0031] When the page mirror data area to be written to disk in the page mirror file includes the first page mirror data area in the storage data area, connect the file header data structure and the first page mirror data area to form a continuous data area;

[0032] Write the continuous data area including the file header data structure and the first page mirror data area to the page mirror file to fill the maximum length of the serial number array in the file header data structure into the file header to be filled.

[0033] Optionally, each search area has the same size; and

[0034] The maximum length of the serial number array is calculated based on the total length of the page mirror file and the length of each search area.

[0035] According to another aspect of the present invention, there is also provided a machine-readable storage medium on which a machine-executable program is stored. When the machine-executable program is executed by a processor, the processing method of any one of the above databases is implemented.

[0036] According to still another aspect of the present invention, there is also provided a computer device, including a memory, a processor, and a machine-executable program stored on the memory and running on the processor. When the processor executes the machine-executable program, the processing method of any one of the above databases is implemented.

[0037] In the processing method of the database of the present invention, a file header is added to the page mirror file. The file header is used to store multiple search data for identifying various position attributes of the stored data area, so as to facilitate searching the content of the page mirror file according to the search data, improving the search efficiency of the content of the page mirror file. In addition, in the processing method of the database of the present invention, after the page mirror file is created, first construct a file header data structure corresponding to the file header to be filled in the page mirror file in the memory of the database, and then fill each search data in the file header data structure according to various position attributes of the stored data area. When the database reaches the preset filling condition, fill the search data in the file header data structure into the file header to be filled, realizing the orderly filling of the content in the file header of the page mirror file, avoiding the problem of excessive additional disk I / O overhead caused by frequent content updates in the file header, and reducing the disk I / O overhead introduced by the content filling of the file header.

[0038] Further, in the processing method of the database of the present invention, multiple preset filling conditions are set, including that the page mirror data area to be written to the disk in the page mirror file includes the first page mirror data area in the stored data area, switching the page mirror file, and the database being idle. Each preset filling condition is set corresponding to at least one search data to realize the refinement of the filling condition, optimizing the filling timing of the file header of the page mirror file, thereby further reducing the disk I / O overhead.

[0039] Those skilled in the art will become more clear about the above and other objects, advantages and features of the present invention according to the following detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings. Brief Description of the Drawings

[0040] Some specific embodiments of the present invention will be described in detail hereinafter with reference to the accompanying drawings in an illustrative rather than restrictive manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0041] Figure 1 is a schematic flowchart of a method for processing a database according to an embodiment of the present invention;

[0042] Figure 2 is a schematic structural diagram of a page mirror file in a method for processing a database according to an embodiment of the present invention;

[0043] Figure 3 is a schematic structural diagram of a transaction log in a method for processing a database according to an embodiment of the present invention;

[0044] Figure 4 is a schematic structural diagram of a page mirror file in a method for processing a database according to another embodiment of the present invention;

[0045] Figure 5 is a schematic structural diagram of a page mirror file in a method for processing a database according to still another embodiment of the present invention;

[0046] Figure 6 is a schematic flowchart of a method for processing a database according to an embodiment of the present invention;

[0047] Figure 7 is a schematic diagram of a machine-readable storage medium according to an embodiment of the present invention; and

[0048] Figure 8 is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Embodiments

[0049] The exemplary embodiments of the present invention will be described in more detail hereinafter with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0050] To solve the above technical problems, an embodiment of the present invention provides a method for processing a database. Figure 1 is a schematic flowchart of a method for processing a database according to an embodiment of the present invention. Figure 2 is a schematic structural diagram of a transaction log in a method for processing a database according to an embodiment of the present invention. As Figure 1As shown, the processing method of the database generally may include:

[0051] Step S102, after the page mirror file is created, construct a file header data structure corresponding to the file header to be filled in the memory of the database. When the page mirror file is created, initialize the file header but do not fill in each search data in the file header.

[0052] Step S104, fill in each search data in the file header data structure according to various position attributes of the storage data area. In this step, the search data in the file header data structure can be filled in one by one, or the search data in the file header data structure can be filled in batches.

[0053] Step S106, when the database reaches a preset filling condition, fill the search data in the file header data structure into the file header to be filled. After part of the search data in the file header data structure is filled, when the database reaches a preset filling condition, the data in the file header data structure can be written to disk into the page mirror file to fill the file header to be filled in the page mirror file.

[0054] In this embodiment, the database pre-creates a page mirror file. As Figure 2 shown, the page mirror file (denoted as the pagedata file) includes a file header and a storage data area. The storage data area includes a valid data area storing multiple page mirror data areas (denoted as pagedata data areas) and an invalid data area for writing page mirror data. The page mirror data area is used to store page mirror data (denoted as page data) separated from the transaction log of the database. The file header is used to store multiple search data identifying various position attributes of the storage data area.

[0055] For the processing method of the database of the present invention, a file header is added to the page mirror file, and the file header is used to store multiple search data identifying various position attributes of the storage data area, so as to facilitate searching the content of the page mirror file according to the search data, and improve the search efficiency of the content of the page mirror file.

[0056] In addition, the processing method of the database of the present invention, after the page mirror file is created, first constructs a file header data structure corresponding to the file header to be filled in the memory of the database, then fills in each search data in the file header data structure according to various position attributes of the storage data area, and when the database reaches a preset filling condition, fills the search data in the file header data structure into the file header to be filled, realizing an orderly filling of the content in the file header of the page mirror file, avoiding the problem of excessive additional disk I / O overhead caused by frequent content updates in the file header, and reducing the disk I / O overhead introduced by the content filling of the file header.

[0057] Specifically, the above step S102 can be specifically executed as follows: Before the first page mirror data area in the page mirror file is separately written to disk or the first page mirror data area and several subsequent page mirror data areas are written to disk together, a file header data structure in memory is constructed.

[0058] In a database, a transaction log refers to an XLOG log, and the XLOG log details the operation process of the service process on the database. During the operation of the database, multiple XLOG logs are generated. The structure of an XLOG log is described below.

[0059] Figure 3 It is a schematic diagram of the structure of the XLOG log in the database processing method according to an embodiment of the present invention. As Figure 3 shown, an XLOG log can include an XLOG log header (denoted as head data) and an XLOG log data area.

[0060] The XLOG log data area can include multiple block data areas (denoted as block) and main data (denoted as main data). Each block can contain page mirror data (denoted as page data) and tuple data (denoted as tupledata). The XLOG log header includes an XLogRecord structure, the header data of each block, and the header data of the main data.

[0061] It should be noted that different types of XLOG logs have different compositions. Each XLOG log contains head data, but not necessarily page data, tupledata, and main data. Some types of XLOG logs have only head data and no XLOG log data area; some other types of XLOG logs have only head data and main data; and some other types of XLOG logs have only head data, page data, and tupledata, and each block may contain both page data and tupledata, or may only contain tupledata. That is to say, an XLOG log may contain page data or may not contain page data.

[0062] During the operating system crash of a database, some operating system pages (e.g., with a page size of 4KB) may not have been written to disk in time, which may cause a database page (e.g., composed of two operating system pages, denoted as a data page) to contain a mixture of old and new data. During the recovery period after the operating system crash of the database, since the information stored in the XLOG log is not complete enough, it is impossible to fully recover the database pages that contain a mixture of old and new data in the database. In the prior art, in order to recover such database pages, after the redo point of the checkpoint (checkpoint log) in the database, when the database page (with a default size of 8KB) is first modified (updated), the entire database page is written to the XLOG log. At this time, the XLOG log contains page mirror data to save the complete database page mirror. It should be noted that the redo point is a special point in the XLOG log. Before this point, all the data in the database is the same as the information reflected by the disk-written XLOG log. During the creation of a checkpoint, a special XLOG log (denoted as a checkpoint log) is created, and a redo point is recorded in this checkpoint log. Only when the XLOG log and user data before this redo point are both written to disk, can the checkpoint be successfully created. When the database crashes, we can find a redo point from the most recent checkpoint log, and then replay the XLOG log starting from this redo point to recover the database. When a broken page is encountered during the database recovery process, the broken page is replaced according to the page mirror data in the XLOG log, thus ensuring the security and stability of the database data. Therefore, the data of the page mirror data in the XLOG log is the data of the database page. That is to say, the page mirror data has its corresponding XLOG log and data page.

[0063] Since the page mirror data in the XLOG log is the data of the database page, the data volume of the XLOG log containing the page mirror data is relatively large. In some embodiments, in order to reduce the data volume of the XLOG log, before the above step S102, the processing method of the database of the present invention may further include the following steps: separating the page mirror data in the XLOG log generated by the database from the XLOG log to obtain the XLOG log after the page mirror data is separated; storing the page mirror data separated from the XLOG log into a page mirror file.

[0064] In this embodiment, the page mirror data in all the XLOG logs generated in the database is separated, which greatly reduces the data volume of the XLOG logs. After the primary and standby database clusters are built, the host only transfers the XLOG logs after page mirror data separation and the XLOG logs without page mirror data to the standby machine, thus greatly reducing the pressure of XLOG log transmission in the streaming replication between the host and the standby machine on the premise of ensuring the security and stability of the database data.

[0065] In some embodiments, as Figure 2 shown, multiple page mirror data areas in the page mirror file are set in one-to-one correspondence with the page mirror data stored in the page mirror file. Each page mirror data area includes page mirror data and related data of the page mirror data.

[0066] The related data of the page mirror data includes the starting position (denoted as lsn) of the XLOG log where the page mirror data is located, the page mirror data header (denoted as page data header), the position information of the data page corresponding to the page mirror data, and the total length of the page mirror data area (denoted as len).

[0067] The lsn corresponding to each page mirror data increases sequentially. The lsn can be used to determine the correspondence between the page mirror data and the XLOG log. The len is used to calculate the starting address of the next page mirror data area in the page mirror file. When traversing the page mirror file, the next page mirror data area can be quickly found according to the len of each page mirror data area. On this basis, one page mirror data corresponds to one XLOG log and one data page.

[0068] The position information of the data page corresponding to the page mirror data includes RelFileNode and BlockNumber. The data page refers to the data disk page. RelFileNode is the position information of the table where the data disk page corresponding to the page mirror data is located, and BlockNumber is the block number of the data disk page corresponding to the page mirror data in its table. RelFileNode and BlockNumber together determine the position of the data disk page on the disk. According to these two data of RelFileNode and BlockNumber, a unique data disk page can be determined on the disk. When searching on the disk, it can be determined whether a version of the target page mirror data is found according to these two values. The page mirror data header is the header information corresponding to the page mirror data.

[0069] In some embodiments, the step of separating the page mirror data and the XLOG log in the XLOG log generated by the database to obtain the XLOG log after separating the page mirror data may include the following steps: In the stage of assembling the XLOG log, separate the page mirror data in the XLOG log to be assembled from the XLOG log to be assembled, to obtain the XLOG log after separating the page mirror data. Specifically, the step of separating the page mirror data in the XLOG log to be assembled from the XLOG log to be assembled may include the following steps: Create a page mirror linked list node for each page mirror data in the XLOG log to be assembled, and sequentially connect the page mirror linked list nodes of all the page mirror data in the XLOG log to be assembled to form a page mirror linked list; Connect the page mirror linked list to the globally pre-created page data linked list; And sequentially connect the remaining nodes in the first assembly linked list except the assembly linked list nodes for storing page mirror data to form a second assembly linked list for storing the XLOG log after separating the page mirror data.

[0070] In this embodiment, the globally pre-created page data linked list includes a plurality of connected page mirror linked lists for storing all the page mirror data separated from the XLOG log in the database. Each page mirror linked list contains at least one page mirror linked list node. Each page mirror linked list node corresponds to a page mirror data, and each page mirror linked list node includes a page mirror data area and a pointer identifier (denoted as next). Next is used to point to the next page mirror linked list node to connect the page mirror linked lists corresponding to two adjacent XLOG logs.

[0071] The processing method of the database in this embodiment, by creating a page mirror linked list node for each page mirror data in the XLOG log to be assembled, sequentially connecting the page mirror linked list nodes of all the page mirror data in the XLOG log to be assembled to form a page mirror linked list, connecting the page mirror linked list to the globally pre-created page data linked list, and sequentially connecting the remaining nodes in the first assembly linked list except the assembly linked list nodes for storing page mirror data to form a second assembly linked list, completes the operation of separating the page mirror data of the XLOG log to be assembled, and further improves the accuracy of the operation of separating the page mirror data from the XLOG log.

[0072] In a specific embodiment, after the step of sequentially connecting the page mirror linked list nodes of all the page mirror data in the XLOG log to be assembled to form a page mirror linked list, the processing method of the database of the present invention may further include the following steps: Connect the page mirror linked list after the remaining nodes in the first assembly linked list except the assembly linked list nodes for storing page mirror data.

[0073] That is to say, after sequentially connecting the page mirror linked list nodes of all page mirror data in the XLOG log to be assembled to form a page mirror linked list, first connect the page mirror linked list to the end of the first assembly linked list, and then remove the page mirror linked list behind the first assembly linked list and connect it to the global page data linked list.

[0074] The processing method of the database in this embodiment can orderly adjust the nodes of the page mirror data, enabling the data to be effectively managed and maintained globally, facilitating data sharing and consistency maintenance, and improving the stability of the database.

[0075] Figure 4 It is a schematic diagram of the structure of a page mirror file in the processing method of a database according to another embodiment of the present invention. Figure 5 It is a schematic diagram of the structure of a page mirror file in the processing method of a database according to yet another embodiment of the present invention. As Figure 4 and Figure 5 shown, the storage data area is pre-divided into multiple search areas, and each search area contains multiple page mirror data areas.

[0076] In this embodiment, as Figure 2 shown, the multiple search data may include the minimum sequence number (denoted as minLsn), the maximum sequence number (denoted as maxLsn), the length value (denoted as totalLen), the sequence number array (denoted as lsn[N]), the starting address array (denoted as ptr[N]), and the maximum length of lsn[N] (denoted as N).

[0077] minLsn is used to identify the log sequence number recorded in the first page mirror data area in the storage data area.

[0078] maxLsn is used to identify the log sequence number recorded in the last page mirror data area in the storage data area.

[0079] It should be noted that the log sequence number recorded in each page mirror data area is the starting position of the XLOG log where the page mirror data in this page mirror data area is located, that is, the lsn recorded in the page mirror data area as Figure 3 shown.

[0080] As Figure 5 described, the lsn of the XLOG log corresponding to the target page mirror data can be denoted as target_lsn. When it is found that target_lsn is between minLsn and maxLsn of the current page mirror file, the target page mirror data can be attempted to be found from the current page mirror file.

[0081] The lsn[N] includes search serial numbers (denoted as lsn[i], where i = 0, 1, …, N - 1) for identifying the log sequence numbers of the first page mirror data area records in each search area of the storage data area. Thus, each search area in the storage data area can be directly located according to lsn[N].

[0082] When searching for target page mirror data in the page mirror file, the search area where the target page mirror data may exist can be directly located according to target_lsn and lsn[N], which greatly improves the search efficiency.

[0083] The ptr[N] includes search start addresses (denoted as ptr[i], where i = 0, 1, …, N - 1) for identifying the start positions of each search area in the page mirror file in the storage data area. That is to say, ptr[i] is the search start address for identifying the start position of the page mirror data area (denoted as pd_i) with the search serial number lsn[i] in the page mirror file. Therefore, as Figure 4 shown, lsn[i] and ptr[i] are in one-to-one correspondence.

[0084] When searching for target page mirror data in the page mirror file, the search area where the target page mirror data may exist can be directly located according to target_lsn and the correspondence between lsn[N] and ptr[N], which greatly improves the search efficiency.

[0085] In a specific embodiment, the sizes of each search area are the same. It should be noted that there may be up to tens of thousands of page mirror data areas in a page mirror file. To accelerate the process of searching for target page mirror data in the page mirror file, as Figure 4 shown, in the storage data area of the page mirror file in the processing method of the database in this embodiment, the lsn value of the page mirror data area record at this position is recorded every fixed distance (denoted as fixed_distance, such as 1MB), and all lsn values are stored in the lsn[N] array. All the page mirror data areas between two adjacent lsn values form a search area. Thus, the sizes of each search area in this embodiment are the same.

[0086] N can also be the maximum lengths of lsn[N] and ptr[N], which is used to identify the number of search areas in the storage data area.

[0087] In a specific embodiment, when the sizes of each search area are the same, N can be calculated based on the total length of the page mirror file and the length of each search area. Specifically, when the sizes of the page mirror file and the interval are fixed, N is a fixed value. Assuming the size of the page mirror file is 64MB and the interval is 1MB, then N = 64MB / 1MB = 64.

[0088] totalLen is used to identify the sum of the length of the file header and the length of the valid data area in the storage data area, that is, the total length of each page mirror data area in the page mirror file plus the length of the file header. totalLen can be calculated cumulatively according to the lengths of the page mirror data areas in the storage data area.

[0089] Specifically, before switching the page mirror file, totalLen can be calculated cumulatively. For example, every time a page mirror data area is written to disk, the current totalLen is added to the length of the page mirror data area written this time to obtain a new totalLen, and the new totalLen is filled into the totalLen field in the file header to update totalLen. Using the above method, the totalLen of the page mirror file can be quickly obtained when switching the page mirror file.

[0090] It should be noted that the content length of the file header is not fixed, but the length of the file header can be calculated based on the content in the file header. Specifically, the length of the file header can be the sum of the sizes of minLsn, maxLsn, N, totalLen, lsn[N], and ptr[N].

[0091] In a specific embodiment, the calculation method of the length of the file header is as follows:

[0092] (1) Determine the variables with fixed lengths: minLsn, maxLsn, N, and totalLen;

[0093] (2) Determine the variables with non-fixed lengths: lsn[N] and ptr[N]. Denote the length of each lsn[] array member as sizeof(lsn[0]), and the length of each ptr[] array member as sizeof(ptr[0]);

[0094] 1) Calculate the length of lsn[N] = sizeof(lsn[0]) * N.

[0095] 2) Calculate the length of ptr[N] = sizeof(ptr[0]) * N.

[0096] (3) Calculate the length of the file header = sizeof(minLsn) + sizeof(maxLsn) + sizeof(N) + sizeof(totalLen) + sizeof(lsn[0]) * N + sizeof(ptr[0]) * N.

[0097] A content filling scheme for the file header includes:

[0098] 1. Fill N: If the total length of the page mirror file (denoted as page_data_file_size) and fixed_distance are determined, then determine N = page_data_file_size / fixed_distance; since page_data_file_size and fixed_distance can be determined during database initialization, N is filled in when each page mirror file is created;

[0099] 2. Fill minLsn: After the first page mirror data area is written to disk, fill the corresponding lsn to minLsn;

[0100] 3. Fill totalLen: After each page mirror data area is written to disk, add the old totalLen to the length of the page mirror data area written this time to get the new totalLen, and fill the new totalLen to the totalLen field in the file header;

[0101] 4. Fill lsn[N] and ptr[N]: After each page mirror data area is written to disk, check whether the data range written this time covers a specific position (the length of the file header + fixed_distance * i); if it covers one or more specific positions, then fill the lsn value in the page mirror data area where the specific position is located into the lsn[N] array; for example: fill the lsn value in the page mirror data area at "the length of the file header + fixed_distance * i" into lsn[i], and at the same time fill the starting address of this page mirror data area into ptr[i];

[0102] 5. Fill maxLsn: The last page mirror data area can only be determined when switching page mirror files, so this value is filled when switching page mirror files.

[0103] In the above content filling scheme for the file header, when filling totalLen, lsn[N] and ptr[N], multiple filling operations are performed. Frequent filling operations will cause a large number of random IOs to be introduced into the page mirror file that was originally written sequentially due to the file header, which will affect the write performance of the XLOG log.

[0104] In order to reduce the disk I / O overhead introduced by the content filling of the file header, an optimized content filling scheme is proposed in the embodiments of the present invention.

[0105] In an embodiment of the present invention, the optimized content filling scheme includes: initializing the file header when creating a page mirror file without filling the N field; performing step S102: after creating the page mirror file, first constructing a file header data structure corresponding to the file header to be filled in the memory of the database; performing step S104: filling each search data in the file header data structure according to various position attributes of the storage data area; performing step S106, when the database reaches a preset filling condition, filling the search data in the file header data structure into the file header to be filled.

[0106] In some embodiments, the above step S104 may include the following steps: calculating the length of the valid data area by cumulative calculation according to the lengths of each page mirror data area of the storage data area; updating the totalLen in the file header data structure according to the length of the file header and the length of the valid data area obtained by cumulative calculation, so as to complete the filling of totalLen in the file header data structure.

[0107] Furthermore, the above step S104 may further include: after constructing the file header data structure in the memory, filling N and minLsn; before switching the page mirror file, accumulating the totalLen value in the memory (updating the totalLen value each time a page mirror data area is written to disk), and gradually filling lsn[N] and ptr[N]; when switching the page mirror file, filling maxLsn and totalLen into the file header data structure.

[0108] In some embodiments, multiple preset filling conditions are set in the database. The preset filling conditions may include one or more of the first page mirror data area in the storage data area, switching the page mirror file, and the database being idle among the page mirror data areas to be written to disk in the page mirror file. The situation that all the page mirror data areas to be written to disk in the page mirror file include the first page mirror data area in the storage data area further includes: the first page mirror data area in the page mirror file is written to disk alone, and the first page mirror data area and several subsequent page mirror data areas are written to disk together.

[0109] It should be noted that each preset filling condition is set according to the characteristics of each search data. And each preset filling condition can be set corresponding to at least one search data.

[0110] The processing method of the database in this embodiment sets multiple preset filling conditions, including the first page mirror data area in the storage data area, switching the page mirror file, and the database being idle among the page mirror data areas to be written to disk in the page mirror file. Each preset filling condition is set corresponding to at least one search data to achieve the refinement of the filling condition, optimize the filling timing of the file header of the page mirror file, and thus further reduce the disk I / O overhead.

[0111] In some embodiments, since N can be determined after the page mirror file is created, and minLsn can be determined when the first page mirror data area in the page mirror file is flushed to disk, N and minLsn can be flushed to disk together with the first page mirror data area in the page mirror file, so as to combine the disk I / O for filling the N field and minLsn of the file header and the disk I / O for filling the storage data area of the page mirror file.

[0112] On this basis, the preset filling condition that the page mirror data area to be flushed in the page mirror file includes the first page mirror data area in the storage data area can correspond to the two search data of N and minLsn.

[0113] Further, the above step S106 may include the following steps: when the page mirror data area to be flushed in the page mirror file includes the first page mirror data area in the storage data area, connect the file header data structure with the first page mirror data area to form a continuous data area; flush the continuous data area including the file header data structure and the first page mirror data area to the page mirror file, so as to fill the minLsn and N in the file header data structure into the file header to be filled.

[0114] In some embodiments, since maxLsn and totalLen can be determined only when the last page mirror data area in the page mirror file is flushed to disk, they can be filled when the page mirror file is switched.

[0115] On this basis, the preset filling condition of switching the page mirror file can correspond to the two search data of maxLsn and totalLen.

[0116] Further, the above step S106 may include the following steps: when the page mirror file is switched, fill the maxLsn and totalLen in the file header data structure into the file header to be filled.

[0117] In some embodiments, since lsn[N] and ptr[N] are search data generated for facilitating the search of the page mirror file, there is generally no urgent filling requirement. Therefore, lsn[N] and ptr[N] can be filled into the file header of the page mirror file when the database is idle.

[0118] On this basis, the preset filling condition that the page mirror data area to be flushed in the page mirror file includes the first page mirror data area in the storage data area can correspond to the two search data of N and minLsn.

[0119] Further, the above step S106 may include the following steps: when the database is idle, fill the lsn[N] and ptr[N] in the file header data structure into the file header to be filled.

[0120] In a specific embodiment, the optimized content filling scheme may specifically include:

[0121] 1. Initialize the file header when creating the page mirror file, but do not fill the N field in the file header;

[0122] 2. Fill N and minLsn: Before the first page mirror data area in the page mirror file is flushed to disk (it is also possible that the first page mirror data area and several subsequent page mirror data areas are flushed together), construct the file header data structure in memory, and fill the N field and minLsn in the file header data structure; splice the file header data structure in front of the first page mirror data area; flush the spliced continuous data together to the new page mirror file to fill N and minLsn into the file header of the page mirror file, which is equivalent to merging the disk I / O for filling the N field of the file header and the disk I / O for filling the storage data area of the page mirror file;

[0123] 3. Fill maxLsn and totalLen: Fill when switching page mirror files; in a preferred solution, before switching page mirror files, accumulate the totalLen value in memory (for example, update the totalLen value every time a page mirror data area is flushed to disk), and when switching page mirror files, the totalLen value can be quickly obtained and filled into the file header of the page mirror file;

[0124] 4. Fill lsn[N] and ptr[N]: Do not fill lsn[N] and ptr[N] during the writing process of the page mirror file; when the database is idle, fill lsn[N] and ptr[N] into the file header of the page mirror file.

[0125] Thus, the content in the file header of the page mirror file is filled securely and efficiently, greatly reducing the disk write I / O of the page mirror file and significantly improving the filling efficiency of the file header of the page mirror file.

[0126] In some embodiments, when searching for target page mirror data in the page mirror file, the following steps may be specifically executed:

[0127] 1. When it is found that the target_lsn of the target page mirror data is between the minLsn and maxLsn of the current page mirror file, then try to find the target page mirror data from this page mirror file;

[0128] 2. Traverse the lsn[N] array from front to back:

[0129] (1) When i < N - 1 and lsn[i] <= target_lsn < lsn[i + 1], stop traversing. The search range is from ptr[i] to ptr[i + 1]. The starting address of the target search area in the page mirror file is denoted as search_begin, and set search_begin = ptr[i]; the ending address of the target search area in the page mirror file is denoted as search_end, and set search_end = ptr[i + 1].

[0130] (2) When i = N - 1 and lsn[i] <= target_lsn < 64MB, stop traversing. The search range is from ptr[i] to totalLen. Denote search_begin = ptr[i]; search_end = totalLen.

[0131] 3. Traverse the page mirror data area one by one starting from search_begin of the page mirror file. Denote the current page mirror data area as pd_cur:

[0132] (1) Determine whether the lsn recorded in pd_cur (denoted as lsn_cur) is equal to target_lsn;

[0133] 1) If yes, jump to (2);

[0134] 2) If no, jump to (3);

[0135] (2) Determine whether the page mirror data position information recorded in pd_cur is consistent with the position information of the target page mirror data; the page mirror data position information recorded in pd_cur is the above-mentioned storage position, including the RelFileNode and BlockNumber corresponding to the page mirror data in pd_cur; the position information of the target page mirror data is the target position, including the RelFileNode and BlockNumber corresponding to the target page mirror data;

[0136] 1) If yes, confirm that the target page mirror data is found and stop traversing;

[0137] 2) If no, jump to (3);

[0138] (3) Check whether the ending position of the current page mirror data area is less than search_end;

[0139] 1) If yes, read the next page mirror data area (denoted as pd_next), set pd_cur = pd_next, and jump to (1);

[0140] 2) If no, stop traversing and confirm that the target page mirror data is not found.

[0141] It should be noted that a single XLOG log may record several page mirror data, and the LSNs corresponding to these page mirror data are the same. Therefore, when the log sequence number is equal to the target sequence number, further determination needs to be made based on the location information of the data pages corresponding to each page mirror data.

[0142] Using the above method, a file header is added to the page mirror file. The file header is used to store the search data for multiple page mirror data areas in the storage data area of the page mirror file, so as to facilitate the search for the content of the page mirror file according to the search data. When searching for the target page mirror data in the page mirror file, first obtain the search data and the location attribute information of the target page mirror data, and then search for the target page mirror data in the page mirror file according to the search data and the location attribute information of the target page mirror data, improving the search efficiency of the content of the page mirror file.

[0143] Figure 6 It is a schematic flowchart of the processing method of a database according to an embodiment of the present invention. The following will be combined with Figure 6 Specifically illustrate the flowchart steps of the processing method of the database in this embodiment.

[0144] Step S602, initialize the file header when the page data file is created.

[0145] Step S604, when the page mirror data area to be written to disk in the page mirror file includes the first page mirror data area in the storage data area, construct the file header data structure in memory and fill in the N field and minLsn in the file header data structure.

[0146] Step S606, splice the file header data structure in front of the first page data data area to form continuous data.

[0147] Step S608, write the spliced continuous data to the page data file together to fill the N and minLsn into the file header of the page data file.

[0148] Step S610, when switching the page data file, fill the maxLsn and totalLen into the file header of the page data file.

[0149] Step S612, when the database is idle, fill the lsn[N] and ptr[N] into the file header to be filled. This process ends.

[0150] Using the above method, after the page mirror file is created, first construct a file header data structure corresponding to the file header to be filled in the page mirror file in the memory of the database, and then fill in each search data in the file header data structure according to various position attributes of the storage data area. When the database reaches the preset filling condition, fill the search data in the file header data structure into the file header to be filled, realizing the orderly filling of the content in the file header of the page mirror file, avoiding the problem of excessive additional disk I / O overhead caused by frequent updates of the content in the file header, and reducing the disk I / O overhead introduced by the filling of the content in the file header.

[0151] This embodiment also provides a machine-readable storage medium and a computer device. Figure 7 It is a schematic diagram of a machine-readable storage medium 80 according to an embodiment of the present invention. Figure 8 It is a schematic diagram of a computer device 90 according to an embodiment of the present invention.

[0152] The machine-readable storage medium 80 stores a machine-executable program 81 thereon. When the machine-executable program 81 is executed by a processor, it implements the processing method of the database in any of the above embodiments.

[0153] The computer device 90 may include a memory 920, a processor 910, and a machine-executable program 81 stored on the memory 920 and running on the processor 910. When the processor 910 executes the machine-executable program 81, it implements the processing method of the database in any of the above embodiments.

[0154] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a fixed sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any machine-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices.

[0155] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For the description of this embodiment, the machine-readable storage medium 80 can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in combination with these instruction execution systems, apparatuses, or devices. The computer device 90 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smart phone. The computer device 90 can include a processor 910 adapted to execute the stored instructions and a memory 920 that provides temporary storage space for the operation of the instructions during operation. The processor 910 can be a single-core processor, a multi-core processor, a computing cluster, or any other number of other configurations. The memory 920 can include random access memory (RAM), read-only memory, flash memory, or any other suitable storage system.

[0156] It should be noted that in some alternative embodiments, the solution of the present invention is applicable to relational databases, and in particular to the KingbaseES database (abbreviated as KES database), which enriches the functions of the database and improves the efficiency of the database. In some other alternative embodiments, the processing method of the database of the present invention can also be applicable to other relational databases.

[0157] In addition, the flowcharts provided in this embodiment are not intended to indicate that the operations of the method will be executed in any specific order, or that all operations of the method are included in every case. In addition, the method can include additional operations. Within the scope of the technical concept provided by the method of this embodiment, additional changes can be made to the above method.

[0158] At this point, those skilled in the art should recognize that although numerous exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications consistent with the principles of the present invention can still be directly determined or derived from the content disclosed in the present invention without departing from the spirit and scope of the present invention. Therefore, the scope of the present invention should be understood and determined to cover all these other variations or modifications.

Claims

1. A processing method for a database, a page mirror file is pre-created in the database, the page mirror file includes a file header and a storage data area, the storage data area includes a valid data area storing a plurality of page mirror data areas and an invalid data area for writing page mirror data, the page mirror data area is used to store the page mirror data separated from the transaction log of the database, and the file header is used to store a plurality of search data identifying various position attributes of the storage data area; And the processing method includes: After the page mirror file is created, a file header data structure corresponding to the file header to be filled in the page mirror file is constructed in the memory of the database; According to various position attributes of the storage data area, each search data is filled in the file header data structure; When the database reaches a preset filling condition, the search data in the file header data structure is filled into the file header to be filled.

2. The processing method of the database according to claim 1, wherein, The preset filling condition includes that the page mirror data area to be written to disk in the page mirror file includes the first page mirror data area in the storage data area, switching the page mirror file, one or more of the database being idle; and Each preset filling condition is correspondingly set with at least one search data.

3. The processing method of the database according to claim 2, wherein, The multiple search data includes a maximum sequence number for identifying the log sequence number of the last page mirror data area record in the storage data area and a length value for identifying the sum of the length of the file header and the length of the valid data area in the storage data area; And When the database reaches a preset filling condition, filling the search data in the file header data structure into the file header to be filled includes: When switching the page mirror file, filling the maximum sequence number and the length value in the file header data structure into the file header to be filled.

4. The processing method of the database according to claim 3, wherein, Filling each search data in the file header data structure according to various position attributes of the page mirror data areas of the storage data area includes: According to the lengths of the respective page mirror data areas of the storage data area, the length of the valid data area is cumulatively calculated; According to the length of the file header and the cumulatively obtained length of the valid data area, the length value in the file header data structure is updated.

5. The processing method of the database according to claim 2, wherein, The multiple search data includes a minimum sequence number for identifying the log sequence number of the first page mirror data area record in the storage data area; And When the database reaches a preset filling condition, filling the search data in the file header data structure into the file header to be filled includes: When the page mirror data area to be written to disk in the page mirror file includes the first page mirror data area in the storage data area, connecting the file header data structure with the first page mirror data area to form a continuous data area; Writing the continuous data area including the file header data structure and the first page mirror data area to disk in the page mirror file, so as to fill the minimum sequence number in the file header data structure into the file header to be filled.

6. The processing method of the database according to claim 2, wherein, The stored data area is pre-divided into a plurality of search areas, and each of the search areas includes a plurality of the page mirror data areas; The plurality of search data includes a sequence number array and a start address array. The sequence number array includes search sequence numbers for identifying the log sequence numbers of the first page mirror data area record in each of the search areas in the stored data area, and the start address array includes search start addresses for identifying the start positions of each of the search areas in the page mirror file in the stored data area; And When the database reaches a preset filling condition, filling the search data in the file header data structure into the file header to be filled includes: When the database is idle, filling the sequence number array and the start address array in the file header data structure into the file header to be filled.

7. The processing method of the database according to claim 6, wherein The plurality of search data includes the maximum length of the sequence number array, and the maximum length of the sequence number array is used to identify the number of the search areas in the stored data area; And When the database reaches a preset filling condition, filling the search data in the file header data structure into the file header to be filled includes: When the page mirror data area to be written to disk in the page mirror file includes the first page mirror data area in the stored data area, connecting the file header data structure and the first page mirror data area to form a continuous data area; Writing the continuous data area including the file header data structure and the first page mirror data area to disk in the page mirror file, so as to fill the maximum length of the sequence number array in the file header data structure into the file header to be filled.

8. The processing method of the database according to claim 7, wherein The size of each of the search areas is the same; and The maximum length of the sequence number array is calculated from the total length of the page mirror file and the length of each of the search areas.

9. A machine-readable storage medium, on which a machine-executable program is stored, and when the machine-executable program is executed by a processor, the processing method of the database according to any one of claims 1 to 8 is implemented.

10. A computer device, including a memory, a processor, and a machine-executable program stored on the memory and running on the processor, and when the processor executes the machine-executable program, the processing method of the database according to any one of claims 1 to 8 is implemented.