A data writing method and a data reading method
By configuring the historical version storage space in the shared storage space of the database system, the read and write nodes store dirty pages as historical version pages, solving the problem of limited dirty brushing speed and achieving more efficient and flexible dirty brushing operations.
Patent Information
- Application Number
- CN202111508815.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-10
AI Technical Summary
In a relational database system that supports one-write and multiple reads, the dirty brushing operation of the read-write node is limited by the read-only node's execution of data read transactions, resulting in a slow brushing speed.
Configure the historical version storage space in the shared storage space of the database system. When the read and write data, the read and write nodes store dirty pages as historical version pages to the historical version storage space without waiting for the read-only node to read the data.
By decoupling dirty brushing operations from read-only node data reading, dirty brushing efficiency and flexibility are improved, allowing the database system to update to the latest version faster.
Smart Images

Figure CN114385584B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computer application technologies, and in particular, to a data writing method and a data reading method. Background Art
[0002] For a relational database system that supports write-once and multiple reads, it generally includes a read-write node and multiple read-only nodes. The multiple nodes share the same storage space (hereinafter referred to as the shared storage space), so as to complete write-once and multiple reads of the relational database in the shared storage space.
[0003] When the read-write node executes a data change transaction, it does not directly write the changed data into the database, but writes the data change instructions included in the data change transaction into the transaction log (redo log) of the shared storage space. Subsequently, when the read-write node meets the parsing conditions, in order to update the data change to the database, it needs to parse the redo log to obtain each data change instruction, read the page where each data is located from the shared storage space, and store both in the memory of the read-write node. Based on the pages stored in the memory, each data change instruction is executed to obtain the page after the data change (also called the dirty page). When the number of dirty pages reaches a certain value or reaches a fixed period, the read-write node synchronizes the data in the dirty pages to the shared storage space (this process is generally called flushing the dirty), so as to complete the data change in the database.
[0004] In addition, a transaction has a sequence identifier for identifying the order of initiation. The database in the shared storage space also has a sequence identifier, which is used to identify which transaction's instruction (the sequence identifier of the instruction is the sequence identifier of the corresponding transaction) the current database is obtained by executing. For the data reading transaction executed by the read-only node, when this transaction is executed, it is required that the read-only node cannot read data with a sequence identifier greater than the data identifier of this transaction (by default, the larger the sequence identifier, the later the initiation order). Then, in order to ensure that this transaction can be executed normally, the sequence identifier of the database in the shared storage space needs to meet: less than the sequence identifier of the transaction with the smallest sequence identifier among the transactions to be executed in each read-only node.
[0005] This limitation makes the flushing operation of the read-write node wait for the read-only node to complete the transaction, resulting in a slow flushing speed. Summary of the Invention
[0006] In view of this, one or more embodiments of this specification provide a data writing method and a data reading method.
[0007] According to the first aspect of one or more embodiments of this specification, a data writing method is proposed, which is applied to the read-write node of a database system. In the shared storage space of the database system, a historical version storage space for storing historical version pages is configured;
[0008] The read-write node performs the following operations to implement data writing:
[0009] According to the data change instruction in the redo log and the page targeted by the data change instruction, a dirty page is obtained;
[0010] When the dirty page flushing condition is met, the obtained dirty page is updated to the database in the shared storage space; the dirty page flushing condition includes at least reaching a specified period or the number of dirty pages exceeding a specified threshold;
[0011] The obtained dirty page is stored as a historical version page in the historical version storage space.
[0012] According to the second aspect of one or more embodiments of this specification, a data reading method is proposed, which is applied to the read-only node of a database system; the data reading method is used to read data based on the historical version pages obtained by the above data writing method; the method includes:
[0013] Determine the sequence identifier of the currently executed data reading transaction;
[0014] When the determined sequence identifier is less than the sequence identifier of the database, determine the target historical version page of the page targeted by the data reading transaction; the sequence identifier of the target historical version page is less than the sequence identifier of the data reading transaction;
[0015] For the page targeted by the data reading transaction, obtain the target data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction;
[0016] Based on the target historical version page, execute the target data change instruction to obtain the data read by the data reading transaction.
[0017] According to the third aspect of one or more embodiments of this specification, a data writing device is proposed, which is applied to the read-write node of a database system. In the shared storage space of the database system, a historical version storage space for storing historical version pages is configured; the device includes:
[0018] The dirty page management module is used to obtain dirty pages according to the data change instructions in the redo log and the pages targeted by the data change instructions, and update the obtained dirty pages to the database in the shared storage space when the dirty page flushing condition is met. The dirty page flushing condition includes at least reaching a specified period or the number of dirty pages exceeding a specified threshold.
[0019] The historical version page management module is used to store the obtained dirty pages as historical version pages in the historical version storage space.
[0020] According to the fourth aspect of one or more embodiments of this specification, a data reading device is proposed, which is applied to a read-only node of a database system. The data reading device is used to read data based on the historical version pages obtained by the above data writing method. The device includes:
[0021] The sequential identifier determination module is used to determine the sequential identifier of the currently executed data reading transaction.
[0022] The target historical version page determination module is used to determine the target historical version page of the page targeted by the data reading transaction when the determined sequential identifier is less than the sequential identifier of the database. The sequential identifier of the target historical version page is less than the sequential identifier of the data reading transaction.
[0023] The target data change instruction determination module is used to obtain, for the page targeted by the data reading transaction, the target data change instructions in the redo log whose sequential identifiers are greater than the sequential identifier of the target historical version page and less than the sequential identifier of the data reading transaction.
[0024] The data reading module is used to execute the target data change instructions based on the target historical version page to obtain the data read by the data reading transaction.
[0025] According to the fifth aspect of one or more embodiments of this specification, a database system is proposed, which includes a read-write node, at least one read-only node, and a shared storage space accessible to all nodes. In the shared storage space of the database system, a historical version storage space configured to store historical version pages is provided.
[0026] The read-write node performs the following operations to implement data writing:
[0027] Obtain dirty pages according to the data change instructions in the redo log and the pages targeted by the data change instructions.
[0028] Update the obtained dirty pages to the database in the shared storage space when the dirty page flushing condition is met. The dirty page flushing condition includes at least reaching a specified period or the number of dirty pages exceeding a specified threshold.
[0029] Store the obtained dirty page as a historical version page in the historical version storage space.
[0030] According to the sixth aspect of one or more embodiments of this specification, a database system is proposed, including a read-write node, at least one read-only node, and a shared storage space accessible to all nodes; any read-only node reads data based on the historical version pages obtained by the above data writing method.
[0031] Any read-only node executes the following steps to complete data reading:
[0032] Determine the sequence identifier of the currently executed data reading transaction.
[0033] When the determined sequence identifier is less than the sequence identifier of the database, determine the target historical version page of the page targeted by the data reading transaction; the sequence identifier of the target historical version page is less than the sequence identifier of the data reading transaction.
[0034] For the page targeted by the data reading transaction, obtain the target data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction.
[0035] Based on the target historical version page, execute the target data change instruction to obtain the data read by the data reading transaction.
[0036] According to the seventh aspect of one or more embodiments of this specification, a computer-readable storage medium is proposed, on which computer instructions are stored, and when the instructions are executed by a processor, the above data writing method or data reading method is implemented.
[0037] According to the eighth aspect of one or more embodiments of this specification, an electronic device is proposed, including:
[0038] A processor;
[0039] A memory for storing processor-executable instructions;
[0040] Wherein, the processor runs the executable instructions to implement the above data writing method or data reading method.
[0041] According to the ninth aspect of one or more embodiments of this specification, a computer program is proposed, and when the computer program is executed by a processor, the above data writing method or data reading method is implemented.
[0042] In one or more embodiments of this specification, a historical version storage space for storing historical version pages is configured in a shared storage space accessible to all nodes of the database system. When a read / write node writes data, it first obtains the data change instructions that have not been executed in the redo log, as well as the pages targeted by the data change instructions, and executes the obtained data change instructions on the basis of the pages to obtain dirty pages. When a specified period is reached, or when the number of dirty pages exceeds a specified threshold, the dirty pages are flushed, and the dirty pages are stored as historical version pages in the historical version storage space. In this way, the read / write node does not need to wait for the read-only node to read data when flushing dirty pages, decoupling the flushing process from the read operation of the read-only node, and improving the flushing efficiency and flexibility.
[0043] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.
[0045] Figure 1A is a schematic structural diagram of a database system in a related art shown in this specification.
[0046] Figure 1B is a schematic structural diagram of a database system shown in this specification according to an exemplary embodiment.
[0047] Figure 1C is a schematic structural diagram of a database system shown in this specification according to another embodiment.
[0048] Figure 2 is a flowchart of a data writing method shown in this specification according to an exemplary embodiment.
[0049] Figure 3 is a flowchart of a data reading method shown in this specification according to an exemplary embodiment.
[0050] Figure 4A is a schematic structural diagram of a read / write node shown in this specification according to a specific embodiment.
[0051] Figure 4B is a schematic structural diagram of a read / write node shown in this specification according to a specific embodiment.
[0052] Figure 5 is a block diagram of a data writing device shown in this specification according to an exemplary embodiment.
[0053] Figure 6It is a block diagram of a data reading device shown in accordance with an exemplary embodiment of this specification.
[0054] Figure 7 It is a hardware structure diagram of an electronic device where a data writing device or a data reading device is located, shown in accordance with an exemplary embodiment of this specification. Detailed implementation manners
[0055] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0056] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0057] First, a data writing method for a relational database will be described. In a database system, to ensure the writing efficiency, when writing data, the data is not directly written into the database. Instead, it is first written into the redo log, and then written into the database through the redo log. In this way, when the database system executes a data change transaction initiated by a user, it can return a successful write to the user after writing the data change instruction into the redo log. Since the data in the redo log is data change instructions sorted by time, the data written into the redo log is written sequentially; while a relational database is stored in the form of an index tree (B+tree), the data written into the relational database is not written sequentially, which makes the speed of writing the redo log faster than directly writing data into the database. Therefore, the above method can improve the speed of returning data change transactions and enhance the user experience. In addition, to prevent the written data from disappearing due to power outages or other reasons, the redo log is placed on the disk instead of in the memory.
[0058] After explaining the redo log of the relational database, the structure of the database system targeted by this specification will be described below. All database systems mentioned in this specification are database systems that support a one-write-multiple-read architecture. To enable a relational database to support a one-write-multiple-read architecture, by separating computing and storage. In other words, in this database system, a read-write node and multiple read-only nodes are configured, as well as a shared storage space accessible to all nodes of the database system, enabling the database system to support simultaneous reading and writing. Among them, the read-write node has the ability to read data from the shared storage space and write data into the shared storage space, and is mainly responsible for executing data change transactions. The read-only node only has the ability to read the shared storage data and does not have the ability to write data into the shared storage space, and is mainly responsible for executing data reading transactions.
[0059] For such a database system that supports a one-write-multiple-read architecture, the specific method of writing data to its read-write node can be specifically referred to the description in the background technology and will not be elaborated here.
[0060] Since there are multiple nodes in the above database system, user requests or transactions will be sent to multiple nodes, and there is a sequence of initiation between different transactions. Then it is required that the transactions initiated earlier cannot be affected by the transactions initiated later. That is to say, the data reading transaction initiated earlier cannot read the database after the data change transaction initiated later, but can only read the database at the time when the data reading transaction is initiated.
[0061] Then the process of the read-only node reading data is generally as follows: when receiving a user's data reading transaction, determine the sequence identifier of the data reading transaction, obtain the page targeted by the data reading transaction in the database, and obtain the data change instructions in the redo log whose data sequence identifier is smaller than the sequence identifier of the data reading transaction and larger than the sequence identifier of the database (the obtained data change instructions are all for the obtained page), and on the basis of the obtained page, execute the obtained data change instructions to obtain the database at the time when the data reading transaction is initiated, thus completing the data reading.
[0062] Due to the above limitations and the above reading method, if the sequence identifier of the database is greater than the sequence identifier of any data reading transaction, it will cause the data reading transaction to be unable to read the data that needs to be read. Then this makes the dirty page flushing condition not only need to meet the conditions mentioned in the background technology, but also need to meet that the sequence identifier of the database after dirty page flushing is smaller than the sequence identifiers of all unexecuted data reading transactions. In this way, since dirty page flushing needs to wait for the execution of data reading transactions, the dirty page flushing speed is slower.
[0063] In addition, for read-only nodes, to ensure the efficiency of data reading, the read-only nodes do not obtain the data change instructions in the redo log when executing data reading transactions. Instead, while executing data reading transactions, they continuously parse the data change instructions from the redo log and organize the data change instructions into a set of instructions in memory. That is, the data change instructions for the same page are stored together, and these instructions are sorted in the set according to the start time of the corresponding data change transaction. In this way, when the read-only node needs to read data change instructions, it does not need to search for the corresponding data change instructions in the redo log, but only needs to obtain them from the instruction set in memory, which improves the speed of obtaining eligible data change instructions and thus improves the user's data writing speed.
[0064] In this case, since all data change instructions with a sequence identifier in the redo log larger than the sequence identifier of the database need to be stored in memory, if the read-write nodes do not flush dirty pages for a long time, the data in memory will continue to increase, eventually causing memory shortage in the read-only node.
[0065] From the above analysis, it can be seen that in the related technology, the flushing speed of the read-write nodes is restricted by the data reading transactions executed by the read-only nodes, and the data stored in the memory of the read-only nodes is affected by the flushing speed of the read-write nodes. In other words, the read-only node and the read-write node restrict each other, which is the contradiction between improving data writing efficiency and flexible flushing.
[0066] Based on this, in order to achieve flexible flushing without affecting data writing efficiency, considering that the main reason for the above mutual restrictions is that the read-only node needs a way to obtain pages with a sequence identifier smaller than the current transaction. Therefore, by storing historical versions of pages (hereinafter referred to as historical version pages) in the shared storage space, the flushing of the read-write nodes can be freed from the constraints of the read-only nodes. In other words, the read-only node can obtain the data to be read through the historical version pages, so the read-write nodes do not need to care about the sequence identifier of the unexecuted data reading transactions of the read-only nodes when flushing dirty pages.
[0067] This specification proposes a data writing method and a data reading method. In a shared storage space accessible to all nodes of a database system, a historical version storage space for storing historical version pages is configured. When a read-write node writes data, it first obtains the data change instructions that have not been executed in the redo log and the pages targeted by the data change instructions, and executes the obtained data change instructions on the pages to obtain dirty pages. When a specified period is reached, or when the number of dirty pages exceeds a specified threshold, the dirty pages are flushed, and the dirty pages are stored as historical version pages in the historical version storage space. In this way, the read-write node does not need to wait for the read-only node to read data when flushing the dirty pages, decoupling the flushing process from the data reading of the read-only node, and improving the flushing efficiency and flexibility of flushing.
[0068] Next, a data writing method provided in this specification will be described in detail.
[0069] The data writing method provided in this specification is applied to the read-write nodes of a database system. In the shared storage space of the database system, a historical version storage space for storing historical version pages is configured.
[0070] First, the database system mentioned in this specification will be introduced. Generally, the structure of a relational database system that supports one write and multiple reads is as Figure 1A shown. Generally, it includes a read-write node, at least one read-only node, and a shared storage space accessible to all nodes. The shared storage space generally stores a database and a redo log. In the embodiments of this specification, compared with the above structure in the related art, in addition to the above parts, as Figure 1B shown, the shared storage space of this specification also stores the historical version pages of each page (stored in the historical version storage space), so that the read-only node can complete data reading through the historical version pages without passing through the pages of the database, enabling the read-write node to flush the dirty pages without being restricted by the data reading of the read-only node.
[0071] Among them, the meaning of configuring a historical version storage space for storing historical version pages can be to reserve a section of storage space in the shared storage space in advance as the historical version storage space, or not to reserve a section of storage space in advance, but only to enable the shared storage space to store the received historical version pages in a preset manner when receiving the historical version pages.
[0072] After describing the database system, the steps executed by the read-write node and how to complete data writing through these steps will be described in detail. As Figure 2 shown, the read-write node executes the following steps to complete data writing:
[0073] Step 201: Obtain dirty pages based on the data change instructions in the redo log and the pages targeted by the data change instructions
[0074] Specifically, in order to complete the dirty page flushing process, the read-write node first needs to obtain dirty pages, that is, the pages after the data change instructions change the data. The process of obtaining dirty pages is as follows: Obtain the unexecuted data change instructions in the redo log, and obtain the pages where the data changed by the data change instructions is located
[0075] ; Execute the obtained data change instructions on the basis of the obtained pages to obtain dirty pages.
[0076] First, explain each noun in Step 201. The above steps can be executed by the dirty page management module of the read-write node in the related technology. In other words, in the embodiments of this specification, the process of obtaining dirty pages and flushing dirty pages in the related technology is not changed to ensure the pluggability of the newly added capabilities, that is, after removing the newly added capabilities, the database system can still work normally.
[0077] As described above, the redo log stores the data change instructions received by the read-write node in the order of initiation. Data change instructions refer to the change instructions executed on the data in the database. For example, it can be performing a certain mathematical operation on a certain data, or adding a data in a certain page of the database, or deleting a certain data in the database, etc. All instructions that can cause changes to the data stored in the database can be called data change instructions. This specification does not limit the specific form of the data change instructions. In addition, the content of the data change instructions and the data change transactions can be the same, or the data change instructions can be parsed from the data change transactions. This specification does not make any limitations on this.
[0078] It should also be noted that the obtained dirty pages are stored in the memory of the read-write node. That is to say, executing the obtained data change instructions on the basis of the obtained pages means storing the obtained pages in the memory and executing the obtained data change instructions on the data in the pages stored in the memory to obtain dirty pages.
[0079] In addition, it should be noted that this specification does not change the process of the read-write node in the related technology for updating the redo log in the shared storage space (writing the received data change instructions into the redo log) to ensure that the time from when the user initiates a transaction to when the transaction is completed does not change, and to ensure the user experience.
[0080] Step 203: Update the obtained dirty pages to the database in the shared storage space when the dirty page flushing condition is met; the dirty page flushing condition includes at least reaching a specified period or the number of dirty pages exceeding a specified threshold.
[0081] Step 205: Store the obtained dirty page as a historical version page in the historical version storage space.
[0082] Next, steps 203 and 205 will be described together. Among them, steps 203 and 205 executed by the read-write node can be executed in parallel or serially, and this specification does not limit this here.
[0083] Specifically, after obtaining the dirty page, in addition to maintaining the original dirty page flushing path to store the dirty page in the database of the shared storage space, it is also necessary to store the historical version page in the shared storage space, so as to decouple the dirty page flushing process of the read-write node and the data reading process of the read-only node, thereby accelerating the dirty page flushing.
[0084] After the overall description of the entire step, the following will describe each sentence and noun in steps 203 and 205.
[0085] First of all, the above-mentioned dirty page flushing conditions at least include reaching a specified period or the number of dirty pages exceeding a specified threshold. In other words, compared with the dirty page flushing conditions in the related art, the dirty page flushing conditions in the embodiments of this specification do not include waiting for the read-only node to complete the reading of the transaction. Then there will be no limitation that the sequence identifier of the database needs to be less than the sequence identifiers of all unexecuted data reading transactions. Then the read-write node can flush the dirty page flexibly.
[0086] In addition, the process of flushing the dirty page and the process of storing the historical version page can be executed simultaneously or not together. For the storage of the historical version page, it can be that when the dirty page flushing is completed, the dirty page copy is incidentally used as the historical version page and stored in the historical version storage space; it can also be that the dirty page flushing is executed by the dirty page management module. While flushing the dirty page, the dirty page management module sends the dirty page copy to the historical version page management module. The historical version page management module determines the storage location of the historical version page in the historical version storage space and stores the historical version page at the corresponding storage location. In this way, the dirty page management module can execute the next task after completing the dirty page flushing without waiting for the historical version page to be stored before executing the next task.
[0087] In addition, if the historical version page management module is responsible for storing the historical version page, the historical version page management module can store it after receiving the historical version page. This can ensure that the historical version page is stored in the shared storage space as soon as possible and enables the read-only node to use the updated historical version page. It can also be that after receiving multiple historical version pages, when the number of historical version pages exceeds the preset page threshold or reaches the preset period, the multiple historical version pages are stored together, which can improve the storage efficiency.
[0088] After explaining the storage process of the historical version page, the historical version page will be described next. To enable the read-only node to select a suitable historical version page from multiple historical version pages of a certain page, each historical version page has a sequence identifier (in addition, the database also has a sequence identifier, each transaction also has a sequence identifier, and the data change instruction also has a sequence identifier. The sequence identifier of the data change instruction is also the sequence identifier of the corresponding data change transaction). The sequence identifier of each historical version page is: the sequence identifier of the last (with the largest sequence identifier) data change instruction executed during the process of obtaining this historical version page. The sequence identifier of the database is similar to that of the historical version page, which is the sequence identifier of the last data change instruction executed during the process of obtaining this database.
[0089] Next, the sequence identifier will be described. The sequence identifier is used to identify the order of initiation time of each transaction, and its existence form can be a timestamp used to identify the initiation time of the transaction. If the sequence identifier is represented by a timestamp, then the earlier the time, the smaller the sequence identifier.
[0090] In addition, the sequence identifier can also be represented by the statement position in the redo log, that is, represented by the Log Sequence Number (LSN).
[0091] If the sequence identifier is represented by the statement position in the redo log, then for the sequence identifiers of the data change transaction and the data change instruction, since the data change instruction needs to be written into the redo log, the sequence identifier of the data change instruction can be represented by its lsn in the redo log, and the data change transaction can use the sequence identifier of its corresponding data change instruction as the sequence identifier of this data change transaction.
[0092] If the sequence identifier is represented by the statement position in the redo log, for the sequence identifier of the data read transaction, since the data read transaction will not be written into the redo log, the method for obtaining the sequence identifier of the data read transaction can be: when the data read transaction is generated, obtain the lsn of the last data change instruction in the current redo log, and use the obtained lsn as the sequence identifier of this data read transaction.
[0093] After explaining the basic concept of the historical version page, the maintenance method of the historical version page will be described next.
[0094] Since the shared storage space is limited, the number of historical version pages for each page cannot be infinite. To ensure that the historical version pages do not occupy too much space in the shared storage space, it is necessary to delete some historical version pages when the number of historical version pages of any page exceeds a preset threshold, so as to maintain the stability of the number of historical version pages.
[0095] In other words, the method further includes: for any page, when the number of historical version pages of the page exceeds the preset threshold of the number of historical version pages, determining the historical version pages to be deleted, and deleting the determined historical version pages. In this way, the maintenance of the historical version pages can be completed, so as to ensure that the number of historical version pages does not occupy too much space in the shared storage space and cause nowhere to store other files.
[0096] Among them, for how to determine the historical version pages to be deleted, several historical version pages can be randomly determined from each historical version page for deletion (it is necessary to ensure that all data reading transactions can be executed, and it is best not to delete the historical version page with the smallest sequence identifier), or the historical version pages to be deleted can be screened according to certain conditions. For example, the difference in the sequence identifiers of each adjacent historical version page after deletion can be made the same. The implementation logics of the above several methods are relatively simple.
[0097] In addition, to ensure the reading efficiency of the read-only nodes, the historical version pages to be deleted can also be screened according to the sequence identifiers of the data reading transactions. Specifically, determining the historical version pages to be deleted includes: determining the sequence identifiers of the unexecuted data reading transactions for the page in each read-only node; for each determined data reading transaction, determining the reference version page of the data reading transaction according to the version number of the data reading transaction; the sequence identifier of the reference version page is less than the sequence identifier of the data reading transaction, and the difference between the sequence identifier of the data reading transaction and the sequence identifier of the reference version page is not greater than the difference between the sequence identifier of the data reading transaction and the sequence identifier of any historical version page; determining the number of times each historical version page is used as a reference version page, and taking the historical version page with the smallest number of times as the historical version page to be deleted.
[0098] In other words, the least needed historical version pages by the data reading transactions are determined and deleted, so as to improve the reading and writing efficiency. For example, if the sequence identifier of a certain data reading transaction is 200, and there are 3 historical version pages for the page targeted by the data reading transaction, and their sequence identifiers are 100, 150, and 250 respectively, then the reference version page of the data reading transaction is the historical version page with the sequence identifier of 150. Through this reference historical version page, it can obtain the data to be read faster (less data change instructions need to be executed).
[0099] Through the method of determining the historical version pages to be deleted, the data reading tasks of read-only nodes are preferentially deleted, thus ensuring the reading efficiency of read-only nodes while ensuring that the shared storage space is not occupied too much.
[0100] In addition, regarding the historical version pages, it should also be noted that in order to ensure the pluggability of the new capabilities, a new file format can be created for the files storing the historical version pages. For example, if the pages of the database are stored in ibd files, then the historical version pages can be stored in files with the file format of mibd. In this way, when this function is not needed, the mibd files can be directly deleted, thus making the new capabilities pluggable.
[0101] Moreover, the above-mentioned maintenance process of the historical version pages can be executed by the historical version page management module mentioned above, which is a new module added in this specification compared with the related technologies. When the maintenance process of the historical version pages is executed by the historical version page management module, the maintenance process of the historical version pages can be triggered when storing the historical version pages, or the maintenance process of the historical version pages and the process of storing the historical version pages can be executed in parallel, that is, the historical version pages are maintained every certain period.
[0102] In addition, in addition to the maintenance of the historical version pages, the redo log will also be maintained. In the related technologies, the redo log that has been executed and the execution result has been updated to the database will be deleted. Since this part of the redo log will no longer be used, it can of course also be deleted some time after it will no longer be used. When deleting, for convenience, generally the entire file is selected for deletion. In the method provided in this specification, since the sequence identifier of the database is smaller than the sequence identifier of the data reading transaction, if the redo log to be deleted is selected according to the sequence identifier of the database, it may cause some data reading transactions to be unable to execute. Therefore, the redo log to be deleted cannot be determined according to the sequence identifier of the database.
[0103] In the embodiments of this specification, the deletion of the redo log can be based on the sequence identifier of the data reading transaction. For example, the redo log that will no longer be used can be deleted, that is, the redo log file with a sequence identifier smaller than all the sequence identifiers of the data reading transactions (that is, the data reading transaction with the largest sequence identifier in this redo log is smaller than all the sequence identifiers of the data reading transactions) can be deleted. It is also mentioned later that there is an instruction set. In the presence of the instruction set, the read-only node can read data according to the instruction set. Then, the redo log files that have been sorted into the instruction set can be deleted.
[0104] After the maintenance method for historical versions is described, another operation in this specification that can improve the data reading efficiency of read-only nodes will be described below.
[0105] Specifically, as Figure 1C shown, in the shared storage space, there is also an instruction storage space for storing instruction sets. For each page in the instruction set, a data change instruction set for that page is stored.
[0106] The instruction set is the instruction set that the original read-only node obtained in memory. By storing the instruction set in the shared storage space, when the read-only node reads data, it can complete data reading based on the historical version page plus the instruction set read from the shared storage space. In this way, the read-only node does not need to reorganize the instructions in the redo log, improving the read and write efficiency and freeing up the memory space of the read-only node.
[0107] In addition, for the convenience of searching, the data change instructions for each page in the instruction set can be stored in the order of the size of the sequence identifiers. The smaller the sequence identifier, the more forward its position.
[0108] In this case, the instruction set is specifically generated by the instruction set management module. Specifically, the read-write node further includes an instruction set management module, which is specifically used for: obtaining the unorganized data change instructions in the redo log that have not been stored in the instruction set; organizing the unorganized data change instructions according to the pages they target, and storing the organized data change instructions in the instruction set of the instruction storage space
[0109] Among them, the instruction set management module can run in parallel with other modules mentioned above, which can improve the processing efficiency.
[0110] This completes the description of the data writing process.
[0111] Next, a data reading method provided in this specification will be described. The data reading method provided in this specification is applied to the read-only node of the database system. The method is used to read data based on the database updated by the above data writing method and the obtained historical version pages.
[0112] The database system in this method is the same as the database system in the above data writing method, which will not be elaborated here. In addition, some nouns in the following text have been described in the description of the data writing method and will not be elaborated later.
[0113] As Figure 3 shown, the data reading method shown in this specification includes the following steps:
[0114] Step 301: Determine the sequence identifier of the currently executed data reading transaction.
[0115] Step 303: When the determined sequence identifier is less than the sequence identifier of the database, determine the target historical version page of the page targeted by the data reading transaction; the sequence identifier of the target historical version page is less than the sequence identifier of the data reading transaction.
[0116] Next, steps 301 - 303 will be described.
[0117] Specifically, a read-only node cannot obtain the data to be read through a future page (i.e., a page with a sequence identifier greater than the currently executed data reading transaction page), so it is necessary to enable the read-only node to obtain a page with a sequence identifier less than the sequence identifier of the currently executed data reading transaction.
[0118] The method for obtaining the target historical version page can be to arbitrarily select a historical version page that meets the requirements as the target historical version page. Further considering that for a read-only node, the smaller the absolute value of the difference between the sequence identifier of the target historical version page and the sequence identifier of the currently executed data reading transaction, the fewer data change instructions the read-only node needs to execute to obtain the read data, which will make the read-write efficiency higher.
[0119] Therefore, the target historical version page can be obtained by the following method: determine all historical version pages of the page targeted by the data reading transaction; determine the target historical version page from all historical version pages; wherein, the difference between the sequence identifier of the data reading transaction and the sequence identifier of the target historical version page is not greater than the difference between the sequence identifier of the data reading transaction and the sequence identifier of any historical version page, and the sequence identifier of the target version page is less than the sequence identifier of the data reading transaction.
[0120] In this way, the read-write efficiency of the read-only node can be increased and the performance can be improved.
[0121] Step 305: For the page targeted by the data reading transaction, obtain the target data change instructions in the redo log whose sequence identifiers are greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction.
[0122] Step 307: On the basis of the target historical version page, execute the target data change instructions to obtain the data read by the data reading transaction.
[0123] Next, steps 305 and 307 will be described.
[0124] Specifically, in order to obtain the data read by the data reading transaction, it is necessary to first obtain the page with the data reading transaction sequence identifier. Then, it is necessary to execute on the basis of the obtained target historical version page: the data change instruction between the sequence identified by the target historical version page sequence identifier and the sequence identified by the data reading transaction sequence identifier (such a data change instruction is hereinafter referred to as the target data change instruction).
[0125] For obtaining the target data change instruction, it can be the same as in the related art, and the target data change instruction is obtained by reading the redo log.
[0126] In addition, on the basis that there is an instruction set in the shared storage space, the target data change instruction can be obtained from the instruction set. In this way, the read-only node does not need to read the redo log and does not need to sort out the data change instructions page by page, so that the memory space can be released. And when executing the data reading instruction, the required instruction can be directly obtained from the instruction set without waiting for the read-only node to sort out the data change instructions, so that the read-only node can improve the data reading efficiency.
[0127] Then, in the case that the shared storage space also stores an instruction set (for each page in the instruction set, a data change instruction set for the page is stored), step 305 may specifically include:
[0128] Step 3051, determining the page targeted by the data reading transaction.
[0129] Step 3052, for the determined page, in the instruction set, obtaining a reference data change instruction whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction. (The above two steps Figure 3 are not shown).
[0130] In addition, the above content introduces the method of reading data according to the historical version page. In the case that the determined sequence identifier is greater than or equal to the sequence identifier of the database, data can also be read through the page in the database. In this case, the following steps can be executed to read the data: obtaining the page targeted by the data reading transaction in the database; for the page targeted by the data reading transaction, obtaining a reference data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the obtained page and less than the sequence identifier of the data reading transaction; executing the obtained reference data change instruction on the basis of the obtained page to obtain the data read by the data reading transaction.
[0131] Through the above data writing method and data reading method, the multi-version pages originally stored in the memory are stored in the shared storage space, and the multi-version pages are provided through the shared storage space, eliminating the dirty page flushing constraint between the read-write node and the read-only node, making the dirty page flushing of the read-write node more flexible, enabling the database to be updated to the latest version faster, improving the performance of the database system, and releasing the memory space of the read-only node. At the same time, through operations such as instruction sets and selecting target historical version pages, the data reading efficiency of the read-only node is improved.
[0132] Next, a specific embodiment will be used to illustrate the data writing method and data reading method provided in this specification.
[0133] The database system includes a read-write node, multiple read-only nodes, and a shared storage space. The read-write node has the permission to read and write data to the shared storage space; the read-only node only has the permission to read the data in the shared storage space and does not have the permission to write data to the shared storage space. A storage engine is installed on the server where the shared storage space is located, responsible for managing the shared storage space and processing requests such as data writing and reading; a computing engine is also installed on the read-only node and the read-write node. In addition, these parts of the database can be on different devices, or different nodes can be isolated by virtualizing multiple virtual machines on the same device. The structure of the computing engine installed on the read-write node is as Figure 4A shown, and the structure of the computing engine installed on the read-only node is as Figure 4B shown.
[0134] First, introduce each part in Figure 4A . Primary refers to the master node, that is, the read-write node. Transaction and B+tree are the transaction module and the index tree module respectively, responsible for processing transactions and managing the index tree. The implementation of the method in this specification does not involve these two modules and will not be elaborated further later. Share storage is the shared storage space. It should be noted that it is not part of the read-write node.
[0135] Redo refers to the transaction log management module, redo hash is the instruction set, persistent redo hash is the instruction set management module, buffer pool refers to the dirty page management module, and Multi-version page store refers to the historical version page management module. Each module can execute in parallel. The instruction set management module and the historical version page management module enclosed by the dashed box are the modules newly added in this specification compared with the related technology.
[0136] Figure 4BThe read-only replicas and the read-write replicas have the same engine installed, but different modules will have functional changes according to the identity of the current node. In other words, Figure 4B The names of the various modules in Figure 4A are the same, but there are differences in functionality.
[0137] Next, the steps executed by each module on the read-write replicas and the read-only replicas will be introduced in detail.
[0138] For the read-write replicas, the functions of the various modules are as described below.
[0139] The transaction log management module is responsible for writing the data change instructions in the data change transaction into the redo log in the shared storage space. It should be noted that the data in the redo log is sorted according to the writing order, and the sequence identifiers of different transactions are characterized by their positions (i.e., lsn) in the redo log. It can be seen that the embodiments of this specification do not change the path of writing the redo log in the related art, ensuring that the user experience will not be affected by the new capabilities.
[0140] The instruction set management module is responsible for reading the data change instructions in the redo log, sorting the data change instructions by page, and then writing them into the instruction set. For the specific implementation, refer to the above description.
[0141] The dirty page management module is responsible for obtaining dirty pages (the steps for obtaining dirty pages are detailed in the above description) and performing the dirty page flushing operation. It is also responsible for sending the dirty pages to the historical version page management module.
[0142] The historical version page management module is responsible for receiving the dirty pages and storing the dirty pages as historical version pages in the shared storage space.
[0143] For the read-only replicas, the functions of the various modules are as described below.
[0144] Log apply, that is, the transaction log management module, is responsible for reading the data change instructions in the redo log and sending the data change instructions to the dirty page management module to update the dirty pages managed by the dirty page management module.
[0145] It should be noted that in the case where the instruction set has already been stored, the reason for retaining this step is, firstly, to make the new capabilities pluggable, and secondly, to retain two paths to prevent the data reading efficiency from decreasing due to untimely reading of the instruction set.
[0146] The historical version page management module is responsible for reading the target historical version page from the shared storage according to the sequence identifier of the data reading transaction (for the method of obtaining the target historical version page, please refer to the above description), and on the basis of the target historical version page, obtaining a version page with the same sequence identifier as the data reading transaction according to the data change instruction, and sending the obtained version page to the dirty page management module.
[0147] The method of obtaining the target historical version page can be seen in the following example: when the sequence identifier of the data reading transaction is 150, the page targeted by the current data reading transaction has 3 historical version pages, namely 100, 125, and 200. Then, according to the selection criteria mentioned in the above method, the historical version page with the sequence identifier of 125 can be selected as the target historical version page.
[0148] The instruction set management module is responsible for reading the instruction set to provide the required target data change instruction for the historical version page according to the sequence identifier of the target historical version page and the sequence identifier of the data reading transaction.
[0149] The method of obtaining the target data change instruction can be seen in the following example: when the sequence identifier of the data reading transaction is 150 and the sequence identifier of the target historical version page is 125, then the target data change instruction is the data change instruction for this page with the sequence identifier between 125 and 150.
[0150] The dirty page management module is responsible for obtaining the data that the user needs to read according to the transaction log management module or according to the version page provided by the historical version page (mentioned above).
[0151] For the dirty page management module, in order to improve the data reading efficiency, it can first judge whether it can obtain the data to be read (the original path) according to the data change instruction provided by the transaction log management module and the pages already existing in the read-only node memory. If not, it then requests the corresponding version page from the historical version page management module.
[0152] After explaining the overall functions of the two nodes, the following will explain some details not mentioned above.
[0153] Regarding the instruction set, first of all, it should be noted that for the convenience of searching, each page in the instruction set is indexed according to the hash value of the identifier of this page and the identifier of the table where this page is located.
[0154] In addition to being available for read-only nodes, the instruction set can also be used by the crash recovery module. That is to say, in the case of data loss in the database, data can be quickly restored through the instruction set in the shared storage space. Compared with restoring based on the redo log, restoring through the instruction set can save the time for finding the instructions for the lost pages and achieve a faster restoration.
[0155] Regarding the correspondence between the instruction set and the redo log, generally one redo log file corresponds to one instruction set file. This setting is mainly to consider the balance between reading efficiency and management efficiency.
[0156] Specifically, if multiple redo log files correspond to 1 instruction set file, then when the read-only node reads data, it can obtain most of the required data change instructions through only 1 instruction set file without having to read multiple files, which improves the data reading efficiency. However, correspondingly, when the redo log file needs to be deleted, the instruction set corresponding to the redo log file needs to be deleted. Then, it is necessary to parse the instruction set to determine the data change instructions that need to be deleted, making the deletion process complicated.
[0157] Correspondingly, if one redo log file corresponds to multiple instruction set files, then when the read-only node reads data, it may need to obtain the corresponding data change instructions from multiple instruction set files. Moreover, it may not know in advance which instructions in the redo log each instruction set file corresponds to. It may be necessary to parse multiple instruction set files to determine which instruction set files need to be read, resulting in a lower data reading efficiency. However, when the redo log file needs to be deleted, it will make the deletion process more flexible.
[0158] Therefore, to balance the reading efficiency and management efficiency, it is chosen that one redo log file corresponds to 1 instruction set file.
[0159] After determining the correspondence between the redo log files and the instruction set files, to facilitate storage space planning, it is also necessary to set how many pages each redo log file corresponds to. After simulation, it is found that for a 1G-sized redo log file, in the case of only write operations, the redo log file will probably contain about 50,000 - 80,000 pages. In the case of both write operations and delete and upgrade operations, the redo log file will probably contain about 100,000 pages. Considering that in the simulation scenario, the fields of the tables used are relatively small, then in the actual user scenario, the number of pages in each redo log will not exceed 100,000. Then, it is possible to first set that each instruction set contains 100,000 pages, so as to design the storage of the instruction set file.
[0160] For the storage design of the instruction set file, it can be designed with Block as the storage unit. Among them, Block0 is used as the Header Block, which records the identifier of the Blocks already allocated in the current file, the identifier of the newly allocated HeaderBlock, as well as the starting sequence identifier and ending sequence identifier of the data change instructions contained, etc. Block1 to Block255 are all Blocks used to store data change instructions. The entire Block is divided into 512 slots. The first 4 slots are used as the Header to store relevant meta-information. The last slot is used as the tail. The heard stores the starting sequence identifier and ending sequence identifier of the data change instructions on this Block (that is, the minimum and maximum values of the sequence identifiers on this Block). When the current block is full, the tail will record the identifier of the next Block of this page. The middle 507 slots are used to store data change instructions. The information stored in each Slot includes the hash index value of the identifiers of the table and the page (the index values of different tables and pages are different), as well as the starting sequence identifier and ending sequence identifier of the recorded data change instructions.
[0161] After a detailed description of the redo log, it is also necessary to describe the process of reading historical version pages. To ensure the pluggability of the new capabilities, the corresponding page in the database needs to be read before reading the historical version page. Since the historical version pages are stored in the same order as the pages in the database, the number of historical version pages for each page is determined. When the number of historical version pages of a page is less than the predetermined number, the corresponding position in the file is left empty. Then, the position of the page to be read in the historical version storage space can be determined based on the position of the corresponding page in the database. For example, if the number of historical version pages for each page is preset to 4, and the starting position of a certain page in the database is A (the distance from the starting point of the first page in the database), then the starting position of this page in the historical version storage space is 4A (the distance from the starting point of the first page in the historical version storage space). In this way, the position of the required historical version page can be quickly determined using the currently read data. Of course, the position of the page to be read in the historical version storage space can also be calculated based on the page identifier.
[0162] Corresponding to the embodiments of the foregoing method, this specification also provides embodiments of an apparatus and a terminal to which the apparatus is applied.
[0163] As Figure 5 shown, Figure 5 FIG. is a data writing apparatus shown in this specification according to an exemplary embodiment, which is applied to a read / write node of a database system. In the shared storage space of the database system, a historical version storage space for storing historical version pages is configured; the apparatus includes:
[0164] A dirty page management module 510, configured to obtain a dirty page according to a data change instruction in the redo log and the page targeted by the data change instruction; and update the obtained dirty page to the database in the shared storage space when the dirty page flushing condition is satisfied; the dirty page flushing condition at least includes reaching a specified period or the number of dirty pages exceeding a specified threshold.
[0165] A historical version page management module 520, configured to store the obtained dirty page as a historical version page in the historical version storage space.
[0166] In addition, in the shared storage space, an instruction storage space for storing an instruction set is further configured. For each page in the instruction set, a data change instruction set for the page is stored; the apparatus further includes an instruction set management module (not shown in the figure), specifically configured to: obtain an unorganized data change instruction in the redo log that has not been stored in the instruction set; organize the unorganized data change instructions according to the pages they target, and store the organized data change instructions in the instruction set in the instruction storage space.
[0167] In addition, the device further includes a historical version page management module, which is further configured to: for the historical version page of any page, when the number of historical version pages of this page exceeds a preset historical version page number threshold, determine the historical version pages to be deleted, and delete the determined historical version pages.
[0168] Among them, the historical version pages determined to be deleted in the historical version page management module include: determining the sequence identifier of the unexecuted data reading transaction for this page in each read-only node; for each determined data reading transaction, determining the reference version page of this data reading transaction according to the version number of this data reading transaction; the sequence identifier of the reference version page is less than the sequence identifier of this data reading transaction, and the difference between the sequence identifier of this data reading transaction and the sequence identifier of the reference version page is not greater than the difference between the sequence identifier of this data reading transaction and the sequence identifier of any historical version page; determining the number of times each historical version page is used as a reference version page, and taking the historical version page with the smallest number of times as the historical version page to be deleted.
[0169] As Figure 6 shown, Figure 6 is a data reading device shown in this specification according to an exemplary embodiment, which is applied to a read-only node of a database system; the data reading device is used to read data based on the historical version pages obtained by the above data writing method; the device includes:
[0170] A sequence identifier determination module 610, configured to determine the sequence identifier of the currently executed data reading transaction.
[0171] A target historical version page determination module 620, configured to determine the target historical version page of the page targeted by the data reading transaction when the determined sequence identifier is less than the sequence identifier of the database; the sequence identifier of the target historical version page is less than the sequence identifier of the data reading transaction.
[0172] A target data change instruction determination module 630, configured to obtain a target data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction for the page targeted by the data reading transaction.
[0173] A data reading module 640, configured to execute the target data change instruction on the basis of the target historical version page to obtain the data read by the data reading transaction.
[0174] Among them, the target historical version page determination module 620 is specifically configured to: determine all historical version pages of the page targeted by the data reading transaction; determine the target historical version page from all historical version pages; wherein, the difference between the sequence identifier of the data reading transaction and the sequence identifier of the target historical version page is not greater than the difference between the sequence identifier of the data reading transaction and the sequence identifier of any historical version page, and the sequence identifier of the target version page is less than the sequence identifier of the data reading transaction.
[0175] In addition, the shared storage space further stores an instruction set; for each page in the instruction set, a data change instruction set for the page is stored; the target data change instruction determination module 630 is specifically configured to: determine the page targeted by the data reading transaction; for the determined page, in the instruction set, obtain a reference data change instruction whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction.
[0176] For the implementation processes of the functions and roles of each module in the above device, specifically refer to the implementation processes of the corresponding steps in the above method, which will not be elaborated here.
[0177] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this specification. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0178] In addition, this specification also provides a database system, including a read-write node, at least one read-only node, and a shared storage space accessible to all nodes; in the shared storage space of the database system, a historical version storage space configured to store historical version pages is provided.
[0179] The read-write node performs the following operations to implement data writing:
[0180] According to the data change instruction in the redo log and the page targeted by the data change instruction, a dirty page is obtained.
[0181] When the dirty page flushing condition is met, the obtained dirty page is updated to the database in the shared storage space; the dirty page flushing condition at least includes reaching a specified period or the number of dirty pages exceeding a specified threshold.
[0182] Store the obtained dirty page as a historical version page in the historical version storage space.
[0183] This specification also provides a database system, including a read-write node, at least one read-only node, and a shared storage space accessible to all nodes; any read-only node reads data based on the historical version pages obtained by the above data writing method; any read-only node performs the following steps to complete data reading:
[0184] Determine the sequence identifier of the currently executed data reading transaction.
[0185] When the determined sequence identifier is less than the sequence identifier of the database, determine the target historical version page of the page targeted by the data reading transaction; the sequence identifier of the target historical version page is less than the sequence identifier of the data reading transaction.
[0186] For the page targeted by the data reading transaction, obtain the target data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction.
[0187] Based on the target historical version page, execute the target data change instruction to obtain the data read by the data reading transaction.
[0188] For the specific steps executed by the above read-only node and read-write node, refer to the description of the data reading method and data writing method in the specification, which will not be elaborated here.
[0189] As Figure 7 shown, Figure 7 shows a hardware structure diagram of a computer device where the data writing device or data reading device of the embodiment is located. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0190] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this specification.
[0191] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store the operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and called and executed by the processor 1010.
[0192] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input devices can include keyboards, mice, touchscreens, microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0193] The communication interface 1040 is used to connect to the communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0194] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0195] It should be noted that although only the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050 are shown in the above device, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification and do not have to include all the components shown in the figure.
[0196] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the data writing method or the data reading method as described above.
[0197] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0198] This specification also provides a computer program, which, when executed, implements the above-mentioned data reading method or data writing method.
[0199] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0200] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A data writing method is applied to a read-write node of a database system. The database system includes a read-write node, at least one read-only node, and a shared storage space accessible to all nodes. In the shared storage space, a historical version storage space for storing historical version pages is configured. The read-write node performs the following operations to implement data writing: Obtain a dirty page according to the data change instruction in the transaction log redo log and the page targeted by the data change instruction. Under the condition of meeting the dirty page flushing condition, update the obtained dirty page to the database in the shared storage space. The dirty page flushing condition at least includes reaching a specified period or the number of dirty pages exceeding a specified threshold. Store the obtained dirty page as a historical version page in the historical version storage space so that the read-only node reads data based on the historical version page.
2. According to the method described in claim 1, in the shared storage space, an instruction storage space for storing an instruction set is further configured. For each page, the instruction set stores a data change instruction set for that page. The method further includes: Obtain the uncollated data change instructions in the redo log that have not been stored in the instruction set. Collate the uncollated data change instructions according to the pages they target and store the collated data change instructions in the instruction set of the instruction storage space.
3. According to the method described in claim 1, the method further includes: For any page, when the number of historical version pages of the page exceeds the preset threshold of the number of historical version pages, determine the historical version pages to be deleted and delete the determined historical version pages.
4. According to the method described in claim 3, the determination of the historical version pages to be deleted includes: Determine the sequence identifier of the unexecuted data reading transaction for the page in each read-only node. For each determined data reading transaction, determine the reference version page of the data reading transaction according to the version number of the data reading transaction. The sequence identifier of the reference version page is less than the sequence identifier of the data reading transaction, and the difference between the sequence identifier of the data reading transaction and the sequence identifier of the reference version page is not greater than the difference between the sequence identifier of the data reading transaction and the sequence identifier of any historical version page. Determine the number of times each historical version page is used as a reference version page, and use the historical version page with the smallest number of times as the historical version page to be deleted.
5. A data reading method is applied to a read-only node of a database system. The data reading method is used to read data based on the historical version pages obtained by the data writing method described in any one of claims 1-4. The method includes: Determine the sequence identifier of the currently executed data reading transaction. When the determined sequence identifier is less than the sequence identifier of the database, determine the target historical version page of the page targeted by the data reading transaction. The sequence identifier of the target historical version page is less than the sequence identifier of the data reading transaction. Obtain a target data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction, for the page targeted by the data reading transaction; Based on the target historical version page, execute the target data change instruction to obtain the data read by the data reading transaction.
6. The method according to claim 5, wherein determining the target historical version page of the page targeted by the data reading transaction comprises: Determine all historical version pages of the page targeted by the data reading transaction; Determine the target historical version page from all historical version pages; wherein, the difference between the sequence identifier of the data reading transaction and the sequence identifier of the target historical version page is not greater than the difference between the sequence identifier of the data reading transaction and the sequence identifier of any historical version page, and the sequence identifier of the target version page is less than the sequence identifier of the data reading transaction.
7. The method according to claim 5, wherein the shared storage space further stores an instruction set; for each page in the instruction set, a data change instruction set for the page is stored; The step of obtaining a reference data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction, for the page targeted by the data reading transaction, comprises: Determine the page targeted by the data reading transaction; For the determined page, in the instruction set, obtain a reference data change instruction whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction.
8. A data writing device is applied to a read-write node of a database system, where the database system includes a read-write node, at least one read-only node, and a shared storage space accessible to all nodes; In the shared storage space, configure a historical version storage space for storing historical version pages; the apparatus comprises: A dirty page management module, configured to obtain dirty pages according to the data change instructions in the redo log and the pages targeted by the data change instructions, and update the obtained dirty pages to the database in the shared storage space when the dirty page flushing condition is met; the dirty page flushing condition at least includes reaching a specified period or the number of dirty pages exceeding a specified threshold; A historical version page management module, configured to store the obtained dirty pages as historical version pages in the historical version storage space, so that the read-only node reads data based on the historical version pages.
9. A data reading apparatus, applied to a read-only node of a database system; the data reading apparatus is configured to read data based on the historical version pages obtained by the data writing method according to any one of claims 1-4; the apparatus comprises: A sequence identifier determination module, configured to determine the sequence identifier of the currently executed data reading transaction; A target historical version page determination module, configured to determine the target historical version page of the page targeted by the data reading transaction when the determined sequence identifier is less than the sequence identifier of the database; the sequence identifier of the target historical version page is less than the sequence identifier of the data reading transaction; A target data change instruction determination module, configured to obtain, for the page targeted by the data reading transaction, a target data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction; A data reading module, configured to execute the target data change instruction based on the target historical version page to obtain the data read by the data reading transaction.
10. A database system, comprising a read-write node, at least one read-only node, and a shared storage space accessible to all nodes; in the shared storage space of the database system, a historical version storage space configured to store historical version pages; The read-write node performs the following operations to implement data writing: Obtain a dirty page according to the data change instruction in the redo log and the page targeted by the data change instruction; When the dirty page flushing condition is met, update the obtained dirty page to the database in the shared storage space; the dirty page flushing condition at least includes reaching a specified period or the number of dirty pages exceeding a specified threshold; Store the obtained dirty page as a historical version page in the historical version storage space, so that the read-only node reads data based on the historical version page.
11. A database system, comprising a read-write node, at least one read-only node, and a shared storage space accessible to all nodes; any read-only node completes data reading based on the historical version page obtained by the data writing method according to any one of claims 1-4; Any read-only node performs the following steps to complete data reading: Determine the sequence identifier of the currently executed data reading transaction; When the determined sequence identifier is less than the sequence identifier of the database, determine the target historical version page of the page targeted by the data reading transaction; the sequence identifier of the target historical version page is less than the sequence identifier of the data reading transaction; For the page targeted by the data reading transaction, obtain a target data change instruction in the redo log whose sequence identifier is greater than the sequence identifier of the target historical version page and less than the sequence identifier of the data reading transaction; Execute the target data change instruction based on the target historical version page to obtain the data read by the data reading transaction.
12. An electronic device, comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor runs the executable instructions to implement the method according to any one of claims 1-4 or 5-8.
13. A computer-readable storage medium, storing computer instructions, where the computer instructions, when executed by a processor, implement the method according to any one of claims 1-4 or 5-8.
14. A computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-4 or 5-8.
Citation Information
Patent Citations
File system page cache writeback method, system and device and storage medium
CN107590287A
Database processing method, device and system
CN110019066A