Data processing method, database cluster and computing device cluster
By obtaining the snapshot read submission sequence number in the read node and reading the copy of the first page from the write node without adding a global lock, the problem of poor database system scalability and poor read and write transaction performance under the shared-page architecture is solved, and efficient data reading and writing is achieved.
Patent Information
- Application Number
- CN202311790247.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
Under the shared-page architecture, the scalability and read and write transaction processing performance of the database system are poor, mainly due to the existence of global page locks.
By obtaining the snapshot read commit sequence number in the read request in the read node, and reading a copy of the first page from the write node without a global lock, the copy of the first page contains the minimum commit sequence number and the maximum commit sequence number, ensuring that the read data meets the requirements.
It realizes data reading and writing without adding global locks, improving the scalability of the database system and read and write transaction processing performance.
Smart Images

Figure CN120196674A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology (IT), and in particular, to a data processing method, a database cluster, and a computing device cluster. Background Art
[0002] In many data-intensive application systems (such as online transaction processing systems (OLTP), online analytical processing systems (OLAP), etc.), the read load is generally significantly higher than the write load. Therefore, separating read and write nodes and providing external read and write services respectively is a commonly used technology. Through read-write separation, the read load can be offloaded to the read node cluster, so that the write node cluster can focus on processing the write load. When the read load increases, only more read nodes need to be added, and the scale and specification of the write nodes do not need to change.
[0003] In the related art, the system architecture of read-write separation can adopt the shared-page architecture. In the shared-page architecture, data is synchronized between write and read nodes through pages. Among them, write and read nodes can organize users' structured records in the form of pages. When the read node provides an external read service, if a page miss occurs on the read node or the page exists but the version is not the latest version, the read node can pull the latest version of the page from the write node and then provide the read service. However, when the read node obtains a page from the write node, it needs to add a global page lock to the global lock manager side. Due to the existence of the global page lock, the scalability and read-write transaction processing performance of the database system under the shared-page architecture are poor. Summary of the Invention
[0004] This application provides a data processing method, a database cluster, a computing device cluster, a computer storage medium, and a computer product, which can realize data reading and writing without adding a global page lock, and improve the scalability and read-write transaction processing performance of the database system under the shared-page architecture.
[0005] In a first aspect, the present application provides a data processing method, which is applied to a read node in a database cluster. The method includes: the read node obtains a read request, and the read request includes: a snapshot read commit sequence number. Among them, the processing flow of the read request includes reading a copy of the first page; the read node reads the copy of the first page from the local or from a write node in the data cluster without adding a global lock. The copy of the first page includes a minimum commit sequence number and a maximum commit sequence number. The minimum commit sequence number is the commit sequence number of the latest write transaction committed on the write node when the read node reads the copy of the first page from the write node, and the maximum commit sequence number is the commit sequence number of the latest write transaction committed on the write node. Among them, the minimum commit sequence number is less than or equal to the maximum commit sequence number; when the snapshot read commit sequence number is less than or equal to the minimum commit sequence number, or when the snapshot read commit sequence number is greater than the minimum commit sequence number and the minimum commit sequence number is the same as the maximum commit sequence number, the read node reads the data based on the copy of the first page.
[0006] In this way, by adding a minimum commit sequence number and a maximum commit sequence number to the copy of the read page, and during the read transaction processing, comparing the snapshot read commit sequence number with the minimum commit sequence number on the copy of the read page, and comparing the minimum commit sequence number and the maximum commit sequence number on the read copy of the page, it is ensured that the data on the read copy of the page meets the requirements, so that data can be read without adding a global lock, avoiding the global page lock overhead during the read processing, and improving the processing performance of read and write transactions.
[0007] In a possible implementation, when the copy of the first page does not exist locally on the read node, or when the snapshot read commit sequence number is greater than the minimum commit sequence number and the minimum commit sequence number is different from the maximum commit sequence number, the read node reads the copy of the first page from the write node. In this way, it can be ensured that the data on the read copy of the page is the latest, thus ensuring the accuracy of the read data.
[0008] In a possible implementation, the method further includes: the read node receives a page invalidation notice broadcast by the write node. Among them, the page invalidation notice includes: the first commit sequence number of the first write transaction completed by the write node, and the first write transaction is at least used to modify the first page. Among them, the first write transaction is executed by the write node without adding a global lock; when the copy of the first page exists locally on the read node, the read node updates the maximum commit sequence number in the copy of the first page to the first commit sequence number. In this way, in the subsequent read processing flow, according to the minimum commit sequence number and the maximum commit sequence number in the copy of the page, it can be known whether the copy of the page stored locally on the read node is the latest and meets the data reading requirements.
[0009] In a possible implementation, the read request further includes: a query condition, where the query condition is a point query. The database index entry queried by the key value of the point query is the first entry, and the first entry is stored on the first page in the database index on the read node. At this time, the method further includes: in the case where there is no copy of the first page locally on the read node, the read node reads a copy of the first page from the write node; in the case where the first entry does not exist on the copy of the first page read by the read node, the read node sequentially reads copies of each page from left to right starting from the first page in the database index on the write node until a copy of the page containing the first entry is read. In this way, it can be ensured that accurate data can still be read after the leaf page in the index splits.
[0010] In a possible implementation, the read request further includes: a query condition, where the query condition is a range query. The database index entry queried by the start key value of the range query is the first entry, and the first entry is stored on the first page in the database index on the read node. The database index entry queried by the end key value of the range query is the second entry, and the second entry is stored on the second page in the database index on the read node. In the database index on the read node, the second page is located to the left of the first page. The pages in the database index on both the read node and the write node include identifiers of the left neighbor page and the right neighbor page of the page. At this time, the method further includes: in the case where there is a copy of the first page locally on the read node, the read node determines that there is no copy of the third page corresponding to the identifier of the left neighbor page included in the copy of the first page locally, or the read node determines that there is a copy of the third page locally and the minimum commit sequence number included in the copy of the third page is different from the maximum commit sequence number included in the copy of the third page; the read node reads the latest copy of the third page from the write node; in the case where the identifier of the right neighbor page included in the latest copy of the third page is not the identifier of the first page, the read node reads the latest copy of the first page from the write node, and sequentially reads copies of each page from right to left starting from the first page in the database index on the write node until a copy of the second page is read. In this way, it can be ensured that accurate data can still be read after the leaf page in the index splits.
[0011] In a possible implementation, the method further includes: in the case where there is a copy of the fourth page that needs to be read from the write node locally on the read node, the read node determines that the minimum commit sequence number included in the copy of the fourth page is the same as the maximum commit sequence number included in the copy of the fourth page; the read node uses the copy of the fourth page locally. In this way, the data transmission between the read and write nodes can be reduced, and the performance of the read operation can be improved.
[0012] In a possible implementation, the method further includes: The read node receives the upstream page invalidation notification broadcast by the write node. The upstream page invalidation notification includes: the identifier of the invalid page identified by the write node in its database index, the identifier of the neighbor page of the invalid page in its database index, and the identifiers of the pages associated with the parent node and the child node of the invalid page respectively; The read node deletes the copies of the invalid page, the neighbor page of the invalid page, and the pages associated with the parent node and the child node of the invalid page that exist locally. In this way, the read node can clean the invalid pages on it.
[0013] In a second aspect, the present application provides a data processing method applied to a write node in a database cluster. The method includes: The write node receives a page copy reading request sent by a read node in the database cluster. The page copy reading request is used to request to read a copy of a first page; The write node sends the copy of the first page to the read node, where the copy of the first page includes a minimum commit sequence number and a maximum commit sequence number. Both the minimum commit sequence number and the maximum commit sequence number are the first commit sequence numbers of the latest first write transaction committed on the write node, where the first write transaction is executed by the write node without taking a global lock.
[0014] In a possible implementation, the method further includes: When the second write transaction is processed and completed, the write node broadcasts a page invalidation notification through a first thread. The page invalidation notification includes: the second commit sequence number of the second write transaction, and the second write transaction is at least used to modify the first page. The page invalidation notification is used to instruct the read node to update the maximum commit sequence number in the copy of the first page to the second commit sequence number, where the second write transaction is executed by the write node without taking a global lock; When the write node does not receive a response related to the page invalidation notification returned by the read node, the write node submits the second write transaction through a second thread. In this way, by asynchronously executing these two operations, the write transaction processing efficiency can be improved.
[0015] In a possible implementation, the method further includes: The write node receives the snapshot read commit sequence numbers sent by all read nodes in the database cluster; The write node deletes the pages or tuples modified by the write transaction associated with the target commit sequence number, where the target commit sequence number is less than the minimum snapshot read commit sequence number. In this way, by delaying the cleaning of invalid pages or tuples, it can be ensured that the read node can read the required data, so that data reading and writing can be achieved without taking a global page lock.
[0016] In a possible implementation, the method further includes: the write node broadcasts an upstream page invalidation notice, where the upstream page invalidation notice includes: the identifier of the invalid page identified by the write node in its database index, the identifier of the neighbor page of the invalid page in its database index, and the identifiers of the pages associated with the parent node and the child node of the invalid page respectively; the write node receives a second message sent by a read node in the database, where the second message is at least used to indicate the successful deletion of the copy of the invalid page; the write node deletes the invalid page on its local side. In this way, the write node can complete the cleaning of the invalid page on it.
[0017] In a third aspect, the present application provides a database cluster, including: at least one read node and at least one write node. Among them, the read node is used to execute the method described in the first aspect or any possible implementation manner of the first aspect; the write node is used to execute the method described in the second aspect or any possible implementation manner of the second aspect.
[0018] In a fourth aspect, the present application provides a computing device cluster, including at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method described in the first aspect or any possible implementation manner of the first aspect, or executes the method described in the second aspect or any possible implementation manner of the second aspect.
[0019] In a fifth aspect, the present application provides a computer-readable storage medium, including computer program instructions. When the computer program instructions are executed by the computing device cluster, the computing device cluster executes the method described in the first aspect or any possible implementation manner of the first aspect, or executes the method described in the second aspect or any possible implementation manner of the second aspect. Among them, the computing device cluster may include one or more computing devices.
[0020] In a sixth aspect, the present application provides a computer program product containing instructions. When the instructions are run by the computing device cluster, the computing device cluster is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect, or execute the method described in the second aspect or any possible implementation manner of the second aspect. Among them, the computing device cluster may include one or more computing devices.
[0021] It can be understood that the beneficial effects of the above third aspect to the sixth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a schematic diagram of the logical architecture of a database cluster provided by an embodiment of the present application;
[0023] Figure 2 It is a schematic diagram of the processing flow of a write transaction provided by an embodiment of the present application;
[0024] Figure 3 It is a schematic diagram of the processing flow of a read transaction provided by an embodiment of the present application;
[0025] Figure 4 It is a schematic diagram of the organizational structure of a page provided by an embodiment of the present application;
[0026] Figure 5 It is a schematic diagram of index splitting provided by an embodiment of the present application;
[0027] Figure 6 It is a schematic diagram of index splitting and reading a page copy provided by an embodiment of the present application. Detailed implementation manners
[0028] In this text, the term "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In this text, the symbol " / " represents that the associated objects are in an "or" relationship, for example, A / B represents A or B.
[0029] In the description of the embodiments of the present application, the terms "first" and "second" in the specification and claims are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.
[0030] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.
[0031] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of elements refers to two or more elements.
[0032] First, some technical terms involved in the embodiments of the present application are introduced.
[0033] (1) page object
[0034] The page on the write node is called a page object, which is different from the page copy on the read node; the data of the page object is the latest.
[0035] (2) page copy
[0036] The page object read by the read node from the write node is called a page copy. In addition to the page object information, the page copy also contains two fields, pminCSN and pmaxCSN, which are used for version detection.
[0037] (3) pminCSN
[0038] When the read node creates a page copy by obtaining a page from a write node, it sets the MaxCommitedTrxCSN on that write node as the pminCSN of the page copy. If the page does not exist on the write node, it fetches the page from the file system and still sets the MaxCommitedTrxCSN on the write node as the pminCSN of the page copy.
[0039] (4) pmaxCSN
[0040] After the read node creates a page copy, the page corresponding to the page copy on the write node will continue to be modified by write transactions. The CSN of the largest committed write transaction that modifies the page is denoted as pmaxCSN. pmaxCSN is placed in the pageinvalidation message and sent by the write node to the read node. Among them, the page copy obtained by the read node contains both pminCSN and pmaxCSN, and pminCSN remains unchanged, while pmaxCSN changes following the CSN of the committed write transaction in the write node. It should be understood that during the subsequent system operation, pmaxCSN will continue to be modified by the write node through the pageinvalidation message. When a write transaction is committed on the write node, the write node will send a pageinvalidation message to all read nodes. This message contains the page modified by the write transaction and the commit sequence number (CommitedTrxCSN) of the write transaction that modified the page. After receiving the page invalidation notification, the read node modifies the pmaxCSN of the page copy to the CommitedTrxCSN of the corresponding page in the notification.
[0041] (5) MaxCommitedTrxCSN
[0042] The commit sequence number (CSN) of the latest committed write transaction on the write node.
[0043] (6) page travel
[0044] The process in which the read node traverses the pages in the BTree / Heap based on the information in the read request (such as query conditions and snapshot read CSN) and finally locates the correct tuple is called page travel. Sometimes, the operation of moving between pages during this process is also called page travel. Exemplarily, the query condition can be a point query or a range query, etc. Among them, a point query is a query operation used to precisely match a certain record according to a unique identifier or key value. A range query is also a query operation used to query data that meets the specified range conditions. The key value in the point query can be used to find an entry in the database index on the read node, and finally one or more database index entries will be located. The range query can include a start key value and an end key value. Both the start key value and the end key value in the range query can be used to find an entry in the database index on the read node.
[0045] Next, the technical solutions provided by the embodiments of the present application will be introduced.
[0046] Exemplarily, Figure 1 FIG. shows a schematic architecture diagram of a database cluster provided by an embodiment of the present application. The database cluster adopts a shared-page architecture. As Figure 1 shown, the database cluster includes: at least one coordinator node (CN) 100, at least one write node (WN) 200, at least one read node (RN) 300, and a global transaction manager (GTM) 400. The coordinator node 100 is mainly responsible for receiving operation requests from the terminal side, such as: operations of reading or writing to the database cluster, etc., and converting the operation request into a write request and transmitting the write request to the write node 200, or converting the operation request into a read request and transmitting the read request to the read node 300. The write node 200 is mainly responsible for the data writing operation in the database cluster. The read node 300 is mainly responsible for the data reading operation in the database cluster. The global transaction manager 400 is mainly responsible for managing global transactions to ensure that all participating nodes can execute transactions according to consistent rules, maintain the consistency of global data, and ensure the stability and reliability of the system, etc. Exemplarily, the database cluster can be, but is not limited to, a cloud database.
[0047] The write node 200 may be configured with a write transaction processing device 210, a page invalidation initiation device 220, a vacuum device 230, and a local buffer pool 240. The write transaction processing device 210 is mainly responsible for write transactions. After receiving a write request transmitted by the coordination node 100, it can insert or update tuples in a page to store user data in the page. In addition, the write transaction processing device 210 can organize pages into a BTree+Heap data structure and maintain multiple versions of user data on the Heap. At the same time, it can also add double links between the neighboring pages of a page and the pages corresponding to the parent and child nodes of the page respectively to construct a new type of BTree structure (see the description below for details). The page invalidation initiation device 220 is mainly responsible for notifying each read node 300 of the commit sequence number (CSN) of the latest committed write transaction and the pages modified by the committed write transaction. The vacuum device 230 is mainly responsible for performing vacuum operations, including identifying invalid tuples and invalid pages, unlinking the links between the invalid page and its neighboring page, and initiating upstream page invalidation notifications to each read node 300. After receiving the responses from the read nodes, it cleans up the invalid pages, etc. Among them, the upstream page invalidation notifications may include: the identifier of the invalid page, the identifier of the neighboring page of the invalid page, and the identifiers of the pages associated with the parent and child nodes of the invalid page respectively. The local buffer pool 240 is mainly responsible for storing pages. Of course, the write node 200 can also perform persistence operations on the pages stored in the local buffer pool 240.
[0048] In the read node 300, a read transaction processing device 310, a page copy pulling device 320, a page invalidation execution device 330, an upstream page invalidation device 340, and a local memory pool 350 can be configured. The read transaction processing device 310 is mainly responsible for read transactions. After obtaining a read request transmitted by the coordination node 100, it can read a page stored in the local memory pool 350 or pull a page from the write node 200 through the page copy pulling device 320, and perform version detection, correction, visibility judgment, etc. on the page. The page copy pulling device 320 is mainly responsible for pulling the pages required by the read node 300 from the write node 200. The page invalidation execution device 330 is mainly responsible for, after obtaining the CSN transmitted by the page invalidation initiation device 200 on the write node 200, if there is a copy of the page modified by the latest committed write transaction on the write node 200 in the local memory pool 350, updating the pmaxCSN on the page copy to the CSN of the latest committed write transaction. Among them, each page in the local memory pool 350 has a pminCSN and a pmaxCSN. The pminCSN of a page is the CSN of the latest committed write transaction on the write node 200 when the page copy pulling device 320 obtains the page from the write node 200. The pmaxCSN of a page is the CSN of the latest committed write transaction on the write node 200 after the page copy pulling device 320 obtains the page from the write node 200, where pminCSN ≤ pmaxCSN. For example, if the write node 200 updates page1 at time t1 and the CSN of this write transaction is CSN1, then when the read node 300 obtains a copy of page1 from the write node 200 at time t2 (t2 > t1) and the write node 200 does not perform an update operation on page1 between t1 and t2, then both the pminCSN and pmaxCSN in the page1 copy on the read node 300 are CSN1. If at time t3 (t3 > t2), the write node 200 updates page1 and the CSN of this write transaction is CSN2, then through page invalidation, it is notified that the pmaxCSN in the page1 copy on the read node 300 will be updated from CSN1 to CSN2. The upstream page invalidation device 340 is mainly responsible for, after receiving the upstream pages invalidation notification transmitted by the invalid data cleaning device 230 in the write node 200, performing a cleaning operation on the copies of the invalid pages stored in the local memory pool 350, the copies of the neighbor pages of the invalid page, and the copies of the pages associated with the parent node and child node of the invalid page respectively. The local memory pool 350 is mainly responsible for storing the copies of the pages pulled from the write node 200.
[0049] Next, the processing procedures for write transactions and read transactions will be introduced separately.
[0050] Exemplarily, Figure 2 FIG. shows a schematic diagram of a processing procedure for a write transaction provided by an embodiment of the present application. As Figure 2 described, the processing procedure for the write transaction may include the following steps:
[0051] S201. A write application (Write APP) sends a data manipulation language request (DMLReq), such as: insert data (insert), update data (update), delete data (delete), etc., to a coordinator node (CN) 100.
[0052] S202. In response to the DMLReq, CN100 starts a new write transaction.
[0053] S203. CN100 obtains an identifier (id) of the write transaction from a global transaction manager (GTM) 400. Among them, CN100 may first send a request to GTM400 to request the id of the write transaction, and then, GTM400 may return the id of the write transaction to CN100.
[0054] S204. CN100 sends a write request (WriteReq) to a write node 200. Among them, the WriteReq may include the data and instructions requested by the DML, such as an insert instruction, a delete instruction, an update instruction, etc. Exemplarily, CN100 may, but is not limited to, transmit the WriteReq to a write transaction processing device 210 in the write node 200.
[0055] S205. In response to the WriteReq, the write node 200 modifies at least one local page, for example, inserts data, deletes data, or updates data, etc. in the page. Exemplarily, this step may be executed by, but is not limited to, the write transaction processing device 210 in the write node 200.
[0056] S206. After processing the WriteReq, the write node 200 returns a message indicating that the WriteReq has been processed to CN100.
[0057] S207. In response to the obtained message, CN100 obtains the CSN1 of the processed write transaction from GTM400. Among them, CN100 may first send a request to GTM400 to request the CSN1 of the processed write transaction, and then, GTM400 may return the CSN1 of the write transaction to CN100.
[0058] After CN100 obtains CSN1 of the processed write transaction, it can initiate a commit command (CommitCmd) of the write transaction to the write node. CSN1 can be included in the CommitCmd. Exemplarily, CN100 can, but is not limited to, transmit the CommitCmd to the write transaction processing device 210 in the write node 200.
[0059] S209. The write node 200 broadcasts a page invalidation notice including CSN1 and the identifier of the modified page to inform the read node 300 which pages are modified and the CSN of this write transaction. Exemplarily, this step can be, but is not limited to, executed by the page invalidation initiation device 220 in the write node 200. Among them, the page invalidation initiation device 220 can initiate the broadcast after receiving the instruction issued by the write transaction processing device 210, and after initiating the broadcast, inform the write transaction processing device 210 that the broadcast has been completed.
[0060] S210. The write node 200 returns a notice of successful commit (CommitOK) of the write transaction to CN100. The CommitOK at least includes PICSN (i.e., the global commit sequence number of the latest page invalidation notice successful transaction). Among them, PICSN can be, but is not limited to, the CSN of the write transaction completed by the write node last time. Exemplarily, CSN1 can also be included in the CommitOK. Exemplarily, this step can be, but is not limited to, executed by the write transaction processing device 210 in the write node 200. Among them, after the page invalidation initiation device 220 successfully initiates the broadcast, it can return its PICSN to the write transaction processing device 210. Exemplarily, S209 and S210 can be executed by different threads on the write node 200, and the write node 200 does not need to wait for the response related to the page invalidation notice returned by the read node when executing S210. Exemplarily, S210 can be understood as the write node 200 committing the write transaction.
[0061] S211. CN100 sends a message including PICSN and CSN1 to GTM400, and obtains the confirmation message of receiving the message returned by GTM400.
[0062] S212. CN100 returns a message of successful execution of the data manipulation language (DML ExecOK) to the Write APP.
[0063] S213. After the read node 300 obtains the page invalidation notification, if there is a copy of the modified page on it, it updates the pmaxCSN in the page copy to CSN1 and keeps the pminCSN unchanged. Exemplarily, this step can be but is not limited to being executed by the page invalidation execution device 330 in the read node 300.
[0064] S214. The read node 300 returns a response message indicating successful page invalidation execution to the write node 200.
[0065] S215. The write node 200 sets the PICSN on it to CSN1.
[0066] In this way, the processing of the write transaction is completed. It can be seen from the above process that after the write node receives the WriteReq, it does not need to add a global page lock to the Global Lock Manager when modifying the local page, that is, a global lock can be not added during the execution of the write transaction. Therefore, this write transaction processing process can avoid the global page lock overhead during the write processing.
[0067] Exemplarily, Figure 3 shows a schematic diagram of a processing process of a read transaction provided by an embodiment of the present application. As Figure 3 described, the processing process of this read transaction can include the following steps:
[0068] S301. The read application (Read APP) sends a data query request (SelectReq) to the CN100.
[0069] S302. After the CN100 receives the SelectReq, it obtains the snapshot read CSN (ReqCSN) that meets the SelectReq from the GTM400. Among them, the CN100 can first send a request to the GTM400 to request the ReqCSN that meets the SelectReq, and then, the GTM400 can return the ReqCSN that meets the SelectReq to the CN100. Exemplarily, the GTM400 can collect all the CICSNs (i.e., the CSNs of the write transactions) and PICSNs reported by all the write nodes, sort the write nodes in ascending order according to their PICSNs to form a write node list; then, starting from the write node with the smallest PICSN in the write node list, if the PICSN of this node = CICSN and this PICSN is less than and not equal to the PICSNs of other nodes, then this node is ignored and the PICSN and CICSN of the next node are judged; otherwise, the PICSN of this node is taken as the ReqCSN.
[0070] S303. After obtaining ReqCSN, CN100 sends a read request (ReadReq) containing ReqCSN to the read node 300. The ReadReq may also contain query conditions, such as point query or range query, etc.
[0071] S304. After obtaining the ReadReq, the read node 300 determines whether there is a page copy containing the data required for the query in its local memory pool. If it exists, S306 is executed; otherwise, S305 is executed. Exemplarily, this step can be, but is not limited to, executed by the read transaction processing device 310 in the read node 300.
[0072] S305. The read node 300 pulls a page copy from the write node 200. The pulled page copy contains pminCSN and pmaxCSN. At this time, both pminCSN and pmaxCSN are the CSNs of the latest committed write transactions on the write node 200. Exemplarily, this step can be, but is not limited to, executed by the page copy pulling device 320 in the read node 300.
[0073] S306. The read node 300 determines whether ReqCSN is less than or equal to the pminCSN of the page copy. If so, it means that the version of this page copy is newer than that required by ReqCSN, that is, the versions of all tuples in the page copy meet the requirements of this snapshot read, so S308 can be executed; otherwise, S307 is executed. Exemplarily, this step can be, but is not limited to, executed by the read transaction processing device 310 in the read node 300.
[0074] S307. The read node 300 determines whether the pminCSN and pmaxCSN of the page copy are the same. If so, it means that the page object on the page copy has not been updated by other transactions since the page copy was created, which means that the content of this page copy is consistent with the page object in the write node 200. Therefore, this page copy also meets the requirements of snapshot read, so S308 can be executed; otherwise, return to execute S305. Exemplarily, this step can be, but is not limited to, executed by the read transaction processing device 310 in the read node 300.
[0075] S308. The read node 300 reads the data queried by the SelectReq based on the page copy. Exemplarily, this step can be, but is not limited to, executed by the read transaction processing device 310 in the read node 300.
[0076] S309. After the read node 300 reads the required data, it returns a read operation response (ReadRsp) of the Read result to the CN100. Among them, the ReadRsp may but is not limited to include the data queried by the SelectReq.
[0077] S310. The CN100 returns the ReadRsp to the Read APP.
[0078] In this way, the processing of the read transaction is completed. By adding pminCSN and pmaxCSN to the page copy, and during the processing of the read transaction, comparing the snapshot read CSN with the pminCSN on the page copy, and comparing the pminCSN and pmaxCSN on the page copy, it is ensured that the data on the page copy meets the requirements, so that data can be read without adding a global lock, avoiding the global page lock overhead during the read processing, and improving the processing performance of the read transaction.
[0079] In some embodiments, the organizational structure of the page can be BTree / B+Tree+Heap. BTree (or B+Tree) is used to implement the index, and index entries can be stored therein; Heap (heap) is used to store actual data records, and table tuples can be stored therein. In the Heap, a table tuple may have multiple versions, and pointers are used to point between the old and new versions of the table tuple to form a multi-version linked list. The head of the linked list is the oldest version of the tuple, and the tail is the latest version. The index entry of the Btree / B+Teee points to the oldest version of the table tuple. For example, as Figure 4 shown, the identifier of the page can be stored in the B+Tree, and the data (i.e., tuple) on the page is stored in the Heap. In Figure 4 , the tuple in HP1 in the Heap is the oldest version, and pointers are used to point between different versions of the tuple to form a linked list; the index entry of P11 in the B+Tree points to HP1 in the Heap. When a read transaction needs to find a table tuple according to the index, first locate a certain index entry of the BTree according to the search information, then locate the oldest version of the table tuple according to the index entry, and then search backward along the version chain. Through version visibility judgment, find the version that the read transaction can see. For example, continue to refer to Figure 4, in a read transaction, if the version of the tuple read by the snapshot read CSN is in HP2, when the read node reads data, it only needs to find HP2 and read the tuple in HP2, without having to read the data after HP2.
[0080] Under the BTree / B+Tree index structure, as new entries are inserted into the index, the BTree / B+Tree will continuously split, grow taller and fatter, and at the same time, the entries will continuously move to the right (as the left page splits, they are continuously moved to the new page on the right). For example, as Figure 5 shown, in Figure 5 (A), entry11 is stored in p13. As new entries are inserted, the Btree splits into the Figure 5 index structure shown in (B). In Figure 5 (B), entry11 is stored in p23. Continuing to refer to Figure 5 , in this case, if the read request obtained by the read node 300 contains a point query, and the database index entry queried by the key value of the point query is entry11. At this time, if the snapshot read CSN in the read request = CSN1, and the version of the root page of the database index on the read node 300 is also CSN1, then the point query process will consider that the version of the root page meets the requirements, and further consider that entry11 is located on P13 according to the content of the root page. At this time, when the read node 300 does not locally store a copy of p13, the read node 300 can read a copy of p13 from the write node. Then, the write node can provide the latest version of p13 on it (i.e., Figure 5Send the p13) under the index shown in (B) to the read node 300. Since the index has split, the read node 300 will not find entry11 in the latest copy of p13, which will result in an error. To avoid this situation, in this embodiment, the read node 300 can read the identifier of the right neighbor page of p13 (i.e., p21) contained in the copy of p13 from the latest copy of p13. Then, the read node 300 can read the copy of p21 from the write node and can detect that entry11 is not stored in the copy of p21. Then, the read node 300 can read the identifier of the right neighbor page of p21 (i.e., p22) contained in the copy of p21 from the copy of p21. Then, the read node 300 can read the copy of p22 from the write node and can detect that entry11 is not stored in the copy of p22. Then, the read node 300 can read the identifier of the right neighbor page of p22 (i.e., p23) contained in the copy of p22 from the copy of p22. Finally, the read node 300 can read the copy of p23 from the write node and can detect that entry11 is stored in the copy of p23. In this way, the copy of the page where the required data is located is read, thus avoiding the occurrence of read errors. It should be understood that each time the read node 300 reads a copy of a page, it can perform version detection on the copy of the page, that is, execute S306 and S307 in the foregoing Figure 3 to ensure that the copy of the page read meets the requirements. In addition, when the read node 300 needs to read a copy of a page, it can first detect whether a copy of the page exists locally. If a copy of the page exists locally, the read node 300 can first determine whether the pminCSN contained in the copy of the page is the same as the pmaxCSN. If they are the same, it means that the content of the page copy is consistent with the page object in the write node. At this time, the read node 300 can use the copy of the page without having to read the copy of the page from the write node again. For example, continue to refer to Figure 5 (B). When the read node 300 learns from the latest read copy of p21 that the left neighbor page of p21 is p22, the read node 300 can first detect whether a copy of p22 exists locally. If it exists and pminCSN = pmaxCSN in the copy of p22, the read node 300 can directly use the local copy of p22 without having to read the copy of p22 from the write node again.
[0081] That is to say, if the query condition included in the read request obtained by the read node 300 is point query, and the database index entry queried by the key value of the point query is entry11, and the snapshot read CSN = CSN1, and entry11 is stored on p13 in the database index on the read node 300. At this time, if the read node 300 does not locally store a copy of p13, the read node 300 can read a copy of p13 from the write node. After the read node 300 obtains the latest copy of p13, when it detects that entry11 does not exist in the copy of p13, the read node 300 can sequentially read the copies of each page from left to right starting from p13 in the database index on the write node until it reads the copy of the page containing entry11. In some embodiments, if the read node 300 locally stores a copy of p13, the read node 300 can read data from the copy of p13 stored locally by it.
[0082] In addition, in this embodiment, in the BTree / B+Tree index structure, pointers (double links) pointing to each other can also be added between the neighbor pages of the page and the pages corresponding to the parent and child nodes respectively. For example, as Figure 6 shown in (A) of, p12 can respectively have pointers pointing to p11, p13, and p1. For example, p12 can contain the identifiers of p11, the identifier of p13, and the identifier of p1, etc. Among them, p11 is the left neighbor page of p12, p13 is the right neighbor page of p12, and p1 is the page corresponding to the parent node of p12. That is to say, the page in the database index can include the identifiers of the left neighbor page and the right neighbor page of the page. Of course, it can also include the identifiers of the pages associated with the parent node and the child node of the page respectively. In this way, after the read node 300 reads a copy of a page, it can, through the information contained in the copy of this page, know the identifiers of the neighbor pages of this page and the pages corresponding to the parent and child nodes respectively, so as to facilitate reading the copies of these pages.
[0083] Further, continue to refer to Figure 6 As Figure 6 shown, on the write node, as the write transaction continues, the index will also continuously split and expand, etc. After submitting the first write transaction, CSN = 1, and the index structure is as Figure 6 shown in (A) of. After submitting the second write transaction, CSN = 2, and the index structure is as Figure 6 shown in (B); in Figure 6 of (B), entry61 is stored on p21, and entry62 is stored on p11. After submitting the third write transaction, CSN = 3, and the index structure is as Figure 6 shown in (C); inFigure 6 In (C) thereof, entry61 is stored at p21 and entry62 is stored at p11. After submitting the fourth write transaction, CSN = 4, and the index structure is as shown in Figure 6 (D); in Figure 6 (D) thereof, entry61 is stored at p21 and entry62 is stored at p11. In this case, if the read request obtained by the read node 300 contains a range query condition, and the database index entry queried by the start key value of the range query is entry61, and the database index entry queried by the end key value of the range query is entry62. At this time, if the snapshot read CSN in the read request = CSN2, that is, ReqCSN = 2, and the version of the root page of the database index on the read node 300 is also CSN2, then the range query process will consider that the version of the root page meets the requirements, and then consider that entry61 is located on P21 and entry62 is located on P11 according to the content of the root page. Among them, this range query can be understood as a reverse scan, that is, scanning sequentially from right to left on the leaf page in the index.. Since the commit sequence number of the latest write transaction on the write node has become CSN4, at this time, if data is continued to be read in the index under the version of CSN = 2, an error will occur. To avoid this situation, in this embodiment, when there is a copy of p21 locally on the read node 300, the read node 300 can detect whether there is a copy of p13 corresponding to the identifier of the left neighbor page included in the copy of p21 locally. If there is no copy of p13 locally on the read node 300, the read node 300 can read the latest copy of p13 from the write node. At this time, pminCSN = pmaxCSN = 4 included in the read copy of p13. In addition, when there is a copy of p13 locally on the read node 300, but pminCSN ≠ pmaxCSN in the copy of p13, it indicates that the write node has executed a new write transaction. At this time, the data in p13 may change, or p13 may split or merge, etc. At this time, if the old version of the copy of p13 under CSN is continued to be used to read data, a read error may occur. Therefore, to avoid this situation, the read node 300 can read the latest copy of p13 from the write node. At this time, pminCSN = pmaxCSN = 4 included in the read copy of p13.
[0084] Further, after the read node 300 reads the latest copy of p13, it can detect whether the identifier of the right neighbor page included in the copy of p13 is the identifier of p21. If not, there is no longer a pointer pointing to each other between p21 under CSN = 2 and p13 under CSN = 4, that is, the double link between the two has failed, and there is a new page between the two at this time. At this time, the read node 300 can read the latest copy of p21 from the write node, and pminCSN = pmaxCSN = 4 in the latest copy of p21. Then, the read node 300 can learn from the latest copy of p21 that the left neighbor page of the latest copy of p21 is p22, and can read the copy of p22 from the write node. Next, the read node 300 can learn from the copy of p22 that the left neighbor page of the copy of p22 is p23, and can read the copy of p23 from the write node. Next, the read node 300 can learn from the copy of p23 that the left neighbor page of the copy of p23 is p13, and can read the copy of p13 from the write node. Next, the read node 300 can learn from the copy of p13 that the left neighbor page of the copy of p13 is p12, and can read the copy of p12 from the write node. Finally, the read node 300 can learn from the copy of p12 that the left neighbor page of the copy of p12 is p11, and can read the copy of p11 from the write node. In this way, the read node 300 reads the copies of all the pages where the required data is located, thus avoiding the occurrence of read errors. It should be understood that each time the read node 300 reads a copy of a page, it can perform version detection on the copy of the page, that is, execute S306 and S307 in the foregoing Figure 3 to ensure that the copy of the page read meets the requirements. In addition, when the read node 300 needs to read a copy of a page, it can first detect whether there is a copy of the page locally. If there is a copy of the page locally, the read node 300 can first determine whether pminCSN included in the copy of the page is the same as pmaxCSN. If they are the same, it means that the content of the page copy is consistent with the page object in the write node. At this time, the read node 300 can use the copy of the page without having to read the copy of the page from the write node again. For example, continue to refer to Figure 6 (E). When the read node 300 learns from the latest read copy of p21 that the left neighbor page of p21 is p22, the read node 300 can first detect whether there is a copy of p22 locally. If there is, and pminCSN = pmaxCSN in the copy of p22, the read node 300 can directly use the local copy of p22 without having to read the copy of p22 from the write node.
[0085] That is to say, continue to refer toFigure 6 , if the query condition included in the read request obtained by the read node 300 is a range query, and the database index entry queried by the start key value of the range query is entry61, the database index entry queried by the end key value of the range query is entry62, and entry61 is stored on p21 in the database index on the read node 300, and entry62 is stored on p11, and the snapshot read CSN = CSN2. At this time, when a copy of p21 is locally stored on the read node 300, if the read node 300 determines that there is no copy of p13 corresponding to the identifier of the left neighbor page included in the copy of p21 locally, or the read node determines that there is a copy of p13 locally and the minimum commit sequence number included in the copy of p13 is different from the maximum commit sequence number included in the copy of p13, then the read node 300 can read the latest copy of p13 from the write node. Then, if the identifier of the right neighbor page included in the latest copy of p13 read by the read node 300 is not the identifier of p21, then the read node 300 can read the latest copy of p21 from the write node, and, in the database index on the write node, read the copies of each page in sequence from right to left starting from p21 until the copy of p11 is read.
[0086] In addition, when the copy of the page read by the read node 300 is a copy of a page of a database table, the process described in the above point query or range query may not be used, but the data may be directly read from the read copy of the page.
[0087] In some embodiments, when a write operation is performed in the database, it sometimes causes fragmentation of the storage space and an increase in unused space. This may be because when the database system performs an update or delete operation, it does not immediately release the occupied storage space, but marks this unused space as recyclable. To improve the utilization rate of the storage space on the write node, a vacuum operation may be performed on the write node to clean up this unused space on the write node and return it to the database system for subsequent write operations. In this embodiment, by delaying the cleaning of invalid pages or tuples on the write node, it can be ensured that the read node can read the required data. Specifically, the write node can poll all read nodes to obtain the minimum snapshot read CSN (MinReqCSN) on each read node, and then obtain the minimum value among these MinReqCSNs, which is MinMinReqCSN. When performing the vacuum operation, the write node only performs the vacuum operation on the tuples / pages inserted by the early write transactions corresponding to the CSNs less than this MinMinReqCSN and later deleted by the later write transactions, that is, deletes the tuples or pages modified by these write transactions, etc.
[0088] In addition, when the write node performs a vacuum operation, it can first mark invalid tuples and invalid pages on it, and unlink the links between the invalid page and its neighboring pages, as well as the pages associated with the parent and child nodes. Then, the write node can initiate an upstream pages invalidation notification to the read node to notify the read node which tuples or pages on the read node need to be cleaned. Exemplarily, the upstream pages invalidation notification may include: the identifier of the invalid page marked by the write node, the identifier of the neighboring page of the invalid page, and the identifiers of the pages associated with the parent and child nodes of the invalid page respectively. After receiving the upstream pages invalidation notification, the read node can delete the copies of the invalid page, the copies of the neighboring pages of the invalid page, and the copies of the pages associated with the parent and child nodes of the invalid page that exist locally. After the read node completes the cleaning, it can send a message to the write node at least indicating the successful deletion of the copies of the invalid page. After receiving the replies from all read nodes, the write node can completely clean the invalid tuples or invalid pages on it. This way of cleaning tuples can avoid the situation where, in a Btree / B+Tree, although no new pages are added or deleted, the records on a page are moved. When walking from the old version page A to the new version page B on the read node, some records in the new version page B are missing (moved to the new version A), resulting in the error of missing data reading. And this way of cleaning pages can avoid the situation where subsequent data in a page cannot be read due to the deletion of a certain page. After the write node cleans a certain page, it can delete the corresponding node of the page in the index structure.
[0089] It can be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, the above-described embodiments can be combined according to actual situations, and the combined solutions are still within the protection scope of the present application.
[0090] Based on the method in the above embodiments, an embodiment of the present application provides a computing device cluster, including at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method executed by the aforementioned read node, or executes the method executed by the aforementioned write node.
[0091] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a computing device cluster including at least one computing device, the computing device cluster is caused to execute the method described in the above embodiments. Exemplarily, the computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc.
[0092] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product containing instructions. When the computer program product runs on a computing device cluster including at least one computing device, the computing device cluster is caused to execute the method in the above embodiments.
[0093] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0094] The method steps in the embodiments of this application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0095] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)), etc.
[0096] It can be understood that the various numerical numbers involved in the embodiments of this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, characterized in that, Applied to a read node in a database cluster, the method includes: The read node obtains a read request, which includes: a snapshot read commit sequence number. Among them, the processing flow of the read request includes reading a copy of the first page; The read node reads the copy of the first page from local or from a write node in the data cluster without adding a global lock. The copy of the first page includes a minimum commit sequence number and a maximum commit sequence number. The minimum commit sequence number is the commit sequence number of the latest committed write transaction on the write node when the read node reads the copy of the first page from the write node. The maximum commit sequence number is the commit sequence number of the latest committed write transaction on the write node. Among them, the minimum commit sequence number is less than or equal to the maximum commit sequence number; When the snapshot read commit sequence number is less than or equal to the minimum commit sequence number, or when the snapshot read commit sequence number is greater than the minimum commit sequence number and the minimum commit sequence number is the same as the maximum commit sequence number, the read node reads data based on the copy of the first page.
2. The method according to claim 1, wherein: When the copy of the first page does not exist locally on the read node, or when the snapshot read commit sequence number is greater than the minimum commit sequence number and the minimum commit sequence number is different from the maximum commit sequence number, the read node reads the copy of the first page from the write node.
3. The method according to claim 1 or 2, characterized in that, It further includes: The read node receives a page invalidation notice broadcast by the write node. Among them, the page invalidation notice includes: a first commit sequence number of a first write transaction completed by the write node. The first write transaction is at least used to modify the first page. Among them, the first write transaction is executed by the write node without adding a global lock; When the copy of the first page exists locally on the read node, the read node updates the maximum commit sequence number in the copy of the first page to the first commit sequence number.
4. The method according to any one of claims 1-3, characterized in that, The read request further includes: a query condition, and the query condition is a point query. The database index item queried by the key value of the point query is the first entry, and the first page in the database index on the read node stores the first entry; The method further includes: When the copy of the first page does not exist locally on the read node, the read node reads the copy of the first page from the write node; When the first entry does not exist on the copy of the first page read by the read node, the read node sequentially reads copies of each page from left to right starting from the first page in the database index on the write node until a copy of the page containing the first entry is read.
5. The method according to any one of claims 1 to 3, characterized in that, The read request further includes: a query condition, where the query condition is a range query. The database index entry queried by the start key value of the range query is the first entry. The first page in the database index on the read node stores the database index of the first entry. The database index entry queried by the end key value of the range query is the second entry. The second page in the database index on the read node stores the second entry. In the database index on the read node, the second page is located to the left of the first page. The pages in the database index on the read node and the database index on the write node both include the identifiers of the left neighbor page and the right neighbor page of the page. The method further includes: When there is a copy of the first page locally at the read node, the read node determines that there is no copy of the third page corresponding to the identifier of the left neighbor page included in the copy of the first page locally, or the read node determines that there is a copy of the third page locally, and the minimum commit sequence number included in the copy of the third page is different from the maximum commit sequence number included in the copy of the third page. The read node reads the latest copy of the third page from the write node. When the identifier of the right neighbor page included in the latest copy of the third page is not the identifier of the first page, the read node reads the latest copy of the first page from the write node, and sequentially reads the copies of each page from right to left starting from the first page in the database index on the write node until the copy of the second page is read.
6. The method according to claim 4 or 5, characterized in that The method further includes: When there is a copy of the fourth page that needs to be read from the write node locally at the read node, the read node determines that the minimum commit sequence number included in the copy of the fourth page is the same as the maximum commit sequence number included in the copy of the fourth page. The read node uses the copy of the fourth page locally.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: The read node receives an upstream page invalidation notice broadcast by the write node. The upstream page invalidation notice includes: the identifier of the invalid page identified by the write node in its database index, the identifiers of the neighbor pages of the invalid page in its database index, and the identifiers of the pages associated with the parent node and the child node of the invalid page respectively. The read node deletes the copy of the invalid page, the copy of the neighbor page of the invalid page, and the copies of the pages associated with the parent node and the child node of the invalid page that exist locally.
8. A data processing method, characterized in that, Applied to a write node in a database cluster, the method includes: The write node receives a page copy read request sent by a read node in the database cluster. The page copy read request is used to request to read a copy of the first page. The write node sends the copy of the first page to the read node, where the copy of the first page includes a minimum commit sequence number and a maximum commit sequence number. Both the minimum commit sequence number and the maximum commit sequence number are the first commit sequence numbers of the latest first write transaction committed on the write node. The first write transaction is executed by the write node without taking a global lock.
9. The method according to claim 8, wherein The method further includes: When the write node completes the second write transaction processing, it broadcasts a page invalidation notice through a first thread. The page invalidation notice includes: a second commit sequence number of the second write transaction, where the second write transaction is at least used to modify the first page, and the page invalidation notice is used to instruct the read node to update the maximum commit sequence number in the copy of the first page to the second commit sequence number. Here, the second write transaction is executed by the write node without taking a global lock. When the write node does not receive a response related to the page invalidation notice returned by the read node, it submits the second write transaction through a second thread.
10. The method according to claim 8 or 9, characterized in that The method further includes: The write node receives snapshot read commit sequence numbers sent by all read nodes in the database cluster. The write node deletes a page or tuple modified by a write transaction associated with a target commit sequence number, where the target commit sequence number is less than the smallest snapshot read commit sequence number.
11. According to the method described in any one of claims 8-10, characterized in that, The method further includes: The write node broadcasts an upstream page invalidation notice. The upstream page invalidation notice includes: an identifier of an invalid page identified by the write node in its upper database index, identifiers of neighbor pages of the invalid page in the upper database index, and identifiers of pages associated with the parent node and child node of the invalid page respectively. The write node receives a second message sent by a read node in the database, where the second message is at least used to indicate successful deletion of a copy of the invalid page. The write node deletes the invalid page locally.
12. A database cluster, characterized in that, Includes: At least one read node for executing the method according to any one of claims 1-7. At least one write node for executing the method according to any one of claims 8-11.
13. A cluster of computing devices, characterized in that, Includes at least one computing device, and each computing device includes a processor and a memory. The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1-7, or the method according to any one of claims 8-11.
14. A computer-readable storage medium stores a computer program. When the computer program runs on a computing device cluster including at least one computing device, it causes the computing device cluster to execute the method according to any one of claims 1-7, or the method according to any one of claims 8-11.
15. A computer program product, characterized in that, When the computer program product runs on a computing device cluster including at least one computing device, it causes the computing device cluster to execute the method according to any one of claims 1-7, or the method according to any one of claims 8-11.