Data processing method, device, electronic device and storage medium
By introducing a security update interval mechanism into the distributed database system, the problem of data inconsistency under network isolation is solved, ensuring that only the DDL master nodes perform data updates, and the database is highly consistent and available.
Patent Information
- Application Number
- CN202510789651.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-12
AI Technical Summary
In the prior art, when a distributed database system undergoes schema changes under network isolation, it is easy to cause data inconsistency. The traditional distributed locking mechanism leads to long-term locking or unusable, and the online change method cannot effectively solve the data inconsistency caused by network isolation.
By introducing a security update interval mechanism in the distributed database system, the first node stores the metadata of the data object and compares the comparison information of the second node. If it is inconsistent, the second node is instructed to go offline and refuses it to perform DDL operations, ensuring that only the DDL master node performs data updates and avoids the occurrence of bi-group split brain situation.
Under network isolation, the user data consistency and availability of the database are guaranteed to the greatest extent, avoid data inconsistency and business interruption, and improve the reliability of the distributed database system.
Smart Images

Figure CN120316116B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of databases, and can be used in the field of distributed databases, for example. Specifically, the present disclosure relates to a data processing method, device, electronic device, and storage medium. Background Art
[0002] Distributed database systems store data objects. To adapt to data changes, optimize query performance, or meet new business requirements, structural changes (i.e., schema changes) are necessary. Traditionally, schema changes are implemented using distributed locks, which can cause the database to be locked for extended periods or become unusable.
[0003] Therefore, the online change method was proposed. Online change refers to executing schema changes without blocking read and write operations. The online change method still stores data inconsistencies caused by network isolation. Summary of the Invention
[0004] The embodiments of the present disclosure provide a data processing method, apparatus, electronic device, and storage medium that can solve the problem of data inconsistency caused by network isolation during online changes in the prior art. The technical solutions provided by the present disclosure are as follows:
[0005] According to one aspect of an embodiment of the present disclosure, there is provided a data processing method, which is applied to a distributed database system, the distributed database system comprising a first node and a second node, the method being executed by the first node;
[0006] The method includes:
[0007] Obtaining a first operation submitted by a second node for metadata of a data object; the first operation including comparison information of the metadata of the data object obtained by the second node;
[0008] If the comparison information is inconsistent with the to-be-matched information of the metadata of the data object stored by the first node, instructing the second node to go offline;
[0009] Among them, the first node stores the security update interval of the data object; the security update interval is used to record the status advancement information of the metadata of the data object and the online time of the data definition language DDL master node in the distributed database system; the security update interval includes at least two timestamps; the information to be matched is the last timestamp in the security update interval; the last timestamp includes the online time of the latest online DDL master node or the start time of the latest status advancement.
[0010] Optionally, the method further includes:
[0011] Obtaining a commit operation of a write operation on a data object submitted by the second node;
[0012] Determining compatible state time information of the data object based on a safe update interval in the latest metadata of the data object; the compatible state time information is determined from a start time of at least one compatible state determined based on the compatibility requirement;
[0013] If the transaction start time of the transaction targeted by the commit operation is not less than the compatible state time information, the commit operation is performed.
[0014] Optionally, the safe update interval of the data object is updated based on the following method:
[0015] If a status advancement start notification of the DDL operation on the data object is detected, a first update operation is performed on the safe update interval;
[0016] The first update operation includes:
[0017] The starting moment of the current state advancement is taken as the first moment;
[0018] Determining a first security update interval corresponding to the current first update operation; the first security update interval includes at least two timestamps;
[0019] The last timestamp in the first security update interval is set as the first moment, and the values of other timestamps in the first security update interval are determined based on the start moments of the respective states between the current state.
[0020] Optionally, the safe update interval of the data object is updated based on the following method:
[0021] If a DDL master node switching operation in the distributed database system is detected, performing a second update operation on the safe update interval;
[0022] The second update operation includes:
[0023] The online time of the DDL master node after the switch in the current DDL master node switch operation is used as the second time;
[0024] Add a timestamp after the last timestamp of the security update interval to be updated, and use the security update interval after the timestamp as the second security update interval corresponding to the current second update operation;
[0025] The last timestamp in the second security update interval is set as the second moment.
[0026] Optionally, the safe update interval of the data object is initialized based on the following method:
[0027] In response to a creation operation on the data object, a start time of a first state advancement in the creation operation is used as a value of each timestamp in a security update interval of the data object.
[0028] Optionally, the distributed database system includes at least two associated nodes associated with the data object; the at least two associated nodes include the first node;
[0029] For each associated node, the associated node stores a mapping structure of the data object, wherein the mapping structure uses the identification information of the data object as a key and the safe update interval of the data object as a value.
[0030] Optionally, the metadata of the data object includes a security update interval of the data object; the at least two associated nodes include a central associated node;
[0031] For other associated nodes except the central associated node, the other associated nodes have registered to monitor the identification information of the data object;
[0032] The safe update interval of the data object in the other associated nodes is updated based on the following method:
[0033] If it is monitored that the security update interval of the data object is updated, the stored security update interval of the data object is replaced with the updated security update interval.
[0034] According to another aspect of an embodiment of the present disclosure, there is provided a data processing device, the device comprising:
[0035] an acquisition module, configured to acquire a first operation submitted by a second node for metadata of a data object; the first operation including comparison information of the metadata of the data object acquired by the second node;
[0036] a comparison module, configured to instruct the second node to go offline if the comparison information is inconsistent with the to-be-matched information of the metadata of the data object stored by the first node;
[0037] Among them, the first node stores the security update interval of the data object; the security update interval is used to record the status advancement information of the metadata of the data object and the online time of the data definition language DDL master node in the distributed database system; the security update interval includes at least two timestamps; the information to be matched is the last timestamp in the security update interval; the last timestamp includes the online time of the latest online DDL master node or the start time of the latest status advancement.
[0038] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above-mentioned data processing methods when executing the program.
[0039] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data processing methods are implemented.
[0040] According to one aspect of an embodiment of the present disclosure, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, the steps of any of the above data processing methods are implemented.
[0041] The technical solutions provided by the embodiments of the present disclosure have the following beneficial effects:
[0042] By comparing the metadata to be matched of the data object stored in the first node with the comparison information obtained by the second node, for the distributed database system, in the event of network isolation, an active and effective judgment is made by using the write to the isolated node, thereby ensuring the consistency and availability of user data in the database to the greatest extent in the event of network isolation.
[0043] When the first node finds inconsistency with the second node after comparing them, it determines that the second node is not the DDL master node and instructs the second node to automatically go offline, thereby refusing to write the execution result of the first operation on the data object performed by the second node. This avoids data inconsistency caused by two nodes executing DDL operations at the same time (i.e., a double-group split-brain situation), and ensures data consistency under network isolation. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments of the present disclosure.
[0045] Figure 1 A flowchart of a data processing method provided in an embodiment of the present disclosure;
[0046] Figure 2 A schematic diagram of the system architecture of a distributed database system with storage and computing separation provided in an embodiment of the present disclosure;
[0047] Figure 3 A schematic diagram of a creation process is provided for an embodiment of the present disclosure;
[0048] Figure 4 A flowchart of a modification operation provided in an embodiment of the present disclosure;
[0049] Figure 5 A flowchart of a deletion operation provided in an embodiment of the present disclosure;
[0050] Figure 6 A schematic diagram of an execution process of a write request provided in an embodiment of the present disclosure;
[0051] Figure 7 A schematic structural diagram of a data processing device provided in an embodiment of the present disclosure;
[0052] Figure 8 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0053] The following describes embodiments of the present disclosure in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions of the embodiments of the present disclosure.
[0054] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a," "an," "said," and "the" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the features, information, data, steps, operations, elements, and / or components presented, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" or "A, B" indicates implementation as "A," or implementation as "B," or implementation as "A and B."
[0055] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.
[0056] Distributed database systems are a more complex type of distributed system. According to the CAP theorem for distributed systems, these systems have three properties: consistency, availability, and partition tolerance. Consistency means that every read operation returns the result of the most recent write operation (data remains consistent across all nodes). Availability means that every request receives a non-error response (though the data is not guaranteed to be the latest). Partition tolerance means that the system can continue to operate even in the event of a network partition.
[0057] The core of the CAP theorem is that it's impossible to simultaneously achieve consistency, availability, and partition tolerance in a distributed system. A distributed system can only guarantee two of these properties, and network isolation is almost inevitable in a distributed system. Data consistency is essential for distributed database systems, while availability is the core competitive advantage of these systems. Therefore, ensuring consistency or availability in all situations becomes a core capability of the services they provide.
[0058] Data objects stored in distributed database systems require structural changes (i.e., schema changes) to adapt to data changes, optimize query performance, or meet new business requirements. Traditionally, schema changes are implemented using distributed locks. This involves applying traditional database user schema changes directly to distributed database nodes, ensuring data consistency through the use of distributed locks.
[0059] The disadvantages of schema changes based on distributed locks include: 1) the implementation of a distributed lock mechanism is relatively complex; 2) schema changes that implement the lock mechanism can impact user services; 3) when a network partition occurs, there can be significant side effects. For example, the node executing the schema change will be stuck waiting for confirmation from the isolated node. Users on other nodes cannot perform normal business activities due to the lock during the loading process. Even with a timeout mechanism, business interruptions can occur during the waiting period.
[0060] To address the shortcomings of schema changes based on distributed locks, an online change solution was proposed. This approach involves making schema changes on one node, with other nodes following the same path. This allows schema changes to be executed without blocking read and write operations. While this approach avoids the long-term database lockout or unavailability caused by distributed locks, it still presents data inconsistencies due to network isolation.
[0061] The data processing method, device, electronic device and storage medium provided in the present disclosure are intended to solve the above technical problems in the prior art.
[0062] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0063] Figure 1 A flow chart of a data processing method provided by an embodiment of the present disclosure is applied to a first node in a distributed database system, such as Figure 1 As shown, the method includes:
[0064] Step S110 : obtaining a first operation submitted by the second node for metadata of a target data object; the first operation includes comparison information of the metadata of the target data object obtained by the second node.
[0065] Specifically, a distributed database system may include multiple nodes. The distributed database system may adopt a storage and computing integrated structure, that is, a node can serve as both a storage node and a computing node; or it may adopt a storage and computing separation structure, that is, the storage node and the computing node are respectively two independent nodes. This is not limited in the embodiments of the present disclosure.
[0066] The data processing method provided by the embodiment of the present disclosure is applied to a first node in a distributed database system. The first node can be used to store data, and the second node can be used to process the data.
[0067] It should be noted that when the distributed database system adopts a storage and computing separation structure, the first node can be a storage node and the second node can be a computing node.
[0068] Under preset conditions, the second node can submit a first operation to the first node, where the first operation can be a state-advancing write operation in a DDL (Data Definition Language) operation. DDL operations are a language used to define and manage data structures in a database. DDL operations include operations such as creating a table, deleting a table, and modifying a table structure.
[0069] When executing a DDL operation (such as adding an index or modifying a table structure), a state machine manages the process. Only the DDL master node (also known as the DDL Owner) can sequentially advance the state machine. The DDL master node is the node used to execute DDL operations in a distributed database system and is determined from multiple nodes through election.
[0070] It should be noted that at the same time, only one node in the distributed database system will be elected as the DDL master node, and only the DDL master node can process DDL tasks in the cluster.
[0071] The first operation may include identification information of the data object targeted by the first operation, that is, the structure of which data object will be changed by the first operation. The data object may include any object such as table, column, index, constraint, view, function / stored procedure, trigger, etc.
[0072] The first operation may further include comparison information of the metadata of the target data object acquired by the second node, wherein the comparison information may be represented as timestamp information.
[0073] When the second node is the DDL master node, when executing the first operation on the data object, the second node can obtain the latest metadata of the data object. The metadata of the data object includes the security update interval of the data object. Based on the security update interval in the metadata of the data object, the second node can determine the comparison information of the metadata of the data object.
[0074] Step S120: If the comparison information is inconsistent with the to-be-matched information of the metadata of the data object stored by the first node, instruct the second node to go offline;
[0075] Among them, the first node stores the security update interval of the data object; the security update interval is used to record the status advancement information of the metadata of the data object and the online time of the data definition language DDL master node in the distributed database system; the security update interval includes at least two timestamps; the information to be matched is the last timestamp in the security update interval; the last timestamp includes the online time of the latest online DDL master node or the start time of the latest status advancement.
[0076] Specifically, the first node may store the latest security update interval of the data object targeted by the first operation. It can be understood that the first node is associated with the data object, that is, the first node may store metadata of the data object.
[0077] The security update interval of a data object can be used to record the time information of the state advancement of the target data object metadata and the online time of the Data Definition Language (DDL) master node in the distributed database system. The security update interval of a data object can include at least two timestamps, where the timestamps can be arranged in numerical order. The last timestamp in the security update interval can include the online time of the most recently online DDL master node or the start time of the most recent state advancement. The last timestamp in the security update interval can be used as the information to be matched.
[0078] It can be understood that the comparison information can be determined based on the security update interval in the metadata of the data object obtained by the second node. That is, the comparison information and the information to be matched are for different versions of metadata. The comparison information is for the metadata of the data object obtained by the second node, and the information to be matched is for the latest metadata of the data object stored by the first node.
[0079] The first node can compare the metadata comparison information obtained by the second node with the information to be matched on the data object. If the two are inconsistent, it indicates that the second node is not the DDL master node and cannot perform DDL operations. The second node is then instructed to automatically go offline. If the two are consistent, it indicates that the second node is the DDL master node and has permission to perform DDL operations. The execution result obtained after the second node performs the first operation on the data object can be stored.
[0080] In the embodiment of the present disclosure, by comparing the to-be-matched information of the metadata of the data object stored in the first node with the comparison information obtained by the second node, an active and effective judgment is made on the writing of the isolated node in the event of network isolation for the distributed database system, thereby ensuring the consistency and availability of the user data of the database to the greatest extent in the event of network isolation.
[0081] When the first node finds inconsistency with the second node after comparing them, it determines that the second node is not the DDL master node and instructs the second node to automatically go offline, thereby refusing to write the execution result of the first operation on the data object performed by the second node. This avoids data inconsistency caused by two nodes executing DDL operations at the same time (i.e., a double-group split-brain situation), and ensures data consistency under network isolation.
[0082] Optionally, the safe update interval of a data object is initialized based on:
[0083] In response to a creation operation on a data object, the start time of the first state advancement in the creation operation is used as the value of each time stamp in the security update interval of the data object.
[0084] Specifically, the initialized security update interval includes at least two timestamps, for example, a start timestamp and an intermediate timestamp.
[0085] During the process of performing a creation operation on a data object, if the first operation corresponds to the first state advancement of the creation operation of the data object, the start time of the first state advancement of the creation operation is obtained, and the start timestamp and the intermediate timestamp are set as the start time to obtain the initialized security update interval.
[0086] Optionally, the safe update interval of a data object is updated based on the following method:
[0087] If a status advancement start notification for a DDL operation on a data object is detected, a first update operation is performed on the safe update interval;
[0088] The first update operation includes:
[0089] The starting moment of the current state advancement is taken as the first moment;
[0090] Determine a first security update interval corresponding to the current first update operation; the first security update interval includes at least two timestamps;
[0091] The last timestamp in the first security update interval is set as the first moment, and the values of other timestamps in the first security update interval are determined based on the start moments of the respective states between the current state.
[0092] Specifically, when the second node starts to perform a state advancement of the DDL operation performed on the data object, a state advancement start notification can be triggered. When the state advancement start notification is detected, a first update operation can be performed on the security update interval of the data object, where the first update operation can be an update operation of the security update interval triggered by the state advancement.
[0093] Determine a first security update interval corresponding to the current first update operation, set the last timestamp in the first security update interval as the first moment, and determine values of other timestamps in the first security update interval based on start times of states between the current state.
[0094] The execution process of the first update operation is described below with reference to an example.
[0095] For example, the start time of the first state of the creation operation of the data object is TSO_200, and the initialized safe update interval is shown in Table 1, where safe-start-ts represents the start timestamp and safe-mid-ts represents the middle timestamp.
[0096] Table 1
[0097]
[0098] Assuming the start time of the second state is TSO_400, a first update operation is performed on the current safe update interval, with the start time of the current state as the first time. The last timestamp of the first safe update interval is set to the first time, that is, safe-mid-ts is set to TSO_400. The first safe update interval includes two timestamps. The start time of the first state before the second state is TSO_200. Therefore, the timestamp before safe-mid-ts (i.e., safe-start-ts) is set to TSO_200.
[0099] The security update interval after the second state starts is shown in Table 2:
[0100] Table 2
[0101]
[0102] In the case shown in Table 2, the comparison information of the metadata of the data object obtained by the second node can be expressed as TSO_400.
[0103] It should be noted that the number of timestamps included in the first security update interval is related to the compatibility requirements of the states. When the compatibility requirement includes two consecutive states that are compatible, two timestamps can be set, thereby dividing the state interval into three states. A timestamp less than the first timestamp represents the previous two states, a timestamp between the first and second timestamps represents the previous state, and a timestamp greater than the second timestamp represents the current state. When the compatibility requirement includes three consecutive states that are compatible, or when multiple specific states are compatible, the number of timestamps can be adaptively adjusted. The embodiments of this disclosure do not specifically limit the setting of the first security interval.
[0104] Optionally, if a DDL master node switching operation in the distributed database system is detected, a second update operation is performed on the safe update interval;
[0105] The second update operation includes:
[0106] The online time of the DDL master node after the switch in the current DDL master node switch operation is used as the second time;
[0107] Add a timestamp after the last timestamp of the security update interval to be updated, and use the security update interval after the timestamp as the second security update interval corresponding to the current second update operation;
[0108] The last timestamp in the second security update interval is set as the second moment.
[0109] Specifically, when a DDL master node switching operation occurs in a distributed database system, a second update operation on the safe update interval of the data object may also be triggered. The second update operation may be an update operation on the safe update interval triggered by the DDL master node switching operation.
[0110] Among them, the second update operation includes taking the online time of the DDL master node after switching in the current DDL master node switching operation as the second time, adding a timestamp after the last timestamp of the security update interval to be updated, and taking the security update interval after the added timestamp as the second security update interval corresponding to the current second update operation, and setting the last timestamp in the second security update interval as the second time.
[0111] The security update interval to be updated may be a security update interval in which the last second update operation has been performed, or may be a security update interval in which the most recent first update operation has been performed.
[0112] Still referring to the above example, the specific process of the second update operation is described below.
[0113] Assume that the second node is the DDL master node for this DDL operation. When the second node is advancing to the third state, network isolation occurs due to network instability. At this time, a new DDL master node will be elected, called the third node. The moment when the third node switches to the DDL master node is TSO_450. Then, it is necessary to perform a second update operation on the security update interval of the data object, add a timestamp at the end of the security update interval to be updated, and set the added timestamp to the second moment.
[0114] Based on the example provided in Table 2, the security update interval is updated as shown in Table 3:
[0115] Table 3
[0116]
[0117] In the case shown in Table 3, the comparison information for the metadata of the data object obtained by the second node is still TSO_400, while the comparison information for the metadata of the data object obtained by the third node is updated to TSO_450, and the information to be matched stored by the first node is also updated to TSO_450.
[0118] If the network connection of the second node is restored after the third node switches successfully, the second node also submits the first operation to advance the third state to the first node. The comparison information of the metadata of the data object obtained by the second node is TSO_400, and the information to be matched stored by the first node has been updated to TSO_450. The two are inconsistent, indicating that the second node is no longer the DDL master node, and the second node is instructed to automatically go offline.
[0119] After the third node goes online, it obtains the latest metadata security update interval. The comparison information of the metadata of the obtained data object is TSO_450, so the third node is the DDL master node.
[0120] When the second node is restarted after network isolation, the comparison information of the metadata of the data object obtained by the second node is the version information of the metadata obtained by the second node before the network isolation occurs. During the network isolation, the switching operation of the DDL master node will trigger the safe update interval of the data object. The first node finds that the comparison information of the second node is inconsistent with the information to be matched, and determines that the second node is not the DDL master node, and thus refuses to write the execution result of the first operation on the data object performed by the second node, avoiding data inconsistency caused by two nodes executing DDL operations at the same time (that is, the double-group brain split situation), and ensuring data consistency under network isolation.
[0121] In the above example, if the third node is isolated due to network isolation and a DDL master node switch occurs again in the distributed database system, and the fourth node goes online at TSO_480, an intermediate timestamp can be added to obtain the updated security update interval as shown in Table 4:
[0122] Table 4
[0123]
[0124] In another example, the fourth node starts to advance the third state, which may trigger a first update operation on the security update interval of the data object.
[0125] The start time of the third state is TSO_600. Based on Table 4, the safe update interval of the data object is shown in Table 5:
[0126] Table 5
[0127]
[0128] It should be noted that the number of timestamps included in the first safe update interval corresponding to the current first update operation is determined based on compatibility requirements. The first safe update interval records the start times of multiple state advancements and does not record the online time of the DDL master node during the state advancement. In other words, when Table 4 is updated to Table 5, the safe-mid-ts-2 and safe-mid-ts3 information in Table 4 can be deleted. It is understood that the safe update interval can record the online time of at least one DDL master node within the time period of a state advancement process. When the state advances to the next state, the structure of the safe update interval (or the number of timestamps) will be updated to the initial state.
[0129] In addition, if a notification of the end of execution of a data definition language DDL operation on a data object is detected, the end time of execution of the DDL operation is obtained, an end timestamp is added after the last timestamp of the latest security update interval, and the end time of execution is used as the value of the end timestamp.
[0130] That is, if an end timestamp is detected in the security update interval and the value of the end timestamp is not empty, it means that the corresponding DDL operation has been completed.
[0131] For example, based on Table 5, if the execution end time of the DDL operation is TSO_700, the safe update interval is updated as Table 6:
[0132] Table 6
[0133]
[0134] Among them, safe-end-ts indicates the end timestamp.
[0135] As an optional embodiment, the method further includes:
[0136] Obtaining a commit operation of a write operation on a data object submitted by the second node;
[0137] Determining compatible state time information of the data object based on a safe update interval in the latest metadata of the data object; the compatible state time information is determined from a start time of at least one compatible state determined based on the compatibility requirement;
[0138] If the transaction start time of the transaction targeted by the commit operation is not less than the compatible state time information, the commit operation is performed.
[0139] Specifically, the second node can execute DML (Data Manipulation Language) operations on the data object. When performing a write operation on the data object, a commit operation needs to be performed to persist it in the database, and the commit operation is sent to the first node that stores the data of the data object. The first node can receive the commit operation for the write operation of the data object, and based on the latest metadata of the data object, obtain the safe update interval of the target object, and determine the compatible state time information from the safe update interval of the target object. When the transaction start time of the transaction targeted by the commit operation is not less than the compatible state time information, the commit operation is performed. In other words, transactions whose transaction start time is before the compatible state time information can be committed, thereby improving the commit success rate of ordinary transactions and improving the execution performance of ordinary transactions.
[0140] The compatible state time information may represent time information of a state that is compatible with the current state (ie, a compatible state), and the compatible state time information may be determined from a start time of at least one compatible state determined based on the compatibility requirement.
[0141] Optionally, the compatible state time information may be the start time of the earliest compatible state among the at least one compatible state.
[0142] For example, when the compatibility requirement is for two consecutive states to be compatible, and the previous state is compatible with the current state, the start time of the previous state can be used as the compatible state time information. For another example, when the compatibility requirement is for three consecutive states to be compatible, and the previous two states and the previous state are both compatible with the current state, the start time of the previous two states can be used as the compatible time information. Those skilled in the art will appreciate that the method for determining compatible time information can be adjusted to suit different compatibility requirements, and this disclosure does not limit this.
[0143] Optionally, when the compatibility requirement is that two consecutive states are compatible, determining the compatible state time information of the data object based on the security update interval in the latest metadata of the data object includes:
[0144] If the security update interval of the data object does not include an end timestamp, the start timestamp in the security update interval of the target data object is used as the compatible state time information;
[0145] If the security update interval of the data object includes an end timestamp, the first intermediate timestamp in the security update interval of the target data object is used as the compatible state time information.
[0146] Specifically, if the safe update interval does not include the end timestamp, it indicates that there is a DDL operation still being executed. In this case, the relationship between the transaction start time of the transaction targeted by the commit operation and the start timestamp in the safe update interval is determined. If the transaction start time is not less than the start timestamp, it indicates that the write is safe, and the commit operation can be performed, that is, the transaction is committed.
[0147] If it is detected that the transaction start time is less than the value of the start timestamp of the latest metadata safe update interval, it means that the metadata version used by the current transaction is incompatible and writing is unsafe, so the transaction is rolled back.
[0148] Taking the security update interval shown in Table 3 as an example, as shown in Table 3, the start time of the second state is TSO_400, then the time interval greater than or equal to TSO_400 corresponds to the time interval of the second state, the time interval greater than or equal to TSO_200 and less than TSO_400 corresponds to the time interval of the first state, and the time interval less than TSO_200 corresponds to the unstarted time interval. Since the two consecutive states are compatible, transactions with a transaction start time not less than the start time of the first state (that is, not less than the start timestamp) can perform the commit operation.
[0149] Taking the security update interval shown in Table 5 as an example, as shown in Table 5, the start time of the third state is TSO_600, then the time interval greater than or equal to TSO_600 corresponds to the time interval of the third state, the time interval greater than or equal to TSO_400 and less than TSO_600 corresponds to the time interval of the second state, and the time interval less than TSO_400 corresponds to the time interval of the first state. Since the two consecutive states are compatible, transactions with a transaction start time not less than the start time of the second state (that is, not less than the start timestamp) can perform the commit operation.
[0150] If the safe update interval includes the end timestamp, it indicates that the DDL operation has ended. To meet compatibility requirements, the first intermediate timestamp can be used as the compatible time information. If the transaction start time is greater than the first intermediate timestamp, the commit operation is performed.
[0151] Taking the safe update interval shown in Table 6 as an example, as shown in Table 6, the end time of the third state is TSO_700. Then, the time interval not less than TSO_700 corresponds to the state interval in which the DDL operation has been completed. The time interval greater than or equal to TSO_600 and less than TSO_700 corresponds to the time interval of the third state. The time interval less than TSO_600 corresponds to the time interval of the second state. When two consecutive states are compatible, transactions whose transaction start time is not less than the start time of the previous state (that is, the first intermediate timestamp) can perform the commit operation.
[0152] In the embodiment of the present disclosure, since the security update interval of the data object records more fine-grained time information during the metadata update process, more detailed compatibility status time information can be obtained from the security update interval of the data object, thereby improving the submission success rate of ordinary transactions and improving the execution performance of ordinary transactions.
[0153] Optionally, the second node submits the first operation on the data object to the first node, which also includes:
[0154] Obtain the latest metadata version information of the data object to be processed, and use the latest version information as reference information;
[0155] If the reference information is inconsistent with the comparison information, it is determined that the preset condition is met.
[0156] Specifically, after restarting, the second node can obtain the latest version information of the metadata of the data object as reference information, and compare the reference information with the comparison information obtained by the second node. If the comparison information is inconsistent with the reference information, it means that a DDL master node has already advanced the status of the metadata of the data object. The second node is not the DDL master node, so the second node can automatically go offline and not submit the first operation to the first node. In this way, when the second node can successfully obtain the latest metadata of the data object, the data inconsistency problem caused by network isolation in the distributed database system is avoided.
[0157] Combined with the example in Table 3, if the network connection of the second node is restored midway, the reference information of the latest metadata of the data object obtained by the second node is TSO_450, which is inconsistent with the comparison information TSO_400 that has been obtained. Then, it can automatically go offline, thereby ensuring that there is always only one DDL master node in the distributed database system to execute DDL operations, thereby avoiding data inconsistency problems caused by network isolation in the distributed database system.
[0158] If the reference information is consistent with the comparison information, it is determined that the preset condition is met, and the second node can submit the first operation to the first node.
[0159] As an optional embodiment, the distributed database system includes at least two associated nodes associated with the data object; the at least two associated nodes include a first node;
[0160] For each associated node, the associated node stores a mapping structure of the data object, in which the identification information of the data object is used as a key and the safe update interval of the data object is used as a value.
[0161] Specifically, in a distributed database system, metadata of a data object may be stored on multiple nodes, and the multiple nodes storing the metadata of the data object are used as multiple associated nodes associated with the data object.
[0162] In addition, the user data of the data object may also be stored on multiple nodes, and the multiple nodes storing the user data of the data object may also be used as multiple associated nodes related to the data object.
[0163] For a data object targeted by the first operation submitted by the second node, the data object may correspond to at least two associated nodes, wherein the at least two associated nodes may include the first node.
[0164] For each associated node of the data object, a mapping structure of the data object can be set in the associated node, for example, a map structure. The map structure is an abstract data structure used to store key-value pairs (Key-ValuePair) and quickly find the corresponding value (Value) through the key (Key).
[0165] In the mapping structure of a data object, the identification information of the data object may be used as a key (ie, key), and the safe update range of the data object may be used as a value (ie, value).
[0166] As an optional embodiment, the metadata of the data object includes a security update interval of the data object; the at least two associated nodes include a central associated node;
[0167] For other associated nodes except the central associated node, other associated nodes have registered to monitor the identification information of the data object;
[0168] The safe update intervals for data objects in other associated nodes are updated based on the following methods:
[0169] If it is detected that the security update interval of the data object is updated, the stored security update interval of the data object is replaced with the updated security update interval.
[0170] Specifically, for each data object in a distributed database system, a safe update interval for the data object can be set in the metadata of the data object. In addition, the metadata of the data object can also store node information of the central associated node of the data object. The metadata can be represented as a data dictionary.
[0171] For the data object, a central associated node may be determined from a plurality of associated nodes of the data object. The central associated node may be an associated node storing a security update interval of the data object.
[0172] For other associated nodes except the central associated node, the other associated nodes can register to listen to the key (i.e., the identification information of the data object) in the mapping structure of the data object of the central associated node. When the value corresponding to the key of the data object in the central associated node (i.e., the security update interval of the data object) is updated, the other associated nodes can synchronously update the security update interval of the data object stored in the other associated nodes based on the watch mechanism, that is, replace the security update interval of the data object stored in the other associated nodes with the updated security update interval of the data object, thereby realizing the synchronous update of the security update interval of the associated nodes associated with the data object when the security update interval in the metadata of the data object is updated.
[0173] Compared to existing methods that use block or region granularity for updates, for example, for a data table, multiple regions must be updated before the table is considered updated. If any of these regions fails to update, a new update is required, resulting in a high failure rate and low efficiency. Furthermore, as the amount of data in a table increases, the number of regions it is divided into also increases, increasing the likelihood of failures during the update process.
[0174] In the embodiment of the present disclosure, the node is used as the smallest unit of update. By registering a watch mechanism for the central associated node through other associated nodes, when the security update interval in the metadata of the data object is updated, the security update interval of the associated nodes associated with the data object is synchronously updated, thereby improving the success rate and efficiency of the update.
[0175] Figure 2 A schematic diagram of the system architecture of a distributed database system with storage and computing separation provided in an embodiment of the present disclosure is shown in FIG. Figure 2 As shown in the figure, the storage node (i.e., storage node) and computing node (i.e., compute node) of a distributed database with storage and computing separation are different nodes.
[0176] Distributed database systems typically distribute user data, such as a table's data, evenly across all nodes according to specific rules (typically based on distribution and access hotspots). This ensures optimal load servicing and horizontal scalability. When new nodes are added, some data is rebalanced across all nodes according to these rules.
[0177] The data processing method provided by the embodiment of the present disclosure is described below using a distributed database system with a storage and computation separation structure.
[0178] For the data dictionary of each data object, a security version information registration structure is added to the data dictionary. The security version information registration structure includes the Leader Watch Point (i.e., the node information of the central associated node) and the security update interval of the data object.
[0179] For each storage node, the storage node stores a map structure of data objects related to the storage node, with the identification information of the data object as the key and the safe update range of the data object as the value. The safe update range can be represented as a JSON object.
[0180] In a distributed database system with separated storage and computing, each data object is independently watched as a key. The storage node where the first shard of each user data is located can be designated as a watch node for other monitoring nodes involving that key value. Monitoring nodes register with the central node to watch for changes to that node. Because the first user data block of each object typically implements a high-reliability protocol like Raft or Paxos, when a network failure occurs, this data block is automatically switched, or the block is migrated to another node due to load balancing. When a migration task is implemented, the Leader Watch Point (i.e., the central associated node) in the metadata is updated. Other corresponding nodes synchronize this migration mechanism through the watch mechanism. If other storage nodes registered with a watch node fail to respond within a period, the storage node is marked as expired, and the leader node in the expired node is migrated to another node.
[0181] Figure 3 A schematic diagram of a creation process is provided for an embodiment of the present disclosure, such as Figure 3 As shown, the execution process of the create operation includes:
[0182] 1. The compute node starts a DDL change transaction, which can include creating a table, adding columns, and adding indexes.
[0183] 2.1 Execute the DDL statement and store the user object definition in the data dictionary according to the format. If it fails, go to step 4, otherwise go to 2.2;
[0184] 2.2 At the storage node where the Leader watch point is located (i.e., the central associated node), update safe-start-ts, safe-mid-ts, and safe-end-ts (i.e., the safe update interval in the data dictionary of the data object), where safe-start-ts and safe-mid-ts are the transaction start values, and safe-end-ts is the transaction commit value;
[0185] 2.3 Determine whether the registration is successful (i.e., whether the security update interval is updated successfully). If yes, proceed to step 5; otherwise, proceed to step 3.
[0186] 3. When watch registration fails, determine what to do based on the returned status and your current operation:
[0187] If 3.1 times out, you can retry and go to step 2.1. If a certain number of retries is exceeded, there may be a hardware failure and the transaction will be rolled back.
[0188] 3.2 If the corresponding key value already exists, it means that other nodes have already been created and the execution is exited;
[0189] 4. Roll back the transaction;
[0190] 5. Commit the transaction.
[0191] Figure 4 A flowchart of a modification operation provided by an embodiment of the present disclosure is shown in FIG. Figure 4 As shown in the figure, the execution process of the modification operation includes:
[0192] 1. Start a DDL change transaction. A DDL change transaction can include modifying the definition of a table or column.
[0193] 2.1 After executing the DDL statement corresponding to the operation of modifying the object definition, the modified object status is stored in the data dictionary according to the format. If it is unsuccessful, go to step 4, otherwise go to 2.2;
[0194] 2.2 The values of safe-start-ts, safe-mid-ts and safe-end-ts are updated as follows:
[0195] a) Set safe-end-ts to commit ts for the transaction;
[0196] b) safe-mid-ts is set to safe-end-ts;
[0197] c) safe-start-ts is set to safe-mid-ts;
[0198] d) When the new DDL master node switches, it obtains the latest data dictionary definition and watch value from the data dictionary. After comparing them, it performs the above update steps again, so that the current ts value is as follows: safe-start-ts, safe-mid-ts-1, safe-mid-ts-2, safe-end-ts (if the network is unstable, there may be multiple mid-ts values in this state);
[0199] If obtaining the data dictionary content fails, or the comparison value is not between the latest safe-mid-ts and safe-end-ts, it means that you are not the owner of the DDL task and the task execution has been taken over by another process. You should actively exit the DDL task execution.
[0200] 2.3 Is the registration successful? If yes, proceed to step 5; otherwise, proceed to step 3.
[0201] 3. When watch registration fails, determine what to do based on the returned status and your current operation:
[0202] If 3.1 times out, you can retry and go to step 2.2. If a certain number of retries is exceeded, there may be a hardware failure and the transaction will be rolled back.
[0203] 3.2 Get the safe-mid-ts and safe-end-ts of the corresponding key value. If they are not within the range, it means that the network is isolated and the task executing the operation has been switched to another process agent and exited automatically.
[0204] 4. Roll back the transaction;
[0205] 5. Commit the transaction.
[0206] Figure 5 A flowchart of a deletion operation provided by an embodiment of the present disclosure is shown as follows: Figure 5 As shown in the figure, the execution process of the delete operation includes:
[0207] 1. Start a DDL change transaction. A DDL change transaction can be to delete a table, delete a column, etc.
[0208] 2.1 After executing the object deletion operation corresponding to the DDL statement, delete the object's data dictionary. If unsuccessful, go to step 4, otherwise go to 2.2;
[0209] 2.2 Complete the watch deletion event of the central associated node and delete the watch value. After discovering the change, other storage nodes will delete the local watch key / value pair accordingly.
[0210] 2.3 Is the deletion successful? If yes, go to step 5, otherwise go to step 3;
[0211] 3. When the watch is not successfully deleted, determine what to do based on the returned status and your current operation:
[0212] If 3.1 times out, you can retry and go to step 2.2. If a certain number of retries is exceeded, there may be a hardware failure and the transaction will be rolled back.
[0213] 3.2 If it is found that the key value no longer exists or the retrieved safe-mid-ts and safe-end-ts are not within the range, it means that the user has been deprived of the execution right due to network isolation and will exit automatically;
[0214] 4. Roll back the transaction;
[0215] 5. Commit the transaction.
[0216] The method provided by the disclosed embodiment can achieve better performance and network-isolated data consistency protection with lower resource usage and higher scalability. By changing the method of synchronizing versions based on the smallest data unit to synchronizing versions based on the deployed storage node, the time cost of version updates can be basically changed from the original proportional relationship with the data unit region to a more moderate proportional relationship after the storage node first converges. This reduces the impact of increased data on performance latency in the overall solution. Through the watch mechanism, protection mechanisms that require synchronous updates can be updated asynchronously to update the storage node status without affecting the consistency judgment of data writes.
[0217] In addition, the impact of the distributed locking mechanism on the business is eliminated. With a layer of security protection, the success rate of users submitting transactions using different versions of metadata can be improved according to different situations and the specific circumstances of the version information. This solution is robust and will not cause DDL statement execution or user transaction submission to be blocked due to the failure of a certain node. At the same time, this solution can also provide metadata scheduling nodes with more information for decision-making based on the situation, further improving the ability of the entire distributed cluster to provide high-availability services to the outside world.
[0218] Figure 6 A schematic diagram of the execution process of a write request provided by an embodiment of the present disclosure is shown as follows: Figure 6 As shown, the execution process of a write request includes:
[0219] 1. Use the latest metadata version information ts of the current computing node;
[0220] 2. Send a commit transaction request to all storage nodes involved in the transaction;
[0221] 3. The storage node only needs to make judgments on write requests;
[0222] (1) If the latest metadata version information ts of the current transaction is less than the corresponding safe-start-ts, it means that the metadata version used by the transaction is incompatible; writing is unsafe, go to step 4 and roll back directly, or go to step 1 and let the computing node update the metadata and resubmit the transaction;
[0223] (2) If the latest metadata version information ts of the current transaction is greater than the corresponding safe-start-ts, it means that the transaction is safe to write, the transaction can be committed, and go to step 5;
[0224] (3) If the latest metadata version information ts of the current transaction is greater than the corresponding safe-end-ts, it means that the current storage node may not be updated in time or network isolation occurs, triggering the storage node to refresh and retry. If the comparison passes, go to step 5; otherwise, the comparison still fails, go to step 1, hand it over to the computing node to update the metadata, and after submitting the corresponding data routing table, resubmit the transaction;
[0225] (4) When some storage nodes involved in this transaction can be submitted but some cannot, retry and prompt these storage nodes to update the synchronization mechanism immediately;
[0226] (5) If the key value of the corresponding data object cannot be found, the object may have been deleted. Go to step 1 and re-determine whether the transaction can be submitted;
[0227] 4. Roll back the transaction;
[0228] 5. Commit the transaction.
[0229] Figure 7 A structural diagram of a data processing device provided in an embodiment of the present disclosure is shown in FIG. Figure 7 As shown, the device of this embodiment may include:
[0230] An acquisition module 210 is configured to acquire a first operation submitted by a second node for metadata of a data object; the first operation includes comparison information of the metadata of the data object acquired by the second node;
[0231] a comparison module 220, configured to instruct the second node to go offline if the comparison information is inconsistent with the to-be-matched information of the metadata of the data object stored by the first node;
[0232] Among them, the first node stores the security update interval of the data object; the security update interval is used to record the status advancement information of the metadata of the data object and the online time of the data definition language DDL master node in the distributed database system; the security update interval includes at least two timestamps; the information to be matched is the last timestamp in the security update interval; the last timestamp includes the online time of the latest online DDL master node or the start time of the latest status advancement.
[0233] As an optional embodiment, the device further includes a general transaction processing module, configured to:
[0234] Obtaining a commit operation of a write operation on a data object submitted by the second node;
[0235] Determining compatible state time information of the data object based on a safe update interval in the latest metadata of the data object; the compatible state time information is determined from a start time of at least one compatible state determined based on the compatibility requirement;
[0236] If the transaction start time of the transaction targeted by the commit operation is not less than the compatible state time information, the commit operation is performed.
[0237] As an optional embodiment, the apparatus further includes a first updating module, configured to:
[0238] If a notification of successful advancement of the data definition language (DDL) operation on the data object is detected, a first update operation is performed on the safe update interval;
[0239] The first update operation includes:
[0240] The starting moment of the current state advancement is taken as the first moment;
[0241] Determining a first security update interval corresponding to the current first update operation; the first security update interval includes at least two timestamps;
[0242] The last timestamp in the first security update interval is set as the first moment, and the values of other timestamps in the first security update interval are determined based on the start moments of the respective states between the current state.
[0243] As an optional embodiment, the apparatus further includes a second updating module, configured to:
[0244] If a DDL master node switching operation in the distributed database system is detected, performing a second update operation on the safe update interval;
[0245] The second update operation includes:
[0246] The online time of the DDL master node after the switch in the current DDL master node switch operation is used as the second time;
[0247] Add a timestamp after the last timestamp of the security update interval to be updated, and use the security update interval after the timestamp as the second security update interval corresponding to the current second update operation;
[0248] The last timestamp in the second security update interval is set as the second moment.
[0249] As an optional embodiment, the device further includes an initialization module, configured to:
[0250] In response to a creation operation on the data object, a start time of a first state advancement in the creation operation is used as a value of each timestamp in a security update interval of the data object.
[0251] As an optional embodiment, the distributed database system includes at least two associated nodes associated with the data object; the at least two associated nodes include the first node;
[0252] For each associated node, the associated node stores a mapping structure of the data object, wherein the mapping structure uses the identification information of the data object as a key and the safe update interval of the data object as a value.
[0253] As an optional embodiment, the metadata of the data object includes a security update interval of the data object; the at least two associated nodes include a central associated node;
[0254] For other associated nodes except the central associated node, the other associated nodes have registered to monitor the identification information of the data object;
[0255] The device further includes a third updating module, configured to:
[0256] If it is monitored that the security update interval of the data object is updated, the stored security update interval of the data object is replaced with the updated security update interval.
[0257] The apparatus of the embodiments of the present disclosure can execute the methods provided by the embodiments of the present disclosure, and their implementation principles are similar and have corresponding technical effects. The actions performed by each module in the apparatus of each embodiment of the present disclosure correspond to the steps in the methods of each embodiment of the present disclosure. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions of the corresponding methods shown above, and will not be repeated here.
[0258] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or portion of a computer program that has a predetermined function and works together with other related components to achieve a predetermined goal. The ... that can be implemented in whole or in part using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0259] In an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method provided in any optional embodiment of the present disclosure. Compared with the prior art, the electronic device can be implemented as follows: by comparing the metadata to be matched of the data object stored by the first node with the comparison information obtained by the second node, in the event of network isolation in the distributed database system, an active and effective judgment is made by utilizing the write to the isolated node, thereby ensuring the consistency and availability of the user data in the database to the greatest extent in the event of network isolation.
[0260] When the first node finds inconsistency with the second node after comparing them, it determines that the second node is not the DDL master node and instructs the second node to automatically go offline, thereby refusing to write the execution result of the first operation on the data object performed by the second node. This avoids data inconsistency caused by two nodes executing DDL operations at the same time (i.e., a double-group split-brain situation), and ensures data consistency under network isolation.
[0261] In an alternative embodiment, an electronic device is provided, such as Figure 8 As shown, Figure 8 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure.
[0262] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.
[0263] Bus 4002 may include a path for transmitting information between the above components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0264] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0265] The memory 4003 is used to store the computer program for executing the embodiments of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiments.
[0266] Among them, electronic devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., as well as fixed terminals such as digital TVs, desktop computers, etc.
[0267] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0268] The embodiments of the present disclosure further provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiments when executed by a processor.
[0269] It should be understood that, although the flowcharts of the embodiments of the present disclosure indicate the various operation steps by arrows, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be performed in other orders as required. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In scenarios where the execution times are different, the order of execution of these sub-steps or stages can be flexibly configured as required, and the embodiments of the present disclosure do not limit this.
[0270] The above description is only an optional implementation method for some implementation scenarios of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of the present disclosure, other similar implementation methods based on the technical ideas of the present disclosure also fall within the protection scope of the embodiments of the present disclosure.
Claims
1. A data processing method, characterized in that: Applied to a distributed database system, the distributed database system includes a first node and a second node, and the method is executed by the first node; The method comprises: Obtaining a first operation submitted by a second node for metadata of a data object; the first operation including comparison information of the metadata of the data object obtained by the second node; If the comparison information is inconsistent with the to-be-matched information of the metadata of the data object stored by the first node, instructing the second node to go offline; Among them, the first node stores the security update interval of the data object; the security update interval is used to record the status advancement information of the metadata of the data object and the online time of the data definition language DDL master node in the distributed database system; the security update interval includes at least two timestamps; the information to be matched is the last timestamp in the security update interval; the last timestamp includes the online time of the latest online DDL master node or the start time of the latest status advancement.
2. The method according to claim 1, characterized in that The method further comprises: Obtaining a commit operation of a write operation on a data object submitted by the second node; Determining compatible state time information of the data object based on a safe update interval in the latest metadata of the data object; the compatible state time information is determined from a start time of at least one compatible state determined based on the compatibility requirement; If the transaction start time of the transaction targeted by the commit operation is not less than the compatible state time information, the commit operation is performed.
3. The method according to claim 1 or 2, characterized in that The safe update interval of the data object is updated based on the following method: If a status advancement start notification of the DDL operation on the data object is detected, a first update operation is performed on the safe update interval; The first update operation includes: The starting moment of the current state advancement is taken as the first moment; Determining a first security update interval corresponding to the current first update operation; the first security update interval includes at least two timestamps; The last timestamp in the first security update interval is set as the first moment, and the values of other timestamps in the first security update interval are determined based on the start moments of the respective states between the current state.
4. The method according to claim 1 or 2, characterized in that The safe update interval of the data object is updated based on the following method: If a DDL master node switching operation in the distributed database system is detected, performing a second update operation on the safe update interval; The second update operation includes: The online time of the DDL master node after the switch in the current DDL master node switch operation is used as the second time; Add a timestamp after the last timestamp of the security update interval to be updated, and use the security update interval after the added timestamp as the second security update interval corresponding to the current second update operation; The last timestamp in the second security update interval is set as the second moment.
5. The method according to claim 1, wherein The safe update interval of the data object is initialized based on the following method: In response to a creation operation on the data object, a start time of a first state advancement in the creation operation is used as a value of each timestamp in a security update interval of the data object.
6. The method according to claim 1, characterized in that The distributed database system includes at least two associated nodes associated with the data object; the at least two associated nodes include the first node; For each associated node, the associated node stores a mapping structure of the data object, wherein the mapping structure uses the identification information of the data object as a key and the safe update interval of the data object as a value.
7. The method according to claim 6, characterized in that The metadata of the data object includes a security update interval of the data object; the at least two associated nodes include a central associated node; For other associated nodes except the central associated node, the other associated nodes have registered to monitor the identification information of the data object; The safe update interval of the data object in the other associated nodes is updated based on the following method: If it is monitored that the security update interval of the data object is updated, the stored security update interval of the data object is replaced with the updated security update interval.
8. A data processing device, characterized in that: Applied to a distributed database system, the distributed database system includes a first node and a second node, the first node includes the data processing device; The data processing device includes: an acquisition module, configured to acquire a first operation submitted by a second node for metadata of a data object; the first operation including comparison information of the metadata of the data object acquired by the second node; a comparison module, configured to instruct the second node to go offline if the comparison information is inconsistent with the to-be-matched information of the metadata of the data object stored by the first node; Among them, the first node stores the security update interval of the data object; the security update interval is used to record the status advancement information of the metadata of the data object and the online time of the data definition language DDL master node in the distributed database system; the security update interval includes at least two timestamps; the information to be matched is the last timestamp in the security update interval; the last timestamp includes the online time of the latest online DDL master node or the start time of the latest status advancement.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Database cluster split-brain prevention method and device, electronic equipment and storage medium
CN117992501A
Database node state monitoring method and device, electronic equipment and storage medium
CN119127625A