Data processing method, device, electronic device and computer readable storage medium
By generating transaction identifiers in a distributed database and judging the database status, it allows modification requests to be performed for migrated data during data migration, which solves the problem of stopping client modification during data migration in the prior art, and achieves data consistency and normal business processing.
Patent Information
- Application Number
- CN202011117925.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-10-19
AI Technical Summary
During the data migration process, existing distributed databases need to stop the client's modification of the data being migrated, resulting in the client's business being affected.
The first transaction identifier is generated by receiving the data migration information sent by the management node, the status of the distributed database is judged, and the data modification request is allowed to be performed on the migration data in the data migration state, and the request is sent to the second node to ensure data consistency.
It realizes the data consistency of the first node and the second node to be migrated during the data migration process, and avoids the impact on client services.
Smart Images

Figure CN114385580B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and more specifically, to a data processing method, device, electronic device, and computer-readable storage medium. Background Art
[0002] Due to changes in data size and data access hotspots, distributed databases may need online expansion (such as online adding nodes), online reduction (such as online deleting nodes), or online load balancing (such as redistribution of data within a distributed cluster). These operations all involve migrating data within the distributed database from one memory database (MDB) node to another.
[0003] At present, when the data of node A (i.e., memory database node) in a distributed database is migrated to node B, it is necessary to stop the client from modifying the migrated data in node A to ensure the consistency of the migrated data of the two nodes, that is, the data received by node B is consistent with the data migrated by node A. It can be seen that during the data migration process of the existing solution, it is necessary to stop the client from modifying the data being migrated, which affects the normal operation of the client business. Summary of the invention
[0004] A first aspect of the present disclosure provides a data processing method, comprising:
[0005] receiving data migration information sent by the management node, and generating a first transaction identifier based on the data migration information;
[0006] When receiving the data modification request, if the distributed database to which the first node belongs is in a data migration state, performing the operation indicated by the data modification request on the data to be migrated in the first node; wherein the state of the distributed database is determined based on the first transaction identifier;
[0007] The data modification request is sent to a second node in the distributed database so that the second node performs the operation indicated by the data modification request on the data to be migrated.
[0008] Optionally, after receiving the data modification request, the method further includes:
[0009] Determine whether the first transaction identifier is greater than a preset threshold;
[0010] If the first transaction identifier is greater than a preset threshold, and the second transaction identifier carried in the data modification request is greater than the first transaction identifier, it is determined that the distributed database is in a data migration state.
[0011] Optionally, after determining whether the first transaction identifier is greater than a preset threshold, the method further includes:
[0012] If the first transaction identifier is not greater than a preset threshold, it is determined that there is no data migration transaction in the distributed database, and the operation indicated by the data modification request is performed on the data to be migrated in the first node;
[0013] If the first transaction identifier is greater than a preset threshold and the second transaction identifier carried in the data modification request is smaller than the first transaction identifier, it is determined that there is a data migration transaction in the distributed database but the data migration state has not been entered, and the operation indicated by the data modification request is performed on the data to be migrated in the first node.
[0014] Optionally, receiving a data modification request includes:
[0015] Receiving a data modification request and distributed database status information sent by a third node in the distributed database;
[0016] The method also includes:
[0017] Based on the distributed database state information, it is determined that the distributed database is in a data migration state.
[0018] Optionally, the data modification request includes a data deletion request, and the second node performs the operation indicated by the data modification request on the data to be migrated, including:
[0019] The second node performs the following operations:
[0020] Searching for data corresponding to the first index based on the first index in the data deletion request;
[0021] If the data corresponding to the first index is found, the data corresponding to the first index is deleted;
[0022] If the data corresponding to the first index is not found, the data position corresponding to the first index in the second node is determined, and a placeholder marker is set for the data position so that after the first node migrates the data corresponding to the first index to the data position, the second node deletes the data corresponding to the first index based on the placeholder marker.
[0023] Optionally, the data modification request includes a data adding request, and the second node performs the operation indicated by the data modification request on the data to be migrated, including:
[0024] The second node performs the following operations:
[0025] If the data corresponding to the second index in the data modification request does not exist in the second node, then the data corresponding to the second index is added;
[0026] If the data corresponding to the second index in the data modification request exists in the second node, detecting whether a placeholder mark is set at the data position corresponding to the second index;
[0027] If no placeholder marker is set, insert the data corresponding to the second index at the data position;
[0028] If a placeholder marker is set, the data corresponding to the second index is inserted at the data position, and a prompt message for canceling the migration of the data corresponding to the second index is sent to the first node to delete the placeholder marker.
[0029] Optionally, the data migration information includes an identifier of the data to be migrated and an identifier of the second node, and before receiving the data modification request, the method further includes:
[0030] Determining the data to be migrated stored in the first node based on the identifier of the data to be migrated;
[0031] Determining, based on the identifier of the second node in the data migration indication information, that the destination node of the data migration is the second node;
[0032] When a data migration instruction is received from the management node, the data to be migrated is migrated from the first node to the second node based on the data migration instruction.
[0033] Optionally, after migrating the data to be migrated from the first node to the second node based on the data migration instruction, the method further includes:
[0034] Modify the first transaction identifier;
[0035] If the modified first transaction identifier is not greater than the preset threshold, it is determined that the distributed database is in a data migration completion state;
[0036] If it is determined that there is no client accessing the original data corresponding to the data to be migrated in the first node, the original data is deleted.
[0037] A second aspect of the present disclosure provides a data processing device, including:
[0038] A first receiving module, configured to receive data migration information sent by the management node, and generate a first transaction identifier based on the data migration information;
[0039] A second receiving module is configured to, upon receiving a data modification request, execute the operation indicated by the data modification request on the data to be migrated in the first node if the distributed database to which the first node belongs is in a data migration state; wherein the state of the distributed database is determined based on the first transaction identifier;
[0040] The sending module is used to send the data modification request to the second node in the distributed database, so that the second node performs the operation indicated by the data modification request on the data to be migrated.
[0041] Optionally, the device further comprises:
[0042] A judging module, used to judge whether the first transaction identifier is greater than a preset threshold;
[0043] The first determination module is used to determine that the distributed database is in a data migration state if the first transaction identifier is greater than a preset threshold and the second transaction identifier carried in the data modification request is greater than the first transaction identifier.
[0044] Optionally, the device further comprises:
[0045] A second determination module is configured to determine that there is no data migration transaction in the distributed database if the first transaction identifier is not greater than a preset threshold, and to perform the operation indicated by the data modification request on the data to be migrated in the first node;
[0046] The third determination module is used to determine that there is a data migration transaction in the distributed database but the data migration state has not been entered if the first transaction identifier is greater than a preset threshold and the second transaction identifier carried in the data modification request is smaller than the first transaction identifier, and to perform the operation indicated by the data modification request on the data to be migrated in the first node.
[0047] Optionally, the second receiving module is specifically configured to:
[0048] Receiving a data modification request and distributed database status information sent by a third node in the distributed database;
[0049] The device also includes a fourth determination module, which is used to determine that the distributed database is in a data migration state based on the distributed database state information.
[0050] Optionally, the data modification request includes a data deletion request, and the second node performs the operation indicated by the data modification request on the data to be migrated, including:
[0051] The second node performs the following operations:
[0052] Searching for data corresponding to the first index based on the first index in the data deletion request;
[0053] If the data corresponding to the first index is found, the data corresponding to the first index is deleted;
[0054] If the data corresponding to the first index is not found, the data position corresponding to the first index in the second node is determined, and a placeholder marker is set for the data position so that after the first node migrates the data corresponding to the first index to the data position, the second node deletes the data corresponding to the first index based on the placeholder marker.
[0055] Optionally, the data modification request includes a data adding request, and the second node performs the operation indicated by the data modification request on the data to be migrated, including:
[0056] The second node performs the following operations:
[0057] If the data corresponding to the second index in the data modification request does not exist in the second node, then the data corresponding to the second index is added;
[0058] If the data corresponding to the second index in the data modification request exists in the second node, detecting whether a placeholder mark is set at the data position corresponding to the second index;
[0059] If a placeholder is set, delete it.
[0060] Optionally, the data migration information includes an identifier of the data to be migrated and an identifier of the second node; and the device further includes:
[0061] a fifth determining module, configured to determine the data to be migrated stored in the first node based on the identifier of the data to be migrated;
[0062] a sixth determining module, configured to determine that the destination node of the data migration is the second node based on the identifier of the second node in the data migration indication information;
[0063] The data migration module is configured to migrate the data to be migrated from the first node to the second node based on the data migration instruction when receiving the data migration instruction sent by the management node.
[0064] Optionally, the device further comprises:
[0065] A modification module, used for modifying the first transaction identifier;
[0066] A seventh determination module, configured to determine that the distributed database is in a data migration completion state if the modified first transaction identifier is not greater than a preset threshold;
[0067] The eighth determination module is configured to delete the original data corresponding to the data to be migrated in the first node if it is determined that there is no client accessing the original data.
[0068] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory and a processor; a computer program is stored in the memory; and the processor is used to execute the method of the first aspect and any one of its optional implementations when running the computer program.
[0069] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method of the first aspect and any one of its optional implementations is implemented.
[0070] The beneficial effects of the technical solution provided by this application are:
[0071] In this embodiment, after receiving the data migration information sent by the management node, a first transaction identifier can be generated based on the data migration information. Then, when a data modification request is received, if the distributed database to which the first node belongs is in a data migration state, that is, the data modification request is received when the distributed database is in a data migration state, then the operation indicated by the data modification request can be performed on the data to be migrated in the first node; wherein, the state of the distributed database is determined based on the first transaction identifier. At the same time, the data modification request can be sent to a second node in the distributed database so that the second node executes the operation indicated by the data modification request on the data to be migrated. In this way, both the first node and the second node can execute the operation indicated by the data modification request on the data to be migrated. Even if there is data modification during the data migration process, the data to be migrated of the first node and the second node are consistent. It can be seen that the present application can ensure the consistency of the data to be migrated of the two nodes during the data migration process, and therefore the client business can also be processed normally during the data migration process. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.
[0073] Figure 1 A schematic diagram of the connection of each node in the distributed data of this application;
[0074] Figure 2 A flowchart of the data processing method of this application;
[0075] Figure 3 Another flowchart of the data processing method of this application is shown below;
[0076] Figure 4 This is a schematic diagram of the structure of the data processing device of this application;
[0077] Figure 5 This is a schematic diagram of the structure of the electronic device of this application. DETAILED DESCRIPTION
[0078] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present invention.
[0079] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "an", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may also be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0080] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0081] First, several terms involved in this application are introduced and explained:
[0082] The management node, first node, second node and third node in the embodiments of the present application all refer to MDB nodes in a distributed database; the management node can manage and monitor other nodes in the distributed database, and the management node can control each node to perform corresponding operations by issuing instructions to each node.
[0083] A transaction identifier can uniquely identify a transaction. The transaction identifier in a distributed database is incremental, that is, the identifier of a transaction generated at a certain moment is greater than the identifier of a transaction generated at a moment before that moment.
[0084] A possible connection relationship between the management node and each MDB node is as follows: Figure 1 As shown, the management node samples a high-availability cluster with one master and multiple backups. There is one master management node and multiple backup management nodes. The management node can monitor each MDB node, and the MDB node can report data to the management node.
[0085] Reference Figure 2 The present application provides a data processing method, which is executed by a first node in a distributed database, wherein the first node refers to a source node for data migration, and data migration is to migrate the data to be migrated in the first node to a second node, and the method includes:
[0086] Step S201: receiving data migration information sent by a management node, and generating a first transaction identifier based on the data migration information;
[0087] The management node determines a data migration plan, generates corresponding data migration information based on the determined data migration plan, and sends the data migration information to each node so that each node (including the first node) can generate its own first transaction identifier based on the data migration information.
[0088] In this embodiment, after each node generates the first transaction identifier, each node can provide services to the client, that is, each node can receive or process client requests, such as receiving a data modification request from the client. The first transaction identifier generated by each node can be different.
[0089] Step S202: upon receiving the data modification request, if the distributed database to which the first node belongs is in a data migration state, performing the operation indicated by the data modification request on the data to be migrated in the first node; wherein the state of the distributed database is determined based on the first transaction identifier;
[0090] As described above, each node has generated a first transaction identifier, and each node can provide services to the client. Then:
[0091] There are two possible situations when the first node receives a data modification request: the first possible situation is that the first node receives a data modification request sent by the client, that is, the client directly accesses the first node. In this case, the state of the distributed database is determined by the first node based on the first transaction identifier generated by the first node; the second possible situation is: the first node receives the data modification request when the first node receives the data modification request sent by other nodes, that is, the client accesses other nodes, and the other nodes can forward the data modification request to the first node. In this case, the state of the distributed database is determined by the other nodes based on the first transaction identifier generated by themselves.
[0092] For the data modification requests received in the above two situations, when the first node determines that the distributed database to which it belongs is in a data migration state (if the client accesses other nodes, the other nodes may determine that the distributed database is in a data migration state and then inform the first node), it means that the data modification request is a request received when the distributed data is in a migration state, and the first node can perform the operation indicated by the data modification request on the data to be migrated in the first node.
[0093] Among them, data modification requests specifically include: data deletion requests, data addition requests and data update requests. Data update requests can actually be divided into data deletion and data addition. The processing method of each node for data update requests can refer to the following processing method for data deletion requests and data addition requests.
[0094] Step S203: Send the data modification request to the second node in the distributed database so that the second node performs the operation indicated by the data modification request on the data to be migrated.
[0095] In this embodiment, the first node can send a data modification request to the second node in the distributed database. The second node refers to the destination node for data migration in the distributed database, so that the second node can also perform the operation indicated by the data modification request on the data to be migrated.
[0096] In this embodiment, both the first node and the second node need to process the data modification request, that is, the first node and the second node need to be dual-written, and the dual writing is serial. After the first node completes processing the data modification request, the first node sends it to the second node to continue processing the data modification request. The processing of the data modification request is performed within the same transaction.
[0097] In this embodiment, after receiving the data migration information sent by the management node, a first transaction identifier can be generated based on the data migration information. Then, when a data modification request is received, if the distributed database to which the first node belongs is in a data migration state, that is, the data modification request is received when the distributed database is in a data migration state, then the operation indicated by the data modification request can be performed on the data to be migrated in the first node; wherein, the state of the distributed database is determined based on the first transaction identifier. At the same time, the data modification request can be sent to a second node in the distributed database so that the second node executes the operation indicated by the data modification request on the data to be migrated. In this way, both the first node and the second node can execute the operation indicated by the data modification request on the data to be migrated. Even if there is data modification during the data migration process, the data to be migrated of the first node and the second node are consistent. It can be seen that the present application can ensure the consistency of the data to be migrated of the two nodes during the data migration process, and therefore the client business can also be processed normally during the data migration process.
[0098] Furthermore, the data migration information specifically includes:
[0099] The identifier of the source node of data migration (i.e., the first node), the identifier of the destination node of data migration (i.e., the second node), and the information of the data to be migrated in the first node. The information of the data to be migrated specifically includes: the table name, column name, and index information of the data to be migrated.
[0100] The index of data can uniquely identify a piece of data. If the data stored in the first node adopts hash partitioning, the index information refers to the hash value of the index field; if it is range partitioning, the index information refers to the start value and end value of the index field.
[0101] It should be noted that, for the case where the first node directly receives the data modification request sent by the client, then:
[0102] After receiving the data modification request in step S203, the method further includes:
[0103] Determine whether the first transaction identifier is greater than a preset threshold;
[0104] If the first transaction identifier is greater than a preset threshold, and the second transaction identifier carried in the data modification request is greater than the first transaction identifier, it is determined that the distributed database is in a data migration state.
[0105] The first node can determine whether the first transaction identifier is greater than a preset threshold. If the first transaction identifier is greater than the preset threshold, the first node can determine that there is a data migration transaction in the distributed database. If the first transaction identifier is not greater than the preset threshold, the first node can determine that there is no data migration transaction in the distributed database. The preset threshold can be set to 0. It should be noted that the first transaction identifier generated by each node after receiving the data migration information is greater than the preset threshold, so that each node can provide services to the client.
[0106] If the first transaction identifier is greater than a preset threshold, and the first node determines based on the migration identifier that the data migration is not completed, the first node may compare the first transaction identifier with the second transaction identifier in the data modification request to further determine whether the distributed database is in a migration state;
[0107] In this embodiment, the first node can determine whether the data migration is completed based on the migration identifier. The migration identifier is variable and can be changed accordingly based on the instructions sent by the management node. If the migration identifier is a preset object, for example, the migration identifier is true, then it is determined that the data migration is completed. If the migration identifier is not a preset object, then it is determined that the data migration is not completed (for example, the migration identifier is false).
[0108] Case A: If the second transaction identifier is greater than the first transaction identifier, it means that the transaction indicated by the second transaction identifier is a transaction after the transaction indicated by the first transaction identifier, that is, the data modification request is a transaction received after the data migration transaction, then it can be determined that the distributed database is in the data migration state, and the data modification request is a request received when the distributed database is in the migration state. In this case, dual writing is required for the first node and the second node, and both the first node and the second node need to process the data modification request.
[0109] Optionally, after determining whether the first transaction identifier is greater than a preset threshold, the method further includes:
[0110] If the first transaction identifier is not greater than a preset threshold, it is determined that there is no data migration transaction in the distributed database, and the operation indicated by the data modification request is performed on the data to be migrated in the first node;
[0111] If the first transaction identifier is greater than a preset threshold and the second transaction identifier carried in the data modification request is smaller than the first transaction identifier, it is determined that there is a data migration transaction in the distributed database but the data migration state has not been entered, and the operation indicated by the data modification request is performed on the data to be migrated in the first node.
[0112] Case B: If the first transaction identifier is not greater than a preset threshold, the first node may execute the operation indicated by the data modification request on the data to be migrated in the first node.
[0113] Case C: If the first transaction identifier is greater than a preset threshold and the migration identifier indicates that the data migration is not completed, then when the second transaction identifier carried in the data modification request is smaller than the first transaction identifier, it means that the transaction indicated by the second transaction identifier is a transaction before the transaction indicated by the first transaction identifier, that is, the data migration transaction needs to be processed after the data modification request is processed. In this case, the first node can determine that there is a data migration transaction in the distributed database but the data migration state has not been entered, and it only needs to perform the operation indicated by the data modification request on the data to be migrated in the first node.
[0114] In the case where the third node receives a data modification request sent by the client and then forwards the data modification request to the first node, then:
[0115] Receiving a data modification request includes:
[0116] Receiving a data modification request and distributed database status information sent by a third node in the distributed database;
[0117] The method further includes:
[0118] Based on the distributed database state information, it is determined that the distributed database is in a data migration state.
[0119] After the third node receives the data modification request sent by the client, the third node can determine the distributed database status information. If the third node determines, based on the distributed database status information, that the data modification request is a request received when the distributed database is in a data migration state, the third node forwards the data modification request to the first node based on the identifier of the source node that performs data migration in the received data migration information, i.e., the first node. At the same time, the third node also sends the distributed database status information to the first node. In this way, the first node can determine that the distributed database is in a data migration state based on the distributed database status information. The first node can process the data modification request. After the first node completes the processing, the data modification request is sent to the second node so that the second node can process the data modification request.
[0120] In this embodiment, the third node is an MDB node in the distributed database that is different from the first node and the second node.
[0121] In this embodiment, the way in which the third node determines that the data modification request is a request received when the distributed database is in a migration state can refer to the way in which the first node determines that the data modification request is a request received when the distributed database is in a data migration state, that is, first determine whether the first transaction identifier is greater than a preset threshold, and then determine based on the migration identifier that the data migration is not completed, and then determine whether the second transaction identifier is greater than the first transaction identifier. The details will not be repeated here.
[0122] Further, the first node performs the operation indicated by the data modification request on the data to be migrated, specifically referring to:
[0123] If the conditions B and C are met, all the data to be migrated are in the first node, and the first node performs the operation indicated by the data modification request on the data to be migrated.
[0124] If situation C is met, the first node stores the original data of the data to be migrated, and the original data is static and will not change due to data migration. The first node can perform the operation indicated by the data modification request on the original data.
[0125] Furthermore, in step S203, the second node performs the operation indicated by the data modification request on the data to be migrated, specifically including:
[0126] One possible situation is: if the data modification request includes a data deletion request, the second node performs the operation indicated by the data modification request on the data to be migrated, including:
[0127] The second node performs the following operations:
[0128] Searching for data corresponding to the first index based on the first index in the data deletion request;
[0129] If the data corresponding to the first index is found, the data corresponding to the first index is deleted;
[0130] If the data corresponding to the first index is not found, the data position corresponding to the first index in the second node is determined, and a placeholder marker is set for the data position so that after the first node migrates the data corresponding to the first index to the data position, the second node deletes the data corresponding to the first index based on the placeholder marker.
[0131] The second node searches for the data corresponding to the first index based on the first index in the data deletion request. If the second node finds the data corresponding to the first index, it means that the data corresponding to the first index has been migrated to the second node, and the second node can directly delete the data corresponding to the first index.
[0132] If the second node fails to search for the data corresponding to the first index, it means that the data corresponding to the first index has not yet been migrated to the second node. The second node can determine the data position corresponding to the first index in the data list of the second node, and the second node sets a placeholder marker for the data position. In this way, after the first node migrates the data corresponding to the first index to the data position in the second node, the second node can delete the data corresponding to the first index based on the pre-set placeholder marker.
[0133] A possible scenario of the present implementation is: during the process of migrating data 1 to 100 to the second node, the first node receives a data deletion request, and the data deletion request indicates to delete data 50. The first node deletes data 50 in the original data based on the data deletion request. Eventually, the data to be migrated in the first node becomes 1 to 49 and 51 to 100. The second node searches the data list for data 50. If so, data 50 is deleted. If not, a placeholder is first set at the data position of data 50, and after waiting for data 50 to be migrated from the first node to the second node, data 50 is deleted. Eventually, the data to be migrated received by the second node are 1 to 49 and 51 to 100, which ensures data consistency during the data migration process between the two nodes. Therefore, the solution of the present application can process client requests during the data migration process.
[0134] It can be seen that the present application processes the client's data deletion request through both the first node and the second node, thereby ensuring the consistency of the data to be migrated during the data migration process between the two nodes, realizing the normal processing of the client's request during the data migration process, and avoiding the impact on the client's business.
[0135] Another possible situation is: if the data modification request includes a data adding request, the second node performs the operation indicated by the data modification request on the data to be migrated, specifically including:
[0136] The second node performs the following operations:
[0137] If the data corresponding to the second index in the data modification request does not exist in the second node, then the data corresponding to the second index is added;
[0138] If the data corresponding to the second index in the data modification request exists in the second node, detecting whether a placeholder mark is set at the data position corresponding to the second index;
[0139] If a placeholder is set, delete it.
[0140] In this embodiment, the second node detects whether the data corresponding to the second index in the data modification request exists in the data list of the second node. If the data corresponding to the second index in the data addition request does not exist, the data corresponding to the second index is inserted into the data list of the second node. Possible reasons why the data corresponding to the second index in the data addition request does not exist in the second node are: the data corresponding to the second index is not included in the data to be migrated; or the data to be migrated includes the data but the first node has not migrated the data to the second node. In this case, the second node can also send a prompt message to the first node to cancel the migration of the data corresponding to the second index.
[0141] If the second node has data corresponding to the second index in the data addition request, the second node further determines whether there is a placeholder mark at the index position of the second index;
[0142] If there is no placeholder marker, the possible situation is: the first node has added data based on the data addition request, and the second node has not had time to add the data based on the data addition request. The first node has migrated the data to the second node. In this case, when dual writing is performed, it will be found that the data already exists in the data list of the second node. At this time, the second node does not need to add the data again.
[0143] If there is a placeholder marker, the possible situation is: the data corresponding to the first index and the data corresponding to the second index are the same data, the second node first performs a deletion operation on the data corresponding to the first index, but because the first node has not yet migrated the data corresponding to the first index to the second node, the second node first sets a placeholder marker at the data position corresponding to the first index. After the first node migrates the data corresponding to the first index to the second node, it has not deleted the data corresponding to the first index. The second node now performs a new operation on the data corresponding to the second index. At this time, the second node will no longer add or delete the data, and the second node can delete the placeholder marker.
[0144] Optionally, the data migration information includes an identifier of the data to be migrated and an identifier of the second node, and before step S202 receives the data modification request, the method further includes:
[0145] Determining the data to be migrated stored in the first node based on the identifier of the data to be migrated;
[0146] Determining, based on the identifier of the second node in the data migration indication information, that the destination node of the data migration is the second node;
[0147] When a data migration instruction is received from the management node, the data to be migrated is migrated from the first node to the second node based on the data migration instruction.
[0148] The first node can determine that the source node for data migration is the first node based on the identifier of the first node in the data migration information, and determine that the destination node for data migration is the second node based on the identifier of the second node in the data migration information. The data to be migrated in the first node can be determined based on the table name, column name and index information of the data to be migrated in the data migration information. When the first node receives the data migration instruction sent by the management node, the first node migrates the data to be migrated from the first node to the second node, and the distributed database to which the first node belongs enters the data migration state.
[0149] The specific process of data migration is: the first node searches for data according to the index, sends the data to the second node, and the second node inserts the data into the corresponding data position in the second node's data list. A fixed number of data rows or a fixed size of data blocks can be set for each migration.
[0150] Optionally, after migrating the data to be migrated from the first node to the second node based on the data migration instruction, the method further includes:
[0151] Modify the first transaction identifier;
[0152] If the modified first transaction identifier is not greater than the preset threshold, it is determined that the distributed database is in a data migration completion state;
[0153] If it is determined that there is no client accessing the original data corresponding to the data to be migrated in the first node, the original data is deleted.
[0154] In this embodiment, the management node can control the first node to modify the first transaction identifier, so that the first node can determine that the distributed database is in a migration completion state based on the modified first transaction identifier. For example, when the modified first transaction identifier is not greater than a preset threshold, it is determined that the distributed database is in a migration completion state. After determining that the data migration is completed, the first node can detect whether all access to the data to be migrated has been completed. If there is no client accessing the original data of the data to be migrated, the original data is deleted.
[0155] In summary, refer to Figure 3, the technical solution of this application may actually include the following steps:
[0156] Step S1: The management node generates data migration information;
[0157] Step S2: the management node sends data migration information to each node, so that each node generates a first transaction identifier based on the data migration information.
[0158] The specific implementation of steps S1 and S2 refers to the relevant description of the above embodiment.
[0159] It should be noted that after each node generates the first transaction identifier, each node may set the first transaction identifier as the transaction identifier in the metadata information of each node.
[0160] At the same time, each node configures other information in the metadata information of each node based on the data migration information, specifically including: the source node of the data migration, the destination node, the information of the data to be migrated and the migration identifier. For the relevant explanation of the information of the data to be migrated and the migration identifier, please refer to the above embodiment.
[0161] Each node can set the first transaction identifier to trx_id_migration_begin. Each node starts a thread to obtain the current active transaction list. When the minimum active transaction number of a node is greater than trx_id_migration_begin, the node notifies the management node that there is no transaction less than trx_id_migration_begin in the current node, that is, all transactions before the start of data migration have been processed. When the management node receives notifications from all nodes, it controls the data migration process to the next step.
[0162] It should be noted that any node must ensure atomicity when obtaining the transaction identifier and the metadata information corresponding to the node configuration.
[0163] Step S3: The management node sends a data migration instruction to the first node, so that the first node starts data migration based on the data migration instruction.
[0164] Step S4: If the first node receives a data modification request from the client, if the first transaction identifier is greater than a preset threshold, and the migration identifier is false, and the second transaction identifier in the second data modification request is greater than the first transaction identifier, then it is determined that the distributed database is in a data migration state, and the first node processes the data modification request and sends the data modification request to the second node at the same time, so that the second node also processes the data modification request.
[0165] The specific implementation of steps S3 and S4 refers to the relevant description of the above embodiment.
[0166] Step S5: The first node sends a notification that data migration is completed to the management node;
[0167] It should be noted that after the migration is completed, the first node notifies the management node. If the migration times out, the management node will actively inquire with the first node. After receiving the migration completion notification, the management node controls the data migration process to the next step.
[0168] Step S6: The management node controls each node to modify the first transaction identifier to ensure that the minimum transaction identifier of each node is greater than the modified first transaction identifier;
[0169] The management node notifies each node to modify the first transaction identifier in the metadata information to trx_id_migration_end and set the migration mark to true.
[0170] Each node starts a thread to obtain the minimum transaction number in the current active transaction list. When the minimum transaction number is greater than trx_id_migration_end, it means that the data has been migrated and the destination node of the data migration is likely to provide correct data. Each node notifies the management node that the minimum transaction number is greater than trx_id_migration_end, so that the management node can control the data migration process to the next step.
[0171] Step S7: The management node determines that all nodes have completed access to the original data of the data to be migrated.
[0172] The management node notifies each node to modify the metadata information of the first node, specifically: modify the first transaction identifier to be no greater than a preset threshold (indicating that there is no data migration, that is, the distributed database is in a data migration completion state), set the migration record to false, and use the second node as the source node for data migration; at this time, the first transaction identifier is no greater than the preset threshold, and the migration identifier is false, so the reset after the migration is completed can be identified.
[0173] Each node applies for a new transaction number, which does not need to be recorded in the metadata information of the first node, and starts a new thread to obtain the minimum transaction identifier of the current active transaction list. When the minimum transaction identifier is greater than the new transaction identifier, it means that all requests to access the original data have ended. Each node notifies the management node so that the management node can control the data migration process to the next step.
[0174] Step S8: The management node notifies the first node to delete the original data and the second node to delete the data with the placeholder marker.
[0175] The migration_delete field may be used to indicate whether a placeholder exists. A value of 0 in the migration_delete field indicates that there is no placeholder, and a value of 1 in the migration_delete field indicates that there is a placeholder.
[0176] The management node starts a transaction to delete the original data on the first node and the row records with placeholders on the second node. If the transaction is successful, the data migration is complete. If it times out, the management node will query the corresponding node.
[0177] Please refer to Figure 4 The present application also provides a data processing device, the device comprising:
[0178] A first receiving module 401 is configured to receive data migration information sent by a management node, and generate a first transaction identifier based on the data migration information;
[0179] The second receiving module 402 is configured to, upon receiving a data modification request, execute the operation indicated by the data modification request on the data to be migrated in the first node if the distributed database to which the first node belongs is in a data migration state; wherein the state of the distributed database is determined based on the first transaction identifier;
[0180] The sending module 403 is used to send the data modification request to the second node in the distributed database, so that the second node performs the operation indicated by the data modification request on the data to be migrated.
[0181] Optionally, the device further comprises:
[0182] A judging module, used to judge whether the first transaction identifier is greater than a preset threshold;
[0183] The first determination module is used to determine that the distributed database is in a data migration state if the first transaction identifier is greater than a preset threshold and the second transaction identifier carried in the data modification request is greater than the first transaction identifier.
[0184] Optionally, after the determination module determines whether the first transaction identifier is greater than a preset threshold, the device further includes:
[0185] A second determination module is configured to determine that there is no data migration transaction in the distributed database if the first transaction identifier is not greater than a preset threshold, and to perform the operation indicated by the data modification request on the data to be migrated in the first node;
[0186] The third determination module is used to determine that there is a data migration transaction in the distributed database but the data migration state has not been entered if the first transaction identifier is greater than a preset threshold and the second transaction identifier carried in the data modification request is smaller than the first transaction identifier, and to perform the operation indicated by the data modification request on the data to be migrated in the first node.
[0187] Optionally, the second receiving module 402 is specifically configured to:
[0188] Receiving a data modification request and distributed database status information sent by a third node in the distributed database;
[0189] The device also includes a fourth determining module, configured to:
[0190] Based on the distributed database state information, it is determined that the distributed database is in a data migration state.
[0191] Optionally, the data modification request includes a data deletion request, and the second node performs the operation indicated by the data modification request on the data to be migrated, including:
[0192] The second node performs the following operations:
[0193] Searching for data corresponding to the first index based on the first index in the data deletion request;
[0194] If the data corresponding to the first index is found, the data corresponding to the first index is deleted;
[0195] If the data corresponding to the first index is not found, the data position corresponding to the first index in the second node is determined, and a placeholder marker is set for the data position so that after the first node migrates the data corresponding to the first index to the data position, the second node deletes the data corresponding to the first index based on the placeholder marker.
[0196] Optionally, the data modification request includes a data adding request, and the second node performs the operation indicated by the data modification request on the data to be migrated, including:
[0197] The second node performs the following operations:
[0198] If the data corresponding to the second index in the data modification request does not exist in the second node, then the data corresponding to the second index is added;
[0199] If the data corresponding to the second index in the data modification request exists in the second node, detecting whether a placeholder mark is set at the data position corresponding to the second index;
[0200] If a placeholder is set, delete it.
[0201] Optionally, the data migration information includes an identifier of the data to be migrated and an identifier of the second node, and the device further includes:
[0202] a fifth determining module, configured to determine the data to be migrated stored in the first node based on the identifier of the data to be migrated;
[0203] a sixth determining module, configured to determine that the destination node of the data migration is the second node based on the identifier of the second node in the data migration indication information;
[0204] The data migration module is configured to migrate the data to be migrated from the first node to the second node based on the data migration instruction when receiving the data migration instruction sent by the management node.
[0205] Optionally, the device further comprises:
[0206] A modification module, used for modifying the first transaction identifier;
[0207] A seventh determination module, configured to determine that the distributed database is in a data migration completion state if the modified first transaction identifier is not greater than a preset threshold;
[0208] The eighth determination module is configured to delete the original data corresponding to the data to be migrated in the first node if it is determined that there is no client accessing the original data.
[0209] The data processing device of this embodiment can execute the data processing method shown in any of the above embodiments of the present application. The implementation principles are similar and will not be repeated here.
[0210] In an alternative embodiment, an electronic device is provided, such as Figure 5 As shown, Figure 5 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0211] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0212] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0213] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.
[0214] The memory 4003 is used to store application code for executing the solution of the present application, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the application code stored in the memory 4003 to implement the content shown in any of the above method embodiments.
[0215] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0216] The above are only some embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A data processing method, It is characterized in that The method is performed by a first node, and includes: receiving data migration information sent by the management node, and generating a first transaction identifier based on the data migration information; When receiving the data modification request, if the distributed database to which the first node belongs is in a data migration state, performing the operation indicated by the data modification request on the data to be migrated in the first node; wherein the state of the distributed database is determined based on the first transaction identifier; Sending the data modification request to a second node in the distributed database so that the second node performs the operation indicated by the data modification request on the data to be migrated; wherein the second node is the node to which the data to be migrated is to be migrated; If the data modification request includes a data deletion request, the second node performing the operation indicated by the data modification request on the data to be migrated includes: Searching for data corresponding to the first index based on the first index in the data deletion request; If the data corresponding to the first index is found, the data corresponding to the first index is deleted; If the data corresponding to the first index is not found, the data position corresponding to the first index in the second node is determined, and a placeholder marker is set for the data position, so that after the first node migrates the data corresponding to the first index to the data position, the second node deletes the data corresponding to the first index based on the placeholder marker.
2. The method according to claim 1, It is characterized in that After receiving the data modification request, the method further includes: Determining whether the first transaction identifier is greater than a preset threshold; If the first transaction identifier is greater than a preset threshold, and the second transaction identifier carried in the data modification request is greater than the first transaction identifier, it is determined that the distributed database is in a data migration state.
3. The method according to claim 2, It is characterized in that After determining whether the first transaction identifier is greater than a preset threshold, the method further includes: If the first transaction identifier is not greater than a preset threshold, determining that there is no data migration transaction in the distributed database, and performing the operation indicated by the data modification request on the data to be migrated in the first node; If the first transaction identifier is greater than a preset threshold and the second transaction identifier carried in the data modification request is smaller than the first transaction identifier, it is determined that there is a data migration transaction in the distributed database but the data migration state has not been entered, and the operation indicated by the data modification request is performed on the data to be migrated in the first node.
4. The method according to claim 1, It is characterized in that The receiving of the data modification request comprises: Receiving a data modification request and distributed database status information sent by a third node in the distributed database; The method further comprises: Based on the distributed database state information, it is determined that the distributed database is in a data migration state.
5. The method according to any one of claims 1 to 4, It is characterized in that If the data modification request includes a data adding request, the second node performs the operation indicated by the data modification request on the data to be migrated, including: The second node performs the following operations: If the data corresponding to the second index in the data modification request does not exist in the second node, then adding the data corresponding to the second index; If the data corresponding to the second index in the data modification request exists in the second node, detecting whether a placeholder mark is set at the data position corresponding to the second index; If the placeholder marker is set, delete the placeholder marker.
6. The method according to any one of claims 1 to 4, It is characterized in that The data migration information includes an identifier of the data to be migrated and an identifier of the second node, and before receiving the data modification request, the method further includes: Determining the data to be migrated stored in the first node based on the identifier of the data to be migrated; Determining, based on the identifier of the second node in the data migration indication information, that the destination node of the data migration is the second node; When receiving the data migration instruction sent by the management node, the data to be migrated is migrated from the first node to the second node based on the data migration instruction.
7. The method according to claim 6, It is characterized in that After migrating the data to be migrated from the first node to the second node based on the data migration instruction, the method further includes: Modifying the first transaction identifier; If the modified first transaction identifier is not greater than a preset threshold, it is determined that the distributed database is in a data migration completion state; If it is determined that no client accesses the original data corresponding to the data to be migrated in the first node, the original data is deleted.
8. A data processing device, It is characterized in that The device is deployed in a first node, and includes: A first receiving module, configured to receive data migration information sent by a management node, and generate a first transaction identifier based on the data migration information; a second receiving module, configured to, upon receiving a data modification request, execute the operation indicated by the data modification request on the data to be migrated in the first node if the distributed database to which the first node belongs is in a data migration state; wherein the state of the distributed database is determined based on the first transaction identifier; A sending module, configured to send the data modification request to a second node in the distributed database, so that the second node performs the operation indicated by the data modification request on the data to be migrated; wherein the second node is the node to which the data to be migrated is to be migrated; If the data modification request includes a data deletion request, the second node performing the operation indicated by the data modification request on the data to be migrated includes: Searching for data corresponding to the first index based on the first index in the data deletion request; If the data corresponding to the first index is found, the data corresponding to the first index is deleted; If the data corresponding to the first index is not found, the data position corresponding to the first index in the second node is determined, and a placeholder marker is set for the data position, so that after the first node migrates the data corresponding to the first index to the data position, the second node deletes the data corresponding to the first index based on the placeholder marker.
9. An electronic device, It is characterized in that The electronic device comprises a memory and a processor; The memory stores a computer program; The processor is configured to execute the method according to any one of claims 1 to 7 when running the computer program.
10. A computer-readable storage medium, It is characterized in that The storage medium stores a computer program, and when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Data migration method, apparatus and medium and electronic apparatus
CN109325016A