Data reading method and device
By updating and synchronizing data version numbers in a distributed storage system, the data reading inconsistency caused by alternating oscillations of OSD nodes is solved, and data consistency and business continuity are achieved.
Patent Information
- Application Number
- CN202510125043.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-23
AI Technical Summary
In distributed storage systems, data reading results are inconsistent due to alternating oscillations of OSD nodes during data reading, which seriously threatens business continuity.
By receiving the data read request, the target placement group PG is determined, and when the single replica read condition is met, data reading is performed based on the PG data mapped by the single replica OSD node. If the PG version number is inconsistent, update the PG version number and synchronize the data in the OSD cluster to ensure data consistency.
It realizes data consistency when reading data, ensures business continuity, and improves data reliability of distributed storage systems.
Smart Images

Figure CN120029548A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more particularly to a data reading method and device. Background Art
[0002] In a distributed storage system, a placement group (PG), as the basic unit of data distribution and replica management in the distributed storage system, is usually mapped to multiple object storage devices (OSDs), so that the multiple OSDs mapped to it are jointly responsible for data storage and access.
[0003] During the data writing process, PG data will be written to each OSD. However, due to various reasons such as software defects, external environmental interference, and hardware failure, some OSDs may successfully write PG data, while the rest of the OSDs may fail to write PG data.
[0004] In this case, even if only one of the multiple OSDs mapped by PG works normally and the writing of PG data in this OSD fails, the system will still read data based on the PG in this OSD. However, when multiple OSD nodes mapped by PG alternately oscillate and some of the OSD nodes are written successfully, if data is read from these OSD nodes next time, the two reading results will be inconsistent, posing a serious threat to business continuity. Summary of the invention
[0005] This specification provides a data reading method and device to partially solve the above problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This specification provides a data reading method, including:
[0008] Receive a first data read request, determine a target placement group PG according to the first data read request, and when it is determined that the object storage device OSD cluster mapped with the target PG meets the single copy read condition, perform data reading according to the PG data in the target PG mapped by the single copy OSD node in the OSD cluster; wherein, if the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, update the PG version number of the target PG, and record the updated PG version number of the target PG in the single copy OSD node; the OSD cluster meeting the single copy read condition means that there is a single copy OSD node in the OSD cluster; the updated PG version number is the maximum version number of the target PG mapped by each OSD node in the OSD cluster;
[0009] When it is detected that other OSD nodes in the OSD cluster switch from a non-operating state to a running state, an authoritative OSD node is selected based on the PG version number of the PG recorded by each OSD node in the OSD cluster, wherein the PG version number of the target PG mapped by the selected authoritative OSD node is the updated PG version number;
[0010] Synchronize the data in the authoritative OSD node to the other OSD nodes.
[0011] Optionally, if the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, the method further includes:
[0012] At a preset time interval, re-judge whether the OSD cluster meets the single copy reading condition, and re-judge whether the PG version number indicated by the first data read request is consistent with the PG version number of the target PG, until a preset judgment number threshold is reached;
[0013] The updating of the PG version number of the target PG specifically includes:
[0014] If the OSD cluster still meets the single copy reading condition after reaching the preset judgment number threshold, and the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, the PG version number of the target PG is updated.
[0015] Optionally, after re-determining whether the OSD cluster meets the single copy reading condition, the method further includes:
[0016] If the OSD cluster currently has multiple OSD nodes in a running state, then determining an authoritative OSD node from the multiple OSD nodes in a running state;
[0017] Synchronize data in the authoritative OSD node determined among the multiple OSD nodes in operation to other OSD nodes currently in operation;
[0018] Data is read according to the PG data in the target PG mapped by the master OSD node in other OSD nodes that are currently in operation.
[0019] Optionally, the method further comprises:
[0020] After synchronizing the data in the authoritative OSD node to the other OSD nodes, if a second data read request is received, data reading is performed based on the PG data in the target PG mapped by the master OSD node in each OSD node in operation.
[0021] Optionally, selecting an authoritative OSD node based on the PG version number of the target PG recorded by each OSD node in the OSD cluster specifically includes:
[0022] Initiate a query message to the replica OSD node in the OSD cluster through the master OSD node in the OSD cluster;
[0023] Receiving, through the master OSD node, a response message returned by each replica OSD node according to the query message; wherein the response message returned by each replica OSD node includes a PG version number of a target PG mapped by the replica OSD node;
[0024] The authoritative OSD node is selected based on the response message returned by each replica OSD node.
[0025] Optionally, the authoritative OSD node is the OSD node with the largest PG version number of the PG mapped among the OSD nodes; or,
[0026] The authoritative OSD node is the OSD node with the largest PG version number of the PG mapped among the OSD nodes and the longest PG data in the mapped PG.
[0027] This specification provides a data reading device, including:
[0028] A reading module, used to receive a first data read request, determine a target placement group PG according to the first data read request, and when it is determined that the object storage device OSD cluster mapped with the target PG meets the single copy read condition, read data according to the PG data in the target PG mapped by the single copy OSD node in the OSD cluster; wherein, if the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, the PG version number of the target PG is updated, and the updated PG version number of the target PG is recorded in the single copy OSD node; the OSD cluster meeting the single copy read condition means that there is a single copy OSD node in the OSD cluster; the updated PG version number is the maximum version number of the target PG mapped by each OSD node in the OSD cluster;
[0029] A selection module is used to select an authoritative OSD node based on the PG version number of the PG recorded by each OSD node in the OSD cluster when detecting that other OSD nodes in the OSD cluster switch from a non-operating state to a running state, wherein the PG version number of the target PG mapped by the selected authoritative OSD node is the updated PG version number;
[0030] A synchronization module is used to synchronize the data in the authoritative OSD node to the other OSD nodes.
[0031] Optionally, the device further comprises:
[0032] A judgment module is used to re-judge whether the OSD cluster meets the single copy reading condition according to a preset time interval, and re-judge whether the PG version number indicated by the first data read request is consistent with the PG version number of the target PG, until a preset judgment number threshold is reached;
[0033] The reading module is specifically used to update the PG version number of the target PG if the OSD cluster still meets the single copy reading condition after reaching a preset judgment number threshold, and the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG.
[0034] Optionally, after determining whether there is a single OSD node in the OSD cluster that is currently in operation, the synchronization module is also used to, if there are multiple OSD nodes in the OSD cluster that are currently in operation, determine an authoritative OSD node among the multiple OSD nodes in operation; synchronize data in the authoritative OSD node determined among the multiple OSD nodes in operation to other OSD nodes that are currently in operation; and read data based on PG data in the target PG mapped by the master OSD node in other OSD nodes that are currently in operation.
[0035] Optionally, the reading module is further used to, after synchronizing the data in the authoritative OSD node to the other OSD nodes, read data according to the PG data in the target PG mapped by the master OSD node in each OSD node in operation if a second data reading request is received.
[0036] Optionally, the synchronization module is specifically used to initiate a query message to the replica OSD nodes in the OSD cluster through the master OSD node in the OSD cluster; receive a response message returned by each replica OSD node based on the query message through the master OSD node; wherein the response message returned by each replica OSD node includes the PG version number of the target PG mapped by the replica OSD node; and select the authoritative OSD node based on the response message returned by each replica OSD node.
[0037] Optionally, the authoritative OSD node is the OSD node with the largest PG version number of the PG mapped in each OSD node; or, the authoritative OSD node is the OSD node with the largest PG version number of the PG mapped in each OSD node and the longest PG data in the mapped PG.
[0038] The present application provides a machine-readable storage medium, which stores machine-executable instructions that can be executed by a processor; wherein the processor is used to execute the machine-executable instructions to implement the above-mentioned data reading method.
[0039] The present application provides a computer program, which is stored in a machine-readable storage medium. When a processor executes the computer program in the machine-readable storage medium, the processor is prompted to implement the above-mentioned data reading method.
[0040] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0041] When this solution meets the single-copy reading conditions during the first data read, it will read the data based on the PG data in the target placement group PG mapped on the single-copy OSD node. When the version number corresponding to the data read request is inconsistent with the version number corresponding to the target PG, the version number of the target PG is updated to the largest version number in the OSD cluster. In this way, the authoritative OSD node can be determined based on the updated version number during the next data read, and data can be read based on the target PG in the authoritative OSD node, ensuring the consistency of the two data reads and further ensuring business continuity. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0043] Figure 1 This is a schematic diagram of an OSD node alternating oscillation process provided in this specification;
[0044] Figure 2 A flowchart of a data reading method provided in this specification;
[0045] Figure 3 A schematic diagram of a data synchronization process provided in this specification;
[0046] Figure 4 A schematic diagram of a data reading process after a retry operation provided in this specification;
[0047] Figure 5 A schematic diagram of a data reading device provided in this specification;
[0048] Figure 6 A method provided in this specification corresponding to Figure 2 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0050] In an OSD cluster, node status changes may include node online, offline, failure, etc. These status changes are critical to cluster stability and data integrity. When the OSD node status changes, the cluster needs to detect and take appropriate measures in a timely manner to ensure data reliability and consistency.
[0051] Therefore, when the state of the OSD node in the cluster changes, it is necessary to read the PG data from the OSD node in the up state, and then read the actual data such as operation log data based on the PG data, so as to replay the log according to the read data, thereby helping the OSD cluster to achieve state synchronization, ensuring that each OSD node in the cluster can process data correctly, and further ensuring the stability and availability of the cluster.
[0052] However, when processing the disk on multiple OSDs mapped by PG, if some OSDs fail to write the new PG data successfully and only one OSD node is in operation (i.e., up state), the system will read the old version of PG data through a single copy. Once the OSD node alternately oscillates, causing the OSD node that successfully wrote the new PG data to come back online (i.e., change from down state to up state), if the new PG data is read from this OSD node, the data read twice will inevitably be inconsistent. This process is as follows: Figure 1 shown.
[0053] Among them, taking the initial PG version number as 1.0 as an example, at epoch1, data is written to three OSD nodes, osd.1, osd.2 and osd.3, respectively. However, due to network reasons or other abnormal conditions, osd.2 fails to write. Assuming that data is successfully written to osd.1 and osd.3 at this time, the PG version number will be updated from 1.0 to 1.1. Since osd.2 fails to write, the PG version number remains 1.0.
[0054] At epoch2, due to a node failure, osd.1 and osd.3, which successfully wrote data, are both in a non-operating state. At this time, only osd.2 is in the up state, so the old data with PG version number 1.0 will be read from it.
[0055] At epoch 3, osd.1 is in the up state, but osd.2 and osd.3 are in the down state. At this time, new data with PG version number 1.1 will be read from osd.1. This scenario leads to inconsistency in the results of the two data reads.
[0056] In order to solve the above problems, the technical solutions provided by the embodiments to be protected by this specification are described in detail below in conjunction with the accompanying drawings.
[0057] Figure 2 A flowchart of a data reading method provided in this specification includes the following steps:
[0058] S201: Receive a first data read request, determine a target placement group PG according to the first data read request, and when it is determined that the object storage device OSD cluster mapped with the target PG meets the single copy read condition, read data according to the PG data in the target PG mapped by the single copy OSD node in the OSD cluster; wherein, if the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, update the PG version number of the target PG, and record the updated PG version number of the target PG at the single copy OSD node; the OSD cluster meets the single copy read condition means that there is a single copy OSD node in the OSD cluster; the updated PG version number is the maximum version number of the target PG mapped by each OSD node in the OSD cluster.
[0059] In this specification, the execution entity used to implement a data reading method can be a device or node such as a server, controller or monitor in an OSD cluster. For the sake of ease of description, the following will only take the server as the execution entity as an example to illustrate a data reading method provided in this specification.
[0060] Among them, the above-mentioned OSD cluster includes multiple OSD nodes mapped with target PGs, and each OSD node corresponds to a physical storage device or virtual storage device in the cluster.
[0061] For each OSD node, one or more PG data is stored in the OSD node. When the client writes data to the OSD node, the data is first converted into objects, and then these objects are mapped to the corresponding PG according to a certain algorithm (such as a hash algorithm). Each object is mapped to a unique PG, and a PG may contain multiple objects, thereby obtaining the data of each PG.
[0062] During this process, each PG containing PG data is mapped to an OSD list. The first OSD node in the list is the primary OSD node, and the remaining OSD nodes are replica OSD nodes.
[0063] When receiving the first data read request, the server may determine the target placement group PG according to the first data read request, and then determine whether the OSD cluster mapped with the target PG meets the single copy read condition.
[0064] If the current OSD cluster meets the single-copy reading condition, the server can read data based on the PG data in the target placement group PG mapped on the single-copy OSD node.
[0065] Among them, the OSD cluster meets the single-copy reading condition means that there is a single-copy OSD node in the OSD cluster, and the single-copy OSD node means that there is a single OSD node in the OSD cluster that is in running state.
[0066] When the OSD cluster meets the single-copy reading condition, the server can first determine whether the PG version number (current_les_epoch) corresponding to the target PG mapped in the OSD node is consistent with the PG version number indicated by the first data read request.
[0067] The first data read request may be triggered and generated by the OSD cluster itself after the state of the OSD node in the OSD cluster changes, and may also be sent by the user through the client.
[0068] If the PG version number indicated by the first data read request is consistent with the PG version number corresponding to the target PG, it means that the latest PG data has been successfully written into the target PG and matches the data read request. Therefore, data can be read directly based on the PG data in the target PG, and the current version number of the target PG can be used as the target version number.
[0069] When the status of each OSD node in the OSD cluster is updated, the server can determine the authoritative OSD node from multiple OSD nodes that map the PG corresponding to the target version number.
[0070] In actual applications, the version number indicated by the data read request is usually greater than or equal to the PG version number in the OSD node. When the status of each OSD node in the OSD cluster is updated, if the PGs mapped in the multiple OSD nodes that are currently in operation are all the target version numbers, it means that the multiple OSD nodes where the target PGs are located are in the up state. Therefore, the authoritative OSD node can be determined among them, and the data in the authoritative OSD node can be synchronized to other OSD nodes.
[0071] In this specification, the authoritative OSD node can be the OSD node with the largest PG version number among multiple OSD nodes. When the PG version numbers mapped by multiple OSD nodes are all the largest version numbers, the authoritative OSD node can be the OSD node with the longest PG data in the mapped PG.
[0072] On this basis, if only one OSD node maps the PG version number as the largest version number, the server can use this OSD node as the authoritative OSD node. If there are multiple OSD nodes whose PG version numbers are all the largest version numbers, the OSD node with the longest PG data in the PGs mapped in these OSD nodes will be used as the authoritative OSD node.
[0073] In addition, if there are multiple longest PG data, the running master OSD node (the first OSD node in the OSD list) will be used as the authoritative OSD node.
[0074] The server can then synchronize the data in the authoritative OSD node to other running OSD nodes.
[0075] For ease of understanding, this specification provides a schematic diagram of the data synchronization process, such as Figure 3 shown.
[0076] Among them, the main OSD node can first initiate a query message to each replica OSD, and then select the authoritative OSD node according to the size of the current_les_epoch value, and give priority to selecting the authoritative OSD node with the largest current_les_epoch value.
[0077] If the current_les_epoch of multiple OSD nodes are equal, the PG data on these OSD nodes are compared to determine the authoritative PG data. The selection of authoritative PG data gives priority to the PG data in the PG with the largest PG version number. At the same time, if the PG version numbers are the same, the PG data with a longer data length is selected as the authoritative PG data. Finally, according to the missing object information on different replica OSDs, the primary OSD pushes these PG data to each replica OSD and persists them to complete data synchronization.
[0078] In this way, when the second data read request is accepted, since the data in the authoritative OSD node has been synchronized to each OSD node in operation before, data reading based on the PG data in the PG mapped by any OSD node can ensure the consistency of the two data reads.
[0079] For example, at epoch 1, osd.1 and osd.3 both successfully wrote PG data, and the PG version numbers were updated from 1.0 to 1.1. However, osd.2 failed to successfully write PG data, and its PG version number remained 1.0.
[0080] At epoch2, osd.1 and osd.2 are turned to down state, and osd.3 is in up state. At this time, the version number corresponding to the data read request is 1.1, which is consistent with the PG version number in osd.3, so data can be read directly from osd.3.
[0081] At epoch3, if osd.1 and osd.3 are both in the up state, osd.3 can be used as the authoritative OSD node and the data in it can be synchronized to osd.1. In this way, the PG version numbers corresponding to the PGs in osd.1 and osd.3 are both 1.1. At this time, whether reading data based on the PG data in osd.1 or osd.3, the data read is consistent with the data read at epoch2.
[0082] If the version number indicated by the first data read request is inconsistent with the PG version number corresponding to the target PG, it means that the OSD currently in the up state has not been successfully written with the latest PG data, and the PG data in the target PG is the historical version of PG data.
[0083] In this case, the server can still read data based on the historical version of PG data, but in order to ensure data consistency the next time data is read, the PG version number corresponding to the target PG data needs to be updated to obtain the updated PG version number, and persist it to the hard disk of the single-copy OSD node to record the updated PG version number of the target PG in the single-copy OSD node. Among them, the updated version number is the maximum version number of the PG mapped by each OSD node in the OSD cluster, that is, it is greater than the PG version number indicated by the current data read request, the PG version number before the update, and the version number of the PG in other OSD nodes. On this basis, even if other OSDs are turned to the up state, the updated PG version number is still the largest PG version number.
[0084] For example, at epoch 1, osd.1 and osd.3 both successfully wrote PG data, and the PG version numbers were updated from 1.0 to 1.1. However, osd.2 failed to successfully write PG data, and its PG version number remained 1.0.
[0085] At epoch2, osd.1 and osd.3 are turned to down state, and osd.2 is in up state. At this time, the PG version number indicated by the data read request is 1.1, but the PG version number in the single copy osd.2 is 1.0. Although the PG data in the PG with version number 1.0 in osd.2 will still be obtained for data reading, its PG version number will be updated to 1.2 and persisted to the hard disk.
[0086] In actual applications, each OSD node may be offline for a short period of time due to situations such as network congestion, process conflict or response timeout, and may still be restored to an operating state after a short period of time. Therefore, the server can first determine whether there is a single OSD node in the OSD cluster that is currently in an operating state. If so, continue to determine whether the PG version number indicated by the current first data read request is consistent with the PG version number of the target PG.
[0087] If the PG version number indicated by the current first data read request is inconsistent with the PG version number of the target PG, a retry process is entered. The retry process can re-judge whether the OSD cluster meets the single copy reading condition at a preset time interval, and re-judge whether the PG version number indicated by the first data read request is consistent with the PG version number of the target PG, until a preset judgment number threshold is reached;
[0088] If the OSD cluster still meets the single-copy reading condition after reaching the preset judgment number threshold, and the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, it means that the OSD node in the down state in the OSD cluster will not be restored to the up state in a short time. Therefore, the PG version number of the target PG can be updated under the single-copy reading condition.
[0089] In addition, if it is detected that there are multiple OSD nodes in running state within the retry number threshold, or the PG version number of the target PG data in the single-copy OSD node is consistent with the PG version number indicated by the first data read request, the PG data is obtained from the OSD node and the data is read through the corresponding multi-copy reading method in step S201.
[0090] For ease of understanding, this manual provides a schematic diagram of the data reading process after a retry operation, such as Figure 4 shown.
[0091] Among them, the server can retry at intervals of 1s. Each retry needs to determine whether the PG version number is consistent with the PG version number indicated by the read op) and whether the PG is a single copy; if more than 5 retries are still unsuccessful, then enter the subsequent processing flow.
[0092] S202: When it is detected that other OSD nodes in the OSD cluster are switched from a non-operating state to a running state, an authoritative OSD node is selected based on the PG version number of the PG recorded by each OSD node in the OSD cluster, wherein the PG version number of the target PG mapped by the selected authoritative OSD node is the updated PG version number;
[0093] S203: Synchronize the data in the authoritative OSD node to the other OSD nodes.
[0094] After detecting that the status of each OSD node in the OSD cluster has been updated, the other OSD nodes in the OSD cluster will switch from a non-running state to a running state. At this time, the server can select an authoritative OSD node based on the PG version number of the PG recorded by each OSD node in the OSD cluster.
[0095] Specifically, since the authoritative node in this specification is the OSD node with the largest PG version number, under the condition of single copy reading, the updated PG version number is the maximum PG version number in the entire OSD cluster. Therefore, the server will use the single copy OSD node that was previously subjected to single copy reading (that is, the OSD node where the PG corresponding to the updated PG version number is located) as the authoritative OSD node, and synchronize the data in the authoritative OSD node to other OSD nodes.
[0096] During data synchronization, the master OSD node in the OSD cluster can initiate a query message to the replica OSD node in the OSD cluster, and then the master OSD node can receive a response message returned by each replica OSD node based on the query message. The response message returned by each replica OSD node contains the PG version number of the target PG mapped by the replica OSD node.
[0097] Afterwards, the authoritative OSD node can be selected from the replica OSD nodes and the primary OSD node based on the PG version number contained in the response message returned by each replica OSD node.
[0098] After synchronizing the data in the authoritative OSD node to the other OSD nodes, if a second data read request is received, since each OSD node currently in operation has completed data synchronization, the PG data in the PG mapped by each OSD node is the same as the PG data in the authoritative OSD node. At this time, no matter which OSD node in operation is used to read the data based on the PG data, the consistency of the data read twice can be guaranteed. In this case, data can be read based on the PG data in the target PG mapped by the main OSD node in each OSD node in operation.
[0099] Continuing with the above example, at epoch 2, data is read from the PG data in the PG with version number 1.0 in osd.2, and the PG version number is updated from 1.0 to 1.2.
[0100] In epoch 3, if osd.2 is in the up state, the largest PG version number in the cluster is the PG version number corresponding to the PG mapped in osd.2, so osd.2 is used as the authoritative OSD node.
[0101] At epoch3, if osd.1, osd.2 and osd.3 are all in the up state, the data in osd.2 will be synchronized to osd.1 and osd.3. At this time, no matter which OSD node's PG data is used to read the data (usually the data is read based on the PG data in the master node osd.1), the data reading result will be consistent with the data reading result of epoch2.
[0102] At epoch3, if only osd.2 is in the up state, data is read directly according to the PG data in the PG mapped by osd.2.
[0103] In addition, at epoch 3, if osd.2 is in the down state, data synchronization cannot be performed. At this time, in order to ensure data consistency, data reading can be prohibited and a read failure message can be returned after receiving the second data read request.
[0104] It can be seen from the above method that when multiple OSDs mapped by PG oscillate alternately and some OSDs write successfully, this scheme ensures the consistency of the two PG data readings by implementing a single-frame data reading and data synchronization method between OSDs, thereby further improving the data reliability of the distributed storage system.
[0105] The above are one or more methods for implementing data reading in this specification. Based on the same idea, this specification also provides a corresponding data reading device, such as Figure 5 shown.
[0106] Figure 5 A schematic diagram of a data reading device provided in this specification includes:
[0107] The reading module 501 is used to receive a first data reading request, determine a target placement group PG according to the first data reading request, and when it is determined that the object storage device OSD cluster mapped with the target PG meets the single copy reading condition, perform data reading according to the PG data in the target PG mapped by the single copy OSD node in the OSD cluster; wherein, if the PG version number indicated by the first data reading request is inconsistent with the PG version number of the target PG, the PG version number of the target PG is updated, and the updated PG version number of the target PG is recorded in the single copy OSD node; the OSD cluster meets the single copy reading condition means that there is a single copy OSD node in the OSD cluster; the updated PG version number is the maximum version number of the target PG mapped by each OSD node in the OSD cluster;
[0108] A selection module 502 is configured to select an authoritative OSD node based on the PG version number of the PG recorded by each OSD node in the OSD cluster when detecting that other OSD nodes in the OSD cluster are switched from a non-operating state to a running state, wherein the PG version number of the target PG mapped by the selected authoritative OSD node is the updated PG version number;
[0109] The synchronization module 503 is used to synchronize the data in the authoritative OSD node to the other OSD nodes.
[0110] Optionally, the device further comprises:
[0111] The judgment module 504 is used to re-judge whether the OSD cluster meets the single copy reading condition according to a preset time interval, and re-judge whether the PG version number indicated by the first data read request is consistent with the PG version number of the target PG, until a preset judgment number threshold is reached;
[0112] The reading module 501 is specifically used to update the PG version number of the target PG if the OSD cluster still meets the single copy reading condition after reaching a preset judgment number threshold, and the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG.
[0113] Optionally, the synchronization module 503 is also used to, if the OSD cluster currently has multiple OSD nodes in running state, determine an authoritative OSD node among the multiple OSD nodes in running state; synchronize data in the authoritative OSD node determined among the multiple OSD nodes in running state to other OSD nodes that are currently in running state; and read data according to PG data in the target PG mapped by the master OSD node in other OSD nodes that are currently in running state.
[0114] Optionally, the reading module 501 is further used to, after synchronizing the data in the authoritative OSD node to the other OSD nodes, read data according to the PG data in the target PG mapped by the master OSD node in each OSD node in operation if a second data reading request is received.
[0115] Optionally, the synchronization module 503 is specifically used to initiate a query message to the replica OSD nodes in the OSD cluster through the master OSD node in the OSD cluster; receive a response message returned by each replica OSD node based on the query message through the master OSD node; wherein the response message returned by each replica OSD node includes the PG version number of the target PG mapped by the replica OSD node; and select the authoritative OSD node based on the response message returned by each replica OSD node.
[0116] Optionally, the authoritative OSD node is the OSD node with the largest PG version number of the PG mapped in each OSD node; or, the authoritative OSD node is the OSD node with the largest PG version number of the PG mapped in each OSD node and the longest PG data in the mapped PG.
[0117] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the data reading method disclosed in the above example of the present application can be implemented.
[0118] The above-mentioned machine-readable storage medium may be any electronic, magnetic, optical or other physical storage device, which may contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as CD, DVD, etc.), or similar storage medium, or a combination thereof.
[0119] This manual also provides Figure 6 The hardware structure diagram of an electronic device is shown in FIG. Figure 6 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 2 The data reading method shown.
[0120] Based on the same application concept as the above method, an embodiment of the present application also provides a computer program product, including a computer program, which implements the above data reading method when executed by a processor.
[0121] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0122] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A data reading method, characterized in that: include: Receive a first data read request, determine a target placement group PG according to the first data read request, and when it is determined that the object storage device OSD cluster mapped with the target PG meets the single copy read condition, perform data reading according to the PG data in the target PG mapped by the single copy OSD node in the OSD cluster; wherein, if the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, update the PG version number of the target PG, and record the updated PG version number of the target PG in the single copy OSD node; the OSD cluster meeting the single copy read condition means that there is a single copy OSD node in the OSD cluster; the updated PG version number is the maximum version number of the target PG mapped by each OSD node in the OSD cluster; When it is detected that other OSD nodes in the OSD cluster switch from a non-operating state to a running state, an authoritative OSD node is selected based on the PG version number of the PG recorded by each OSD node in the OSD cluster, wherein the PG version number of the target PG mapped by the selected authoritative OSD node is the updated PG version number; Synchronize the data in the authoritative OSD node to the other OSD nodes.
2. The method according to claim 1, characterized in that If the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, the method further includes: At a preset time interval, re-judge whether the OSD cluster meets the single copy reading condition, and re-judge whether the PG version number indicated by the first data read request is consistent with the PG version number of the target PG, until a preset judgment number threshold is reached; The updating of the PG version number of the target PG specifically includes: If the OSD cluster still meets the single copy reading condition after reaching the preset judgment number threshold, and the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, the PG version number of the target PG is updated.
3. The method according to claim 2, characterized in that After re-determining whether the OSD cluster meets the single copy reading condition, the method further includes: If the OSD cluster currently has multiple OSD nodes in a running state, then determining an authoritative OSD node from the multiple OSD nodes in a running state; Synchronize data in the authoritative OSD node determined among the multiple OSD nodes in operation to other OSD nodes currently in operation; Data is read according to the PG data in the target PG mapped by the master OSD node in other OSD nodes that are currently in operation.
4. The method according to any one of claims 1 or 3, characterized in that: The method further comprises: After synchronizing the data in the authoritative OSD node to the other OSD nodes, if a second data read request is received, data reading is performed based on the PG data in the target PG mapped by the master OSD node in each OSD node in operation.
5. The method according to claim 1, characterized in that The selecting an authoritative OSD node based on the PG version number of the target PG recorded by each OSD node in the OSD cluster specifically includes: Initiate a query message to the replica OSD node in the OSD cluster through the master OSD node in the OSD cluster; Receiving, through the master OSD node, a response message returned by each replica OSD node according to the query message; wherein the response message returned by each replica OSD node includes a PG version number of a target PG mapped by the replica OSD node; The authoritative OSD node is selected based on the response message returned by each replica OSD node.
6. The method according to claim 5, characterized in that The authoritative OSD node is the OSD node with the largest PG version number of the PG mapped among the OSD nodes; or, The authoritative OSD node is the OSD node with the largest PG version number of the PG mapped among the OSD nodes and the longest PG data in the mapped PG.
7. A data reading device, characterized in that: include: A reading module, used to receive a first data read request, determine a target placement group PG according to the first data read request, and when it is determined that the object storage device OSD cluster mapped with the target PG meets the single copy read condition, read data according to the PG data in the target PG mapped by the single copy OSD node in the OSD cluster; wherein, if the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG, the PG version number of the target PG is updated, and the updated PG version number of the target PG is recorded in the single copy OSD node; the OSD cluster meeting the single copy read condition means that there is a single copy OSD node in the OSD cluster; the updated PG version number is the maximum version number of the target PG mapped by each OSD node in the OSD cluster; A selection module is used to select an authoritative OSD node based on the PG version number of the PG recorded by each OSD node in the OSD cluster when detecting that other OSD nodes in the OSD cluster switch from a non-operating state to a running state, wherein the PG version number of the target PG mapped by the selected authoritative OSD node is the updated PG version number; A synchronization module is used to synchronize the data in the authoritative OSD node to the other OSD nodes.
8. The data reading device according to claim 7, characterized in that: The device also includes: A judgment module is used to re-judge whether the OSD cluster meets the single copy reading condition according to a preset time interval, and re-judge whether the PG version number indicated by the first data read request is consistent with the PG version number of the target PG, until a preset judgment number threshold is reached; The reading module is specifically used to update the PG version number of the target PG if the OSD cluster still meets the single copy reading condition after reaching a preset judgment number threshold, and the PG version number indicated by the first data read request is inconsistent with the PG version number of the target PG.
9. The data reading device according to claim 8, characterized in that: After determining whether the OSD cluster currently has a single OSD node in operation, the synchronization module is further configured to, if the OSD cluster currently has multiple OSD nodes in operation, determine an authoritative OSD node from the multiple OSD nodes in operation; Synchronize data in the authoritative OSD node determined among the multiple OSD nodes in operation to other OSD nodes currently in operation; Data is read according to the PG data in the target PG mapped by the master OSD node in other OSD nodes that are currently in operation.
10. The data reading device according to any one of claims 7 or 9, characterized in that: The reading module is further used to, after synchronizing the data in the authoritative OSD node to the other OSD nodes, read data according to the PG data in the target PG mapped by the master OSD node in each OSD node in operation if a second data reading request is received.
11. The data reading device according to claim 7, characterized in that: The synchronization module is specifically used to initiate a query message to the replica OSD nodes in the OSD cluster through the master OSD node in the OSD cluster; receive a response message returned by each replica OSD node based on the query message through the master OSD node; wherein the response message returned by each replica OSD node contains the PG version number of the target PG mapped by the replica OSD node; and select the authoritative OSD node based on the response message returned by each replica OSD node.
12. The data reading device according to claim 11, characterized in that: The authoritative OSD node is the OSD node with the largest PG version number of the PG mapped among the OSD nodes; or, the authoritative OSD node is the OSD node with the largest PG version number of the PG mapped among the OSD nodes and the longest PG data in the mapped PG.