OSD switching method, system and device in distributed storage pool and medium
By identifying faulty OSDs and their physical fault domains in a distributed storage pool, obtaining the number of global and failed fault domains, and controlling the switching of faulty OSDs, the issues of service reliability and data security during OSD fault switching are resolved, achieving higher system reliability and data security.
Patent Information
- Application Number
- CN202511072501.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
In distributed storage architectures, existing technologies struggle to guarantee service reliability and data security during OSD failover, especially in failover scenarios where situations can easily occur outside the fault domain, leading to service interruption and data loss.
By identifying the faulty OSD and its physical fault domain in the distributed storage pool, the number of globally allowed fault domains and the number of already failed fault domains are obtained. The faulty OSD is marked as down only when the number of already failed fault domains is less than the number of globally allowed fault domains, ensuring that the number of allowed fault domain failures is not exceeded during the handover process.
It improves the service reliability and data security of distributed storage in failover scenarios, avoids service interruption and data loss due to fault domain overflow, and enhances the overall reliability and data durability of the system.
Smart Images

Figure CN120909519A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of distributed storage, and more particularly to an OSD switching method, system, device and medium in a distributed storage pool. BACKGROUND
[0002] In a distributed storage architecture, physical storage resources are organized and managed through logical storage pools. Specifically, a plurality of server nodes deploying distributed storage software form a cluster, each node containing a number of disks, and a storage pool integrates node resources as a virtual container, manages disks through an object storage daemon (OSD), such as a single disk corresponding to a single OSD, and implements data redundancy based on a placement group (PG) mechanism.
[0003] When an OSD service is unexpectedly stopped, such as server downtime, OSD process crash, or severe network isolation, the faulty OSD must be promptly marked as down so that the cluster can quickly switch the PG primary and secondary roles on the OSD to other active OSDs, maintaining the continuity of client read and write services. However, after detecting a faulty OSD and marking it as down, the situation may exceed the failure domain, reducing the service reliability and data security of distributed storage in a failure switching scenario.
[0004] In summary, how to improve the service reliability and data security of distributed storage in a failure switching scenario is a problem that needs to be solved by those skilled in the art. SUMMARY
[0005] The purpose of the present application is to provide an OSD switching method in a distributed storage pool, which can solve the technical problem of how to improve the service reliability and data security of distributed storage in a failure switching scenario to some extent. The present application also provides an OSD switching system in a distributed storage pool, an electronic device and a computer readable storage medium.
[0006] To achieve the above purpose, the present application provides the following technical solutions:
[0007] An OSD switching method in a distributed storage pool, comprising:
[0008] Determining a faulty OSD in a distributed storage pool;
[0009] Determining the physical failure domain of the target storage pool in which the faulty OSD is located;
[0010] Obtaining the global number of allowed failure domains of the physical failure domain;
[0011] determining a failed failure domain number of the physical failure domain;
[0012] in response to the failed failure domain number being less than the global allowed failed failure domain number, marking a state of the failed OSD as down;
[0013] the global allowed failed failure domain number comprises a minimum value of allowed failure domains of a target storage pool set, the target storage pool set comprising a set of all storage pool combinations sharing the physical failure domain.
[0014] In an example embodiment, the obtaining the global allowed failed failure domain number of the physical failure domain comprises:
[0015] In the distributed storage, screening other storage pools sharing the physical failure domain with the target storage pool;
[0016] combining the target storage pool and the other storage pools into the target storage pool set;
[0017] generating, according to a redundancy rule, a failed failure domain number allowed by each storage pool in the target storage pool set;
[0018] taking the failed failure domain number with the minimum value as the global allowed failed failure domain number of the physical failure domain.
[0019] In an example embodiment, the generating, according to a redundancy rule, a failed failure domain number allowed by each storage pool in the target storage pool set comprises:
[0020] for each storage pool in the target storage pool set, if the storage pool is a replica pool, generating a difference value between a replica number and a minimum safe replica number, and taking the difference value as the failed failure domain number allowed by the storage pool; if the storage pool is an erasure code pool, taking a check block number as the failed failure domain number allowed by the storage pool.
[0021] In an example embodiment, the determining a failed failure domain number of the physical failure domain comprises:
[0022] In the physical failure domain, determining other OSDs other than the failed OSD;
[0023] traversing states of the other OSDs, and each time an other OSD in a down or out state is traversed, determining that a failure domain to which the other OSD belongs has failed;
[0024] counting a number of the failed failure domains to obtain the failed failure domain number of the physical failure domain.
[0025] In an example embodiment, the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domains further comprises:
[0026] generating a difference between the global allowed failed failure domains and the number of failed failure domains, and using the difference as a failed failure domain failure remaining quota;
[0027] detecting whether the failure domain to which the failed OSD belongs is a failed failure domain;
[0028] in response to the failure domain to which the failed OSD belongs being a failed failure domain, marking the status of the failed OSD as down and keeping the failed failure domain failure remaining quota unchanged;
[0029] in response to the failure domain to which the failed OSD belongs not being a failed failure domain, detecting whether the value of the failed failure domain failure remaining quota is greater than zero;
[0030] in response to the value of the failed failure domain failure remaining quota being greater than zero, marking the status of the failed OSD as down.
[0031] In an example embodiment, the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domains further comprises:
[0032] determining a minimum safe read-write replica number of the target storage pool configuration;
[0033] determining a target PG carried by the failed OSD;
[0034] determining a real-time replica number of a primary replica or a read-write member participating replica in the target PG;
[0035] detecting whether the real-time replica number is greater than the minimum safe read-write replica number;
[0036] if the real-time replica number is less than or equal to the minimum safe read-write replica number, prohibiting switching of the failed OSD;
[0037] if the real-time replica number is greater than the minimum safe read-write replica number, performing the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domains.
[0038] In an example embodiment, the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domains further comprises:
[0039] detecting whether the failed OSD is marked as fast switch prohibited;
[0040] if the failed OSD is marked as fast switch prohibited, prohibiting switching to the failed OSD;
[0041] if the failed OSD is not marked as fast switch prohibited, performing the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global number of allowed failed failure domains.
[0042] An OSD switching system in a distributed storage pool, comprising:
[0043] a failed OSD determining module configured to determine a failed OSD in the distributed storage pool;
[0044] a physical failure domain determining module configured to determine a physical failure domain of a target storage pool in which the failed OSD is located;
[0045] a global number of failed failure domains obtaining module configured to obtain a global number of allowed failed failure domains of the physical failure domain;
[0046] a number of failed failure domains determining module configured to determine a number of failed failure domains of the physical failure domain;
[0047] an OSD switching module configured to mark the status of the failed OSD as down in response to the number of failed failure domains being less than the global number of allowed failed failure domains;
[0048] the global number of allowed failed failure domains comprises a minimum value of a number of allowed failed failure domains of a target storage pool set, the target storage pool set comprising a set of all storage pool combinations sharing the physical failure domain.
[0049] An electronic device, comprising:
[0050] a memory configured to store a computer program;
[0051] a processor configured to implement the steps of the OSD switching method in a distributed storage pool according to any one of the above embodiments when the computer program is executed.
[0052] A computer readable storage medium, the computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps of the OSD switching method in a distributed storage pool according to any one of the above embodiments.
[0053] The application provides an OSD switching method in a distributed storage pool. A faulty OSD is determined in the distributed storage pool. A physical failure domain of a target storage pool where the faulty OSD is located is determined. A global allowed failure domain number of the physical failure domain is obtained. A failed failure domain number of the physical failure domain is determined. In response to the failed failure domain number being less than the global allowed failure domain number, the state of the faulty OSD is marked as down. The global allowed failure domain number is the minimum value of allowed failure domains of a target storage pool set, and the target storage pool set comprises a set of all storage pool combinations sharing the physical failure domain. In the application, the global allowed failure domain number is the minimum value of allowed failure domains of the target storage pool set, and the target storage pool set comprises a set of all storage pool combinations sharing the physical failure domain of the target storage pool. Therefore, the global allowed failure domain number represents the minimum failure domain number allowed when the physical failure domain is shared by all storage pools. In this way, the state of the faulty OSD is marked as down only when the failed failure domain number is less than the global allowed failure domain number, so that the latest failed failure domain number is at most equal to the minimum failure domain number, thereby avoiding the latest failed failure domain number exceeding the minimum failure domain number, and improving the service reliability and data security of the distributed storage in a failure switching scenario. The OSD switching system in the distributed storage pool, the electronic device and the computer readable storage medium provided by the application also solve the corresponding technical problems. BRIEF DESCRIPTION OF DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of the provided drawings.
[0055] Figure 1 A flowchart of the OSD switching method in the distributed storage pool provided by the embodiments of the present application;
[0056] Figure 2 A determination diagram of the global allowed failure domain number;
[0057] Figure 3 A switching diagram of the faulty OSD;
[0058] Figure 4 A diagram of the Monitor processing the faulty OSD;
[0059] Figure 5 A structure diagram of the OSD switching system in the distributed storage pool provided by the embodiments of the present application;
[0060] Figure 6 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1.
[0061] Figure 7 Another structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION
[0062] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort belong to the scope of protection of the present application.
[0063] Please refer to Figure 1 , Figure 1 A flowchart of an OSD switching method in a distributed storage pool provided by an embodiment of the present application is shown in FIG. 3.
[0064] The OSD switching method in the distributed storage pool provided by an embodiment of the present application can include the following steps.
[0065] Step S101: determining a faulty OSD in the distributed storage pool.
[0066] In actual application, the faulty OSD can be determined in the distributed storage pool first, so as to switch the faulty OSD according to the present application. The faulty OSD can be a faulty OSD detected by external fast fault detection, such as CTDB, or a faulty OSD detected by an OSD heartbeat timeout mechanism, etc.
[0067] It should be noted that the OSD heartbeat timeout mechanism relies on the heartbeat messages exchanged periodically between OSD nodes. If a certain OSD fails to respond to the heartbeat request of its neighbor OSD within a preset timeout period (60s), the neighbor OSD will identify this OSD as a faulty OSD. The external fast fault detection introduces an independent cluster fault detection service, which identifies the faulty OSD through methods such as lightweight heartbeat between nodes, network link detection, etc.
[0068] Step S102: determining a physical failure domain of a target storage pool where the faulty OSD is located.
[0069] Step S103: obtaining a global number of allowed failure failure domains of the physical failure domain; the global number of allowed failure failure domains includes the minimum value of the allowed failure failure domains of a target storage pool set, and the target storage pool set includes a set of all storage pool combinations sharing the physical failure domain.
[0070] In actual applications, each OSD belongs to a storage pool, and the physical failure domains can be shared between the storage pools, and the number of allowed failure domains set between the storage pools can be different. Thus, after a faulty OSD in a storage pool is down, the actual number of failure domains of other storage pools sharing the physical failure domain can exceed the number of allowed failure domains set. In order to avoid this situation, the physical failure domain of the target storage pool in which the faulty OSD is located can be determined. The physical failure domain (Fault Domain) is a core fault-tolerant and high-availability design concept, which refers to a group of storage nodes or components that can fail simultaneously due to the same underlying hardware, software, network, or environment failure. The set of storage pools sharing the physical failure domain with the target storage pool is determined, and the minimum value of the number of allowed failure domains of the storage pools in the set is taken as the global number of allowed failure domains, so as to determine whether to switch the faulty OSD later by means of the global number of allowed failure domains.
[0071] In an example embodiment, according to the definition of the global number of allowed failure domains, in the process of obtaining the global number of allowed failure domains of the physical failure domain, as shown in FIG. 1, the other storage pools sharing the physical failure domain with the target storage pool can be filtered out in the distributed storage; the target storage pool and the other storage pools are combined into a target storage pool set; according to the redundancy rule, the number of allowed failure domains of each storage pool in the target storage pool set is generated, and the number of allowed failure domains can be the maximum number of allowed failure domains of the storage pool. Figure 2 For example, the number of allowed failure domains of storage pool A is 3, the number of allowed failure domains of storage pool B is 5, and the number of allowed failure domains of storage pool C is 8. Thus, the global number of allowed failure domains is 3. In other words, the global number of allowed failure domains represents how many new failure domains the failure domain level can withstand under the premise of meeting the most stringent redundancy requirements of all associated storage pools.
[0072] In a specific application scenario, due to different redundancy rules, the types of storage pools can be different, that is, the types of storage pools correspond to the redundancy rules. Therefore, in the process of generating the number of allowed failure domains of each storage pool in the target storage pool set according to the redundancy rule, the number of allowed failure domains of the storage pool can be determined according to the type of the storage pool, that is, for each storage pool in the target storage pool set, if the storage pool is a replica pool, the difference between the number of replicas and the minimum number of safe replicas is generated, and the difference is taken as the number of allowed failure domains of the storage pool; if the storage pool is an erasure code pool, the number of check blocks is taken as the number of allowed failure domains of the storage pool.
[0073] Step S104: Determine the number of failed failure domains of the physical failure domain.
[0074] Step S105: in response to the number of failed failure domains being less than the global allowed failed failure domains, marking the status of the failed OSD as down.
[0075] In actual application, the number of failed failure domains of the physical failure domain can be determined after the global allowed failed failure domains of the physical failure domain is acquired. Since the status of the failed OSD is marked as down, the value of the number of failed failure domains will increase. If the value of the number of failed failure domains is greater than the global allowed failed failure domains, the number of failed failure domains of a storage pool will exceed the set value. Therefore, the status of the failed OSD needs to be marked as down only when the number of failed failure domains is less than the global allowed failed failure domains. Since the status of the failed OSD is marked as down only when the number of failed failure domains is less than the global allowed failed failure domains, the value of the number of failed failure domains will at most increase to equal the global allowed failed failure domains, and the number of failed failure domains will not exceed the global allowed failed failure domains. Accordingly, in response to the number of failed failure domains being equal to or greater than the global allowed failed failure domains, the status of the failed OSD is prohibited from being marked as down. In this way, high-risk switching operations that may cause violation of the failure domain constraint are effectively filtered out, thereby significantly improving the overall service reliability and data security of the distributed storage system in the fast failure switching scenario.
[0076] In an exemplary embodiment, in the process of determining the number of failed failure domains of the physical failure domain, if the status of an OSD is down or out, the failure domain to which the OSD belongs has failed. Therefore, the number of failed failure domains can be screened according to the status of the OSD, that is, in the physical failure domain, other OSDs except the failed OSD are determined. The status of the other OSDs is traversed. If each other OSD is in the down or out state, it is determined that the failure domain to which the other OSD belongs has failed, until the traversal ends. The number of failed failure domains is counted to obtain the number of failed failure domains of the physical failure domain.
[0077] In the example embodiment, considering that the difference between the global allowed failed failure domain number and the failed failure domain number represents the remaining failed number that the physical failure domain can allow, in the process of marking the state of the failed OSD as down in response to the failed failure domain number being less than the global allowed failed failure domain number, the difference between the global allowed failed failure domain number and the failed failure domain number can be generated, and the difference is taken as the failure domain failed remaining quota; it is detected whether the failure domain to which the failed OSD belongs belongs to the failed failure domain; in response to the failure domain to which the failed OSD belongs belonging to the failed failure domain, the state of the failed OSD is marked as down, and the failure domain failed remaining quota is kept unchanged, because the failure domain to which the failed OSD belongs has been deducted, and after the failed OSD is marked as down, no new consumption quota is added, so that the latest failed failure domain number does not increase to exceed the global allowed failed failure domain number; in response to the failure domain to which the failed OSD belongs not belonging to the failed failure domain, it is detected whether the value of the failure domain failed remaining quota is greater than zero; in response to the value of the failure domain failed remaining quota being greater than zero, the state of the failed OSD is marked as down.
[0078] In the example embodiment, considering that the switching of the failed OSD will affect the normal copy number of the corresponding PG, if the normal copy number is too low, it will cause the PG to be abnormal, that is, the switching of the failed OSD will affect the normality of the PG, in order to avoid that the switching of the failed OSD causes the PG to be abnormal, before the state of the failed OSD is marked as down in response to the failed failure domain number being less than the global allowed failed failure domain number, the minimum safe read-write copy number configured by the target storage pool can also be determined; the target PG carried by the failed OSD is determined; the real-time copy number of the primary copy or the read-write member participating in the target PG is determined; it is detected whether the real-time copy number is greater than the minimum safe read-write copy number; if the real-time copy number is less than or equal to the minimum safe read-write copy number, the switching of the failed OSD is prohibited; if the real-time copy number is greater than the minimum safe read-write copy number, the step of marking the state of the failed OSD as down in response to the failed failure domain number being less than the global allowed failed failure domain number is executed.
[0079] In the example embodiment, the switching of the failed OSD can also be controlled by the user, in order to meet the needs of the user, before the state of the failed OSD is marked as down in response to the failed failure domain number being less than the global allowed failed failure domain number, it is also detected whether the failed OSD is marked as prohibited from fast switching; if the failed OSD is marked as prohibited from fast switching, the switching of the failed OSD is prohibited; if the failed OSD is not marked as prohibited from fast switching, the step of marking the state of the failed OSD as down in response to the failed failure domain number being less than the global allowed failed failure domain number is executed.
[0080] It should be noted that the three schemes in the present application for determining whether to switch the failed OSD according to the remaining failure quota of the failure domain, according to the number of primary replicas or member replicas participating in read and write in the target PG, and according to the fast switching prohibition flag can be applied alone or in combination with two or three of them, wherein the priority of determining whether to switch the failed OSD according to the number of primary replicas or member replicas participating in read and write in the target PG can be set to be the highest, the priority of determining whether to switch the failed OSD according to the fast switching prohibition flag can be set to be the second, and the priority of determining whether to switch the failed OSD according to the remaining failure quota of the failure domain can be set to be the lowest, which is not specifically limited in the present application.
[0081] In the exemplary embodiment, the switching process of the OSD can be as shown in Figure 3 The w represents the remaining failure quota of the failure domain, including the following processes: determining the minimum safe read and write replica number of the target storage pool configuration; determining the target PG carried by the failed OSD; determining the real-time replica number of the primary replica or the member participating in read and write in the target PG; detecting whether the real-time replica number is greater than the minimum safe read and write replica number; if the real-time replica number is less than or equal to the minimum safe read and write replica number, the switching of the failed OSD is prohibited; if the real-time replica number is greater than the minimum safe read and write replica number, the difference between the global allowed failure domain number and the failed failure domain number is generated, and the difference is taken as the remaining failure quota of the failure domain; detecting whether the failure domain to which the failed OSD belongs belongs to the failed failure domain; in response to the failure domain to which the failed OSD belongs belonging to the failed failure domain, the state of the failed OSD is marked as down, and the remaining failure quota of the failure domain is kept unchanged; in response to the failure domain to which the failed OSD belongs not belonging to the failed failure domain, it is detected whether the value of the remaining failure quota of the failure domain is greater than zero; in response to the value of the remaining failure quota of the failure domain being greater than zero, the state of the failed OSD is marked as down, and in response to the value of the remaining failure quota of the failure domain being less than zero, the switching of the failed OSD is prohibited.
[0082] In this way, in the process of switching the faulty OSD, the application can first detect whether the real-time number of replicas in the PG is greater than the minimum safe read-write replica number. If the real-time number of replicas is less than or equal to the minimum safe read-write replica number, switching the faulty OSD is prohibited, which can avoid the situation that switching the faulty OSD causes the target PG to be abnormal. If the real-time number of replicas is greater than the minimum safe read-write replica number, in the case that the faulty domain to which the faulty OSD belongs belongs to the failed faulty domain, the state of the faulty OSD is marked as down, which can avoid the situation that switching the faulty OSDs in the same failed faulty domain causes the number of failed faulty domains to increase in theory but not in fact, ensures the real effectiveness of the number of failed faulty domains, and thus ensures the accurate switching of the faulty OSD. In the case that the faulty domain to which the faulty OSD belongs does not belong to the failed faulty domain, whether the value of the failed faulty domain failure remaining amount is greater than zero is detected. In response to the value of the failed faulty domain failure remaining amount being greater than zero, the state of the faulty OSD is marked as down. In response to the value of the failed faulty domain failure remaining amount being less than zero, switching the faulty OSD is prohibited. Thus, the faulty OSD can be switched in a real and effective manner without exceeding the number of time-out faulty domains, and the robustness of the switching of the faulty OSD is improved.
[0083] It should be further noted that after the faulty OSD is marked as down, the corresponding PG master replica switching, i.e., the PG peering process, can be triggered. At the same time, an alarm information with a "fast switching trigger" identifier can be generated and reported to the management software when the faulty OSD is switched quickly, and the alarm content can include the affected abnormal node IP, the downed OSD list, the timestamp, and other key information, which facilitates subsequent auditing and fault root cause analysis.
[0084] This application provides a method for OSD switching in a distributed storage pool, which involves: identifying a faulty OSD in the distributed storage pool; determining the physical fault domain of the target storage pool where the faulty OSD is located; obtaining the globally allowed number of fault domains in the physical fault domain; determining the number of already failed fault domains in the physical fault domain; and, in response to the fact that the number of already failed fault domains is less than the globally allowed number of fault domains, marking the state of the faulty OSD as down. The globally allowed number of fault domains includes the minimum value of the set of target storage pools that allows fault domain failures, and the set of target storage pools includes the set of all storage pool combinations that share a physical fault domain. In this application, the globally allowed number of fault domains is the minimum number of fault domains allowed to fail in the target storage pool set. The target storage pool set includes the set of all storage pool combinations that share physical fault domains. Therefore, the globally allowed number of fault domains represents the minimum number of fault domains allowed to fail when physical fault domains are shared by all storage pools. Thus, if the state of a faulty OSD is marked as down only when the number of failed fault domains is less than the globally allowed number of fault domains, the latest number of failed fault domains will at most be equal to the minimum number of failed fault domains. This avoids the situation where the latest number of failed fault domains exceeds the minimum number of failed fault domains, thereby improving the service reliability and data security of distributed storage in failover scenarios.
[0085] It should be noted that the OSD switching method in the distributed storage pool of this application can also be used during the switching process between CTDB and the heartbeat timeout mechanism, such as... Figure 4 As shown, the following processes may be included:
[0086] When CTDB detects node anomalies, such as network interruptions or system crashes, it sends a quick switch command containing a list of abnormal node IPs to the monitoring service (Monitor).
[0087] Based on this instruction, the Monitor queries the current OSD map of the cluster and collects all storage pool information, including: storage pool ID, configured fault domain topology (such as rack level), and a list of OSD IDs corresponding to the IP of each node under that topology;
[0088] Monitor identifies all storage pools that share the same physical fault domain definition with the storage pool, such as all storage pools that use the same rack as the fault domain boundary, forming a "shared fault domain storage pool set"; it iterates through each storage pool in the storage pool set and calculates the number of allowed fault domain failures based on its redundancy rules.
[0089] The Monitor takes the minimum number of allowed failure domains calculated by all storage pools in the set as the global allowed failure domain number w under the physical failure domain topology;
[0090] Monitor traverses all node IPs except the abnormal IP under the current fault domain topology, checks the state of each OSD under it: if a certain OSD is found to be in the down state but not out, it is considered that the physical failure domain to which the OSD belongs has failed, and the information of these failed failure domains is recorded, and the previously calculated global allowed failure failure domain number N is reduced by 1;
[0091] For each abnormal IP, Monitor determines the physical failure domain to which it belongs, and checks and decides each OSD corresponding to the abnormal IP under the failure domain: if the OSD current state is down or out, skip processing; if the OSD is manually marked by the administrator as prohibited from fast switching, or the number of member replicas in the PG it carries, the primary and secondary replicas or the number of read-write member replicas, is less than or equal to the minimum safe read-write replica number (usually min_size) configured by the storage pool, the OSD is prohibited from performing this fast switching operation; if the failure domain to which the OSD belongs has been recorded as failed, the OSD can be safely added to the allowed down list; if the failure domain is not recorded as failed, that is, it is newly failed, and the current global allowed failure failure domain number N>0, that is, there is still a quota available, the OSD is allowed to be added to the allowed down list, and the failure domain is marked as "this newly failed domain", and N is reduced by 1.
[0092] For the OSD on the abnormal node that fails to enter the allowed down list, Monitor does not perform fast switching, but uses the traditional inter-OSD heartbeat timeout mechanism to handle it.
[0093] In this way, by introducing a dynamic evaluation mechanism of global failure domain safety quota (w) and real-time PG state checking, after the external failure detection triggers the fast switching instruction, the OSD subset allowed to perform the down operation is accurately selected; at the same time, a decision model based on the existing failure domain quota occupation and the newly added domain quota consumption is constructed, and the high-risk OSD is automatically downgraded to the traditional heartbeat timeout process. This scheme completely inherits the millisecond-level response capability of external services, and basically eliminates the systematic risk of "replica aggregation exceeding the failure domain" that may be caused by the fast switching mechanism, strictly guarantees the dynamic compliance of the cross-failure domain redundancy strategy, and significantly improves the data persistence and business high-availability protection of the distributed storage in extreme scenarios such as machine room-level failures and network partitioning.
[0094] Please refer to Figure 5 , Figure 5 A structure diagram of an OSD switching system in a distributed storage pool provided by an embodiment of the present application.
[0095] An OSD switching system in a distributed storage pool provided by an embodiment of the present application can include:
[0096] The fault OSD determination module 101 is configured to determine a fault OSD in a distributed storage pool.
[0097] The physical fault domain determination module 102 is configured to determine a physical fault domain of a target storage pool in which the fault OSD is located.
[0098] The global failure number acquisition module 103 is configured to acquire a global allowed failure fault domain number of the physical fault domain.
[0099] The failed number determination module 104 is configured to determine a failed fault domain number of the physical fault domain.
[0100] The OSD switching module 105 is configured to, in response to the failed fault domain number being less than the global allowed failure fault domain number, mark a state of the fault OSD as down.
[0101] The global allowed failure fault domain number includes a minimum value of allowed failure fault domains of a target storage pool set, and the target storage pool set includes a set of all storage pool combinations sharing the physical fault domain.
[0102] The OSD switching system in the distributed storage pool provided by the embodiment of the present application includes a global failure number acquisition module.
[0103] The storage pool screening unit is configured to screen other storage pools sharing a physical fault domain with a target storage pool in a distributed storage.
[0104] The storage pool combination unit is configured to combine the target storage pool and the other storage pools into a target storage pool set.
[0105] The failure number generation unit is configured to generate an allowed failure fault domain number of each storage pool in the target storage pool set according to a redundancy rule.
[0106] The global failure number acquisition unit is configured to take a minimum value of the failure fault domain numbers as a global allowed failure fault domain number of the physical fault domain.
[0107] The failure number generation unit can be configured to, for each storage pool in the target storage pool set, if the storage pool is a replica pool, generate a difference value between a replica number and a minimum safe replica number, and take the difference value as the allowed failure fault domain number of the storage pool; and if the storage pool is an erasure code pool, take a check block number as the allowed failure fault domain number of the storage pool.
[0108] The failed number determination module can include an OSD determination unit.
[0109] The OSD determination unit is configured to determine other OSDs in the physical fault domain except the fault OSD.
[0110] traverse units for traversing states of other OSDs, and determining that a failure domain to which an other OSD belongs has failed when each other OSD is in a down or out state;
[0111] statistical units for counting the number of failed failure domains to obtain the number of failed failure domains of the physical failure domain.
[0112] The OSD switching system provided by the embodiment of the application can comprise:
[0113] remaining quota generation units for generating a difference between the global allowed failed failure domain number and the failed failure domain number, and taking the difference as a failure domain failure remaining quota;
[0114] failure domain detection units for detecting whether a failure domain to which a failed OSD belongs belongs to a failed failure domain; in response to the failure domain to which the failed OSD belongs belonging to the failed failure domain, marking a state of the failed OSD as down and keeping the failure domain failure remaining quota unchanged; in response to the failure domain to which the failed OSD belongs not belonging to the failed failure domain, detecting whether a value of the failure domain failure remaining quota is greater than zero; and in response to the value of the failure domain failure remaining quota being greater than zero, marking the state of the failed OSD as down.
[0115] The OSD switching system provided by the embodiment of the application can further comprise:
[0116] minimum replica number determination units for determining a minimum safe read-write replica number configured by a target storage pool before the OSD switching module marks the state of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domain number.
[0117] PG determination units for determining a target PG borne by the failed OSD;
[0118] real-time replica number determination units for determining a real-time replica number of a primary replica or a read-write member participating in the target PG;
[0119] replica number detection units for detecting whether the real-time replica number is greater than the minimum safe read-write replica number; if the real-time replica number is less than or equal to the minimum safe read-write replica number, prohibiting switching of the failed OSD; and if the real-time replica number is greater than the minimum safe read-write replica number, performing the step of marking the state of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domain number.
[0120] The OSD switching system provided by the embodiment of the application can further comprise:
[0121] The switching detection module is configured to, before the OSD switching module marks the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed number of failed failure domains, detect whether the failed OSD is marked as fast switching prohibited; if the failed OSD is marked as fast switching prohibited, switching of the failed OSD is prohibited; and if the failed OSD is not marked as fast switching prohibited, the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed number of failed failure domains is performed.
[0122] The application further provides an electronic device and a computer readable storage medium, both of which have the corresponding effects of the OSD switching method in the distributed storage pool provided by the embodiments of the application. Please refer to Figure 6 , Figure 6 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the application.
[0123] The electronic device provided by the embodiment of the application comprises a memory 201 and a processor 202, the memory 201 stores a computer program, and the processor 202 implements the steps of the OSD switching method in the distributed storage pool described in any of the above embodiments when executing the computer program.
[0124] Please refer to Figure 7 , the other electronic device provided by the embodiment of the application can further comprise: an input port 203 connected to the processor 202, configured to transmit a command input by the outside to the processor 202; a display unit 204 connected to the processor 202, configured to display the processing result of the processor 202 to the outside; and a communication module 205 connected to the processor 202, configured to realize communication between the electronic device and the outside. The display unit 204 can be a display panel, a laser scanning display, etc.; the communication mode adopted by the communication module 205 includes but is not limited to Mobile High-Definition Link (MHL), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connection: WIreless Fidelity (WiFi), Bluetooth communication technology, low-power Bluetooth communication technology, and IEEE 802.11s-based communication technology.
[0125] The computer readable storage medium provided by the embodiment of the application stores a computer program, and the computer program is executed by a processor to implement the steps of the OSD switching method in the distributed storage pool described in any of the above embodiments.
[0126] The computer readable storage medium involved in the present application includes random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art.
[0127] The computer program product provided by the embodiment of the present application comprises computer programs / instructions, which, when executed by a processor, implement the steps of the OSD switching method in the distributed storage pool as described in any of the above embodiments.
[0128] The related parts of the OSD switching system, electronic device, computer readable storage medium and computer program product provided by the embodiment of the present application are described in detail in the corresponding part of the OSD switching method provided by the embodiment of the present application, and will not be described here. In addition, the parts of the above technical solutions provided by the embodiment of the present application which are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail, so as not to be too verbose.
[0129] It should also be noted that, in this document, relational terms such as first and second, and the like, are used solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the stated elements.
[0130] The above description of disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for OSD switching in a distributed storage pool, the method comprising: The method comprises the following steps: determining a failed OSD in a distributed storage pool; determining a target storage pool of the failed OSD; determining a physical failure domain of the target storage pool; acquiring a global allowed failed failure domain number of the physical failure domain; determining a failed failure domain number of the physical failure domain; in response to the failed failure domain number being less than the global allowed failed failure domain number, marking a state of the failed OSD as down.
2. The method of claim 1, wherein, The global allowed failed failure domain number comprises a minimum value of allowed failure domain failure of a target storage pool set, the target storage pool set comprising a set of all storage pool combinations sharing the physical failure domain. The acquiring the global allowed failed failure domain number of the physical failure domain comprises: filtering other storage pools sharing the physical failure domain with the target storage pool in the distributed storage; combining the target storage pool and the other storage pools into the target storage pool set; generating an allowed failure domain failure number of each storage pool in the target storage pool set according to a redundancy rule; 3. The method of distributed storage pool OSD switchover according to claim 2, wherein, taking the minimum value of the failure domain failure number as the global allowed failed failure domain number of the physical failure domain. The generating the allowed failure domain failure number of each storage pool in the target storage pool set according to the redundancy rule comprises:
4. The method of distributed storage pool OSD switchover according to claim 1, wherein, for each storage pool in the target storage pool set, if the storage pool is a replica pool, generating a difference value between a replica number and a minimum safe replica number, and taking the difference value as the allowed failure domain failure number of the storage pool; if the storage pool is an erasure code pool, taking a check block number as the allowed failure domain failure number of the storage pool. The determining the failed failure domain number of the physical failure domain comprises: determining other OSDs in the physical failure domain except the failed OSD; traversing states of the other OSDs, and determining that a failure domain to which each other OSD belongs has failed if each other OSD is in a down or out state; 5. The method of distributed storage pool OSD switchover according to claim 4, wherein, counting a number of the failed failure domains to obtain the failed failure domain number of the physical failure domain. The marking the state of the failed OSD as down in response to the failed failure domain number being less than the global allowed failed failure domain number comprises: generating a difference value between the global allowed failed failure domain number and the failed failure domain number, and taking the difference value as a failure domain failure remaining amount; detecting whether a failure domain to which the failed OSD belongs belongs to the failed failure domain; in response to the failure domain to which the failed OSD belongs belonging to the failed failure domain, marking the state of the failed OSD as down, and keeping the failure domain failure remaining amount unchanged; in response to the failure domain to which the failed OSD belongs not belonging to the failed failure domain, detecting whether a value of the failure domain failure remaining amount is greater than zero; 6. The method of Claim 1, wherein, in response to the value of the failure domain failure remaining amount being greater than zero, marking the state of the failed OSD as down. Before the marking the state of the failed OSD as down in response to the failed failure domain number being less than the global allowed failed failure domain number, the method further comprises: determining a minimum safe read-write replica number configured by the target storage pool; determining a target PG carried by the failed OSD. determining a number of primary replicas or read-write member participating replicas in the target PG; detecting whether the real-time number of replicas is greater than the minimum safe read-write replica number; if the real-time number of replicas is less than or equal to the minimum safe read-write replica number, prohibiting switching of the failed OSD; if the real-time number of replicas is greater than the minimum safe read-write replica number, performing the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domain number.
7. The method of claim 1, wherein, The method further comprises, before the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domain number: detecting whether the failed OSD is marked as prohibiting fast switching; if the failed OSD is marked as prohibiting fast switching, prohibiting switching of the failed OSD; if the failed OSD is not marked as prohibiting fast switching, performing the step of marking the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domain number.
8. An OSD switching system in a distributed storage pool, characterized in that, The method comprises: a failed OSD determining module configured to determine a failed OSD in a distributed storage pool; a physical failure domain determining module configured to determine a physical failure domain of a target storage pool in which the failed OSD is located; a global failed number obtaining module configured to obtain a global allowed failed failure domain number of the physical failure domain; a failed number determining module configured to determine a number of failed failure domains of the physical failure domain; an OSD switching module configured to mark the status of the failed OSD as down in response to the number of failed failure domains being less than the global allowed failed failure domain number. The global allowed failed failure domain number comprises a minimum value of allowed failure domain numbers of a target storage pool set, the target storage pool set comprising a set of all storage pool combinations sharing the physical failure domain.
9. An electronic device, comprising: The method comprises: a memory configured to store a computer program; a processor configured to implement the steps of the OSD switching method in the distributed storage pool according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer program is stored in the computer readable storage medium and is executed by the processor to implement the steps of the OSD switching method in the distributed storage pool according to any one of claims 1 to 7.