A Solid State Drive Management Method, Device, Equipment, Medium and Product
By implementing two-layer monitoring and preset fault detection strategies in distributed storage systems, the problem of failure of solid-state drives cannot be identified and redundant protection in time when solid-state drive failure is solved, and timely identification and rapid redundant protection of the healthy state of solid-state drives are achieved, avoiding data loss.
Patent Information
- Application Number
- CN202510200150.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The prior art is difficult to identify the healthy state of solid-state drives in a timely and accurate manner, resulting in the inability to quickly and efficiently protect redundantly when the SSD disk fails, increasing the risk of data loss.
By implementing dual-layer monitoring of the service layer and storage layer in a distributed storage system, failure detection of abnormal solid-state drives is performed using preset fault detection strategies, and target protection strategies are determined based on the detection results, so as to achieve redundant protection of abnormal solid-state drives.
It realizes timely and accurate health status identification of solid-state drives, ensuring rapid and efficient redundant protection when a failure occurs, effectively avoiding data loss problems.
Smart Images

Figure CN119724313B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of storage technologies, and in particular, to a method for managing a solid-state drive, and also relates to a solid-state drive management device, an electronic device, a non-volatile storage medium, and a computer program product. Background Art
[0002] Currently, a solid-state drive (SSD) usually deploys multiple DB (DataBase) partitions for OSD services. If the SSD fails, multiple OSDs will fail simultaneously, and the impact will be much greater than that of a normal disk failure.
[0003] In related technologies, when there are consecutive SSD failures in a short period of time in a storage system, or when an SSD failure occurs in a scenario close to the super-fault domain, since the number of failed OSDs is large, if the data water level of the cluster storage pool itself is not low, then the data on all failed OSDs needs to be reconstructed, and the amount of reconstructed data is large. If the data on a batch of OSDs that failed first has not been reconstructed yet, and then an SSD failure occurs, it is equivalent to a superimposed failure occurring in a scenario where the data redundancy of the storage pool has already decreased, which will cause a further decrease in the data redundancy. This scenario will cause the risk of storage data loss to increase rapidly, and may even directly lead to the system entering the super-fault domain and causing data loss. For this scenario, the common current solution is usually to regularly detect the disk status and data redundancy status of the system, and give an alarm about the system health status when disk anomalies are detected. However, this solution has high requirements for the accuracy of disk detection and the timeliness of staff handling. If the disk failure cannot be identified and processed in time, the data loss problem cannot be avoided.
[0004] Therefore, how to timely and accurately identify the health status of a solid-state drive, so as to achieve fast and efficient redundant protection of a faulty disk and avoid data loss problems is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for managing a solid-state drive, which can timely and accurately identify the health status of the solid-state drive, helps to achieve fast and efficient redundant protection of a faulty disk, and effectively avoids data loss problems; another purpose of the present invention is to provide a solid-state drive management device, an electronic device, a non-volatile storage medium, and a computer program product, all of which have the above beneficial effects.
[0006] In a first aspect, the present invention provides a solid-state drive management method applied to a distributed storage system. The distributed storage system includes a service layer and a storage layer. Multiple object storage devices are deployed in the storage layer, and multiple solid-state drives are deployed in each object storage device. The method includes:
[0007] Monitor the service layer in the distributed storage system to determine whether there is an abnormality in the service layer;
[0008] When there is an abnormality in the service layer, determine the abnormal solid-state drive and the type of solid-state drive abnormality in the storage layer; the type of solid-state drive abnormality includes hard disk self-abnormality and / or hard disk life expiration;
[0009] When the type of solid-state drive abnormality is hard disk self-abnormality, perform a fault detection on the abnormal solid-state drive using a preset fault detection strategy to obtain a fault detection result;
[0010] Determine a target protection strategy using the type of solid-state drive abnormality and the fault detection result, and perform redundant protection on the abnormal solid-state drive using the target protection strategy.
[0011] Among them, monitoring the service layer in the distributed storage system to determine whether there is an abnormality in the service layer includes:
[0012] Monitor the read / write process of the object storage device in the service layer of the distributed storage system to determine whether an input / output block event is monitored;
[0013] If an input / output block event is monitored, it is determined that there is an abnormality in the service layer;
[0014] If no input / output block event is monitored, it is determined that there is no abnormality in the service layer.
[0015] Among them, monitoring the service layer in the distributed storage system to determine whether there is an abnormality in the service layer includes:
[0016] Monitor the read / write process of the object storage device in the service layer of the distributed storage system to determine whether an error code is monitored;
[0017] If no error code is monitored, it is determined that there is an abnormality in the service layer;
[0018] If an error code is monitored, determine the abnormal disk partition according to the error code;
[0019] If the storage medium to which the abnormal disk partition belongs is a solid-state drive, it is determined that there is an abnormality in the service layer;
[0020] If the storage medium to which the abnormal disk partition belongs is not a solid-state drive, it is determined that there is no abnormality in the service layer.
[0021] Among them, monitoring the service layer in the distributed storage system to determine whether there is an abnormality in the service layer includes:
[0022] Monitoring the startup process of the object storage device in the service layer of the distributed storage system to determine whether a mounting failure event is monitored; among them, the mounting failure event includes a disk partition mounting failure event and / or a file system mounting failure event;
[0023] If the mounting failure event is monitored, it is determined that there is an abnormality in the service layer;
[0024] If the mounting failure event is not monitored, it is determined that there is no abnormality in the service layer.
[0025] Among them, monitoring the service layer in the distributed storage system to determine whether there is an abnormality in the service layer includes:
[0026] Monitoring the startup process and / or the read / write process of the object storage device in the service layer of the distributed storage system to determine whether a hard disk removal event is monitored;
[0027] If the hard disk removal event is monitored, it is determined that there is an abnormality in the service layer;
[0028] If the hard disk removal event is not monitored, it is determined that there is no abnormality in the service layer.
[0029] Among them, monitoring the service layer in the distributed storage system to determine whether there is an abnormality in the service layer includes:
[0030] Monitoring the startup process and / or the read / write process of the object storage device in the service layer of the distributed storage system to determine whether a hard disk life expiration signal is monitored;
[0031] If the hard disk life expiration signal is monitored, it is determined that there is an abnormality in the service layer;
[0032] If the hard disk life expiration signal is not monitored, it is determined that there is no abnormality in the service layer.
[0033] Among them, using a preset fault detection strategy to perform fault detection on the abnormal solid-state drive to obtain a fault detection result includes:
[0034] Collect parameters of the abnormal solid-state drive to obtain drive parameters; among them, the drive parameters include one or a combination of multiple of power-on time, wear level, and data write volume;
[0035] If all the drive parameters do not exceed the corresponding parameter thresholds, determine that the fault detection result is that there is no fault in the abnormal solid-state drive;
[0036] If any of the drive parameters exceeds the corresponding parameter threshold, determine that the fault detection result is that there is a fault in the abnormal solid-state drive
[0037] Among them, using a preset fault detection strategy to perform fault detection on the abnormal solid-state drive to obtain a fault detection result, including:
[0038] Obtain the drive log corresponding to the abnormal solid-state drive in the system log;
[0039] If there is no drive error log in the drive log, determine that the fault detection result is that there is no fault in the abnormal solid-state drive;
[0040] If there is the drive error log in the drive log, determine that the fault detection result is that there is a fault in the abnormal solid-state drive.
[0041] Among them, using a preset fault detection strategy to perform fault detection on the abnormal solid-state drive to obtain a fault detection result, including:
[0042] Determine whether an abnormal error message about the abnormal solid-state drive is received;
[0043] If no abnormal error message about the abnormal solid-state drive is received, determine that the fault detection result is that there is no fault in the abnormal solid-state drive;
[0044] If an abnormal error message about the abnormal solid-state drive is received, determine that the fault detection result is that there is a fault in the abnormal solid-state drive.
[0045] Among them, the solid-state drive management method further includes:
[0046] When the fault detection result is that there is a fault in the abnormal solid-state drive, determine the faulty object storage device to which the abnormal solid-state drive belongs;
[0047] Control the faulty object storage device to stop running.
[0048] Among them, using the solid-state drive abnormal type and the fault detection result to determine a target protection strategy, and using the target protection strategy to perform redundancy protection on the abnormal solid-state drive, including:
[0049] Determine the abnormal scenario corresponding to the abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result; wherein, the abnormal scenario includes a scenario of continuous hard disk failures within a short period of time and / or a scenario of hard disk expiration in a state approaching the super fault domain; the scenario of continuous hard disk failures within a short period of time indicates that a preset number of solid-state drives in the distributed storage system have failed within a preset time period; the scenario of hard disk expiration in a state approaching the super fault domain indicates that the number of data reconstruction members in the object storage device to which the solid-state drive with expired lifespan belongs is not less than the minimum redundancy number of the object storage device.
[0050] Determine the target protection policy according to the abnormal scenario corresponding to the abnormal solid-state drive, and perform redundancy protection on the abnormal solid-state drive by using the target protection policy.
[0051] Among them, determining the abnormal scenario corresponding to the abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result includes:
[0052] If the abnormal type of the solid-state drive is the hard disk itself abnormal, and the fault detection result is that the abnormal solid-state drive has a fault, then obtain the abnormal time node of the abnormal solid-state drive.
[0053] If the time interval between the abnormal time node of the abnormal solid-state drive and the abnormal time node of the historical abnormal solid-state drive does not exceed the first time period, then determine that the abnormal scenario is the scenario of continuous hard disk failures within a short period of time.
[0054] Among them, determining the abnormal scenario corresponding to the abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result includes:
[0055] If the abnormal type of the solid-state drive is the hard disk lifespan expiration, then determine the expired object storage device to which the abnormal solid-state drive belongs.
[0056] If there are placement groups in the expired object storage device, then determine that the abnormal scenario is the scenario of hard disk expiration in a state approaching the super fault domain.
[0057] Among them, when the abnormal scenario is the scenario of continuous hard disk failures within a short period of time, performing redundancy protection on the abnormal solid-state drive by using the target protection policy includes:
[0058] Set an abnormal label for all the object storage devices in the distributed storage system to block all write operations in all the object storage devices, and output a first alarm prompt.
[0059] Among them, when the abnormal scenario is the scenario of hard disk expiration in a state approaching the super fault domain, performing redundancy protection on the abnormal solid-state drive by using the target protection policy includes:
[0060] Determine the abnormal object storage device to which the abnormal solid-state drive in the distributed storage system belongs;
[0061] Set an abnormal label for the abnormal object storage device to block all write operations in the abnormal object storage device and output a second warning prompt.
[0062] Among them, determining the abnormal scenario corresponding to the abnormal solid-state drive according to the solid-state drive abnormal type and the fault detection result includes:
[0063] Determine the time node when the fault detection result is obtained;
[0064] When the interval duration between the time node and the current time node reaches a second duration, determine the abnormal scenario corresponding to the abnormal solid-state drive according to the solid-state drive abnormal type and the fault detection result.
[0065] In a second aspect, the present invention also discloses a solid-state drive management device, which is applied to a distributed storage system. The distributed storage system includes a service layer and a storage layer. Multiple object storage devices are deployed in the storage layer, and multiple solid-state drives are deployed in each object storage device. The device includes:
[0066] A monitoring module for monitoring the service layer in the distributed storage system to determine whether there is an abnormality in the service layer;
[0067] A determination module for determining an abnormal solid-state drive and a solid-state drive abnormal type in the storage layer when there is an abnormality in the service layer; the solid-state drive abnormal type includes hard disk self-abnormality and / or hard disk life expiration;
[0068] A detection module for, when the solid-state drive abnormal type is the hard disk self-abnormality, performing a fault detection on the abnormal solid-state drive by using a preset fault detection strategy to obtain a fault detection result;
[0069] A protection module for determining a target protection strategy by using the solid-state drive abnormal type and the fault detection result, and performing redundant protection on the abnormal solid-state drive by using the target protection strategy.
[0070] In a third aspect, the present invention also discloses an electronic device, including:
[0071] A memory for storing a computer program;
[0072] A processor for implementing the steps of any one of the above-mentioned solid-state drive management methods when executing the computer program.
[0073] Fourthly, the present invention also discloses a non-volatile storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned solid-state drive management methods are implemented.
[0074] Fifthly, the present invention also discloses a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of any one of the above-mentioned solid-state drive management methods are implemented.
[0075] The present invention provides a solid-state drive management method, which is applied to a distributed storage system. The distributed storage system includes a service layer and a storage layer. A plurality of object storage devices are deployed in the storage layer, and a plurality of solid-state drives are deployed in each object storage device. The method includes: monitoring the service layer in the distributed storage system to determine whether there is an abnormality in the service layer; when there is an abnormality in the service layer, determining an abnormal solid-state drive and a solid-state drive abnormality type in the storage layer; the solid-state drive abnormality type includes hard disk self-abnormality and / or hard disk life expiration; when the solid-state drive abnormality type is the hard disk self-abnormality, performing a fault detection on the abnormal solid-state drive by using a preset fault detection strategy to obtain a fault detection result; determining a target protection strategy by using the solid-state drive abnormality type and the fault detection result, and performing redundancy protection on the abnormal solid-state drive by using the target protection strategy.
[0076] By applying the technical solution provided by the present invention, first, an abnormality detection is performed on the service layer in the distributed storage system to determine whether there may be an abnormality in the solid-state drive in the distributed storage system. When there is an abnormality in the service layer, an abnormal solid-state drive and a solid-state drive abnormality type are determined in the storage layer. Thus, the preset solid-state drive fault detection strategy can be used to perform a fault detection on the abnormal solid-state drive again to determine whether the abnormal solid-state drive actually has a fault. Then, a target protection strategy corresponding to the abnormal solid-state drive is determined according to the solid-state drive abnormality type and the fault detection result, and the redundancy protection of the abnormal solid-state drive is realized based on the target protection strategy. It can be seen that the technical solution realizes double-layer monitoring of the service layer and the storage layer in the distributed storage system, takes into account various situations such as hard disk life expiration, hard disk real fault, and hard disk abnormal scenarios, can timely and accurately identify the health state of the solid-state drive, helps to realize the redundancy protection of the faulty disk quickly and efficiently, and further avoids the problem of data loss.
[0077] The solid-state drive management device, electronic device, non-volatile storage medium, and computer program product provided by the present invention also have the above technical effects, and the present invention will not be elaborated herein. Description of the Drawings
[0078] To more clearly illustrate the technical solutions in the prior art and the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the prior art and the embodiments of the present invention. Of course, the following drawings related to the embodiments of the present invention only describe a part of the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings, and the other obtained drawings also fall within the protection scope of the present invention.
[0079] Figure 1 A schematic structural diagram of a distributed storage system provided by an embodiment of the present invention;
[0080] Figure 2 A schematic flowchart of a solid-state drive management method provided by an embodiment of the present invention;
[0081] Figure 3 A schematic flowchart of another solid-state drive management method provided by an embodiment of the present invention;
[0082] Figure 4 A timing diagram of an interaction process in the scenario of continuous failures of solid-state drives within a short time provided by an embodiment of the present invention;
[0083] Figure 5 A schematic flowchart of an algorithm processing process of a Monitor in the scenario of continuous failures of solid-state drives within a short time provided by an embodiment of the present invention;
[0084] Figure 6 A schematic flowchart of an algorithm processing process of an OSD service in the scenario of continuous failures of solid-state drives within a short time due to IO errors provided by an embodiment of the present invention;
[0085] Figure 7 A schematic flowchart of an algorithm processing process of an OSD service in the scenario of continuous failures of solid-state drives within a short time due to IO jams provided by an embodiment of the present invention;
[0086] Figure 8 A schematic flowchart of an algorithm processing process of an OSD service in the scenario of continuous failures of solid-state drives within a short time due to disk partition mounting failures provided by an embodiment of the present invention;
[0087] Figure 9 A schematic flowchart of an algorithm processing process of an OSD service in the scenario of continuous failures of solid-state drives within a short time due to DB mounting failures provided by an embodiment of the present invention;
[0088] Figure 10 A schematic flowchart of an algorithm processing process of a Monitor in the scenario of the expiration of the solid-state drive life in the near super-failure scenario provided by an embodiment of the present invention;
[0089] Figure 11 A schematic structural diagram of a solid-state drive management device provided by an embodiment of the present invention;
[0090] Figure 12 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0091] The core of the present invention is to provide a solid-state drive management method, which can timely and accurately identify the health status of the solid-state drive, contribute to realizing the redundant protection of the faulty drive quickly and efficiently, and effectively avoid the problem of data loss; another core of the present invention is to provide a solid-state drive management device, an electronic device, a non-volatile storage medium, and a computer program product, all of which have the above beneficial effects.
[0092] In order to describe the technical solutions in the embodiments of the present invention more clearly and completely, the following will introduce the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0093] An embodiment of the present invention provides a solid-state drive management method.
[0094] First, please refer to Figure 1 , Figure 1 A schematic structural diagram of a distributed storage system provided by an embodiment of the present invention. The distributed storage system may include a storage layer, a service layer, and an interface layer. A plurality of object storage devices are deployed in the storage layer, and a plurality of solid-state drives are deployed in each object storage device; among them, the interface layer is used to dock with each user device, the storage layer is used to implement distributed data storage, and the service layer docks with the interface layer and the storage layer, and is used to execute the user services received by the interface layer by using the stored data in the storage layer. The solid-state drive management method provided by the embodiment of the present invention is implemented based on this distributed storage system.
[0095] Further, please refer to Figure 2 , Figure 2 A flowchart of a solid-state drive management method provided by an embodiment of the present invention. The solid-state drive management method may include the following S101 to S104.
[0096] S101: Monitor the service layer in the distributed storage system to determine whether there is an abnormality in the service layer.
[0097] This step aims to achieve service layer monitoring in a distributed storage system to determine whether there are any abnormalities in its service layer. In the specific implementation process, it can mainly monitor the input / output service and / or storage object service running in the service layer. Among them, the input / output service is the read / write service regarding the storage layer, and the storage object service is the startup service and running service regarding the storage layer. It can be understood that when there are abnormalities in the input / output service and storage object service in the storage system, these abnormalities may be caused by solid-state drive failures. Therefore, the input / output service and / or storage object service in the distributed storage system can be used to determine whether there may be solid-state drive abnormalities in the storage system. Based on this, the service layer of the distributed storage system can be monitored to obtain the input / output service and / or storage object service therein, thereby realizing service layer abnormality judgment.
[0098] Furthermore, the following embodiments provide several different service layer abnormality judgment methods. It can be understood that in actual application scenarios, the following several judgment methods can exist simultaneously, or any one or part of them can be selected for execution. The present invention does not make any limitations in this regard.
[0099] In an embodiment of the present invention, monitoring the service layer in a distributed storage system to determine whether there are abnormalities in the service layer may include: monitoring the read / write process of object storage devices in the service layer of the distributed storage system to determine whether an input / output blockage event is monitored; if an input / output blockage event is monitored, it is determined that there are abnormalities in the service layer; if no input / output blockage event is monitored, it is determined that there are no abnormalities in the service layer. That is to say, during the read / write process of each object storage device in the storage layer, if an input / output blockage event (i.e., an IO stuck problem) is triggered, it can be considered that there are abnormalities in the service layer, otherwise, it is considered that there are no abnormalities in the service layer.
[0100] In an embodiment of the present invention, monitoring the service layer in a distributed storage system to determine whether there are abnormalities in the service layer includes: monitoring the read / write process of object storage devices in the service layer of the distributed storage system to determine whether an error code is monitored; if no error code is monitored, it is determined that there are abnormalities in the service layer; if an error code is monitored, the abnormal disk partition is determined according to the error code; if the storage medium to which the abnormal disk partition belongs is a solid-state drive, it is determined that there are abnormalities in the service layer; if the storage medium to which the abnormal disk partition belongs is not a solid-state drive, it is determined that there are no abnormalities in the service layer. That is to say, during the read / write process of each object storage device in the storage layer, if a read / write operation returns an error code (an error code less than 0), and the abnormal disk partition corresponding to this read / write operation belongs to a solid-state drive (because it may also belong to other types of storage media), it can be considered that there are abnormalities in the service layer, otherwise, it is considered that there are no abnormalities in the service layer.
[0101] In an embodiment of the present invention, the service layer in the distributed storage system is monitored to determine whether there is an abnormality in the service layer, including: monitoring the startup process of the object storage device in the service layer of the distributed storage system to determine whether a mounting failure event is monitored; wherein, the mounting failure event includes a disk partition mounting failure event and / or a file system mounting failure event; if a mounting failure event is monitored, it is determined that there is an abnormality in the service layer; if no mounting failure event is monitored, it is determined that there is no abnormality in the service layer. That is to say, during the startup process of each storage object device in the distributed storage system, if a mounting failure event is triggered, such as a disk partition mounting failure event and / or a file system mounting failure event, etc., it can be considered that there is an abnormality in the service layer, otherwise, it is considered that there is no abnormality in the service layer.
[0102] In an embodiment of the present invention, the service layer in the distributed storage system is monitored to determine whether there is an abnormality in the service layer, including: monitoring the startup process and / or the read / write process of the object storage device in the service layer of the distributed storage system to determine whether a hard disk removal event is monitored; if a hard disk removal event is monitored, it is determined that there is an abnormality in the service layer; if no hard disk removal event is monitored, it is determined that there is no abnormality in the service layer. That is to say, during the startup process of each storage object device in the distributed storage system, and / or, during the read / write process of each object storage device, if a hard disk removal event (udev remove event) is triggered, it can be considered that there is an abnormality in the service layer, otherwise, it is considered that there is no abnormality in the service layer.
[0103] In an embodiment of the present invention, the service layer in the distributed storage system is monitored to determine whether there is an abnormality in the service layer, including: monitoring the startup process and / or the read / write process of the object storage device in the service layer of the distributed storage system to determine whether a hard disk life expiration signal is monitored; if a hard disk life expiration signal is monitored, it is determined that there is an abnormality in the service layer; if no hard disk life expiration signal is monitored, it is determined that there is no abnormality in the service layer. That is to say, during the startup process of each storage object device in the distributed storage system, and / or, during the read / write process of each object storage device, if a hard disk life expiration signal is detected, it can be considered that there is an abnormality in the service layer, otherwise, it is considered that there is no abnormality in the service layer.
[0104] S102: When there is an abnormality in the service layer, determine the abnormal solid-state drive and the solid-state drive abnormality type in the storage layer; the solid-state drive abnormality type includes hard disk self-abnormality and / or hard disk life expiration.
[0105] This step aims to determine the abnormal solid-state drive and the type of solid-state drive abnormality. That is to say, when it is determined that there is an abnormality in the service layer, it is possible to further determine the abnormal solid-state drive in the storage layer within the distributed storage system (i.e., the solid-state drive with an abnormality, and the number of solid-state drives in the distributed storage system is not unique) and the type of abnormality of the abnormal solid-state drive. Among them, the types of solid-state drive abnormalities include hard disk self-abnormality and hard disk life expiration. Hard disk self-abnormality means that there are abnormal problems with the solid-state drive itself, and hard disk life expiration means that the life of the solid-state drive has expired and a new solid-state drive needs to be replaced.
[0106] S103: When the type of solid-state drive abnormality is the hard disk self-abnormality, use a preset fault detection strategy to perform a fault detection on the abnormal solid-state drive to obtain a fault detection result.
[0107] This step aims to implement the fault detection of the abnormal solid-state drive. It can be understood that even though it is possible to judge whether there is an abnormal solid-state drive in the storage layer based on whether there is an abnormality in the service layer of the distributed storage system and determine the abnormal solid-state drive existing in the distributed storage system, it is impossible to determine whether the abnormal solid-state drive really has a fault problem. For example, in the disk replacement scenario, the solid-state drive abnormality will also be triggered, but this does not mean that the solid-state drive has a fault. Therefore, when it is determined that the type of abnormality of the abnormal solid-state drive is the hard disk self-abnormality, a preset fault detection strategy can also be used to perform a fault detection on the abnormal solid-state drive to determine whether the abnormal solid-state drive really has a fault problem. It can be understood that the fault detection result is that the abnormal solid-state drive has a fault or the abnormal solid-state drive does not have a fault.
[0108] Furthermore, the following embodiments provide several different implementation methods for using a preset fault detection strategy to perform a fault detection on the abnormal solid-state drive. It can be understood that in the actual application scenario, the following several detection methods can exist simultaneously, or one or some of them can be selected and executed, and the present invention does not limit this.
[0109] In one embodiment of the present invention, a preset fault detection strategy is used to detect faults in an abnormal solid-state drive, and a fault detection result can be obtained, which may include: collecting parameters of the abnormal solid-state drive to obtain hard disk parameters; wherein, the hard disk parameters include one or a combination of the power-on time, wear degree, and data write volume; if all the hard disk parameters do not exceed the corresponding parameter thresholds, it is determined that the fault detection result is that there is no fault in the abnormal solid-state drive; if any one of the hard disk parameters exceeds the corresponding parameter threshold, it is determined that the fault detection result is that there is a fault in the abnormal solid-state drive. That is to say, when the value of any type of hard disk parameter in the abnormal solid-state drive does not meet the respective corresponding threshold ranges (i.e., the above parameter thresholds), it can be determined that there is a fault in the abnormal solid-state drive, otherwise it can be considered that there is no fault in the abnormal solid-state drive. It should be noted that the specific values of the respective parameter thresholds are determined by the performance of the solid-state drive itself, which does not affect the implementation of this technical solution, and the present invention does not limit this.
[0110] In one embodiment of the present invention, a preset fault detection strategy is used to detect faults in an abnormal solid-state drive, and a fault detection result can be obtained, which may include: obtaining the hard disk log corresponding to the abnormal solid-state drive in the system log; if there is no hard disk error log in the hard disk log, it is determined that the fault detection result is that there is no fault in the abnormal solid-state drive; if there is a hard disk error log in the hard disk log, it is determined that the fault detection result is that there is a fault in the abnormal solid-state drive. That is to say, when a hard disk error log appears in the system log of the storage system, it can be determined that there is a fault in the abnormal solid-state drive, otherwise it can be considered that there is no fault in the abnormal solid-state drive. It can be understood that this process can be achieved directly by retrieving the system log for log analysis.
[0111] In one embodiment of the present invention, a preset fault detection strategy is used to detect faults in an abnormal solid-state drive, and a fault detection result can be obtained, which may include: determining whether an abnormal error message about the abnormal solid-state drive is received; if an abnormal error message about the abnormal solid-state drive is not received, it is determined that the fault detection result is that there is no fault in the abnormal solid-state drive; if an abnormal error message about the abnormal solid-state drive is received, it is determined that the fault detection result is that there is a fault in the abnormal solid-state drive. That is to say, in an actual operation scenario, if the execution entity (the main control module of the storage system) receives an abnormal error message about the abnormal solid-state drive from any source, it can be determined that there is a fault in the abnormal solid-state drive, otherwise it can be considered that there is no fault in the abnormal solid-state drive.
[0112] In one embodiment of the present invention, the solid-state drive management method may further include: when the fault detection result is that there is a fault in the abnormal solid-state drive, determining the faulty object storage device to which the abnormal solid-state drive belongs; controlling the faulty object storage device to stop running.
[0113] It can be understood that when it is determined that the abnormal solid-state drive (SSD) actually has a fault, since the abnormal SSD can no longer provide storage services, the corresponding faulty object storage device of the abnormal SSD can be directly controlled to stop running, so as to effectively avoid data loss problems caused by the generation of new data but the inability to store it.
[0114] S104: Determine the target protection strategy based on the SSD exception type and the fault detection result, and perform redundant protection on the abnormal SSD using the target protection strategy.
[0115] This step aims to achieve redundant protection of the abnormal SSD. Specifically, different protection strategies can be preset for different SSD exception types and fault detection results. Therefore, after determining the exception type and fault detection result of the abnormal SSD, the corresponding target protection strategy can be determined by combining the two, so as to achieve redundant protection of the abnormal SSD using the target protection strategy.
[0116] It can be seen that the SSD management method provided by the embodiments of the present invention first determines whether there may be an SSD exception problem in the distributed storage system by performing exception detection on the service layer in the distributed storage system. When there is an exception in the service layer, the abnormal SSD and the SSD exception type are determined in the storage layer. Thus, the preset SSD fault detection strategy can be used to re-detect the fault of the abnormal SSD to determine whether the abnormal SSD actually has a fault, and then the target protection strategy corresponding to the abnormal SSD can be determined according to the SSD exception type and the fault detection result, and redundant protection of the abnormal SSD is achieved based on the target protection strategy. It can be seen that this technical solution realizes double-layer monitoring of the service layer and the storage layer in the distributed storage system, taking into account various situations such as the expiration of the hard disk life, the real fault of the hard disk, and the hard disk exception scenario, and can timely and accurately identify the health status of the SSD, which helps to achieve fast and efficient redundant protection of the faulty disk and further avoid data loss problems.
[0117] Based on the above embodiments:
[0118] In one embodiment of the present invention, a target protection policy is determined based on the abnormal type of the solid-state drive and the fault detection result, and the abnormal solid-state drive is redundantly protected using the target protection policy, including: determining the abnormal scenario corresponding to the abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result; wherein, the abnormal scenario includes a scenario of continuous hard disk failures within a short period of time and / or a scenario of hard disk expiration in a state close to the super fault domain; the scenario of continuous hard disk failures within a short period of time indicates that a preset number of solid-state drives in the distributed storage system have failed within a preset time period; the scenario of hard disk expiration in a state close to the super fault domain indicates that the number of data reconstruction members in the object storage device to which the solid-state drive with expired lifespan belongs is not less than the minimum redundancy number of the object storage device; determining the target protection policy according to the abnormal scenario corresponding to the abnormal solid-state drive, and redundantly protecting the abnormal solid-state drive using the target protection policy.
[0119] Specifically, based on the abnormal type of the solid-state drive and the fault detection result, the abnormal scenario in which the abnormal solid-state drive is currently located can be determined. This abnormal scenario mainly includes a scenario of continuous hard disk failures within a short period of time and a scenario of hard disk expiration in a state close to the super fault domain. Thus, the determination of the target protection policy can be achieved according to different abnormal scenarios (different abnormal scenarios correspond to different protection policies). Further, the following two embodiments respectively provide identification methods for the scenario of continuous hard disk failures within a short period of time and the scenario of hard disk expiration in a state close to the super fault domain.
[0120] In one embodiment of the present invention, determining the abnormal scenario corresponding to the abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result includes: if the abnormal type of the solid-state drive is a hard disk self-abnormality and the fault detection result is that there is a fault in the abnormal solid-state drive, obtaining the abnormal time node of the abnormal solid-state drive; if the time interval between the abnormal time node of the abnormal solid-state drive and the abnormal time node of the historical abnormal solid-state drive does not exceed the first time period, determining the abnormal scenario as a scenario of continuous hard disk failures within a short period of time. As described above, the scenario of continuous hard disk failures within a short period of time is a problem that multiple (two in this embodiment, and of course, more can be used) solid-state drive failures occur in the distributed storage system within a short period of time. Therefore, the identification of the scenario of continuous hard disk failures within a short period of time can be achieved by referring to the detection time interval between two adjacent abnormal solid-state drives. It should be noted that the specific value of the above-mentioned first time period does not affect the implementation of this technical solution and can be set by those skilled in the art according to the actual situation. For example, it can be set to 3 days, and the present invention does not limit this.
[0121] In an embodiment of the present invention, determining an abnormal scenario corresponding to an abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result includes: if the abnormal type of the solid-state drive is that the hard disk life expires, determining the expired object storage device to which the abnormal solid-state drive belongs; if there is a placement group in the expired object storage device, determining the abnormal scenario as the hard disk expiration scenario under the state of approaching the super fault domain. That is to say, when the life of the solid-state drive expires, if there is a placement group in any storage object device to which it belongs, it can be considered that the current abnormal scenario is the hard disk expiration scenario under the state of approaching the super fault domain. Because when the life of the solid-state drive expires, if there is a placement group in the storage object device to which it belongs, it means that the number of members for data reconstruction in this storage object device has reached its corresponding minimum redundancy (determined by the performance of the corresponding object storage device). If there is another solid-state drive abnormality or the solid-state drive with expired life is not replaced, it will inevitably lead to the overall abnormality of the distributed storage system and cause data loss problems.
[0122] Furthermore, the following two embodiments respectively provide target protection strategies for the scenarios of continuous hard disk failures within a short period of time and hard disk expiration scenarios under the state of approaching the super fault domain, in order to achieve redundant protection of abnormal solid-state drives. Specifically, it may include:
[0123] When the abnormal scenario is the scenario of continuous hard disk failures within a short period of time, using the target protection strategy to perform redundant protection on the abnormal solid-state drive may include: setting an abnormal label for all object storage devices in the distributed storage system to block all write operations in all object storage devices and output a first alarm prompt.
[0124] When the abnormal scenario is the hard disk expiration scenario under the state of approaching the super fault domain, using the target protection strategy to perform redundant protection on the abnormal solid-state drive may include: determining the abnormal object storage device to which the abnormal solid-state drive belongs in the distributed storage system; setting an abnormal label for the abnormal object storage device to block all write operations in the abnormal object storage device and output a second alarm prompt.
[0125] It should be noted that the above write operations mainly include business write operations and data reconstruction operations in the distributed storage system.
[0126] In an embodiment of the present invention, determining an abnormal scenario corresponding to an abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result may include: determining the time node when the fault detection result is obtained; when the interval duration between the time node and the current time node reaches a second duration, determining the abnormal scenario corresponding to the abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result.
[0127] Specifically, after obtaining the fault detection result, it is possible to enter the waiting state starting from the moment when the fault detection result is obtained and continue for a preset duration, that is, the above-mentioned second duration, and then execute the step of determining the abnormal scenario corresponding to the abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result. It can be understood that through this implementation method, a time window can be left for normal disk replacement and operation and maintenance operations, so as to effectively distinguish disk replacement operations and hard disk fault scenarios, avoid mis-triggering the redundant protection mechanism for abnormal solid-state drives, and ensure the accuracy of solid-state drive management. It should be noted that the specific value of the above-mentioned second duration does not affect the implementation of the technical solution of the present invention, and can be set by those skilled in the art according to the actual situation. For example, it can be set to 10 minutes, and the present invention does not limit this.
[0128] An embodiment of the present invention provides another solid-state drive management method.
[0129] First of all, the solid-state drive management method provided by the present invention mainly includes three modules: a solid-state drive health status detection module, a solid-state drive abnormal scenario recognition module, and a post-abnormal processing module. The specific implementation functions of each module are as follows:
[0130] 1. Solid-state drive health status detection module: By calling this module, the specified solid-state drive is detected to determine whether there are any fault phenomena in the solid-state drive and whether it meets the standard of disk failure.
[0131] 2. Solid-state drive abnormal scenario recognition module: This module is implemented at the service layer of the distributed storage system. It mainly senses whether a specified abnormal scenario has occurred in the relevant processes of the business input / output service and the storage object service. These abnormal scenarios may all be caused by solid-state drive failures. Therefore, after detecting the occurrence of these scenarios, it is necessary to call the solid-state drive health status detection module to detect the specified solid-state drive. If it is detected that the solid-state drive has indeed failed, relevant command lines can be called to report to the distributed storage system monitoring service (Monitor).
[0132] 3. Post-abnormal processing module: This module is mainly implemented in Monitor and the storage layer. It determines the specific abnormal situation according to the reported information in 2 and executes the corresponding processing logic.
[0133] Further, please refer to Figure 3 , Figure 3 which is a schematic flowchart of another solid-state drive management method provided by an embodiment of the present invention. Based on the above functional modules, the specific implementation process of the solid-state drive management method provided by the present invention is as follows:
[0134] 1. During the startup process and operation process (mainly referring to the read / write process) of the storage object device, the following abnormal detections will be performed:
[0135] (1) During the startup process, whether there are problems such as disk partition mounting failure and file system mounting failure;
[0136] (2) During the running process, whether there is an IO stuck problem;
[0137] (3) During the running process, whether there is an IO error problem;
[0138] (4) During the running process, whether there is a disk drop or disk removal scenario, triggering the udev remove event;
[0139] (5) During the running process, whether there is a problem that the solid state drive reaches the end of its lifespan.
[0140] 2. If an abnormality is detected according to Step 1, it is necessary to determine whether the abnormality is an abnormality of the solid state drive itself (i.e., (1), (2), (3), (4) in Step 1), or the solid state drive reaches the end of its lifespan (i.e., (5) in Step 1); if it is an abnormality of the solid state drive itself, call the solid state drive health status detection module of the underlying software to detect the corresponding solid state drive, and confirm whether the solid state drive is actually faulty at present. If the solid state drive reaches the end of its lifespan, it can be directly reported to the Monitor.
[0141] 3. If the detection result of calling the solid state drive health status detection module returns that the solid state drive is faulty, relevant information (such as the information of the solid state drive itself, the information of the affiliated object storage device, etc.) can be reported to the Monitor, and the storage object service will exit after the reporting. Among them, the functional implementation of the solid state drive health status detection module mainly depends on:
[0142] (1) Determine whether hard disk parameters such as the power-on time, wear level, and data write volume of the solid state drive exceed the limit;
[0143] (2) Determine whether there is relevant log information about hardware information error reporting in the system log of the solid state drive;
[0144] (3) Determine whether there is abnormal error reporting information in the obtained solid state drive information.
[0145] 4. After receiving the reported information, the Monitor records and updates the abnormal solid state drive, and judges:
[0146] (1) Whether there is a continuous and complete failure of the solid-state drive within a short period: Set a short-time threshold (such as 3 days). If the solid-state drive is continuously and completely damaged and cannot be repaired within this time range, and the services on the solid-state drive cannot be normally restarted, then this failure scenario is met. If this failure scenario is triggered, it is usually due to abnormal failure reasons in the solid-state drive itself (such as software bugs in the solid-state drive, etc.), resulting in the current abnormal damage frequency of the solid-state drive. Usually when such reasons exist, other solid-state drives in the storage system may also have the risk of continuous damage in a short period, which may further lead to the object storage device exceeding the fault domain and causing data loss problems.
[0147] (2) Whether there is a situation where the solid-state drive reaches the end of its life when the system is near the state of exceeding the fault domain: The object storage device is currently in a state near exceeding the fault domain, that is, some data in the current cluster has no redundancy, and any further data damage will cause the fault domain to be exceeded (scenarios such as: several solid-state drives in the cluster have been damaged, but data reconstruction has not been completed). Usually, the end of the solid-state drive's life is a prediction. When the standard for reaching the end of the life is met, the solid-state drive will not immediately fail, and the services on the drive will not immediately malfunction. However, when an alarm indicating that the solid-state drive is about to fail appears in the state where the cluster has no redundancy, it also means that the cluster currently has a risk of exceeding the fault domain, which is essentially the same as the risk in failure scenario (1).
[0148] 5. If it is a scenario of continuous solid-state drive failures, then mark all object storage devices in the storage system with an abnormal solid-state drive label, block all write operations of the object storage devices, including business writes and data reconstruction, and report an alarm; if it is a scenario where the solid-state drive reaches the end of its life when the system is near the state of exceeding the fault domain, then mark the relevant object storage devices of the solid-state drive that reaches the end of its life (i.e., the above abnormal object storage devices) with an abnormal solid-state drive label, block the write operations of these object storage devices, including business writes and data reconstruction, and report an alarm.
[0149] Further, the following is a detailed introduction to the processing flows for the two failure scenarios respectively:
[0150] I. Continuous solid-state drive failures within a short period.
[0151] First, please refer to Figure 4 , Figure 4 which is the interaction process timing diagram for the scenario of continuous solid-state drive failures within a short period provided by the embodiment of the present invention. Based on Figure 4It can be seen that the triggering action of the overall detection mechanism for consecutive failures of solid-state drives within a short period is performed by the object storage service. When the object storage service perceives an anomaly at the business process level, it invokes the solid-state drive health status detection module of the underlying software, and the solid-state drive health status detection module is always in a state of being passively invoked. Subsequently, based on the output result of the solid-state drive health status detection module, it is decided whether to report relevant information to the Monitor, and then the Monitor decides whether to block the write operation within the distributed storage system and report an alarm. In the entire process, the mechanism is triggered by the object storage service, the underlying software conducts specific inspections on the hardware, and finally the Monitor makes the final decision.
[0152] Further, please refer to Figure 5 , Figure 5 which is a schematic diagram of the algorithm processing flow of a Monitor in the scenario of consecutive failures of solid-state drives within a short period provided by an embodiment of the present invention, and its implementation process is as follows:
[0153] (1) After receiving the reported information, the Monitor records the information, including the reported information and the time node when the reported information is received.
[0154] (2) After recording the information, wait for 10 minutes and then perform a failure scenario judgment. The purpose is to leave a time window for normal disk replacement operation and maintenance to distinguish between disk replacement operation and hard disk failure scenario. If it is a disk replacement operation, when the solid-state drive is removed, a udev remove event will be triggered to report the solid-state drive anomaly; and when the solid-state drive is inserted back, a udev add event will be triggered to report the elimination of the solid-state drive anomaly, and the abnormal solid-state drive caused by the disk removal just recorded by the Monitor will be deleted. In this way, it is possible to avoid accidentally triggering the redundancy protection mechanism during the disk removal process of the normal operation and maintenance disk replacement process.
[0155] (3) Judge the time interval between the reports of abnormal solid-state drives. If it is the first reported solid-state drive, the time interval is 0; if it is not the first solid-state drive, calculate the time interval from the reporting time of the previous abnormal solid-state drive and judge whether it meets the time threshold for consecutive failures of solid-state drives within a short period (assumed to be 3 days).
[0156] (4) If the judgment criterion for consecutive failures of solid-state drives within a short period is met, mark all object storage devices within the distributed storage system with a solid-state drive anomaly label to block all write operations and report an alarm.
[0157] Finally, please refer to Figures 6 to 9 , Figure 6 which is a schematic diagram of the algorithm processing flow of an OSD service in the scenario of consecutive failures of solid-state drives within a short period due to IO errors provided by an embodiment of the present invention.Figure 7 It is a schematic diagram of the algorithm processing flow of an OSD service provided by an embodiment of the present invention in the scenario of consecutive failures of a solid-state drive within a short period due to IO jamming. Figure 8 It is a schematic diagram of the algorithm processing flow of an OSD service provided by an embodiment of the present invention in the scenario of consecutive failures of a solid-state drive within a short period due to disk partition mounting failure. Figure 9 It is a schematic diagram of the algorithm processing flow of an OSD service provided by an embodiment of the present invention in the scenario of consecutive failures of a solid-state drive within a short period due to DB mounting failure. As can be seen from Figures 6 to 9 it, the detection logic of the OSD service in the scenario of consecutive failures of a solid-state drive within a short period is mainly as follows:
[0158] (1) Identify abnormal scenarios (IO error reporting, IO jamming, mounting failure, etc.).
[0159] (2) Obtain the current storage medium through different strategies in different abnormal scenarios and determine whether the current storage medium is a solid-state drive.
[0160] (3) If it is a solid-state drive, call the solid-state drive health status detection module to confirm whether the solid-state drive is indeed damaged.
[0161] (4) If the solid-state drive is damaged, report relevant information to the Monitor and exit the process.
[0162] II. The lifespan of the solid-state drive expires in the scenario of approaching the super-failure scenario.
[0163] Please refer to Figure 10 , Figure 10 It is a schematic diagram of the algorithm processing flow of a Monitor provided by an embodiment of the present invention in the scenario where the lifespan of the solid-state drive expires in the scenario of approaching the super-failure scenario. The implementation process is as follows:
[0164] (1) After the OSD service detects that the lifespan of the solid-state drive has expired, it reports the information of the solid-state drive with the expired lifespan to the Monitor, but the process does not exit immediately.
[0165] (2) After the Monitor receives the reported message from the OSD service, it obtains the object storage device list corresponding to the solid-state drive (that is, which object storage devices are related to the solid-state drive, which can be determined through the ID information) and the device information of the object storage device where the solid-state drive is located.
[0166] (3) Traverse whether there is a PG on these object storage devices. If it exists, the current storage system is already in the state of approaching the super-failure domain; if it does not exist, the current storage system is not in the state of approaching the super-failure domain.
[0167] (4) If it is determined that the distributed storage system is in a state approaching a super failure domain, then mark the relevant object storage devices with a solid-state drive anomaly label, block all write operations of the relevant object storage devices, and report an alarm.
[0168] (5) The relevant information of the solid-state drives whose lifespan has expired recorded by the Monitor is cleared after the data reconstruction is completed and all PG data redundancy is restored and the PG status is restored to the active (active) + clean (clean) state.
[0169] Thus, it can be seen that the solid-state drive management method provided by the embodiments of the present invention first determines whether there may be a solid-state drive anomaly problem in the distributed storage system by performing anomaly detection on the service layer in the distributed storage system, and when there is an anomaly in the service layer, determines the abnormal solid-state drive and the type of solid-state drive anomaly in the storage layer. Thus, the preset solid-state drive fault detection strategy can be used to re-detect the abnormal solid-state drive to determine whether the abnormal solid-state drive actually has a fault, so as to determine the target protection strategy corresponding to the abnormal solid-state drive according to the type of solid-state drive anomaly and the fault detection result, and implement redundant protection of the abnormal solid-state drive based on the target protection strategy. It can be seen that this technical solution realizes double-layer monitoring of the service layer and the storage layer in the distributed storage system, takes into account various situations such as the expiration of the hard disk lifespan, the real fault of the hard disk, and the hard disk anomaly scenario, can timely and accurately identify the health status of the solid-state drive, helps to achieve redundant protection of the faulty disk quickly and efficiently, and further avoids the problem of data loss.
[0170] The embodiments of the present invention provide a solid-state drive management device.
[0171] Please refer to Figure 11 , Figure 11 , which is a schematic structural diagram of a solid-state drive management device provided by the present invention. The solid-state drive management device is applied to a distributed storage system, and the distributed storage system includes a service layer and a storage layer. Multiple object storage devices are deployed in the storage layer, and multiple solid-state drives are deployed in each object storage device. It may include:
[0172] Monitoring module 1, configured to monitor the service layer in the distributed storage system to determine whether there is an anomaly in the service layer;
[0173] Determination module 2, configured to determine the abnormal solid-state drive and the type of solid-state drive anomaly in the storage layer when there is an anomaly in the service layer; the type of solid-state drive anomaly includes hard disk self-anomaly and / or hard disk lifespan expiration;
[0174] Detection module 3, configured to perform a fault detection on the abnormal solid-state drive by using a preset fault detection strategy to obtain a fault detection result when the type of solid-state drive anomaly is a hard disk self-anomaly;
[0175] The protection module 4 is configured to determine a target protection policy based on the SSD exception type and the fault detection result, and perform redundant protection on the abnormal SSD by using the target protection policy.
[0176] It can be seen that for the SSD management device provided by the embodiment of the present invention, first, by performing anomaly detection on the service layer in the distributed storage system, it is determined whether there may be an SSD anomaly problem in the distributed storage system. When there is an anomaly in the service layer, the abnormal SSD and the SSD anomaly type are determined in the storage layer. Thus, the preset SSD fault detection policy can be used to perform fault detection on the abnormal SSD again to determine whether the abnormal SSD actually has a fault. Then, based on the SSD anomaly type and the fault detection result, the target protection policy corresponding to the abnormal SSD is determined, and redundant protection of the abnormal SSD is realized based on the target protection policy. It can be seen that this technical solution realizes double-layer monitoring of the service layer and the storage layer in the distributed storage system, takes into account various situations such as the hard disk reaching the end of its life, the hard disk having a real fault, and the hard disk being in an abnormal scenario, can identify the health status of the SSD in a timely and accurate manner, helps to achieve redundant protection of the faulty disk quickly and efficiently, and further avoids the problem of data loss.
[0177] In an embodiment of the present invention, the above-mentioned determination module 1 may be specifically configured to monitor the read / write process of the object storage device in the service layer of the distributed storage system to determine whether an input / output blockage event is monitored; if an input / output blockage event is monitored, it is determined that there is an anomaly in the service layer; if no input / output blockage event is monitored, it is determined that there is no anomaly in the service layer.
[0178] In an embodiment of the present invention, the above-mentioned determination module 1 may be specifically configured to monitor the read / write process of the object storage device in the service layer of the distributed storage system to determine whether an error code is monitored; if no error code is monitored, it is determined that there is an anomaly in the service layer; if an error code is monitored, the abnormal disk partition is determined according to the error code; if the storage medium to which the abnormal disk partition belongs is an SSD, it is determined that there is an anomaly in the service layer; if the storage medium to which the abnormal disk partition belongs is not an SSD, it is determined that there is no anomaly in the service layer.
[0179] In an embodiment of the present invention, the above-mentioned determination module 1 may be specifically configured to monitor the startup process of the object storage device in the service layer of the distributed storage system to determine whether a mount failure event is monitored; where the mount failure event includes a disk partition mount failure event and / or a file system mount failure event; if a mount failure event is monitored, it is determined that there is an anomaly in the service layer; if no mount failure event is monitored, it is determined that there is no anomaly in the service layer.
[0180] In one embodiment of the present invention, the above-mentioned determination module 1 may specifically be used to monitor the startup process of the object storage device in the service layer of the distributed storage system and / or the read / write process of the object storage device to determine whether a hard disk removal event is monitored; if a hard disk removal event is monitored, it is determined that there is an abnormality in the service layer; if no hard disk removal event is monitored, it is determined that there is no abnormality in the service layer.
[0181] In one embodiment of the present invention, the above-mentioned determination module 1 may specifically be used to monitor the startup process of the object storage device in the service layer of the distributed storage system and / or the read / write process of the object storage device to determine whether a hard disk life expiration signal is monitored; if a hard disk life expiration signal is monitored, it is determined that there is an abnormality in the service layer; if no hard disk life expiration signal is monitored, it is determined that there is no abnormality in the service layer.
[0182] In one embodiment of the present invention, the above-mentioned detection module 3 may specifically be used to collect parameters of the abnormal solid-state drive to obtain hard disk parameters; wherein, the hard disk parameters include one or a combination of the power-on time, wear level, data write volume, etc.; if all hard disk parameters do not exceed the corresponding parameter thresholds, it is determined that the fault detection result is that the abnormal solid-state drive has no fault; if any hard disk parameter exceeds the corresponding parameter threshold, it is determined that the fault detection result is that the abnormal solid-state drive has a fault.
[0183] In one embodiment of the present invention, the above-mentioned detection module 3 may specifically be used to obtain the hard disk log corresponding to the abnormal solid-state drive in the system log; if there is no hard disk error log in the hard disk log, it is determined that the fault detection result is that the abnormal solid-state drive has no fault; if there is a hard disk error log in the hard disk log, it is determined that the fault detection result is that the abnormal solid-state drive has a fault.
[0184] In one embodiment of the present invention, the above-mentioned detection module 3 may specifically be used to determine whether an abnormal error message about the abnormal solid-state drive is received; if no abnormal error message about the abnormal solid-state drive is received, it is determined that the fault detection result is that the abnormal solid-state drive has no fault; if an abnormal error message about the abnormal solid-state drive is received, it is determined that the fault detection result is that the abnormal solid-state drive has a fault.
[0185] In one embodiment of the present invention, the solid-state drive management device may further include a control module, which is used to determine the faulty object storage device to which the abnormal solid-state drive belongs when the fault detection result is that the abnormal solid-state drive has a fault; and control the faulty object storage device to stop running.
[0186] In one embodiment of the present invention, the above-mentioned protection module 4 may include:
[0187] A first determination unit is configured to determine an abnormal scenario corresponding to an abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result; wherein, the abnormal scenario includes a scenario of continuous hard disk failures within a short period of time and / or a scenario of hard disk expiration in a state approaching the super fault domain; the scenario of continuous hard disk failures within a short period of time indicates that a preset number of solid-state drives in the distributed storage system have failed within a preset time period; the scenario of hard disk expiration in a state approaching the super fault domain indicates that the number of data reconstruction members in the object storage device to which the solid-state drive with expired lifespan belongs is not less than the minimum redundancy number of the object storage device.
[0188] A second determination unit is configured to determine a target protection policy according to the abnormal scenario corresponding to the abnormal solid-state drive, and perform redundancy protection on the abnormal solid-state drive by using the target protection policy.
[0189] In an embodiment of the present invention, the above-mentioned first determination unit may be specifically configured to, if the abnormal type of the solid-state drive is a hard disk self-abnormality and the fault detection result is that there is a fault in the abnormal solid-state drive, obtain the abnormal time node of the abnormal solid-state drive; if the time interval between the abnormal time node of the abnormal solid-state drive and the abnormal time node of the historical abnormal solid-state drive does not exceed the first duration, determine that the abnormal scenario is a scenario of continuous hard disk failures within a short period of time.
[0190] In an embodiment of the present invention, the above-mentioned first determination unit may be specifically configured to, if the abnormal type of the solid-state drive is hard disk expiration, determine the expired object storage device to which the abnormal solid-state drive belongs; if there is a placement group in the expired object storage device, determine that the abnormal scenario is a scenario of hard disk expiration in a state approaching the super fault domain.
[0191] In an embodiment of the present invention, when the abnormal scenario is a scenario of continuous hard disk failures within a short period of time, the above-mentioned second determination unit may be specifically configured to set an abnormal label for all object storage devices in the distributed storage system to block all write operations in all object storage devices, and output a first warning prompt.
[0192] In an embodiment of the present invention, when the abnormal scenario is a scenario of hard disk expiration in a state approaching the super fault domain, the above-mentioned second determination unit may be specifically configured to determine the abnormal object storage device to which the abnormal solid-state drive in the distributed storage system belongs; set an abnormal label for the abnormal object storage device to block all write operations in the abnormal object storage device, and output a second warning prompt.
[0193] In an embodiment of the present invention, the above-mentioned first determination unit may be specifically configured to determine the time node when the fault detection result is obtained; when the interval duration between the time node and the current time node reaches the second duration, determine the abnormal scenario corresponding to the abnormal solid-state drive according to the abnormal type of the solid-state drive and the fault detection result.
[0194] For the introduction of the device provided in the embodiments of the present invention, please refer to the above method embodiments, and the present invention will not be elaborated herein.
[0195] Embodiments of the present invention provide an electronic device.
[0196] Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of an electronic device provided by the present invention. The electronic device may include:
[0197] A memory 11 for storing computer programs;
[0198] A processor 10, which can implement the steps of any of the above solid-state drive management methods when executing the computer program.
[0199] As Figure 12 shown, it is a schematic diagram of the composition structure of an electronic device. The electronic device may include: a processor 10, a memory 11, a communication interface 12, and a communication bus 13. The processor 10, the memory 11, and the communication interface 12 all complete mutual communication through the communication bus 13.
[0200] In the embodiments of the present invention, the processor 10 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices, etc.
[0201] The processor 10 can call the program stored in the memory 11. Specifically, the processor 10 can execute the operations in the embodiments of the solid-state drive management method.
[0202] The memory 11 is used to store one or more programs. The program may include program codes, and the program codes include computer operation instructions. In the embodiments of the present invention, the memory 11 stores at least programs for implementing the following functions:
[0203] Monitor the service layer in the distributed storage system to determine whether there is an abnormality in the service layer;
[0204] When there is an abnormality in the service layer, determine the abnormal solid-state drive and the type of solid-state drive abnormality in the storage layer; the type of solid-state drive abnormality includes hard disk self-abnormality and / or hard disk life expiration;
[0205] When the type of solid-state drive abnormality is hard disk self-abnormality, use a preset fault detection strategy to perform fault detection on the abnormal solid-state drive to obtain a fault detection result;
[0206] Determine a target protection strategy based on the type of solid-state drive abnormality and the fault detection result, and use the target protection strategy to perform redundant protection on the abnormal solid-state drive.
[0207] In a possible implementation, the memory 11 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store data created during use.
[0208] In addition, the memory 11 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device or other volatile solid-state storage devices.
[0209] The communication interface 12 may be an interface of a communication module for connecting to other devices or systems.
[0210] Of course, it should be noted that Figure 12 The structure shown does not limit the electronic device in the embodiments of the present invention. In practical applications, the electronic device may include more or fewer components than Figure 12 those shown, or combine certain components.
[0211] The embodiments of the present invention provide a non-volatile storage medium.
[0212] The computer program stored on the non-volatile storage medium provided by the embodiments of the present invention can implement the steps of any of the above-mentioned solid-state drive management methods when executed by a processor.
[0213] Among them, the non-volatile storage medium may be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more integrated available media. For example, it may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape, etc.), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc., which can store computer program code.
[0214] For the introduction of the non-volatile storage medium provided by the embodiments of the present invention, please refer to the above method embodiments, and the present invention will not be elaborated here.
[0215] The embodiments of the present invention provide a computer program product.
[0216] The computer program product provided by the embodiments of the present invention includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, they can implement the steps of any of the above-mentioned solid-state drive management methods.
[0217] Specifically, in the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0218] Among them, the computer program product may include one or more computer programs / instructions. When the computer program / instructions are loaded and executed on a computer, they may wholly or partly generate the processes or functions described in the embodiments of the present invention. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a non-volatile storage medium or transmitted from one non-volatile storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line, etc.) or wirelessly (such as infrared, wireless, microwave, etc.).
[0219] For the introduction of the computer program product provided in the embodiments of the present invention, please refer to the above method embodiments, and the present invention will not be elaborated herein.
[0220] The embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.
[0221] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0222] The steps of the methods or algorithms described in combination with the embodiments disclosed in this document can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0223] The above has introduced the technical solution provided by the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A solid state hard disk management method, characterized in that: Applied to a distributed storage system, the distributed storage system includes a service layer and a storage layer, the storage layer is deployed with multiple object storage devices, each of the object storage devices is deployed with multiple solid-state hard disks, the method includes: Monitoring the service layer in the distributed storage system to determine whether there is an abnormality in the service layer; When an abnormality exists in the service layer, an abnormal solid-state hard disk and an abnormal type of the solid-state hard disk are determined in the storage layer; the abnormal type of the solid-state hard disk includes an abnormality of the hard disk itself and / or expiration of the hard disk life; When the abnormal type of the solid state hard disk is the abnormality of the hard disk itself, a preset fault detection strategy is used to perform fault detection on the abnormal solid state hard disk to obtain a fault detection result; Determine the abnormal scenario corresponding to the abnormal solid-state hard disk according to the abnormal type of the solid-state hard disk and the fault detection result; wherein the abnormal scenario includes a scenario of continuous hard disk failure in a short period of time and / or a scenario of hard disk expiration in a state close to a super fault domain; the scenario of continuous hard disk failure in a short period of time indicates that a preset number of solid-state hard disks in the distributed storage system fail within a preset time period; the scenario of hard disk expiration in a state close to a super fault domain indicates that the number of data reconstruction members in the object storage device to which the solid-state hard disk with expired life belongs is not less than the minimum redundancy number of the object storage device; A target protection strategy is determined according to the abnormal scenario corresponding to the abnormal solid state drive, and the target protection strategy is used to perform redundant protection on the abnormal solid state drive.
2. The solid state drive management method according to claim 1, characterized in that: Monitoring the service layer in the distributed storage system to determine whether the service layer has an abnormality includes: Monitoring the read and write process of the object storage device of the service layer in the distributed storage system to determine whether an input and output congestion event is monitored; If the input and output blocking event is monitored, it is determined that an abnormality exists in the service layer; If the input / output congestion event is not monitored, it is determined that there is no abnormality in the service layer.
3. The solid state drive management method according to claim 1, characterized in that: Monitoring the service layer in the distributed storage system to determine whether the service layer has an abnormality includes: Monitoring the read and write process of the object storage device of the service layer in the distributed storage system to determine whether an error code is monitored; If the error code is not monitored, it is determined that an abnormality exists in the service layer; If the error code is monitored, determining the abnormal disk partition according to the error code; If the storage medium to which the abnormal disk partition belongs is a solid state drive, it is determined that an abnormality exists in the service layer; If the storage medium to which the abnormal disk partition belongs is not a solid-state hard disk, it is determined that there is no abnormality in the service layer.
4. The solid state drive management method according to claim 1, characterized in that: Monitoring the service layer in the distributed storage system to determine whether the service layer has an abnormality includes: The object storage device startup process of the service layer in the distributed storage system is monitored to determine whether a mount failure event is monitored; wherein the mount failure event includes a disk partition mount failure event and / or a file system mount failure event; If the mount failure event is monitored, it is determined that an abnormality exists in the service layer; If the mount failure event is not monitored, it is determined that there is no abnormality in the service layer.
5. The solid state drive management method according to claim 1, characterized in that: Monitoring the service layer in the distributed storage system to determine whether the service layer has an abnormality includes: Monitoring the object storage device startup process and / or the object storage device read and write process of the service layer in the distributed storage system to determine whether a hard disk removal event is monitored; If the hard disk removal event is monitored, it is determined that an abnormality exists in the service layer; If the hard disk removal event is not monitored, it is determined that there is no abnormality in the service layer.
6. The solid state drive management method according to claim 1, characterized in that: Monitoring the service layer in the distributed storage system to determine whether the service layer has an abnormality includes: Monitoring the object storage device startup process and / or the object storage device read and write process of the service layer in the distributed storage system to determine whether a hard disk life expiration signal is monitored; If the hard disk life expiration signal is monitored, it is determined that the service layer is abnormal; If the hard disk life expiration signal is not monitored, it is determined that there is no abnormality in the service layer.
7. The solid state drive management method according to claim 1, characterized in that: Performing fault detection on the abnormal solid state drive using a preset fault detection strategy to obtain a fault detection result includes: Collecting parameters of the abnormal solid state hard disk to obtain hard disk parameters; wherein the hard disk parameters include a combination of one or more of power-on time, wear degree, and data writing amount; If all the hard disk parameters do not exceed the corresponding parameter thresholds, determining that the fault detection result is that the abnormal solid state hard disk has no fault; If any of the hard disk parameters exceeds the corresponding parameter threshold, the fault detection result is determined to be that the abnormal solid state hard disk is faulty.
8. The solid state drive management method according to claim 1, characterized in that: Performing fault detection on the abnormal solid state drive using a preset fault detection strategy to obtain a fault detection result includes: Obtaining a hard disk log corresponding to the abnormal solid state hard disk in the system log; If there is no hard disk error log in the hard disk log, determining that the fault detection result is that the abnormal solid state hard disk has no fault; If the hard disk error log exists in the hard disk log, it is determined that the fault detection result is that the abnormal solid state hard disk has a fault.
9. The solid state drive management method according to claim 1, characterized in that: Performing fault detection on the abnormal solid state drive using a preset fault detection strategy to obtain a fault detection result includes: Determining whether abnormal error information about the abnormal solid state hard disk is received; If no abnormal error information about the abnormal solid state hard disk is received, determining that the fault detection result is that the abnormal solid state hard disk has no fault; If abnormal error information about the abnormal solid state hard disk is received, it is determined that the fault detection result is that the abnormal solid state hard disk is faulty.
10. The solid state drive management method according to claim 1, characterized in that: Also includes: When the fault detection result indicates that the abnormal solid state hard disk is faulty, determining the fault object storage device to which the abnormal solid state hard disk belongs; Control the faulty object storage device to stop running.
11. The solid state drive management method according to claim 1, characterized in that: Determining an abnormal scenario corresponding to the abnormal solid state drive according to the abnormal type of the solid state drive and the fault detection result includes: If the abnormal type of the solid state hard disk is that the hard disk itself is abnormal, and the fault detection result is that the abnormal solid state hard disk is faulty, then obtaining the abnormal time node of the abnormal solid state hard disk; If the time interval between the abnormal time node of the abnormal solid state hard disk and the abnormal time node of the historical abnormal solid state hard disk does not exceed the first time length, it is determined that the abnormal scenario is a scenario of continuous hard disk failure within the short period of time.
12. The solid state drive management method according to claim 1, characterized in that: Determining an abnormal scenario corresponding to the abnormal solid state drive according to the abnormal type of the solid state drive and the fault detection result includes: If the abnormal type of the solid state hard disk is that the hard disk has expired, determining the expired object storage device to which the abnormal solid state hard disk belongs; If there is a placement group in the expired object storage device, it is determined that the abnormal scenario is a hard disk expiration scenario in the near super fault domain state.
13. The solid state drive management method according to claim 1, characterized in that: When the abnormal scenario is a scenario of continuous hard disk failures within a short period of time, redundant protection is performed on the abnormal solid-state hard disk using the target protection strategy, including: An abnormal tag is set for all the object storage devices in the distributed storage system to block all write operations in all the object storage devices and output a first alarm prompt.
14. The solid state drive management method according to claim 1, characterized in that: When the abnormal scenario is a hard disk expiration scenario in the state close to the super fault domain, using the target protection strategy to perform redundancy protection on the abnormal solid state hard disk includes: Determine the abnormal object storage device to which the abnormal solid state hard disk in the distributed storage system belongs; An abnormal tag is set for the abnormal object storage device to block all write operations in the abnormal object storage device and output a second alarm prompt.
15. The solid state drive management method according to claim 1, characterized in that: Determining an abnormal scenario corresponding to the abnormal solid state drive according to the abnormal type of the solid state drive and the fault detection result includes: Determine a time point for obtaining the fault detection result; When the interval duration between the time node and the current time node reaches a second duration, the abnormal scenario corresponding to the abnormal solid state drive is determined according to the abnormal type of the solid state drive and the fault detection result.
16. A solid state hard disk management device, characterized in that: Applied to a distributed storage system, the distributed storage system includes a service layer and a storage layer, the storage layer is deployed with multiple object storage devices, each of the object storage devices is deployed with multiple solid-state hard disks, and the apparatus includes: A monitoring module, used to monitor the service layer in the distributed storage system to determine whether there is an abnormality in the service layer; A determination module, configured to determine an abnormal solid-state drive and an abnormal type of the solid-state drive at the storage layer when an abnormality exists at the service layer; the abnormal type of the solid-state drive includes an abnormality of the hard drive itself and / or expiration of the hard drive life; A detection module, used for, when the abnormal type of the solid state hard disk is the abnormality of the hard disk itself, performing fault detection on the abnormal solid state hard disk using a preset fault detection strategy to obtain a fault detection result; A protection module is used to determine the abnormal scenario corresponding to the abnormal solid-state hard disk according to the abnormal type of the solid-state hard disk and the fault detection result; determine the target protection strategy according to the abnormal scenario corresponding to the abnormal solid-state hard disk, and use the target protection strategy to perform redundant protection on the abnormal solid-state hard disk; wherein, the abnormal scenario includes a scenario of continuous hard disk failure in a short period of time and / or a scenario of hard disk expiration in a state close to a super fault domain; the scenario of continuous hard disk failure in a short period of time indicates that a preset number of solid-state hard disks in the distributed storage system fail within a preset time period; the scenario of hard disk expiration in a state close to a super fault domain indicates that the number of data reconstruction members in the object storage device to which the solid-state hard disk whose life has expired belongs is not less than the minimum redundancy number of the object storage device.
17. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the solid state drive management method as described in any one of claims 1 to 15 when executing the computer program.
18. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the solid state drive management method according to any one of claims 1 to 15 are implemented.
19. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the solid state drive management method according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Distributed storage solid state disk processing method and device
CN116414661A