Hard disk management method and device, equipment and storage medium
Through the hard disk fault management system, the hard disk status data is automatically collected and parsed, and the fault status is determined and automatically repaired. The problems of low efficiency of hard disk troubleshooting and high operation and maintenance costs in cloud storage systems are solved, and efficient and automated hard disk management is achieved.
Patent Information
- Application Number
- CN202311461932.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
In existing cloud storage systems, hard disk troubleshooting is inefficient, manual operation and maintenance costs are high, and the perception of hard disk failure depends on the coupling relationship between storage services and business modules, resulting in waste of resources.
The hard disk failure management system is adopted to collect and parse the hard disk status data of the storage server, automatically determine the fault status of the hard disk, and automatically repair the hard disk according to the management strategy of the bound storage application.
It improves the accuracy and efficiency of hard disk troubleshooting, reduces manual operation and maintenance costs, reduces resource waste, and realizes the automated management of cloud storage systems.
Smart Images

Figure CN119937898A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a hard disk management method, device, equipment and storage medium. Background Art
[0002] A cloud storage system is a system that provides online storage functions. Specifically, a server cluster managed by a third party provides storage services for users. The storage service implements an access interface through a client to store files uploaded by users in the hard disk of the corresponding server. In order to meet the storage needs of users, the third party needs to ensure that the server cluster has stable storage capacity to provide storage services.
[0003] In the related art, in order to maintain the stability of the cloud storage system, manual operation and maintenance is usually adopted. When the business module in the storage service fails to operate normally, resulting in storage service abnormalities, the operating status of the hard disk is checked, the abnormal hard disk is found and repaired.
[0004] However, due to the huge number of hard disks in the server cluster, the above method has extremely high manpower costs and low operation and maintenance efficiency. In addition, the perception of abnormal hard disks depends on the degree of coupling between the storage service and the business module. If the coupling relationship is too tight or too loose, it will lead to waste of resources.
[0005] For example, if the coupling relationship between the storage service and the business module is too close, a slight jitter of the hard disk may cause the storage service to be abnormal, which may then incorrectly initiate the manual troubleshooting process and lead to a waste of manual resources.
[0006] For another example, if the coupling relationship between the storage service and the business module is too loose, even if the hard disk is completely unusable, it may result in the abnormal hard disk not being detected for a long time, thereby wasting hard disk resources.
[0007] Therefore, it is now urgent to redesign a hard disk management method and overcome the above-mentioned defects. Summary of the invention
[0008] The embodiments of the present application provide a hard disk management method, apparatus, device and storage medium to improve the accuracy of hard disk fault troubleshooting, reduce the resources consumed in hard disk fault troubleshooting and repair, and improve hard disk management efficiency.
[0009] In a first aspect, an embodiment of the present application provides a hard disk management method, which is applied to a hard disk fault management system. The hard disk fault management system is used to manage the hard disk operation status of each storage server in a cloud storage system, and records the binding relationship between each hard disk and the storage application that is allowed to access, as well as the management strategy of each storage application for the hard disk. The method includes:
[0010] For a storage server, hard disk status data of a target hard disk is collected; the storage server belongs to a server cluster for providing cloud storage services, and the hard disk status data represents: the usage of the target hard disk when reading and writing relevant data generated by the cloud storage service;
[0011] Parsing the hard disk status data to obtain a target parsing result, wherein the target parsing result indicates whether the target hard disk is in a normal working state;
[0012] When it is determined based on the target analysis result that the target hard disk is currently in a preset fault state, the hard disk identifier of the target hard disk is used to obtain a target management policy of a target storage application bound to the target hard disk;
[0013] Based on the target management strategy, a hard disk repair solution corresponding to the preset fault state is obtained, and the target hard disk is repaired using the hard disk repair solution.
[0014] In a second aspect, an embodiment of the present application provides a hard disk management device, which is applied to a hard disk fault management system. The hard disk fault management system is used to manage the hard disk operation status of each storage server in a cloud storage system, and records the binding relationship between each hard disk and the storage application that each hard disk is allowed to access, as well as the management strategy of each storage application for the hard disk. The device includes:
[0015] A data collection unit is used to collect hard disk status data of a target hard disk for a storage server; the storage server belongs to a server cluster for providing cloud storage services, and the hard disk status data represents: usage of the target hard disk when reading and writing relevant data generated by the cloud storage service;
[0016] A data analysis unit, used to analyze the hard disk status data to obtain a target analysis result, wherein the target analysis result indicates whether the target hard disk is in a normal working state;
[0017] A fault confirmation unit, configured to, when it is determined based on the target analysis result that the target hard disk is currently in a preset fault state, use the hard disk identifier of the target hard disk to obtain a target management policy of a target storage application bound to the target hard disk;
[0018] The hard disk repair unit is used to obtain a hard disk repair plan corresponding to the preset fault state based on the target management strategy, and use the hard disk repair plan to repair the target hard disk.
[0019] Optionally, the fault confirmation unit is specifically used to:
[0020] Based on the hard disk identifier, a preset number of historical analysis results are obtained from the hard disk status record;
[0021] When the target analysis result and the obtained historical analysis results meet the preset read-write abnormality condition, it is determined that the target hard disk is currently in a bad disk state; the read-write abnormality condition represents: the degree of read-write abnormality of the target hard disk when performing read-write data processing;
[0022] When the target analysis result and the historical analysis results meet the preset recognition abnormality condition, it is determined that the target hard disk is currently in a disk drop state; the recognition abnormality condition indicates that: when the target hard disk is recognized by the device where it is located, a recognition abnormality occurs;
[0023] When the target analysis result and the historical analysis results meet the preset required bandwidth, and the average response time of the target hard disk exceeds the preset duration, it is determined that the target hard disk is currently in a slow disk state.
[0024] Optionally, the preset fault state is a slow disk state or a bad disk state;
[0025] The hard disk repair unit is specifically used for:
[0026] Reading target data in the target hard disk, migrating the target data to a first other hard disk, and repairing the target hard disk using a hard disk repair solution corresponding to the dropped disk state or the bad disk state;
[0027] After the target hard disk is repaired by the hard disk repair solution, the method further includes:
[0028] Acquire storage data from a second other hard disk, and write the storage data into the repaired target hard disk, wherein the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
[0029] Optionally, the preset fault state is a slow disk state or a bad disk state;
[0030] The hard disk repair unit is specifically used for:
[0031] If the target data in the target hard disk cannot be read, then determine the associated data of the target data, and obtain the associated data from the associated hard disk where the associated data is located, the target data and the associated data belong to the same storage file, and are encoded in a preset encoding method and stored in the target hard disk and the associated hard disk respectively;
[0032] Decoding and restoring the associated data based on the preset encoding method to obtain the storage file, migrating the target data in the storage file to a first other hard disk, and using a hard disk repair solution corresponding to the dropped disk state or the bad disk state to repair the target hard disk;
[0033] The hard disk repair unit is also used for:
[0034] Acquire storage data from a second other hard disk, and write the storage data into the repaired target hard disk, wherein the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
[0035] Optionally, the preset fault state is a disk drop state;
[0036] The hard disk repair unit is also used for:
[0037] Obtaining a target business in a running state on the device where the target hard disk is located, and based on a custom logic in the target business, initiating a restart prompt message to the target object, wherein the target business is used to provide relevant services to the target object;
[0038] Then the method of using the hard disk repair solution to repair the target hard disk includes:
[0039] Restart the device where the target hard disk is located.
[0040] Optionally, the preset fault state is a disk drop state or a bad disk state;
[0041] The hard disk repair unit is also used for:
[0042] Adjusting a data collection time interval for collecting the hard disk status data from a first time interval to a second time interval, wherein the second time interval is shorter than the first time interval;
[0043] The hard disk repair unit is also used for:
[0044] Collect and analyze the hard disk status data of the repaired target hard disk to obtain the corresponding status analysis result;
[0045] When it is determined based on the status analysis result that the repaired target hard disk is currently in a normal working state, the data collection time interval is adjusted from the second time interval to the first time interval.
[0046] Optionally, the data parsing unit is further used to:
[0047] Based on the obtained operation data of each hard disk in the cloud storage system, statistical analysis is performed to obtain the hard disk operation status of each hard disk in the cloud storage system, wherein the hard disk operation status includes: hard disk basic status and system operation status;
[0048] The basic information of the hard disk includes: the machine model, whether it is within the warranty period, the fault type, and at least one of the business it belongs to;
[0049] The system operation status includes: at least one of the following: the number of newly failed hard disks per month, the failure rate, the warranty expiration rate, and the processing time of each process step in the hard disk repair process.
[0050] Optionally, the hard disk repair unit is further used to:
[0051] Comparing the hard disk identification of the repaired target hard disk with the hard disk identification of the target hard disk, and determining that the comparison result meets the comparison condition, the hard disk repair is successful;
[0052] Wherein, the contrast condition characterizes;
[0053] When the preset fault state is a slow disk state or a bad disk state, a comparison result of the hard disk identification of the repaired target hard disk and the hard disk identification of the target hard disk is that the identifications are different;
[0054] When the fault state is a disk drop state, a comparison result of the hard disk identification of the repaired target hard disk and the hard disk identification of the target hard disk is that the identifications are the same.
[0055] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any one of the hard disk management methods in the first aspect.
[0056] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program is run on a computer device, the computer program is used to enable the computer device to execute any one of the hard disk management methods in the first aspect.
[0057] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; when a processor of a computer device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computer device executes any one of the hard disk management methods in the above-mentioned first aspect.
[0058] The beneficial effects of this application are as follows:
[0059] The hard disk management method, apparatus, computer equipment and storage medium provided in the embodiments of the present application adopt a hard disk fault management system to manage the operation status of the hard disks in each storage server in the cloud storage system, so as to timely discover and automatically repair hard disk failures. The hard disk fault management system also records the binding relationship between each hard disk and the storage application that each is allowed to access, as well as the management strategy of each storage application for the hard disk, so that the management party of the storage application can customize the management strategy of the hard disk corresponding to the storage application, and adopt a management strategy that better meets the needs to manage the hard disk when the hard disk fails.
[0060] Specifically, for each target hard disk in the storage server, the hard disk status data generated by the target hard disk during read and write processing is collected. In this way, the automatically collected hard disk status data can accurately represent the usage and operating status of the target hard disk. Then, by analyzing the hard disk status data, the target analysis result of whether the target hard disk is in a normal working state can be accurately and quickly obtained.
[0061] Among them, in the hard disk fault management system, for each target storage application, a target management strategy of the target storage application and a hard disk repair plan under the target management strategy under each preset fault state are set. If the target analysis result indicates that the target hard disk is in an abnormal working state, the target storage application bound to the target hard disk and the preset fault state can be automatically confirmed, and then the target management strategy and the hard disk repair plan for the preset fault state are obtained, and the hard disk repair plan is used to repair the target hard disk.
[0062] Compared with the related art, which uses manual operation and maintenance to operate and maintain the cloud storage system, resulting in low operation and maintenance efficiency, requiring extremely high labor costs, and poor accuracy in detecting hard disk failures, the present application can automatically collect the hard disk status data of the target hard disk, automatically parse the target hard disk status data, obtain the target parsing result, and further automatically obtain the preset fault state of the target hard disk based on the target parsing result analysis, determine the hard disk repair plan based on the preset fault state of the target hard disk and the target management strategy of the bound target storage application, and repair the target hard disk. In this way, the method of automatically collecting data and statistically analyzing is adopted to free up the manpower required for manual operation and maintenance, reduce labor costs, reduce errors caused by manual troubleshooting, improve the accuracy of hard disk fault detection, and improve the operation and maintenance efficiency of the cloud storage system, so as to realize the automated management of the cloud storage system.
[0063] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0065] Figure 1 An optional schematic diagram of an application scenario provided in an embodiment of the present application;
[0066] Figure 2 A flowchart of a hard disk management method provided in an embodiment of the present application;
[0067] Figure 3 A flowchart of a data analysis method provided in an embodiment of the present application;
[0068] Figure 4 A flowchart of a data migration method in a hard disk repair process provided by an embodiment of the present application;
[0069] Figure 5 A flowchart of a method for obtaining target data when target data in a target hard disk is damaged provided in an embodiment of the present application;
[0070] Figure 6 A flowchart of a hard disk repair method provided in an embodiment of the present application;
[0071] Figure 7 A schematic diagram of the system architecture of a hard disk failure management system provided in an embodiment of the present application;
[0072] Figure 8 A flowchart of a hard disk management method provided in an embodiment of the present application;
[0073] Fig. 9 A schematic diagram of an overall situation page of a cloud storage system provided in an embodiment of the present application;
[0074] Fig.10 A schematic diagram of a bad disk distribution page in a cloud storage system provided in an embodiment of the present application;
[0075] Fig.11 A schematic diagram of a task processing timeliness status page in a cloud storage system provided in an embodiment of the present application;
[0076] Fig.12A schematic diagram of the changes in the number of work orders (fault orders corresponding to hard disk repair tasks) and manpower input, as well as the changes in the coverage rate of toad skin intervention and the length of work order processing time after adopting a hard disk identification reconciliation method provided in an embodiment of the present application;
[0077] Fig.13 A schematic diagram of a hard disk management device provided in an embodiment of the present application;
[0078] Fig.14 A schematic diagram of a hardware structure of a computer device to which an embodiment of the present application is applied;
[0079] Fig.15 A schematic diagram of the hardware composition structure of another computer device to which the embodiment of the present application is applied. DETAILED DESCRIPTION
[0080] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.
[0081] It is understandable that in the following specific implementations of the present application, when hard disk status data, storage data and other related data are involved, when the various embodiments of the present application are applied to specific products or technologies, relevant licenses or consents need to be obtained, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, when it is necessary to obtain relevant data, relevant volunteers can be recruited and relevant agreements on volunteer authorization data can be signed, and then the data of these volunteers can be used for implementation; or, by implementing within the scope of the authorized organization, the following implementation methods can be implemented by using the data of internal members of the organization to make relevant recommendations to internal members; or, the relevant data used in the specific implementation are all simulated data, such as simulated data generated in a virtual scene.
[0082] To facilitate understanding of the technical solutions provided in the embodiments of the present application, some key terms used in the embodiments of the present application are explained here:
[0083] Cloud Object Storage (COS), also known as object-based storage, is a general term used to describe methods of solving and processing discrete units. These discrete units are called objects. Just like files, objects contain data, but unlike files, objects no longer have a hierarchical structure in a layer structure. Each object is at the same level in a flat address space called a storage pool. One object does not belong to the next level of another object. It is a distributed storage service with no directory hierarchy, no data format restrictions, can accommodate massive data, and supports HTTP / HTTPS protocol access. COS has no capacity limit for bucket space and does not require partition management. It is suitable for a variety of scenarios such as CDN data distribution, data processing, or data lakes for big data computing and analysis.
[0084] Cloud Infinite (CI) is a data processing platform based on COS, covering image processing, content review, media processing, AI recognition, document preview and other functions, providing users with a one-stop professional data processing solution to meet the needs of various business scenarios.
[0085] EC coding (Erasure Code) is a data coding technology used to improve the reliability and availability of data. It divides the data into multiple blocks and encodes each block to generate multiple redundant blocks, thereby achieving redundant backup and recovery of data. For example, n copies of original data can be encoded and m copies of redundant data can be added. The original data can be restored by any n copies of data from the obtained n+m copies.
[0086] Iostat is used to report statistics about the central processing unit (CPU) and input / output related data of the entire system, adapters, tty devices, disks, and CD-ROMs. Specifically, the iostat command is used to observe the current system's input / output device load, CPU load, and hard disk load, and to generate statistics and reports based on the observed related information. The first report generated by the iostat command provides statistics about the time since the system was booted. Each subsequent report contains the time since the last report. Each time the iostat command is run, all statistics are reported. The report consists of a tty and CPU header line followed by a line of tty and CPU statistics. On a multi-processor system, CPU statistics are calculated system-wide and are the average of all processors. The report generated by the iostat command can be used to change the system configuration to better balance the input / output load between physical disks and adapters.
[0087] Host Bus Adapter (HBA) is a card that can be inserted into a computer or mainframe. Fibre Channel HBA is an essential device for connecting a host to a FC network.
[0088] Fibre Channel (FC) is a high-speed network technology standard (T11) mainly used in storage networks.
[0089] Data reliability refers to the degree to which data will not be lost or damaged during data storage and transmission. It is usually expressed as a percentage, for example, 99.999% reliability means that only one out of every million data may be lost or damaged. Data reliability is critical to the security and integrity of data. If data is lost or damaged, it may lead to information loss, business interruption, financial loss and other problems. Therefore, data reliability is one of the important indicators of storage IAAS data management and storage security.
[0090] In order to ensure the reliability of data, a variety of methods are usually adopted, such as backup, redundant storage, data verification, etc. In addition, storage devices need to be regularly inspected and maintained to ensure their normal operation and data integrity.
[0091] Infrastructure as a Service (IaaS) refers to a service model that provides IT infrastructure as a service through the Internet and charges users based on their actual usage or occupation of resources. In this service model, ordinary users do not need to build a data center or other hardware facilities by themselves, but instead obtain computer infrastructure services, including servers, storage, and network services, from IaaS service providers through leasing through the Internet.
[0092] Due to the natural wear and tear of hardware and business use, hard disk failures may occur. The timeliness of hard disk failure processing and the magnitude of simultaneous failures directly affect data reliability. Each bad disk needs to be manually detected and a bill of lading issued after follow-up on-site repair before it can continue to be used.
[0093] The technical solutions of the embodiments of the present application involve computer technology and cloud storage technology. Computer technology refers to the technical methods and technical means used in the computer field, or refers to its hardware technology, software technology and application technology. Computer technology is the technology used by a complete system. There are mainly system structure technology, system management technology, system maintenance technology and system application technology. Among them, in system management technology, computer system management automation is realized by the operating system. The basic purpose of the operating system is to make the most effective use of the software and hardware resources of the computer to improve the throughput capacity and problem-solving time of the machine, facilitate operation and use, improve the reliability of the system, and reduce the cost of solving problems. In system maintenance technology, the technology of automatic maintenance and diagnosis of computer systems. The main software for implementing maintenance and diagnosis automation is the function check program and the automatic diagnosis program. The function check program tests all the microscopic functions of various components of the computer system with strict data graphics or action retries and compares the correctness of the results to determine whether the components are working properly.
[0094] In practice, the application of computer technology is very extensive. On the basis of computer technology, cloud computing technology and cloud storage technology can be further realized. Cloud storage is a new concept extended and developed from the concept of cloud computing. Distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.
[0095] At present, the storage method of the storage system is: create a logical volume, and when creating a logical volume, allocate physical storage space for each logical volume, and the physical storage space may be composed of a disk of a storage device or several storage devices. The client stores data on a logical volume, that is, stores the data on the file system. The file system divides the data into many parts, each of which is an object. The object not only contains data but also contains additional information such as data identification (ID, ID entity). The file system writes each object into the physical storage space of the logical volume, and the file system records the storage location information of each object, so that when the client requests to access the data, the file system can allow the client to access the data according to the storage location information of each object. The cloud storage system involved in the present application can be composed of a storage server cluster including at least one storage server, so as to provide cloud storage services to the user objects based on the storage server cluster. In the process of providing cloud storage services, the user objects can access the hard disk (disk) in the storage server to achieve data writing and data reading. Then, the hard disk fault management system can manage each hard disk in the cloud storage system by observing the operation of the hard disk.
[0096] The following is a brief introduction to the design concept of the embodiment of the present application:
[0097] The COS of the cloud storage system has a huge amount of data, reaching the EB level. EB (Exabyte) is a unit of computer storage capacity, which means ten to the 18th power of bytes. This data needs to be carried by millions of storage hard disks. Therefore, the storage servers in the storage server cluster are generally customized storage servers. A large number of hard disks are set in the customized storage servers to obtain a larger amount of storage data and provide good cloud storage services. Correspondingly, how to operate and maintain hard disks with such a huge amount of data has become an urgent problem to be solved.
[0098] At present, in the related technology, manual operation and maintenance is usually adopted to operate and maintain the hard disks in the cloud storage system. When the cloud storage service generates an abnormal alarm, each hard disk corresponding to the relevant business module that generates the abnormal alarm can be checked to find the abnormal hard disk and repair it. It can be seen that in the actual implementation of this method, the operation and maintenance personnel need to have a certain understanding of the cloud storage system and need to be proficient in the relevant knowledge of hard disk technology. In the face of the operation and maintenance of such a large number of hard disks, it is a relatively difficult problem to gather enough operation and maintenance personnel, which takes up a lot of manpower and material resources. In addition, manual operation and maintenance are bound to make mistakes, resulting in low accuracy and efficiency of operation and maintenance.
[0099] In addition, the degree of coupling between the cloud storage service and each business module is different; if the coupling between the cloud storage service and the business module is too tight, then even a slight jitter of the hard disk corresponding to the business module will cause the cloud storage service to generate an abnormal alarm, resulting in unnecessary resource consumption; if the coupling between the cloud storage service and the business module is too loose, then even if the hard disk corresponding to the business module fails, the cloud storage service may not be able to generate an alarm in time, resulting in a decline in the quality of the cloud storage service.
[0100] In view of this, the embodiment of the present application provides a hard disk management method, device, computer equipment and storage medium, which adopts a hard disk fault management system, and collects the hard disk status data generated by the target hard disk during the read and write processing for each target hard disk in each storage server in the storage server cluster of the cloud storage system. In this way, the automatically collected hard disk status data can accurately represent the usage and operation status of the target hard disk, and the hard disk status data can be analyzed to accurately and quickly obtain the target analysis result of whether the target hard disk is in a normal working state. Among them, for each storage application, the hard disk fault management system is provided with the management strategy of the storage application for the hard disk, and the hard disk repair solution under each preset fault state contained in the management strategy. If the target analysis result represents that the target hard disk is in an abnormal working state, the target storage application bound to the target hard disk and the preset fault state in which it is can be automatically confirmed, and then the target management strategy and the hard disk repair solution under the preset fault state are obtained, and the hard disk repair solution is used to repair the target hard disk.
[0101] Compared with the related art, which uses manual operation and maintenance to operate and maintain the cloud storage system, resulting in low operation and maintenance efficiency, extremely high labor costs, and poor accuracy in detecting hard disk failures, the present application can automatically collect the hard disk status data of the target hard disk, automatically parse the target hard disk status data, obtain the target parsing result, and further automatically obtain the preset fault state of the target hard disk based on the target parsing result analysis, determine the hard disk repair plan based on the target management policy of the target storage application to which the target hard disk belongs, and the preset fault state of the target hard disk, and repair the target hard disk. In this way, the method of automatically collecting data and statistically analyzing is adopted to free up the manpower required for manual operation and maintenance, reduce labor costs, reduce errors caused by manual troubleshooting, improve the accuracy of hard disk fault detection, and improve the operation and maintenance efficiency of the cloud storage system, so as to realize automated management of the cloud storage system.
[0102] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application may be combined with each other if there is no conflict.
[0103] like Figure 1 As shown, it is a schematic diagram of an application scenario of an embodiment of the present application. The application scenario diagram includes any terminal device 110 among multiple terminal devices, any server 120 (storage server) among multiple servers, and any server 130 among multiple servers.
[0104] In the embodiment of the present application, the terminal device 110 includes but is not limited to mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, smart home appliances, car terminals and other devices; a client related to the cloud storage service may be installed on the terminal device, and the client may be software (such as a browser, communication software, etc.), or a web page, a small program, etc. The server 120 is a background server corresponding to the software or web page, small program, etc., or a background server specifically used to provide cloud storage services to the client, which is not specifically limited in the present application. The server 130 corresponds to the server 120 and is used to manage the hard disk in the server 120, including statistical analysis of hard disk operation-related data, analysis of hard disk working status, and repair of faulty hard disks, etc., which is not specifically limited in the present application. The server 120 and the server 130 may be independent physical servers, or a server cluster or distributed system composed of multiple physical servers, or may provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Networks), etc. Network, CDN), as well as cloud servers for basic cloud computing services such as big data and artificial intelligence platforms.
[0105] It should be noted that the hard disk management method in the embodiment of the present application can be executed by the server 130. The server 130 is provided with storage applications to which each hard disk belongs, and management policies of each storage application. The management policy of each storage application is provided with a hard disk repair plan corresponding to each preset fault state of the hard disk bound to it. The hard disk status data of the target hard disk in the server 120 is collected, and the hard disk status data is parsed to obtain a target analysis result representing whether the target hard disk is in a normal working state. Based on the target analysis result, when it is determined that the target hard disk is currently in a preset fault state, the hard disk identifier of the target hard disk is used to obtain the target management policy of the target storage application bound to the target hard disk. Based on the target management policy, the hard disk repair plan corresponding to the preset fault state is obtained, and the hard disk repair plan is used to repair the target hard disk.
[0106] It should be noted that Figure 1 What is shown is just an example. In fact, the number and communication mode of terminal devices and servers are not limited and are not specifically limited in the embodiments of the present application.
[0107] In addition, the hard disk management method of the embodiment of the present application can be applied to various scenarios, such as hard disk management in a cloud storage system corresponding to reading software, hard disk management in a cloud storage system corresponding to video software, hard disk management in a cloud storage system corresponding to a file database, and other scenarios.
[0108] The hard disk management method provided by the exemplary embodiment of the present application is described below in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present application, and the implementation methods of the present application are not limited in this regard.
[0109] See also Figure 2 As shown, it is a flow chart of the hard disk management method provided in the embodiment of the present application. Here, the server is used as an example for explanation. It is applied to the hard disk fault management system. The hard disk fault management system is used to manage the hard disk operation status of each storage server in the cloud storage system. It records the binding relationship between each hard disk and the storage application that each storage application is allowed to access, as well as the management strategy of each storage application for the hard disk. The specific implementation process of the method is as follows:
[0110] Step 201: For a storage server, collect hard disk status data of a target hard disk; the storage server belongs to a server cluster for providing cloud storage services, and the hard disk status data represents: usage of the target hard disk when reading and writing relevant data generated by the cloud storage service;
[0111] In one embodiment, a cloud storage system is operated based on a storage server cluster and is used to provide cloud storage services to users. The storage server cluster includes at least one storage server, and the storage server includes at least one hard disk. The user can use the cloud storage service to read data from the hard disk of the storage server, or use the cloud storage service to store data in the hard disk of the storage server. Each hard disk in each storage server can be a target hard disk.
[0112] In one embodiment, an operating system is installed in the storage server, and the operating system contains data generated by the hard disk during operation, such as hard disk identification information, slot information of data stored in the hard disk, related information about the hard disk being read and written, physical tracks and logical tracks of the hard disk, etc. The aforementioned related data in the operating system can be collected as hard disk status data and reported to the hard disk fault management system.
[0113] In one embodiment, the hard disk fault management system may collect hard disk status data from the operating system of the storage server by injecting an Agent into the cloud storage system.
[0114] In one embodiment, the hard disk fault management system may also install part of the data collection subsystem into the storage server of the cloud storage system to collect hard disk status data.
[0115] In one embodiment, the usage of the target hard disk may be: whether the target hard disk fails, and whether the relevant data of the hard disk failure meets the preset failure state. If the target hard disk meets the preset failure state, the target hard disk has a failure corresponding to the preset failure state.
[0116] In one embodiment, the cloud storage system can provide cloud storage services to multiple storage applications, and the multiple hard disks of each storage server can be bound to their respective storage applications. When the user uses the cloud storage service of the storage application, the user can read and write at least one hard disk bound to the storage application. Each hard disk in each storage server can be a target hard disk.
[0117] Step 202: parse the hard disk status data to obtain a target parsing result, which indicates whether the target hard disk is in a normal working state.
[0118] In one embodiment, hard disk status data such as hard disk identification information, slot information of data stored in the hard disk, related information of the hard disk being read and written, and the status of the physical track and logical track of the hard disk are parsed to obtain a target parsing result, wherein the target parsing result may indicate a read and write anomaly, or the hard disk cannot be identified, or the hard disk identification does not match the hard disk slot number in the storage server, etc.
[0119] In one embodiment, the parsing method used to parse the hard disk status data can be a hard disk fault management system, a unified parsing method for the hard disk status data of all hard disks in the cloud storage system, or a parsing method corresponding to each storage application. The object corresponding to the parsing method (hard disk fault management system or storage application, etc.) is not specifically limited here. Taking the parsing method corresponding to each storage application as an example, after obtaining the hard disk status data, the parsing method of the hard disk status data can be determined according to the target storage application to which the target hard disk of the hard disk status data belongs. For example, the parsing method of target storage application 1 includes: determining whether the target hard disk has a read and write error based on the read and write data of the target hard disk, or determining the average response time of the target hard disk read and write based on the read and write data of the target hard disk. The parsing method of target storage application 2 includes: determining whether the target hard disk has a read and write error based on the read and write data of the target hard disk, and determining the average response time of the target hard disk read and write based on the read and write data of the target hard disk, etc. The specific parsing method of each target storage application is not limited here. If the hard disk fault management system uses the same analysis method for all hard disks in the cloud storage system, then after obtaining the hard disk status data, the hard disk status data can be directly analyzed. Similarly, the analysis method here is the same as mentioned above, which can be based on the read and write data of the target hard disk to determine whether the target hard disk has a read or write error, based on the read and write data of the target hard disk, determine the average response time of the target hard disk reading and writing, etc.
[0120] In one embodiment, the hard disk status data is parsed to obtain a target parsing result. It is also possible to obtain the weekly failure rate of the target hard disk (the hard disk status data of the target hard disk is collected at a set time interval, and the parsing result of the hard disk status data collected within a week is the ratio of the target hard disk failure to the total number of hard disk status data collection within a week), the monthly failure rate (the hard disk status data of the target hard disk is collected at a set time interval, and the parsing result of the hard disk status data collected within a month is the ratio of the target hard disk failure to the total number of hard disk status data collection within a month), the daily failure rate, the ratio of each failure type (the parsing result is the number of parsing results of bad disk (or slow disk, or dropped disk, etc.) failures / target parsing result + the number of historical parsing results obtained = the ratio of bad disk (or slow disk, or dropped disk, etc.) failure types), etc.
[0121] Step 203: when it is determined based on the target analysis result that the target hard disk is currently in a preset fault state, the hard disk identifier of the target hard disk is used to obtain a target management policy of a target storage application bound to the target hard disk.
[0122] In one embodiment, there may be multiple preset fault states, and different target analysis results may correspond to different preset fault states. For example, if the target analysis result is that the target hard disk cannot be identified, then the corresponding preset fault state is a dropped disk. For another example, if the target analysis result is that the read and write data of the target hard disk is abnormal, then the corresponding preset fault state is a bad disk, etc. For another example, if the target analysis result is that the average response time of the read and write of the target hard disk exceeds the set time, then the corresponding preset fault state is a slow disk.
[0123] In one embodiment, based on the target analysis result, determining that the target hard disk is currently in a preset fault state may include: obtaining a fixed number of historical analysis results, and determining that the target hard disk is currently in a preset fault state based on the target analysis result and the fixed number of historical analysis results, or obtaining historical analysis results within a preset time length before the target analysis result, and determining that the target hard disk is currently in a preset fault state based on the target analysis result and the historical analysis results within a preset time length, etc. There is no specific restriction on the method of obtaining the historical analysis results.
[0124] In one embodiment, a hard disk fault management system is provided with a management strategy of each storage application, and the management strategies of different storage applications may be the same or different. For example, the management strategy 1 of the target storage application 1 may be that if the average response time of reading and writing the target hard disk 1 exceeds 5s, it is confirmed that the target hard disk 1 needs to be replaced; the management strategy 2 of the target storage application 2 may be that if the average response time of reading and writing the target hard disk 2 exceeds 5s, it is confirmed that the target hard disk 2 needs to be replaced; the management strategy 3 of the target storage application 3 may be that if the average response time of reading and writing the target hard disk 3 exceeds 5s but does not exceed 10s, it is confirmed that the target hard disk 3 does not need to be replaced.
[0125] In one embodiment, an observation period is set in the hard disk fault management system, that is, when it is determined based on the target analysis result that the target hard disk is currently in a preset fault state, the timing is triggered. During the observation period, if the target hard disk obtained from the cloud storage system is still in the preset fault state, the hard disk identifier of the target hard disk is used to obtain the target management policy of the target storage application bound to the target hard disk, and the target hard disk is repaired based on the corresponding hard disk repair plan under the target management policy, so that a brief disk jitter triggers the repair process.
[0126] Based on the method flow, the present application embodiment provides a data analysis method, such as Figure 3 As shown, in step 203, based on the target analysis result, it is determined that the target hard disk is currently in a preset fault state, including:
[0127] Step 301: Based on the hard disk identification, a preset number of historical analysis results are obtained from the hard disk status record.
[0128] In one embodiment, each time the hard disk fault management system obtains an analysis result, the analysis result and its corresponding hard disk identification are recorded in the hard disk status record. Then, the fault management system obtains the hard disk status data and parses it to obtain the target analysis result. In order to accurately obtain the preset fault state that the target hard disk is currently in, a preset number of historical analysis results are obtained from the hard disk status record.
[0129] In one embodiment, a preset number of historical analysis results may be obtained by obtaining a preset number of historical analysis results closest to the current moment within a set time period before the current moment, or by evenly dividing the set time period before the current moment into a preset number of portions to obtain a preset number of divided time periods, and randomly selecting a historical analysis result from each divided time period, or by obtaining all historical analysis results within the set time period before the current moment, etc. There is no specific limitation on the method of obtaining the preset number of historical analysis results here, and it can be set as needed.
[0130] Step 3021: When the target analysis result and the obtained historical analysis results meet the preset read / write abnormality condition, it is determined that the target hard disk is currently in a bad disk state; the read / write abnormality condition represents: the degree of read / write abnormality of the target hard disk when performing read / write data processing.
[0131] In one embodiment, the read / write abnormality condition represents: the degree of read / write abnormality of the target hard disk when performing read / write data processing;
[0132] The degree of read / write anomaly may be the proportion of the analysis results that are read / write anomaly results in all the analysis results. Assuming that the number of target analysis results and historical analysis results is 10, and the number of analysis results that are read / write anomaly results is 8, the degree of read / write anomaly is 80%;
[0133] Alternatively, the read / write anomaly degree can be: obtain the proportion of all analysis results that are read / write anomaly results, and calculate the obtained proportion + current traffic value / traffic weight adjustment parameter = read / write anomaly degree. The method for obtaining the read / write anomaly degree is not specifically limited here and can be set as needed.
[0134] In one embodiment, when the target analysis result and the historical analysis results obtained meet the preset read-write abnormality condition, it is determined that the target hard disk is currently in a bad disk state. In other words, if the read-write abnormality degree obtained by the target analysis result and the historical analysis results obtained exceeds the read-write abnormality degree in the read-write abnormality condition, it is determined that the target hard disk is currently in a bad disk state.
[0135] Step 3022: When the target analysis result and each historical analysis result meet the preset recognition abnormality condition, it is determined that the target hard disk is currently in a disk drop state; the recognition abnormality condition indicates that: when the target hard disk is recognized by the device where it is located, a recognition abnormality occurs;
[0136] In one embodiment, the recognition abnormality condition is characterized by: when the target hard disk is recognized by the device where it is located, a recognition abnormality occurs;
[0137] The situation where the recognition anomaly occurs can be the proportion of the unrecognizable results in all the analysis results. Assuming that the number of target analysis results and historical analysis results is 10, and the number of unrecognizable analysis results is 8, it is confirmed that the recognition anomaly occurs;
[0138] Alternatively, the situation where an abnormal identification occurs can be: among all the analysis results, the proportion of the analysis results that cannot be identified is obtained, and the obtained proportion + current flow value / flow weight adjustment parameter = X, X exceeds X', then the abnormal identification is confirmed. The specific method for confirming the abnormal identification is not limited here, and can be set as needed.
[0139] In one embodiment, when the target analysis result and the historical analysis results obtained meet the preset abnormal identification conditions, it is determined that the target hard disk is currently in a disk-dropped state. In other words, if the identification conditions obtained when the target analysis result and the historical analysis results obtained meet the abnormal identification conditions, it is determined that the target hard disk is currently in a disk-dropped state.
[0140] Step 3023: When the target analysis result and each historical analysis result meet the preset required bandwidth, and the average response time of the target hard disk exceeds the preset duration, it is determined that the target hard disk is currently in a slow disk state.
[0141] In one embodiment, if the access traffic corresponding to the target hard disk is less than the preset required bandwidth, it is confirmed that the bandwidth of the target hard disk is relatively sufficient and the bandwidth will not affect the access traffic of the target hard disk. However, in this case, the average response time of the target hard disk still exceeds the preset duration, then it is determined that the target hard disk is currently in a slow disk state.
[0142] In one embodiment, if the access traffic corresponding to the target hard disk is not less than the preset required bandwidth, it is confirmed that the bandwidth of the target hard disk affects the access traffic of the target hard disk. In this case, even if the target hard disk is in a healthy state, the average response time may exceed the preset duration. Therefore, when the access traffic is less than the preset required bandwidth, it is possible to re-judge whether the preset fault state of the target hard disk is a slow disk state.
[0143] In one embodiment, when the target resolution result and each historical resolution result meet the preset required bandwidth, and the average response time of the target hard disk exceeds the preset duration, it is determined that the target hard disk is currently in a slow disk state. For example, if the iostat command detects that the target hard disk has an average response time of the program util of the target hard disk that exceeds 97% of the normal response time for 15% of the time in a day, but the iops is less than 30 or the bandwidth is less than 10MB, it is considered a slow disk. The normal response time can be obtained based on the historical response time of the target hard disk in the normal state, or it can be set based on experience, and there is no limitation here.
[0144] In one embodiment, a preset duration of the traffic peak period can also be set based on the average response time of a healthy target hard disk when the access traffic is large. If the access traffic corresponding to the target hard disk is not less than the preset required bandwidth, and the average response time of the target hard disk still exceeds the preset duration of the traffic peak period, it is determined that the target hard disk is currently in a slow disk state.
[0145] Step 204: Based on the target management strategy, a hard disk repair solution corresponding to the preset fault state is obtained, and the target hard disk is repaired using the hard disk repair solution.
[0146] In one embodiment, a management strategy corresponding to each storage application is set in the hard disk fault management system. Different storage applications can correspond to the same or different management strategies. The target storage application bound to the target hard disk has its corresponding target management strategy, and under the target management strategy of the target storage application, the target hard disk bound to the target storage application has hard disk repair solutions corresponding to different preset fault states.
[0147] For example, based on the target analysis results, when it is determined that the target hard disk is in a dropped state, the target storage application to which it belongs is determined based on the hard disk identifier of the target hard disk, and the target management policy of the target storage application is further obtained. According to the hard disk repair plan in the dropped state according to the target management policy (assuming that the storage server of the target hard disk is restarted), a restart will be performed.
[0148] In one embodiment, the hard disk repair solution includes: restarting the storage server where the target hard disk is located, replacing the target hard disk, etc.
[0149] Based on the above method flow, the embodiment of the present application provides a method for data migration during hard disk repair, such as Figure 4 As shown, the preset fault state is a slow disk state or a bad disk state; in step 204, a hard disk repair solution is used to repair the target hard disk, including:
[0150] Step 401: read target data in a target hard disk, migrate the target data to a first other hard disk, and use a hard disk repair solution corresponding to a dropped disk state or a broken disk state to repair the target hard disk;
[0151] In step 204, after the target hard disk is repaired using the hard disk repair solution, the method further includes:
[0152] Step 402: Acquire storage data from the second other hard disk, and write the storage data into the repaired target hard disk. The first other hard disk and the second other hard disk may be the same hard disk or different hard disks.
[0153] In one embodiment, if the target hard disk is determined to be in a slow disk state or a bad disk state, the hard disk repair solution is to replace the target hard disk. Accordingly, in order to ensure that the target data stored in the target hard disk is not lost, the target data is migrated to the first other hard disk, and then the target hard disk is replaced with a new target hard disk. The first other hard disk is a healthy hard disk other than the target hard disk in the cloud storage.
[0154] In one embodiment, if it is determined that the target hard disk is in a slow disk state or a bad disk state, the target hard disk is tested to determine the damaged components in the target hard disk, such as if the distance between the magnetic head and the target hard disk is too large, then the target hard disk needs to be repaired, and the hard disk repair solution is to adjust the distance between the magnetic head and the target hard disk, then the target data is migrated to the first other hard disk, and then the distance between the target hard disk and the magnetic head is adjusted. The first other hard disk is a healthy hard disk other than the target hard disk in the cloud storage.
[0155] In one embodiment, after obtaining the repaired target hard disk, the stored data in the second other hard disk is written into the repaired target hard disk, and the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
[0156] In one embodiment, the first other hard disk may be a hard disk that stores a smaller amount of data or has a faster average response time among the hard disks of the cloud storage system, and the second other hard disk may be a hard disk that stores a larger amount of data or has a longer average response time among the hard disks of the cloud storage system.
[0157] In one embodiment, before reading the target data in the target hard disk and migrating the target data to the first other hard disk, the target hard disk may be tested for status. If the target hard disk is in a healthy state, the repair process of the target hard disk will not be continued, so as to avoid unnecessary repair processes on a large scale due to software problems, manpower problems, etc. For example, if the target hard disk is in a slow disk state, a dropped disk state, or a broken disk state due to a software update, the target hard disk will automatically return to a normal state after the software update is completed. If the target hard disk repeatedly issues an alarm that it is in a slow disk state or a broken disk state due to a lack of manpower to handle hard disk replacement during holidays, the target hard disk that has issued an alarm may be repaired after the holidays when there is sufficient manpower.
[0158] Compared with the related art, the above method uses manual operation and maintenance to operate and maintain the cloud storage system, which leads to low operation and maintenance efficiency, requires extremely high labor costs, and has poor accuracy in detecting hard disk failures. The above method can automatically collect the hard disk status data of the target hard disk, automatically parse the target hard disk status data, obtain the target parsing result, and further automatically obtain the preset fault state of the target hard disk based on the target parsing result analysis, determine the hard disk repair plan based on the target management policy of the target storage application to which the target hard disk belongs, and the preset fault state of the target hard disk, and repair the target hard disk. In this way, the method of automatically collecting data and statistically analyzing it can free up the manpower required for manual operation and maintenance, reduce labor costs, reduce errors caused by manual troubleshooting, improve the accuracy of hard disk fault detection, and improve the operation and maintenance efficiency of the cloud storage system, thereby realizing automated management of the cloud storage system.
[0159] Based on the above method flow, the embodiment of the present application provides a method for obtaining target data when the target data in the target hard disk is damaged, such as Figure 5 As shown, the preset fault state is a slow disk state or a bad disk state;
[0160] In step 204, the target hard disk is repaired using a hard disk repair solution, including:
[0161] Step 501: If the target data in the target hard disk cannot be read, the associated data of the target data is determined, and the associated data is obtained from the associated hard disk where the associated data is located. The target data and the associated data belong to the same storage file and are encoded using a preset encoding method and stored in the target hard disk and the associated hard disk respectively.
[0162] Step 502: decode and restore the associated data based on a preset encoding method to obtain a storage file, migrate the target data in the storage file to the first other hard disk, and use a hard disk repair solution corresponding to the disk drop state or the bad disk state to repair the target hard disk.
[0163] In step 204, after the target hard disk is repaired using the hard disk repair solution, the following steps are also included:
[0164] Step 503: Acquire storage data from the second other hard disk, and write the storage data into the repaired target hard disk. The first other hard disk and the second other hard disk may be the same hard disk or different hard disks.
[0165] In one embodiment, if the target data in the target hard disk cannot be read, the target data may be damaged. In order to ensure the integrity of the target data, associated data belonging to the same storage file as the target data is obtained, and the associated data is read from the associated hard disk. A preset encoding method (EC encoding or data repair program, etc., in the related technology, the method of decoding and restoring partial data after encoding to obtain all data is currently very mature and will not be repeated here) is used to decode and restore the associated data to obtain the storage file, extract the target data from the storage file, and migrate the target data to the first other hard disk. After obtaining the repaired target hard disk, the storage data obtained from the second other hard disk is written to the repaired target hard disk.
[0166] Based on the above method flow, the present application embodiment provides a hard disk repair method, such as Figure 6 As shown, the preset fault state is a disk drop state; in step 204, before the hard disk repair solution is used to repair the target hard disk, it includes:
[0167] Step 601: obtaining a target service in operation on the device where the target hard disk is located, and sending a restart prompt message to the target object based on the custom logic in the target service, wherein the target service is used to provide related services to the target object;
[0168] Then the method of using the hard disk repair solution to repair the target hard disk includes:
[0169] Step 602: Restart the device where the target hard disk is located.
[0170] In one embodiment, the device where the target hard disk is located is a storage server or a corresponding storage service device in the storage server, etc., and at least one business logic (including the target business) is running in the device where the target hard disk is located, such as the business logic for providing read services and the business logic for write services. When it is determined that the device where the target hard disk is located needs to be restarted, for the target business that is running on the device where the target hard disk is located, based on the custom logic in the target business, a restart prompt message is sent to the target object to ensure that the user object obtains the restart information, stops using the related services in the cloud storage service such as the read service or write service provided by the target business, and stops accessing the target hard disk, so as to ensure that the user object has a good experience in using the cloud storage service. Based on the above Figure 2 In the method flow, the embodiment of the present application provides a data collection method, and the preset fault state is a disk drop state or a bad disk state;
[0171] In step 204, before the target hard disk is repaired using the hard disk repair solution, the method further includes:
[0172] Adjusting a data collection time interval for collecting hard disk status data from a first time interval to a second time interval, wherein the second time interval is shorter than the first time interval;
[0173] After the target hard disk is repaired by the hard disk repair solution, the method further includes:
[0174] Collect and analyze the hard disk status data of the repaired target hard disk to obtain the corresponding status analysis result;
[0175] When it is determined based on the status analysis result that the repaired target hard disk is currently in a normal working state, the data collection time interval is adjusted from the second time interval to the first time interval.
[0176] In one embodiment, if based on the target analysis result, it is obtained that the preset fault state of the target hard disk is a disk drop state or a bad disk state, in order to quickly confirm whether the target hard disk is indeed in a fault state, the hard disk status data collection time interval of the target hard disk can also be shortened to increase the data collection frequency. For example, if under normal circumstances, the first time interval for collecting hard disk status data is 1 day, based on the target analysis result, it is confirmed that the preset fault state of the target hard disk is a disk drop state or a bad disk state, and the first time interval for collecting hard disk status data is adjusted to a second time interval of 10 minutes.
[0177] In one embodiment, correspondingly, after the target hard disk is repaired, the hard disk status data of the repaired target hard disk continues to be collected, and the collected hard disk status data is analyzed to obtain a status analysis result of the hard disk status data of the repaired target hard disk. Further, if based on the status analysis result, or based on the status analysis result and the historical status analysis result of the hard disk status data of the repaired target hard disk (the analysis method in the data collection here can be the same as the analysis method involved in the above step 203, which will not be repeated), when it is determined that the repaired target hard disk is currently in a normal working state, the data collection time interval is adjusted from the second time interval to the first time interval. Based on the above Figure 2-6 The method flow in this application embodiment provides a system architecture of a hard disk fault management system, such as Figure 7 As shown in the figure, the architecture of the hard disk fault management system is divided into four layers:
[0178] The first layer: fault discovery, which is divided into passive discovery (the cloud storage system's own fault monitoring system and hard disk detection tools) and manual entry - manual active discovery and entry of relevant data.
[0179] The second layer: application product access, configuration access is performed according to the characteristics of different businesses, including data statistical analysis, management policy configuration, hard disk registration program inspection and hard disk repair task scheduling, etc.
[0180] The third layer: fault handling, the automated processing flow of hard disk repair tasks for faulty hard disks.
[0181] The fourth layer: data operation, building multi-dimensional and multi-indicator fault processing data, and visualizing the health of business hard disks and machines. It should be noted that the various parts of the above hard disk fault management system are only one implementation method and do not limit the specific implementation of this solution. For example, the fault processing layer can also add a data collection time interval adjustment module.
[0182] The functions of each part are as follows:
[0183] The hard disk inspection tool of the hard disk fault management system is set in the cloud storage system (eg, injecting Agent as a hard disk inspection tool in the cloud storage system), which is used to collect hard disk status data from the operating system of the storage server and upload the hard disk status data to the hard disk fault management system.
[0184] Among them, hard disk status data includes whether the drive letter drifts, whether the slot information jumps, hard disk IO, glist (logical track) and plist (physical track) conditions, etc. After the hard disk status data of all hard disks in the storage server cluster in the cloud storage system are aggregated, the hard disk fault management system directly calls the corresponding interface to obtain the hard disk status data of each hard disk, and perceives the changes in the hard disk status in real time, and then establishes the automatic processing task of the faulty hard disk.
[0185] The hard disk registration program check of the application product connected to the hard disk fault management system records the binding relationship of the storage application bound to each hard disk in the cloud storage system (the application ID of the storage application corresponding to the hard disk ID), and also records the management strategy of the storage application and the hard disk repair plan corresponding to the management strategy in the management policy configuration. The hard disk status record is also recorded in the data statistical analysis, including the analysis results of the hard disk status data of the target hard disk each time. A task queue is set in the hard disk repair task scheduling. Each time a hard disk repair task is confirmed, the hard disk repair task is placed in the queue and waits for fault processing in turn.
[0186] The fault processing layer in the hard disk fault management system: obtains the hard disk repair task from the hard disk repair task scheduling queue, and performs the following processing flow for the obtained hard disk repair task:
[0187] For target hard disks that are in slow or bad state:
[0188] Step a, silent confirmation, such as silent confirmation: before triggering the disk replacement process for a hard disk repair task, an observation period is required. When the hard disk fault management system detects a hard disk repair task, it needs to observe for 2 hours. If the hard disk status has remained unchanged during these 2 hours, then proceed to the next step b. The 2-hour observation period is to prevent unnecessary fault processing caused by short-term jitter of the hard disk status. Among them, during this observation period, the data collection time interval can also be adjusted from the first time interval to the second time interval.
[0189] Step b, data migration, after silent confirmation, when it is determined that the hard disk is in a bad disk state or a slow disk state, first check whether the health status of the target hard disk is healthy. If healthy, the task returns failure. If the target data of the target hard disk can be read, the target data is migrated to the healthy hard disk (the first other hard disk). If the target data cannot be read, the data is repaired through EC coding, the target data is calculated inversely, and the target data is migrated to the first other hard disk.
[0190] Step c, on-site disk replacement. After data migration is completed, you can check again whether the target data still exists in the target hard disk. If it still exists, continue with data migration. After confirming that the target data has been migrated, an order will be sent to the on-site staff, and the staff will queue up for maintenance and exception handling on the on-site side to repair the target hard disk (in slow disk state or bad disk state).
[0191] Step d, hard disk acceptance. After the on-site maintenance is completed, the hard disk fault management system needs to conduct acceptance. The hard disk fault management system checks the health status of the repaired target hard disk (whether there is a bad disk or a dropped disk), whether the disk head can be read successfully, whether the on-site hard disk is not replaced or repaired and directly completes the task processing, and if the target hard disk needs to be replaced, confirm whether it is the target hard disk that was replaced before, etc. If the repaired target hard disk drops a disk, a new hard disk repair task is regenerated. Among them, when the repaired target hard disk is restored, the relevant services of the corresponding target hard disk need to be added back.
[0192] Step d, data balancing, data balancing means breaking up the data, writing part of the storage data of the second other hard disk to the repaired target hard disk (which can be the original target hard disk after repair, or the replaced hard disk, that is, the hard disk inserted into the slot of the original target hard disk) to achieve data balancing. Before data balancing, it will be checked whether the repaired target hard disk has the same fixed asset information as the disk on the current machine to prevent the same disk from being put into multiple locations after floating, that is, whether the basic hard disk information (including hard disk identification) of the repaired target hard disk is consistent in the cloud storage system and the hard disk fault management system.
[0193] For the target hard disk in the dropped state:
[0194] Step (1), silent confirmation, such as silent confirmation: before triggering the restart process for a dropped disk, an observation period is required. When the hard disk fault management system detects a dropped disk, it needs to be observed for 2 hours. If the hard disk status has remained unchanged during these 2 hours, then proceed to the next step (2). The 2-hour observation period is to prevent unnecessary fault processing caused by short-term jitters in the hard disk status. Among them, during this observation period, the data collection time interval can also be adjusted from the first time interval to the second time interval.
[0195] Step (2), issue a restart order. If the machine needs to be restarted to recover from a lost disk, a restart prompt is required for the business logic (target business) deployed on the machine before restarting the machine. The configuration is performed according to the customized logic of the target business, and the interface of the hard disk fault management system is called to complete the sending of the restart prompt information to the user.
[0196] Step (3), recovery confirmation, collecting hard disk status data based on the third time interval, confirming that the analysis result corresponding to the hard disk status data is that the hard disk is in a healthy state, and then ending the process. The third time interval is less than the first time interval, and can be greater than the second time interval or equal to the second time interval, or less than the second time interval. After confirming that the repaired hard disk is in a healthy state, the data collection time interval is adjusted to the first time interval.
[0197] Among them, in the above-mentioned processing flow of dropped disks, bad disks and slow disks, if the target management policy of the target storage application of the target hard disk indicates that the number of hard disks currently available is very small and the storage capacity is very small, which is insufficient to meet the storage disaster recovery conditions, the automatic processing task of the hard disk will not be started.
[0198] In addition, frequency control related: In order to ensure the safety of automated processing, a maximum of 100 hard disk repair tasks for bad and slow disks can be initiated in 1 day; only one disk replacement task is allowed to be initiated on the same machine in 1 day; when a machine has 3 or more disks lost, based on operation and maintenance experience, the probability of simultaneous disk loss is low, and an HBA card fault ticket is directly submitted to check the backplane and HBA card; the SLA time for a bad disk processing task is 7 days; an alarm is triggered when the frequency control is exceeded, and manual processing is required.
[0199] Abnormal failure: In the automated processing tasks, there is a corresponding failure type between each step. When an automated processing task fails abnormally, the abnormal failure can be pushed to manual processing. Then, manual processing only needs to deal with failed tasks every day, and failed tasks can be retried and terminated on the hard disk fault management system. This optimization reduces the hard disk failure tasks that need to be processed daily from 100+ to only 0-5 abnormal failure tasks, greatly reducing labor costs.
[0200] Based on the architecture of the hard disk fault management system and the processing flow of the two faulty hard disks, the embodiment of the present application provides a hard disk management method, such as Figure 8 As shown,
[0201] Step 801, confirming the hard disk repair task. Determine whether the target hard disk of the hard disk repair task is a lost disk, if so, execute steps 802-804, if not, execute steps 805-809.
[0202] Step 802: Confirmation of disk drop.
[0203] Step 803: Execute the hard disk repair task and restart the device where the target hard disk is located.
[0204] Step 804: After the restart is completed, the target hard disk is restored and confirmed.
[0205] Step 805: Slow disk or bad disk confirmation.
[0206] Step 806: Perform data migration.
[0207] Step 807: After the data migration is completed, replace the hard disk.
[0208] Step 808: Data balancing.
[0209] Step 809: After the hard disk is replaced, restore confirmation is performed.
[0210] based on Figure 2-6 and Figure 8 The method flow in this application embodiment provides a data statistical analysis method, which, after collecting the hard disk status data of the target hard disk, also includes:
[0211] Based on the obtained operation data of each hard disk in the cloud storage system, statistical analysis is performed to obtain the hard disk operation status of each hard disk in the cloud storage system, and the hard disk operation status includes: hard disk basic status and system operation status;
[0212] The basic information of the hard disk includes: the machine model, whether it is within the warranty period, the fault type, and at least one of the business it belongs to;
[0213] The system operation status includes: at least one of the following: the number of newly failed hard disks per month, the failure rate, the warranty expiration rate, and the processing time of each process step in the hard disk repair process.
[0214] In one embodiment, based on the above Figure 7 The data operation layer in the hard disk fault management system collects the hard disk operation data of each hard disk in each storage server of the cloud storage system (including the hard disk status data of each hard disk), and then statistically analyzes the hard disk operation status of the hard disks in the cloud storage system. The statistical analysis can be divided into the hard disk basic status dimension and the system operation status dimension to obtain multi-dimensional analysis data, which is convenient for the operation and maintenance of the cloud storage system.
[0215] In one embodiment, if Fig. 9 As shown, a cloud storage system overall situation page provided in an embodiment of the present application can display relevant data obtained by performing data collection and statistical analysis multiple times in a set time period.
[0216] The overall status page of the cloud storage system can be a page composed of multiple indicators and dimensions, including the total number of hard disks, hard disk warranty rate, hard disk monthly failure rate, current number of failed disks, unavailable hard disk capacity, and unprocessed failed disk capacity. The total number and distribution of hard disks can observe the changes in the scale of the entire product. The hard disk failure rate reflects the health quality of the hard disk, the warranty rate reflects the service status of the hard disk (overwarranty: hard disks that have exceeded the shelf life, the higher the warranty rate, the higher the failure rate), and the unavailable hard disk capacity and unprocessed failed hard disk capacity reflect whether there is a backlog of hard disk repair tasks.
[0217] The specific data displayed on the cloud storage system overall status page may include:
[0218] Dimensions: machine entry year, power-on time, model, manufacturer, fault type, whether it is out of warranty, cluster to which it belongs, etc.
[0219] Indicators: total number of hard disks / machines, warranty expiration rate, hard disk failure rate, current failure count, daily new hard disk failure rate, average processing time and processing time for each step, hard disk health score, overtime fault ticket statistics, etc. Among them, the health score can be obtained based on the average response time of the hard disk, the number of failures or frequency in the historical time period, and other related data.
[0220] Expiration rate: calculated based on the year the machine was put into storage. Machines that have been stored for more than three years are considered to be out of warranty.
[0221] Unavailable hard disk capacity: The total capacity of hard disks that have completed data migration and are undergoing on-site repair.
[0222] Unprocessed hard disk failure capacity: The total capacity of hard disks that have been found to have failed but have not been processed due to task queuing and other reasons.
[0223] Fault type: Hard disk fault types are divided into bad disk, slow disk and dropped disk.
[0224] Multiple dimensions can help segment fault types and identify concentrated problems over a period of time; multiple indicators can measure faults from multiple angles and more directly reflect the health of hard disks, cloud storage systems, and corresponding target storage systems.
[0225] In one embodiment, if Fig.10As shown. A bad disk distribution page in a cloud storage system provided by an embodiment of the present application can be displayed, which can display the current number of faulty hard disks obtained by statistical analysis, the number of newly added faulty hard disks per day, and the number of processed bad disks in a set historical time period. The bad disk situation can also be counted separately according to the year the machine was put into storage, the power-on time, the model, the manufacturer, the fault type, whether it is out of warranty, the cluster to which it belongs, etc. For example, the total number of faulty hard disks in the year the machine was put into storage in 2019, the total number of faulty hard disks of a set model, the total number of faulty hard disks of a certain manufacturer, the total number of faulty hard disks that are out of warranty, etc. The total number of faulty hard disks and the proportion of the number of faulty hard disks in each dimension to the total number of bad disks can be counted in the form of a pie chart. A line graph can also be drawn with time as the vertical axis and the current number of faulty hard disks or the number of newly added faulty hard disks per day as the vertical axis to display the bad disk situation.
[0226] The current number of faulty hard disks reflects the efficiency of automatic processing of faulty hard disks. If the current number of faulty hard disks exceeds the set value, it is necessary to explore which step is stuck and causes the accumulation. The number of newly added faulty hard disks per day reflects the health quality of the hard disks in the current cloud storage system. If the number of newly added bad disks suddenly increases several times a day, it is possible to determine through multiple dimensions whether they are concentrated in a certain manufacturer, model or cluster, and then locate the detailed cause, such as a batch problem of hard disks.
[0227] In one embodiment, if Fig.11 As shown, a task processing time efficiency status page in a cloud storage system provided by an embodiment of the present application may include the average processing time of a hard disk repair task and the average processing time of each collected step. A time trend graph may be drawn to provide feedback on the recent processing time efficiency of each step. If each step takes too long, targeted optimization may be performed. For example, the speed of data migration and data balancing may be optimized from a software level, and on-site personnel may be encouraged to improve maintenance time efficiency from a human level. A pie chart of the average processing time of a hard disk repair task may be drawn, and the proportion of the average processing time of each step may be plotted therein. In this way, the pie chart may provide feedback on the proportion of each step in the entire life cycle of a bad disk task. A line graph of the average processing time of a hard disk repair task may also be drawn with time as the horizontal axis and processing time efficiency as the vertical axis.
[0228] In one embodiment, a change and trend page for indicators related to faulty hard drives in a cloud storage system can be designed, including: statistics on hard drive failure rate, daily number of newly faulty hard drives and other indicators for each dimension (year of out-of-warranty hard drives, manufacturer, model, etc.). In this way, the root cause of the faulty hard drive can be analyzed from multiple dimensions (year of entry, manufacturer, abnormality type and model). For example, in 2022, it was found that the increased failure rate of a storage model was due to the firmware version; taking the task step as the dimension, the number of hard drive repair tasks is counted, and the current number of tasks and the number of timed tasks of each step are counted every day to observe the daily task accumulation, and help quickly locate which link the task accumulation occurs in. For example, the increase in fault handling time and the accumulation of fault tickets at the end of 2022 were due to the fact that most employees took sick leave during the flu period.
[0229] In one embodiment, if any step in the automatic processing of a faulty hard disk exceeds the processing time limit, an alarm message can be automatically generated and pushed to the corresponding mailbox or session, etc., to promptly notify the relevant processing object and perform system maintenance in a timely manner, wherein the alarm message can include the processing time of the step, various dimensional data of the faulty hard disk, and various indicator data, etc. For example, the unprocessed hard disks that have not been processed for more than 7 days are urged to be processed, and are pushed regularly every day, and the relevant processing objects need to follow up and claim them, and follow up in time with the detailed reasons for the timeout task.
[0230] In one embodiment, a page for viewing hardware single-machine fault status in a cloud storage system can be designed, which can display the details of the hard disk repair task that has not been completed for the faulty hard disk, and display the detailed information of the fault ticket of each hard disk repair task: the cluster to which it belongs, the fault type, the current processing step, the processing time, whether the machine is out of warranty, etc. The statistics of the dead nodes display the target business that currently needs to be shut down due to hardware failure (ajs and ping both execute normally) or is completely dead (ajs and ping both execute abnormally). For modules that have been dead for more than 7 days, the reason for the unfinished processing needs to be clarified. If it cannot be repaired for a long time, the relevant hardware will be replaced or removed from the shelves.
[0231] Based on the above Figure 2-6 and Figure 8 In the method flow, the embodiment of the present application provides a method for restoring confirmation after a faulty hard disk is repaired. In step 204, after the target hard disk is repaired using the hard disk repair solution, the method further includes:
[0232] Compare the hard disk identification of the repaired target hard disk with the hard disk identification of the target hard disk, and if the comparison result meets the comparison condition, the hard disk repair is successful;
[0233] Among them, contrast condition characterization;
[0234] When the preset fault state is a slow disk state or a bad disk state, a comparison result of the hard disk identification of the repaired target hard disk and the hard disk identification of the target hard disk is that the identifications are different;
[0235] When the fault state is a disk drop state, a comparison result of the hard disk identifier of the repaired target hard disk and the hard disk identifier of the target hard disk is that the identifiers are the same.
[0236] In one embodiment, generally, the repair solution for a bad or full hard disk is to replace the hard disk. After the target hard disk is repaired, it is confirmed whether the replaced hard disk has been replaced. Then, the hard disk identifiers of the replaced target hard disk (the repaired target hard disk) and the target hard disk before replacement (the repaired target hard disk) are inconsistent. If they are inconsistent, it proves that the hard disk replacement is successful, otherwise, the hard disk replacement fails.
[0237] Generally, the solution to repair a lost hard disk is to reinstall the device. After the target hard disk is repaired, confirm whether the restarted hard disk has completed the restart. Then, confirm whether the hard disk identifier of the target hard disk after the restart (the repaired target hard disk) is consistent with the target hard disk before the restart (the repaired target hard disk). If they are consistent, it proves that the device restarted successfully, otherwise, the restart failed.
[0238] The above-mentioned hard disk identification reconciliation (recovery confirmation method after the faulty hard disk is repaired) is aimed at frequent on-site misoperation, selecting the wrong hard disk when changing disks, causing data security risks, and establishing minute-level SN information (solid asset information, including hard disk identification) reconciliation to prevent human errors. Each storage server records the SN information of each hard disk. The hard disk fault management system deploys a hard disk detection tool Agent in each storage server to detect the hardware status. It initiates an SN information reconciliation every minute. It obtains the SN information of the hard disk in the storage server from the hard disk fault management system and compares it with the hard disk identification and other related information in the storage server. After reconciliation, if the SN is inconsistent, an alarm is immediately pushed, and the on-site staff is manually contacted to replace the original hard disk. Because the underlying storage of COS is multi-copy and EC encoding, a machine will only store one column of data for an object, so the on-site staff will not lose data if they change the wrong disk for a storage server. However, for COS hard disks at the million level, the frequent occurrence of such problems will pose a greater data security risk. Minute-level reconciliation and real-time alarms add another layer of protection for data security.
[0239] Using the above hard disk identification reconciliation method, the COS team optimized and transformed the entire process of automated processing of server fault management, and gradually connected it in 2021. Fig.12As shown, after adopting a hard disk identification reconciliation method provided in an embodiment of the present application, the changes in the amount of work orders (fault orders corresponding to hard disk repair tasks) and manpower input, as well as the toad skin intervention coverage rate and the changes in work order processing time, when the COS stock size continued to increase and the operating failure rate was 6.5%, the total number of work orders was reduced by half in two years, the average work order processing time was reduced to 1 / 3 of the original, and the maintenance manpower was greatly reduced. Currently, less than 10% of the manpower input is required to manage hundreds of thousands of servers and millions of hard disks, and zero manpower has been achieved for processing faulty hard disks.
[0240] At the same time, we optimize decision-making through data measurement, build a dashboard for faulty hard drives and hardware failure data, and visualize the health of the entire product hardware from multiple dimensions (such as year of entry, manufacturer, model, fault type, etc.) and multiple indicators (such as failure rate, health score, processing time, warranty rate, number of new failures per month, etc.). We continuously reduce the failure rate and increase the coverage rate of automated order creation. The current coverage rate has reached 95%. Subsequently, we will conduct regular reviews of the 5% of hardware failures that cannot be proactively discovered to improve our ability to discover faults.
[0241] Based on the same concept, the present application embodiment provides a hard disk management device 1300, such as Fig.13 As shown, it is applied to a hard disk fault management system, which is used to manage the hard disk operation status of each storage server in the cloud storage system, and records the binding relationship between each hard disk and the storage application that each is allowed to access, as well as the management strategy of each storage application for the hard disk. The device includes:
[0242] The data collection unit 1301 is used to collect hard disk status data of a target hard disk for a storage server; the storage server belongs to a server cluster for providing cloud storage services, and the hard disk status data represents: usage of the target hard disk when reading and writing relevant data generated by the cloud storage service;
[0243] The data analysis unit 1302 is used to analyze the hard disk status data to obtain a target analysis result, wherein the target analysis result indicates whether the target hard disk is in a normal working state;
[0244] The fault confirmation unit 1303 is used to obtain the target management policy of the target storage application bound to the target hard disk by using the hard disk identifier of the target hard disk when it is determined that the target hard disk is currently in a preset fault state based on the target analysis result;
[0245] The hard disk repair unit 1304 is used to obtain a hard disk repair solution corresponding to the preset fault state based on the target management policy, and use the hard disk repair solution to repair the target hard disk.
[0246] Optionally, the fault confirmation unit 1303 is specifically configured to:
[0247] Based on the hard disk identifier, a preset number of historical analysis results are obtained from the hard disk status record;
[0248] When the target analysis result and the obtained historical analysis results meet the preset read-write abnormality condition, it is determined that the target hard disk is currently in a bad disk state; the read-write abnormality condition represents: the degree of read-write abnormality of the target hard disk when performing read-write data processing;
[0249] When the target analysis result and the historical analysis results meet the preset recognition abnormality condition, it is determined that the target hard disk is currently in a disk drop state; the recognition abnormality condition indicates that: when the target hard disk is recognized by the device where it is located, a recognition abnormality occurs;
[0250] When the target analysis result and the historical analysis results meet the preset required bandwidth, and the average response time of the target hard disk exceeds the preset duration, it is determined that the target hard disk is currently in a slow disk state.
[0251] Optionally, the preset fault state is a slow disk state or a bad disk state;
[0252] The hard disk repair unit 1304 is specifically used for:
[0253] Reading target data in the target hard disk, migrating the target data to a first other hard disk, and repairing the target hard disk using a hard disk repair solution corresponding to the dropped disk state or the bad disk state;
[0254] The hard disk repair unit 1304 is also used for:
[0255] Acquire storage data from a second other hard disk, and write the storage data into the repaired target hard disk, wherein the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
[0256] Optionally, the preset fault state is a slow disk state or a bad disk state;
[0257] The hard disk repair unit 1304 is specifically used for:
[0258] If the target data in the target hard disk cannot be read, then determine the associated data of the target data, and obtain the associated data from the associated hard disk where the associated data is located, the target data and the associated data belong to the same storage file, and are encoded in a preset encoding method and stored in the target hard disk and the associated hard disk respectively;
[0259] Decoding and restoring the associated data based on the preset encoding method to obtain the storage file, migrating the target data in the storage file to a first other hard disk, and using a hard disk repair solution corresponding to the dropped disk state or the bad disk state to repair the target hard disk;
[0260] The hard disk repair unit 1304 is also used for:
[0261] Acquire storage data from a second other hard disk, and write the storage data into the repaired target hard disk, wherein the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
[0262] Optionally, the preset fault state is a disk drop state;
[0263] The hard disk repair unit 1304 is also used for:
[0264] Obtaining a target business in a running state on the device where the target hard disk is located, and based on a custom logic in the target business, initiating a restart prompt message to the target object, wherein the target business is used to provide relevant services to the target object;
[0265] The hard disk repair unit 1304 is specifically used for:
[0266] Restart the device where the target hard disk is located.
[0267] Optionally, the preset fault state is a disk drop state or a bad disk state;
[0268] The hard disk repair unit 1304 is also used for:
[0269] Adjusting a data collection time interval for collecting the hard disk status data from a first time interval to a second time interval, wherein the second time interval is shorter than the first time interval;
[0270] After the target hard disk is repaired by the hard disk repair solution, the method further includes:
[0271] Collect and analyze the hard disk status data of the repaired target hard disk to obtain the corresponding status analysis result;
[0272] When it is determined based on the status analysis result that the repaired target hard disk is currently in a normal working state, the data collection time interval is adjusted from the second time interval to the first time interval.
[0273] Optionally, the data parsing unit 1302 is further configured to:
[0274] Based on the obtained operation data of each hard disk in the cloud storage system, statistical analysis is performed to obtain the hard disk operation status of each hard disk in the cloud storage system, wherein the hard disk operation status includes: hard disk basic status and system operation status;
[0275] The basic information of the hard disk includes: the machine model, whether it is within the warranty period, the fault type, and at least one of the business it belongs to;
[0276] The system operation status includes: at least one of the following: the number of newly failed hard disks per month, the failure rate, the warranty expiration rate, and the processing time of each process step in the hard disk repair process.
[0277] Optionally, the hard disk repair unit is further used to:
[0278] Comparing the hard disk identification of the repaired target hard disk with the hard disk identification of the target hard disk, and determining that the comparison result meets the comparison condition, the hard disk repair is successful;
[0279] Wherein, the contrast condition characterizes;
[0280] When the preset fault state is a slow disk state or a bad disk state, a comparison result of the hard disk identification of the repaired target hard disk and the hard disk identification of the target hard disk is that the identifications are different;
[0281] When the fault state is a disk drop state, a comparison result of the hard disk identification of the repaired target hard disk and the hard disk identification of the target hard disk is that the identifications are the same.
[0282] The hard disk management method and the above method embodiment are based on the same inventive concept. In an embodiment of the present application, a computer device is also provided. In one embodiment, the computer device may be a server, such as Figure 1 The server 120 and the server 130 shown in FIG. 1 are shown in FIG. 1 . In this embodiment, the structure of the computer device can be as follows: Fig.14 As shown, it includes a memory 1401, a communication module 1403 and one or more processors 1402.
[0283] The memory 1401 is used to store computer programs executed by the processor 1402. The memory 1401 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and programs required for running the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0284] The memory 1401 may be a volatile memory, such as a random-access memory (RAM); the memory 1401 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 1401 may be any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1401 may be a combination of the above memories.
[0285] The processor 1402 may include one or more central processing units (CPU) or a digital processing unit, etc. The processor 1402 is used to implement the above hard disk management method when calling the computer program stored in the memory 1401 .
[0286] The communication module 1403 is used to communicate with terminal devices and other servers.
[0287] The specific connection medium between the memory 1401, the communication module 1403 and the processor 1402 is not limited in the embodiment of the present application. Fig.14 In the embodiment, the memory 1401 and the processor 1402 are connected via a bus 1404. Fig.14 The connections between the other components are only for illustration and are not intended to be limiting. The bus 1404 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Fig.14 The diagram shows that only one thick line is used, but this does not mean that there is only one bus or only one type of bus.
[0288] The memory 1401 stores a computer storage medium, and the computer storage medium stores computer executable instructions, and the computer executable instructions are used to implement the hard disk management method of the embodiment of the present application. The processor 1402 is used to execute the above-mentioned hard disk management method, such as Figure 2-6 and Figure 8 shown.
[0289] In another embodiment, the computer device may also be other computer devices, such as Figure 1 The terminal device 110 shown in FIG. 1 is a terminal device 110 shown in FIG. 1 . In this embodiment, the structure of the computer device can be as follows: Fig.15As shown, it includes: a communication component 1510, a memory 1520, a display unit 1530, a camera 1540, a sensor 1550, an audio circuit 1560, a Bluetooth module 1570, a processor 1580 and other components.
[0290] The communication component 1510 is used to communicate with the server. In some embodiments, a wireless fidelity (WiFi) module may be included. The WiFi module belongs to a short-range wireless transmission technology. The computer device can help the user to send and receive information through the WiFi module.
[0291] The memory 1520 can be used to store software programs and data. The processor 1580 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1520. The memory 1520 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. The memory 1520 stores an operating system that enables the terminal device 110 to run. In the present application, the memory 1520 can store an operating system and various application programs, and can also store a computer program for executing the hard disk management method in the embodiment of the present application.
[0292] The display unit 1530 can also be used to display information input by the user or information provided to the user and a graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1530 may include a display screen 1532 disposed on the front of the terminal device 110. The display screen 1532 may be configured in the form of a liquid crystal display, a light emitting diode, etc. The display unit 1530 can be used to display the user interface of the hard disk management method in the embodiment of the present application, etc.
[0293] The display unit 1530 can also be used to receive input digital or character information and generate signal input related to user settings and function control of the terminal device 110. Specifically, the display unit 1530 may include a touch screen 1531 arranged on the front of the terminal device 110, which can collect user touch operations on or near it, such as clicking a button, dragging a scroll box, etc.
[0294] The touch screen 1531 can be covered on the display screen 1532, or the touch screen 1531 and the display screen 1532 can be integrated to realize the input and output functions of the terminal device 110, and the integrated display screen can be referred to as a touch display screen. In this application, the display unit 1530 can display the application program and the corresponding operation steps.
[0295] The camera 1540 can be used to capture still images, and the user can post comments on the images taken by the camera 1540 through the application. The camera 1540 can be one or more. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the processor 1580 to convert it into a digital image signal.
[0296] The terminal device may further include at least one sensor 1550, such as an acceleration sensor 1551, a distance sensor 1552, a fingerprint sensor 1553, and a temperature sensor 1554. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0297] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the terminal device 110. The audio circuit 1560 can transmit the electrical signal converted from the received audio data to the speaker 1561, which is converted into a sound signal for output. The terminal device 110 can also be configured with a volume button for adjusting the volume of the sound signal. On the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560 and converted into audio data, and then the audio data is output to the communication component 1510 to be sent to, for example, another terminal device 110, or the audio data is output to the memory 1520 for further processing.
[0298] The Bluetooth module 1570 is used to exchange information with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable computer device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1570 to exchange data.
[0299] The processor 1580 is the control center of the terminal device. It uses various interfaces and lines to connect various parts of the entire terminal. It executes various functions of the terminal device and processes data by running or executing software programs stored in the memory 1520 and calling data stored in the memory 1520. In some embodiments, the processor 1580 may include one or more processing units; the processor 1580 may also integrate an application processor and a baseband processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the baseband processor mainly processes wireless communications. It is understandable that the above-mentioned baseband processor may not be integrated into the processor 1580. In the present application, the processor 1580 can run the operating system, application programs, user interface display and touch response, and the hard disk management method of the embodiment of the present application. In addition, the processor 1580 is coupled to the display unit 1530.
[0300] In some possible implementations, various aspects of the hard disk management method provided in the present application may also be implemented in the form of a program product, which includes a computer program. When the program product is run on a computer device, the computer program is used to enable the computer device to execute the steps of the hard disk management method according to various exemplary embodiments of the present application described above in this specification. For example, the computer device may execute the following steps: Figure 2-6 In any method flow or Figure 8 Steps shown in the method flow.
[0301] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0302] The program product of the embodiment of the present application may adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and can be run on a computer device. However, the program product of the present application is not limited thereto, and in this document, a readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with a command execution system, apparatus, or device.
[0303] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable computer program is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.
[0304] The computer program embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0305] The computer program for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The computer program may be executed entirely on the user's computer device, partially on the user's device, as a separate software package, partially on the user's computer device and partially on a remote computer device, or entirely on a remote computer device or server. In the case of a remote computer device, the remote computer device may be connected to the user's computer device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer device (e.g., via the Internet using an Internet service provider).
[0306] It should be noted that, although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided into multiple units to be embodied.
[0307] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0308] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0309] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0310] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0311] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0312] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A hard disk management method, characterized in that: Applied to a hard disk fault management system, the hard disk fault management system is used to manage the hard disk operation status of each storage server in the cloud storage system, and records the binding relationship between each hard disk and the storage application that is allowed to access, as well as the management strategy of each storage application for the hard disk. The method includes: For a storage server, hard disk status data of a target hard disk is collected; the storage server belongs to a server cluster for providing cloud storage services, and the hard disk status data represents: the usage of the target hard disk when reading and writing relevant data generated by the cloud storage service; Parsing the hard disk status data to obtain a target parsing result, wherein the target parsing result indicates whether the target hard disk is in a normal working state; When it is determined based on the target analysis result that the target hard disk is currently in a preset fault state, the hard disk identifier of the target hard disk is used to obtain a target management policy of a target storage application bound to the target hard disk; Based on the target management strategy, a hard disk repair solution corresponding to the preset fault state is obtained, and the target hard disk is repaired using the hard disk repair solution.
2. The method according to claim 1, characterized in that The determining, based on the target analysis result, that the target hard disk is currently in a preset fault state includes: Based on the hard disk identifier, a preset number of historical analysis results are obtained from the hard disk status record; When the target analysis result and the obtained historical analysis results meet the preset read-write abnormality condition, it is determined that the target hard disk is currently in a bad disk state; the read-write abnormality condition represents: the degree of read-write abnormality of the target hard disk when performing read-write data processing; When the target analysis result and the historical analysis results meet the preset recognition abnormality condition, it is determined that the target hard disk is currently in a disk drop state; the recognition abnormality condition indicates that: when the target hard disk is recognized by the device where it is located, a recognition abnormality occurs; When the target analysis result and the historical analysis results meet the preset required bandwidth, and the average response time of the target hard disk exceeds the preset duration, it is determined that the target hard disk is currently in a slow disk state.
3. The method according to claim 1, characterized in that The preset fault state is a slow disk state or a bad disk state; The step of repairing the target hard disk by using the hard disk repair solution includes: Reading target data in the target hard disk, migrating the target data to a first other hard disk, and repairing the target hard disk using a hard disk repair solution corresponding to the dropped disk state or the bad disk state; After the target hard disk is repaired by the hard disk repair solution, the method further includes: Acquire storage data from a second other hard disk, and write the storage data into the repaired target hard disk, wherein the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
4. The method according to claim 1, characterized in that The preset fault state is a slow disk state or a bad disk state; The step of repairing the target hard disk by using the hard disk repair solution includes: If the target data in the target hard disk cannot be read, then determine the associated data of the target data, and obtain the associated data from the associated hard disk where the associated data is located, the target data and the associated data belong to the same storage file, and are encoded in a preset encoding method and stored in the target hard disk and the associated hard disk respectively; Decoding and restoring the associated data based on the preset encoding method to obtain the storage file, migrating the target data in the storage file to a first other hard disk, and using a hard disk repair solution corresponding to the dropped disk state or the bad disk state to repair the target hard disk; After the target hard disk is repaired by the hard disk repair solution, the method further includes: Acquire storage data from a second other hard disk, and write the storage data into the repaired target hard disk, wherein the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
5. The method according to claim 1, characterized in that The preset fault state is a disk drop state; Before the hard disk repair solution is used to repair the target hard disk, the method includes: Obtaining a target business in a running state on the device where the target hard disk is located, and based on a custom logic in the target business, initiating a restart prompt message to the target object, wherein the target business is used to provide relevant services to the target object; Then the method of using the hard disk repair solution to repair the target hard disk includes: Restart the device where the target hard disk is located.
6. The method according to claim 1, characterized in that The preset fault state is a disk drop state or a broken disk state; Before the hard disk repairing solution is used to repair the target hard disk, the method further includes: Adjusting a data collection time interval for collecting the hard disk status data from a first time interval to a second time interval, wherein the second time interval is shorter than the first time interval; After the target hard disk is repaired by the hard disk repair solution, the method further includes: Collect and analyze the hard disk status data of the repaired target hard disk to obtain the corresponding status analysis result; When it is determined based on the status analysis result that the repaired target hard disk is currently in a normal working state, the data collection time interval is adjusted from the second time interval to the first time interval.
7. The method according to any one of claims 1 to 6, characterized in that: After collecting the hard disk status data of the target hard disk, the method further includes: Based on the obtained operation data of each hard disk in the cloud storage system, statistical analysis is performed to obtain the hard disk operation status of each hard disk in the cloud storage system, wherein the hard disk operation status includes: hard disk basic status and system operation status; The basic information of the hard disk includes: the machine model, whether it is within the warranty period, the fault type, and at least one of the business it belongs to; The system operation status includes: at least one of the following: the number of newly failed hard disks per month, the failure rate, the warranty expiration rate, and the processing time of each process step in the hard disk repair process.
8. The method according to any one of claims 1 to 6, characterized in that: After the target hard disk is repaired by the hard disk repair solution, the method further includes: Comparing the hard disk identification of the repaired target hard disk with the hard disk identification of the target hard disk, and determining that the comparison result meets the comparison condition, the hard disk repair is successful; Wherein, the contrast condition characterizes; When the preset fault state is a slow disk state or a bad disk state, a comparison result of the hard disk identification of the repaired target hard disk and the hard disk identification of the target hard disk is that the identifications are different; When the fault state is a disk drop state, a comparison result of the hard disk identification of the repaired target hard disk and the hard disk identification of the target hard disk is that the identifications are the same.
9. A hard disk management device, characterized in that: Applied to a hard disk fault management system, the hard disk fault management system is used to manage the hard disk operation status of each storage server in the cloud storage system, and records the binding relationship between each hard disk and the storage application that each hard disk is allowed to access, as well as the management strategy of each storage application for the hard disk. The device includes: A data collection unit is used to collect hard disk status data of a target hard disk for a storage server; the storage server belongs to a server cluster for providing cloud storage services, and the hard disk status data represents: usage of the target hard disk when reading and writing relevant data generated by the cloud storage service; A data analysis unit, used to analyze the hard disk status data to obtain a target analysis result, wherein the target analysis result indicates whether the target hard disk is in a normal working state; A fault confirmation unit, configured to, when it is determined based on the target analysis result that the target hard disk is currently in a preset fault state, use the hard disk identifier of the target hard disk to obtain a target management policy of a target storage application bound to the target hard disk; The hard disk repair unit is used to obtain a hard disk repair plan corresponding to the preset fault state based on the target management strategy, and use the hard disk repair plan to repair the target hard disk.
10. The device according to claim 9, characterized in that The fault confirmation unit is specifically used for: Based on the hard disk identifier, a preset number of historical analysis results are obtained from the hard disk status record; When the target analysis result and the obtained historical analysis results meet the preset read / write abnormality condition, it is determined that the target hard disk is currently in a bad disk state; The read / write abnormality condition represents: the degree of read / write abnormality of the target hard disk when performing read / write data processing; When the target analysis result and the historical analysis results meet the preset recognition abnormality condition, it is determined that the target hard disk is currently in a disk drop state; the recognition abnormality condition indicates that: when the target hard disk is recognized by the device where it is located, a recognition abnormality occurs; When the target analysis result and the historical analysis results meet the preset required bandwidth, and the average response time of the target hard disk exceeds a preset duration, it is determined that the target hard disk is currently in a slow disk state.
11. The device according to claim 9, characterized in that The preset fault state is a slow disk state or a bad disk state; The hard disk repair unit is specifically used for: Reading target data in the target hard disk, migrating the target data to a first other hard disk, and repairing the target hard disk using a hard disk repair solution corresponding to the dropped disk state or the bad disk state; After the target hard disk is repaired by the hard disk repair solution, the method further includes: Acquire storage data from a second other hard disk, and write the storage data into the repaired target hard disk, wherein the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
12. The device according to claim 9, characterized in that The preset fault state is a slow disk state or a bad disk state; The hard disk repair unit is specifically used for: If the target data in the target hard disk cannot be read, then determine the associated data of the target data, and obtain the associated data from the associated hard disk where the associated data is located, the target data and the associated data belong to the same storage file, and are encoded in a preset encoding method and stored in the target hard disk and the associated hard disk respectively; Decoding and restoring the associated data based on the preset encoding method to obtain the storage file, migrating the target data in the storage file to a first other hard disk, and using a hard disk repair solution corresponding to the dropped disk state or the bad disk state to repair the target hard disk; The hard disk repair unit is also used for: Acquire storage data from a second other hard disk, and write the storage data into the repaired target hard disk, wherein the first other hard disk and the second other hard disk are the same hard disk or different hard disks.
13. A computer-readable non-volatile storage medium, characterized in that: The computer-readable non-volatile storage medium stores a program, and when the program is run on a computer, the computer is enabled to implement the method according to any one of claims 1 to 8.
14. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to call a computer program stored in the memory, and execute the method according to any one of claims 1 to 8 according to the obtained program.
15. A computer program product, characterized in that The method comprises a computer program, wherein the computer program is stored in a computer-readable storage medium; when a processor of a computer device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computer device executes the method according to any one of claims 1 to 8.