Data processing method, device, server and storage medium for solid state hard drive failure

By iteratively monitoring and automatically collecting data from multiple solid-state drives, the problem of low efficiency in diagnosing solid-state drive failures in the existing technology is solved, and efficient and accurate fault diagnosis is achieved.

CN120469845BActive Publication Date: 2025-09-09INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510955103.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-09
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

The existing technology is inefficient in diagnosing solid-state drive failures, especially those failure scenarios that cause the operating system to report system-level errors but can still be identified by the host. The cause is difficult to locate, resulting in low diagnostic efficiency.

Method used

By iteratively monitoring multiple solid-state drives, enabling multiple processes to collect operating data and save it to different types of files, regularly checking whether system-level errors occur, obtaining the latest file if no errors occur, and automatically obtaining data before and after the failure until a fault is detected. System-level error information and device information are saved.

Benefits of technology

It improves the efficiency and accuracy of data acquisition, reduces manual intervention, and improves the efficiency of diagnosing solid-state drive failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469845B_ABST
    Figure CN120469845B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, device, server and storage medium for solid-state hard drive failures, which relates to the technical field of solid-state hard drives. Through iterative monitoring, multiple processes are enabled to collect the operating data of multiple solid-state hard drives and save them in different types of files. In each iterative cycle, each solid-state hard drive is regularly detected to see if a system-level error occurs. If not, the latest file of each type of file is obtained until a system-level error is detected in any solid-state hard drive, and the iteration is stopped. After a system-level error occurs, the operating data of multiple solid-state hard drives continues to be collected within a preset time and saved to the latest file of each type; the system-level error information and device information of the solid-state hard drive are saved to a log save file. Automatically obtaining data before and after the system-level error improves the efficiency of data acquisition, thereby improving the efficiency of diagnosing solid-state hard drive failures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of solid-state hard disks, and in particular to a method, device, server, and storage medium for processing data related to solid-state hard disk failures. Background Art

[0002] In the information society, data collection and application are becoming increasingly important. High-performance, large-capacity, and compact solid-state drives (SSDs) are becoming increasingly popular. As demand for SSDs increases, so too does the number of SSD failures. One failure scenario involves an SSD causing the operating system to report a system-level error, yet the SSD remains recognized by the host. This failure scenario is less likely to occur, but the cause can be difficult to pinpoint.

[0003] Currently, related technologies reproduce the failure scenario, manually obtain data from the failure scenario during the reproduction process, and then analyze it. If the cause of the failure is located, the reproduction and acquisition process stops; if the cause of the failure is not located, the reproduction is repeated until the cause of the failure is located. However, this method is inefficient in obtaining data from the failure scenario, which reduces the efficiency of diagnosing SSD failures. Summary of the Invention

[0004] The present application provides a method, device, server, and storage medium for processing data on solid-state drive failures, to at least solve the problem of low efficiency in diagnosing solid-state drive failures in related technologies.

[0005] This application provides a method for processing data from a solid-state drive failure, including:

[0006] Monitor multiple SSDs and iteratively perform the following steps:

[0007] In the current iteration cycle, multiple processes are enabled to collect operating data from multiple solid-state drives, and the operating data is saved in multiple types of files; the operating data collected by different processes are saved in different types of files;

[0008] During a preset timing period, each solid-state drive is regularly checked for system-level errors;

[0009] If no system-level error occurs on each SSD, the latest file of each file type is obtained;

[0010] Iteration stops until a system-level error is detected on any solid-state drive;

[0011] Save system-level error information and device information of any solid-state drive to a log save file;

[0012] Within the preset time, the operation data of multiple solid-state drives continues to be collected and saved to the latest file of each type of file; the latest file of each type of file and the log save file are used to diagnose the fault of any solid-state drive.

[0013] The present application also provides a data processing device for a solid-state hard drive failure, comprising:

[0014] The monitoring module is used to monitor multiple solid-state drives and iteratively execute the following steps:

[0015] The monitoring module includes:

[0016] The collection submodule is used to enable multiple processes to collect the operating data of multiple solid-state drives in the current iteration cycle, and save the operating data into multiple types of files; the operating data collected by different processes are saved into different types of files;

[0017] The detection submodule is used to regularly detect whether each solid-state drive has a system-level error within a preset timing period;

[0018] A determination submodule, configured to obtain the latest file of each type if no system-level error occurs in each solid-state drive;

[0019] The stop iteration submodule is used to stop iteration until a system-level error is detected in any solid-state drive;

[0020] A first saving module is used to save the system-level error information and device information of any solid-state drive to a log saving file;

[0021] The second saving module is used to continue collecting the operating data of multiple solid-state drives within a preset time and save the operating data to the latest file of each type of file; the latest file of each type of file and the log save file are used to diagnose faults of any solid-state drive.

[0022] The present application also provides a server, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned methods for processing data of a solid-state hard drive failure when executing the computer program.

[0023] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for processing data of a solid-state hard drive failure are implemented.

[0024] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned methods for processing data of solid-state hard drive failures when the computer program is executed by a processor.

[0025] Through this application, multiple solid-state drives are iteratively monitored, and multiple processes are enabled to collect operational data from multiple SSDs and save it to different types of files. During each iteration, each SSD is periodically checked for system-level errors. If not, the latest file of each file type is obtained until a system-level error is detected on any SSD, indicating that the SSD has failed, and iteration stops. The latest file of each file type contains the data before the SSD failed. After a system-level error occurs, operational data from multiple SSDs is continuously collected within a preset time and saved to the latest file of each type. The system-level error information and device information of the SSD are saved to a log file. At this point, the data after the failure includes the data in the latest file of each file type and the data in the log file. Without the need to reproduce the failure scenario, during the monitoring process of multiple SSDs, each SSD is periodically checked for system-level errors, and data before and after the failure is automatically obtained, improving the efficiency and accuracy of data acquisition, thereby improving the efficiency of diagnosing SSD failures. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0027] Figure 1 A schematic diagram of a scenario of a data processing method for a solid-state drive failure provided in an embodiment of the present application;

[0028] Figure 2 A flowchart of a method for processing data related to a solid-state drive failure provided in an embodiment of the present application;

[0029] Figure 3 A schematic diagram of the structure of a data processing device for a solid-state hard drive failure provided in an embodiment of the present application;

[0030] Figure 4 A schematic diagram of the structure of the server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0032] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0033] In the information society, data collection and application are becoming increasingly important, and high-performance, large-capacity, and compact solid-state drives are widely used. As the demand for solid-state drives increases, the number of faulty solid-state drives also increases. One type of fault scenario is that the solid-state drive causes the operating system to report a system-level error, but the solid-state drive can still be recognized by the host. The probability of this fault scenario is low, but the cause of the fault is difficult to locate. Currently, related technologies reproduce the fault scenario, and during the reproduction process, the data of the fault scenario is manually obtained and then analyzed. If the cause of the fault is located, the reproduction and acquisition actions are stopped; if the cause of the fault is not located, the reproduction is repeated until the cause of the fault is located. However, this method is inefficient in obtaining data for fault scenarios, resulting in a reduced efficiency in diagnosing solid-state drive faults.

[0034] To address the technical problems in the related art, the present application proposes the following technical concept: Considering that the related art requires manual data acquisition through reproducing the fault scenario, the inventors devised a method to eliminate the need for reproducing the fault scenario. Instead, during the monitoring process of multiple SSDs, the inventors periodically detect whether each SSD has experienced a system-level error, automatically acquiring data before and after the fault. Specifically, the inventors iteratively monitor multiple SSDs, collect operational data from each SSD, and save the data to different types of files. During each iteration, the inventors periodically detect whether each SSD has experienced a system-level error. If not, the inventors retrieve the latest file of each file type until a system-level error is detected on any SSD, indicating a fault on that SSD, and the iteration process ceases. At this point, the latest file of each file type represents the data before the SSD fault. After the iteration process ceases, the inventors continue to collect operational data from the multiple SSDs within a preset time period and save it to the latest file of each file type. The inventors also save the system-level error information and device information of each SSD to a log file. In this case, the data after the fault occurs includes the data in the latest file of each file type and the data in the log file. The system can automatically obtain data before and after a fault occurs, thereby improving the efficiency of data acquisition and thus improving the efficiency of diagnosing solid-state hard drive faults.

[0035] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0036] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the data processing method for a solid state drive failure depends, the specific application environment architecture or specific hardware architecture is described herein.

[0037] refer to Figure 1 , Figure 1 A schematic diagram of a scenario of a data processing method for a solid state drive failure provided in an embodiment of the present application, such as Figure 1 As shown, it includes: a server 101 and multiple solid-state drives 102.

[0038] It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the method for processing data in the event of a SSD failure. In other feasible implementations of this application, the above architecture may include more or fewer components than shown, or may combine or split certain components, or arrange the components differently. The specific configuration can be determined based on the actual application scenario and is not limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0039] In the specific implementation process, the server 101 monitors multiple solid-state drives 102. During the monitoring process, the server 101 enables multiple threads to collect the operating data of multiple solid-state drives 102 and save the operating data to different types of files. Regularly detect whether each solid-state drive 102 has a system-level error. If no system-level error occurs, the latest file of each type of file is obtained; if any solid-state drive 102 fails, the operating data of multiple solid-state drives 102 is continuously collected within a preset time, saved to the latest file of each type of file, and the system-level error information and device information of the solid-state drive 102 are saved to the log save file. Finally, the latest file of each type of file and the log save file are obtained; the server 101 uses the latest file of each type of file and the log save file to diagnose the fault of the solid-state drive 102 with a system-level error.

[0040] In addition, the network architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0041] Figure 2 A flow chart of a method for processing data of a solid state drive failure provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, an embodiment of the present application provides a method for processing data of a solid state drive failure, and the method is described in detail as follows:

[0042] S201: Monitor multiple solid-state drives and iteratively perform the following steps:

[0043] In this embodiment, the solid-state drive is a solid-state drive based on the storage interface standard (Non-Volatile Memory Express, NVMe) protocol.

[0044] Before executing step S201, an initialization operation needs to be performed first, wherein the initialization operation steps include S2001 to S2004:

[0045] S2001: Create a log save file.

[0046] S2002: Obtain drive letter names of multiple solid-state hard drives.

[0047] In this embodiment, the drive letter names of all solid-state hard disks in the operating system are obtained.

[0048] S2003: Create an empty solid-state drive list.

[0049] S2004: Save the drive letter names of the multiple solid-state hard drives to an empty solid-state hard drive list to obtain a pre-built solid-state hard drive list.

[0050] Optionally, the initialization operation further includes: initializing variables under the operating system, such as a timer variable; saving the dmesg log under the operating system to a log storage file; and clearing the dmesg log under the operating system so that when a new dmesg log is subsequently generated, only the new dmesg log is analyzed. The initialization iteration cycle begins with the first iteration cycle.

[0051] dmesg is a command used in Linux systems to display startup information, hardware detection results, driver loading status, and kernel runtime error or warning messages. The command output is saved in the dmesg log.

[0052] Specifically, step S201 includes S2011 to S2014:

[0053] S2011: In the current iteration cycle, multiple processes are enabled to collect operating data of multiple solid-state drives, and the operating data is saved in multiple types of files; wherein the operating data collected by different processes are saved in different types of files.

[0054] In this embodiment, four processes are enabled in parallel: the iostat process, the sar process, the interrupt statistics collection process, and the blktrace process. iostat is a tool used in Linux to monitor CPU usage and disk I / O performance and is part of the sysstat toolkit. sar is a performance monitoring tool in Linux and a core component of the sysstat toolkit. blktrace is a tool used in Linux to trace the entire lifecycle of block device I / O operations and is a core component of the block device debugging toolchain.

[0055] In this embodiment, the iostat process collects data such as the number of input / output operations per second (IOPS) and latency, as well as CPU utilization, from multiple SSDs and saves it to a first type of file. The sar process collects multi-dimensional performance data from the operating system layer, including CPU, memory, SSD IOPS, network, and process data, and saves it to a second type of file. The interrupt statistics collection process collects interrupt event statistics from multiple SSDs at the operating system layer and saves it to a third type of file. The blktrace process obtains the drive letter names of multiple SSDs, collects data on the entire lifecycle of IO operations for each SSD based on its drive letter name, and saves it to a fourth type of file.

[0056] In this embodiment, the file name of each type of file is determined according to the file type and the current iteration cycle number.

[0057] For example, different numbers represent different types of files, such as 1 for the first type, 2 for the second type, 3 for the third type, and 4 for the fourth type. If the current iteration cycle is the first iteration cycle, the file name of the first type file is 1-1, where the first 1 represents the file type and the second 1 represents the current iteration cycle number; the file name of the second type file is 2-1, where 2 represents the file type and 1 represents the current iteration cycle number; the file name of the third type file is 3-1, where 3 represents the file type and 1 represents the current iteration cycle number; and the file name of the fourth type file is 4-1, where 4 represents the file type and 1 represents the current iteration cycle number.

[0058] In this embodiment, the file names of the first type of files are 1-N, the file names of the second type of files are 2-N, the file names of the third type of files are 3-N, and the file names of the fourth type of files are 4-N, where N represents the number of iteration cycles.

[0059] Optionally, the collected operating data of multiple solid-state hard drives may be saved in different types of folders.

[0060] S2012: Within a preset timing period, periodically detecting whether each solid-state drive has a system-level error.

[0061] In this embodiment, system-level errors include timeout errors and slow disks. A timeout error occurs when an input / output operation fails to complete within the specified time due to network delays, insufficient resources, or excessive operation time, thereby triggering a timeout error. A slow disk error occurs when a solid-state drive (SSD) responds slowly to read / write requests.

[0062] In this embodiment, a timer is started to periodically detect whether there is any system-level error information of each solid-state drive in the dmesg log within a preset timing period.

[0063] In this embodiment, the preset timing period may be 20 minutes, and the timer detection time is 2 seconds. For example, within 20 minutes, the dmesg log is checked every 2 seconds to see if there is any system-level error information of each solid-state drive.

[0064] In this embodiment, the preset timing period and the time of the timing detection can be adjusted according to actual needs.

[0065] S2013: If no system-level error occurs in each solid-state drive, the latest file of each type of file is obtained.

[0066] Specifically, the number of files of each type is obtained; if the number of files is a preset value, the latest file of each type is obtained according to the file name of each type of file; wherein the file name of each type of file is determined according to the file type and the current iteration cycle number.

[0067] For example, in the second iteration cycle, the operating data of multiple solid-state drives collected by four processes are saved in four types of files, and the file names are 1-2, 2-2, 3-2, and 4-2. Previously, in the first iteration cycle, the operating data of multiple solid-state drives collected by four processes were also saved in four types of files, and the file names were 1-1, 2-1, 3-1, and 4-1. At this time, for the first type of file, it includes 1-1 and 1-2; for the second type of file, it includes 2-1 and 2-2; for the third type of file, it includes 3-1 and 3-2; for the fourth type of file, it includes 4-1 and 4-2. The number of files of each type is 2.

[0068] In this embodiment, the preset number is 2.

[0069] Specifically, according to the file names of the files of each type, the steps of obtaining the latest files of each type include Sa~Sd:

[0070] Sa: Get the number of iteration cycles in the file name of each type of file.

[0071] Exemplarily, in the second iteration cycle, for the first type of file, the corresponding iteration cycle numbers are 1 and 2 respectively; for the second type of file, the corresponding iteration cycle numbers are 1 and 2 respectively; for the third type of file, the corresponding iteration cycle numbers are 1 and 2 respectively; for the fourth type of file, the corresponding iteration cycle numbers are 1 and 2 respectively.

[0072] Sb: For each type of file, obtain the file name with the largest number of iteration cycles.

[0073] For example, for the first type of file, the maximum number of iteration cycles is 2, and the corresponding file name is 1-2; for the second type of file, the maximum number of iteration cycles is 2, and the corresponding file name is 2-2; for the third type of file, the maximum number of iteration cycles is 2, and the corresponding file name is 3-2; for the fourth type of file, the maximum number of iteration cycles is 2, and the corresponding file name is 4-2.

[0074] Sc: The file corresponding to the file name with the largest number of iteration cycles is determined as the latest file of each type.

[0075] For example, the file names with the largest number of iteration cycles are 1-2, 2-2, 3-2, and 4-2.

[0076] Sd: For each file type, delete all files except the latest file to obtain the latest file of each type.

[0077] For example, for the first type of files, only 1-2 are retained; for the second type of files, only 2-2 are retained; for the third type of files, only 3-2 are retained; and for the fourth type of files, only 4-2 are retained.

[0078] In this embodiment, the purpose of obtaining the latest file of each type is to delete the files of each type in the previous iteration cycle in each iteration cycle and only retain the files of each type in the current iteration cycle.

[0079] Optionally, if the number of files is not a preset value, the current iteration cycle number is increased by one, and the next iteration is entered until a system-level error is detected in any solid-state drive, and the iteration is stopped.

[0080] In this embodiment, only in the first iteration cycle does the number of files not equal 2. At this point, the current iteration cycle number is incremented by one, and the second iteration cycle begins. Multiple processes collect operating data from multiple solid-state drives and save it to newly created files of different types, namely, 1-2, 2-2, 3-2, and 4-2.

[0081] In this embodiment, before obtaining the latest file of each type of file, it is necessary to determine the remaining space for storing the latest file of each type of file to ensure that there is enough space for storing the latest file of each type of file.

[0082] Specifically, the steps of determining the remaining space for storing the latest file of each type of file include Se~Sh:

[0083] Se: Check in turn whether the operating data and manufacturer log of each solid-state drive have been backed up.

[0084] Specifically, step Se includes Se1 to Se5:

[0085] Se1: traverse the drive letter name of each solid-state drive from the pre-built solid-state drive list in sequence; the drive letter name carries a backup mark.

[0086] In this embodiment, the pre-built solid-state drive list includes drive letter names of multiple solid-state drives.

[0087] Se2: During the traversal process, based on the backup mark, determine whether the operating data and manufacturer log of the solid-state drive corresponding to the current drive letter name have been backed up.

[0088] In this embodiment, the backup mark includes backed up and not backed up.

[0089] Se3: If the backup mark is backed up, it is determined that the operating data and manufacturer log of the solid-state drive corresponding to the current drive letter name have been backed up.

[0090] In this embodiment, if the backup flag indicates "not backed up," a determination is made as to whether the operating data and manufacturer logs of the solid-state drive corresponding to the current drive letter need to be backed up. If backup is required, the backup is performed; if not, the next solid-state drive's drive letter is traversed. The case where the backup flag indicates "not backed up" is described in subsequent embodiments.

[0091] Se4: Traverse the drive letter name of the next solid state drive.

[0092] Se5: Based on the backup mark corresponding to the drive letter name of the next solid-state drive, determine whether the operation data and manufacturer log of the next solid-state drive have been backed up, until the drive letter names of each solid-state drive in the pre-built solid-state drive list are traversed.

[0093] Sf: If the operating data and manufacturer logs of each solid-state drive have been backed up, the backup data is obtained; the path for saving the backup data is the same as the path for saving various types of files.

[0094] Sg: Get the remaining space parameters of the current directory for various types of files.

[0095] In this embodiment, the backup data and the various types of files are saved in the same path, that is, the backup data is also saved in the current directory of the various types of files.

[0096] Sh: If the remaining space parameter is not less than the preset space parameter, the step of obtaining the latest file of each type of file is executed.

[0097] In this embodiment, the preset spatial parameter may be 5G. The preset spatial parameter may be adjusted according to actual needs.

[0098] In this embodiment, if the remaining space parameter is not less than the preset space parameter, the step of obtaining the latest file of each type of file is executed. If the remaining space parameter is less than the preset space parameter, the collection of operating data of multiple solid-state drives is terminated, and a prompt is output that the current directory has insufficient remaining space.

[0099] In this embodiment, during each iteration, a determination is made as to whether the remaining space parameter is not less than a preset space parameter. If the remaining space parameter is less than the preset space parameter, it indicates insufficient remaining space. The collection of operating data from the multiple solid-state drives is terminated to prevent the operating system from crashing due to insufficient remaining space.

[0100] S2014: Iteration stops until a system-level error is detected in any solid-state drive.

[0101] In this embodiment, the condition for stopping iteration is that a system-level error message of any solid-state drive is detected from the dmesg log, and the solid-state drive is determined to be faulty.

[0102] In this embodiment, the backup data is data backed up when each solid-state drive has a new error count. The backup data is data before the system-level error occurs and is used for diagnosis when any solid-state drive fails.

[0103] S202: Save system-level error information and device information of any solid-state drive to a log save file.

[0104] In this embodiment, the system-level error information of any solid-state drive is obtained from the dmesg log, and the device information of any solid-state drive is obtained through a command.

[0105] In this embodiment, the system-level error information and device information of any solid-state drive are data after any solid-state drive fails.

[0106] S203: within a preset time, continue to collect the operating data of multiple solid state drives, and save the operating data to the latest file of each type of file; wherein the latest file of each type of file and the log save file are used to diagnose the fault of any solid state drive.

[0107] In this embodiment, the preset time can be 30 seconds. When any solid-state drive fails, the operating data of the multiple solid-state drives continues to be collected within 30 seconds. The operating data obtained at this time is the data after the failure of the solid-state drive. The operating data of the multiple solid-state drives collected within 30 seconds is saved to the latest file of each type of file. At this time, the latest file of each type of file saves the data before and after the failure of the solid-state drive.

[0108] In this embodiment, the preset time can be adjusted according to actual needs.

[0109] In this embodiment, the fault diagnosis is performed on any solid-state hard disk that has failed using the latest files of each type and log save files.

[0110] In summary, multiple SSDs are iteratively monitored, with multiple processes enabled to collect operational data from them and save it to different types of files. During each iteration, each SSD is periodically checked for system-level errors. If not, the latest file of each file type is retrieved. This process continues until a system-level error is detected on any SSD, indicating a failure on that SSD. The latest file of each file type contains the data before the failure of that SSD. After a system-level error occurs, operational data from multiple SSDs is continuously collected within a preset timeframe and saved to the latest file of each type. The system-level error information and device information for that SSD are saved to a log file. Post-failure data now includes data from the latest file of each file type and data from the log file. This eliminates the need to replicate the failure scenario. During the monitoring process of multiple SSDs, system-level errors are periodically detected on each SSD, automatically acquiring both pre-failure and post-failure data. This improves data acquisition efficiency and accuracy, thereby enhancing the efficiency of SSD failure diagnosis. In addition, firstly, backup data is obtained. The backup data is the data before the system fails, and is also used to diagnose any solid-state hard drive, playing an auxiliary diagnosis role, further improving the efficiency of diagnosing solid-state hard drive failures; secondly, the remaining space parameters of the current directory of multiple types of files are obtained, and execution will only continue when the remaining space parameters are not less than the preset space parameters, ensuring that there is enough remaining space to enter the next iteration; thirdly, at the end of each iteration, only the latest files of each type are retained to prevent insufficient remaining space due to repeated accumulation of files; fourthly, there is no need to manually obtain data, reducing labor costs.

[0111] Based on the above embodiment, in this embodiment, the situation where the operating data and manufacturer log of the solid-state drive corresponding to the current drive letter name are not backed up is introduced, as detailed below:

[0112] S301: If the backup mark is not backed up, obtain the manufacturer log of the solid state drive corresponding to the current drive letter name.

[0113] In this embodiment, the manufacturer log stores key data such as the underlying hardware operating status and error information.

[0114] S302: Obtain error counts of preset key fields in the manufacturer's log.

[0115] In this embodiment, the preset key fields include an error check code and a command timeout, etc.

[0116] S303: Determine whether there is a new error count.

[0117] In this embodiment, the current error count is compared with the error count in the last backup data to determine whether there is a new count.

[0118] S304: If it is determined that the error count has a new count, the operation data of the solid-state hard disk corresponding to the current drive letter name is obtained from the operation data of the plurality of solid-state hard disks.

[0119] In this embodiment, if the error count increases, it indicates that the corresponding solid-state drive may have an abnormality in the current iteration cycle. For example, if the error count of the error check code increases, it may indicate that the flash memory chip is damaged; if the error count of the command timeout increases, it may indicate that the controller firmware is abnormal.

[0120] S305: Back up the manufacturer log of the solid state drive corresponding to the current drive letter name and the operating data of the solid state drive corresponding to the current drive letter name.

[0121] S306: Modify the backup mark corresponding to the current drive letter name to backed up.

[0122] In this embodiment, the backup mark corresponding to the current drive letter name is modified to backed up in order to avoid repeated backing up of the operating data and manufacturer log of the solid-state drive corresponding to the current drive letter name in the current iteration cycle.

[0123] S307: If it is determined that the error count has no new count, the manufacturer log of the solid state drive corresponding to the current drive letter name and the operating data of the solid state drive corresponding to the current drive letter name are not backed up.

[0124] In this embodiment, if it is determined that the error count has no new count, it means that the corresponding solid-state drive has no abnormality in the current iteration cycle.

[0125] It should be noted that after the current iteration cycle ends, the backup flag corresponding to each drive letter name needs to be changed to "not backed up". In the next iteration cycle, when the backup flag corresponding to each drive letter name determines that the operating data and manufacturer log of the SSD corresponding to each drive letter name have not been backed up, steps S301 to S07 will be executed. The backup will only be performed if the error count of the preset key field has increased. If there is no increase in the error count, the backup data that has already been backed up will be used.

[0126] In summary, if the operating data and manufacturer logs of each solid-state drive have not been backed up, the manufacturer log of the solid-state drive corresponding to the current drive letter name is obtained; based on the error counts of the preset key fields in the manufacturer log, determine whether a backup is required. If there are new counts in the error counts, it means that the solid-state drive may have an abnormality in the current iteration cycle, and the operating data and manufacturer logs of the solid-state drive corresponding to the current drive letter name need to be backed up; if there are no new counts, it means that the solid-state drive has no abnormality in the current iteration cycle, and no backup is performed. If there are new counts in the error counts of the preset key fields, it means that the current solid-state drive may fail. Therefore, only backing up when there are new counts in the error counts of the preset key fields can retain key data and save remaining space.

[0127] In the above embodiment, the data processing method for SSD failures requires that the operating system sends services to the SSD, meaning that the SSD receives read and write requests from the upper layer. If no services are sent to the SSD at the operating system level, the operating system will not report a system-level SSD error.

[0128] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0129] Figure 3 This is a schematic diagram of the structure of the data processing device for solid state hard disk failure provided in the embodiment of the present application. Figure 3 As shown, an embodiment of the present application further provides a data processing device for a solid-state drive failure, comprising: a monitoring module 301, a first storage module 302, and a second storage module 303. The monitoring module 301 comprises: a collection submodule 3011, a detection submodule 3012, a determination submodule 3013, and a stop iteration submodule 3014.

[0130] The monitoring module 301 is used to monitor multiple solid-state drives and iteratively perform the following steps:

[0131] The monitoring module 301 includes:

[0132] The collection submodule 3011 is used to enable multiple processes to collect operating data from multiple solid-state drives in the current iteration cycle, and save the operating data into multiple types of files; the operating data collected by different processes are saved into different types of files;

[0133] The detection submodule 3012 is used to periodically detect whether a system-level error occurs in each solid-state drive within a preset timing period;

[0134] The determination submodule 3013 is configured to obtain the latest file of each type if no system-level error occurs in each solid-state drive;

[0135] The stop iteration submodule 3014 is configured to stop iteration until a system-level error is detected in any solid-state drive;

[0136] A first saving module 302 is configured to save system-level error information and device information of any solid-state drive to a log saving file;

[0137] The second saving module 303 is used to continue collecting the operating data of multiple solid-state drives within a preset time and save the operating data to the latest file of each type of file; wherein the latest file of each type of file and the log save file are used to diagnose faults of any solid-state drive.

[0138] In a possible implementation, the determination submodule 3013 includes:

[0139] A first obtaining unit is used to obtain the number of files of each type;

[0140] The second acquisition unit is used to acquire the latest file of each type according to the file name of each type of file if the number of files is the preset value; wherein the file name of each type of file is determined according to the type of file and the current iteration cycle number.

[0141] In a possible implementation, the second acquiring unit includes:

[0142] The first acquisition subunit is used to obtain the number of iteration cycles in the file name of each type of file;

[0143] The second acquisition subunit is used to obtain the file name with the largest number of iteration cycles for each type of file;

[0144] A determination subunit, configured to determine the file corresponding to the file name with the largest number of iteration cycles as the latest file of each type of file;

[0145] The third acquisition subunit is used to delete all files except the latest file for each type of file to obtain the latest file of each type of file.

[0146] In a possible implementation, the monitoring module 301 further includes a judgment submodule, which includes:

[0147] a judgment unit, configured to sequentially judge whether the operating data and manufacturer log of each solid-state drive have been backed up;

[0148] A third acquisition unit is configured to acquire backup data if the operating data and manufacturer log of each solid-state drive have been backed up; wherein the backup data is stored in the same path as that of the various types of files;

[0149] A fourth obtaining unit is used to obtain free space parameters of a current directory of files of various types;

[0150] The execution unit is configured to execute the step of obtaining the latest file of each type of file if the remaining space parameter is not less than the preset space parameter.

[0151] In a possible implementation, the judging unit includes:

[0152] The first traversal sub-unit is used to traverse the drive letter name of each solid-state drive from the pre-built solid-state drive list in sequence; wherein the drive letter name carries a backup mark;

[0153] The judgment subunit is used to judge whether the operation data and manufacturer log of the solid-state drive corresponding to the current drive letter name have been backed up according to the backup mark during the traversal process;

[0154] a determination subunit, configured to determine, if the backup mark is backed up, whether the operating data and the manufacturer log of the solid-state drive corresponding to the current drive letter name have been backed up;

[0155] The second traversal subunit is used to traverse the drive letter name of the next solid state drive;

[0156] The third traversal sub-unit is used to determine whether the operation data and manufacturer log of the next solid-state drive have been backed up according to the backup mark corresponding to the drive letter name of the next solid-state drive, until the drive letter name traversal of each solid-state drive in the pre-built solid-state drive list is completed.

[0157] In a possible implementation, the judgment unit further includes a backup subunit, and the backup subunit includes:

[0158] A first acquisition component is used to obtain the manufacturer log of the solid-state drive corresponding to the current drive letter name if the backup mark is not backed up;

[0159] The second acquisition component is used to obtain the error count of the preset key field in the manufacturer log;

[0160] A judgment component is used to determine whether there is a new count in the error count;

[0161] A first determination component is configured to obtain, from the operation data of the plurality of solid-state hard disks, the operation data of the solid-state hard disk corresponding to the current drive letter name if it is determined that the error count has a new count;

[0162] A backup component is used to back up the manufacturer log of the solid-state drive corresponding to the current drive letter name and the operating data of the solid-state drive corresponding to the current drive letter name;

[0163] The modification component is used to change the backup mark corresponding to the current drive letter name to backed up;

[0164] The second determination component is configured to not back up the manufacturer log of the solid-state drive corresponding to the current drive letter name and the operating data of the solid-state drive corresponding to the current drive letter name if it is determined that the error count has no new count.

[0165] In a possible implementation, the data processing device for a solid state drive failure further includes a building module, the building module including:

[0166] First, create a submodule to create a log save file;

[0167] Get submodule, used to get the drive letter names of multiple solid-state drives;

[0168] The second creation submodule is used to create an empty solid-state drive list;

[0169] The save submodule saves the drive letter names of multiple solid-state drives to an empty solid-state drive list to obtain a pre-built solid-state drive list.

[0170] For the description of the features in the embodiment corresponding to the data processing device for solid state hard disk failure, please refer to the relevant description of the embodiment corresponding to the data processing method for solid state hard disk failure, and no further details will be given here.

[0171] Figure 4 This is a schematic diagram of the structure of the server provided in the embodiment of the present application. Figure 4 As shown, the server provided in this embodiment includes: at least one processor 401 and a memory 402. Optionally, the server also includes a communication component 403. The processor 401, the memory 402 and the communication component 403 are connected via a bus.

[0172] In a specific implementation process, at least one processor 401 executes the computer-executable instructions stored in the memory 402 , so that the at least one processor 401 executes the above-mentioned embodiment of the method for processing data of a solid-state drive failure.

[0173] The specific implementation process of the processor 401 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0174] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0175] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0176] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0177] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned embodiments of the method for processing data of a solid-state hard drive failure when running.

[0178] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0179] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned data processing methods for solid-state hard drive failures are implemented.

[0180] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps of any of the above-mentioned solid-state hard drive failure data processing method embodiments.

[0181] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0182] The above is a detailed introduction to the data processing method, device, server and storage medium for a solid-state hard drive failure provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for processing data of a solid state hard disk failure, characterized in that: include: Monitor multiple SSDs and iteratively perform the following steps: In a current iteration cycle, multiple processes are enabled to collect operating data of the multiple solid-state drives, and the operating data is saved in multiple types of files; The operation data collected by different processes are saved in different types of files; During a preset timing period, each solid-state drive is regularly checked for system-level errors; If the system-level error does not occur in any of the solid-state drives, obtaining the latest file of each type of file; The iteration stops until the system-level error is detected in any solid-state drive; Saving system-level error information and device information of any solid-state drive to a log save file; within a preset time, continuing to collect the operating data of the plurality of solid-state drives, and saving the operating data to the latest file of each type of file; The latest files of each type of files and the log save files are used to perform fault diagnosis on any solid state drive.

2. The method according to claim 1, characterized in that The method of obtaining the latest files of various types includes: Get the number of files of each type; If the number of files is a preset value, the latest file of each type of file is obtained according to the file name of each type of file; wherein the file name of each type of file is determined according to the type of file and the current iteration cycle number.

3. The method according to claim 2, characterized in that The acquiring of the latest file of each type of file according to the file name of each type of file includes: Obtain the number of iteration cycles in the file name of each type of file; For each type of file, obtain the file name with the largest number of iteration cycles; Determine the file corresponding to the file name with the largest number of iteration cycles as the latest file of each type; For each type of file, all files except the latest file are deleted to obtain the latest file of each type of file.

4. The method according to claim 1, wherein Before obtaining the latest file of each type of file, the method further includes: Determining in sequence whether the operating data and manufacturer logs of each solid-state drive have been backed up; If the operation data and manufacturer logs of each solid-state drive have been backed up, the backup data is obtained; wherein the path for storing the backup data is the same as the path for storing the multiple types of files; Obtaining free space parameters of the current directory for the multiple types of files; If the remaining space parameter is not less than the preset space parameter, the step of obtaining the latest file of each type of file is executed.

5. The method according to claim 4, characterized in that The step of sequentially determining whether the operating data and manufacturer logs of each solid-state drive have been backed up includes: Traversing the drive letter names of the solid-state hard disks in sequence from a pre-built solid-state hard disk list; wherein the drive letter names carry a backup mark; During the traversal process, based on the backup mark, it is determined whether the operating data and manufacturer log of the solid-state drive corresponding to the current drive letter name have been backed up; If the backup mark is backed up, it is determined that the operating data and manufacturer log of the solid-state drive corresponding to the current drive letter name have been backed up; Traverse the drive letter name of the next solid state drive; According to the backup mark corresponding to the drive letter name of the next solid-state drive, determine whether the operating data and manufacturer log of the next solid-state drive have been backed up until the drive letter names of each solid-state drive in the pre-built solid-state drive list are traversed.

6. The method according to claim 5, characterized in that After determining whether the operating data and manufacturer log of the solid-state drive corresponding to the current drive letter name have been backed up according to the backup mark, the method further includes: If the backup mark is not backed up, obtaining the manufacturer log of the solid-state drive corresponding to the current drive letter name; Obtaining error counts of preset key fields in the manufacturer's log; Determine whether there is a new count in the error count; If it is determined that the error count has a new count, obtaining the operating data of the solid-state hard disk corresponding to the current drive letter name from the operating data of the plurality of solid-state hard disks; Backing up the manufacturer log of the solid-state hard disk corresponding to the current drive letter name and the operating data of the solid-state hard disk corresponding to the current drive letter name; Modify the backup mark corresponding to the current drive letter name to "backed up"; If it is determined that the error count has no new count, the manufacturer log of the solid state drive corresponding to the current drive letter name and the operating data of the solid state drive corresponding to the current drive letter name are not backed up.

7. The method according to claim 5, characterized in that Before monitoring the multiple solid state drives, the method further includes: Create a log save file; Get the drive letter names of multiple solid-state drives; Create an empty SSD list; The drive letter names of the multiple solid-state hard drives are saved to the empty solid-state hard drive list to obtain a pre-built solid-state hard drive list.

8. A data processing device for a solid state hard disk failure, characterized in that: include: The monitoring module is used to monitor multiple solid-state drives and iteratively execute the following steps: Wherein, the monitoring module includes: The collection submodule is configured to enable multiple processes to collect the operating data of the multiple solid-state drives in the current iteration cycle, and save the operating data into multiple types of files; wherein the operating data collected by different processes are saved into different types of files; The detection submodule is used to regularly detect whether each solid-state drive has a system-level error within a preset timing period; a determination submodule, configured to obtain the latest file of each type of file if none of the solid state drives has the system-level error; The stop iteration submodule is configured to stop iteration until the system-level error is detected in any solid-state drive; A first saving module, configured to save the system-level error information and device information of any solid-state drive to a log saving file; The second saving module is used to continue collecting the operating data of the multiple solid-state hard drives within a preset time, and save the operating data to the latest file of each type of file; wherein the latest file of each type of file and the log saving file are used to perform fault diagnosis on any of the solid-state hard drives.

9. A server, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for processing data of a solid state drive failure as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for processing data of a solid-state hard disk failure according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Solid state disk management method, device and equipment, medium and product

    CN119724313A

  • Fault prediction method and device and baseboard management controller

    CN119883843A