A method, server, medium, and program product for fast detection of SSD data anomalies

The method uses machine learning to create SSD type adaptation models with adjustable anomaly thresholds, addressing performance disparities and environmental factors for precise SSD data anomaly detection, enhancing detection accuracy and reliability.

CN119811467BActive Publication Date: 2025-07-15SHENZHEN G-BONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510289350.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-15
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

In the prior art, due to the differences in the performance parameters of SSDs of different manufacturers and models, fixed or unified threshold parameters are difficult to accurately detect SSD data abnormalities, and it is impossible to efficiently adapt accurately on different SSDs.

Method used

By obtaining performance data of different types of SSDs, using machine learning algorithms to train the SSD type adaptation model, dynamically adjust the abnormal threshold parameters, and combining the working environment and service life data acquired by the sensor, correct the abnormality detection model to achieve fast and accurate detection of SSD data abnormalities.

Benefits of technology

It realizes fast and accurate data abnormality detection of different SSD models, reduces misjudgment, timely isolates faulty SSDs, and ensures the stable operation of the server and data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811467B_ABST
    Figure CN119811467B_ABST
Patent Text Reader

Abstract

The present application provides a method for quickly detecting SSD data anomalies, a server, a medium, and a program product, relating to the technical field of solid-state drive detection. The method includes obtaining performance data of different types of SSDs as samples, training an SSD type adaptation model through a machine learning algorithm for determining the type of the current SSD. Then, after obtaining new performance data of a new SSD, inputting the new performance data into the model to determine the SSD type. Next, adjusting the anomaly threshold parameter of the SSD data anomaly determination model according to the SSD type, and inputting the new performance data into the model to determine the data anomaly situation of the SSD. Implementing this method can promptly detect data anomalies of different types of SSDs, thereby ensuring the security and stability of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of solid-state drive (SSD) detection, and particularly to a method for quickly detecting SSD data anomalies, a server, a medium, and a program product. Background Art

[0002] With the development of information technology, the usage of various servers has been increasing continuously. Due to advantages such as small volume and fast read speed, solid-state drives (SSDs) are widely used in servers. To ensure the normal operation of the servers, it is necessary to regularly detect the SSDs in the servers to find out existing data anomalies.

[0003] The prior art detects all SSDs by setting a fixed threshold for judging abnormal data. The specific approach is to first collect normal and abnormal data of different SSDs, and set threshold parameters for determining anomalies based on this data. Then, when detecting a new SSD, directly compare the preset threshold parameters with the data of the new SSD to determine whether it is abnormal.

[0004] However, due to differences in performance parameters among SSDs of different manufacturers and different models, using fixed or unified threshold parameters for detection will result in poor detection effects for data anomalies of some models of SSDs, making it difficult to meet the requirement of accurately judging anomalies. Therefore, it is difficult for the existing methods to achieve efficient detection of abnormal data on different SSDs based on accurate adaptation to various unknown SSDs. Summary of the Invention

[0005] This application provides a method for quickly detecting SSD data anomalies, a server, a medium, and a program product, which are used to accurately and quickly detect data anomalies of different types of SSDs when the prior art is difficult to adapt to performance differences of different SSDs.

[0006] In a first aspect, this application provides a method for quickly detecting SSD data anomalies, which is applied to a server. The method includes: obtaining performance data of different types of SSDs as sample data; training an SSD type adaptation model through a machine learning algorithm based on the sample data, where the SSD type adaptation model is used to determine the type of the current SSD; after obtaining new performance data of a new SSD, inputting the new performance data into the SSD type adaptation model to determine the SSD type; adjusting the abnormal threshold parameters corresponding to the SSD data anomaly determination model according to the SSD type, where the SSD data anomaly determination model is a model constructed based on multiple normal and abnormal data of SSDs obtained in advance; and inputting the new performance data into the SSD data anomaly determination model to determine the data anomaly situation of the SSD.

[0007] By adopting the above technical solutions, it is possible to quickly and accurately determine the type of the current SSD by using the SSD type adaptation model trained by machine learning algorithms. Then, the abnormal threshold parameter of the SSD data abnormality determination model is adjusted according to the SSD type, and the new performance data is input into the model, so as to realize the rapid detection of the SSD data abnormality. This method improves the detection efficiency and accuracy, can timely detect SSD data abnormalities, and ensures the security and stability of data.

[0008] Combined with some embodiments of the first aspect, in some embodiments, after the step of inputting the new performance data into the SSD data abnormality determination model to determine the SSD data abnormality situation, it further includes: if the data abnormality situation is that the new performance data is abnormal, sending an SSD data abnormality alarm message to the management end; after determining the abnormal new performance data as abnormal data, feeding back the abnormal data to the SSD data abnormality determination model for updating.

[0009] By adopting the above technical solutions, relevant personnel can be notified in time for handling, reducing the losses caused by data abnormalities. At the same time, feeding back the abnormal data to the SSD data abnormality determination model for updating helps to improve the accuracy and adaptability of the model, enabling it to better detect future SSD data abnormality situations.

[0010] Combined with some embodiments of the first aspect, in some embodiments, the step of sending an SSD data abnormality alarm message to the management end if the data abnormality situation is that the new performance data is abnormal specifically includes: determining the log information when the new performance data sends an abnormality; determining the association rule between the data abnormality situation and the log information; when the log information meets the association rule corresponding to the data abnormality situation, sending an SSD data abnormality alarm message to the management end in advance.

[0011] By adopting the above technical solutions, potential problems can be discovered more timely, improving the accuracy and timeliness of fault warning, and further ensuring the stable operation of the server.

[0012] Combined with some embodiments of the first aspect, in some embodiments, after the step of adjusting the abnormal threshold parameter corresponding to the SSD data abnormality determination model according to the SSD type, it further includes: obtaining the working environment data of the new SSD through a sensor, where the working environment data at least includes the temperature and humidity in the working state of the new SSD; inputting the working environment data into the environment impact model to obtain SSD performance impact data, and the environment impact model is a model established according to the impact of different temperature and humidity environments on SSD performance data obtained in advance; correcting the abnormal threshold parameter in the SSD data abnormality determination model according to the SSD performance impact data.

[0013] By adopting the above technical solution, it is possible to more accurately determine whether the SSD data is abnormal, reduce misjudgment caused by environmental factors, and improve the reliability of detection.

[0014] Combined with some embodiments of the first aspect, in some embodiments, after the step of correcting the SSD data anomaly determination model for the anomaly threshold parameter according to the SSD performance affecting the data, the method further includes: obtaining the service life data of the new SSD; and readjusting the anomaly threshold parameter in the SSD data anomaly determination model according to the service life data.

[0015] By adopting the above technical solution, it is possible to consider the impact of the service life of the SSD on its performance, make the anomaly threshold parameter more reasonable, improve the accuracy and pertinence of detection, and better ensure the normal operation of the SSD.

[0016] Combined with some embodiments of the first aspect, in some embodiments, after the step of inputting the new performance data into the SSD data anomaly determination model to determine the data anomaly situation of the SSD, the method further includes: when the data anomaly situation is that the new performance data continuously matches the anomaly threshold parameter for multiple times, determining that the new SSD has a fault; obtaining the current operating state data of the server; inputting the operating state data into a fault impact model, and the fault impact model determines the type and degree of impact of the fault on the server performance and data integrity according to the operating state data; when it is determined that the fault impact type meets the set conditions, marking the physical location of the new SSD; and determining the isolation method of the new SSD according to the physical location.

[0017] By adopting the above technical solution, it is possible to isolate the faulty SSD in time, reduce its impact on the server, and ensure the stable operation and data security of the server.

[0018] Combined with some embodiments of the first aspect, in some embodiments, after the step of inputting the new performance data into the SSD data anomaly determination model to determine the data anomaly situation of the SSD, the method further includes: sending a check prompt message to the management end according to the data anomaly situation; after receiving the check result feedback from the management end, determining the abnormal cause of the new performance data according to the check result; and determining an optimization plan according to the abnormal cause, where the optimization plan is used to eliminate the factors causing the new performance data to be abnormal, so that the new SSD returns to the normal working state.

[0019] By adopting the above technical solution, it is possible to take measures in time to eliminate the factors causing data anomalies, make the new SSD return to the normal working state, and improve the reliability and stability of the SSD.

[0020] In a second aspect, the present application provides a server, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the server to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0021] In a third aspect, the present application provides a computer-readable storage medium, including instructions that, when running on a server, cause the server to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0022] In a fourth aspect, the present application provides a computer program product that, when running on a server, causes the server to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0023] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0024] 1. Since the technical means of obtaining performance data of different types of SSDs as sample data, training an SSD type adaptation model through a machine learning algorithm, and adjusting the abnormal threshold parameters corresponding to the SSD data abnormality determination model according to the SSD type are adopted, the problem of low efficiency and easy error in manually identifying the SSD type in the prior art is effectively solved, the SSD type can be quickly and accurately identified, and then the technical effect of accurately detecting the SSD data abnormality is realized.

[0025] 2. Since the technical means of obtaining the working environment data of a new SSD through a sensor and inputting it into an environment impact model to obtain SSD performance impact data, and then correcting the abnormal threshold parameters in the SSD data abnormality determination model according to this data are adopted, the technical problem of misjudgment in SSD data abnormality detection due to the failure to consider environmental factors in the prior art is effectively solved, and then the technical effect of more accurately judging whether the SSD data is abnormal is realized.

[0026] 3. Since the technical means of determining that a new SSD has a fault when the new performance data continuously matches the abnormal threshold parameters multiple times, obtaining the current running state data of the server, inputting it into a fault impact model to determine the type and degree of the impact of the fault on the server performance and data integrity, and then marking the physical location of the new SSD and determining the isolation method are adopted, the technical problem of being unable to isolate the faulty SSD in time to reduce its impact on the server in the prior art is effectively solved, and then the technical effect of timely isolating the faulty SSD, ensuring the stable operation of the server and data security is realized. Description of the Drawings

[0027] Figure 1 It is a schematic flowchart of a method for quickly detecting SSD data anomalies in an embodiment of the present application;

[0028] Figure 2 It is another schematic flowchart of a method for quickly detecting SSD data anomalies in an embodiment of the present application;

[0029] Figure 3 It is a schematic structural diagram of an entity device of a server in an embodiment of the present application. Detailed implementation manners

[0030] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items.

[0031] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0032] For ease of understanding, an application scenario in a related method is introduced below.

[0033] In a large data center, a vast amount of data is stored, and solid-state drives (SSDs) from different manufacturers and different models are used. To ensure the security and stability of the data, the administrator of the data center adopts the unified threshold detection method in the prior art to monitor the data anomalies of the SSDs. One day, a batch of high-performance SSDs from different manufacturers was newly introduced into the data center. During operation, the performance parameters of these new SSDs are very different from those of the previously used SSDs. When detecting according to the fixed threshold parameters, some of the new SSDs that are actually in a normal working state are misjudged as having data anomalies and frequently issue alarms, making the administrator overwhelmed. At the same time, some new SSDs that actually have data anomalies are not detected in time, posing a great risk to data security.

[0034] It can be seen that the existing method obviously cannot meet the requirements of accurately adapting to various unknown SSDs and achieving efficient detection of abnormal data.

[0035] By adopting the SSD data anomaly fast detection method in the embodiments of the present application, by obtaining the performance data of different types of SSDs as sample data, and training an SSD type adaptation model through a machine learning algorithm based on the sample data, the type of the current SSD can be determined quickly and accurately. Moreover, the anomaly threshold parameters corresponding to the SSD data anomaly determination model can be adjusted according to the SSD type, and the new performance data is input into the model, so as to accurately detect the data anomaly situation of the SSD.

[0036] For the convenience of understanding, the method provided in this embodiment will be described in terms of its process below in combination with the above content. Please refer to Figure 1 which is a schematic flow diagram of the SSD data anomaly fast detection method in the embodiments of the present application.

[0037] S101. Obtain the performance data of different types of SSDs as sample data;

[0038] The server first needs to collect the performance data of different types of SSDs as sample data. Specifically, the server can connect to SSDs produced by various manufacturers, and then read and collect the performance parameters of these SSDs, such as read and write speeds, I / O response times, life cycles and other data. To obtain sufficient and comprehensive sample data, the server needs to connect as many different manufacturers and different models of SSDs as possible, at least covering the mainstream products of mainstream manufacturers.

[0039] S102. Train an SSD type adaptation model through a machine learning algorithm based on the sample data, and the SSD type adaptation model is used to determine the type of the current SSD;

[0040] The server inputs the rich SSD performance sample data collected into a machine learning algorithm for training to obtain an SSD type adaptation model. This model can accurately identify the characteristics of different types of SSDs and is used to quickly determine the category and model of a new SSD. The machine learning algorithm can adopt a neural network. The server can construct a deep neural network including an input layer, multiple hidden layers and an output layer. The number of nodes in the input layer corresponds to the number of features of the sample data, and the number of nodes in the output layer corresponds to the number of SSD types. The hidden layers can be set in multiple layers stacked, and the number of nodes gradually decreases in each layer, increasing the network's ability to extract and abstract the features of the sample data.

[0041] The server feeds the rich SSD performance sample data collected into the neural network for iterative training, adjusts the weights of the network edges, so that it fits the complex mapping relationship between the SSD performance data and the SSD type, and obtains a model that can accurately identify the SSD type of the newly input SSD performance data. This model contains both the SSD performance data characteristics and the SSD type output, and the new SSD type adaptation model can quickly determine its category.

[0042] S103. After obtaining the new performance data of the new SSD, input the new performance data into the SSD type adaptation model to determine the SSD type.

[0043] After the server connects to the new SSD, it will obtain the latest performance data of the SSD. The server can obtain multi-faceted performance data such as the read and write speeds, IOPS, and average response time of the SSD through collection commands. The server inputs these newly obtained performance data into the already trained SSD type adaptation model for inference calculation. The model analyzes the characteristics of these data and outputs that the new SSD is likely to be of the 512GB type of a certain brand of mobile phone.

[0044] Through the accurate judgment of the model here, the server can quickly know the detailed category information of the new SSD, laying a foundation for subsequent customized anomaly detection.

[0045] S104. Adjust the anomaly threshold parameters corresponding to the SSD data anomaly determination model according to the SSD type. The SSD data anomaly determination model is a model constructed based on multiple normal data and abnormal data of the SSD obtained in advance.

[0046] The server will adjust the parameters of the model used to detect the data anomaly situation of the new SSD accordingly according to the type of the new SSD.

[0047] The SSD data anomaly detection model is pre-trained by the server based on a large amount of normal and abnormal data of the collected SSDs. For different types and specifications of SSDs, their performance data performances are also different. Therefore, the detection model needs to dynamically adjust the threshold parameters for judging data anomalies according to the SSD category and specifications in order to achieve accurate anomaly detection.

[0048] For example, for a PCIe SSD of the 512GB type of a certain brand of mobile phone determined, according to its read and write speed ranges, the server can determine that its read speed is lower than 3000MB / s as abnormal; its write speed is lower than 1500MB / s as abnormal; its average read latency is higher than 100μs as abnormal, and so on. After adjusting the thresholds of multiple performance parameters, the model can achieve accurate judgment for this specific model of SSD.

[0049] S105. Input the new performance data into the SSD data anomaly determination model to determine the data anomaly of the SSD.

[0050] After obtaining the parameterized anomaly detection model adjusted for the new SSD, the server can use this model to perform anomaly detection on the SSD.

[0051] The server will regularly obtain the latest performance data of the new SSD and input it into the adjusted model for calculation to determine whether the data of each performance indicator is abnormal. For example: read speed 3200MB / s → normal; write speed 1400MB / s → normal; average read latency 120μs → abnormal... By inputting into the parameterized detection model, the server can quickly and automatically determine whether there are abnormal situations in the data of the new SSD. Once an abnormality is detected, feedback measures can be quickly taken to ensure data security. In addition, the detection results can also be used to add abnormal data samples of the new SSD and feedback them to the model to enhance its detection ability. And as the model processes more new SSDs, its generalization performance will continue to improve, enabling it to adapt to more unknown SSDs, thus achieving efficient and accurate SSD data anomaly detection.

[0052] In the embodiments of the present application, since the technical means of obtaining the performance data of different types of SSDs as samples, training an SSD type adaptation model through machine learning algorithms, and dynamically adjusting the parameters of the SSD data anomaly determination model according to the SSD type are adopted, the problem that it is difficult to adapt to the performance differences of different SSDs with fixed threshold parameters in the prior art is effectively solved, and the SSD type can be accurately identified and a data anomaly detection model can be customized according to the model, thereby achieving the technical effect of quickly and accurately detecting data anomalies for different SSD models.

[0053] In some embodiments, when the server detects that the performance data of the new SSD is abnormal, it is necessary to promptly send an alarm message to the management end so that the administrator can process it as soon as possible. At the same time, it is also necessary to feedback the abnormal data to the SSD data anomaly detection model to improve the model accuracy. The sending of the alarm message can be optimized in combination with the log information to achieve a more intelligent and accurate early warning.

[0054] Specifically, when the server detects abnormal performance data, it will immediately extract the relevant log information of the SSD before the abnormality occurs, such as the interface statistics log, which records information such as the interface response time and transmission speed. Then, the server will use the association rule algorithm to determine the correlation between the detected performance anomaly situation and the interface log information. For example, if a read speed anomaly is detected, analyze whether the read speed data item in the interface log also shows an abnormal value; if an interface response timeout is detected, analyze whether the response time in the performance data is abnormal.

[0055] When the log information meets the association rules corresponding to the current performance anomaly situation, it is confirmed that there is a high correlation between the two, that is, the log information can verify and predict the occurrence of performance anomalies. At this time, the server will send a performance anomaly warning to the management end in advance. For example, the interface log of a certain PCIe SSD showed multiple sudden drops in read speed in the previous period, but the performance data was still normal. The detection model matches the interface log with the association rule of "read speed reduction", which conforms to the corresponding relationship. Then the server will determine that a read speed anomaly may be brewing and timely send a warning of "SSD read speed degradation warning" to the management end, so that the administrator has time to optimize the configuration and avoid the occurrence of faults.

[0056] In this way, through the association analysis with log information, more intelligent and accurate performance anomaly warnings can be realized, ensuring that the management personnel have enough time for fault prevention, avoiding the expansion of data anomalies into faults, and improving the system reliability.

[0057] At the same time, when the new performance data is determined to be abnormal data, the server will also feedback it to the SSD anomaly detection model to enhance the model's ability to judge abnormal data. The model can learn new types and new patterns of anomalies and continuously optimize the accuracy of the model.

[0058] In some embodiments, the server regularly monitors the performance data of all SSDs, such as read and write speeds, IOPS, response times, etc.

[0059] Compare the latest performance data of the new SSD with the normal threshold parameters. The abnormal threshold parameters can be set in advance or dynamically generated by analyzing historical data. If the latest performance data of the new SSD matches the abnormal threshold parameters continuously for multiple times (for example, 5 times), it is determined that the new SSD has a fault. Obtain the current running state data of the server, including the usage rates, workloads, IO request numbers, etc. of subsystems such as CPU, memory, and network. Input the running state data into a pre-trained fault impact model. This model can adopt algorithms such as neural networks or decision trees, and according to the running state data, predict the types and degrees of the impact of the fault on the server performance and data integrity.

[0060] The impact types can be classified into: performance degradation, data loss, service unavailability, etc. The impact degree can be expressed by probability or score. According to the output of the fault impact model, if it is determined that the fault will cause data loss or service unavailability, it meets the set severe impact conditions. Mark the physical installation location of the new SSD. For example, rack number, server number, slot number, etc. The location information can be obtained from the device asset database. If the new SSD is installed in a certain chassis of a rack server, we can directly turn off the power switch of this chassis, cut off the power supply of the new SSD, and isolate it from the system. In this way, without opening the chassis, the faulty SSD can be simply and quickly isolated by turning off the power, avoiding affecting other running servers. For a server using RAID, multiple hard disks are combined into a disk array. If the new SSD is a member of the array, we can delete or remove it from the array through the RAID management interface, removing it from the current disk array. In this way, the faulty new SSD can be isolated. At the same time, since it is removed through software logic and there is no physical power-off, the impact on the server operation is relatively small.

[0061] Through the above embodiments, the automated and intelligent management capabilities can be fully utilized to accurately locate the cause of the SSD fault, evaluate the fault impact, and finally take isolation and backup measures to ensure the reliable operation of the system.

[0062] In some embodiments, when the performance data of the new SSD is continuously abnormal, the server can push a prompt message of the new SSD fault to the management end through the alarm system. The information contains key information of the new SSD such as time, serial number, abnormal data indicators, etc. After the management end receives the prompt, the server needs to provide necessary interfaces and permissions to allow the management end to remotely log in and check the detailed status of the new SSD through tools. After the relevant personnel of the management end complete the status check of the new SSD, the results are fed back to the server. The server needs to analyze the inspection results to determine the cause of the performance abnormality, such as firmware problems, interface incompatibility, improper parameter settings, etc. According to the abnormal cause, the server can determine corresponding optimization measures, such as upgrading the SSD firmware, adjusting interface parameters, compatibility settings, etc. The goal is to eliminate the factors causing the abnormality and make the new SSD return to normal. The entire optimization process requires the server to actively control and provide status data as needed, forming a closed loop with the management end to quickly locate the problem and verify the optimization effect, so that the new SSD returns to the normal working state.

[0063] In some scenarios, in the data center server room, if the temperature or humidity environment fluctuates greatly, it will affect the SSD performance. Directly applying the preset anomaly detection model may cause false alarms. Moreover, after the SSD has been used for a period of time, its performance will show a downward trend. If the preset anomaly detection model is directly applied, it may misjudge the already degraded performance.

[0064] After combining the above scenarios, the following is a further and more specific process description of the method provided in this embodiment. Please refer to Figure 2 which is another process schematic diagram of the SSD data anomaly rapid detection method in the embodiment of the present application.

[0065] S201. Obtain the working environment data of the new SSD through a sensor. The working environment data at least includes the temperature and humidity under the working state of the new SSD;

[0066] Staff can set temperature and humidity sensors in the cabinet in advance to detect the working environment temperature and humidity of the new SSD in real time. Considering that the temperature distribution and humidity distribution may be different, the sensors need to be set close to the new SSD to directly monitor it. For example, if the new SSD is set in the middle layer of the cabinet, temperature and humidity sensors are set above and below it to obtain the temperature and humidity data of this area.

[0067] The server needs to regularly read the temperature and humidity data collected by the sensor. For example, read the temperature and humidity values at the location of the new SSD every 5 minutes and record them in the environmental data database. In this way, the real-time working environment temperature and humidity data of the new SSD can be continuously obtained.

[0068] S202. Input the working environment data into the environmental impact model to obtain SSD performance impact data. The environmental impact model is a model established based on the impact of different temperature and humidity environments on SSD performance data obtained in advance;

[0069] The server will collect a large amount of SSD performance data under different environmental conditions (temperature, humidity) as samples and input them into a machine learning algorithm to train the environmental impact model.

[0070] This model can input temperature and humidity parameters and output the performance data of the SSD under such environmental conditions, such as changes in read and write speeds, changes in failure rates, etc. The model training can use algorithms such as linear regression to learn the mapping relationship between temperature, humidity and SSD performance data.

[0071] For example, the server has collected the following sample data: temperature 20°C, humidity 50% --> average SSD read speed 3500MB / s, temperature 30°C, humidity 60% --> average SSD read speed 3200MB / s, temperature 40°C, humidity 70% --> average SSD read speed 3000MB / s...

[0072] Through training with a large number of such samples, an environmental impact model is obtained. Then, the current temperature and humidity data of the new SSD are input into the model in real time, and the impact data such as the reduction of the SSD read speed by xx MB / s and the increase of the failure rate by xx% can be quickly predicted. For example, when the current environmental temperature of the new SSD is 32°C and the humidity is 65%, after inputting into the model, the output is that the predicted read speed is reduced by 100MB / s, the Writes are reduced by 80MB / s, and the failure rate is increased by 2%, etc. These are the SSD performance impact data.

[0073] S203. Modify the abnormal threshold parameter in the SSD data abnormality determination model according to the SSD performance impact data;

[0074] After obtaining the SSD performance impact data output by the environmental impact model, the server will apply it to the parameters of the original SSD anomaly detection model for correction. For example, the original read speed threshold in the model is 3000MB / s to be judged as abnormal, while the environmental impact model predicts that the read speed will be reduced by 100MB / s in this environment. Then the server needs to correct the original threshold to 2900MB / s to avoid false alarms caused by environmental changes.

[0075] Similarly, it is also necessary to correct the thresholds of parameters such as the write rate and the failure rate according to the output data of the environmental impact model, so that the thresholds used for the final abnormal judgment are set more reasonably, thereby improving the detection effect.

[0076] In this way, the server can dynamically adjust the parameters of its anomaly detection model based on the monitoring of the new SSD environment, and solve the problem of false alarms caused by environmental changes.

[0077] S204. Obtain the service life data of the new SSD;

[0078] The server needs to track and record the service life of the new SSD in real time to consider the impact of the SSD usage time on its performance.

[0079] Specifically, when the new SSD is connected, the server can record information such as its factory date, model, and expected life. During use, the server can statistically collect data such as the working time, read and write times, and wear degree of the SSD in real time. And calculate the proportion of the used life of each SSD to the total life, that is, the service life percentage. For example, 60% of the life has been used. The life data can be continuously collected at a certain period, such as once a day or once a week to obtain the latest usage situation. And record it in the SSD service life database for subsequent model adjustment reference.

[0080] S205. Adjust the abnormal threshold parameter in the SSD data abnormality determination model again according to the service life data.

[0081] As the SSD is used for a longer time, its performance will show a downward trend. Therefore, the server needs to re-adjust the anomaly detection model based on the new service life data. Specifically, a mapping relationship between the service life and the performance degradation can be established in advance. For example, when 50% of the service life is used, the read and write speeds decrease by 10%, etc. When it is obtained that the SSD has been used for 40% of its service life, the data anomaly threshold for its read and write speed parameters can be calculated to be adjusted accordingly. It is also possible to directly train an aging impact model by collecting the performance data of the aging SSD in real time. Input the service life and output the impact value on the performance, which is used to adjust the detection threshold.

[0082] In this way, as the SSD is used for a longer and longer time, the judgment threshold in the anomaly detection model will also be dynamically adjusted, avoiding misjudging its anomaly due to the decrease in SSD performance, thereby improving the detection accuracy and ensuring the security of server data.

[0083] In the embodiments of the present application, by adopting the technical means of obtaining the temperature and humidity data of the working environment of the new SSD, inputting them into the environment impact model to obtain the performance impact value, and then correcting the parameters of the anomaly detection model, the problem of false alarms of SSD data anomalies caused by environmental changes in the prior art is effectively solved, and thus the technical effects of improving the accuracy and adaptability of SSD data anomaly detection are achieved.

[0084] The server in the embodiments of the present invention application will be described from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of the server in the embodiments of the present application.

[0085] It should be noted that Figure 3 The structure of the server shown is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention.

[0086] As Figure 3 shown, the server includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 302 or the program loaded from the storage section 308 into the Random Access Memory (RAM) 303, such as executing the methods described in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other through a bus 304. The Input / Output (I / O) interface 305 is also connected to the bus 304.

[0087] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, a button switch, etc.; an output section 307 including a liquid crystal display (LCD), an audio output device, an indicator light, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. The drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.

[0088] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the present invention are executed.

[0089] It should be noted that specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or component.

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings.

[0091] Specifically, the server in this embodiment includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the SSD data exception fast detection method provided in the above embodiment is implemented.

[0092] On the other hand, the present invention also provides a computer-readable storage medium, which may be included in the server described in the above embodiment; or it may exist separately and not be assembled into the server. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of the server, the server implements the SSD data exception fast detection method provided in the above embodiment.

[0093] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.

[0094] As used in the above embodiments, depending on the context, the term "when..." may be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" may be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".

[0095] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. These processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disk that can store program codes.

Claims

1. A method for rapid detection of SSD data anomalies, applied to a server, characterized in that, The method includes: Obtaining performance data of different types of SSDs as sample data; Training an SSD type adaptation model through a machine learning algorithm according to the sample data, where the SSD type adaptation model is used to determine the type of the current SSD; After obtaining the new performance data of the new SSD, inputting the new performance data into the SSD type adaptation model to determine the SSD type; Adjusting the abnormal threshold parameter corresponding to the SSD data abnormality determination model according to the SSD type, where the SSD data abnormality determination model is a model constructed based on a plurality of normal data and abnormal data of the SSD obtained in advance; Inputting the new performance data into the SSD data abnormality determination model to determine the data abnormality situation of the SSD; After the step of adjusting the abnormal threshold parameter corresponding to the SSD data abnormality determination model according to the SSD type, it further includes: Obtaining the working environment data of the new SSD through a sensor, where the working environment data at least includes the temperature and humidity under the working state of the new SSD; Inputting the working environment data into the environment impact model to obtain SSD performance impact data, where the environment impact model is a model established based on the impact of different temperature and humidity environments on SSD performance data obtained in advance; Correcting the abnormal threshold parameter in the SSD data abnormality determination model according to the SSD performance impact data; After the step of correcting the abnormal threshold parameter in the SSD data abnormality determination model according to the SSD performance impact data, it further includes: Obtaining the service life data of the new SSD; Adjusting the abnormal threshold parameter in the SSD data abnormality determination model again according to the service life data.

2. The method according to claim 1, wherein After the step of inputting the new performance data into the SSD data abnormality determination model to determine the data abnormality situation of the SSD, it further includes: If the data abnormality situation is that the new performance data is abnormal, sending an SSD data abnormality alarm message to the management end; After determining the abnormal new performance data as abnormal data, feeding back the abnormal data to the SSD data abnormality determination model for updating.

3. The method according to claim 2, characterized in that, The step of sending an SSD data abnormality alarm message to the management end if the data abnormality situation is that the new performance data is abnormal specifically includes: Determining the log information when the new performance data sends an abnormality; Determining the association rule between the data abnormality situation and the log information; When the log information meets the association rule corresponding to the data abnormality situation, sending an SSD data abnormality alarm message to the management end in advance.

4. The method according to claim 1, characterized in that, After the step of inputting the new performance data into the SSD data abnormality determination model to determine the data abnormality situation of the SSD, it further includes: When the data abnormality situation is that the new performance data continuously matches the abnormal threshold parameter multiple times, determining that the new SSD has a fault; Obtaining the current operating state data of the server; Input the operation status data into a fault impact model, which determines the impact type and degree of the fault on the server performance and data integrity based on the operation status data; When it is determined that the fault impact type meets the set conditions, mark the physical location of the new SSD; Determine the isolation method of the new SSD according to the physical location.

5. The method according to claim 1, wherein After the step of inputting the new performance data into the SSD data anomaly determination model to determine the data anomaly situation of the SSD, it further includes: Send a check prompt message to the management end according to the data anomaly situation; After receiving the check result feedback from the management end, determine the abnormal cause leading to the new performance data according to the check result; Determine an optimization plan according to the abnormal cause, and the optimization plan is used to eliminate the factors causing the new performance data to be abnormal, so that the new SSD returns to the normal working state.

6. A server, characterized in that, The server includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the server to execute the method according to any one of claims 1-5.

7. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the server, it causes the server to execute the method according to any one of claims 1-5.

8. A computer program product, characterized in that, When the computer program product runs on the server, it causes the server to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Abnormal power failure test method and device during wear leveling of solid state disk

    CN113094222A

  • Hard disk fault prediction method and device, electronic equipment and storage medium

    CN115904916A