Solid state disk state prediction model construction method

By building a solid-state drive state prediction model combining hard disk physical characteristics and multi-dimensional characteristics, the problem of insufficient prediction accuracy and reliability caused by the existing model failing to fully consider the physical characteristics of the hard disk, achieving higher fault prediction accuracy and reliability.

CN120179484AActive Publication Date: 2025-06-20CIVIL AVIATION UNIV OF CHINA

Patent Information

Application Number
CN202510671857.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing solid-state drive failure prediction model fails to fully consider the physical characteristics of the hard disk, resulting in insufficient prediction accuracy and reliability.

Method used

By obtaining the initial feature data set including time series data, image data and basic information, generating adversarial sample data, and adding it to the training set, a solid-state drive state prediction model that can be trained using hard disk physical characteristics and multi-dimensional features.

Benefits of technology

Improve the accuracy and reliability of SSD failure prediction, and enhance the context understanding ability of the model by combining the physical characteristics of the hard disk and multi-dimensional feature data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179484A_ABST
    Figure CN120179484A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of solid state disk prediction, in particular to a solid state disk state prediction model construction method, which is characterized in that real sample data and simulated fault sample data are utilized for training, and the sample data comprise log data, image data and basic information of a solid state disk. The basic information is combined with the time sequence data, more context information can be provided for the model, and therefore the accuracy and reliability of fault prediction of the solid state disk can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of solid state drive prediction, and particularly to a method for constructing a solid state drive state prediction model. Background Art

[0002] With the rapid development of cloud storage technology and data centers, solid state drives (SSDs) have gradually replaced traditional hard disk drives (HDDs) and become a core component of modern storage systems. SSDs are widely used in cloud storage, big data processing, and high-performance computing due to their high performance, low power consumption, anti-vibration, and fast response time. However, although SSDs have significant performance advantages over HDDs, their failure modes are also somewhat complex, and as the usage time increases, the failure rate of SSDs shows a relatively special growth trend, which makes the failure prediction of SSDs a key issue in modern storage systems. Patent document (CN113778766A) discloses a method for establishing a hard disk failure prediction model based on multi-dimensional features. This document simultaneously uses the SMART information of the hard disk, the firmware version information, and the event log information of the device where the hard disk is located as the feature data for hard disk failure prediction. Among the feature data, the values of each numerical type data item are cumulative values, and the model is trained and tested based on the time series of the feature data. Although it can ensure the training effect of the model to a certain extent based on multi-dimensional feature information, it does not consider the influence of information such as the physical characteristics of the hard disk on the hard disk state, such as the hard disk model, supplier, wear level, usage duration, etc. These physical characteristics are of great significance for determining whether the SSD is in a failure state. For example, the usage years and the number of program / erase (P / E) cycles of the hard disk are important factors determining the health status of the SSD. Especially in an environment with high load operation, the failure risk of the hard disk may increase with the change of these physical characteristics. In addition, in this document, the positive samples are sampled in a per-time-period sampling manner, and the number of negative samples is expanded, but this method usually brings problems such as overfitting, data loss, or important features being ignored. Therefore, there is a need to provide a failure prediction solution that can improve the accuracy and reliability of solid state drive failure prediction. Summary of the Invention

[0003] For the above technical problems, the technical solution adopted by the present invention is as follows: According to a first aspect of the present invention, there is provided a method for constructing a solid state drive state prediction model, the method comprising the following steps: S100, obtain the initial feature dataset; wherein, each initial feature dataset includes the status label and feature data of the corresponding solid-state drive within a set time period, the feature data includes time series data, image data, and basic information; the time series data is a time series formed by the numerical values of the key parameters in the SMART log data, the image data is the image data obtained based on the time series data, the basic information is used to characterize the physical characteristics of the solid-state drive, and the status label includes a first label indicating that the solid-state drive is in a normal state and a second label indicating that the solid-state drive is in a faulty state.

[0004] S200, use the feature data with the status label of the first label in the initial feature dataset as the data to be processed, and generate the corresponding adversarial sample data based on the data to be processed.

[0005] S300, add the generated adversarial sample data to the initial feature dataset as the sample dataset, and divide the sample dataset into a training set and a test set.

[0006] S300, use the feature data in the sample dataset as the input information and the status label as the label information, and use the training set to train the initial solid-state drive status prediction model to obtain the trained solid-state drive status prediction model, and use the test set to test the trained solid-state drive status prediction model.

[0007] The present invention has at least the following beneficial effects: The method for constructing a solid-state drive status prediction model provided by the embodiment of the present invention is trained using real sample data and simulated faulty sample data. The sample data includes the log data, image data, and basic information of the solid-state drive. These basic information combined with the time series data can provide more context information for the model, thereby improving the accuracy and reliability of the solid-state drive fault prediction.

[0008] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1Flowchart of the method for constructing a solid - state drive state prediction model provided by an embodiment of the present invention. Detailed implementation manners

[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific implementation manners, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0013] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but there can also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, sub - program, etc.

[0014] An embodiment of the present invention provides a method for constructing a solid - state drive state prediction model, as Figure 1 shown, the method may include the following steps: S100, obtain an initial feature data set.

[0015] In an embodiment of the present invention, each initial feature data set includes the state label and feature data of the corresponding solid - state drive within a set time period, that is, an initial feature data set is the data of a solid - state drive within a set time period. Among them, the feature data includes time - series data, image data, and basic information. Among them, the time - series data is a time series formed by the numerical values of key parameters in the SMART log data.

[0016] In an embodiment of the present invention, the time - series data of each solid - state drive can be the data obtained by smoothing the SMART log data monitored within a set time period. Those skilled in the art know that any data obtained by smoothing the SMART log data monitored within a set time period to obtain the corresponding time - series data falls within the protection scope of the present invention.

[0017] In the embodiments of the present invention, the duration of the set time period can be set according to actual needs, such as several months. The time series data is the data obtained after normalizing the original data. The information recorded in the SMART log data includes: disk_id, date, model, label, smart_n. Among them, disk_id is the unique identification number / serial number of the hard disk, date is the generation date of the log data, model is the model of the hard disk, label is a tag indicating the status of the hard disk, and its value is 0 or 1. 0 indicates that the hard disk is operating healthily, i.e., the normal state, and 1 indicates that the hard disk has a fault. n in smart_n is a numerical value used to represent different SMART metrics. In a schematic embodiment, the SMART metrics can be as shown in Table 1 below: Table 1

[0018] In the embodiments of the present invention, the image data is the image data obtained based on the time series data. Existing methods for converting time series into images can be used to obtain the image data. For example, the Gramian Angular Field method can be used to convert the time series data in the SMART log data into an image representation. The Gramian Angular Field method maps the numerical value of each time point to the pixel value of the image by calculating the angular representation of the time series, and can effectively capture the long-term trends and patterns in the time series. This conversion not only preserves the change trend of the time series data, but also can further extract potential fault features through image processing techniques.

[0019] In the embodiments of the present invention, the basic information is used to characterize the physical characteristics of the solid-state drive. The basic information may include the hard disk model, supplier, wear level, and power-on time. Traditional methods mostly rely on the numerical data in the SMART log and ignore other key information of the hard disk (such as the hard disk model, supplier, usage duration, etc.), and these factors are also crucial for the fault prediction of SSDs. The present invention converts the basic information of the hard disk (such as the hard disk model, supplier, wear level, power-on time, etc.) into text features. After normalizing these metadata, they are input into the model to ensure that these key information are effectively integrated into the prediction model. The texturized basic information of the hard disk combined with other numerical features can provide more context information for the model.

[0020] In the embodiments of the present invention, the status label includes a first label indicating that the solid-state drive is in a normal state and a second label indicating that the solid-state drive is in a fault state. The first label can be represented by 0, and the second label can be represented by 1.

[0021] Furthermore, in the embodiments of the present invention, the key parameters can be obtained through the following steps: S1. Based on the original dataset, which includes status labels and SMART log data of the solid-state drive within a set time period, use N similarity calculation methods respectively to obtain the similarities between Z parameters in the SMART log data and the status of the solid-state drive, resulting in Z similarity sets; N≥2.

[0022] In an embodiment of the present invention, the similarity calculation methods may include, for example, Pearson correlation coefficient calculation method, Spearman correlation coefficient calculation method, J-Index coefficient calculation method, and XGBoost model. The XGBoost model provides feature importance calculation based on decision trees. First, use the original dataset to train the XGBoost model for the training set, and then extract the gain value of each parameter as the similarity with the status of the solid-state drive.

[0023] S2. Obtain the first weight Wu of the u-th parameter among the Z parameters u =∑ N v=1 P uv , resulting in Z first weights; P uv is the position of the u-th parameter in the v-th similarity set. For example, if the position of the u-th parameter in the v-th similarity set is 3, then P uv =3. The value range of u is from 1 to Z, and the value range of v is from 1 to N.

[0024] S3. Sort the Z first weights in descending order, and use the n1 parameters corresponding to the first n1 weights among the sorted Z first weights as candidate parameters, where n1<Z. The candidate parameters can be set according to actual needs and can be obtained through the following steps: Step 1. Use the original data corresponding to the current intermediate parameter set to train the current AI model to obtain the status prediction result of the current solid-state drive; the initial value of the current intermediate parameter set is the first parameter among the Z intermediate parameters, the Z intermediate parameters are the Z parameters corresponding to the sorted Z first weights, and the current AI model is the initialized AI model; Step 2. Based on the status prediction result of the current solid-state drive and the corresponding true result, obtain the prediction accuracy of the current AI model and add it to the current accuracy set. If the difference between any two prediction accuracies among the last Q prediction accuracies in the current accuracy set is less than the set difference, use the parameters of the current intermediate parameter set as candidate parameters; otherwise, add the intermediate parameter after the last parameter in the current intermediate parameter set among the Z intermediate parameters to the current intermediate parameter set and execute Step 1.

[0025] In the embodiments of the present invention, the initial value of the current accuracy rate set is empty. The AI model can be an existing neural network model. The set difference can be an empirical value, for example, a value infinitely close to 0. In one illustrative embodiment, n1 = 20.

[0026] S4. Use the original data corresponding to the n1 candidate parameters as training data to train the random forest model, and obtain the second weights of the n1 candidate parameters.

[0027] S5. Sort the n1 second weights in descending order, and use the n2 candidate parameters corresponding to the first n2 weights among the sorted n1 second weights as the key parameters, where n2 < n1.

[0028] In the embodiments of the present invention, n2 can be an empirical value. In one illustrative embodiment, n2 = 15.

[0029] S200. Use the feature data with the status label of the first label in the initial feature dataset as the data to be processed, and generate corresponding adversarial sample data based on the data to be processed.

[0030] In S200, the adversarial sample data can be generated based on a trained adversarial sample generation model. Among them, the data input into the trained adversarial sample generation model includes the image data and text fusion data in the data to be processed, and the text fusion data is the data obtained by splicing the time series data and text data in the data to be processed.

[0031] In the embodiments of the present invention, the adversarial sample generation model includes a first feature encoding module, a second feature encoding module, a feature merging module, a generator, and a discriminator.

[0032] Among them, the first feature encoding module is used to convert the text data in the received text fusion data into numerical form features, obtain a text fusion data feature matrix, and input it to the feature merging module. The second feature encoding module is used to convert the received image data into numerical form features, obtain an image data feature matrix, and input it to the feature merging module. In one illustrative embodiment, the second feature encoding module can be a Gaussian mixture model.

[0033] The feature merging module is used to multiply the received text fusion data feature matrix and image data feature matrix to obtain a merged matrix, and input it to the generator.

[0034] The generator is used to add random noise to the received merging matrix, generate adversarial sample data based on the merging matrix with added random noise, and input it to the discriminator. In an embodiment of the present invention, the random noise is noise sampled from a standard normal distribution. The random noise vector is gradually converted into a simulated fault image through multi-layer transposed convolution operations. The generated fault image not only has randomness but is also affected by the fault patterns learned during the training process, thus being closer to the distribution of real fault images.

[0035] In an embodiment of the present invention, the generator may include a noise addition module, a fully connected layer, a first deconvolution layer, a second deconvolution layer, and a third deconvolution layer connected in sequence.

[0036] The discriminator is used to discriminate the authenticity of the received adversarial sample data to obtain a corresponding discrimination value. The goal of the discriminator is to identify whether the input data comes from real fault samples and at the same time feedback to the generator to help it continuously improve the authenticity of the generated samples. For real samples, it can judge that they are real data, that is, output 1, and for generated samples, it can judge that they are generated samples, that is, output 0.

[0037] In an embodiment of the present invention, the discriminator includes a first convolution layer, a second convolution layer, a third convolution layer, a fourth convolution layer, and a fully connected layer connected in sequence.

[0038] Further, during the training process of the adversarial sample generation model, when the initial score between the adversarial sample data and the real sample data is greater than the set score threshold, and the initial distance between the adversarial sample data and the real sample data is less than the set distance threshold, the model training ends.

[0039] S300, add the generated adversarial sample data to the initial feature dataset as the sample dataset, and divide the sample dataset into a training set and a test set.

[0040] S300, use the feature data in the sample dataset as input information and the status label as label information, and use the training set to train the initial solid-state drive status prediction model to obtain the trained solid-state drive status prediction model, and use the test set to test the trained solid-state drive status prediction model.

[0041] Those skilled in the art know that any method of using the training set to train the initial solid-state drive status prediction model to obtain the trained solid-state drive status prediction model and using the test set to test the trained solid-state drive status prediction model falls within the protection scope of the present invention.

[0042] In an embodiment of the present invention, a focal loss function can be used as the loss function of the solid-state drive state prediction model. The focal loss is used to focus on difficult-to-classify failure samples, reduce the false alarm rate, and improve the recall rate of failure prediction. In addition, methods such as grid search or Bayesian optimization can be used to tune the hyperparameters of the solid-state drive state prediction model to ensure that the model is trained under optimal configurations, thereby improving the performance of the model.

[0043] In an embodiment of the present invention, the solid-state drive state prediction model includes a time-series data feature encoding module, an image encoder, a text encoder, a feature fusion module, and a prediction output module. Among them, the time-series data feature encoding module, the image encoder, and the text encoder are respectively connected to the feature fusion module, and the feature fusion module is connected to the prediction output module. The time-series data feature encoding module is used to perform feature encoding on time-series data to obtain corresponding data feature vectors as the first-modal feature vectors; the image encoder is used to perform feature encoding on image data to obtain corresponding image feature vectors as the second-modal feature vectors; the text encoder is used to perform feature encoding on basic information to obtain corresponding text feature vectors as the third-modal feature vectors; the feature fusion module is used to fuse the received first-modal feature vectors, second-modal feature vectors, and third-modal feature vectors to obtain a modal fusion feature vector and input it to the prediction output module. The prediction output module is used to perform a linear transformation on the received modal fusion feature vector to obtain a corresponding linear transformation value, and perform a mapping process on the linear transformation value to obtain a state prediction value of the solid-state drive. The state prediction value is used to characterize whether the solid-state drive is likely to fail.

[0044] Further, the feature fusion module is specifically used to perform the following operations: S10, perform a dot product of the received i-th modal feature vector with the first transformation weight matrix, the second transformation weight matrix, and the third transformation weight matrix corresponding to the i-th modal feature vector respectively to obtain the first transformation feature vector, the second transformation feature vector, and the third transformation feature vector corresponding to the i-th modal feature vector; the value of i ranges from 1 to 3; the first transformation weight matrix, the second transformation weight matrix, and the third transformation weight matrix are all learnable weight matrices, that is, they will be updated based on the loss of the solid-state drive state prediction model during the training process of the solid-state drive state prediction model.

[0045] S20, perform a dot product of the first transformation feature vector corresponding to the i-th modal feature vector with the second transformation feature vector of the j-th modal feature vector respectively to obtain the first correlation coefficient vector R1 corresponding to the i-th modal feature vector i = (R1 i1 , R1 i2 , R1 i3 ); R1 irFor R i is the correlation coefficient of the r-th dimension in i , where r ranges from 1 to 3.

[0046] S30. Based on R1 i , obtain the second correlation coefficient vector R2 of the i-th modal feature vector i = (R2 i1 , R2 i2 , R2 i3 ); R2 ir = R1 ir / d r , where d r is the dimension of the second transformed feature vector corresponding to the r-th modal feature vector.

[0047] S40. Based on R2 i , obtain the third correlation coefficient vector R3 of the i-th modal feature vector i = (R3 i1 , R3 i2 , R3 i3 ); R3 ir = exp(R2 ir ) / ∑ 3 h=1 exp(R2 ih ), where R2 ih is the h-th correlation coefficient in R2 i , h ranges from 1 to 3, and exp() is the exponential function.

[0048] S50. Obtain the output feature vector F of the i-th modal feature vector i = (F i1 , F i2 , F i3 ), where F ir is the r-th output feature in F i , and F ir = R3 i1 × F3 1r + R3 i2 × F3 2r + R3 i3 × F3 3r , where F3 sr is the r-th feature in the third transformed feature vector corresponding to the s-th modal feature vector.

[0049] S60. Add the i-th modal feature vector and the corresponding output feature vector to obtain the final feature vector of the i-th modal feature vector, which is used as the final feature vector of the i-th modality; obtain the fusion feature matrix M, the size of M is q × 3, the r-th row data of M is the final feature vector of the i-th modality, and q is the dimension size of the final feature vector of each modality.

[0050] S70, based on M to obtain the modal fusion feature vector FM = (FM1, FM2, ……, FM x , ……, FM q ), where FM x is the x-th feature in FM, x ranges from 1 to q, and FM x is the average value of the x-th column data in M.

[0051] In the embodiment of the present invention, the prediction output module can be a fully connected layer. Specifically, the state prediction value satisfies the following conditions: S = (1 + e -y ) -1 ; S is the state prediction value, e is the natural number, y is the linear conversion value, y = FM × W + b, W is the learnable weight matrix, and b is the bias term.

[0052] The method for constructing the solid-state drive state prediction model provided by the embodiment of the present invention is trained using real sample data and simulated fault sample data. The sample data includes the log data, image data, and basic information of the solid-state drive. These basic information combined with the time series data can provide more context information for the model, thereby improving the accuracy and reliability of the solid-state drive fault prediction.

[0053] Another embodiment of the present invention provides a method for predicting the state of a solid-state drive, the method includes: Obtain the current feature data of the solid-state drive.

[0054] Input the obtained feature data into the solid-state drive state prediction model constructed by the foregoing method to predict the current state of the solid-state drive.

[0055] The embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the embodiment of the present invention.

[0056] The embodiment of the present invention also provides a computer-readable storage medium, storing computer-executable instructions, and the computer instructions are used to execute the method described in the embodiment of the present invention.

[0057] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and no limitation is made herein.

[0058] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for constructing a solid-state drive status prediction model, characterized in that, The method includes the following steps: S100. Obtain an initial feature data set. Each initial feature data set includes a status label and feature data of the corresponding solid-state drive within a set time period. The feature data includes time series data, image data, and basic information. The time series data is a time series formed by the values of key parameters in the SMART log data. The image data is image data obtained based on the time series data. The basic information is used to characterize the physical characteristics of the solid-state drive. The status label includes a first label indicating that the solid-state drive is in a normal state and a second label indicating that the solid-state drive is in a faulty state; S200. Use the feature data with the status label being the first label in the initial feature data set as the data to be processed, and generate corresponding adversarial sample data based on the data to be processed; S300. Add the generated adversarial sample data to the initial feature data set as a sample data set, and divide the sample data set into a training set and a test set; S300. Use the feature data in the sample data set as input information and the status label as label information, and train the initial solid-state drive status prediction model using the training set to obtain a trained solid-state drive status prediction model, and test the trained solid-state drive status prediction model using the test set.

2. The method according to claim 1, characterized in that, The solid-state drive status prediction model includes a time series data feature encoding module, an image encoder, a text encoder, a feature fusion module, and a prediction output module. Among them, the time series data feature encoding module, the image encoder, and the text encoder are respectively connected to the feature fusion module. The feature fusion module is connected to the prediction output module. The time series data feature encoding module is used to perform feature encoding on the time series data to obtain corresponding data feature vectors as the first modality feature vectors. The image encoder is used to perform feature encoding on the image data to obtain corresponding image feature vectors as the second modality feature vectors. The text encoder is used to perform feature encoding on the basic information to obtain corresponding text feature vectors as the third modality feature vectors. The feature fusion module is used to fuse the received first modality feature vectors, second modality feature vectors, and third modality feature vectors to obtain a modality fusion feature vector and input it to the prediction output module. The prediction output module is used to perform a linear transformation on the received modality fusion feature vector to obtain a corresponding linear transformation value, and perform a mapping process on the linear transformation value to obtain a status prediction value of the solid-state drive.

3. The method according to claim 2, characterized in that, The feature fusion module is specifically used to perform the following operations: S10. Take the dot product of the received i-th modality feature vector with the first transformation weight matrix, the second transformation weight matrix, and the third transformation weight matrix corresponding to the i-th modality feature vector respectively to obtain the first transformation feature vector, the second transformation feature vector, and the third transformation feature vector corresponding to the i-th modality feature vector. The value of i ranges from 1 to 3. The first transformation weight matrix, the second transformation weight matrix, and the third transformation weight matrix are all learnable weight matrices; S20. Take the dot product of the first transformed eigenvector corresponding to the $i$-th modal eigenvector and the second transformed eigenvector of the $j$-th modal eigenvector respectively to obtain the first correlation coefficient vector $R1$ corresponding to the $i$-th modal eigenvector. i = (R1 i1 , R1 i2 , R1 i3 ); $R1$ ir is the correlation coefficient between the $i$-th modal eigenvector and the $r$-th modal eigenvector, and the value range of $r$ is from 1 to 3. S30, based on R1 i , obtain the second correlation coefficient vector R2 of the i-th modal feature vector i = (R2 i1 , R2 i2 , R2 i3 ); R2 ir = R1 ir / d r , where d r is the dimension of the second transformed feature vector corresponding to the r-th modal feature vector; S40, based on R2 i , obtain the third correlation coefficient vector R3 of the i-th modal feature vector i = (R3 i1 , R3 i2 , R3 i3 ); R3 ir = exp(R2 ir ) / ∑ 3 h=1 exp(R2 ih ), where R2 ih is the h-th correlation coefficient in R2 i , h ranges from 1 to 3, and exp() is the exponential function; S50, obtain the output feature vector F of the i-th modal feature vector i = (F i1 , F i2 , F i3 ), where F ir is the r-th output feature in F i , F ir = R3 i1 × F3 1r + R3 i2 × F3 2r + R3 i3 × F3 3r , and F3 sr is the r-th feature in the third transformed feature vector corresponding to the s-th modal feature vector; S60. Add the i-th modal feature vector and the corresponding output feature vector to obtain the final feature vector of the i-th modal feature vector, which serves as the final feature vector of the i-th modality; obtain a fused feature matrix M. The size of M is q×3. The r-th row data of M is the final feature vector of the i-th modality, and q is the dimensionality size of the final feature vector of each modality; S70, obtaining the modal fusion feature vector FM = (FM1, FM2, ……, FM x , ……, FM q ), where FM x is the x-th feature in FM, x ranges from 1 to q, and FM x is the average value of the x-th column data in M.

4. The method according to claim 2, characterized in that, The state prediction value satisfies the following conditions: S = (1 + e -y ); S is the state prediction value, e is the natural number, y is the linear conversion value, y = FM × W + b, W is the learnable weight matrix, and b is the bias term. -1 ; S is the state prediction value, e is the natural number, y is the linear conversion value, y = FM × W + b, W is the learnable weight matrix, b is the bias term.

5. The method according to claim 1, characterized in that, In S200, the adversarial sample data is generated based on the trained adversarial sample generation model. Among them, the data input into the trained adversarial sample generation model includes the image data and text fusion data in the data to be processed. The text fusion data is the data obtained by concatenating the time series data and text data in the data to be processed; Among them, the adversarial sample generation model includes a first feature encoding module, a second feature encoding module, a feature merging module, a generator, and a discriminator. The first feature encoding module is used to convert the text data in the received text fusion data into numerical form features to obtain a text fusion data feature matrix and input it to the feature merging module; the second feature encoding module is used to convert the received image data into numerical form features to obtain an image data feature matrix and input it to the feature merging module; the feature merging module is used to multiply the received text fusion data feature matrix and image data feature matrix to obtain a merged matrix and input it to the generator; the generator is used to add random noise to the received merged matrix and generate adversarial sample data based on the merged matrix with random noise added and input it to the discriminator; the discriminator is used to discriminate the authenticity of the received adversarial sample data to obtain the corresponding discrimination value.

6. The method according to claim 5, characterized in that, The generator includes a noise addition module, a fully connected layer, a first transposed convolutional layer, a second transposed convolutional layer, and a third transposed convolutional layer connected in sequence. The discriminator includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fully connected layer connected in sequence; Among them, during the training process of the adversarial sample generation model, when the initial score between the adversarial sample data and the real sample data is greater than the set score threshold, and the initial distance between the adversarial sample data and the real sample data is less than the set distance threshold, the model training ends.

7. The method according to claim 1, characterized in that, The key parameters are obtained through the following steps: S1. Based on the original dataset, the original dataset includes the state labels and SMART log data of the solid-state drive within a set time period. Respectively use N similarity calculation methods to obtain the similarity between Z parameters in the SMART log data and the state of the solid-state drive, and obtain Z similarity sets; N≥2; S2, obtain the first weight W of the u-th parameter among the Z parameters u = ∑ N v=1 P uv , and obtain Z first weights; P uv is the position of the u-th parameter in the v-th similarity set, where u ranges from 1 to Z and v ranges from 1 to N; S3. Sort the Z first weights in descending order, and use the n1 parameters corresponding to the first n1 weights in the sorted Z first weights as candidate parameters, where n1<Z; S4. Use the original data corresponding to the n1 candidate parameters as training data to train the random forest model to obtain the second weights of the n1 candidate parameters; S5. Sort the n1 second weights in descending order, and use the first n2 weights among the sorted n1 second weights to correspond to the n2 candidate parameters as the key parameters, where n2 < n1.

8. The method according to claim 7, characterized in that, The candidate parameters are obtained through the following steps: Step 1: Use the original data corresponding to the current intermediate parameter set to train the current AI model to obtain the state prediction result of the current solid-state drive; the initial value of the current intermediate parameter set is the first parameter among the Z intermediate parameters, the Z intermediate parameters are the Z parameters corresponding to the Z first weights after sorting, and the current AI model is the initialized AI model; Step 2: Based on the state prediction result of the current solid-state drive and the corresponding real result, obtain the prediction accuracy of the current AI model and add it to the current accuracy set. If the difference between any two prediction accuracies among the last Q prediction accuracies in the current accuracy set is less than the set difference, use the parameters of the current intermediate parameter set as candidate parameters; otherwise, add the intermediate parameter after the last parameter in the current intermediate parameter set among the Z intermediate parameters to the current intermediate parameter set, and execute Step 1.

9. The method according to claim 1, characterized in that, The basic information includes the hard disk model, supplier, wear level, and boot time.

Citation Information

Patent Citations

  • Disk fault prediction method, apparatus and device, and storage medium

    CN111782491A

  • Magnetic disk SMART data expansion method and system and electronic equipment

    CN115543762A

  • Enterprise-level solid state disk fault early warning method based on multi-instance adversarial learning

    CN115658401A

  • Disk failure prediction method and device, electronic equipment and storage medium

    CN116680602A

  • Equipment fault prediction method and device, equipment and computer readable storage medium

    CN118606826A

Cited By

  • Solid state disk state information prediction model training method, prediction method and device

    CN120821647A