Solid-state storage hard disk data recovery method and system
Through differential analysis and metadata repair, backtracking reasoning and data engraving algorithms, data loss of solid-state storage hard disks is solved, and data loss location and recovery problems are achieved, high-precision data correction and secure migration are achieved, and the accuracy of data recovery and system fault tolerance are improved.
Patent Information
- Application Number
- CN202510815339.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The prior art cannot effectively locate and recover data loss caused by misoperation, abnormal attacks and environmental interference in solid-state storage hard disks, resulting in the risk of data coverage, mapping chain breaks and permanent loss of business data.
Data loss events are identified through differential analysis, metadata repair, backtracking reasoning and data engraving algorithms are used to deal with data loss caused by human error operations, abnormal attacks and environmental interference, and optimize recovery mechanisms through verification and evaluation, combining the mapping regulations framework to achieve high-precision correction and secure migration of data.
Classification and structured repair of lost data is realized, targeted, accurate and automated data recovery is improved, data integrity guarantee and fault tolerance capabilities are enhanced, and data integrity guarantee and storage system are ensured to ensure security and consistency after data migration.
Smart Images

Figure CN120371611A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data processing module, and more particularly, to a method and system for recovering data of a solid-state storage hard disk. Background Art
[0002] Solid-state storage hard disk (SSD) data refers to all information stored inside the SSD with flash chips as the medium, mainly including user-level data, system-level data, file system metadata, flash translation layer mapping tables maintained by the controller, cache data, and log information, etc. These data jointly support the read and write operations of the SSD under high performance and high reliability conditions.
[0003] Recovering data of a solid-state storage hard disk is to ensure data integrity, business continuity, and system stability. Due to the use of a flash translation layer and a garbage collection mechanism inside the SSD, data has the characteristics of non-sequential physical storage, dynamically changing mapping, and being vulnerable to anomalies. It is easy to cause data loss, mapping table damage, or page data being unable to be located due to various factors such as misoperation, abnormal attacks, and environmental interference. If different recovery methods cannot be implemented for different loss reasons, it will lead to a single recovery strategy, unable to accurately locate the position and structural relationship of the lost data, resulting in serious consequences such as data overwriting, mapping chain breakage, unverifiable recovered data, or permanent loss of business data, thus threatening data security and availability.
[0004] In view of the problems in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] In view of the problems in the related art, the present invention provides a method and system for recovering data of a solid-state storage hard disk to overcome the above technical problems existing in the prior related art.
[0006] For this purpose, the specific technical solutions adopted by the present invention are as follows: According to one aspect of the present invention, a method for recovering data of a solid-state storage hard disk is provided, and the method includes: Performing a difference analysis on the historical log state characteristics and the health state characteristics stored in the solid-state hard disk, and identifying a data loss event of the solid-state hard disk based on the result of the difference analysis, and the data loss event includes a human misoperation event, an abnormal attack event, and an environmental interference event; Constructing a data recovery mechanism based on the data loss abnormal event, and performing a recovery process on the lost data of the solid-state hard disk based on the data recovery mechanism; Performing a verification evaluation on the recovered data, and optimizing the data recovery mechanism according to the result of the verification evaluation.
[0007] Preferably, constructing a data recovery mechanism based on the data loss abnormal event, and performing a recovery process on the lost data of the solid-state hard disk based on the data recovery mechanism includes: Based on human error operation events, use metadata repair technology to repair the lost data caused by human error operations; Based on abnormal attack events, use the backtracking inference algorithm to inversely infer the attack source, and correct the lost data caused by abnormal attacks according to the results of the inverse inference; Based on environmental interference events, use the data carving algorithm to reconstruct the flash translation layer mapping table, and reconstruct the lost data caused by environmental interference through the flash translation layer mapping table.
[0008] Preferably, based on abnormal attack events, using the backtracking inference algorithm to inversely infer the attack source, and correcting the lost data caused by abnormal attacks according to the results of the inverse inference includes: Collect the abnormal access nodes of the solid-state drive stored data when the abnormal attack event is triggered, and construct a monitoring path matrix representing the correlation strength of the access path according to the mapping relationship between the abnormal access nodes and the original data blocks; Set the initial parameters of the backtracking inference algorithm, search for the access node most relevant to the abnormal attack event in the monitoring path matrix based on the predefined correlation coefficient as the candidate attack source, and construct the initial support set of the backtracking inference in combination with the candidate attack source; Iteratively calculate the strength estimation value of each candidate attack source leading to the original data block, re-screen the candidate attack source with the greatest contribution based on the strength estimation value, and update the initial support set to obtain the backtracking inference support set; Use the sparse regression residual function to calculate the residual vector of each candidate attack source in the backtracking inference support set, and compare the residual vector with the preset threshold; If the residual vector is less than the preset threshold, terminate the backtracking inference and select the candidate attack source with the smallest residual vector in the backtracking inference support set as the final attack source, otherwise, continue the iterative calculation; Calculate the sparse coefficient vector of the final attack source, substitute it into the constructed data correction function to generate an approximate reconstruction result of the original data block of the solid-state drive, and encrypt and migrate the approximate reconstruction result to the secure data block.
[0009] Preferably, iteratively calculate the strength estimation value of each candidate attack source leading to the original data block, re-screen the candidate attack source with the greatest contribution based on the strength estimation value, and update the initial support set to obtain the backtracking inference support set includes: Use the regularization function to process the feature data of the candidate attack sources in the initial support set to obtain the regularization feature vector of each attack source; Sort the regularization feature vectors in descending order, select the attack source indexes corresponding to the first several regularization feature vectors to construct a new support set, and use the new support set as the preliminary backtracking support set for the current iteration round; Extract the corresponding columns according to the attack source index of the preliminary backtracking support set and construct a sub-matrix, and use the least squares method to fit the pre-sub-matrix to output the intensity estimation value vector; Sort the intensity estimation value vectors, select the attack source index corresponding to the intensity estimation value vector with the largest component among the first several ones to update the preliminary backtracking support set, and obtain the backtracking inference support set.
[0010] Preferably, encrypting and migrating the approximate reconstruction result to the secure data block includes: Slice the corrected approximate reconstruction result according to a preset size to generate a number of logical data blocks, and each logical data block is represented in the form of a key-value pair with the block number and the block content; Use the key-value pairs of the logical data blocks as input and distribute them to the mapping nodes in the map-reduce framework. Each mapping node independently encrypts the assigned block content and retains the block number; After each mapping node completes encryption, it outputs the encrypted key-value pairs, and distributes the encrypted key-value pairs as input to the reduce nodes in the map-reduce framework. The reduce nodes sort the block numbers after receiving the encrypted key-value pairs to restore the order of the original data blocks; Integrate the block content corresponding to the encrypted key according to the order of the original data blocks to obtain the encrypted data block, and replace the original data block with the encrypted data block and migrate it to the secure data block storage area.
[0011] Preferably, based on the environmental interference event, use the data carving algorithm to reconstruct the flash translation layer mapping table, and reconstruct the lost data caused by the environmental interference through the flash translation layer mapping table, including: Collect the physical block interference area of the solid-state drive when the environmental interference event is triggered, and extract the effective mapping table fragment from the undamaged physical block interference area; Use the data carving algorithm to extract the context features of the effective mapping table fragment as input, train a conditional generative adversarial network, construct the mapping relationship from the logical data block to the physical block through the generator, and evaluate the consistency of the mapping relationship through the discriminator; Train the conditional generative adversarial network using the Wasserstein loss and gradient penalty, and output the flash translation layer mapping table through the conditional generative adversarial network; Use the reconstructed flash translation layer mapping table to locate the original logical page data features, and reconstruct the missing page content through the context data blocks of the original logical page data features.
[0012] Preferably, using the data carving algorithm to extract the context features of the effective mapping table fragment includes: Extract the original logical pages from the effective mapping table fragment, construct a sparse matrix for the mapping relationship from the logical data block to the physical block, and construct an enhanced matrix factorization model based on the context features of each sparse matrix; Fuse the context features of the sparse matrix to form a context feature vector, and use an enhanced matrix decomposition model to score the confidence of the context feature vector to screen out high-confidence context features.
[0013] Preferably, using the enhanced matrix decomposition model to score the confidence of the context feature vector and screen out high-confidence context features includes: Construct a scoring candidate segment from the context feature vector in the original logical page, and use the enhanced matrix decomposition model to calculate the local confidence score and global importance score corresponding to the random context feature vector in the context feature vector of the scoring candidate segment respectively; Take the context feature vector with the local confidence score higher than the global importance score as local noise deviating from the global pattern, and remove it from the scoring candidate segment; Take the remaining scoring candidate segments as high-confidence scoring candidate segments, and use the high-confidence scoring candidate segments as the enhanced matrix decomposition model to screen out high-confidence context features through the enhanced matrix decomposition model.
[0014] Preferably, the expression of the Wasserstein loss is: ; In the formula, W represents the Wasserstein distance; E represents taking the average over the entire data distribution; represents the sample from the data distribution generated by the generator; represents the sample of the true data distribution; represents the output value of the discriminator; represents the generated sample score; represents the true sample score.
[0015] According to another embodiment of the present invention, there is also provided a solid-state storage hard disk data recovery system, which includes: A data loss analysis module, which is used to perform a difference analysis on the historical log state features and health state features stored in the solid-state hard disk, and identify data loss events of the solid-state hard disk based on the difference analysis results, and the data loss events include human misoperation events, abnormal attack events, and environmental interference events; A data recovery processing module, which is used to construct a data recovery mechanism based on the data loss abnormal event, and perform a recovery process on the lost data of the solid-state hard disk based on the data recovery mechanism; A data verification and evaluation module, which is used to verify and evaluate the recovered data, and optimize the data recovery mechanism according to the verification and evaluation results.
[0016] The beneficial effects of the present invention are as follows: 1. By respectively targeting three typical data loss scenarios of human error operation, abnormal attack, and environmental interference, and combining differential technical means such as metadata repair, backtracking reasoning, and data carving, the present invention realizes the classification and positioning, cause identification, and structured repair of lost data, effectively improving the pertinence, accuracy, and automation of data recovery. It not only enhances the adaptability of the recovery process to multiple types of failure mechanisms but also significantly improves the data integrity guarantee level and the fault tolerance ability of the storage system.
[0017] 2. The present invention gradually and accurately locates the attack source through the backtracking reasoning algorithm. On the basis of constructing a monitoring path matrix, screening the initial support set, iteratively calculating the intensity estimation value, introducing regularization and the least squares method to optimize the support set update, and combining the sparse regression residual to judge the optimal attack source, it realizes the high-precision correction of tampered or lost data. At the same time, after encrypting the reconstruction result, it completes the sorting, integration, and secure migration of logical data blocks through the mapping reduction framework, which not only improves the accuracy of the recovery process but also enhances the security and consistency after data migration, and has good practicability and security guarantee ability.
[0018] 3. The present invention extracts the effective mapping table fragments and their context features in the undamaged physical blocks through the data carving algorithm, constructs a sparse matrix using the context features and fuses them to form a context feature vector, further introduces an enhanced matrix decomposition model for confidence scoring to screen out high-confidence features, effectively avoiding local noise interference. On the basis of the generator constructing the mapping relationship between logical data and physical blocks and the discriminator evaluating its consistency, it realizes the highly stable reconstruction of the mapping table by combining the Wasserstein loss and gradient penalty, and finally realizes the accurate positioning of the logical page data features and the context reconstruction of the missing page content, significantly improving the recovery accuracy, fault tolerance ability, and model generalization ability in the scenarios of physical damage or page table loss. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 is a flowchart of a method for data recovery of a solid-state storage hard disk according to an embodiment of the present invention; Figure 2 is a principle block diagram of a data recovery system for a solid-state storage hard disk according to an embodiment of the present invention; Figure 3It is a schematic flowchart of correcting lost data caused by abnormal attacks in a solid - state storage hard - disk data recovery method according to an embodiment of the present invention.
[0021] In the figure: 1. Data loss analysis module; 2. Data recovery processing module; 3. Data verification and evaluation module. Specific implementation manners
[0022] To further illustrate the embodiments, the present invention provides accompanying drawings. These drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be combined with the relevant descriptions in the specification to explain the operation principle of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.
[0023] According to an embodiment of the present invention, a solid - state storage hard - disk data recovery method and system are provided.
[0024] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners. As Figure 1 shown, the solid - state storage hard - disk data recovery method according to an embodiment of the present invention includes: S1. Perform a difference analysis on the historical log status characteristics and the health status characteristics stored in the solid - state drive, and identify data loss events of the solid - state drive based on the difference analysis results. The data loss events include human - error operation events, abnormal attack events, and environmental interference events.
[0025] It should be noted that the historical log status characteristics and the current health status characteristics. The historical log status characteristics include time - series characteristics such as logical block access frequency, write / erase times, system IO logs, operation instruction sequences, historical temperature, voltage, current, etc.; the health status characteristics mainly come from current internal status monitoring indicators, such as the number of bad block pages, the number of ECC error corrections failed, the current temperature, the number of power outages, the consistency of the mapping table, etc.
[0026] Perform difference analysis on the historical log status characteristics and the current health status characteristics through time - series alignment, window aggregation, and structured processing, calculate indicators such as the change rate, difference value, and cumulative deviation value of each characteristic, form a difference feature vector, and use an unsupervised anomaly detection algorithm to model the difference features to identify data points that deviate from the normal mode.
[0027] Among them, in the human - error operation event, it is manifested as block erasure, accidental deletion, etc. after a legal IO operation; the attack behavior is often manifested as illegal instruction access, frequent remapping, a sharp increase in ECC errors, and a mutation in access entropy; environmental interference is often accompanied by temperature, voltage, power fluctuations, and a wide range of abnormal health parameters.
[0028] S2. Build a data recovery mechanism based on the data loss exception event, and recover the lost data of the solid-state drive based on the data recovery mechanism.
[0029] Among them, building a data recovery mechanism based on the data loss exception event and recovering the lost data of the solid-state drive based on the data recovery mechanism includes: Based on the human error operation event, use the metadata repair technology to repair the lost data caused by the human error operation.
[0030] It should be noted that identify the characteristics of the misoperation behavior; extract key metadata information such as the inode table of the file system and the directory structure from the log system or metadata backup within this time window, and restore the version history of the damaged data structure; locate the remaining positions of the deleted data in the storage medium through metadata residue analysis techniques (such as block redundancy comparison, address reference counting, deleted block scanning); establish an index relationship between the restored file metadata structure and the actual data block mapping, mount it to the temporary file system for verification and export, and achieve effective repair and recovery of the misoperation data.
[0031] Based on the abnormal attack event, use the backward reasoning algorithm to reverse-infer the attack source, and correct the lost data caused by the abnormal attack according to the reverse-inference result.
[0032] Among them, as Figure 3 shown, based on the abnormal attack event, using the backward reasoning algorithm to reverse-infer the attack source and correcting the lost data caused by the abnormal attack according to the reverse-inference result includes: Collect the abnormal access nodes of the solid-state drive storage data when the abnormal attack event is triggered, and build a monitoring path matrix representing the association strength of the access path according to the mapping relationship between the abnormal access nodes and the original data blocks.
[0033] It should be noted that accurately capture the diffusion mode and propagation path of the attack behavior in the storage access path, and reveal the potential correlation between the attacker's operation logic and the underlying data structure; secondly, quantify the concentration and persistence of the abnormal access through the strength information in the path matrix, so as to assist in judging whether the attack has systematicness and destructiveness, and improve the recognition ability of complex behaviors such as multi-hop attacks, persistent access, and camouflage operations.
[0034] Set the initial parameters of the backward reasoning algorithm, search for the access node most relevant to the abnormal attack event in the monitoring path matrix based on the predefined correlation coefficient as the candidate attack source, and build the initial support set of the backward reasoning in combination with the candidate attack source; Iteratively calculate the strength estimation values of each candidate attack source leading to the original data block, re-screen the candidate attack sources with the greatest contribution based on the strength estimation values, and update the initial support set to obtain the backtracking inference support set.
[0035] Among them, iteratively calculating the strength estimation values of each candidate attack source leading to the original data block, re-screening the candidate attack sources with the greatest contribution based on the strength estimation values, and updating the initial support set to obtain the backtracking inference support set includes: Use a regularization function to process the feature data of the candidate attack sources in the initial support set to obtain the regularized feature vectors of each attack source.
[0036] It should be noted that using a regularization function to process the feature data of the candidate attack sources in the initial support set to obtain the regularized feature vectors of each attack source includes: Extract the original feature data of each candidate attack source from the initial support set to form the original feature vector of the attack source; use min-max normalization to perform normalization preprocessing on different feature dimensions to eliminate the influence of different dimensions on subsequent analysis; introduce a regularization function (such as L1 norm or L2 norm) to perform regularization processing on the preprocessed feature vector; use the cost function minimization method for constraint solving during the regularization optimization process, and finally solve to obtain the regularized feature vectors of each candidate attack source, and its expression is: ; In the formula, R ( x ) represents the selected regularization term, λ represents the regularization strength hyperparameter; J ( x ) represents the objective function; x represents the original attack source feature vector; represents the ideal reference value of the attack source feature vector.
[0037] Sort the regularized feature vectors in descending order, select the attack source indices corresponding to the first several regularized feature vectors to construct a new support set, and use the new support set as the preliminary backtracking support set for the current iteration round; Extract the corresponding columns according to the attack source indices of the preliminary backtracking support set and construct a submatrix, and use the least squares method to fit the pre-submatrix to output the strength estimation value vector.
[0038] It should be noted that according to the attack source indices identified in the preliminary backtracking support set, extract the corresponding columns from the original monitoring path matrix to construct a submatrix. Let the original matrix be A ∈ R m×n , and the submatrix composed of the columns corresponding to the support set indices is As ∈ Rm×k ; Analyze the target observation vector b ∈ R m , where the target observation vector represents the actual collected path access intensity or the degree of influence of abnormal behavior. Based on this, the least squares method is used to solve A s x = b for the approximate solution, where x ∈R k is the weight vector of the attack source intensity to be estimated; the least squares solution is obtained by optimizing the objective function , that is, the analytical solution is obtained through the pseudo-inverse equation ; if the matrix condition is poor, regularization (ridge regression) form can be introduced to enhance the stability of the solution; the finally output vector x is the intensity estimation value of each preliminary backtracking attack source under the current access path.
[0039] Sort the intensity estimation value vector, select the intensity estimation value vector of the largest component in the top several, and update the preliminary backtracking support set with the corresponding attack source index to obtain the backtracking inference support set.
[0040] Use the sparse regression residual function to calculate the residual vector of each candidate attack source in the backtracking inference support set, and compare the residual vector with the preset threshold.
[0041] It should be noted that the sparse regression residual function is a function used to measure the fitting ability of feature vectors under sparse modeling conditions, usually constructed as , where y represents the observation vector; A s represents the feature submatrix composed of candidate attack sources; x represents the sparse coefficient solution obtained through sparse regression. The residual vector r reflects the insufficient part of the explanatory ability of the current candidate attack source combination for abnormal observations. By comparing the residual vector with the preset error threshold, candidate attack sources with poor fitting ability and high reconstruction error can be effectively removed, and key attack sources with high contribution and strong structural fitting in the current abnormal behavior can be screened out. At the same time, the interference of pseudo-correlation or collinear features can be suppressed, and the accuracy and interpretability of attack path tracking can be improved.
[0042] If the residual vector is less than the preset threshold, terminate the backtracking inference and select the candidate attack source with the smallest residual vector in the backtracking inference support set as the final attack source. Otherwise, continue the iterative calculation; Calculate the sparse coefficient vector of the final attack source, substitute it into the constructed data correction function to generate an approximate reconstruction result of the original data block of the solid-state drive, and encrypt and migrate the approximate reconstruction result to the secure data block.
[0043] The encryption and migration of the approximate reconstruction result to the secure data block includes: The corrected approximate reconstruction result is divided into a preset size to generate a number of logical data blocks, each of which is represented by a block number and a block content as a key-value pair; The key-value pairs of the logical data blocks are distributed as input to the mapping nodes in the mapping specification framework. Each mapping node independently encrypts the content of the block assigned to it and retains the block number; After each mapping node completes encryption, it outputs an encrypted key-value pair and distributes the encrypted key-value pair as input to the specification node in the mapping specification framework. After receiving the encrypted key-value pair, the specification node sorts the block numbers to restore the order of the original data blocks. The block contents corresponding to the encryption key values are integrated according to the order of the original data blocks to obtain the encrypted data blocks, and the encrypted data blocks replace the original data blocks and are migrated to the secure data block storage area.
[0044] It should be noted that by dividing the corrected approximate reconstruction results into logical data blocks and introducing distributed processing in the mapping reduction framework, parallel encryption and orderly reconstruction of data blocks are achieved, thereby significantly improving data processing efficiency and system throughput.
[0045] At the same time, each mapping node independently encrypts the content of its assigned data block and only retains the block number, which enhances the isolation and local security of data during the processing process and reduces the overall risk brought by single-node leakage; the specification node sorts according to the block number and reconstructs the original order, effectively ensuring data integrity and structural correctness, and avoiding order disorder introduced by encrypted distribution.
[0046] Based on environmental interference events, a data carving algorithm is used to reconstruct a flash translation layer mapping table, and lost data caused by environmental interference is reconstructed through the flash translation layer mapping table.
[0047] It should be noted that in environmental interference events (such as sudden power outages and sudden failures of solid-state drives), the original mapping information is very easy to be lost or damaged. Reconstructing the flash translation layer mapping table using a data carving algorithm can extract residual information from physical pages that have not been covered or erased in the flash medium, and restore the mapping relationship from logical to physical addresses by analyzing valid data fragments, logical page number identifiers, write sequence information, etc. in the physical blocks. The beneficial effect of this combination is that it can bypass the failed controller cache or metadata storage area and directly reverse restore the core mapping structure of the FTL from the underlying physical data, thereby achieving recovery of critical data in extreme cases where the original mapping table is not available or is severely damaged.
[0048] Among them, based on the environmental interference event, the flash translation layer mapping table is reconstructed using the data carving algorithm. Reconstructing the lost data caused by environmental interference through the flash translation layer mapping table includes: Collect the physical block interference area of the solid-state drive when the environmental interference event is triggered, and extract the effective mapping table fragments from the undamaged physical block interference area; Use the data carving algorithm to extract the context features of the effective mapping table fragments as the input, train a conditional generative adversarial network, construct the mapping relationship from logical data blocks to physical blocks through the generator, and evaluate the consistency of the mapping relationship through the discriminator.
[0049] It should be noted that the data carving algorithm is a technology that directly analyzes the residual original physical data in the storage medium (such as in-page identifiers, redundancy checks, writing order, LBA residual bits, etc.) to reverse-infer data structures, logical relationships, and mapping information in the absence of original metadata or mapping information.
[0050] Use the data carving algorithm to extract the context features of the effective mapping table fragments as the input of the conditional generative adversarial network, so as to train a generator with context awareness ability to automatically complete the missing mapping relationship from logical data blocks to physical blocks. At the same time, the discriminator scores the rationality and consistency of the generated mapping, and then realizes the mapping reconstruction in the case of environmental interference or mapping damage. The effect is to improve the accuracy of mapping reconstruction, maintain logical consistency, and significantly improve the integrity and intelligence level of data recovery.
[0051] Among them, extracting the context features of the effective mapping table fragments using the data carving algorithm includes: Extract the original logical pages from the effective mapping table fragments, construct a sparse matrix for the mapping relationship from logical data blocks to physical blocks, and construct an enhanced matrix factorization model based on the context features of each sparse matrix; Fuse the context features of the sparse matrix to form a context feature vector, and use the enhanced matrix factorization model to score the confidence of the context feature vector, and screen out the high-confidence context features.
[0052] Among them, using the enhanced matrix factorization model to score the confidence of the context feature vector and screen out the high-confidence context features includes: Construct a scoring candidate segment for the context feature vector in the original logical page, and use the enhanced matrix factorization model to calculate the local confidence score and global importance score corresponding to the random context feature vector in the scoring candidate segment in the context respectively; Take the context feature vector with the local confidence score higher than the global importance score as the local noise deviating from the global pattern, and remove it from the scoring candidate segment; Take the remaining scoring candidate segments as high-confidence scoring candidate segments, and use the high-confidence scoring candidate segments as an enhanced matrix factorization model to screen high-confidence context features through the enhanced matrix factorization model.
[0053] Train a conditional generative adversarial network using the Wasserstein loss and gradient penalty, and output a flash translation layer mapping table through the conditional generative adversarial network.
[0054] It should be noted that a training dataset is constructed, where each sample includes a known logical page number, its possible corresponding physical page number (PPN) segment, and context structure features, as the input of the conditional generative adversarial network; in the training phase, the generator generates the mapping relationship from logical data blocks to physical data blocks guided by conditional features, while the discriminator receives real and generated mapping pairs and judges their authenticity and structural consistency; to improve training stability and generation quality, the Wasserstein loss function is introduced: ; In the formula, W represents the Wasserstein distance; E represents taking the average over the entire data distribution; represents the sample from the data distribution generated by the generator; represents the sample of the real data distribution; represents the output value of the discriminator; represents the generated sample score; represents the real sample score.
[0055] Replace the logarithmic loss of the traditional GAN to solve the training instability problem, and add a gradient penalty term: ; In the formula, GP represents the gradient penalty term; λ represents the gradient penalty coefficient; represents at the real sample x and the generated sample the sample linearly interpolated between; represents the gradient of the discriminator output with respect to the generated sample ; The generator and the discriminator are alternately trained, and through multiple rounds of iterative optimization, the generator gradually learns and outputs an FTL mapping table structure that more conforms to the real distribution; after training is completed, the generator is used to directly output the mapping information from logical page numbers to physical page numbers under new input conditions, reconstruct the missing or damaged FTL mapping table, and provide structural support and automatic mapping capabilities for data reconstruction caused by environmental interference.
[0056] Locate the original logical page data features using the reconstructed flash translation layer mapping table, and reconstruct the missing page content through the context data blocks of the original logical page data features.
[0057] It should be noted that by accurately locating the corresponding relationship of the logical page in physical storage through the reconstructed flash translation layer mapping table, the original logical page data features can be restored, and content reconstruction is carried out in combination with its context data blocks, making full use of logical continuity, access patterns, and data adjacency to improve the reconstruction accuracy of missing pages, thereby enhancing the data integrity recovery ability, especially applicable to the scenario of page-level data loss caused by abnormal power failure or media damage.
[0058] S3. Perform verification and evaluation on the restored data, and optimize the data recovery mechanism according to the verification and evaluation results.
[0059] It should be noted that perform verification and evaluation on the restored data to identify errors, missing, or redundant content in the restored data; quantify the verification results into an evaluation index vector, record the data recovery quality of each recovery path or topology node, and identify paths or nodes with lower recovery accuracy through clustering, score ranking, etc.; combine the access logs, mapping reconstruction records, and error distribution characteristics during the recovery process to construct a topological dependency graph of the data recovery process, and mark the recovery credibility of each node; use the feedback information to adjust the key parameters in the topological mechanism to strengthen the information of high-trust paths and suppress the influence of low-quality paths, thereby realizing the dynamic optimization and adaptive adjustment of the data recovery topology structure.
[0060] According to another aspect of the present invention, as Figure 2 shown, there is also provided a solid-state storage hard disk data recovery system, which includes: A data loss analysis module 1, configured to perform a difference analysis on the historical log state characteristics and health state characteristics stored in the solid-state hard disk, and identify the data loss events of the solid-state hard disk based on the difference analysis results, and the data loss events include human misoperation events, abnormal attack events, and environmental interference events; A data recovery processing module 2, configured to construct a data recovery mechanism based on the data loss abnormal event, and perform a recovery process on the lost data of the solid-state hard disk based on the data recovery mechanism; A data verification and evaluation module 3, configured to perform verification and evaluation on the restored data, and optimize the data recovery mechanism according to the verification and evaluation results.
[0061] Wherein, the data loss analysis module 1 is connected to the data recovery processing module 2, and the data recovery processing module 2 is connected to the data verification and evaluation module 3.
[0062] In summary, by means of the above technical solutions of the present invention, the present invention realizes the classification and positioning, cause identification and structured repair of lost data by combining differential technical means such as metadata repair, backtracking reasoning and data carving for three typical data loss scenarios of human error operation, abnormal attack and environmental interference, effectively improving the pertinence, accuracy and automation of data recovery. It not only enhances the adaptability of the recovery process to multiple types of failure mechanisms, but also significantly improves the data integrity guarantee level and the fault tolerance ability of the storage system. The present invention gradually and accurately locates the attack source through the backtracking reasoning algorithm. Based on constructing a monitoring path matrix, screening an initial support set, iteratively calculating the intensity estimation value, introducing regularization and the least squares method to optimize the support set update, and combining the sparse regression residual to judge the optimal attack source, it realizes the high-precision correction of the tampered or lost data. At the same time, after encrypting the reconstruction result, it completes the sorting, integration and secure migration of the logical data blocks through the mapping reduction framework, not only improving the accuracy of the recovery process, but also enhancing the security and consistency after data migration, and having good practicability and security guarantee ability. The present invention extracts the effective mapping table fragments and their context features in the undamaged physical blocks through the data carving algorithm, constructs a sparse matrix using the context features and fuses them to form a context feature vector, further introduces an enhanced matrix decomposition model for confidence scoring to screen out high-confidence features, effectively avoiding local noise interference. Based on the generator constructing the mapping relationship between logical data and physical blocks and the discriminator evaluating its consistency, it realizes the highly stable reconstruction of the mapping table by combining the Wasserstein loss and gradient penalty, and finally realizes the accurate positioning of the logical page data features and the context reconstruction of the missing page content, significantly improving the recovery accuracy, fault tolerance ability and model generalization ability in the scenarios of physical damage or page table loss.
[0063] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for recovering data of a solid-state storage hard disk, characterized in that, The method includes: Performing a difference analysis on the historical log state features and health state features stored in the solid-state drive, and identifying data loss events of the solid-state drive based on the results of the difference analysis, and the data loss events include human error operation events, abnormal attack events, and environmental interference events; Constructing a data recovery mechanism based on the data loss abnormal events, and performing a recovery process on the lost data of the solid-state drive based on the data recovery mechanism; Performing a verification evaluation on the recovered data, and optimizing the data recovery mechanism according to the results of the verification evaluation; The constructing a data recovery mechanism based on the data loss abnormal events and performing a recovery process on the lost data of the solid-state drive based on the data recovery mechanism includes: Based on the abnormal attack event, using the backtracking inference algorithm to inversely infer the attack source, and correcting the lost data caused by the abnormal attack according to the results of the inverse inference; Based on the environmental interference event, using the data carving algorithm to reconstruct the flash translation layer mapping table, and reconstructing the lost data caused by the environmental interference through the flash translation layer mapping table.
2. A method for recovering data of a solid-state storage hard disk according to claim 1, characterized in that, Before the based on the abnormal attack event, using the backtracking inference algorithm to inversely infer the attack source, and correcting the lost data caused by the abnormal attack according to the results of the inverse inference, further includes: Based on the human error operation event, using the metadata repair technology to repair the lost data caused by the human error operation.
3. A method for recovering data of a solid-state storage hard disk according to claim 1, characterized in that, The based on the abnormal attack event, using the backtracking inference algorithm to inversely infer the attack source, and correcting the lost data caused by the abnormal attack according to the results of the inverse inference includes: Collecting the abnormal access nodes of the data stored in the solid-state drive when the abnormal attack event is triggered, and constructing a monitoring path matrix representing the association strength of the access path according to the mapping relationship between the abnormal access nodes and the original data blocks; Setting the initial parameters of the backtracking inference algorithm, searching for the access node most relevant to the abnormal attack event in the monitoring path matrix based on the predefined correlation coefficient as the candidate attack source, and constructing the initial support set of the backtracking inference in combination with the candidate attack source; Iteratively calculating the strength estimation values of each candidate attack source leading to the original data block, re-screening the candidate attack source with the greatest contribution based on the strength estimation values, and updating the initial support set to obtain the backtracking inference support set; Using the sparse regression residual function to calculate the residual vector of each candidate attack source in the backtracking inference support set, and comparing the residual vector with a preset threshold; If the residual vector is less than the preset threshold, then terminate the backtracking inference and select the candidate attack source with the smallest residual vector in the backtracking inference support set as the final attack source, otherwise, continue the iterative calculation; Calculating the sparse coefficient vector of the final attack source, substituting it into the constructed data correction function to generate an approximate reconstruction result of the original data block of the solid-state drive, and encrypting and migrating the approximate reconstruction result to the secure data block.
4. A method for recovering data of a solid - state storage hard disk according to claim 3, wherein, The iteratively calculating the strength estimation values of each candidate attack source leading to the original data block, re-screening the candidate attack source with the greatest contribution based on the strength estimation values, and updating the initial support set to obtain the backtracking inference support set includes: Using the regularization function to process the feature data of the candidate attack sources in the initial support set to obtain the regularization feature vector of each attack source; Sort the regularized feature vectors in descending order, select the attack source indices corresponding to the first several regularized feature vectors to construct a new support set, and use the new support set as the preliminary backtracking support set for the current iteration round; Extract the corresponding columns according to the attack source indices of the preliminary backtracking support set to construct a submatrix, and use the least squares method to fit the pre-submatrix to output the intensity estimation value vector; Sort the intensity estimation value vector, select the attack source indices corresponding to the intensity estimation value vectors with the largest components among the first several to update the preliminary backtracking support set, and obtain the backtracking inference support set.
5. A method for recovering data of a solid-state storage hard disk according to claim 4, characterized in that The encrypting and migrating the approximate reconstruction result to the secure data block includes: Slice the corrected approximate reconstruction result according to a preset size to generate several logical data blocks, and each logical data block is represented in the form of a key-value pair with the block number and the block content; Use the key-value pairs of the logical data blocks as inputs and distribute them to the mapping nodes in the map-reduce framework. Each mapping node independently encrypts the block content assigned to it and retains the block number; After each mapping node completes the encryption, it outputs the encrypted key-value pairs, and uses the encrypted key-value pairs as inputs to distribute them to the reduce nodes in the map-reduce framework. The reduce nodes sort the block numbers after receiving the encrypted key-value pairs to restore the order of the original data blocks; Integrate the block contents corresponding to the encrypted keys according to the order of the original data blocks to obtain the encrypted data block, and replace the original data block with the encrypted data block and migrate it to the secure data block storage area.
6. The method for recovering data of a solid-state storage hard disk according to claim 1, wherein The reconstructing the flash translation layer mapping table by using the data carving algorithm based on the environmental interference event and reconstructing the lost data caused by the environmental interference through the flash translation layer mapping table includes: Collect the physical block interference areas of the solid-state drive when the environmental interference event is triggered, and extract the effective mapping table fragments from the undamaged physical block interference areas; Use the data carving algorithm to extract the context features of the effective mapping table fragments as inputs, train a conditional generative adversarial network, construct the mapping relationship from the logical data block to the physical block through the generator, and evaluate the consistency of the mapping relationship through the discriminator; Train the conditional generative adversarial network by using the Wasserstein loss and gradient penalty, and output the flash translation layer mapping table through the conditional generative adversarial network; Use the reconstructed flash translation layer mapping table to locate the original logical page data features, and reconstruct the missing page content through the context data blocks of the original logical page data features.
7. A method for recovering data of a solid-state storage hard disk according to claim 6, characterized in that, The extracting the context features of the effective mapping table fragments by using the data carving algorithm includes: Extract the original logical pages from the effective mapping table fragments, construct a sparse matrix for the mapping relationship from the logical data block to the physical block, and construct an enhanced matrix factorization model based on the context features of each sparse matrix; Fuse the context features of the sparse matrix to form a context feature vector, and use the enhanced matrix factorization model to score the confidence of the context feature vector and screen the high-confidence context features.
8. A method for recovering data of a solid-state storage hard disk according to claim 7, characterized in that, The scoring the confidence of the context feature vector by using the enhanced matrix factorization model and screening the high-confidence context features includes: Construct scoring candidate segments from the context feature vectors in the original logical page, and use the enhanced matrix factorization model to calculate the local confidence scores and global importance scores corresponding to the random context feature vectors in the scoring candidate segments in the context respectively; Regard the context feature vectors with local confidence scores higher than the global importance scores as local noises deviating from the global pattern, and remove them from the scoring candidate segments; Regard the remaining scoring candidate segments as high-confidence scoring candidate segments, and use the high-confidence scoring candidate segments as the enhanced matrix factorization model to screen high-confidence context features through the enhanced matrix factorization model.
9. A method for recovering data of a solid-state storage hard disk according to claim 6, characterized in that, The expression of the Wasserstein loss is as follows: ; In the formula, W represents the Wasserstein distance; E represents taking the average over the entire data distribution; represents the sample from the data distribution generated by the generator; represents the sample of the true data distribution; represents the output value of the discriminator; represents the generated sample score; represents the true sample score.
10. A solid-state storage hard disk data recovery system, characterized in that, For implementing the solid-state storage hard disk data recovery method described in any one of claims 1-9, the system includes: A data loss analysis module, configured to perform a difference analysis on the historical log status features and health status features stored in the solid-state hard disk, identify the data loss events of the solid-state hard disk based on the difference analysis results, and the data loss events include human misoperation events, abnormal attack events, and environmental interference events; A data recovery processing module, configured to construct a data recovery mechanism based on the data loss abnormal event, and perform a recovery process on the lost data of the solid-state hard disk based on the data recovery mechanism; A data verification and evaluation module, configured to perform verification and evaluation on the recovered data, and optimize the data recovery mechanism according to the verification and evaluation results.
Citation Information
Patent Citations
Failure hard disk recovery method and system
CN115904820A
Method and device for testing data storage security of solid state disk
CN119416278A
Cited By
Industrial equipment data recovery method, system, equipment, medium and product
CN120653641A
Computer hard disk data recovery method
CN121029491A
Error correction storage method and system based on polynomial interpolation mixed verification
CN121807239A
An error correction storage method and system based on polynomial interpolation hybrid check
CN121807239B
Intelligent data fragment recombination and repair method for file damage recovery
CN122309215A