Data recovery method and device, electronic equipment and program product

By directly reading the original data and using a pre-trained network model to reconstruct the mapping table between logical addresses and physical block addresses when an SSD fails, the problem of lost mapping relationships during data recovery is solved, achieving efficient data recovery.

CN121996475APending Publication Date: 2026-05-08SLICONGO MICROELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SLICONGO MICROELECTRONICS INC
Filing Date
2026-01-08
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In the event of controller failure, firmware corruption, or physical damage to a solid-state drive (SSD), existing technologies struggle to recover the data because the dynamic mapping table is stored in volatile memory or a vulnerable firmware area, resulting in the loss of the mapping relationship between physical block addresses and logical addresses.

Method used

The raw data is read directly through the target data interface. A pre-trained network model is used to reconstruct the mapping table between logical addresses and physical block addresses. Based on this mapping table, the target data is determined and stored in a fault-free disk.

Benefits of technology

It improves the data recovery integrity rate and enables efficient data recovery in the event of SSD failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996475A_ABST
    Figure CN121996475A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and provides a data recovery method and device, electronic equipment and a program product. According to the embodiment of the invention, in response to a received data recovery instruction, a first target disk corresponding to the data recovery instruction is determined, a target data interface is called, original data is acquired from the first target disk, and user data and metadata are processed based on a pre-trained network model. And obtaining a target mapping relation table of the logic address and the physical block address, determining target data based on the target mapping relation table, and storing the target data to a second target disk. The original data in the first target disk is directly read through the target data interface under the condition that the first target disk fails, and the mapping relation table of the logic address and the physical block address is reconstructed based on the pre-trained network model, so that the recovered target data is obtained and stored, and the integrity rate of data recovery is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and in particular relates to a data recovery method, apparatus, electronic device and program product. Background Technology

[0002] Solid State Drives (SSDs) are hard drive devices that store data using NAND flash memory technology, offering advantages such as high-speed read / write speeds and low power consumption. In cases of controller failure, firmware corruption, or physical damage to an SSD, data recovery is necessary to transfer the stored data to a known working storage device to prevent data loss.

[0003] Data recovery methods typically rely on the SSD controller's interface to directly access data from the SSD. However, because dynamic mapping tables are usually stored in the SSD's volatile memory or easily corrupted firmware, the retrieved data often suffers from a loss of the mapping between physical block addresses and logical addresses, making data recovery difficult. Summary of the Invention

[0004] In view of this, embodiments of this application provide a data recovery method, apparatus, electronic device, and program product that can directly read the original data in the disk through the target data interface and reconstruct the mapping relationship table between logical address and physical block address based on a pre-trained network model, thereby obtaining the recovered target data and storing it, thus improving the data recovery integrity rate.

[0005] A first aspect of this application provides a data recovery method, including: In response to receiving a data recovery instruction, a first target disk corresponding to the data recovery instruction is determined; wherein, the first target disk includes a plurality of physical blocks; The target data interface is invoked to obtain raw data from the first target disk; wherein, the raw data includes at least one of user data, metadata, and redundancy check data in each of the physical blocks; The user data and metadata are processed based on a pre-trained network model to obtain a target mapping table; wherein, the target mapping table includes the mapping relationship between the logical address and the physical block address of the user data; Based on the target mapping table, the target data is determined; The target data is stored in the second target disk.

[0006] A second aspect of this application provides a data recovery apparatus, comprising: A first determining module is configured to, in response to receiving a data recovery instruction, determine a first target disk corresponding to the data recovery instruction; wherein the first target disk includes a plurality of physical blocks; The data acquisition module is used to call the target data interface to acquire raw data from the first target disk; wherein, the raw data includes at least one of user data, metadata and redundancy check data in each of the physical blocks; The processing module is used to process the user data and the metadata based on a pre-trained network model to obtain a target mapping table; wherein, the target mapping table includes the mapping relationship between the logical address and the physical block address of the user data; The second determining module is used to determine target data based on the target mapping relationship table; A storage module is used to store the target data to a second target disk.

[0007] A third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the data recovery method provided in the first aspect.

[0008] A fourth aspect of this application provides a computer program product, characterized in that, when the computer program product is run on an electronic device, it causes the electronic device to execute the steps of the data recovery method provided in the first aspect.

[0009] The data recovery method provided in the first aspect of this application, in response to receiving a data recovery instruction, determines a first target disk corresponding to the data recovery instruction, calls a target data interface to obtain raw data from the first target disk, processes user data and metadata based on a pre-trained network model to obtain a target mapping table of logical addresses and physical block addresses, determines target data based on the target mapping table, and stores the target data in a second target disk. This method enables the direct reading of raw data from the first target disk through the target data interface in the event of a first target disk failure, and the reconstruction of the mapping table of logical addresses and physical block addresses based on a pre-trained network model, thereby obtaining and storing the recovered target data, thus improving the data recovery integrity rate.

[0010] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is one of the flowcharts illustrating the data recovery method provided in the embodiments of this application; Figure 2 This is a second schematic flowchart of the data recovery method provided in the embodiments of this application; Figure 3 This is the third flowchart illustrating the data recovery method provided in the embodiments of this application; Figure 4 This is the fourth flowchart illustrating the data recovery method provided in the embodiments of this application; Figure 5 This is the fifth flowchart illustrating the data recovery method provided in the embodiments of this application; Figure 6 This is the sixth flowchart illustrating the data recovery method provided in the embodiments of this application; Figure 7 This is the seventh flowchart illustrating the data recovery method provided in the embodiments of this application; Figure 8 This is a schematic diagram of the data recovery device provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0014] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0015] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0016] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0017] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0018] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0019] Solid-state drives (SSDs) are hard drive devices that store data using NAND flash memory technology, offering advantages such as high-speed read / write speeds and low power consumption. In the event of controller failure, firmware corruption, or physical damage to an SSD, it is necessary to recover the data stored on the SSD and transfer it to a known working storage device to prevent data loss.

[0020] Data recovery methods typically rely on the SSD controller's interface to directly retrieve data from the SSD. However, because dynamic mapping tables are usually stored in the SSD's volatile memory or easily corrupted firmware, the retrieved data often suffers from a loss of the mapping relationship between physical block addresses and logical addresses, making data recovery impossible.

[0021] To address the aforementioned problems, embodiments of this application provide a data recovery method, apparatus, electronic device, and program product. In response to a received data recovery instruction, a first target disk corresponding to the instruction is determined. A target data interface is invoked to retrieve raw data from the first target disk. User data and metadata are processed based on a pre-trained network model to obtain a target mapping table between logical addresses and physical block addresses. Based on this mapping table, target data is determined and stored on a second target disk. This allows for direct reading of raw data from the first target disk via the target data interface in the event of a first target disk failure. The mapping table between logical addresses and physical block addresses is then reconstructed based on the pre-trained network model, resulting in the recovered target data, which is then stored, thus improving the data recovery completeness rate.

[0022] It should be noted that the data recovery method, apparatus, electronic device and program product provided in this application can be used in the field of electronic devices, or in any field other than electronic devices. This application does not limit the application field of the data recovery method, apparatus, electronic device and program product.

[0023] The aforementioned electronic devices may include personal computers (PCs), smartphones, and other electronic devices that support data recovery functions. This application embodiment does not impose any special restrictions on the specific type of the electronic device.

[0024] The following explanation will primarily use a personal computer (PC) as an example to illustrate this application.

[0025] like Figure 1 As shown in the embodiments of this application, the data recovery method includes the following steps: Step S101: In response to receiving a data recovery instruction, determine a first target disk corresponding to the data recovery instruction; wherein the first target disk includes multiple physical blocks.

[0026] In the application, upon receiving a data recovery instruction, the instruction is parsed to determine the first target disk corresponding to the instruction. The first target disk is a solid-state drive (SSD) that has failed (e.g., controller failure, firmware corruption, or physical damage) and can no longer store data. The first target disk comprises multiple physical blocks. Each physical block stores original data, including but not limited to user data, metadata, and redundancy check data. The data recovery instruction is an instruction or request to perform data recovery processing on the original data in the first target disk.

[0027] Among them, the data recovery command can be a command or request sent by the user through the user terminal; or it can be a command or request automatically generated based on the user's operation on the current electronic device (such as pressing external buttons, clicking or dragging the screen of the current electronic device).

[0028] Step S102: Call the target data interface to obtain raw data from the first target disk; wherein the raw data includes at least one of user data, metadata and redundancy check data in each physical block.

[0029] In the application, the target data interface is called to read the raw data in the NAND flash memory directly without going through the controller. The raw data includes, but is not limited to, at least one of the following: user data, metadata, and redundancy check data stored in each physical block of the NAND flash memory.

[0030] The target data interface refers to the physical interface of the NAND flash memory that is adapted to the first target disk. For example, the Open NAND Flash Interface (ONFI) interface and the Toggle / Async conversion interface.

[0031] As an example and not a limitation, the current electronic device establishes a physical connection with the first target disk through the target data interface, and invokes the target data interface by sending a data acquisition command to achieve the operation of reading raw data directly from the NAND flash memory without going through the original controller of the first target disk.

[0032] Understandably, user data and metadata are in an undecoded state when directly extracted from the raw data in NAND flash memory. Redundancy check data can be used to decode the user data and metadata, yielding the decoded user data and metadata. Based on this decoded user data and metadata, a target mapping table can then be constructed.

[0033] As an example and not a limitation, a hardware reading device can be pre-built before invoking the target data interface. This hardware reading device includes, but is not limited to, a signal adaptation module, a voltage regulation module, a data buffer module, and a control module. The signal adaptation module is used to determine the communication protocol of the first target disk and perform protocol conversion based on the aforementioned communication protocol to convert the acquired raw data into data under the communication protocol supported by the current electronic device. The voltage regulation module is used to acquire the operating voltage of the first target disk and perform voltage regulation processing based on the aforementioned operating voltage to adapt to the actual voltage data of the first target disk. The control module is used to read raw data from the NAND flash memory based on the control acquisition command when a data acquisition command is received. The data buffer module is used to store the acquired raw data.

[0034] In some embodiments, the acquired raw data can be stored in the form of a binary data stream in the temporary storage medium of the current electronic device to facilitate data recovery processing.

[0035] By using a pre-defined target data interface, it is possible to directly read the raw data in the underlying NAND without going through the SSD controller or being affected by SSD controller or firmware failure, which facilitates data recovery operations and is applicable to various types of NAND SSDs.

[0036] Step S103: Process the user data and the metadata based on the pre-trained network model to obtain a target mapping relationship table; wherein, the target mapping relationship table includes the mapping relationship between the logical address and the physical block address of the user data.

[0037] In the application, the user data and metadata in the original data are processed based on the pre-trained network model to obtain the target mapping table, which includes the mapping relationship between the logical address and physical block address of the user data.

[0038] In this context, the logical address refers to the storage address under the operating system of the first target disk. The physical block address, also known as the physical block address, describes the location of the physical block storing user data. Based on the mapping relationship between logical addresses and physical block addresses, the corresponding storage content (i.e., the corresponding physical block) can be accessed to retrieve the corresponding user data.

[0039] Step S104: Determine the target data based on the target mapping table.

[0040] In the application, based on the target mapping table, a logically continuous user data stream is constructed to obtain the target data.

[0041] Step S105: Store the target data to the second target disk.

[0042] In the application, user data from the target data is written to and stored sequentially on the second target disk according to the logical address order. The second target disk refers to a fault-free, undamaged solid-state disk with data storage capabilities.

[0043] As is understood, the dynamic mapping table refers to the Flash Translation Layer (FTL) in an SSD, which stores the mapping relationship between logical addresses and physical block addresses. The physical block address of user data in the flash memory chip can be determined based on the logical address. When an SSD experiences controller failure, firmware corruption, or physical damage, the FTL is affected, causing the mapping relationship in the dynamic mapping table to be lost. This makes it impossible to accurately locate the physical block address of user data, thus preventing data repair operations. This application embodiment uses a pre-trained network model to construct, repair, and complete the mapping relationship between logical addresses and physical block addresses in the first target disk, obtaining a target mapping table. This table can determine the corresponding physical block address based on the logical address, thereby obtaining the corresponding target data and improving the probability and integrity of data repair.

[0044] like Figure 2 As shown, in one embodiment, step S103 includes the following steps: Step S1031: Cluster the user data and metadata to determine the physical block addresses with the same logical address, and construct a first mapping relationship; wherein, the first mapping relationship is the mapping relationship between the logical address and the physical block address.

[0045] In the application, the user data and metadata in the acquired raw data are clustered to determine the physical block addresses that belong to the same logical file (i.e., have the same logical address). Based on the clustering results and the newness level of the physical blocks, a complete mapping relationship between logical addresses and physical block addresses is constructed, which is the first mapping relationship.

[0046] Clustering algorithms include, but are not limited to, unsupervised K-means clustering or density-based spatial clustering (DBSCAN) algorithms.

[0047] For example, by inputting raw data, and based on the semantic features of the raw data (e.g., file header identification information, data type feature information, etc.), the raw data is clustered to determine the physical block addresses belonging to the same logical file. This is then combined with the structural features of the file system in the first target disk and the new / old level of the physical blocks to construct a first mapping relationship between the logical addresses and the physical block addresses. The file system includes, but is not limited to, the default file system of the Windows operating system (New Technology File System, NTFS) or the default file system of the Linux system (Fourth Extended File System, EXT4).

[0048] Optionally, before clustering, the original data can be parsed to extract metadata related to the dynamic mapping table, and then the data recovery operation can be performed. This metadata includes, but is not limited to, at least one of the following: block allocation table, wear count, and logical address markers.

[0049] Step S1032: Identify the erroneous mapping relationships in the first mapping relationship; wherein, the erroneous mapping relationships include a first erroneous mapping relationship with a logical address error and a second erroneous mapping relationship with a missing logical address.

[0050] In the application, data reading operations are performed based on each logical address and its corresponding physical block address in the first mapping relationship. Based on the reading results, the system identifies whether there is a first erroneous mapping relationship in the first mapping relationship where the logical address does not match (or the logical address is incorrect); and whether there is a second erroneous mapping relationship in the first mapping relationship where the logical address is missing, thus obtaining the erroneous mapping relationship identification result.

[0051] Step S1033: Input the error mapping relationship, the user data, and the metadata into the pre-trained network model for completion processing to obtain the target mapping relationship.

[0052] In the application, there may be incorrect mapping relationships in the first mapping relationship. The first mapping relationship between logical address and physical block address, as well as the user data and metadata in the original data, are input into the pre-trained network model to correct and complete the incorrect mapping relationships in the first mapping relationship, so as to obtain the target mapping relationship between logical address and physical block address.

[0053] Understandably, the first target SSD suffers from firmware corruption, resulting in mapping jumps due to bad physical blocks and incorrect metadata markings caused by NAND bit flips, leading to incomplete logical addresses. This results in the initial mapping relationship between logical addresses and physical block addresses constructed from the original data potentially containing errors such as incorrect logical address mappings or missing logical address mappings. A pre-trained network model can correct these errors in the initial mapping relationship, yielding a more accurate target mapping relationship and thus improving the accuracy of the mapping.

[0054] In one embodiment, the pre-trained network model includes, but is not limited to, a pre-trained recurrent neural network model and a pre-trained sequence recognition model.

[0055] Optionally, the output of the pre-trained recurrent neural network model is connected to the input of the pre-trained sequence recognition model.

[0056] Step S1034: Based on the first mapping relationship and the target mapping relationship, construct the target mapping relationship table.

[0057] In application, the target mapping relationship includes: a second mapping relationship obtained after correcting the first erroneous mapping relationship, and a third mapping relationship obtained after correcting the second erroneous mapping relationship. The first mapping relationship and the target mapping relationship are transformed to obtain a corresponding target mapping relationship table, which stores each logical address and its corresponding physical block address. This facilitates finding the physical block address corresponding to the logical address based on the target mapping relationship table, thereby obtaining the corresponding user data and the repaired target data.

[0058] By using a pre-trained network model to repair abnormal data and correct erroneous mapping relationships, the target mapping relationship table between logical addresses and physical block addresses can be reconstructed, thereby improving the data recovery integrity rate and the accuracy of the target mapping relationship table.

[0059] like Figure 3 As shown, in one embodiment, step S1033 includes the following steps: Step S10331: Input the user data, metadata, and the first error mapping relationship into the pre-trained recurrent neural network model for relationship correction processing to obtain the second mapping relationship.

[0060] In the application, the first erroneous mapping relationship, each user data, and each metadata in the first mapping relationship are input into a pre-trained recurrent neural network model (RNN). The pre-trained recurrent neural network model corrects the first erroneous mapping relationship (e.g., a logical address error jump mapping relationship) to obtain the second mapping relationship between the logical address and the physical block address.

[0061] Step S10332: Identify the starting address of each of the user data and each of the metadata; wherein the starting address is used to indicate the starting position of the physical block address.

[0062] In the application, the firmware type of the first target disk is identified, and a metadata feature library of the first target disk is constructed based on the firmware type. The starting address of each user data and metadata is identified based on the metadata feature library. The starting address is used to indicate the starting position of the physical block address.

[0063] It is understandable that the first target disk, the SSD, uses the FTL to construct a dynamic mapping relationship between the starting address and the logical address (LBA), thus obtaining a dynamic mapping table. The mapping relationship between the starting address and the logical address can also be called the mapping relationship between the physical block address and the logical address.

[0064] Step S10333: Input the starting addresses, user data, metadata, and the second error mapping relationship into the pre-trained sequence recognition model for logical address supplementation processing to obtain the third mapping relationship.

[0065] In the application, each metadata, each starting address, each user data, and the second error mapping relationship are input into a pre-trained sequence recognition model. The pre-trained sequence recognition model identifies the second error mapping relationship caused by abnormal metadata and user data (e.g., abnormal data such as incomplete data and blurred data due to NAND bit flipping causing metadata to have incorrect markings). Logical address supplementation processing is then performed to obtain the third mapping relationship between the logical address and the physical block address.

[0066] By identifying and completing abnormal data through a pre-trained network model and correcting erroneous mapping relationships, the mapping relationship between logical addresses and physical block addresses can be quickly reconstructed, thereby achieving data recovery efficiently and accurately.

[0067] like Figure 4 As shown, in one embodiment, the following steps are included before step S1031: Step S1035: When encrypted data is detected in the original data, the corresponding key dataset is determined based on historical key feature information; Step S1036: Decrypt the encrypted data based on the key dataset to obtain decrypted data.

[0068] In applications, when encrypted data is detected within the acquired raw data, it is determined that a target mapping table cannot be directly constructed based on the encrypted data. Historical key characteristic information can be obtained, and based on this information, the corresponding key dataset can be determined. The encrypted data is then decrypted using this key dataset to obtain decrypted data. This decrypted data facilitates the establishment of a mapping relationship between logical addresses and physical block addresses, resulting in the target mapping table. The historical key characteristic information refers to the characteristic information of the key data used to decrypt the raw data before the current time value.

[0069] Optionally, the key dataset can be determined using a preset neural network model. For example, historical key feature information and encrypted data are input into a preset neural network model to obtain the key dataset output by the model. This preset neural network model includes, but is not limited to, any pre-trained network model with key recognition capabilities, such as a Recurrent Neural Network (RNN) or a Convolutional Neural Network (CNN).

[0070] By constructing a key dataset and decrypting the encrypted original data, a data foundation is laid for establishing the mapping relationship between logical addresses and physical block addresses, thereby improving the integrity rate of data recovery.

[0071] like Figure 5 As shown, in one embodiment, step S104 includes the following steps: Step S1041: Based on the target mapping relationship table, the user data is concatenated to obtain an initial data stream; Step S1042: Detect the initial data stream and obtain a detection result; wherein the detection result is used to indicate the integrity of the initial data stream; Step S1043: Add a marker to the initial data stream based on the detection results to obtain the target data.

[0072] In the application, based on the logical addresses in the target mapping table, the user data stored in the physical block corresponding to the mapped physical block address is obtained. The user data is then concatenated according to the logical address order to obtain an initial data stream with consecutive logical addresses. Integrity verification is performed on the initial data stream to obtain the corresponding detection result. Based on the detection result, a corresponding tag is added to the initial data stream to obtain the target data.

[0073] The detection results are used to indicate the integrity of the initial data stream.

[0074] For example, when the detection result indicates that the initial data stream is complete, a data stream completeness marker can be added to the initial data stream. This allows the second target disk to determine that the data recovery operation is complete based on the marker, and then query and retrieve the data. When the detection result indicates that the initial data stream is incomplete, a data stream incompleteness marker can be added to the initial data stream. This allows the second target disk to determine that some data has not been repaired based on the marker, reducing the need for querying and retrieving unrepaired data.

[0075] As an example rather than a limitation, the initial data stream can be input into a pre-trained convolutional neural network (CNN), the accuracy of the data can be verified by the pre-trained CNN, and the file structure integrity of the data can be verified based on the file system of the first target disk, and the corresponding detection results can be output.

[0076] As an example rather than a limitation, when the detection result indicates that the initial data stream is incomplete, the corrupted data in the initial data stream (e.g., missing data, ambiguous data, etc.) can be automatically marked. The corrupted data and context in the initial data stream can be repaired and input into a pre-trained generative adversarial network (GAN). The pre-trained GAN performs byte completion processing on the initial data stream to supplement missing and ambiguous data, thereby improving the integrity of the initial data stream.

[0077] For example, data stored in a fault-free, undamaged SSD can be pre-trained into a convolutional neural network (CNN) to learn the integrity features of the data and the file structure integrity features, resulting in a pre-trained CNN. Similarly, data stored in the fault-free, undamaged SSD can be pre-trained into a generative adversarial network (GAN) to learn the feature information of the undamaged data and the contextual feature information of the undamaged data, resulting in a pre-trained GAN. The fault-free, undamaged SSD has the same firmware type as the first target disk (e.g., custom firmware, generic firmware).

[0078] In some embodiments, the initial data stream can also be displayed on the screen of the current electronic device to facilitate the user's detection of the integrity of the initial data stream and determination of the corresponding detection result. Alternatively, the initial data stream can be sent to the user terminal to facilitate the user's integrity verification of the initial data stream and the acquisition of the detection result returned by the user terminal.

[0079] like Figure 6 As shown, in one embodiment, the data recovery method further includes the following steps: Step S201: Obtain training data; wherein, the training data includes a logical address to be trained and a physical block address to be trained that maps to the logical address to be trained. Step S202: Input the data to be trained into the recurrent neural network model for training to obtain the pre-trained recurrent neural network model.

[0080] In this application, data from an undamaged SSD is obtained as training data. This training data includes the logical addresses to be trained and the physical block addresses that are mapped to these logical addresses. The training data is then input into a recurrent neural network model for training, allowing the model to learn the mapping relationship between the logical addresses and the physical block addresses, resulting in a pre-trained recurrent neural network model.

[0081] as well as, Step S203: Input the data to be trained into the sequence recognition model for training to obtain the pre-trained sequence recognition model.

[0082] In the application, the data to be trained is input into the sequence recognition model for training, so that the sequence recognition model learns the mapping relationship between the logical address to be trained and the physical block address to be trained, thus obtaining the pre-trained sequence recognition model.

[0083] In some embodiments, the training data described above may also be referred to as a dynamic mapping table in an undamaged solid-state drive (SSD).

[0084] like Figure 7 As shown, in one embodiment, step S105 further includes the following steps: Step S1051: Determine the logical address corresponding to each piece of user data in the target data; Step S1052: According to the logical address, each user data is stored sequentially in the second target disk.

[0085] In the application, the logical addresses corresponding to each user data in the target data are identified, and the corresponding user data are stored in the second target disk in the order of the logical addresses, so that the second target disk stores the user data and the logical addresses mapped to the user data, thus completing the data recovery operation.

[0086] In some embodiments, after the user data is stored in a second target disk, the original data stored in the temporary storage medium may be erased.

[0087] In some embodiments, a physical connection may be pre-established with the physical interface of the second target disk. The physical interface mentioned above includes, but is not limited to, a Serial Advanced Technology Attachment (SATA) interface and a Non-Volatile Memory Express (NVMe) interface.

[0088] In some embodiments, a data writing device can be pre-built. When a write operation is determined to be performed, a data write instruction is sent to the data writing device to write target data to the second target disk through the physical interface. This data writing device includes, but is not limited to, a driver adaptation module, a write control module, and a data processing module. The data processing module performs format conversion on the target data, changing its format to a data format supported by the second target disk, and concatenates the corresponding user data according to the logical address order to obtain the target data. The driver adaptation module determines the protocol supported by the physical interface and adjusts the interface protocol to ensure compatibility with the physical interface, facilitating data transmission. The data writing device receives the data write instruction and, based on the data write instruction, sequentially writes the corresponding user data to the second target disk according to the logical address order.

[0089] By directly reading the raw data from the disk through the target data interface, a mapping table between logical addresses and physical block addresses is reconstructed based on a pre-trained network model. The recovered target data is then determined based on the mapping table and stored in the second target disk, thus realizing data recovery operations and avoiding problems such as data loss.

[0090] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0091] This application also provides a data recovery apparatus for performing the steps described in the data recovery method embodiments above. The data recovery apparatus can be a virtual appliance within an electronic device, run by the electronic device's processor, or it can be the electronic device itself. Figure 8 As shown, the data recovery apparatus 800 provided in this application embodiment includes: The first determining module 801 is configured to determine a first target disk corresponding to the data recovery instruction in response to receiving the data recovery instruction; wherein the first target disk includes a plurality of physical blocks; The data acquisition module 802 is used to call the target data interface to acquire raw data from the first target disk; wherein, the raw data includes at least one of user data, metadata and redundancy check data in each of the physical blocks; The processing module 803 is used to process the user data and the metadata based on a pre-trained network model to obtain a target mapping relationship table; wherein, the target mapping relationship table includes the mapping relationship between the logical address and the physical block address of the user data; The second determining module 804 is used to determine target data based on the target mapping relationship table; Storage module 805 is used to store the target data to a second target disk.

[0092] In some embodiments, the processing module includes: A clustering processing unit is used to perform clustering processing on each of the user data and each of the metadata, determine the physical block addresses with the same logical address, and construct a first mapping relationship; wherein, the first mapping relationship is the mapping relationship between the logical address and the physical block address; An identification unit is used to identify erroneous mapping relationships in the first mapping relationship; wherein, the erroneous mapping relationship includes a first erroneous mapping relationship with a logical address error and a second erroneous mapping relationship with a missing logical address; The model processing unit is used to input the error mapping relationship, the user data, and the metadata into the pre-trained network model for completion processing to obtain the target mapping relationship; The mapping table construction unit is used to construct the target mapping relationship table based on the first mapping relationship and the target mapping relationship.

[0093] In some embodiments, the pre-trained network model includes a pre-trained recurrent neural network model and a pre-trained sequence recognition model; the model processing unit is specifically used for: The user data, metadata, and the first erroneous mapping relationship are input into the pre-trained recurrent neural network model for relationship correction processing to obtain the second mapping relationship; Identify the starting address of each of the user data and each of the metadata; wherein the starting address is used to indicate the starting position of the physical block address; The starting addresses, user data, metadata, and the second error mapping relationship are input into the pre-trained sequence recognition model for logical address supplementation processing to obtain the third mapping relationship.

[0094] In some embodiments, the processing module further includes: A key determination unit is used to determine the corresponding key dataset based on historical key feature information when encrypted data is detected in the original data. The decryption unit is used to decrypt the encrypted data based on the key dataset to obtain decrypted data.

[0095] In some embodiments, the second determining module includes: The data splicing unit is used to splice the user data based on the target mapping relationship table to obtain an initial data stream; A data detection unit is used to detect the initial data stream and obtain a detection result; wherein the detection result is used to indicate the integrity of the initial data stream; The initial data stream is tagged based on the detection results to obtain the target data.

[0096] In some embodiments, the data recovery apparatus further includes: The training data acquisition module is used to acquire training data; wherein, the training data includes a logical address to be trained and a physical block address to be trained that maps to the logical address to be trained. The first training module is used to input the data to be trained into the recurrent neural network model for training, so as to obtain the pre-trained recurrent neural network model. as well as, The second training module is used to input the data to be trained into the sequence recognition model for training, so as to obtain the pre-trained sequence recognition model.

[0097] In some embodiments, the storage module includes: The address determination unit is used to determine the logical address corresponding to each piece of user data in the target data; The storage unit is used to sequentially store each of the user data into the second target disk according to the logical address.

[0098] In applications, the modules in a data recovery device can be software program modules, or they can be implemented through different logic circuits integrated in a processor, or they can be implemented through multiple distributed processors.

[0099] like Figure 9 As shown, this application embodiment also provides an electronic device 900, including: at least one processor 901 ( Figure 9 The diagram shows only one processor, memory 902, and computer program 903 stored in memory 902 and executable on at least one processor 901. When processor 901 executes computer program 903, it implements the steps in any of the above method embodiments.

[0100] In applications, electronic devices may include, but are not limited to, processors and memory. Those skilled in the art will understand that... Figure 9 This is merely an example of an electronic device and does not constitute a limitation on electronic devices. It may include more or fewer components than shown in the illustration, or combinations of certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0101] Optionally, the electronic device also includes a hardware data interface and an intermediate connector. The hardware data interface is used to enable data transmission between different components, and the intermediate connector is used to enable connection between different components.

[0102] In applications, the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0103] In applications, memory can be an internal storage unit of an electronic device in some embodiments, such as a hard drive or RAM. In other embodiments, memory can be an external storage device of the electronic device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal and external storage units of the electronic device. Memory is used to store operating systems, applications, bootloaders, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.

[0104] It should be noted that the information interaction and execution process between the above-mentioned devices / modules are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0105] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The functional modules in the embodiments can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules can be implemented in hardware or as software functional modules. Furthermore, the specific names of the functional modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the modules in the above-described device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0106] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps described in the various method embodiments above.

[0107] This application provides a computer program product that, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.

[0108] If an integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0109] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0110] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0111] In the embodiments provided in this application, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or modules may be electrical, mechanical, or other forms.

[0112] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0113] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A data recovery method, characterized in that, include: In response to receiving a data recovery instruction, a first target disk corresponding to the data recovery instruction is determined; wherein, the first target disk includes a plurality of physical blocks; The target data interface is invoked to obtain raw data from the first target disk; wherein, the raw data includes at least one of user data, metadata, and redundancy check data in each of the physical blocks; The user data and metadata are processed based on a pre-trained network model to obtain a target mapping table; wherein, the target mapping table includes the mapping relationship between the logical address and the physical block address of the user data; Based on the target mapping table, the target data is determined; The target data is stored in the second target disk.

2. The method as described in claim 1, characterized in that, The pre-trained network model processes the user data and the metadata to obtain a target mapping table, including: Clustering is performed on each of the user data and each of the metadata to determine the physical block addresses with the same logical address, and a first mapping relationship is constructed; wherein, the first mapping relationship is the mapping relationship between the logical address and the physical block address; Identify erroneous mapping relationships in the first mapping relationship; wherein, the erroneous mapping relationship includes a first erroneous mapping relationship with a logical address error and a second erroneous mapping relationship with a missing logical address; The error mapping relationship, the user data, and the metadata are input into the pre-trained network model for completion processing to obtain the target mapping relationship; Based on the first mapping relationship and the target mapping relationship, the target mapping relationship table is constructed.

3. The method as described in claim 2, characterized in that, The pre-trained network model includes a pre-trained recurrent neural network model and a pre-trained sequence recognition model; the step of inputting the error mapping relationship, the user data, and the metadata into the pre-trained network model for completion processing to obtain the target mapping relationship includes: The user data, metadata, and the first erroneous mapping relationship are input into the pre-trained recurrent neural network model for relationship correction processing to obtain the second mapping relationship; Identify the starting address of each of the user data and each of the metadata; wherein the starting address is used to indicate the starting position of the physical block address; The starting addresses, user data, metadata, and the second error mapping relationship are input into the pre-trained sequence recognition model for logical address supplementation processing to obtain the third mapping relationship.

4. The method as described in claim 3, characterized in that, Before performing clustering processing on each of the user data and each of the metadata to determine the physical block addresses with the same logical address and constructing the first mapping relationship, the method further includes: When encrypted data is detected in the original data, the corresponding key dataset is determined based on historical key feature information; The encrypted data is decrypted based on the key dataset to obtain decrypted data.

5. The method as described in claim 1, characterized in that, The step of determining the target data based on the target mapping table includes: The user data is concatenated based on the target mapping table to obtain an initial data stream; The initial data stream is inspected to obtain an inspection result; wherein the inspection result is used to indicate the integrity of the initial data stream; Based on the detection results, tags are added to the initial data stream to obtain the target data.

6. The method according to any one of claims 1 to 5, characterized in that, Before the pre-trained network model processes the user data and the metadata to obtain the target mapping table, the process includes: Acquire training data; wherein the training data includes a logical address to be trained and a physical block address to be trained that maps to the logical address to be trained; The data to be trained is input into the recurrent neural network model for training to obtain the pre-trained recurrent neural network model; as well as, The data to be trained is input into the sequence recognition model for training to obtain the pre-trained sequence recognition model.

7. The method according to any one of claims 1 to 5, characterized in that, The step of storing the target data to the second target disk includes: Determine the logical address corresponding to each piece of user data in the target data; According to the logical address, each user data is stored sequentially in the second target disk.

8. A data recovery device, characterized in that, include: A first determining module is configured to, in response to receiving a data recovery instruction, determine a first target disk corresponding to the data recovery instruction; wherein the first target disk includes a plurality of physical blocks; The data acquisition module is used to call the target data interface to acquire raw data from the first target disk; wherein, the raw data includes at least one of user data, metadata and redundancy check data in each of the physical blocks; The processing module is used to process the user data and the metadata based on a pre-trained network model to obtain a target mapping table; wherein, the target mapping table includes the mapping relationship between the logical address and the physical block address of the user data; The second determining module is used to determine target data based on the target mapping relationship table; A storage module is used to store the target data to a second target disk.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the data recovery method according to any one of claims 1 to 7.

10. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the steps of the data recovery method according to any one of claims 1 to 7.