A method and system for database content recovery

By scanning the disk and checking the deep neural network, obtaining and reorganizing the deleted database file node information blocks, the problem of failing to effectively analyze multiple historical files and logging blocks in the prior art is solved, and the maximum recovery of database content is achieved.

CN114968663BActive Publication Date: 2025-05-30CHENGDU YIWO TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210552412.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-30
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

The existing database recovery technology only analyzes a single database file or log file, and fails to effectively analyze multiple historical existing files and log record blocks identified by the file system driver as deleted, resulting in the content of the deleted data block not being analyzed, and the recovery effect is not ideal.

Method used

By scanning the disk, the original file node information blocks identified as deleted are obtained, their recoverability weight is calculated, and the file node information blocks are checked and reorganized using the deep neural network to form the restored database content and merge them with the existing database content.

Benefits of technology

Recover deleted database content to the maximum extent. In theory, as long as the data is still stored on the hard disk, it can be scanned and restored, overcoming the disadvantage of only targeting a single file or log in traditional technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114968663B_ABST
    Figure CN114968663B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for database content recovery. The method includes the following steps: obtaining disk partition information corresponding to the storage of database files and log files; finding the log files according to the disk partition information, and scanning the disk to obtain database file node information blocks with deletion marks; calculating the recoverability weights of the file node information blocks, sorting and de-duplicating the file node information blocks according to the recoverability weights; building a deep neural network, using the deep neural network to perform correctness verification on the file node information blocks, removing the incorrect file point information blocks, and reorganizing the remaining file point information blocks after removing the errors to form the recovered database content; merging the recovered database content with the existing database content. By identifying, reorganizing, and ensuring the correctness through deep learning of this part of the deleted file node information blocks, the database content can be recovered to the greatest extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of disk recovery, and particularly relates to a method and system for database content recovery. Background Art

[0002] Database content (including information such as indexes and logs) ultimately exists in the form of files on the disk, and these files are managed by the operating system. When there are significant changes in the database content (such as deleting tables, clearing data, creating new table columns, etc.), the files stored on the hard disk also change. Among them, the deleted data blocks are removed from the cluster chain of the original files, and the cluster chain of the existing file data blocks is reconstructed.

[0003] Existing database recovery technologies are based on analyzing and collecting evidence from the database files (real files stored on the hard disk) or database log files after deletion. However, after the database content is deleted, the database files have changed significantly compared to before deletion (such as deleting database table columns or entire tables or performing compression operations, etc.). Since only a single real existing database file or log file is analyzed, and multiple historical existing files and log record blocks marked as deleted in the hard disk storage space by the file system driver are not analyzed, the content of the deleted data blocks is not analyzed, resulting in unsatisfactory recovery effects. Summary of the Invention

[0004] In view of this, the present invention provides a method and system for database content recovery. After scanning the disk to obtain the original file nodes marked as deleted, the database is recovered by analyzing the file node information blocks, so as to recover the deleted database content to the greatest extent.

[0005] To solve the above technical problems, the technical solution of the present invention is to adopt a method for database content recovery, including:

[0006] Obtain the disk partition information corresponding to the storage of the database file and the log file;

[0007] Find the log file according to the disk partition information, and scan the disk to obtain the database file node information blocks with deletion marks;

[0008] Calculate the recoverability weight of the file node information blocks, and sort and deduplicate the file node information blocks according to the recoverability weight;

[0009] Build a deep neural network, use the deep neural network to perform correctness verification on the file node information blocks, remove the incorrect file node information blocks, and reorganize the remaining file node information blocks after removing the errors to form the recovered database content;

[0010] Merge the recovered database content with the existing database content.

[0011] As an improvement, the disk partition information corresponding to the storage of the database file and the log file is obtained by reading the database configuration file or the registry.

[0012] As a further improvement, the method for scanning the disk according to the disk partition information to obtain the database file node information block with a deletion flag includes:

[0013] Quickly scan some partitions of the hard disk using the log file analysis result to quickly obtain the file node information block before some databases are deleted;

[0014] Deeply scan all partitions of the hard disk using the magic value of the database file page feature to obtain the remaining file node information block before the database is deleted.

[0015] As another further improvement, the file node information block includes:

[0016] The data information block for storing user data, the index information block for quickly finding user data, and the log information block for recording the operation logs of database administrators.

[0017] As an improvement, the method for calculating the recoverability weight of the file node information block includes:

[0018] For the data information block, obtain the serial number of the data information block referenced by the index information block or the log information block. The correctness weight of the data information block referenced by the index information block or the log information block needs to be increased, and the correctness weight of the data information block not referenced by the index information block and the log information block needs to be decreased;

[0019] For the index information block, obtain the serial numbers of the parent block and the child block of the index information block. The correctness weights of the parent block and the child block of the index information block need to be increased;

[0020] For the log information block, obtain the specific operations of the log information block. The correctness weights of the log information blocks for inserting data and updating data need to be increased, and the correctness weights of the log information blocks for deleting data and deleting tables need to be decreased.

[0021] As an improvement, the file information blocks are sorted in descending order according to the required value of the correctness weight.

[0022] As an improvement, the construction of the deep neural network includes:

[0023] Train the deep neural network using the training data result set;

[0024] Use the trained deep neural network to perform deep learning on the existing database content to output the result set;

[0025] Merge the learned result set with the training data result set to update the training data result set.

[0026] As an improvement, the method for using a deep neural network to perform correctness verification on file node information blocks includes:

[0027] Successively call several rules in the training data result set to perform correctness verification on each file node information block one by one;

[0028] If any one of the rules is satisfied, the file node information block is considered correct.

[0029] As an improvement, the method for merging the restored database content and the existing database content includes:

[0030] Compare the restored database content with the table content of the existing database. For tables with the same name, if there is content in the restored database content that does not exist in the existing database, add this part of the content to the corresponding table in the existing database; for tables that do not exist in the existing database, create a new table with the same name in the existing database and import the corresponding data in the restored database content into the newly created table with the same name.

[0031] The present invention also provides a database content recovery system, including:

[0032] A disk partition information acquisition module, configured to acquire disk partition information corresponding to the storage of database files and log files;

[0033] A file node information block scanning module, configured to find a log file according to the disk partition information and scan the disk to obtain database file node information blocks with deletion marks;

[0034] A file node information block sorting module, configured to calculate the recoverability weight of file node information blocks, sort and deduplicate the file node information blocks according to the recoverability weight;

[0035] A file point information block correctness verification module, configured to build a deep neural network, use the deep neural network to perform correctness verification on file node information blocks, remove incorrect file point information blocks, and reorganize the remaining file point information blocks after removing errors to form the restored database content;

[0036] A database content merging module, configured to merge the restored database content and the existing database content.

[0037] The advantages of the present invention are as follows: The present invention overcomes the drawback that traditional database recovery only targets the recovery of a single database file or log information, while the actually deleted data still exists on physical storage media such as hard disks and fails to be recognized and recovered. By identifying, reorganizing, and ensuring the correctness through deep learning of this part of the deleted file node information blocks, the database content can be recovered to the greatest extent. In theory, as long as the data is still stored on the hard disk, it can be scanned and recovered. Description of the Drawings

[0038] Figure 1 It is a flowchart of the present invention.

[0039] Figure 2 It is a structural schematic diagram of the present invention. Detailed Implementation Manner

[0040] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below in conjunction with the specific implementation manner.

[0041] The database content (including information such as indexes and logs) ultimately exists in the form of files on the disk, and these files are managed by the operating system. When the database content undergoes huge changes (such as deleting tables, emptying data, creating new table columns, etc.), the files stored on the hard disk also change. Among them, the deleted data blocks will be removed from the cluster chain of the original file, and the cluster chain of the existing file data blocks will be reconstructed. The file system (NTFS, Refs, EXT, etc.) will generate new MFT or iNode file node information according to its strategy, resulting in the coexistence of the original file nodes marked as deleted and the information of the existing file nodes after deletion on the hard disk. At the same time, the file system driver will also record information such as the starting position and size of the corresponding deleted and modified data blocks in its log management file.

[0042] The present invention utilizes the original file nodes marked as deleted on the hard disk to recover the database content to the greatest extent.

[0043] As Figure 1 shown, the present invention provides a method for recovering database content, including:

[0044] S1 Obtain the disk partition information corresponding to the storage of the database file and the log file; in this embodiment, the disk partition information corresponding to the storage of the database file and the log file is obtained by reading the database configuration file or the registry. The disk partition information includes the drive letter, size, file system type, etc.

[0045] S2 Locate the log file according to the disk partition information, and scan the disk to obtain the database file node information block with a deletion flag; in this embodiment, the file node information block before database deletion is obtained through two scanning methods: fast scanning and deep scanning.

[0046] The so-called fast scanning is to use the analysis result of the log file to scan some partitions of the hard disk to quickly obtain some file node information blocks before database deletion; through the log file, the data block distribution position and size of the deleted file node information block can be analyzed, and scanning can be targeted based on the above information, so the speed is relatively fast.

[0047] The so-called deep scanning is to use the magic value of the database file page feature to deeply scan all partitions of the hard disk to obtain the remaining file node information blocks before database deletion. Different database file pages have different constants at the page start position to distinguish. For example, the constant of the first page of the Sqlite database is "SQLite format 3", and this constant is also called the magic value. Using the magic value can comprehensively scan the disk to obtain all file node information blocks on the hard disk.

[0048] The reason for using the combination of the two scans is that although the deep scan is more comprehensive, its scanning speed is slow, and the time consumption can be up to ten times that of the fast scan. To reduce the waiting time, the fast scan is used to quickly obtain a part of the file node information blocks for subsequent operations, and at the same time, the remaining file node information blocks are scanned out through the deep scan, which saves more time and reduces the system waiting.

[0049] S3 Calculate the recoverability weight of the file node information block, sort and deduplicate the file node information block according to the recoverability weight; the file node information block includes: a data information block for storing user data, an index information block for quickly finding user data, and a log information block for recording the operation logs of database administrators. There are different weight calculation methods for different information blocks:

[0050] For the data information block, obtain the serial number of the data information block referenced by the index information block or the log information block, and the correctness weight of the data information block referenced by the index information block or the log information block needs to be increased, while the correctness weight of the data information block not referenced by the index information block and the log information block needs to be decreased;

[0051] For the index information block, obtain the serial numbers of the parent block and the child block of the index information block, and the correctness weights of the parent block and the child block of the index information block need to be increased;

[0052] For the log information block, for the specific operation of obtaining the log information block, the correctness weight of the log information block for inserting data and updating data needs to be increased, and the correctness weight of the log information block for deleting data and deleting tables needs to be decreased.

[0053] After the correctness weight of each file node information block needs to be determined, the file node information blocks are sorted in descending order according to the required value of the correctness weight, and the file node information blocks with high weights are presented in the front row for selection.

[0054] S4 Build a deep neural network, use the deep neural network to perform correctness verification on the file node information blocks, remove the incorrect file node information blocks, and reorganize the remaining file node information blocks after removing the errors to form the restored database content.

[0055] The specific correct verification methods include:

[0056] Successively call several rules in the training data result set to perform correctness verification on each file node information block one by one;

[0057] If any one of the rules is satisfied, the file node information block is considered correct.

[0058] In addition, during the construction process of the deep neural network, it can be trained and upgraded while learning. After the initial prototype of the deep neural network is initially built, the training data result set is used to train the deep neural network; the initial training data result set is generally self-assembled, including artificial intelligence algorithm sets such as name correctness judgment, date validity, and transaction content legality judgment.

[0059] Use the trained deep neural network to perform deep learning on the existing database content to output a result set; the result set output here is the result after the above-mentioned correctness verification.

[0060] Merge the learned result set with the training data result set to update the training data result set. Through continuous learning and training, the deep neural network is iteratively upgraded, and the result of its correctness verification is more accurate.

[0061] S5 Merge the restored database content and the existing database content. The specific merging method is: compare the restored database content with the table content of the existing database. For tables with the same name, if there is content in the restored database that does not exist in the existing database, add this part of the content to the corresponding table in the existing database; for tables that do not exist in the existing database, create a new table with the same name in the existing database, and import the corresponding data in the restored database content into the newly created table with the same name.

[0062] The content of the merged database is equivalent to the union of the existing database content and the restored database content, restoring to the situation before deletion to the greatest extent possible.

[0063] As Figure 2 shown, the present invention also provides a database content restoration system, including:

[0064] A disk partition information acquisition module, configured to acquire the disk partition information corresponding to the storage of database files and log files;

[0065] A file node information block scanning module, configured to find the log file according to the disk partition information, and scan the disk to obtain the database file node information blocks with deletion marks;

[0066] A file node information block sorting module, configured to calculate the recoverability weights of the file node information blocks, sort and deduplicate the file node information blocks according to the recoverability weights;

[0067] A file point information block correctness verification module, configured to build a deep neural network, use the deep neural network to verify the correctness of the file node information blocks, remove the incorrect file point information blocks, and reorganize the remaining file point information blocks after removing the errors to form the restored database content;

[0068] A database content merging module, configured to merge the restored database content and the existing database content.

[0069] The above are only the preferred embodiments of the present invention. It should be noted that the above preferred embodiments should not be regarded as limiting the present invention. The protection scope of the present invention should be subject to the scope defined by the claims. For those of ordinary skill in the art of this technology, without departing from the spirit and scope of the present invention, several improvements and refinements can also be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for restoring database content, characterized in that it includes: Obtaining the disk partition information corresponding to the storage of the database file and the log file; Finding the log file according to the disk partition information, and scanning the disk to obtain the database file node information block with a deletion flag; The file node information block includes: a data information block for storing user data, an index information block for quickly finding user data, and a log information block for recording the operation logs of database administrators; Calculating the recoverability weight of the file node information block, sorting and de-duplicating the file node information block according to the recoverability weight; The method for calculating the recoverability weight of the file node information block includes: for the data information block, obtaining the serial number of the data information block referenced by the index information block or the log information block, increasing the recoverability weight of the data information block referenced by the index information block or the log information block, and decreasing the recoverability weight of the data information block not referenced by the index information block and the log information block; for the index information block, obtaining the serial numbers of the parent block and the sub-block of the index information block, and increasing the recoverability weight of the parent block and the sub-block of the index information block; for the log information block, obtaining the log information block, increasing the recoverability weight of the log information block for inserting data and updating data, and decreasing the recoverability weight of the log information block for deleting data and deleting tables; Building a deep neural network, using the deep neural network to perform correctness verification on the file node information block, removing the incorrect file node information block, and reorganizing the remaining file node information block after removing the errors to form the restored database content; Merging the restored database content with the existing database content.

2. A method for restoring database content according to claim 1, characterized in that: Obtaining the disk partition information corresponding to the storage of the database file and the log file by reading the database configuration file or the registry.

3. A method for restoring database content according to claim 1, characterized in that The method for scanning the disk according to the disk partition information to obtain the database file node information block with a deletion flag includes: using the analysis result of the log file to quickly scan some partitions of the hard disk to quickly obtain some file node information blocks before the database is deleted; Using the magic value of the database file page feature to deeply scan all partitions of the hard disk to obtain the remaining file node information blocks before the database is deleted.

4. A method for restoring database content according to claim 1, characterized in that: Sorting the file information blocks in descending order according to the required value of the correctness weight.

5. A method for restoring database content according to claim 1, characterized in that Building the deep neural network includes: training the deep neural network using the training data result set; Using the trained deep neural network to perform deep learning on the input result set of the existing database content; Merging the learned result set with the training data result set to update the training data result set.

6. A method for restoring database content according to claim 1, characterized in that The method for verifying the correctness of file node information blocks using a deep neural network includes: sequentially calling a number of rules in the training data result set to verify the correctness of each file node information block one by one; If any one of the rules is satisfied, the file node information block is considered correct.

7. A method for restoring database content according to claim 1, characterized in that The method for merging the restored database content with the existing database content includes: comparing the restored database content with the table content of the existing database. For tables with the same name, if there is content in the restored database that does not exist in the existing database, the content that does not exist in the existing database is added to the corresponding table in the existing database; for tables that do not exist in the existing database, a new table with the same name is created in the existing database, and the corresponding data in the restored database content is imported into the newly created table with the same name.

8. A database content restoration system, characterized in that It includes: A disk partition information acquisition module for acquiring the disk partition information corresponding to the storage of database files and log files; A file node information block scanning module for finding the log file according to the disk partition information and scanning the disk to obtain database file node information blocks with deletion marks; The file node information block includes: a data information block for storing user data, an index information block for quickly searching user data, and a log information block for recording the operation logs of database administrators; A file node information block sorting module for calculating the recoverability weight of file node information blocks, sorting and de-duplicating the file node information blocks according to the recoverability weight; the method for calculating the recoverability weight of file node information blocks includes: for the data information block, obtaining the serial number of the data information block referenced by the index information block or the log information block, increasing the recoverability weight of the data information block referenced by the index information block or the log information block, and decreasing the recoverability weight of the data information block not referenced by the index information block and the log information block; for the index information block, obtaining the serial numbers of the parent block and the sub-block of the index information block, and increasing the recoverability weight of the parent block and the sub-block of the index information block; for the log information block, obtaining the log information block, increasing the recoverability weight of the log information block for inserting data and updating data, and decreasing the recoverability weight of the log information block for deleting data and deleting tables; A file node information block correctness verification module for building a deep neural network, using the deep neural network to verify the correctness of file node information blocks, removing incorrect file node information blocks, and reorganizing the remaining file node information blocks after removing errors to form the restored database content; A database content merging module for merging the restored database content with the existing database content.

Citation Information

Patent Citations

  • Caching method and system for single disk restoration of disk array

    CN106294032A

  • Database data recovery system and method

    US20040267835A1