Data recovery method and device, electronic equipment and computer readable medium

By analyzing and filtering the database log files, generating and executing data query statements, the problems of low data recovery efficiency and waste of storage resources in the existing technology are solved, and efficient data recovery and storage resources are achieved.

CN120144355APending Publication Date: 2025-06-13MULTIPOINT LIFE (CHENGDU) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311717019.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art when the database source instance is lost, data recovery efficiency is low, redundant data is generated, resulting in waste of storage resources, low database performance and poor user experience.

Method used

By obtaining and sorting database log file sequences, filtering target log file sequences in response to user requests, parsing and filtering log events, generating data query statements, performing recovery and storing data.

Benefits of technology

Improve data recovery efficiency, reduce waste of storage resources, improve user experience, and reduce database performance burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144355A_ABST
    Figure CN120144355A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data recovery method and device, electronic equipment and a computer readable medium. A specific embodiment of the method comprises the steps of obtaining a database log file sequence; screening at least one database log file from the database log file sequence to obtain a target log file sequence; analyzing the target log file sequence to obtain an analyzed log event sequence; filtering the analyzed log event sequence to obtain a filtered log event sequence; generating a data query statement information sequence; executing the data query statement information sequence; and storing the recovered to-be-recovered data set into a verification data table in a preset storage database, and generating a data recovery execution file and a user operation execution file. According to the embodiment, the data can be recovered by analyzing the log file of the local database under the condition that the source database instance is lost, so that the data recovery efficiency can be improved, the waste of storage resources is reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to data recovery methods, apparatuses, electronic devices, and computer-readable media. Background Art

[0002] Due to the update and iteration of database systems and user operations, it is easy to cause the loss of database source instances and data, reducing the security and stability of database systems. For data recovery in the case of the loss of a database source instance, the commonly used method is to send the sequence of database log files stored in disk files to the binlog2sql software, restore the sequence of database log files into SQL (Structured Query Language) statements, perform semantic inverse operations on the SQL statements to obtain inverse SQL statements, execute the inverse SQL statements to recover the data to be recovered, and store the recovered data to be recovered.

[0003] However, the inventors have found that when using the above method to recover data, the following technical problems often exist:

[0004] First, since restoring the sequence of database log files into SQL statements and performing semantic inverse operations on the SQL statements are step-by-step inverse operations on the data in the database until the data that the user wants to recover is traced back, some data may be data that the user does not want to recover, resulting in redundant data, as well as low data recovery efficiency and long recovery time, leading to waste of database storage resources, low database performance, and low user experience.

[0005] Second, the existing method for executing inverse SQL statements is to execute SQL statements after anomaly detection based on machine learning. Since machine learning requires a large amount of high-quality data for training and calculating the similarity between each query statement and other query statements, the detection efficiency is low, and a large amount of storage resources and computing resources are required. In the case of less training data or low-quality training data, false alarms and missed alarms are likely to occur, resulting in the leakage of database data and the damage of database servers.

[0006] Third, the existing technology for storing the recovered data to be recovered is to directly perform a full backup on the recovered data. There may be data in the current full backup that is not frequently accessed, and there is a large amount of redundant data that users do not need in the full backup, resulting in waste of storage resources and low user experience.

[0007] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to ordinary skilled in the art in this country. Summary of the Invention

[0008] This disclosure is in part for introducing concepts in a concise form, which will be described in detail in the following Detailed Description section. This disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0009] Some embodiments of the present disclosure provide a data recovery method, apparatus, electronic device, and computer-readable medium to solve one or more of the technical problems mentioned in the above Background section.

[0010] In a first aspect, some embodiments of the present disclosure provide a data recovery method, including: obtaining a database log file sequence, where the database log file sequence is a log file sequence sorted according to the corresponding time order; in response to receiving a user request information set, screening out at least one database log file that meets the request start time and request end time included in the user request information set from the database log file sequence to obtain a target log file sequence; performing parsing processing on the target log file sequence to obtain a parsed log event group sequence; filtering the parsed log event group sequence according to a user request information subset to obtain a filtered log event sequence, where the user request information subset is a set obtained by removing the request start time and the request end time from the user request information set; generating a data query statement information sequence according to the filtered log event sequence; executing the data query statement information sequence to recover a data set to be recovered corresponding to the data query statement information sequence; storing the recovered data set to be recovered in a verification data table in a preset storage database, generating a data recovery execution file according to the data query statement information sequence, and generating a user operation execution file according to the filtered log event sequence.

[0011] Second aspect, some embodiments of the present disclosure provide a data recovery device, including: an acquisition unit configured to acquire a database log file sequence, where the database log file sequence is a log file sequence sorted according to the corresponding chronological order; a screening unit configured to, in response to receiving a user request information set, screen out at least one database log file that satisfies the request start time and the request end time included in the user request information set from the database log file sequence to obtain a target log file sequence; a parsing and processing unit configured to perform parsing and processing on the target log file sequence to obtain a parsed log event group sequence; a filtering and processing unit configured to perform filtering and processing on the parsed log event group sequence according to a user request information subset to obtain a filtered log event sequence, where the user request information subset is a set obtained by removing the request start time and the request end time from the user request information set; a generating unit configured to generate a data query statement information sequence according to the filtered log event sequence; an execution unit configured to execute the data query statement information sequence to recover a data set to be recovered corresponding to the data query statement information sequence; a storage unit configured to store the recovered data set to be recovered in a verification data table in a preset storage database, generate a data recovery execution file according to the data query statement information sequence, and generate a user operation execution file according to the filtered log event sequence.

[0012] Third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device having stored thereon one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.

[0013] Fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having stored thereon a computer program, where the computer program, when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0014] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: In some embodiments of the present disclosure, in the case of the loss of the source database instance, the data recovery method can recover data by parsing the local database log file, which can improve the data recovery efficiency, reduce the waste of storage resources, and improve the user experience. Specifically, the reasons for the waste of relevant database storage resources, low database performance, and low user experience are as follows: Since restoring the database log file sequence into SQL statements and performing semantic inverse operations on the SQL statements is to perform step-by-step inverse operations on the data in the database until the data that the user wants to recover is traced back. There may be some data that the user does not want to recover, resulting in redundant data, as well as low data recovery efficiency and long recovery time, leading to waste of database storage resources, low database performance, and low user experience. Based on this, the data recovery method of some embodiments of the present disclosure can first obtain the database log file sequence, where the above-mentioned database log file sequence is a log file sequence sorted according to the corresponding time order. Here, the database log file sequence is the data log file sequence in the locally called file, which can reduce the waste of communication resources. Secondly, in response to receiving the user request information set, at least one database log file that meets the request start time and request end time included in the above-mentioned user request information set is filtered out from the above-mentioned database log file sequence to obtain the target log file sequence. Here, the amount of data to be recovered can be reduced, and the waste of computing resources can be reduced. Thirdly, the above-mentioned target log file sequence is parsed to obtain the parsed log event group sequence. Here, since the parsing is performed on the local target log file sequence, the number of target log files to be parsed can be reduced, improving data integrity and parsing efficiency. Then, according to the user request information subset, the above-mentioned parsed log event group sequence is filtered to obtain the filtered log event sequence, where the above-mentioned user request information subset is a set obtained by removing the above-mentioned request start time and the above-mentioned request end time from the above-mentioned user request information set. Here, the filtering process can more quickly and accurately locate the data that needs to be recovered, reduce the amount of data, and improve the data recovery efficiency. Subsequently, according to the above-mentioned filtered log event sequence, a data query statement information sequence is generated. Here, through the more accurate log event sequence after filtering, a more accurate data query statement information sequence can be generated, which is convenient for generating the subsequent data recovery execution file. Then, the above-mentioned data query statement information sequence is executed to recover the data set to be recovered corresponding to the above-mentioned data query statement information sequence. Here, the data recovery efficiency can be improved. Finally, the recovered data set to be recovered is stored in the verification data table in the preset storage database, a data recovery execution file is generated according to the above-mentioned data query statement information sequence, and a user operation execution file is generated according to the above-mentioned filtered log event sequence.Here, the obtained verification data table, data recovery execution file, and user operation execution file facilitate the user to determine whether the recovered data is the data the user wants, which can improve the user experience and reduce the waste of storage resources. Moreover, if it is the data the user wants, the data recovery execution file can be executed to improve the data recovery efficiency. Thus, the data recovery method can, in the case of the loss of the source database instance, parse and filter the local database log file through the user request information set, quickly locate the log event the user wants to recover, and the generated verification data table, data recovery execution file, and user operation execution file can quickly determine whether the recovered data is the data the user wants. Additionally, directly executing the data recovery execution file can quickly recover the data, improve the data recovery efficiency, reduce the waste of storage resources and communication resources, and improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0016] Figure 1 is a flowchart of some embodiments of the data recovery method according to the present disclosure;

[0017] Figure 2 is a schematic structural diagram of some embodiments of the data recovery apparatus according to the present disclosure;

[0018] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0020] It should also be noted that, for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0021] It should be noted that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0022] It should be noted that the modifications of "one" and "plural" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0023] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes, and are not used to limit the scope of these messages or information.

[0024] The following will detail this disclosure with reference to the drawings and in conjunction with embodiments.

[0025] Figure 1 Flow 100 of some embodiments of the data recovery method according to this disclosure is shown. The data recovery method includes the following steps:

[0026] Step 101, obtain a database log file sequence.

[0027] In some embodiments, the execution subject (such as an electronic device) of the above data recovery method can obtain the database log file sequence from the local file system through a wired connection or a wireless connection. Among them, the above database log file sequence can be a log file sequence sorted in chronological order. The database logs in the above database log file sequence can be files used to record change operations on the database. For example, the database log files in the above database log file sequence can be binlog (binary log) files. The format of the above binlog file is the row format.

[0028] Step 102, in response to receiving a user request information set, filter out at least one database log file from the database log file sequence that satisfies the request start time and request end time included in the user request information set, and obtain a target log file sequence.

[0029] In some embodiments, the above-mentioned execution entity may, in response to receiving a user request information set, screen out at least one database log file from the above-mentioned database log file sequence that meets the request start time and request end time included in the above-mentioned user request information set, and obtain a target log file sequence. Among them, the user request information in the above-mentioned user request information set may be the request information for the user to screen the above-mentioned database log file sequence. For example, the above-mentioned user request information set may include, but is not limited to, at least one of the following: the name of the starting database log file of the database log file sequence to be parsed, the start time and end time of the database log file sequence to be parsed, the data table name, and the type of operation on the database to be restored.

[0030] In some optional implementation manners of some embodiments, the above-mentioned step of, in response to receiving a user request information set, screening out at least one database log file from the above-mentioned database log file sequence that meets the request start time and request end time included in the above-mentioned user request information set, and obtaining a target log file sequence may include the following steps:

[0031] First, determine the log creation time of each database log file in the above-mentioned database log file sequence to obtain a log creation time sequence. Among them, the above-mentioned log creation time may be the generation time of the database log file. The above-mentioned log creation time may be a time format including year, month, day, hour, minute, and second. In practice, the above-mentioned execution entity may execute a log query statement for the database log file to determine the log creation time of each database log file in the above-mentioned database log file sequence and obtain a log creation time sequence.

[0032] Second, respectively compare the above-mentioned request start time and the above-mentioned request end time with the above-mentioned log creation time sequence to obtain a first comparison result sequence and a second comparison result sequence. Among them, the above-mentioned first comparison result sequence may be the result sequence obtained by comparing the request start time with the log creation time sequence. The above-mentioned second comparison result sequence may be the result sequence obtained by comparing the request end time with the log creation time sequence.

[0033] Third, select at least one first comparison result from the above-mentioned first comparison result sequence that represents the request start time being greater than or equal to the log creation time to obtain a first comparison result subsequence.

[0034] Fourth, determine the database log file corresponding to the first comparison result at the termination position in the above-mentioned first comparison result subsequence as the starting database log file.

[0035] Step 5: Select at least one second comparison result from the above second comparison result sequence that represents the request end time being greater than or equal to the log creation time, to obtain a second comparison result subsequence.

[0036] Step 6: Determine the database log file corresponding to the second comparison result at the termination position in the above second comparison result subsequence as the termination database log file.

[0037] Step 7: Determine the database log file subsequence corresponding to the log creation time subsequence within the time range of the log creation time corresponding to the above start database log file and the log creation time corresponding to the above termination database log file as the target log file sequence.

[0038] Step 103: Perform parsing processing on the target log file sequence to obtain a parsed log event group sequence.

[0039] In some embodiments, the above execution entity may perform parsing processing on the above target log file sequence to obtain a parsed log event group sequence. Among them, the parsed log events in the above parsed log event group sequence may be information recording the details of each modification operation on the database and information converting a binary file into a preset format. The above preset format may be a format that is preset and understandable by users. For example, the above preset format may be a text format. The parsed log events in the above parsed log event group sequence may include, but are not limited to, at least one of the following: the event operation type, operation object, operation time, and operation data for operating on the database. The above operation type may include: write operation type write_rows_event, modification operation type update_rows_event, deletion operation type delete_rows_event, and rows_query_event for recording SQL statements. The above operation object may be a data table in the database. The above operation data may be the data in the data table. The target log files in the above target log file sequence may include multiple parsed log events.

[0040] As an example, the above execution entity may parse the above target log file sequence by means of table lookup to obtain a parsed log event sequence.

[0041] Step 104: Filter the parsed log event group sequence according to the user request information subset to obtain a filtered log event sequence.

[0042] In some embodiments, the above-mentioned execution entity may filter the parsed log event sequence according to a subset of user request information to obtain a filtered log event group sequence. Among them, the above-mentioned subset of user request information may be a set obtained by removing the above-mentioned request start time and the above-mentioned request end time from the above-mentioned user request information set.

[0043] As an example, the above-mentioned execution entity may filter the parsed log event group sequence according to the data sets and data table names included in the subset of user request information to obtain a parsed log event sequence.

[0044] In some optional implementation manners of some embodiments, the above-mentioned filtering the parsed log event group sequence according to a subset of user request information to obtain a filtered log event sequence may include the following steps:

[0045] First step, extract the fields representing data tables in the above-mentioned parsed log event group sequence to obtain a data table key-value pair set. Among them, the data table key-value pairs in the above-mentioned data table key-value pair set may include: data table fields and data table attribute values corresponding to the data table fields. The above-mentioned data table attribute value may be a specific data table name.

[0046] Second step, perform a matching process on each data table attribute value in the data table attribute value set and each data table string in the data table string set included in the above-mentioned subset of user request information to generate a data table matching result group set. In practice, the above-mentioned matching process may be to use the KMP (Knuth-Morris-Pratt, string search algorithm) matching algorithm to perform a matching process on each data table attribute value in the data table attribute value set and each data table string in the data table string set included in the above-mentioned user request information set to generate a data table matching result group set.

[0047] Third step, screen out each parsed log event corresponding to at least one data table matching result representing successful data table matching from the above-mentioned data table matching result group set as a data table log event set.

[0048] Fourth step, sort the above-mentioned data table log event set to obtain a data table log event sequence. Among them, the above-mentioned sorting may be an ascending order sorting according to the operation time included in the data table log event set.

[0049] Fifth step, filter the above-mentioned data table log event sequence according to the operation event type string set included in the above-mentioned subset of user request information to obtain a filtered log event sequence.

[0050] In some alternative implementations of some embodiments, filtering the data table log event sequence according to the set of operation event type strings included in the user request information set to obtain a filtered log event sequence may include the following steps:

[0051] First, extract the fields representing the operation event type from the data table log event sequence to obtain a set of operation event type key-value pairs. Among them, the operation event type key-value pairs in the set of operation event type key-value pairs may include: an operation event type field and an operation event type attribute value corresponding to the operation event type field. The operation event type attribute value may be a specific event operation type. For example, the operation event type attribute value may be, but is not limited to, at least one of the following: write_rows_event, update_rows_event, delete_rows_event, and rows_query_event.

[0052] Second, perform a matching process on each operation event type attribute value in the set of operation event type attribute values and each operation event type string in the set of operation event type strings to generate a set of operation type matching result groups.

[0053] Third, screen out each data table log event corresponding to at least one operation type matching result indicating successful operation type matching from the set of operation type matching result groups as the filtered log event set.

[0054] Fourth, sort the filtered log event set to obtain a filtered log event sequence. Among them, the sorting may be an ascending order sorting of the operation events included in each data table log event.

[0055] Step 105, generate a data query statement information sequence according to the filtered log event sequence.

[0056] In some embodiments, the above-mentioned execution entity may generate a data query statement information sequence according to the above-mentioned filtered log event sequence. Among them, the data query statement information in the above-mentioned data query statement information sequence may be an SQL statement for restoring the data to be restored. The above-mentioned data query statement information sequence may be obtained by sorting the data query statement information set according to the operation time included in the filtered log event. As an example, the above-mentioned execution entity may first determine the event operation type of the above-mentioned filtered log event sequence to obtain an event operation type sequence. Secondly, extract the operation data included in the above-mentioned filtered log event sequence to obtain an operation data group sequence. Then, input the operation data group corresponding to the event operation type in the above-mentioned event operation type sequence and the above-mentioned operation data group sequence into the database query statement format to generate a target data query statement information sequence. Among them, the above-mentioned database query statement format may be the basic syntax statement of the database statement. Finally, perform inverse semantic processing on each target data query statement information in the above-mentioned target data query statement information sequence to obtain a data query statement information sequence. Among them, the above-mentioned inverse semantic processing may be an inverse operation processing on the semantics of the target data query statement information. For example, the above-mentioned target data query statement information may be to write data into a data table. The data query statement information after inverse semantic processing may be to delete the data written by the above-mentioned target data query statement information from the data table.

[0057] In some optional implementation manners of some embodiments, the above-mentioned generating a data query statement information sequence according to the above-mentioned filtered log event group sequence may include the following steps:

[0058] First step, determine the event type of each filtered log event in the above-mentioned filtered log event sequence to obtain an event type set. In practice, the above-mentioned execution entity may determine the event type of each filtered log event in the above-mentioned filtered log event sequence by querying the event operation type field included in the above-mentioned filtered log event to obtain an event type set.

[0059] Second step, according to the above-mentioned event type set, determine the statement type set of the data query statement information set. Among them, the statement types in the above-mentioned statement type set may be operation types for the database. The above-mentioned statement type set may include: delete type, modify type, query type, write type.

[0060] As an example, the above-mentioned execution entity may determine the statement operation field set of the above-mentioned data query statement information set through the above-mentioned event type set. Through the above-mentioned statement operation field set, determine the statement type set of the above-mentioned data query statement information set.

[0061] Third step, for each statement type in the above-mentioned statement type set, perform the following statement construction steps:

[0062] The first sub-step, in response to determining that the above statement type is a deletion-type statement type, constructs a statement for the filtered log event corresponding to the above statement type to obtain write data statement information. In practice, the above execution entity can first extract the deletion data set of the filtered log event corresponding to the above statement type information. Then, input the above statement type information and the extracted deletion data set into the write query statement format to obtain write data statement information. Among them, the above write query statement format can be the basic syntax statement of the write statement of the database.

[0063] The second sub-step, in response to determining that the above statement type is a write-type statement type, constructs a statement for the filtered log event corresponding to the above statement type to obtain delete data statement information. In practice, the above execution entity can first extract the write data set of the filtered log event corresponding to the above statement type information. Then, input the above statement type information and the extracted write data set into the delete query statement format to obtain delete data statement information. Among them, the above delete query statement format can be the basic syntax statement of the delete statement of the database.

[0064] The third sub-step, in response to determining that the above statement type is a modification-type statement type, constructs a statement for the filtered log event corresponding to the above statement type to obtain reverse modification data statement information. In practice, the above execution entity can first extract the modification data set of the filtered log event corresponding to the above statement type information. Then, swap the positions of each piece of pre-modification data in the pre-modification data set included in the above modification data set and the post-modification data corresponding to the pre-modification data in the post-modification data set to obtain the swapped modification data set. Among them, the above post-modification data can be the modification data after the set keyword of the modification query statement. The above pre-modification data can be the modification data after the where keyword of the modification query statement. The above pre-modification data and post-modification data can be the data sets corresponding to a modification query statement. Finally, input the above statement type information and the above swapped modification data set into the modification query statement format to generate reverse modification data statement information. Among them, the above modification query statement format can be the basic syntax statement of the modification statement of the database.

[0065] The fourth step, through the operation time series included in the above filtered log event sequence, combines and sorts the obtained write data statement information set, the obtained delete data statement information set, and the obtained reverse modification data statement information set to obtain a data query statement information sequence.

[0066] As an example, the above execution entity may first determine the set of operation times included in the filtered log event set corresponding to the above write data statement information set to obtain the write operation time set. Secondly, determine the set of operation times included in the filtered log event set corresponding to the above delete data statement information set to obtain the delete operation time set. Thirdly, determine the set of operation times included in the filtered log event set corresponding to the above reverse modification data statement information set to obtain the reverse modification operation time set. Then, perform a descending order sorting on the above write operation time set, delete operation time set, and reverse modification operation time set to obtain an operation time sequence, which is used as the statement operation time sequence. Finally, determine the data statement information corresponding to each statement operation time in the above statement operation time sequence to obtain the data query statement information sequence. Among them, the data statement information corresponding to the above statement operation time may be the data statement information in the write data statement information set, the delete data statement information set, and the reverse modification data statement information set.

[0067] Step 106: Execute the data query statement information sequence to restore the data set to be restored corresponding to the data query statement information sequence.

[0068] In some embodiments, the above execution entity may execute the above data query statement information sequence to restore the data set to be restored corresponding to the above data query statement information sequence.

[0069] In some optional implementation manners of some embodiments, the above execution of the data query statement information sequence to restore the data set to be restored corresponding to the data query statement information sequence may include the following steps:

[0070] First step: Perform syntactic word segmentation processing on each data query statement information in the above data query statement information sequence to generate a query word segmentation group to obtain a query word segmentation group sequence. Among them, the above execution entity may use a lexical analysis algorithm to perform syntactic word segmentation processing on each data query statement information in the above data query statement information sequence to obtain a query word segmentation group sequence.

[0071] Second step: Perform syntactic analysis on each query word segmentation group in the above query word segmentation group sequence to generate an abstract syntax tree to obtain an abstract syntax tree sequence. Among them, the above abstract syntax tree may represent the syntax tree formed by the execution order of the database query optimizer executing the data query statement. In practice, use the above syntax parser to perform syntactic analysis on each query word segmentation group in the above query word segmentation group sequence to generate an abstract syntax tree to obtain an abstract syntax tree sequence.

[0072] In the third step, perform data cleaning on the above abstract syntax tree sequence to obtain the cleaned abstract syntax tree sequence. Among them, the above data cleaning can be to replace the non-SQL keyword characters corresponding to the leaf nodes in the abstract syntax tree with predefined characters to reduce data redundancy. The above predefined characters can be preset characters. For example, the above predefined character can be a question mark character.

[0073] In the fourth step, perform feature extraction on the above cleaned abstract syntax tree sequence to obtain a query feature vector sequence. Among them, the query feature vectors in the above query feature vector sequence represent the feature information of the logical execution order of the data tables included in the query statement information. In practice, the above execution entity can use BOW (Bag Of Words) to perform feature extraction on the above cleaned abstract syntax tree sequence to obtain a query feature vector sequence.

[0074] In the fifth step, perform hash binary encoding on each query feature vector in the above query feature vector sequence to generate a binary encoding vector, and obtain a binary encoding vector sequence. Among them, the above binary encoding vector can be a hash representation vector that maps the above query feature vector to a hash value of a fixed length.

[0075] As an example, the above execution entity can use the principal component analysis method to perform data dimensionality reduction on each query feature vector group in the above query feature vector sequence to obtain a dimensionality-reduced feature vector sequence. Then, use the equal variance hashing method to perform hash binary encoding on the above dimensionality-reduced feature vector sequence to generate a binary encoding vector, and obtain a binary encoding vector sequence.

[0076] In the sixth step, perform anomaly detection on the above binary encoding vector sequence to obtain an anomaly detection result sequence. In practice, the above execution entity can use the Hamming sorting algorithm and hash table query to perform anomaly detection on the above binary encoding vector sequence to obtain an anomaly detection result set.

[0077] In the seventh step, determine the information of each data query statement corresponding to at least one anomaly detection result indicating an abnormal detection result in the above anomaly detection result sequence as an abnormal query statement information set.

[0078] In the eighth step, screen out at least one abnormal query statement information that meets the preset conditions from the above abnormal query statement information set to obtain a target query statement information set. Among them, the above preset condition can be a condition that the abnormal query statement affects the query performance of the database.

[0079] In the ninth step, query optimization is performed on the above-mentioned target query statement information set to obtain an optimized query statement information set. Among them, the optimized query statement information in the above-mentioned optimized query statement information set can be query statement information that changes the execution logic order of the data tables included in the query statement information. In practice, the above-mentioned execution entity can perform query optimization on the above-mentioned target query statement information by methods such as merging multiple query conditions, optimizing join operations, and eliminating redundant subqueries to obtain an optimized query statement information set.

[0080] In the tenth step, the above-mentioned target query statement information set is removed from the above-mentioned abnormal query statement information set to obtain a query statement information set after removal.

[0081] In the eleventh step, the above-mentioned query statement information set after removal is modified to obtain a modified query statement information set.

[0082] In practice, for each query statement information after removal in the above-mentioned query statement information set after removal, the above-mentioned execution entity can perform the following modification steps: First, determine the abnormal type of the above-mentioned query statement information after removal. Then, in response to determining that the abnormal type is a semantic abnormality, format the above-mentioned query statement information after removal to obtain formatted query statement information, which is used as the modified query statement information. In response to determining that the abnormal type is a permission abnormality, add a user permission restriction statement to the above-mentioned query statement information after removal to obtain an added query statement, which is used as the modified query statement information set.

[0083] In the twelfth step, each data query statement information corresponding to at least one anomaly detection result indicating detection pass in the above-mentioned anomaly detection result sequence, the above-mentioned optimized query statement information set, and the above-mentioned modified query statement information set are determined as the data query statement information set to be restored.

[0084] In the thirteenth step, the above-mentioned data query statement information set to be restored is executed to restore the data set to be restored.

[0085] The above first step to the thirteenth step and their related content are an inventive point of the embodiment of the present disclosure, which solves the second technical problem mentioned in the background art: "For the existing method of executing reverse SQL statements, anomaly detection based on machine learning is performed first and then SQL statements are executed. Since machine learning requires a large amount of high-quality data for training and calculates the similarity between each query statement and other query statements, the detection efficiency is low, and a large amount of storage resources and computing resources are required. In the case of less training data or low-quality training data, false positives and false negatives are likely to occur, resulting in the leakage of database data and the damage of the database server." The factors leading to the leakage of database data and the damage of the database server are often as follows: For the existing method of executing reverse SQL statements, anomaly detection based on machine learning is performed first and then SQL statements are executed. Since machine learning requires a large amount of high-quality data for training and calculates the similarity between each query statement and other query statements, the detection efficiency is low, and a large amount of storage resources and computing resources are required. In the case of less training data or low-quality training data, false positives and false negatives are likely to occur. If the above factors are solved, the effect of reducing the leakage of database data and the damage of the database server can be achieved. To achieve this effect, the present disclosure first performs word segmentation, syntax analysis, and data cleaning on the data query statement information sequence, which can reduce the number of similar data query statements, reduce redundant data, and reduce the waste of computing resources. Secondly, feature extraction and anomaly detection based on the hash algorithm are performed on the cleaned data query statements. By mapping the cleaned data query statements to hash values and determining abnormal data by comparing the hash values, the detection efficiency and detection accuracy of anomaly detection can be improved, and the situation of missed detection and false detection can be reduced. Then, modifying and optimizing the data query statements of different abnormal types can improve the accuracy of modifying abnormal data query statements and improve the security of the modified data query statements. Finally, executing the information set of the data query statements to be restored to restore the data set to be restored can improve the security and stability of the database server, and reduce data leakage and the damage of the database service.

[0086] Step 107: Store the restored data set to be restored in the verification data table in the preset storage database, generate a data recovery execution file according to the data query statement information sequence, and generate a user operation execution file according to the filtered log event sequence.

[0087] In some embodiments, the above-mentioned execution entity may store the restored dataset to be restored in the verification data table in the preset storage database, generate a data recovery execution file according to the above data query statement information sequence, and generate a user operation execution file according to the above filtered log event sequence. Among them, the above-mentioned preset storage database may be a database for storing the restored dataset to be restored. For example, the above-mentioned preset storage database may be a MySQL database. The above-mentioned verification data table may be a data table for storing the restored dataset to be restored. The above-mentioned user operation execution file may be a file formed by the SQL statement information of the user operation database corresponding to the data to be restored. The above-mentioned data recovery execution file may be a file obtained by storing the data query statement information sequence in a text file.

[0088] As an example, the above-mentioned execution entity may first store the above data query statement information sequence in a text file to obtain a data recovery execution file. Secondly, in response to determining that there is a log event of a preset event type in the above database log file sequence, determine at least one log event corresponding to the preset event type as the target log event sequence. Among them, the above-mentioned preset event type may be rows_query_event. Thirdly, extract the user operation data query statements included in the above target log event sequence to obtain a user operation data query statement sequence. Then, determine at least one user operation data query statement corresponding to the above filtered log event sequence as the target user operation data query statement sequence. Finally, store the above target user operation data query statement sequence in a text file to obtain a user operation execution file.

[0089] In some optional implementation manners of some embodiments, the above-mentioned generation of the data recovery execution file according to the above data query statement information sequence may include the following steps:

[0090] First step, classify the above data query statement information sequence according to the query data table fields and data fields included in the data query statement information to obtain a classified query statement information group sequence. Among them, the above-mentioned data fields may be the fields corresponding to each column of data included in the data table. The classified query statement information included in the above classified query statement information group may be the data query statement information with the same query data table fields and data fields.

[0091] As an example, select at least one data query statement information group with the same data table fields and data fields from the above data query statement information sequence as the classified query statement information group sequence.

[0092] Step 2: For each classified query statement information group in the above sequence of classified query statement information groups, determine the classified query statement information at the initial position in the above classified query statement information group as the initial query statement information, and delete each classified query statement information in the above classified query statement information group except the above initial query statement information.

[0093] Step 3: Sort the above obtained initial query statement information to obtain an initial query statement information sequence. Among them, the above sorting can be an ascending order sorting based on the log operation time included in the initial query statement information.

[0094] Step 4: Determine the log operation time, event operation type, data query statement, and log offset of each initial query statement information in the above initial query statement information sequence to obtain a log operation time sequence, an event operation type sequence, a data query statement sequence, and a log offset sequence. Among them, the above log offset can represent the number of bytes from the start position of the binary file to the position of the initial query statement information where the initial query statement information is located in the binary file. In practice, the above execution entity can first determine the parsed log event corresponding to each initial query statement information in the above initial query statement information sequence as the initial log event to obtain an initial log event sequence. Then, determine the log operation time, event operation type, and log offset of each initial log event in the above initial log event sequence to obtain a log operation time sequence, an event operation type sequence, and a log offset sequence.

[0095] Step 5: Store the above log operation time sequence, the above event operation type sequence, the above data query statement sequence, and the above log offset sequence in a predetermined format to obtain a data recovery execution file. Among them, the above predetermined format can be a storage format in the order of operation time, event operation type, data query statement, and log offset.

[0096] Optionally, after 107, the above execution entity can also perform the following steps:

[0097] Step 1: Obtain a historical query statement information set and a historical data recovery information set. Among them, the historical query statement information in the above historical query statement information set includes an access data table group. The historical data recovery information in the above historical data recovery information set includes a recovery data table group. The above historical query statement information can be an SQL statement for performing corresponding operations on the data in the database before the current time. The above historical data recovery information can be information for performing a data recovery operation on the data in the database before the current time.

[0098] Step 2: Determine the access frequency of each access data table in the access data table set included in the above historical query statement information to obtain an access frequency set. Among them, the above access frequency can represent the degree of operation of the user on the data tables in the database. The access data tables in the above access data table set can be the data tables included in the historical query statement information. In practice, the above execution entity can count the number of queries of each access data table in the access data table set included in the above historical query statement information as the access frequency to obtain an access frequency set.

[0099] Step 3: Determine the recovery frequency of each recovery data table in the recovery data table set included in the above historical data recovery information set to obtain a recovery frequency set. Among them, the above recovery frequency can represent the number of times the recovery data table appears in the historical data recovery information set. The recovery data tables in the above recovery data table set can be the data tables included in the historical data recovery information. In practice, the above execution entity can count the number of occurrences of each recovery data table in the recovery data table set included in the above historical data recovery information set as the recovery frequency to obtain a recovery frequency set.

[0100] Step 4: Perform a deduplication process on the above access data table set and the above recovery data table set to obtain a deduplicated data table set.

[0101] Step 5: According to the above access frequency set and the above recovery frequency set, determine the loss degree value of each deduplicated data table in the above deduplicated data table set to obtain a loss degree value set. Among them, the above loss degree value can represent the degree of loss and recovery of the deduplicated data table.

[0102] As an example, the above execution entity can perform the following determination steps for each deduplicated data table in the above deduplicated data table set: First, determine the access frequency and recovery frequency of each deduplicated data table in the above deduplicated data table set as the target access frequency and the target recovery frequency. Then, determine the sum of the product of the above target access frequency and the first preset weight and the product of the above target recovery frequency and the second preset weight as the loss degree value. Among them, the sum of the above first preset weight and the second preset weight is 1. The above first preset weight and the second preset weight can be preset weight values, and the specific values can be determined according to specific situations.

[0103] Step 6: Screen out a preset number of loss degree values from the above loss degree value set as the target loss degree value sequence. Among them, the above preset number can be a preset number. For example, the above preset number can be 50. The above target loss degree value sequence can be the first preset number of loss degree values selected after sorting the above loss degree values from large to small.

[0104] Step 7: Determine the association relationship information set among the deduplicated data tables included in the deduplicated data table set corresponding to the above target loss degree value sequence. Among them, the association relationship information in the above association relationship information set can be the relationship between the primary key and the foreign key of the deduplicated data table. In practice, the above execution entity can determine the association relationship information set among the deduplicated data tables included in the deduplicated data table set corresponding to the above target loss degree value sequence through the primary key and the foreign key of each deduplicated data table.

[0105] Step 8: According to the above association relationship information set, perform data deduplication processing on the loss data set included in the above loss data table set to obtain the deduplicated loss data set.

[0106] As an example, the above execution entity can first, through the above association relationship information set, determine the association fields and association data between each loss data table in the above loss data table set and each loss data table other than the loss data table in the above loss data table set, to obtain the association field set and the association data set. Secondly, perform deduplication processing on the above association field set and the above association data set to obtain the deduplicated association fields and the deduplicated association data set. Then, remove the association field subset and the association data subset corresponding to the loss data table from each loss data table in the above loss data table set to obtain the loss data table removal set. Finally, determine the above loss data table removal set, the above deduplicated association fields, and the above deduplicated association data set as the deduplicated loss data set.

[0107] Step 9: Perform a hashing operation on the above association relationship information set and the above deduplicated loss data set to obtain a data hash value set. Among them, the data hash values in the above data hash value set can be the data hash value set obtained by substituting the above association relationship information set and the above deduplicated loss data set into the hashing algorithm. The above hashing algorithm can be one of the MD5 (Message-Digest Algorithm 5) algorithm, the SHA-3 (Secure Hash Algorithm 3) algorithm, the SHA-512 (Secure Hash Algorithm 512) algorithm, and the RIPEMD160 (RACE Integrity Primitives Evaluation Message Digest) algorithm.

[0108] Step 10: Determine the block node address information set of each data hash value in the above data hash value set. Among them, the block node address information in the above block node address information set can be the address information of the blocks in the blockchain for backup. In practice, the above execution entity can use the consistent hashing algorithm to determine the block node address information set of each data hash value in the above data hash value set.

[0109] In the eleventh step, send the above-mentioned associated relationship information set and the deduplicated loss data set to the data storage layer of the block nodes corresponding to the above-mentioned block node address information set for storage, and send the above-mentioned data hash value to the network consensus layer of each block node in the block node set corresponding to the above-mentioned block node address information set for storage.

[0110] The above steps from the first step to the eleventh step and their related content are an inventive point of the embodiments of the present disclosure, which solve the third technical problem mentioned in the background technology: "The existing technology for storing the data to be restored after restoration directly performs a full backup on the restored data, which may result in that the data currently fully backed up is not frequently accessed, and there are a large amount of redundant data that users do not need in the full backup, leading to a waste of storage resources and a low user experience." The factors that lead to a waste of storage resources and a low user experience are usually as follows: The existing technology for storing the data to be restored after restoration directly performs a full backup on the restored data, which may result in that the data currently fully backed up is not frequently accessed, and there are a large amount of redundant data that users do not need in the full backup. If the above factors are solved, the effect of reducing the waste of storage resources and improving the user experience can be achieved. To achieve this effect, the present disclosure first determines the access frequency and restoration frequency of the data in each data table included in the database through the historical query statement information and the historical data restoration information set. Secondly, different weights are set for the access frequency and restoration frequency to determine the degree of frequent access and restoration of each data table, and the data loss degree can be obtained, which can reduce the storage of redundant data in subsequent backups and reduce the waste of storage resources. Then, determine the associated relationship between the data tables included in the data table set with a higher data loss degree, and further deduplicate the redundant data through the associated relationship to reduce the existence of redundant data. Finally, perform a hash operation on the data tables with a higher data loss degree, and determine the corresponding block storage layer and consensus layer, which can reduce the amount of backup data, reduce the blockchain storage resources, and improve the recognition efficiency of the consensus algorithm of the blockchain, improving the backup efficiency and data security.

[0111] The above embodiments of the present disclosure have the following beneficial effects: In some embodiments of the present disclosure, in the case of the loss of the source database instance, the data recovery method can recover data by parsing the local database log file, which can improve the data recovery efficiency, reduce the waste of storage resources, and improve the user experience. Specifically, the reasons for the waste of relevant database storage resources, low database performance, and low user experience are as follows: Since restoring the database log file sequence into SQL statements and performing semantic inverse operations on the SQL statements is to perform step-by-step inverse operations on the data in the database until the data that the user wants to recover is traced back. It is possible that some data is data that the user does not want to recover, which will generate redundant data, as well as low data recovery efficiency and long recovery time, resulting in waste of database storage resources, low database performance, and low user experience. Based on this, the data recovery method of some embodiments of the present disclosure can first obtain the database log file sequence, where the above database log file sequence is a log file sequence sorted according to the corresponding time order. Here, the database log file sequence is the data log file sequence in the locally called file, which can reduce the waste of communication resources. Secondly, in response to receiving the user request information set, at least one database log file that meets the request start time and request end time included in the above user request information set is filtered out from the above database log file sequence to obtain the target log file sequence. Here, the amount of data to be recovered can be reduced, and the waste of computing resources can be reduced. Thirdly, the above target log file sequence is parsed to obtain the parsed log event group sequence. Here, since the parsing is performed on the local target log file sequence, the number of target log files to be parsed can be reduced, and the data integrity and parsing efficiency can be improved. Then, according to the user request information subset, the above parsed log event group sequence is filtered to obtain the filtered log event sequence, where the above user request information subset is a set obtained by removing the above request start time and the above request end time from the above user request information set. Here, the filtering process can more quickly and accurately locate the data that needs to be recovered, reduce the amount of data, and improve the data recovery efficiency. Subsequently, according to the above filtered log event sequence, a data query statement information sequence is generated. Here, through the more accurate log event sequence after filtering, a more accurate data query statement information sequence can be generated, and it is convenient to generate a data recovery execution file subsequently. Then, the above data query statement information sequence is executed to recover the data set to be recovered corresponding to the above data query statement information sequence. Here, the data recovery efficiency can be improved. Finally, the recovered data set to be recovered is stored in the verification data table in the preset storage database, and according to the above data query statement information sequence, a data recovery execution file is generated, and according to the above filtered log event sequence, a user operation execution file is generated.Here, the obtained verification data table, data recovery execution file, and user operation execution file facilitate the user to determine whether the restored data is the data the user wants, which can improve the user experience and reduce the waste of storage resources. Moreover, if it is the data the user wants, the data recovery execution file can be executed to improve the data recovery efficiency. Thus, the data recovery method can, in the case of the loss of the source database instance, parse and filter the local database log file through the user request information set, quickly locate the log event the user wants to recover, and generate the verification data table, data recovery execution file, and user operation execution file, which can quickly determine whether the restored data is the data the user wants, directly execute the data recovery execution file, quickly recover the data, improve the data recovery efficiency, reduce the waste of storage resources and communication resources, and improve the user experience.

[0112] Further reference is made to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a data recovery device. These device embodiments correspond to Figure 1 the method embodiments shown, and the data recovery device can be specifically applied to various electronic devices.

[0113] As Figure 2As shown in the figure, a data recovery device 200 includes: an acquisition unit 201, a screening unit 202, an analysis processing unit 203, a filtering processing unit 204, a generation unit 205, an execution unit 206, and a storage unit 207. Among them, the acquisition unit 201 is configured to: acquire a database log file sequence, where the database log file sequence is a log file sequence sorted according to the corresponding time sequence. The screening unit 202 is configured to: in response to receiving a user request information set, screen out at least one database log file from the database log file sequence that satisfies the request start time and request end time included in the user request information set, and obtain a target log file sequence. The analysis processing unit 203 is configured to: perform analysis processing on the target log file sequence to obtain an analyzed log event group sequence. The filtering processing unit 204 is configured to: perform filtering processing on the analyzed log event group sequence according to a user request information subset, and obtain a filtered log event sequence, where the user request information subset is a set obtained by removing the request start time and the request end time from the user request information set. The generation unit 205 is configured to: generate a data query statement information sequence according to the filtered log event sequence. The execution unit 206 is configured to: execute the data query statement information sequence to recover a data set to be recovered corresponding to the data query statement information sequence. The storage unit 207 is configured to: store the recovered data set to be recovered in a verification data table in a preset storage database, generate a data recovery execution file according to the data query statement information sequence, and generate a user operation execution file according to the filtered log event sequence.

[0114] It can be understood that the various units described in the data recovery device 200 correspond to the respective steps in the method described in the reference Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the data recovery device 200 and the units included therein, and will not be elaborated here.

[0115] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device (for example, an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0116] As Figure 3As shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0117] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 3 an electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 3 Each block shown in may represent a device or, as needed, multiple devices.

[0118] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are executed.

[0119] It should be noted that, in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0120] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0121] The above computer-readable medium may be included in the above electronic device; or it may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a database log file sequence, where the database log file sequence is a log file sequence sorted according to the corresponding chronological order; in response to receiving a user request information set, filter out at least one database log file from the database log file sequence that satisfies the request start time and the request end time included in the user request information set to obtain a target log file sequence; perform parsing processing on the target log file sequence to obtain a parsed log event group sequence; filter the parsed log event group sequence according to a user request information subset to obtain a filtered log event sequence, where the user request information subset is a set obtained by removing the request start time and the request end time from the user request information set; generate a data query statement information sequence according to the filtered log event sequence; execute the data query statement information sequence to restore a data set to be restored corresponding to the data query statement information sequence; store the restored data set to be restored in a verification data table in a preset storage database, generate a data recovery execution file according to the data query statement information sequence, and generate a user operation execution file according to the filtered log event sequence.

[0122] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0124] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a screening unit, an analysis and processing unit, a filtering and processing unit, a generation unit, an execution unit, and a storage unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring the sequence of database log files".

[0125] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0126] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A data recovery method, including: obtaining a database log file sequence, where the database log file sequence is a log file sequence sorted in the corresponding chronological order; in response to receiving a user request information set, screening at least one database log file that meets the request start time and request end time included in the user request information set from the database log file sequence to obtain a target log file sequence; performing parsing processing on the target log file sequence to obtain a parsed log event group sequence; filtering the parsed log event group sequence according to a user request information subset to obtain a filtered log event sequence, where the user request information subset is a set obtained by removing the request start time and the request end time from the user request information set; generating a data query statement information sequence according to the filtered log event sequence; executing the data query statement information sequence to recover a data set to be recovered corresponding to the data query statement information sequence; storing the recovered data set to be recovered in a verification data table in a preset storage database, generating a data recovery execution file according to the data query statement information sequence, and generating a user operation execution file according to the filtered log event sequence.

2. The method according to claim 1, wherein, the generating a data query statement information sequence according to the filtered log event sequence includes: determining the event type of each filtered log event in the filtered log event sequence to obtain an event type set; determining a statement type set of a data query statement information set according to the event type set; for each statement type in the statement type set, performing the following statement construction steps: in response to determining that the statement type is a delete type statement type, performing statement construction on the filtered log event corresponding to the statement type to obtain write data statement information; in response to determining that the statement type is a write type statement type, performing statement construction on the filtered log event corresponding to the statement type to obtain delete data statement information; in response to determining that the statement type is a modification type statement type, performing statement construction on the filtered log event corresponding to the statement type to obtain reverse modification data statement information; combining and sorting the obtained write data statement information set, the obtained delete data statement information set, and the obtained reverse modification data statement information set through an operation time sequence included in the filtered log event sequence to obtain a data query statement information sequence.

3. The method according to claim 1, wherein, the screening at least one database log file that meets the request start time and request end time included in the user request information set from the database log file sequence in response to receiving the user request information set to obtain a target log file sequence includes: determining the log creation time of each database log file in the database log file sequence to obtain a log creation time sequence; Compare the request start time and the request end time with the log creation time series respectively to obtain a first comparison result series and a second comparison result series; Select at least one first comparison result indicating that the request start time is greater than or equal to the log creation time from the first comparison result series to obtain a first comparison result subsequence; Determine the database log file corresponding to the first comparison result at the termination position in the first comparison result subsequence as the start database log file; Select at least one second comparison result indicating that the request end time is greater than or equal to the log creation time from the second comparison result series to obtain a second comparison result subsequence; Determine the database log file corresponding to the second comparison result at the termination position in the second comparison result subsequence as the termination database log file; Determine the database log file subsequence corresponding to the log creation time subsequence within the time range of the log creation time corresponding to the start database log file and the log creation time corresponding to the termination database log file as the target log file sequence.

4. The method according to claim 1, wherein, The filtering the parsed log event group sequence according to the user request information subset to obtain a filtered log event sequence includes: Extracting the fields representing the data table in the parsed log event group sequence to obtain a data table key-value pair set, wherein the data table key-value pairs in the data table key-value pair set include: data table fields and data table attribute values corresponding to the data table fields; Performing a matching process on each data table attribute value in the data table attribute value set and each data table string in the data table string set included in the user request information subset to generate a data table matching result group to obtain a data table matching result group set; Selecting each parsed log event corresponding to at least one data table matching result indicating successful data table matching from the data table matching result group set as the data table log event set; Sorting the data table log event set to obtain a data table log event sequence; Filtering the data table log event sequence according to the operation event type string set included in the user request information subset to obtain a filtered log event sequence.

5. The method according to claim 4, wherein, The filtering the data table log event sequence according to the operation event type string set included in the user request information subset to obtain a filtered log event sequence includes: Extracting the fields representing the operation event type in the data table log event sequence to obtain an operation event type key-value pair set, wherein the operation event type key-value pairs in the operation event type key-value pair set include: operation event type fields and operation event type attribute values corresponding to the operation event type fields; Performing a matching process on each operation event type attribute value in the operation event type attribute value set and each operation event type string in the operation event type string set to generate an operation type matching result group to obtain an operation type matching result group set; Filter out each data table log event corresponding to at least one operation type matching result indicating successful operation type matching from the set of operation type matching result sets as the filtered log event set; Sort the filtered log event set to obtain a filtered log event sequence.

6. The method according to claim 1, wherein, generating a data recovery execution file according to the data query statement information sequence includes: Classify the data query statement information sequence according to the query data table fields and data fields included in the data query statement information to obtain a classified query statement information group sequence; For each classified query statement information group in the classified query statement information group sequence, determine the classified query statement information at the initial position in the classified query statement information group as the initial query statement information, and delete each classified query statement information in the classified query statement information group except the initial query statement information; Sort the obtained initial query statement information to obtain an initial query statement information sequence; Determine the log operation time, event operation type, data query statement, and log offset of each initial query statement information in the initial query statement information sequence to obtain a log operation time sequence, an event operation type sequence, a data query statement sequence, and a log offset sequence; Store the log operation time sequence, the event operation type sequence, the data query statement sequence, and the log offset sequence in a predetermined format to obtain a data recovery execution file.

7. A data recovery device, comprising: An acquisition unit configured to acquire a database log file sequence, where the database log file sequence is a log file sequence sorted in the corresponding time order; A screening unit configured to, in response to receiving a user request information set, screen out at least one database log file from the database log file sequence that satisfies the request start time and request end time included in the user request information set to obtain a target log file sequence; An analysis processing unit configured to perform analysis processing on the target log file sequence to obtain an analyzed log event group sequence; A filtering processing unit configured to perform filtering processing on the analyzed log event group sequence according to a user request information subset to obtain a filtered log event sequence, where the user request information subset is a set obtained by removing the request start time and the request end time from the user request information set; A generating unit configured to generate a data query statement information sequence according to the filtered log event sequence; An execution unit configured to execute the data query statement information sequence to recover the data set to be recovered corresponding to the data query statement information sequence; A storage unit configured to store the recovered data set to be recovered in a verification data table in a preset storage database, generate a data recovery execution file according to the data query statement information sequence, and generate a user operation execution file according to the filtered log event sequence.

8. An electronic device, comprising: One or more processors; A storage device storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

9. A computer-readable medium storing a computer program, wherein, When the computer program is executed by a processor, it implements the method according to any one of claims 1-6.