Spare library repairing method and device, electronic equipment, medium and computer program product
By automatically matching keywords in the library error log to obtain bad table information, and stopping the data interaction of the main and backup database based on this information, generating files to be backed up to repair the bad table, the automation missing problem of single table bad block detection and repair of backup databases in the existing technology is solved, improving repair efficiency and reducing business risks.
Patent Information
- Application Number
- CN202411865451.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-06
AI Technical Summary
The existing technology lacks a mechanism for automatically detecting and repairing single-table bad blocks of database backup databases, resulting in delays in library copying and application impacts.
By obtaining the preset keywords and the backup library error log, the bad table information is automatically obtained, the bad tables need to be repaired, the main and backup library data interaction is stopped, the files to be backed up are generated synchronously, and the bad tables are repaired based on these files.
It realizes automatic detection and repair of bad tables in the backup library, shortens the detection and repair time of bad tables, and reduces the impact on database business and the risk of data loss.
Smart Images

Figure CN119938399A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of database technology, and in particular relates to a backup database repair method, device, electronic device, medium and computer program product. Background Art
[0002] Currently, there is a lack of a mechanism to detect single-table corruption in the standby database. The detection of single-table corruption in the standby database is mainly achieved through passive discovery methods such as maintenance personnel based on Structured Query Language (SQL) thread error information or query errors, which affects the overall application of the database. Summary of the invention
[0003] Embodiments of the present application provide a backup database repair method, device, electronic device, medium, and computer program product.
[0004] The present application provides a method for repairing a standby database, the method comprising:
[0005] Obtain preset keywords, match the keywords with the error log of the standby database, and automatically obtain information about the bad table; wherein the bad table indicates a damaged table in the standby database; the standby database is used to back up data in the main database;
[0006] Determining a first bad table that needs to be repaired according to the bad table information;
[0007] Stop the data interaction between the standby database and the primary database, and synchronously generate the to-be-backed-up files corresponding to the primary database;
[0008] The first bad table is repaired based on the file to be backed up.
[0009] In some embodiments, before determining the first bad table that needs to be repaired based on the information of the bad table, the method also includes: automatically reporting the information of the bad table to the target storage space; defining a structure; the structure is used to determine preset information that needs to be obtained; the preset information is a part of the information of the bad table; determining the first bad table that needs to be repaired based on the information of the bad table includes: obtaining preset information in the target storage space based on the structure, and determining the first bad table that needs to be repaired based on the preset information.
[0010] It can be seen that by storing the automatically acquired bad table information in the target storage space and acquiring preset information in the target storage space based on the defined structure, it is helpful to automatically acquire the bad table information that meets the requirements and repair the bad table.
[0011] In some embodiments, before obtaining preset information in the target storage space based on the structure, the method further includes: obtaining, in real time, information of bad tables reported to the target storage space based on a controller in a Kubernetes cluster.
[0012] It can be seen that obtaining the bad table information of the target storage space in real time through the controller helps to obtain the reported bad table information in time and realize timely repair of the bad table.
[0013] In some embodiments, the first bad table represents a single table having bad blocks in the standby database, and the repairing the first bad table based on the file to be backed up includes: determining a first identifier corresponding to the primary database and a second identifier corresponding to the first bad table; wherein the first identifier represents an identifier of a latest transaction corresponding to the file to be backed up, and the second identifier represents an identifier of a latest transaction received by the first bad table before the moment of stopping data interaction between the standby database and the primary database; and repairing the first bad table based on the first identifier, the second identifier, and the file to be backed up.
[0014] It can be seen that after stopping the data interaction between the main database and the standby database, the difference between the file to be backed up currently exported in the main database and the file to be backed up received by the first bad table of the standby database before stopping the data interaction can be obtained through the first identifier and the second identifier. Therefore, the first bad table is repaired through the first identifier, the second identifier, and the file to be backed up, which is conducive to ensuring the consistency of the data to be backed up in the first bad table and the data in the file to be backed up exported by the main database.
[0015] In some embodiments, the repairing of the first bad table based on the first identifier, the second identifier, and the file to be backed up includes: importing the file to be backed up in the backup database to obtain a first backup file; determining the first identifier as a start point for rollback, and determining the second identifier as an end point for rollback; based on the start point and the end point, performing a rollback operation on the first backup file, and using the first backup file after the rollback operation as data in the repaired first bad table.
[0016] It can be seen that rolling back the first backup file through the first identifier and the second identifier is conducive to ensuring the consistency of the first backup file used for repair and the backup file originally executed in the first bad table, which is conducive to reducing the error probability of repairing the first bad table in the standby database and realizing rapid repair of the first bad table.
[0017] In some embodiments, before performing a rollback operation on the first backup file based on the starting point and the end point, the method further includes: determining a third identifier; the third identifier represents an identifier of a transaction completed by the first bad table backup; performing a rollback operation on the first backup file based on the starting point and the end point includes: when the second identifier is the same as the third identifier, performing a rollback operation on the first backup file based on the starting point and the end point.
[0018] It can be seen that by rolling back the first backup file when the second identifier is the same as the third identifier, it can be ensured that the current backup task is completed in the first bad table before repairing it, which is beneficial to ensure the consistency of the standby database data, reduce the probability of database standby database errors, and improve the efficiency of repairing the first bad table.
[0019] In some embodiments, before repairing the first bad table based on the first identifier, the second identifier, and the file to be backed up, the method further includes: stopping log generation in the standby database.
[0020] It can be seen that before repairing the first bad table, stopping the standby database from generating logs, that is, not recording the repair operation in the standby database log, helps to achieve consistency of the logs before and after the repair operation in the standby database after the repair is completed, thereby reducing the probability of database errors.
[0021] In some embodiments, matching the keyword with the error log of the standby database includes: matching the error log of the standby database with the keyword based on a regular detection algorithm.
[0022] It can be seen that the use of regular detection algorithms to detect error logs of the standby database helps to improve the accuracy of information extraction of bad tables in the error logs of the standby database.
[0023] The embodiment of the present application also provides a standby database repair device, the device comprising:
[0024] An acquisition module is used to acquire preset keywords, match the keywords with the error log of the standby database, and automatically acquire information about the bad table; wherein the bad table indicates a damaged table in the standby database; the standby database is used to back up data in the main database;
[0025] A determination module, used for determining a first bad table that needs to be repaired according to the information of the bad table;
[0026] The processing module is used to stop the data interaction between the standby database and the main database, and synchronously generate a to-be-backed-up file corresponding to the main database; and repair the first bad table based on the to-be-backed-up file.
[0027] An embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory for storing a computer program that can be run on the processor; wherein:
[0028] The processor is used to run the computer program to execute any one of the above-mentioned standby database repair methods.
[0029] An embodiment of the present application provides a computer storage medium on which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned standby database repair methods is implemented.
[0030] An embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the computer program implements any of the above-mentioned standby database repair methods.
[0031] The embodiments of the present application provide a backup database repair method, device, electronic device, medium and computer program product. Through the backup database repair method provided in the embodiments of the present application, it is possible to automatically obtain information about bad tables in the backup database, which helps to repair the bad tables in a timely manner through the bad table information, improve the delay problem caused by the bad tables in the backup database, and help reduce the risk of database business interruption. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A flow chart of a backup database repair method provided in an embodiment of the present application;
[0033] Figure 2 A schematic diagram of a complete process of repairing a single table bad block in a standby database provided in an embodiment of the present application;
[0034] Figure 3 A schematic diagram of the structure of a standby database repair device provided in an embodiment of the present application;
[0035] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] In a database system, the master and standby databases play different roles and work together to ensure high availability and consistency of data. The master database is the core component in the database system and is responsible for processing all write operations and updates. The standby database is a copy of the master database and is used for read operations or data backup. The master and standby databases of the database work together to ensure high availability and consistency of data. While the master database processes write operations, the standby database maintains data consistency with the master database through a replication mechanism and takes over the work of the master database when needed. This architecture plays an important role in improving database performance and reliability. Single-table corruption in the standby database is a problem that may be encountered in database operation and maintenance. Single-table corruption in the standby database means that in the storage area occupied by a specific table in the standby database of the database, there are one or more blocks that cannot read or write data normally due to physical damage or logical errors. These blocks are called "bad blocks". Single-table corruption in the standby database may be caused by a variety of reasons, including but not limited to physical damage to the disk, errors in the file system, defects in the database management system, and operational errors.
[0037] Currently, there is a lack of a mechanism to detect single-table corruption in the standby database. The discovery of single-table corruption in the standby database is mainly achieved through passive discovery methods such as SQL thread error information or query error reporting by maintenance personnel, which causes replication delays in the standby database and affects the standby database application. The general method for handling single-table corruption in the standby database is to redo the standby database, that is, to eliminate the original standby database, then back up the data in full from the primary database, restore the data on the new standby database, and re-establish the primary-standby relationship. In this processing flow, the redo process needs to be manually triggered. This method takes a long time to repair the standby database. During the process of rebuilding the standby database, if the primary database is unavailable due to reasons such as downtime, the database business will be unavailable for a long time, and there is even a risk of data loss.
[0038] In order to overcome the problems existing in the related art, the embodiments of the present application provide a backup database repair method, device, electronic device, medium and computer program product. The backup database repair method provided in the embodiments of the present application can automatically detect and obtain information about bad tables in the backup database, and promptly determine the bad tables that need to be repaired, automatically repair the bad tables in the backup database, improve the detection efficiency of the bad tables in the backup database, and reduce the risk of data loss and business interruption caused by bad tables in the backup database.
[0039] The following is a further detailed description of the embodiments of the present application in conjunction with the accompanying drawings and examples. It should be understood that the embodiments provided herein are only used to explain the embodiments of the present application and are not intended to limit the embodiments of the present application. In addition, the embodiments provided below are partial embodiments for implementing the present application, rather than providing all embodiments for implementing the present application. In the absence of conflict, the technical solutions recorded in the embodiments of the present application can be implemented in any combination.
[0040] It should be noted that, in the embodiments of the present application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a method or device including a series of elements includes not only the elements explicitly recorded, but also includes other elements not explicitly listed, or also includes elements inherent to the implementation of the method or device. In the absence of further restrictions, an element defined by the sentence "includes..." does not exclude the presence of other related elements (such as steps in the method or units in the device, such as a unit in the device may be a part of a circuit, a part of a processor, a part of a program or software, etc.) in the method or device including the element.
[0041] The backup database repair method provided in the embodiment of the present application includes a series of steps, but the backup database repair method provided in the embodiment of the present application is not limited to the recorded steps. Similarly, the backup database repair device provided in the embodiment of the present application includes a series of modules, but the device provided in the embodiment of the present application is not limited to including the modules explicitly recorded, and can also include modules required to obtain relevant information or perform processing based on information.
[0042] The present application embodiment provides a method for repairing a standby database, such as Figure 1 As shown, Figure 1 A flow chart of a backup database repair method is shown. Figure 1 The following standby database repair methods are shown:
[0043] Step 101: Obtain preset keywords, match the keywords with the error log of the standby database, and automatically obtain information about the bad table; wherein the bad table indicates a damaged table in the standby database; the standby database is used to back up data in the main database.
[0044] The error log of the standby database usually contains a variety of key information, such as errors and warnings encountered during the operation of the standby database, abnormal performance of the standby database, commit, rollback and failure of transactions, etc. The error log provides detailed context information when the problem occurs.
[0045] In order to accurately extract the information of the bad table in the error log, you can preset keywords, which are used to match the information in the error log to extract the information of the bad table in the error log. For example, the table name, table identifier and other information in the standby database can be used as keywords, and table-related error prompt information such as "table damage", "inaccessible table", "page damage pagecorruption" can also be set as keywords, and the error code known to the database system can also be used as a keyword. By matching the keywords with the information in the error log, the information of the bad table contained in the error log is obtained. Here, the information of the bad table includes but is not limited to the table name, storage engine type and other information.
[0046] In actual applications, you can use the grep command to extract the line containing the preset keyword in the error log, and extract the table name and database information of the bad table in the line. You can also set a preset time interval to match the preset keyword with the error log to automatically obtain the bad table information in the error log. The preset time interval here can be controlled by parameters and the default value can be set. Since the probability of generating a single table bad block is not high and this detection has little effect on database performance, the preset time interval can be freely set according to needs.
[0047] In some embodiments, matching the error log of the standby database with the keyword includes: matching the error log of the standby database with the keyword based on a regular detection algorithm.
[0048] A regular expression may be constructed based on the above preset keywords, and information matching the keywords in the error log may be obtained based on the regular expression.
[0049] Step 102: Determine the first bad table that needs to be repaired according to the bad table information.
[0050] After obtaining the bad table information, the first bad table that needs to be repaired can be determined according to the bad table information according to actual needs. Here, the first bad table can be any bad table determined according to the bad table information, and the first bad table can also be all bad tables that need to be repaired contained in the bad table information.
[0051] Step 103: Stop the data interaction between the standby database and the primary database, and synchronously generate the to-be-backed files corresponding to the primary database.
[0052] After determining the first bad table that needs to be repaired, first stop the data interaction between the standby database and the primary database, for example, stop the read / write (Input / Output, I / O) thread of the standby database. Specifically, you can execute the "STOP SLAVE" command on the standby database to stop the I / O thread and SQL thread, thereby stopping the data interaction between the standby database and the primary database.
[0053] After stopping the data interaction between the standby database and the primary database, the standby database will no longer actively receive the data or files that need to be backed up in the synchronization primary database. After stopping the data interaction between the standby database and the primary database, the synchronization generates the backup files in the primary database. Because there is a master-slave delay or replication delay between the primary database and the standby database, that is, the time required for the data changes on the primary database to be reflected on the standby database, when the data interaction between the standby database and the primary database is stopped, the backup files generated by the synchronization primary database usually include some data or transactions that do not exist in the standby database.
[0054] Here, the backup file generated by the main database may be an SQL file, a memory snapshot file DUMP file, a backup (BAK) file or other database backup file. This application does not limit the format type of the backup file.
[0055] Step 104: Repair the first bad table based on the file to be backed up.
[0056] First, determine the situation of the standby database where the first bad table is located, and ensure that the environment of the standby database is the same or similar to that of the main database. Use the import tool or command provided by the database management system (such as MySQL's mysqlimport or LOAD DATA INFILE) to import the file to be backed up into the standby database to repair the first bad table.
[0057] The embodiment of the present application provides a backup database repair method, which can automatically obtain information about bad tables in error logs and automatically repair bad tables, thereby shortening the bad table detection and repair time, reducing the impact on database services, and reducing the risk of data loss.
[0058] In practical applications, steps 101 to 104 can be implemented based on a processor, and the processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a CPU, a controller, a microcontroller, and a microprocessor.
[0059] In some embodiments, before determining the first bad table that needs to be repaired based on the information of the bad table, the method further includes: automatically reporting the information of the bad table to the target storage space; defining a structure; the structure is used to determine preset information that needs to be obtained; the preset information is a part of the information of the bad table; determining the first bad table that needs to be repaired based on the information of the bad table includes: obtaining the preset information in the target storage space based on the structure, and determining the first bad table that needs to be repaired based on the preset information.
[0060] This embodiment provides a method for automatically reporting and acquiring bad table information. Based on the method provided in the above embodiment, after acquiring bad table information in the error log, the acquired bad table information can be uniformly reported to the target storage space, so as to uniformly read the bad table information from the target storage space.
[0061] A structure can be predefined, and the structure is used to determine which part of the information to obtain from the multiple pieces of bad table information. For example, by defining a structure, preset information of a bad table related to bad table repair can be obtained, including but not limited to database type, database version, table name (Tablename), and table storage engine (Engine). In actual applications, in order to obtain database access and modification permissions, a user name with database access and modification permissions can also be preset in the target storage space.
[0062] For example, a structure can be defined as follows:
[0063]
[0064]
[0065] Taking the structure defined in the above example as an example, the defined structure can extract Username, Tablename and Engine from the information of multiple bad tables in the target storage space. Username is used to log in to the backup database where the bad table is located and obtain the repair permission for the bad table; Tablename is used to locate the bad table that needs to be repaired, that is, the first bad table; Engine is used to determine the repair method and formulate an appropriate recovery strategy.
[0066] The preset information is obtained from the target storage space through the structure, and the first bad table to be repaired can be determined based on the preset information. For example, after obtaining multiple Usernames, Tablenames, and Engines from the target storage space through the structure, the specific bad table to be repaired can be determined as the first bad table based on Tablename and Engine. For example, based on the number of bad tables and the current repair processing capacity, a part of the multiple bad tables can be used as the first bad table to be repaired, or all the bad tables can be used as the first bad table to be repaired; or, according to the current database usage priority, one or more specific bad tables among the multiple bad tables in the target storage space can be determined as the first bad table to be repaired.
[0067] This embodiment provides a method for automatically obtaining bad table information and determining the first bad table that needs to be repaired. By reporting the bad table information to the target storage space, it is helpful to achieve unified management and processing of the bad table information. By defining a structure, it is helpful to extract the bad table information required for repair from the information of multiple bad tables, which helps to determine the first bad table that needs to be repaired, and improve the bad table repair efficiency.
[0068] Those skilled in the art will appreciate that, based on the method provided in this embodiment, in a specific application, it is also possible to call and filter the information of the bad table by setting a calling script and writing a related program.
[0069] In some embodiments, before obtaining preset information in the target storage space based on the structure, the method further includes: obtaining information of bad tables reported to the target storage space in real time based on a controller in the Kubernetes cluster.
[0070] Kubernetes (K8s) is an open source, distributed automated container management platform with complete cluster management capabilities. It has a built-in intelligent load balancer and powerful fault detection and self-healing functions. The controller is a control loop in K8s that can continuously obtain the status of the cluster. The controller ensures that the actual state of the cluster is consistent with the expected state.
[0071] When the controller in the K8s cluster is used to obtain the bad table information reported to the target storage space in real time, the target storage space can be the storage space inside the K8s cluster. Based on the method provided in this embodiment, the bad table information can be stored and processed in the K8s cluster, and the bad table information can be obtained in time.
[0072] Specifically, the structure given in the above embodiment can be defined based on the K8s cluster. After obtaining the information of the bad table based on the error log, the information of the bad table is reported to the target storage space in the K8s cluster, and the controller method customized by K8s is used to obtain the new operations automatically reported to the target storage space in real time through the controller, obtain the information of the bad table in real time for the new operations, and obtain the preset information in combination with the defined structure, determine the first bad table, and trigger the repair process.
[0073] When the database in the embodiment of the present application is a cloud database, since the cloud database (MySQL) is a database service based on a cloud computing platform, the cloud database is usually designed to run in a K8s environment. K8s provides flexible deployment and expansion options for cloud databases through its powerful container management capabilities. Therefore, by determining the information of bad tables in the database standby library in K8s, especially the information of bad tables in the cloud database, K8s can provide powerful containerized management, automated deployment, elastic scaling, high availability, and resource optimization functions for the cloud database, so that the cloud database can better meet the needs of actual applications.
[0074] In some embodiments, the first bad table represents a single table with bad blocks in the standby database, and the repairing of the first bad table based on the file to be backed up includes: determining a first identifier corresponding to the primary database and a second identifier corresponding to the first bad table; wherein the first identifier represents an identifier of the latest transaction corresponding to the file to be backed up, and the second identifier represents an identifier of the latest transaction received by the first bad table before the moment of stopping data interaction between the standby database and the primary database; and repairing the first bad table based on the first identifier, the second identifier, and the file to be backed up.
[0075] In a database, a transaction identifier is an identifier used to uniquely identify a transaction. Specifically, a transaction identifier can be a transaction identifier (Transaction Identifier, TID) or a global transaction identifier (Global Transaction Identifier, GTID). For example, when the database is a cloud database, the transaction identifier can be a GTID. In most database systems, such as MySQL, transaction identifiers are usually incremented.
[0076] Since there is a master-slave delay or replication delay between the primary and backup databases, when the data interaction between the backup and primary databases is stopped, the backup files generated by the primary database are synchronized, which usually include some data or transactions that do not exist in the backup database. Therefore, the data difference between the primary and backup databases can be obtained through transaction identifiers.
[0077] In order to repair a single-table corruption, it is necessary to first determine the data difference between the primary database and the secondary database, that is, determine the difference in transactions that have been executed between the primary database and the secondary database. Based on the second identifier and the first identifier, the difference between the transaction received by the secondary database and the transaction currently executed by the primary database at the moment when the data interaction between the secondary database and the primary database is stopped can be obtained, that is, the difference between the data in the primary database and the backup data received by the secondary database can be obtained.
[0078] After determining the data difference between the primary database and the backup database based on the first identifier and the second identifier, the file to be processed can be processed based on the file to be processed by using the first identifier and the second identifier, so that the data in the processed file to be processed is the same as the data obtained after processing the transaction corresponding to the second identifier, that is, the transaction corresponding to the file to be processed is intercepted, so that the data in the file finally backed up in the first bad table is the same as the data of the transaction corresponding to the second identifier that is executed in sequence.
[0079] In actual applications, when the I / O thread of the standby database is stopped, the Retrieved_Gtid_Set information of the standby database can be obtained synchronously. This information is the second Gtid position in the binlog log (binary log file) of the main database received by the standby database at this time, and this second Gtid position is the second identifier. The structure and data of the bad block table on the standby database can be synchronously exported from the main database through the mysqldump backup tool to generate an SQL file. Get the GTID_PURGED information of the main database, through which the first Gtid position when the main database exports the file to be processed can be obtained, and the first Gtid is the first identifier. And comment out the line where the first Gtid position is located, so that the Gtid information synchronized with the standby database will not be modified during the subsequent import of the file to be backed up.
[0080] For example, suppose there are transaction 1, transaction 2, transaction 3, transaction 4, and transaction 5. Assume that at the moment of stopping the I / O thread of the standby database, the transactions received by the standby database include transaction 1, transaction 2, and transaction 3. Due to the master-slave delay or replication delay between the master database and the standby database, the backup file corresponding to the master database generated synchronously at this moment already includes transaction 1, transaction 2, transaction 3, transaction 4, and transaction 5, that is, the backup file generated by the master database includes the data obtained after executing transaction 5. At this time, the identifier of transaction 5 is the first identifier, and the identifier of transaction 3 is the second identifier. After determining the first identifier and the second identifier, and obtaining the backup file, the data in the backup file is rolled back from event 5 to transaction 3, and the first bad table is repaired based on the backup file after the rollback process.
[0081] It can be seen that through the method provided in this embodiment, the to-be-backed-up file that is consistent with the backup task currently executed by the standby database can be obtained, which is conducive to ensuring the consistency of the to-be-backed-up file used for repairing the bad table and the transaction corresponding to the file currently processed by the bad table, and realizes the repair of only a single bad table, without the need to restore data of the entire standby database, thereby improving the repair efficiency.
[0082] In some embodiments, the above-mentioned repairing of the first bad table based on the first identifier, the second identifier, and the file to be backed up includes: importing the file to be backed up in the backup database to obtain a first backup file; determining the first identifier as the starting point of the rollback, and determining the second identifier as the end point of the rollback; based on the starting point and the end point, performing a rollback operation on the first backup file, and using the first backup file after the rollback operation as the data in the repaired first bad table.
[0083] After determining the first identifier and the second identifier, the first identifier is used as the rollback start point and the second identifier is used as the rollback end point to perform a rollback operation on the first backup file imported into the standby database. In actual applications, the first bad table can be deleted first, and the reverse SQL file between the two GTIDs can be generated by the open source binlog2sql tool based on the Retrieved_Gtid_Set and GTID_PURGED information and the binlog of the main database. The rollback operation is performed based on the reverse SQL file, and the first backup file after the rollback operation is performed to obtain the repaired first bad table, that is, a new table.
[0084] In some embodiments, before performing the rollback operation on the first backup file based on the starting point and the end point, the method further includes: determining a third identifier; the third identifier represents an identifier of a transaction completed by the first bad table backup; performing the rollback operation on the first backup file based on the starting point and the end point includes: when the second identifier is the same as the third identifier, performing the rollback operation on the first backup file based on the starting point and the end point.
[0085] Stopping the I / O thread of the standby database can stop the data interaction between the main database and the standby database. However, for the standby database itself, when there is a backup task being executed in the standby database, even if the I / O thread of the standby database is stopped, the standby database will still execute the backup task that has been received. Based on the example given in the above embodiment, it is assumed that before stopping the I / O thread of the standby database, the transactions received by the standby database include transaction 1, transaction 2, and transaction 3, and the transactions that have been executed in the first bad table of the standby database include transaction 1 and transaction 2, that is, for the first bad table, transaction 3 has not been executed.
[0086] Generally, if the backup task in the first bad table is in progress, directly terminating the current task may cause data inconsistency or loss. In this case, the repair operation is usually performed after the current backup task is completed. Therefore, in order to ensure the consistency of the standby database data, before repairing the first bad table, it is necessary to ensure that all backup tasks in the first bad table are completed, that is, after the transaction 3 received by the standby database is completed, the repair operation can be performed.
[0087] Based on the examples given in the above embodiments, in actual applications, it is possible to continuously determine whether the values of Retrieved_Gtid_Set and Executed_Gtid_Set of the standby database are consistent, where the Gtid in Executed_Gtid_Set is the Gtid position information of the binlog of the SQL thread of the standby database executing synchronization with the main database, that is, the information of the transaction being executed by the standby database. If they are inconsistent, wait for consistency and perform cyclic detection.
[0088] The method provided in this embodiment can ensure that the bad table repair process is performed after the standby database completes the current backup task, which is conducive to ensuring the consistency of data in the first bad table and realizing the repair of the first bad table.
[0089] In some embodiments, before repairing the first bad table based on the first identifier, the second identifier, and the file to be backed up, the method further includes: stopping log generation in the standby database.
[0090] Before repairing, by stopping log generation in the standby database, that is, not recording the repair operation in the standby database log, it is beneficial to ensure the consistency of log record information in the standby database when backing up data after the repair, avoid log confusion, and enable the repaired standby database to continue to perform backup tasks.
[0091] In the specific implementation process, you can set SQL_LOG_BIN=0 at the session level of the session control on the standby database to prevent subsequent repair operations from generating binlog logs on the standby database.
[0092] Based on the method given in the above embodiment, after the first bad table is repaired, the standby database I / O thread is restarted, and the primary database data is continued to be backed up in the standby database. Based on the example of the above embodiment, after the standby database I / O thread is restarted, the standby database will start from transaction 4 and continue to back up the primary database data.
[0093] The embodiment of the present application provides a method for automatically discovering and automatically repairing a single table bad block in a database standby database. When a single table bad block appears in the standby database, the bad table can be actively discovered based on the error log, and the bad table can be repaired in time without redoing the standby database. By only restoring the single table that produces the bad block, the single table bad block can be repaired in a short time, minimizing the impact on the database business and reducing the risk of data loss. The embodiment of the present application solves the risk of data loss and business interruption caused by bad blocks in the standby database through mechanisms such as regular detection, automatic reporting, and automatic triggering. During the repair process, only a single table needs to be repaired. Compared with the current standby database redo process, there is no need to restore the entire database, which not only avoids the resource consumption of the backup on the main database in the redo process and minimizes the impact on the main database, but also greatly shortens the standby database repair time and avoids the risk of data loss and business interruption caused by other unexpected factors in the standby database redo process. A method for realizing the recovery of a single table in a database to an arbitrary transaction identifier position at the logical level.
[0094] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0095] In combination with the standby database repair method proposed in the above embodiment, Figure 2 The following is a schematic diagram showing the complete process of repairing a single table corrupt block in a standby database. Figure 2 As shown, based on the repair method provided in the embodiment of the present application, the repair of a single table corrupt block in a standby database includes:
[0096] Step 201: Regular testing.
[0097] Based on the preset keywords, the error logs in the standby database are checked regularly to extract the information of the bad tables in the error logs. Specifically, the Application Programming Interface Server (APIServer) and Informer in the K8s cluster can receive and process the information of the bad tables in the error logs, and store the information of the bad tables in the target storage space in the K8s cluster.
[0098] Step 202: Obtain information.
[0099] Based on the structure defined in the above embodiment, preset information is obtained in the target storage space to determine the first bad table to be repaired.
[0100] Step 203: Stop the I / O thread of the standby database and record the Retrieved_Gtid_Set information of the standby database.
[0101] The Retrieved_Gtid_Set information recorded here includes the second identifier.
[0102] Step 204: Export the table structure and data from the main database, obtain the GTID_PURGED information, and comment the line.
[0103] Export the table structure and data from the main database to obtain the file to be backed up, and obtain the first identifier by obtaining the GTID_PURGED information.
[0104] Step 205: Generate a reverse SQL statement based on the Retrieved_Gtid_Set and GTID_PURGED information and the binlog of the master database.
[0105] The rollback starting point is determined based on the first identifier, the rollback end point is determined based on the second identifier, and a reverse SQL statement for executing the rollback operation is generated in combination with the binlog of the master database.
[0106] Step 206: Determine whether the Retrieved_Gtid_Set and Executed_Gtid_Set of the standby database are consistent.
[0107] When the Retrieved_Gtid_Set and Executed_Gtid_Set of the standby database are consistent, it is considered that the standby database has completed the current backup task and the damaged table repair operation is performed; if the Retrieved_Gtid_Set and Executed_Gtid_Set of the standby database are inconsistent, it is considered that there is still a backup task being executed on the standby database at this time, and the damaged table repair operation is performed after the backup task is completed.
[0108] Step 207: The standby database sets SQL_LOG_BIN=0, deletes the bad block table, and executes the SQL file and the reverse SQL file generated by mysqldump in sequence.
[0109] Before performing the bad table repair, first set the standby database SQL_LOG_BIN = 0, that is, stop the standby database log record. Delete the bad block table, execute the SQL file generated by the main database, that is, the file to be backed up, get the first backup file, and then execute the reverse SQL file on the first backup file to ensure that the data in the first backup file in the standby database is consistent with the data in the backup file of the bad table before the repair. The data in the first backup file after the rollback operation is used as the backup data in the bad table after the repair.
[0110] Step 208: Start the standby database I / O thread.
[0111] After the repair is complete, start the I / O thread of the standby database and enable synchronous backup of the primary and standby databases.
[0112] Based on the backup database repair method proposed in the above embodiment, the present application embodiment also provides a backup database repair device, such as Figure 3 As shown, the standby database repair device includes:
[0113] The acquisition module 301 is used to obtain preset keywords, match the keywords with the error log of the standby database, and automatically obtain the information of the bad table; wherein the bad table means a damaged table in the standby database; the standby database is used to back up the data in the main database.
[0114] The determination module 302 is used to determine the first bad table that needs to be repaired according to the bad table information.
[0115] The processing module 303 is used to stop the data interaction between the standby database and the primary database, and synchronously generate the to-be-backed-up files corresponding to the primary database; and repair the first bad table based on the to-be-backed-up files.
[0116] In practical applications, the acquisition module 301, the determination module 302, and the processing module 303 can be implemented based on a processor and a communication device.
[0117] In some embodiments, before determining the first bad table that needs to be repaired based on the information of the bad table, the processing module 303 is also used to automatically report the information of the bad table to the target storage space; define a structure; the structure is used to determine the preset information that needs to be obtained; the preset information is part of the information of the bad table; the determination module 302 is specifically used to obtain the preset information in the target storage space based on the structure, and determine the first bad table that needs to be repaired based on the preset information.
[0118] In some embodiments, before obtaining preset information in the target storage space based on the structure, the obtaining module 301 is further used to obtain information of bad tables reported to the target storage space in real time based on the controller in the Kubernetes cluster.
[0119] In some embodiments, the first bad table represents a single table with bad blocks in the standby database, and the processing module 303 is specifically used to determine a first identifier corresponding to the main database and a second identifier corresponding to the first bad table; wherein the first identifier represents an identifier of the latest transaction corresponding to the file to be backed up, and the second identifier represents an identifier of the latest transaction received by the first bad table before the moment of stopping data interaction between the standby database and the main database; based on the first identifier, the second identifier, and the file to be backed up, the first bad table is repaired.
[0120] In some embodiments, the processing module 303 is specifically used to import the files to be backed up in the backup database to obtain a first backup file; determine the first identifier as the starting point of the rollback, and determine the second identifier as the end point of the rollback; based on the starting point and the end point, perform a rollback operation on the first backup file, and use the first backup file after the rollback operation as the data in the repaired first bad table.
[0121] In some embodiments, before performing a rollback operation on the first backup file based on the starting point and the end point, the processing module 303 is also used to determine a third identifier; the third identifier represents the identifier of the transaction completed by the first bad table backup; the processing module 303 is specifically used to perform a rollback operation on the first backup file based on the starting point and the end point when the second identifier is the same as the third identifier.
[0122] In some embodiments, based on the first identifier, the second identifier, and the file to be backed up, before repairing the first bad table, the processing module 303 is further configured to stop generating logs in the standby database.
[0123] In some embodiments, the acquisition module 301 is specifically configured to match the error log of the standby database with keywords based on a regular detection algorithm.
[0124] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the same method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0125] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable a computer device (which can be a terminal, a server, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0126] An embodiment of the present application also provides an electronic device. Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the electronic device 40 may include:
[0127] The memory 401 is used to store executable instructions.
[0128] The processor 402 is configured to implement any one of the above-mentioned standby database repair methods when executing the executable instructions stored in the memory 401.
[0129] The processor 402 may be at least one of an ASIC, a DSP, a DSPD, a PLD, a FPGA, a CPU, a controller, a microcontroller, and a microprocessor.
[0130] The above-mentioned computer-readable storage medium or memory 401 can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and other memories; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0131] An embodiment of the present application further provides a computer storage medium, on which computer executable instructions are stored, and the computer executable instructions are used to implement any one of the standby database repair methods provided in the above embodiments.
[0132] Correspondingly, an embodiment of the present application further provides a computer program product, wherein the computer program product includes computer executable instructions, and the computer executable instructions are used to implement any one of the standby database repair methods provided in the above embodiments.
[0133] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0134] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0135] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0136] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0137] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0138] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0139] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A method for repairing a standby database, characterized in that: The method comprises: Obtain preset keywords, match the keywords with the error log of the standby database, and automatically obtain information about the bad table; wherein the bad table indicates a damaged table in the standby database; the standby database is used to back up data in the main database; Determining a first bad table that needs to be repaired according to the bad table information; Stop the data interaction between the standby database and the primary database, and synchronously generate the to-be-backed-up files corresponding to the primary database; The first bad table is repaired based on the file to be backed up.
2. The method according to claim 1, characterized in that: Before determining the first bad table to be repaired according to the bad table information, the method further includes: Automatically reporting the bad table information to the target storage space; A structure is defined; the structure is used to determine the preset information that needs to be obtained; the preset information is a part of the information of the bad table; The determining, according to the information of the bad table, a first bad table that needs to be repaired includes: Preset information is obtained in the target storage space based on the structure, and a first bad table that needs to be repaired is determined based on the preset information.
3. The method according to claim 2, characterized in that Before acquiring preset information in the target storage space based on the structure, the method further includes: Based on the controller in the Kubernetes cluster, the information of the bad table reported to the target storage space is obtained in real time.
4. The method according to claim 1, characterized in that: The first bad table indicates a single table having bad blocks in the standby database, and the repairing of the first bad table based on the file to be backed up includes: Determine a first identifier corresponding to the primary database and a second identifier corresponding to the first bad table; wherein the first identifier indicates an identifier of a latest transaction corresponding to the file to be backed up, and the second identifier indicates an identifier of a latest transaction received by the first bad table before the data interaction between the standby database and the primary database is stopped; The first bad table is repaired based on the first identifier, the second identifier, and the file to be backed up.
5. The method according to claim 4, characterized in that The repairing of the first bad table based on the first identifier, the second identifier, and the file to be backed up includes: Importing the file to be backed up into the backup database to obtain a first backup file; Determine the first identifier as the start point of the rollback, and determine the second identifier as the end point of the rollback; Based on the starting point and the end point, a rollback operation is performed on the first backup file, and the first backup file after the rollback operation is performed is used as data in the repaired first bad table.
6. The method according to claim 5, characterized in that Before performing the rollback operation on the first backup file based on the starting point and the end point, the method further includes: Determine a third identifier; the third identifier represents an identifier of the transaction completed by the backup of the first bad table; The performing a rollback operation on the first backup file based on the starting point and the end point includes: When the second identifier is the same as the third identifier, a rollback operation is performed on the first backup file based on the starting point and the end point.
7. The method according to claim 4, characterized in that Before repairing the first bad table based on the first identifier, the second identifier, and the file to be backed up, the method further includes: Stop generating logs in the standby database.
8. The method according to claim 1, characterized in that The matching of the error log of the standby database with the keyword includes: The error log of the standby database is matched with the keyword based on a regular detection algorithm.
9. A standby database repair device, characterized in that: The device comprises: An acquisition module is used to acquire preset keywords, match the keywords with the error log of the standby database, and automatically acquire information about the bad table; wherein the bad table indicates a damaged table in the standby database; the standby database is used to back up data in the main database; A determination module, used for determining a first bad table that needs to be repaired according to the information of the bad table; The processing module is used to stop the data interaction between the standby database and the main database, and synchronously generate a to-be-backed-up file corresponding to the main database; and repair the first bad table based on the to-be-backed-up file.
10. An electronic device, characterized in that: The electronic device comprises a processor and a memory for storing a computer program that can be run on the processor; wherein, The processor is configured to run the computer program to perform the method according to any one of claims 1 to 8.
11. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 8 when executed by a processor.