Database backup automatic recovery verification method and system based on large language model
By parsing user natural language input through a large language model, combining historical data sets to generate backup and recovery scripts and automatically perform verification, it solves the problems of large manpower and material resources and lack of professional knowledge in existing technologies, and realizes efficient and low-threshold database backup and recovery verification.
Patent Information
- Application Number
- CN202510966850.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-14
AI Technical Summary
The existing database backup and recovery verification work requires a lot of manpower and material resources, takes up a lot of developer time, and some users lack professional knowledge and are unable to perform backup verification on their own, resulting in low efficiency and high communication costs.
Using a method based on a large language model, a user interaction interface is built through the RAG platform to receive natural language input, and a backup and recovery script is generated by combining historical data sets and prompt words. The recovery server is automatically selected and verification is performed, and the recovery strategy is monitored and optimized in real time, forming a closed loop for the entire process.
It significantly reduces dependence on database expertise, reduces manual operation costs, improves the efficiency and success rate of backup and recovery verification, shortens recovery time, and lowers the usage threshold.
Smart Images

Figure CN120803814A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of database operation and maintenance and intelligent information processing, and particularly relates to a database backup automatic recovery verification method and system based on a large language model. BACKGROUND
[0002] In the digital economy era, data has become the core strategic asset of enterprise operation and development, and its security and availability are directly related to business continuity, user experience and stable operation of social public service system. With the wide application of cloud computing, AI and other technologies, the data scale is growing exponentially, and the reliability of database systems as the core infrastructure carrying key business data is facing unprecedented challenges. Under this background, database backup and recovery technology is not only the basic defense line to deal with hardware failure, human error, network attack and other risks, but also an important cornerstone to support enterprise digital transformation and realize data lifecycle management.
[0003] Database backup verification is an important part of digital work and a core link to ensure enterprise data security. It ensures the effectiveness of backup files by restoring backup data regularly. However, the daily large number of recovery verification operations require a lot of manpower and resources to operate. Currently, a single database recovery requires a total time of 127-134 hours, which means that 5-6 working days are needed to verify the monthly database backup. In addition, because of the possibility of recovery failure, manual follow-up is required, which occupies a lot of manual intervention time. Furthermore, the backup and recovery verification work requires a lot of parameter configuration, which requires certain professional knowledge, making it difficult for users lacking such knowledge to perform database backup verification.
[0004] The existing solutions mainly use the following methods for database backup recovery: 1) Manual + automated combination for processing, which involves manual processing of test requirements and checking of recovery status, and automated calling of API scripts to perform recovery operations and verify data integrity based on logs and requirements. However, this method requires continuous manual follow-up and operation, and requires certain professional knowledge. Currently, according to the total amount of internal databases, the recovery estimate requires a large amount of manpower. 2) Using mature middleware for processing, such as OEM or third-party middleware, to perform recovery and automated recovery verification. However, this method requires professional knowledge for configuration information, which increases communication costs for non-professional personnel. As can be seen, most existing solutions are based on automated program construction for the entire backup verification scheme, which requires users to be familiar with database backup and recovery related parameters and to configure a large number of parameters to implement automated recovery processes. This is not suitable for users who are not familiar with database backup and recovery related parameters to perform recovery verification.
[0005] In summary, in order to reduce labor costs, minimize manual work, and not occupy a large number of developers and time, and not increase the communication cost between departments, an efficient, accurate and full-coverage method is needed to optimize the database backup verification work and improve the overall competitiveness and efficiency of enterprises. SUMMARY
[0006] To solve the problems of requiring a large amount of manpower and material resources, occupying a large number of developers and time, and users lacking professional knowledge being unable to implement database backup verification, the present application provides a database backup automatic recovery verification method based on a large language model, which can effectively reduce the cost of manual operation, reduce the workload of users on database backup recovery operation, and reduce the use threshold of users on backup recovery verification, thereby improving the efficiency of backup recovery verification. The present application also relates to a database backup automatic recovery verification system based on a large language model.
[0007] The technical scheme of the present application is as follows:
[0008] A database backup automatic recovery verification method based on a large language model, characterized in that it comprises the following steps:
[0009] A user interaction interface construction step: a user interaction interface is constructed using an RAG platform, and database backup recovery requirement information input by a user in natural language form is received through the user interaction interface;
[0010] A historical data set construction step: script files and log records used in historical backup recovery operations of a number of databases of an enterprise are collected; the script files and log records are converted into text format and stored in a knowledge base of an RAG platform to form a historical data set;
[0011] A user requirement analysis step: a large language model is used to analyze database backup recovery requirement information input by a user in natural language form in combination with prompt words to generate target positioning information that the user needs to recover;
[0012] A backup recovery script generation step: the historical data set is queried to obtain historical backup recovery records matching the target positioning information that the user needs to recover, and a backup recovery script containing recovery resource configuration parameters is generated based on the historical backup recovery records by using a large language model in combination with prompt words;
[0013] A recovery server selection step: backup cluster server performance data stored in the historical data set is queried, a recovery server that meets performance requirements determined based on the performance data and the estimated backup data volume in the backup recovery script is selected from the backup cluster servers by using a large language model;
[0014] The recovery verification execution step: automatically executing the generated backup recovery script on the selected recovery server to perform recovery verification on the backup file of the target database specified by the user, monitoring the execution process of the backup recovery script in real time and recording the execution log, when an error in the execution of the backup recovery script task is monitored, the error execution log is fed back to the large language model for cause analysis, and according to the analysis result of the large language model, the backup recovery script is re-executed or a new backup recovery script is generated after adjusting the recovery resource configuration parameters;
[0015] The result feedback and optimization step: when the execution of the backup recovery script task is monitored, the complete execution log of this recovery process is stored in the historical data set, and the user is notified of the task execution result through the user interaction interface, and the automatic recovery verification of the database backup is completed.
[0016] Preferably, in the backup recovery script generation step, a backup recovery verification interface is also built based on the FastAPI framework developed by Python, the backup recovery verification interface includes a backup query interface and a parameter query interface; the backup query interface is used to call the RMAN API to query the backup dataset file name corresponding to the target positioning information in the historical data set, based on the backup dataset file name, the recovery resource configuration parameters required for the backup file of the target database are analyzed by the large language model combined with the prompt word, and then the backup recovery script is constructed;
[0017] In the recovery server selection step, the backup cluster server performance data stored in the historical data set is queried through the parameter query interface, the performance data includes current running environment parameters and physical server performance indicators, based on the estimated backup data volume in the backup recovery script, and through the large language model, the recovery server that meets the storage space demand and memory capacity demand in the current running environment parameters, and the processor performance indicator demand in the physical server performance indicator is selected from the backup cluster server; the running environment parameters include memory configuration state, storage path information and version characteristics.
[0018] Preferably, the backup recovery verification interface further includes a log callback interface, a backup execution interface and a backup state query interface, in the recovery verification execution step, the recorded execution log is fed back to the large language model in real time through the log callback interface for analysis, and the recovery strategy is dynamically adjusted based on the analysis result of the large language model; the backup recovery script is called through the backup execution interface to perform recovery verification on the backup file of the target database; the current backup file recovery state is obtained by querying the execution log through the backup state query interface; the recovery strategy includes parameter optimization, script correction and server reselection.
[0019] Preferably, the backup query interface further receives a query request containing a target database identifier and a recovery time range, and returns backup file metadata in a standardized JSON format.
[0020] Preferably, in the user interaction interface construction step, the database backup recovery requirement information includes a target database identifier and a desired recovery time point.
[0021] In the historical data set construction step, the script file includes backup recovery commands and parameter configurations, and the log record includes script execution results and error information.
[0022] In the user requirement analysis step, the target positioning information includes a target database network address, a database system identifier, and a time node to which the user expects to recover.
[0023] Preferably, in the backup recovery script generation step, the recovery resource configuration parameters include system global memory area configuration parameters, program global memory area configuration parameters, and recovery operation target instruction sets. The system global memory area configuration parameters include memory allocation values of database buffers, shared pools, and redo log buffers. The program global memory area configuration parameters include memory limits of sorting areas, hash areas, and private SQL areas. The recovery operation target instruction sets include backup set paths, recovery termination points, and data file redirection rules.
[0024] A database backup automatic recovery verification system based on a large language model, characterized by comprising a user interaction interface construction module, a historical data set construction module, a user requirement analysis module, a backup recovery script generation module, a recovery server selection module, a recovery verification execution module, and a result feedback and optimization module connected in sequence,
[0025] The user interaction interface construction module uses the RAG platform to construct a user interaction interface, and receives database backup recovery requirement information input by a user in natural language form through the user interaction interface.
[0026] The historical data set construction module collects script files and log records used in historical backup recovery operations of a plurality of databases of an enterprise, converts the script files and log records into text format, and stores them in a knowledge base of the RAG platform to form a historical data set.
[0027] The user requirement analysis module analyzes database backup recovery requirement information input by a user in natural language form through a large language model and in combination with prompt words, and generates target positioning information required by the user for recovery.
[0028] The backup recovery script generation module queries the historical data set, obtains a historical backup recovery record matched with target positioning information required by the user to be recovered, and then generates a backup recovery script containing recovery resource configuration parameters based on the historical backup recovery record through a large language model and in combination with prompt word analysis.
[0029] The recovery server selection module queries backup cluster server performance data stored in the historical data set, selects a recovery server that meets performance requirements determined based on the performance data from the backup cluster servers based on the estimated backup data volume in the backup recovery script, and through a large language model.
[0030] The recovery verification execution module automatically executes the generated backup recovery script on the selected recovery server, performs recovery verification on the backup file of the target database specified by the user, monitors the execution process of the backup recovery script in real time and records execution logs, feeds error execution logs back to the large language model for cause analysis when monitoring task execution errors of the backup recovery script, and selects to re-execute the backup recovery script or generate a new backup recovery script after adjusting the recovery resource configuration parameters according to the analysis result of the large language model.
[0031] The result feedback and optimization module stores the complete execution logs of this recovery process to the historical data set when monitoring the completion of the backup recovery script task execution, sends a task execution result notification to the user through the user interaction interface, and completes the automatic recovery verification of the database backup.
[0032] Preferably, the backup recovery script generation module also builds a backup recovery verification interface based on the FastAPI framework developed by Python, which includes a backup query interface and a parameter query interface; the backup query interface calls the RMAN API to query the backup data set file name corresponding to the target positioning information in the historical data set, and based on the backup data set file name, the large language model and the prompt word analysis are used to analyze the recovery resource configuration parameters required for the backup file of the target database, and then the backup recovery script is constructed.
[0033] In the recovery server selection module, the backup cluster server performance data stored in the historical data set is queried through the parameter query interface, the performance data includes current running environment parameters and physical server performance indicators, and based on the estimated backup data volume in the backup recovery script, a recovery server that meets the storage space requirement and memory capacity requirement in the current running environment parameters and the processor performance indicator requirement in the physical server performance indicators is selected from the backup cluster servers through the large language model; the running environment parameters include memory configuration state, storage path information and version characteristics.
[0034] Preferably, the backup recovery verification interface further comprises a log callback interface, a backup execution interface and a backup state query interface, in the recovery verification execution module, the recorded execution log is fed back to the large language model in real time through the log callback interface for analysis, and the recovery strategy is dynamically adjusted based on the analysis result of the large language model; the backup recovery script is called through the backup execution interface to recover and verify the backup file of the target database; the execution log is queried through the backup state query interface to obtain the current backup file recovery state; the recovery strategy comprises parameter optimization, script correction and server reselection.
[0035] Preferably, the backup query interface further receives a query request containing a target database identifier and a recovery time range, and returns backup file metadata in a standardized JSON format.
[0036] The technical effects of the present application are as follows:
[0037] The application provides a database backup automatic recovery verification method based on a large language model. First, a user interaction interface is constructed using an RAG platform, and database backup recovery requirement information input by a user in natural language form is received through the user interaction interface, effectively reducing the user operation threshold, enabling non-professionals to describe requirements in natural language without needing to master professional database command syntax, improving interaction efficiency, and avoiding format errors or parameter omissions that may occur in traditional command line input. Then, script files and log records used in historical database backup recovery operations of an enterprise are collected; the script files and log records are converted into text format and stored in a knowledge base of the RAG platform to form a searchable historical data set (knowledge base), and the RAG platform searches the knowledge base in real time, matches user input with historical cases, effectively improves the accuracy of subsequent large language model analysis, provides data support for subsequent automated decision-making, and realizes fast matching through structured storage, avoiding the low efficiency of manually reviewing historical records. Then, the database backup recovery requirement information input by the user in natural language form is analyzed by a large language model combined with prompt words to generate target positioning information required by the user for recovery, accurately extracts user intent, avoids subjective errors in manual analysis, effectively reduces the user's professional knowledge requirement threshold, supports multiple language inputs, and adapts to the needs of internationalized enterprise environments. Then, based on historical backup recovery records and a large language model, and combined with prompt words, a backup recovery script containing recovery resource configuration parameters is analyzed and generated, the backup recovery API is automatically called through the script, and the recovery operation is automatically executed, skipping traditional manual coding and avoiding memory shortages or resource waste caused by manual settings, replacing the traditional manual and inefficient spot checks, and effectively improving the recovery success rate by inheriting historical successful experience. According to the script requirements and large language model analysis, a matching recovery server is selected, dynamic resource allocation is realized, the server performance is matched with the recovery task requirements, and recovery failures or performance bottlenecks caused by insufficient resources are avoided. Then, the backup recovery script is automatically executed to recover and verify the target database backup file and real-time monitoring is performed, the large language model is triggered for analysis in case of an exception, real-time fault detection and self-repair are performed, the need for manual intervention is reduced, the root cause is quickly located through log analysis, and the fault handling time is shortened. Finally, the execution log is stored and the user is notified of the task execution result, the automatic verification of the database backup is completed, a closed-loop learning is formed, the manual operation cost is effectively reduced, the workload of the user on the database backup recovery operation is reduced, the use threshold of the user on the backup recovery verification is reduced, and the efficiency of the backup recovery verification is improved.
[0038] The application directly analyzes user natural language input through a large language model, understands user intent, significantly reduces dependence on database professional knowledge, realizes "zero configuration" operation, greatly reduces the use threshold of user backup recovery verification, and improves the efficiency of backup recovery verification. In other words, it can be understood as understanding user semantics through an AIGC large model (i.e., a generative artificial intelligence large language model capable of semantic understanding and content generation based on user input natural language), automatically performing parameter analysis and script calling, which reduces the user's professional knowledge requirement threshold compared to the current scheme. And build a structured knowledge base for historical backup scripts and log records, combined with RAG (retrieval augmented generation) technology, so that the large model can dynamically recommend the optimal recovery strategy (such as selecting servers and risk early warning) based on historical data, avoiding repeated manual analysis of historical logs, automatically matching similar scenarios through a knowledge base (historical data set), effectively improving the recovery success rate and reducing the trial and error cost. In addition, the application integrates FastAPI interface and automated scripts to realize the whole process closed loop from demand analysis to script generation, recovery execution, log monitoring, and failure callback (such as automatic retry or email notification). The traditional scheme requires manual step-by-step operation and cannot correct in real time, while the application compresses the single recovery time from 127+ hours to unattended completion through automatic tracking and callback mechanism, effectively improving the efficiency and success rate of recovery verification, and greatly saving the labor cost.
[0039] The application compresses the traditional manual operation of 5-6 people days to a fully automated process, with an efficiency improvement of more than 90%. The dependence on professional DBAs is reduced through a historical knowledge base and a large language model, and the labor cost is reduced by 70%. The abnormal automatic processing mechanism improves the recovery success rate from 80% of manual operation to more than 98%. The standardized process avoids human operation differences and ensures enterprise data security.
[0040] Further, the backup recovery script generation step also builds a backup recovery verification interface based on the FastAPI framework developed by Python, which includes a backup query interface and a parameter query interface; the backup query interface calls RMAN API to query the backup dataset file name corresponding to the target positioning information in the historical dataset, and the queried backup dataset file name is injected as a key parameter into the generation process of the backup recovery script; and in the recovery server selection step, the current running environment parameters and physical server performance indicators are obtained through the parameter query interface, and the obtained current running environment parameters and physical server performance indicators are used as important basis for recovery server selection. Backup information is obtained through the official RMAN API, avoiding manual input errors, ensuring the accuracy of the backup set path, and eliminating the manual operation link of querying parameters and backup files, so that the single recovery preparation time is compressed from 30 minutes to 10 seconds.
[0041] Further, the backup recovery verification interface further comprises a log callback interface, a backup execution interface and a backup state query interface. The recorded execution log is fed back to the large language model in real time through the log callback interface for analysis, and the recovery strategy is dynamically adjusted based on the analysis result of the large language model, realizing real-time monitoring of the backup / recovery process, ensuring that any abnormality or error can be captured in time. Combined with the powerful semantic understanding and reasoning ability of the large language model, the system can automatically identify potential problems and patterns in the log, thereby improving the overall recovery efficiency and success rate. The backup recovery script is called for backup recovery through the backup execution interface. The current backup recovery state is obtained by querying the log through the backup state query interface. The backup recovery verification interface system integrates intelligent analysis, automatic execution and visual monitoring three core capabilities, realizes the whole process closed loop from log collection, intelligent analysis, strategy adjustment to task execution, state feedback, improves the stability, flexibility and intelligent level of the backup recovery system, significantly reduces the manual operation and maintenance cost, enhances the self-adaptation ability and fault tolerance ability of the system in the face of complex scenes, and at the same time supports unified backup recovery management under multi-tenant, multi-business line and multi-environment, has good scalability and deployment adaptability.
[0042] The present invention also relates to a database backup automatic recovery verification system based on a large language model. This system corresponds to the above-mentioned database backup automatic recovery verification method based on a large language model and can be understood as a system that implements the above-mentioned database backup automatic recovery verification method based on a large language model. The system includes a user interaction interface construction module, a historical data set construction module, a user demand analysis module, a backup recovery script generation module, a recovery server selection module, a recovery verification execution module, and a result feedback and optimization module, which are connected in sequence. Each module works in conjunction with each other, directly parsing the user's natural language input through the large language model and understanding the user's intention, significantly reducing the dependence on database expertise, greatly lowering the user's usage threshold for backup recovery verification, and improving the efficiency of backup recovery verification. The historical backup scripts and log records are constructed into a structured knowledge base, combined with RAG technology, so that the large model can dynamically recommend the optimal recovery strategy based on historical data, avoiding repeated manual analysis of historical logs, and automatically matching similar scenarios through historical data sets, effectively improving the recovery success rate and reducing trial and error costs. In addition, the present invention integrates the FastAPI interface with automated scripts to achieve a closed loop of the entire process from demand analysis → script generation → recovery execution → log monitoring → failure callback (such as automatic retry or email notification), which effectively improves the efficiency and success rate of recovery verification and greatly saves labor costs. The system improves the time for manual intervention in recovery verification operations, while increasing the communication efficiency of the development business department for backup recovery verification, greatly saving labor costs and digital efficiency. According to calculations, compared with traditional methods, the overall manual intervention time has been reduced by more than 90%, and the overall communication efficiency has been improved by more than 50%. After adopting this system, the backup verification work can save more than 20% of labor costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a flow chart of the database backup automatic recovery verification method based on the large language model of the present invention.
[0044] Figure 2 This paper provides a schematic diagram for implementing the automatic recovery verification method for database backup based on a large language model. DETAILED DESCRIPTION
[0045] The present invention will be described below with reference to the accompanying drawings.
[0046] The application relates to a database backup automatic recovery verification method based on a large language model, which obtains a user recovery target database, analyzes and constructs recovery related parameters, and performs a database recovery operation. The AIGC large language model understands user semantics, accurately identifies user requirements, extracts target database IP, SID, and automatically analyzes database size, allocates to a backup recovery environment that meets the requirements, configures corresponding database parameters, analyzes possible risk points, constructs a backup recovery script, automatically calls the script on the backup recovery environment within the plan range, and tracks the recovery log. If an error occurs, the cause is analyzed, the script is appropriately recalled for recovery, and if it cannot be completed, an email is automatically sent to notify the user. The flowchart of the method is shown in Figure 1 The method comprises the following steps in sequence:
[0047] I. User interaction interface construction step: a user interaction interface is constructed by using an RAG platform, and database backup recovery demand information input by a user in a natural language form is received through the user interaction interface.
[0048] Specifically, as shown in Figure 2 , first, an open source RAG platform (such as FastGPT and Dify) is used to construct a user interaction interface supporting natural language input (such as “please restore the production database to the backup of last night”), so as to provide a dialogue between the user and the large language model and subsequent low-code operation. After the user initiates a natural language request through a PC terminal, the RAG platform receives database backup recovery demand information input by the user in a natural language form through the user interaction interface. The database backup recovery demand information includes a target database identifier and a time point expected to be recovered.
[0049] II. Historical data set construction step: script files and log records used in historical backup recovery operations of a plurality of databases of an enterprise are collected; the script files and log records are converted into a text format and stored in a knowledge base of the RAG platform to form a historical data set. The historical data set is an important and indispensable link for the large model to improve the knowledge understanding ability, and the purpose is to let the large language model understand the meaning in the backup recovery script through a series of processing steps.
[0050] Specifically, first, script files and log records used in historical backup recovery operations of each database of the enterprise in the past period (such as the past few years) are collected, wherein the script files include backup recovery commands and parameter configurations; the log records include script execution results and error information. Then, the script files and the log records are converted into TXT text format and stored in the knowledge base of the RAG platform (that is, stored in the target database DB of the RAG platform, such as the temporary database VDB and the relational database RDBMS for verification), forming a historical data set for calling by the RAG platform. By constructing the historical backup script and the log record into the historical data set (which can also be referred to as the knowledge base), combined with the RAG (retrieval enhancement generation) technology, the large model can dynamically recommend the optimal recovery strategy (such as selecting the server and risk early warning) based on the historical data, avoiding repeated manual analysis of historical logs, automatically matching similar scenarios through the knowledge base, greatly improving the recovery success rate and reducing the trial and error cost.
[0051] III. User demand analysis step: The database backup recovery demand information input by the user in the natural language form is analyzed by a large language model combined with a prompt word to generate target positioning information required by the user to be recovered.
[0052] Specifically, the database backup recovery demand information input by the user in the natural language form (such as “restore database A to May 1, 2024”) is analyzed by a large language model (such as DeepSeek-R1:32B model) using a prompt word (the analysis includes text analysis, parameter generation, and log analysis) to generate target positioning information required by the user to be recovered. The DeepSeek-R1:32B model is a medium-sized locally deployed model that can realize high-precision text generation and semantic understanding ability under the demand of small amount of computing power, and its text generation process can be simply regarded as a black box, only more attention is paid to the text input to the model and the knowledge base and the output result of the text. Preferably, the target positioning information includes the network address of the target database (target database IP), the database system identifier (SID), and the time node expected to be recovered by the user.
[0053] IV. Backup recovery script generation step: The historical data set is queried to obtain the historical backup recovery record matched with the target positioning information required by the user to be recovered, and then a backup recovery script containing recovery resource configuration parameters is generated by analyzing a large language model combined with a prompt word based on the historical backup recovery record.
[0054] Specifically, first, a backup recovery verification interface is built based on the FastAPI framework developed by Python (the development tool is PyCharm, and the API service interface transmission format is JSON), which can include a backup query interface, a parameter query interface, a log callback interface, a backup execution interface, and a backup state query interface. The backup query interface is used to call the RMAN API to query historical data sets, and obtain historical backup recovery records (such as backup data set file names) in the historical data sets that match the target positioning information required by the user to recover. Based on the historical backup recovery records (backup data set file names), the required recovery resource configuration parameters of the backup file of the target database are analyzed by using a large language model and using prompt words, and then a backup recovery script is built. In addition, the backup query interface also receives a query request containing a target database identifier and a recovery time range, and returns backup file metadata in a standardized JSON format. The recovery resource configuration parameters include system global memory area (SGA) configuration parameters, program global memory area (PGA) configuration parameters, and a recovery operation target instruction set. The system global memory area (SGA) configuration parameters include memory allocation values of the database buffer, the shared pool, and the redo log buffer. The program global memory area (PGA) configuration parameters include memory limits of the sorting area, the hash area, and the private SQL area. The recovery operation target instruction set includes a backup set path, a recovery termination point, and a data file redirection rule.
[0055] Five, recovery server selection step: query the backup cluster server (the backup cluster server includes multiple recovery servers) performance data (i.e. database current running environment parameters and physical server performance indicators) stored in the historical data set through the parameter query interface, based on the estimated backup data volume in the backup recovery script, and select a recovery server that meets the performance requirements determined based on the performance data (such as meeting the storage space requirement and memory capacity requirement in the database current running environment parameters, and the processor performance indicator in the physical server performance indicators) from the backup cluster server by using appropriate prompt words in the large language model, that is, according to the size of the recovery data, and determine which recovery server that meets the server performance requirements by using appropriate prompt words in the large language model. For example, based on the estimated backup data volume in the backup recovery script, the following matching decisions are made by the large language model:
[0056] 1) Filter server nodes with available storage space ≥ 120% of the backup volume.
[0057] 2) Select physical hosts with free memory capacity ≥ 110% of the total SGA / PGA requirement of the script.
[0058] 3) Preferentially match processor resources with no overload record in recent period.
[0059] The final output satisfies the comprehensive performance requirement of the recovery server identification. The current database running environment parameters are the memory configuration state, storage path information, and version characteristics obtained by querying the database running parameter view. The physical server performance indicators are the storage remaining capacity, memory available space, and processor real-time load.
[0060] Six, recovery verification execution step: automatically executing the generated backup recovery script on the selected recovery server, performing recovery verification on the backup file of the target database, monitoring the execution process of the backup recovery script in real time and recording the execution log, when the backup recovery script task execution error is monitored, the error execution log is fed back to the large language model for cause analysis, and according to the analysis result of the large language model, the backup recovery script is selected to be re-executed or a new backup recovery script is generated after adjusting the recovery resource configuration parameters.
[0061] Specifically, first, the generated backup recovery script is automatically executed on the selected recovery server, the backup recovery script is called through the backup execution interface to perform recovery verification on the backup file of the target database, the execution process of the backup recovery script is monitored in real time and the execution log is recorded, on one hand, the current backup recovery state is obtained by querying the log through the backup state query interface, on the other hand, the recorded execution log is fed back to the large language model in real time through the log callback interface for analysis, and the recovery strategy is dynamically adjusted based on the analysis result of the large language model, when the backup recovery script task execution error is monitored, the error execution log is fed back to the large language model through the log callback interface for cause analysis, and according to the analysis result of the large language model, the backup recovery script is selected to be re-executed or a new backup recovery script is generated after adjusting the recovery resource configuration parameters, or the user is notified of the recovery failure and the reason for the failure through multiple ways. In addition, the recovery strategy can also be dynamically adjusted based on the analysis result of the large language model, the recovery strategy includes parameter optimization, script correction and server reselection. The recovery resource configuration parameters include system global memory area configuration parameters, program global memory area configuration parameters and recovery operation target instruction set, the system global memory area configuration parameters include memory allocation values of database buffer, shared pool and redo log buffer; the program global memory area configuration parameters include memory limits of sorting area, hash area and private SQL area; the recovery operation target instruction set includes backup set path, recovery termination point and data file redirection rule.
[0062] Seven, result feedback and optimization step: when the backup recovery script task execution is completed, the complete execution log of this recovery process is stored to the historical data set, and the task execution result (the task execution result includes the recovery success state or the failure reason analysis) is sent to the user through the user interaction interface to notify the user, completing the automatic recovery verification of the database backup.
[0063] Figure 2The implementation principle diagram of the database backup automatic recovery verification method based on a large language model shows the whole process implementation principle of user demand large model analysis automatic execution closed loop feedback. The core RAG platform involves dialogue interaction, interface calling, and data bus, which can be understood as corresponding user interface construction, user demand analysis, and result feedback. The dialogue interaction receives user natural language demand and transmits it to the large language model analysis. The interface calling connects the backup recovery verification interface, schedules execution, log callback, and query operations. The data bus integrates historical data sets, server information, and execution logs to support the whole process data flow. The large language model service involves text analysis, parameter generation, and log analysis, which can be understood as corresponding user demand analysis, backup recovery script generation, and recovery verification execution (error analysis). The text analysis analyzes user natural language demand and extracts target database IP, SID, and recovery time positioning information. The parameter generation generates SGA, PGA, and other recovery resource configuration parameters in combination with historical data sets to construct a backup recovery script. The log analysis monitors the execution log, automatically analyzes the reason when an error is identified, and decides whether to retry or adjust the script.
[0064] Embodiment:
[0065] After a natural language request instruction of "please help me restore the CORIS database data of yesterday" is input by a user interaction interface built by a database administrator of an enterprise through an RAG platform, the RAG platform transmits the request instruction to a large language model center through a data bus, a model communication module of the model center schedules a large language model service after receiving the request, the large language model analyzes the natural language request, and accurately extracts the unique identification information of the target database, including the network protocol address (IP: 192.168.1.100) of the database server, the database system identifier (SID: ORCL_CORIS) and the specific recovery time node (such as 2025-7-1023:00:00); then the query interface in the backup recovery verification interface is queried to match out 5 historical backup recovery records of the database in the past three months, the large language model generates a resource configuration scheme (backup recovery script) containing the system global area memory (SGA) 12GB and the program global area memory (PGA) 6GB based on the historical backup recovery records and the prompt word analysis, and estimates that 500GB of storage space is required for this recovery; the system continues to query the server performance data in the historical data set, selects the physical recovery server Node-12 with a current load rate lower than 40% and available storage space exceeding 600GB from the backup cluster server as the recovery node; the generated recovery script is automatically executed on the Node-12 server, and when the real-time monitoring "ORA-19809: insufficient space for specified recovery area" error is detected, the error log is immediately fed back to the large language model, the large language model dynamically adjusts the recovery area parameter from 400GB to 650GB after analysis, and continues to execute; finally, after successful verification, the system updates all configuration parameters, execution indicators and optimization records of this recovery to the historical data set, and sends a notification of "ORCL_PROD database backup verification success, time-consuming 8 hours and 23 minutes" to the administrator, and completes the closed-loop automation process from demand input to verification completion.
[0066] The application also relates to a database backup automatic recovery verification system based on a large language model, which corresponds to the above-mentioned database backup automatic recovery verification method based on a large language model and can be understood as a system for realizing the above-mentioned method, and comprises a user interaction interface building module, a historical data set building module, a user demand analysis module, a backup recovery script generation module, a recovery server selection module, a recovery verification execution module and a result feedback and optimization module connected in sequence.
[0067] The user interaction interface building module builds a user interaction interface by using an RAG platform, and receives database backup recovery demand information input by a user in a natural language form through the user interaction interface.
[0068] The historical data set construction module collects script files and log records used in historical backup recovery operations of the databases of the enterprise, converts the script files and the log records into a text format, and stores the script files and the log records into a knowledge base of the RAG platform to form the historical data set;
[0069] The user demand analysis module analyzes database backup recovery demand information input by a user in a natural language form through a large language model and in combination with a prompt word, and generates target positioning information required to be recovered by the user.
[0070] The backup recovery script generation module queries the historical data set, acquires historical backup recovery records matched with the target positioning information required to be recovered by the user, and analyzes and generates a backup recovery script containing recovery resource configuration parameters based on the historical backup recovery records through a large language model and in combination with a prompt word.
[0071] The recovery server selection module queries backup cluster server performance data stored in the historical data set, selects a recovery server meeting performance requirements determined based on the performance data and the backup data volume estimated in the backup recovery script from backup cluster servers through a large language model, and selects a recovery server meeting performance requirements determined based on the performance data and the backup data volume estimated in the backup recovery script from backup cluster servers through a large language model.
[0072] The recovery verification execution module automatically executes the generated backup recovery script on the selected recovery server, performs recovery verification on backup files of the target database specified by the user, monitors an execution process of the backup recovery script in real time and records an execution log, feeds back error execution logs to a large language model for cause analysis when an error in the execution of the backup recovery script task is monitored, and selects to execute the backup recovery script again or generates a new backup recovery script after adjusting recovery resource configuration parameters according to an analysis result of the large language model.
[0073] The result feedback and optimization module stores complete execution logs of the current recovery process to the historical data set when the execution of the backup recovery script task is completed, sends a task execution result notification to the user through a user interactive interface, and completes automatic recovery verification of the database backup.
[0074] Preferably, the backup recovery script generation module further constructs a backup recovery verification interface based on a FastAPI framework developed by Python, the backup recovery verification interface includes a backup query interface and a parameter query interface, the backup query interface is used to query backup dataset filenames corresponding to the target positioning information in the historical data set through RMANAPI, and recovery resource configuration parameters required for backup files of the target database are analyzed through a large language model and in combination with a prompt word based on the backup dataset filenames, so as to construct the backup recovery script.
[0075] In the recovery server selection module, the backup cluster server performance data stored in the historical data set is queried through a parameter query interface, the performance data including current running environment parameters and physical server performance indicators, and based on the estimated backup data volume in the backup recovery script, a recovery server is selected from the backup cluster server through a large language model to meet the storage space requirement and memory capacity requirement in the current running environment parameters and the processor performance requirement in the physical server performance indicators; the running environment parameters including memory configuration state, storage path information and version characteristics.
[0076] Preferably, the backup recovery verification interface further includes a log callback interface, a backup execution interface and a backup state query interface, in the recovery verification execution module, the recorded execution log is fed back to the large language model in real time through the log callback interface for analysis, and the recovery strategy is dynamically adjusted based on the analysis result of the large language model; the backup recovery script is called through the backup execution interface to recover and verify the backup file of the target database; the backup state query interface is used to query the execution log to obtain the current backup file recovery state; the recovery strategy includes parameter optimization, script correction and server reselection.
[0077] Preferably, the backup query interface further receives a query request containing a target database identifier and a recovery time range, and returns backup file metadata in a standardized JSON format.
[0078] The present application provides an objective and scientific database backup automatic recovery verification method and system based on a large language model, which directly analyzes user natural language input through a large language model, understands user intent, significantly reduces dependence on database professional knowledge, greatly reduces the user's use threshold for backup recovery verification, and improves the efficiency of backup recovery verification. And the historical backup script and log record are constructed into a structured knowledge base, combined with RAG technology, so that the large model can dynamically recommend the optimal recovery strategy based on historical data, avoiding repeated manual analysis of historical logs, automatically matching similar scenarios through historical data sets, effectively improving the recovery success rate and reducing the trial and error cost. In addition, the present application integrates FastAPI interface and automatic script to realize the whole process closed loop from demand analysis → script generation → recovery execution → log monitoring → failure callback (such as automatic retry or email notification), effectively improving the efficiency and success rate of recovery verification, greatly saving the labor cost.
[0079] It should be noted that the above detailed description of the specific embodiments of the present application is not intended to limit the present application in any way. Thus, while the present application has been described in detail with reference to specific embodiments thereof, it will be apparent to those skilled in the art that various modifications and changes can be made thereto without departing from the spirit and scope of the present application.
Claims
1. A database backup automatic recovery verification method based on a large language model, characterized in that: The following steps are involved: User interaction interface construction step: Use the RAG platform to build a user interaction interface, and receive database backup and recovery requirement information input by the user in natural language through the user interaction interface; Historical dataset construction steps: Collect the script files and log records used in historical backup and recovery operations of several enterprise databases; convert the script files and log records into text format and store them in the knowledge base of the RAG platform to form a historical dataset; User demand analysis step: The database backup and recovery demand information entered by the user in natural language is analyzed using a large language model and combined with prompt words to generate the target location information that the user needs to restore; Backup and recovery script generation step: querying the historical data set to obtain historical backup and recovery records that match the target location information that the user needs to restore, and then generating a backup and recovery script containing recovery resource configuration parameters based on the historical backup and recovery records through a large language model and combined with prompt word analysis; Recovery server selection step: querying the backup cluster server performance data stored in the historical data set, based on the estimated backup data volume in the backup recovery script, and using a large language model to select a recovery server from the backup cluster servers that meets the performance requirements determined based on the performance data; Recovery verification execution steps: Automatically execute the generated backup and recovery script on the selected recovery server, perform recovery verification on the backup file of the user-specified target database, monitor the execution process of the backup and recovery script in real time and record the execution log. When an execution error of the backup and recovery script task is detected, the error execution log is fed back to the large language model for cause analysis. Based on the analysis results of the large language model, the backup and recovery script is re-executed or the recovery resource configuration parameters are adjusted to generate a new backup and recovery script. Result feedback and optimization steps: When the backup and recovery script task is monitored to be completed, the complete execution log of the recovery process is stored in the historical data set. At the same time, a task execution result notification is sent to the user through the user interaction interface to complete the automatic recovery verification of the database backup.
2. The database backup automatic recovery verification method based on a large language model according to claim 1 is characterized in that: In the backup and recovery script generation step, a backup and recovery verification interface is also constructed based on the FastAPI framework developed in Python. The backup and recovery verification interface includes a backup query interface and a parameter query interface; The backup query interface calls the RMAN API to query the backup dataset file name corresponding to the target location information in the historical dataset. Based on the backup dataset file name, the large language model and prompt words are used to analyze the recovery resource configuration parameters required to restore the backup file of the target database, and then a backup recovery script is constructed. In the recovery server selection step, the backup cluster server performance data stored in the historical data set is queried through the parameter query interface. The performance data includes the current operating environment parameters and physical server performance indicators. Based on the estimated backup data volume in the backup recovery script, a recovery server that meets the storage space requirements and memory capacity requirements in the current operating environment parameters and the processor performance indicator requirements in the physical server performance indicators is selected from the backup cluster servers through a large language model; the operating environment parameters include memory configuration status, storage path information and version characteristics.
3. The database backup automatic recovery verification method based on a large language model according to claim 2 is characterized in that: The backup and recovery verification interface also includes a log callback interface, a backup execution interface, and a backup status query interface. During the recovery verification execution step, the recorded execution log is fed back to the large language model in real time for analysis through the log callback interface, and the recovery strategy is dynamically adjusted based on the analysis results of the large language model; the backup execution interface is used to call the backup recovery script to perform recovery verification on the backup file of the target database; and the backup status query interface is used to query the execution log to obtain the current backup file recovery status; The recovery strategy includes parameter optimization, script modification and server reselection.
4. The database backup automatic recovery verification method based on a large language model according to claim 2 is characterized in that: The backup query interface also receives a query request including a target database identifier and a recovery time range, and returns backup file metadata in a standardized JSON format.
5. The database backup automatic recovery verification method based on a large language model according to claim 1 is characterized in that: In the user interaction interface construction step, the database backup and recovery requirement information includes a target database identifier and a desired recovery time point; In the historical data set construction step, the script file includes backup and recovery commands and parameter configurations; the log record includes script execution results and error information; In the user demand analysis step, the target location information includes the target database network address, the database system identifier, and the time point to which the user expects to restore.
6. The database backup automatic recovery verification method based on a large language model according to claim 1 is characterized in that: In the backup and recovery script generation step, the recovery resource configuration parameters include system global memory area configuration parameters, program global memory area configuration parameters and a recovery operation target instruction set, wherein the system global memory area configuration parameters include memory allocation values for a database buffer, a shared pool and a redo log buffer; the program global memory area configuration parameters include memory limits for a sort area, a hash area and a private SQL area; and the recovery operation target instruction set includes a backup set path, a recovery termination point and data file redirection rules.
7. A database backup automatic recovery verification system based on a large language model, characterized in that: It includes a user interaction interface construction module, a historical data set construction module, a user demand analysis module, a backup and recovery script generation module, a recovery server selection module, a recovery verification execution module, and a result feedback and optimization module. The user interaction interface construction module uses the RAG platform to construct a user interaction interface, and receives database backup and recovery requirement information input by the user in natural language through the user interaction interface; The historical data set construction module collects script files and log records used in historical backup and recovery operations of several databases of the enterprise; converts the script files and log records into text format and stores them in the knowledge base of the RAG platform to form a historical data set; The user demand analysis module analyzes the database backup and recovery demand information input by the user in natural language using a large language model and combined with prompt words to generate the target location information that the user needs to restore; The backup and recovery script generation module queries the historical data set to obtain historical backup and recovery records that match the target location information that the user needs to restore, and then generates a backup and recovery script containing recovery resource configuration parameters based on the historical backup and recovery records using a large language model and combined with prompt word analysis; The recovery server selection module queries the backup cluster server performance data stored in the historical data set, and selects a recovery server from the backup cluster servers that meets the performance requirements determined based on the performance data based on the estimated backup data volume in the backup recovery script and using a large language model; The recovery verification execution module automatically executes the generated backup recovery script on the selected recovery server, performs recovery verification on the backup file of the user-specified target database, monitors the execution process of the backup recovery script in real time and records the execution log. When an execution error of the backup recovery script task is detected, the error execution log is fed back to the large language model for cause analysis. Based on the analysis results of the large language model, the module chooses to re-execute the backup recovery script or adjust the recovery resource configuration parameters to generate a new backup recovery script; The result feedback and optimization module, when monitoring the completion of the backup and recovery script task, stores the complete execution log of the recovery process in the historical data set, and sends the task execution result notification to the user through the user interaction interface to complete the automatic recovery verification of the database backup.
8. The database backup automatic recovery verification system based on a large language model according to claim 7 is characterized in that: In the backup and recovery script generation module, a backup and recovery verification interface is also constructed based on the FastAPI framework developed in Python. The backup and recovery verification interface includes a backup query interface and a parameter query interface; The backup query interface calls the RMAN API to query the backup dataset file name corresponding to the target location information in the historical dataset. Based on the backup dataset file name, the large language model and prompt words are used to analyze the recovery resource configuration parameters required to restore the backup file of the target database, and then a backup recovery script is constructed. In the recovery server selection module, the backup cluster server performance data stored in the historical data set is queried through the parameter query interface. The performance data includes the current operating environment parameters and physical server performance indicators. Based on the estimated backup data volume in the backup recovery script, a recovery server that meets the storage space requirements and memory capacity requirements in the current operating environment parameters and the processor performance indicator requirements in the physical server performance indicators is selected from the backup cluster servers through a large language model; the operating environment parameters include memory configuration status, storage path information and version characteristics.
9. The database backup automatic recovery verification system based on a large language model according to claim 8 is characterized in that: The backup and recovery verification interface also includes a log callback interface, a backup execution interface, and a backup status query interface. In the recovery verification execution module, the recorded execution log is fed back to the large language model in real time for analysis through the log callback interface, and the recovery strategy is dynamically adjusted based on the analysis results of the large language model; the backup execution interface calls the backup recovery script to perform recovery verification on the backup file of the target database; and the backup status query interface queries the execution log to obtain the current backup file recovery status; The recovery strategy includes parameter optimization, script modification and server reselection.
10. The database backup automatic recovery verification system based on a large language model according to claim 8, characterized in that: The backup query interface also receives a query request including a target database identifier and a recovery time range, and returns backup file metadata in a standardized JSON format.
Citation Information
Patent Citations
Automatic recovery verification scheduling method and device for database
CN117743031A
Auto point in time data restore for instance copy
US20200073763A1
Cloud-based application performance management and automated recovery
US20240259280A1